[BidClub_]
The Cognitive Revolution · · 54 min

Aaron Levie, CEO of Box, on Box AI, Enterprise Enthusiasm, and the Evolution of SaaS

Nathan LabenzAaron Levie

YouTube
TL;DR
  • Enterprise AI demand is running far ahead of cloud-era enthusiasm, but production remains in “very early innings.” Cloud required reluctant companies to surrender physical infrastructure and trust unfamiliar vendors; AI instead has executives proposing “almost as many use cases as possible,” sometimes more than practical. Levie sees a relatively narrow window in which large technology companies, enterprise-software vendors, and new agent startups can capture that demand.

  • AI changes IT from a software-enablement function into an operator of digital labor. Business units will ask IT not merely to deploy CRM or HR systems, but to provision agents that run sales campaigns, review contracts and invoices, or execute onboarding. Borrowing Jensen at NVIDIA’s framing, “the IT department becomes the HR department of AI,” requiring much deeper business knowledge and strategic authority.

  • Box’s RAG advantage came from architecture it built before ChatGPT, partly through luck. Box Hubs lets users curate authoritative documents without copying or changing permissions, reducing the risk that draft-heavy corporate repositories contaminate retrieval when users query a bounded hub. Because enterprises lack the public web’s PageRank-like authority signals, Levie argues this human curation makes Box’s RAG “at least a hundred times better” than indiscriminate search across all company data.

  • The strategic destination is a new “system of intelligence” that combines structured control with probabilistic judgment. Box’s current agents include model access, tools, skills, instructions, and enterprise data, though Levie concedes many would have been called assistants two years ago. The larger prize is agents that review content, make context-dependent decisions, route work, and coordinate with Salesforce, ServiceNow, Microsoft, or humans.

  • Being human-equivalent or 40% cheaper is insufficient to overcome enterprise adoption friction. A buyer still faces 17 competing projects, AI-council approval, testing that might take six months, privacy and security concerns, and workforce transition; consequential workflows may also demand “99.99999% reliability,” not 98%. Levie’s commercial threshold is an order-of-magnitude gain in cost, quality, or capability, especially for work that was never automated before.

  • Incumbents should retain their natural application domains, while startups win in cross-platform or previously unserved workflows. An “AI-first CRM” must assume Salesforce becomes AI-first too, just as Workday and ServiceNow will defend their domains; thin layers over an incumbent or OpenAI are therefore vulnerable. Independent compliance, agent red-teaming, cross-application workflows, and products requiring substantial non-AI interfaces remain credible startup territory.

  • AI may reopen SaaS pricing while improving productivity before the aggregate statistics clearly register it. Levie expects outcome pricing, compute consumption, and subscriptions all to coexist after two decades dominated by per-seat SaaS; Box itself is still testing. Internally, coding-tool gains range from 5–10% for some engineers to possibly 50% for a new hire, but Levie would take the over on economy-wide productivity forecasts if the clock starts in a few years.

Digest · the substance, structured for research

1. Enterprise enthusiasm has inverted the cloud-adoption pattern

  • Levie would not pin down AGI because its definition remains amorphous and he is “downstream of whatever Ilya, Sam, or Greg are talking about.” Conditional on the current pace continuing, he points to reasoning gains in math, logic, and coding; continued benchmark improvement could provide building blocks for models that learn on their own, making “some form of whatever we would have previously defined AGI as” plausible within a few years.

  • Personal adoption offers one small demand signal: Levie added at least five AI applications to his home screen in six months, versus perhaps one meaningful addition every year or two over the preceding decade. He regularly talks to Gemini Voice and ChatGPT voice and video, uses Perplexity, plays with xAI and Grok, and built a prototype with Artifacts and Claude.

  • Cloud’s first years brought skepticism over moving infrastructure out of company data centers, trusting unfamiliar vendors such as Amazon, and replacing physical control with APIs, dashboards, and audit reports. At the same point in AI’s cycle, customers are inventing use cases “probably in many cases more than is actually practical.”

  • The caveat is deployment: excitement has not yet translated into much at-scale enterprise usage. Still, Levie sees a “relatively narrow window” in which large technology companies such as Microsoft, Oracle, and Google, enterprise-software vendors such as Box, Salesforce, and ServiceNow, and entirely new agent companies can address problems enterprises previously could not attack.

2. IT becomes the operating department for AI labor

  • Traditional IT selected, deployed, secured, and managed systems while business functions such as sales, finance, marketing, and HR remained responsible for execution. A CRM made salespeople productive; it did not itself own the sales campaign.

  • AI “flips that on its head.” A sales leader might ask IT to spin up agents for a campaign, while operations teams request agents to review invoices, process contracts, or manage client onboarding—the technology organization now helps perform the work, not merely support it.

  • Jensen at NVIDIA’s formulation captures the organizational shift: “the IT department becomes the HR department of AI.” IT must understand operating processes, model capabilities, agent vendors, and the wider ecosystem well enough to provision digital labor, making the function more strategic but forcing dramatic transformation.

  • Box, which Levie says serves about 15,000 customer companies and stores well over 100 billion files, is positioning enterprise content as that labor’s substrate. Box AI preserves permissions, privacy, security, and access controls while connecting models to documents, extracting structured metadata, and eventually enabling agents to operate on contracts, financial records, research, media, and product plans.

3. Authoritative curation—not embeddings alone—makes enterprise RAG work

  • Labenz’s pushback concerns the familiar “trough of disillusionment”: many RAG systems fail because vector search retrieves the wrong source before generation even begins. Similar embeddings cannot reliably distinguish an authoritative earnings report from “Earnings_Final,” “Earnings_Draft_1,” and “Earnings_Draft_1_Sally_Edits” variants accumulated over years.

  • Box happened to begin Hubs roughly a year before ChatGPT. Its many-to-many architecture lets one canonical document appear in 20 topic-specific hubs without moving it, duplicating it, or changing permissions; updating the source automatically updates every hub. Levie’s candid assessment: “In this case we got totally lucky.”

  • Public search benefits from PageRank-like evidence that one article or site is more authoritative than another; messy corporate content generally has no comparable signal. By choosing what belongs in a sales, product, regional, or HR hub, employees identify both the trusted corpus and the questions appropriate to it.

  • That bounded, curated retrieval layer is the claimed breakthrough. Rather than ask generic questions across 100 million heterogeneous files, a user queries authoritative sales material inside a sales hub; Levie estimates the resulting service is “at least a hundred times better” than broad-based RAG across everything.

4. Agents evolve from branded assistants into probabilistic workflows

  • Box deliberately uses a broad definition of agent to avoid confronting customers with “17 different versions of a thing.” An agent combines one or more models, platform tools, underlying skills or capabilities, system prompts, proprietary architecture, and permissioned access to enterprise data.

  • The first generation is modest: users can converse with one document, query many files, generate content, extract metadata, or create a sales agent with specialized language and instructions. Levie acknowledges that “many of these agents would be what we would have called assistants two years ago.”

  • The destination is closer to Labenz’s autonomy test: multi-step workflows with meaningful decision-making discretion. An agent reviews a contract, identifies risky clauses, determines what subsequent process to trigger, and routes the result to another agent or human; browser operation could bridge systems where clean APIs do not exist.

  • Levie maps the shift across three eras: systems of record deterministically changed database rows; systems of engagement made human collaboration fluid and less structured; systems of intelligence combine record-like control with engagement-like adaptability. Because most real work “does require judgment,” probabilistic workflows expand what enterprise software can digitize.

5. Reliability matters, but enterprise inertia demands a 10x result

  • Levie agrees that organizations underuse capabilities already available. Even while immersed in AI, he must remind himself that instead of asking somebody to handle a task, he should “go and try” creating it with AI; the awareness gap outside the industry is likely much larger.

  • He nevertheless rejects 98% reliability for consequential processes: nobody accepts arriving at an airport to discover that 2% of booked tickets do not exist. Enterprises may require “99.99999% reliability,” and even a 0.01% error rate is disqualifying when applied to a billion financial transactions.

  • Cost is another constraint. Customers may want 10,000 agents aimed at a problem, then discover they cannot yet afford the required AI; privacy, security, workforce transitions, and retraining add more friction. Levie therefore expects a minimum decade-long change for enterprises to become genuinely AI-first.

  • Labenz’s sharper challenge: use deterministic code whenever an explicit algorithm exists because it is faster, cheaper, and more reliable; reserve intelligence for fuzzy work, where a well-contextualized model may already match a human. Levie agrees that an additional commercial obstacle is that “incrementally better, incrementally cheaper, incrementally faster” rarely clears organizational prioritization.

6. Startups need orthogonal markets, not AI replicas of incumbents

  • A pitch offering the existing process at 40% lower cost sounds compelling in an economics model, yet may rank ninth behind 17 other projects. The buyer must secure AI-council approval and perhaps run a six-month validation, so Levie believes vendors need one-tenth the cost, 10x the quality, or some comparable order-of-magnitude improvement.

  • Net-new augmentation faces less resistance than replacement. Copilot and Cursor work because users keep coding while receiving immediate productivity gains. Levie separately argues that a 10-times-better result—such as more effective cancer discovery—can justify adoption; work that was not automated before can add capability without first dismantling a staffed process.

  • Levie’s “timeless” startup rule is to pursue what incumbents cannot naturally absorb. Thin layers over OpenAI or Salesforce are bad bets: “you should anticipate that Salesforce is an AI-for-CRM system,” while Workday and ServiceNow will build AI products for their domains.

  • Openings remain where workflows cross applications, require extensive non-AI interfaces, or sit orthogonally to incumbent strategy. Levie says independent agent safety, compliance, and red-teaming can be viable, while a vendor spanning Salesforce, HR, ERP, and Box can occupy a layer no single application necessarily owns.

7. Pricing and productivity will change, but diffusion sets the clock

  • On reports that Klarna shut down a couple of systems of record, Levie’s view is deliberately mixed: the account may be overplayed, but the technical claim is not impossible. Building a homegrown AI Workday replacement to save a few hundred thousand dollars is “super provocative,” yet he doubts 90% of corporations would prioritize it.

  • AI’s proximity to completed outcomes reopens pricing. A lead-generation agent could charge per lead, per unit of compute consumed to create 10,000 leads, or through a fixed subscription that absorbs volume variation; Box is testing models rather than offering universal guidance after more than 20 years of per-seat SaaS convention.

  • Inside Box, a new employee learning to sell the product can query the sales hub “like you’re talking to a top expert in the company,” available 24/7. Coding gains reportedly range from 5–10% to possibly 50% for a new hire; the practical KPI is simply shipping more software.

  • Levie would take the “over” on AI eventually adding more than half a percentage point to annual productivity if the clock starts in a few years. Diffusion through “human-mediated environments” will take longer than expected, even as instant access to roughly “90th-percentile expertise on any topic” makes the underlying technological break unmistakable.

Aaron Levie

The rate of change that we're seeing and the rate of just exponential improvement that we're seeing from AI models is incredible. I would assume that if we keep that pace up, these systems will increasingly be able to perform any kind of general task.

Jensen at Nvidia put it the best: effectively, the IT department becomes the HR department of AI. That just opens up so many new questions about what the future of IT looks like. I think we're entering a new era with systems of intelligence that let us combine data, AI, and underlying enterprise software to automate really anything about our business.

There's going to be a tremendous amount of AI startup opportunity, but it will not come from just doing an AI-first CRM system, because you should anticipate that Salesforce is an AI-for-CRM system.

Nathan Labenz

Aaron Levie, founder and CEO of Box, welcome to The Cognitive Revolution.

Aaron Levie

Thank you. Good to be here. A lot of stuff going on in AI land right now.

Nathan Labenz

How are you doing?

Aaron Levie

Never a dull moment, that is for sure.

Nathan Labenz

Let's do a quick warm-up, and then we'll get into what's going on with AI in the enterprise and what you guys are bringing to the enterprise with your latest AI features. For starters, what are you doing with AI in your personal life, and what is your AI worldview, particularly as it relates to whether we're going to get AGI soon? What are you expecting over the next couple of years?

Aaron Levie

It's hard to say, obviously, based on the very amorphous definition of AGI. I'm downstream of whatever Ilya, Sam, or Greg are talking about, so I have no better ability to predict where that's going than anybody else right now.

All I would say is that the rate of change and the rate of exponential improvement we're seeing from AI models is incredible. I would assume that if we keep that pace up, these systems will increasingly be able to perform any kind of general task that you give them. With these reasoning models, we're seeing incredible results in math, very complex logic, and obviously coding.

Once you have those basic foundations, if we can continue to improve more and more on the benchmarks, that gives you the building blocks for models that can continue to learn on their own and get as advanced as we need them to be for basically any task that you'd want to give them. It seems like within the next few years, we could be getting some form of whatever we would have previously defined AGI as.

Nathan Labenz

Yeah, not that much further to go, I would say. Certainly, the checkpoints seem to be coming, if anything, more densely than they were not long ago. How about your day-to-day use? Do you have favorite use cases? Do you feel like you're properly challenging AI Pro, or are you doing more basic stuff?

Aaron Levie

Partly because I always get enamored with new technology, I try everything out. One thing I was reflecting on recently is that I have more new apps on my home screen in the past 6 months than probably at any other time in the past decade or decade and a half.

My home screen used to be, “Okay, you added Uber, then you added Spotify, maybe one social app, and then WhatsApp.” Only every 1 or 2 years did something get to the home page. Recently, I have at least 5 new apps that I've added into the mix, which to me is a proxy for just how much infusion of AI has already occurred in our personal lives.

I'm talking to Gemini Voice and OpenAI with ChatGPT voice and video pretty regularly, using Perplexity to get different kinds of answers, and playing with xAI and Grok. It's incredibly exciting to see the range of what these AI models can do. I built a prototype with Artifacts and Claude. Everywhere in my personal and professional life, AI is getting added into the mix.

Nathan Labenz

You recently put out a LinkedIn post that I thought was pretty interesting, where you basically said that this is the most energized you've seen enterprise companies about a new technology at any point in your career. Obviously, you made your company on the cloud wave, so that's another big wave that people were pretty enthused about.

For people who aren't in the room and aren't hearing these conversations with senior leadership at enterprise companies, what is the vibe like? What are they doing? Are they hands-on? What are you seeing?

Aaron Levie

Given that you brought up cloud, it's actually a really interesting comparison between the two. In the early days of the cloud, I would not describe the energy as super excited, super animated, and aggressively pursuing moving to the cloud. Most of the conversations we had with enterprises involved some degree of skepticism, resistance, and friction. There wasn't a lot of pure excitement with no friction around it.

The reasons were that this was a very big shift for enterprises. They had to move their infrastructure from the data centers they managed to the cloud. They had to trust these new vendors that they hadn't worked with before. If you were an enterprise, you'd never worked with Amazon as an enterprise vendor, which was obviously very atypical. You also had to change your whole notion about data privacy and compliance, from a world of physical management of hardware to a world where you get a set of APIs and maybe some dashboards and audit reports, and that's your only way of controlling these systems.

Enterprises were very reluctant early on to move to the cloud. If I compare that to today, and I snap the line at 2 or 3 years into cloud versus 2 or 3 years into AI, the environments couldn't be more different. Enterprises are looking for almost as many use cases as possible in which they can deploy AI, probably in many cases more than is actually practical.

There's a sense of creativity, excitement, and innovation that didn't necessarily exist with the cloud. With the cloud, it was very pragmatic. It was, “I could take this piece of infrastructure and move it into a virtualized environment.” There's nothing particularly exciting about that. It's very utilitarian.

With AI, people are saying, “What if we could actually solve this business problem that we've never been able to attack before?” Or, “What if I could deploy human resources against much more interesting problems than where my talent is currently dedicated?” All of a sudden, the creativity, energy, and opportunity are totally different for these enterprises.

There's a big asterisk, which is that we're still very early. The actual amount of at-scale deployment in the enterprise is still in the very early innings in terms of where this technology has been deployed. But the excitement level is totally different.

To me, as someone in enterprise software, this means there's going to be an insane amount of opportunity. We're going to see opportunity for bigger companies like Microsoft, Oracle, and Google. There's going to be opportunity for the software stack, like Box, Salesforce, and ServiceNow. There's going to be a lot of opportunity for brand-new startups to build AI agents that solve a whole new set of problems for the enterprise.

I think there's a relatively narrow window of opportunity right now. You're going to see a tremendous amount of change and a tremendous number of new startups emerge, and enterprises are very much ready and excited to pursue that.

Nathan Labenz

What do you think that change looks like in practice? One of the comments that you made in that post was that you expect IT departments, in a way, to undergo almost a reversal from the cloud.

In the cloud, it was, “You aren't going to manage this physical stuff anymore. More of that responsibility is going to shift to the vendor in terms of making sure we have uptime, and we'll just consume that and build apps on top of it.” But if I understand you correctly, you're saying almost the opposite when you say IT departments are going to go from supporting the work to actually doing more of the work.

What does that mean, and how many enterprise IT departments are up to that challenge today? How much are they going to have to transform to take on that challenge?

Aaron Levie

I think enterprise IT departments are going to have to transform pretty dramatically. The history of IT was that, many times in partnership with—and even proactively with—the business, the IT department would work with the business. By “the business,” I mean the marketing team, sales team, finance team, or somebody else in the business who had a particular need.

They might say, “I want to have a CRM tool so I can track all my leads,” or, “I want an HR system so, as we hire more people globally, we can make sure we're managing all the local laws from an HR standpoint.” They would go to the IT team, review a set of vendors, and then basically hand it off to IT to manage the system, deploy the technology, pick the vendor, and enable the business.

That relationship has been well-defined and codified for a few decades as we've had modern IT environments. But it was still the responsibility of the business to make use of the technology, be productive, and drive execution in the company. The sales team still owned all of the execution, and they were using a CRM system from the IT team to be productive.

AI flips that on its head. Now it's not just the people in the business who are using the technology to enable the business. The business is going to go to IT and say, “I need you to deploy AI labor against different kinds of problems that I'm dealing with.”

You might have a future where the head of sales goes to the IT team and says, “I need to spin up a new sales campaign,” if you have AI agents that can help do that. Or a back-office processing team might go to the IT team and say, “Do you have AI agents that can review all my contracts, review all these invoices, or help with all my client-onboarding workflows?”

This is a totally different era for IT: being in the position of actually solving the business problem, not just deploying the software to help the business solve its problem. This means you're going to have to understand the business much more if you're in IT. You're going to have to become even more strategic for the company. You're certainly going to have to understand all the trends happening in AI and the ecosystem, including all the surrounding companies that produce AI agents and the software around them.

Jensen at NVIDIA put it the best: effectively, the IT department becomes the HR department of AI. That opens up so many new questions about what the future of IT looks like, all of which are much more exciting than the past. But we are in for quite a bit of change in this space.

Nathan Labenz

If I torture a metaphor here, if IT departments are the HR departments of AI, then Box is applying for, or looking for, a promotion in a lot of these organizations. You've just been through what you might call an upskilling, with a bunch of new generative AI product and feature releases.

Tell us what's new as we enter 2025 at Box.

Aaron Levie

We work with about 15,000 customer companies, and what they use Box for is to secure, collaborate on, and automate workflows around their content. That could be their contracts, financial documents, marketing assets, research data, or product plans. We store all those documents, media assets, digital images, and other content; help secure it; automate workflows around it; and enable companies to collaborate on that data.

AI is the next frontier of what we can do with all of this unstructured data in the enterprise. We built a layer called Box AI that connects AI models to enterprise content in a secure way. That lets the enterprise continue to have its security, privacy, access controls, and permissions. We handle all of that, and then connect AI models with this abstraction layer.

What we've been building out is a set of capabilities to let people transform and unleash the power of their data. The first big use case is being able to talk to your data for the first time. We store well over 100 billion files in Box, and you can imagine that every single one has incredible insights and value inside of it. But most of the time, you don't know what's inside your data because until you search for it and look at it, you don't know what's actually out there.

Now, for the first time, you can just talk to your data. I could take 100 sales presentations and ask, “What's the best practice for selling this product to a new customer?” Our system will read through all those documents and produce exactly the right answer for whatever you're doing with that customer.

You could take a bunch of medical research or life sciences documents and research and talk to all that data to find interesting trends or patterns in that information. That first use case is basically retrieval-augmented generation over many documents, with an AI layer that powers it.

The next big use case, which we've just started working on, is the ability to read through any kind of content and extract the most important structured data from that content. Take a contract: you want to pull out the renewal date or the names of the parties, and then store that in a structured database. That way, you can query it later, create dashboards, and automate workflows. That's the next big use case we're working on, and it's rolling out imminently.

Where this is all going is eventually to agentic workflows within our platform or connected to other systems. The very big prize in this space is taking all of the information in your enterprise and having agents operate on that data.

I could have a contract agent, a marketing agent, or a sales agent that performs tasks to make me more productive. It could review a contract, pull out the riskiest clauses, and route it to the right person. I could design an agent to perform those workflows across my entire data set and then connect it to another system.

I might want to connect it to Salesforce, ServiceNow, or Microsoft. Box becomes a layer for managing intelligent content, and then we connect to other technologies that handle the AI in their particular part of the workflow.

This is the future we're building out. I think it's the biggest set of changes and the most we've built in the history of the company. It's unbelievably exciting because on a daily basis, you're seeing new things that were never possible before with our data.

Nathan Labenz

Let me ask 2 questions, digging into the retrieval and then the agents, one by one. Retrieval-augmented generation has obviously been a huge trend, and a lot of people and companies have rushed out and implemented a version of it. I've heard over and over again that it hasn't worked super well for people.

When I get under the hood and explore why it isn't giving people the answer they want, it seems like the most common failure is that the vector-embedding-mediated search isn't retrieving the right content in the first place. If you don't have the right sources, you don't get the right answers. What have you done to deal with that? Perhaps structuring the unstructured data is part of it, but how do you address that problem, which seems to pop up so often?

Aaron Levie

That's the first trough of disillusionment for so many people. Let me ask you a question: when you see that trough of disillusionment, is it often with data sets that are quite broad or heterogeneous?

We've seen this a lot with the use case of, “I have my email, my data, and my calendar, and I want to search across all of it with some kind of AI system.” Are you seeing it in those kinds of scenarios, or in other ones?

Nathan Labenz

Honestly, I feel like it almost always pops up, even if it isn't a crazy complicated situation.

Aaron Levie

I'd love to attribute this to pure genius, but we got lucky. We were building a product about a year before ChatGPT launched. It's now called Hubs, and it was the ability to organize content on a topic-by-topic basis so you could share or search that content on a per-topic basis.

The novel idea was that within your Box account, you manage files inside folders, but you could create a hub that points to content within your Box account. The whole power of it was to create a many-to-many relationship. I could have the same documents show up in 20 different hubs without ever moving them or changing their permissions. If I made one update to that document, it would propagate to all the hubs where it was accessed.

We created this architecture because we saw that a lot of people wanted to create a sales hub that pointed to sales data. Then they would say, “I want a sales hub for the Japan region, and I want to create a sales hub for a particular product line. I don't want to change the documents, fork them, and keep track of why the one in the Japan hub is updated but the one in the product hub isn't.”

The idea was to create virtual pointers back to the same source of truth, with virtual hubs that you could create. We were working on that about a year before ChatGPT, unrelated to AI, with search and discovery as common use cases. As soon as ChatGPT launched, we thought, “Obviously, the next big thing we should do is let you talk to all the data in a hub.”

What we discovered was that this was a breakthrough architecture for a RAG use case. Where RAG runs into problems is when you have, for example, 100 million files inside your enterprise and somebody asks, “What was the last revenue figure in our last quarterly earnings?”

The challenge is that you might have a document called “Earnings_Final,” another called “Earnings_Draft_1,” and another called “Earnings_Draft_1_Sally_Edits.” The vector embeddings on all of those documents look very relevant to producing the answer, but the accuracy and authoritativeness of any one of those versions might be totally wrong based on where it was in the editing process.

Now multiply that by years of data and signals from other systems, such as email or other data sources. All of a sudden, it becomes very hard to ask any generic question of your internal enterprise data.

This is less of a problem for public services like Perplexity because you get the benefit of almost a PageRank algorithm for public data. You can look at a curve and say, “This CNN article is more authoritative than this random blog,” or, “This blog is more authoritative than this random website that was created 2 days ago.” You get an authoritative score on the public internet in a way that corporate data doesn't really have.

Corporate data is much messier. It tends not to have a PageRank element, so it's very hard to know the ultimate source of truth. Hubs basically solve this problem because users tell us what the authoritative copy of the data is that they're putting into a hub.

If I created a sales hub with sales presentations and product information, I would put only the authoritative records into that hub. When you're asking a question in that hub, you're asking questions about sales. You aren't going to ask the sales hub an HR question; you're going to go to the HR hub.

We've been able to get the user to ask the right types of questions where the data set is authoritative and isn't as messy as doing a broad-based RAG environment over all your data. Those are 2 or 3 different hacks we've locked into and doubled down on, and they've made our particular RAG service, I think, at least 100 times better than just doing a broad-based deployment across all your data.

Nathan Labenz

That's interesting. I've heard so many stories about how somebody just happened to be building the right thing that was perfectly complemented by AI, and then the whole business shifted as a result of those things coming together.

Aaron Levie

We pinch ourselves. If we hadn't been building that, I get very scared. It was about a year to a year and a half of deep architecture work. There was no way to build it any faster. You had to create a file system that could have a virtual sharing component, and for the complexity of our platform, that took a year or more.

If we hadn't already been a year or a year and a half into it, I fear we would have felt that it was too daunting based on how fast AI was moving. I'm not sure we would have landed on this exact architecture. We might have done a less optimal version. In this case, we got totally lucky but ultimately built the exact right thing for the exact right moment.

Nathan Labenz

On the agents: 2025 is the year of agents. We thought it might be 2024, but it turns out to be 2025. What does an agent mean to you?

I have my own sense of what the definition is, and I contrast an agent with what I tend to call an intelligent workflow based on how much autonomy or decision-making discretion the AI ultimately has. How do you think about it, and how much autonomy or discretion are you giving the agents in your platform?

Aaron Levie

I like that definition, and I would fully subscribe to it. In our platform, we've taken probably the broadest definition because we didn't want there to be 17 different versions of a thing, with the user having to understand all the differences. We've defined an agent as a mix of an AI model or many AI models, a set of tool use within the platform, and underlying skills or capabilities.

Those capabilities are some mix of system prompts, a proprietary architectural change, and access to your data. That whole collection of capabilities creates an agent.

Within our platform, you can do very simple things with an agent. You can talk to a single document, and you're using an agent, but it isn't doing anything agentic. You're talking to either the core Box AI agent or a custom agent.

You could create a sales agent with a custom prompt and custom instructions that contain information about how your sales workflows work or the kind of language you should use. You can create custom agents to let you talk to your data, extract metadata from documents, talk to many files at once, or create content.

That's our first era of agents. Many of these agents are what we would have called assistants 2 years ago, but we've created a universal language for this. Where it's really going is much closer to your definition: agentic workflows where the agent performs many sequential tasks and there are some probabilistic elements to those tasks. For example, “I want to review a document, and based on how I review that document, I want to kick off another process, either to another agent or to a human.” Or, “I want to take a lot of data, collate it, and produce something.”

These are much more multistep flows that we expect to make very agentic in the future. Right now, the limitation of software is that it's very good at deterministic workflows, but the vast majority of work is not deterministic. It requires judgment, and you change your answer based on other inputs or insights.

The majority of work in the future will be nondeterministic, judgment-oriented, agentic work. That's an incredible opportunity for enterprise software because that's what we can finally digitize.

The way I've been thinking about this recently is in terms of the eras of enterprise software. Forty years ago, we had the initial wave of systems of record, such as CRM and ERP systems. This was the definition of the most deterministic technology possible. You're effectively changing rows inside databases, and based on that, it kicks off a process. That was 100% deterministic.

Then we moved to systems of engagement. This was the idea of collaborative systems—Slack, Box, and other tools—where things are a lot less structured and deterministic. The workflows are more fluid and can adapt a bit more because they're human-to-human and a little messy.

Now we're in a new era of systems of intelligence. This is the era in which AI automates those workflows. These systems have all the properties and benefits of being structured like systems of record, but they have all the flexibility of a system of engagement. They can adapt and change because that's what AI can do.

I think we're entering a new era with systems of intelligence that let us combine data, AI, and underlying enterprise software to automate really anything about our business.

Nathan Labenz

I often say that intelligence has been debated for a long time, and I don't pretend to be the final word, but my working definition of intelligence is the ability to do useful work when there is no explicit algorithm that tells you what every step ought to be.

Aaron Levie

I think that's a great, very practical definition, especially for enterprise use cases. You don't want to have a predetermined or prewired workflow because it's going to change. You need it to adapt to new data. Maybe there aren't even APIs available for the thing you're trying to do.

That's where even browser access gets very exciting. There are a lot of future workflows where I want to automate 5 systems and have them talk to each other, but there aren't clean APIs for making them all communicate.

Nathan Labenz

I find myself in my AI-assisted coding workflows, when they're working well, basically copy-pasting things around most of the time. When they're not working well, I start to have to troubleshoot. The smooth thing is that I'm just the copy-and-paste monkey gluing these other things together. It's a weird experience when it's working well.

What would you say are the biggest bottlenecks that enterprise companies face today in realizing value? I have a thesis that the AIs are good enough to do a lot more than people are actually deploying them to do. Tyler Cowen recently said, in front of an audience, “All of you humans, you are the bottlenecks.” What do you see as the practical barriers that people are still struggling to get over?

Aaron Levie

I 100% agree that AI can do vastly more than most enterprises either think or have the near-term appetite to deploy. To some extent, there's more technology available at the moment than most people realize.

Even I have to remind myself: that task I would normally have asked somebody to work on—I'm reminding myself more and more, “No, go and try to create it in Artifacts or use AI to do this.” Even right in the center of AI, watching everything happen around the world, I have to trigger my brain to remember how much these things can do. I can only imagine that if you're not in it every day, that gap is probably fairly massive.

That being said, I don't want to let the AI model providers off the hook. One thing that prevents AI deployment is that you can't have an enterprise workflow, particularly in a regulated industry, that works 98% of the time.

You wouldn't find it acceptable if 98% of the flights you scheduled were successful, but 2% of the time you showed up at the airport and didn't actually have a ticket. Enterprises need 99.99999% reliability on almost anything that's important.

If it's a creative task—write a marketing campaign, review a blog, write a blog post, or send an email on this topic—you have some room for hallucination, or you can edit it later. But if you're going to have a billion financial transactions submitted every week or month and AI gets 0.01% of them wrong, that's a nonstarter for deploying those systems.

What we need from the model providers—and you can see it in the benchmarks—is for accuracy to keep getting higher. We need to get to the point where all the evaluations have to be completely reset because everything has hit 100%.

I think we're still a little bit technology-dependent and AI-model-dependent. We also need the cost of AI to continue to come down. The excitement from customers is real: they'll say, “I'd love to deploy 10,000 AI agents against this problem,” but then they'll look at the cost and say, “I still can't afford what that would look like in my business.” Maybe we have to take it in a more stepwise fashion until costs come down.

We need AI to be cheaper and the performance of these models to get even higher. From there, you're running into Tyler's point: all of the classic, human-based change-management difficulties.

Those range from privacy and security concerns, in some cases, to the very real issue of, “I still have humans doing that thing, so until I can transition them to a different role or teach them a new skill, we're not going to be able to automate that particular workflow.”

All of that is still what we're in for in the enterprise, and it's going to take years. This is going to be, at a minimum, a decade-long change in how enterprises become more AI-first.

But I think we have a roadmap because we did it with cloud. We also have an increasingly clear vision because we're starting to understand what this world could look like: What would an AI-first enterprise look like? How would agents be part of the workforce? How do we get better insights from all our data? How do we automate almost any workflow in our enterprise?

I think you see an increasing understanding of what an AI-first future could look like in most organizations.

Nathan Labenz

To push a little harder on this, I feel like a corollary to my definition of intelligence is that you should only use intelligence where there is no explicit algorithm. If you have a way to do something with traditional code, you probably should. It will be faster, cheaper, and more reliable.

If we bring AI back into the domain of all these fuzzy things for which we don't have algorithms, and ask how good AI is at doing those things compared with a human, my belief is that with some elbow grease—making sure you have the right context, putting a few examples together, and applying the best practices—you can most of the time get to the point where AI can do as well as a human at a much lower cost and much faster.

We've seen that in medical diagnosis recently. If you buy that, doesn't it suggest that something else is going on? Maybe it's a fallacy—or not necessarily a fallacy, but an attitude—that we need AIs to be not just on par but to have a 10-times-lower defect rate or something. That seems to be the case with self-driving.

Aaron Levie

I think you just nailed it. You can't go to a company and say, “I can do exactly what you're doing today, and you'll save 40%.” An economist would say, “Everybody would do that deal all day long,” but once that meets real life, the person has 17 other projects.

There is an incredible number of people, priorities, and demands all competing for their time. If you could wave a magic wand and make something 40% cheaper, you would totally do it, but of all the things you have to do, that might be number 9 on the list.

To your point, there's something else going on. That person now has to go to the AI council and get approval. They have to run a full 6-month test to make sure that it actually is 40% cheaper. They have to weigh that against everything else they're doing in their business.

I think we need AI to produce multiples-better improvements over the status quo. That's how you compel motivation. You can't be incrementally better, cheaper, or faster. You have to be an order of magnitude better on one of those dimensions.

If you could go to a company and say, “I can be literally one-tenth the cost of what you do today to review your contracts, review your invoices, or automate this back-end supply-chain process,” then you're talking. You're saying, “I could save you millions of dollars.”

Or you could say, “I can do a 10-times-better job than your human-based workflow today. We'll discover cancer more effectively, or we'll target even more automation across your enterprise.”

This often explains AI's biggest opportunities. I think the flavor of what you're saying is that you go after work that isn't automated today. You're not even replacing something that already exists; you're layering onto an existing workflow and making it better.

I would say this largely explains the breakthrough in Copilot or Cursor. I get to do exactly the same thing I'm doing now, but I see incremental productivity gains immediately without really changing any behavior.

The more AI can solve those problems—net-new use cases that add incremental productivity—the easier it is on the change-management front. Anything that replaces an existing process and saves only a little money is a much harder problem to go after in the enterprise.

Nathan Labenz

What do you make of the debates around the future of enterprise software? We've heard conflicting narratives because everything is happening so fast.

On the one hand, it was, “It's never been a better time to start a startup.” Then it was, “Actually, this technology is so easy to implement that incumbents will probably capture most of the value because they'll be able to roll it out to their existing customers.” They might have a little bit of a go-to-market lead, but incumbents already have the customers, so they'll beat you if you're trying to start up rather than having them figure it out and deploy it to their current customer base.

Then we have the Klarna narrative, where they've allegedly or reportedly shut down a couple of systems of record. There's also the pricing debate: do we still charge by seats, or do we have to move to per-outcome-based pricing? There's a lot there that could probably consume the rest of our time. What do you make of all that?

Aaron Levie

Exactly. If I had heard that when I was just starting, I would have been way too stressed to start a company.

There are definitely a lot of variables in flux. I compare that with our early days, when we had plenty of other problems but not the fundamental variables of the company's business model.

I think you're right. You have to ask which spaces give incumbents a natural advantage, whether the underlying billing model is seat-based or outcome-based, and whether there's a future of some kind of AGI-light that makes some software irrelevant because you don't need the software in the first place.

All of those things will happen. I would lean more toward timeless lessons of competitive strategy. If you're a brand-new startup, go after things that aren't easy for the incumbent to pursue.

If all you're doing is building a thin layer on top of OpenAI, that's a bad idea. If you're building a thin layer on top of Salesforce with AI, that's a bad idea. Salesforce is very competent; it will build the CRM AI product. Workday will build the HR AI product. ServiceNow will build the ServiceNow AI product.

Equally, if you're just finding a slight gap in what OpenAI does today, you have some risk of them moving up the stack, or the model getting better and eventually bringing that capability into the model layer.

That being said, I can think of a number of things that would be unnatural for OpenAI to do because they might involve a lot of non-AI interface and workflow work to solve a particular problem. You could be doing things inside the sales, HR, or IT service-management worlds that the incumbents equally aren't going to do.

Maybe it's cross-platform AI workflows that aren't natural for any one of those players to pursue. Maybe it's building AI agents that are so orthogonal to the normal strategy of Salesforce, Workday, or ServiceNow that those companies wouldn't think to pursue them. You can get enough traction fast enough to create some degree of a moat.

I think there's going to be a tremendous amount of AI startup opportunity, but it won't come from just doing an AI-first CRM system. You should anticipate that Salesforce is an AI-for-CRM system. We're seeing startups all the time that are finding those windows of opportunity right now.

Nathan Labenz

What do you make of Klarna? I've heard every take on it, from “It's all hype; they're not really doing it,” to “Maybe, but it's the exception that proves the rule.”

Aaron Levie

Right now, it's the exception. As I've seen the reports, I'm more in the camp that maybe it's overplayed a little bit, but nothing about it is impossible.

Given that it's not impossible to do what they've said, maybe they've chosen this as something that will differentiate them as a company. It's still different from every company on the planet to do it in a homegrown way.

My understanding was that they were going to build their own Workday system with AI, and that's just not a priority. Again, where are you in the prioritization stack? Most companies simply aren't focused on building their own HR system to save a couple hundred thousand dollars.

I think what they're doing is super provocative and super interesting, but not translatable to the broader economy. It's fun to watch. I invite as many companies as possible to try that experience and share their lessons along the way. It makes the conversation in the ecosystem much more interesting and dynamic.

But I'm not convinced that 90% of corporations would ever do what they're doing.

Nathan Labenz

Any pricing guidance?

Aaron Levie

I don't really have any guidance because we're testing all the business models ourselves. In general, one of the biggest benefits of AI is that you can get much closer to software solving an outcome.

The conclusion of that theory would be that your pricing model should be closer to that outcome. That could be the outcome very literally: you pay an AI agent to generate leads, and therefore you pay per lead. Or it could be the consumption that goes into that outcome: “I want the AI to generate 10,000 leads, and that takes a certain amount of compute capacity,” so I'm paying for the consumption of that compute capacity.

Then there are traditional subscription models: “I want an ongoing license that will roughly do this much volume for me. Sometimes it's a little more expensive, sometimes it's a little more volume, sometimes it's a little less, but I pay the same fixed rate the entire time.”

I think we'll see every version of these business models pursued. I put it in the category of something that's intellectually interesting because we're in such a dynamic period. We haven't had open questions about business models in software for 2 decades.

Salesforce, perhaps with a couple of other companies, essentially invented the idea that you pay per seat on a subscription basis. That's been the business model of SaaS for more than 20 years. Now we have a chance to say, “There are other business models that will begin to emerge.”

I find that incredibly fascinating, but each company has to develop its own understanding of what its customers are looking to pay for. That determines which business model makes the most sense.

Nathan Labenz

One niche that's really interesting to me, in terms of things people might want to buy separately, is safety or compliance. Maybe they're going to continue being a Salesforce customer and get all the agents or whatever, but perhaps they might want to buy something separately from somebody who red-teams the agents. Do you see that as viable?

Aaron Levie

Absolutely.

Nathan Labenz

That's interesting, because the other option would be for Salesforce to deliver that as another feature.

Aaron Levie

It's really important to understand the layers of the stack. If you think about an IT stack—Salesforce, HR, an ERP system, Box—which things cut across all of those? That's where you need a new vendor independent of any one of those players.

Some things make sense to exist in one of those systems. Other things are more likely to be technology that works across all the apps in your enterprise. That determines whether you could be a startup or whether it's really an incumbent game in that market.

Nathan Labenz

That's a good perspective. One last question: it's been famously said that we see the impact of computers everywhere except in the productivity statistics. It seems like AI may still be in that zone.

What are you seeing internally at Box when it comes to AI-enabled productivity boosts? Is that something you can measure, or is it something you believe in and encourage on faith right now? What do you have for other leaders who want to make sure they're getting the productivity that's promised?

Aaron Levie

We're 100% committed to being an AI-first enterprise and company. That's particularly important because we sell AI technology to enterprises. We need to be the first to understand where this is all going, but I also think it's going to be a way to run a better company in the future.

It's showing up in a handful of ways already. With Box AI, this is how we work with our unstructured data. If you're a new employee and want to learn how to sell our product, you go to our sales hub and ask it any kind of question. It gives you an answer back, and it's basically like talking to a top expert in the company.

You're getting all the value of talking to the smartest existing employee, but now you can do it 24/7. You don't have to wait for somebody to respond to your Slack message. That's probably less measurable because it permeates everything we do and simply improves productivity.

In other areas, it's anecdotal, but we've deployed AI coding tools. I'll get ranges of feedback: somebody will say they were 5% or 10% more productive, while a new hire might be 50% more productive because they can ramp up so much faster.

I think the biggest way this will show up is that we'll ship more software. That will be the measure of productivity we care about internally. I think AI will certainly be the first distinct technology category to show up in the long-term GDP graph that people talk about as driving productivity gains from technology.

Nathan Labenz

So, more than 0.5% a year, which is what Tyler Cowen recently said? You'd take the over?

Aaron Levie

If you let us start the clock in a few years, I think so. You still have diffusion of the technology across the economy, and that takes much longer than I ever think it should. But that's just life in human-based, human-mediated environments.

Nathan Labenz

I'm with you on that. I always underestimate the timelines. Anything else you want to touch on or leave with the audience before we break?

Aaron Levie

We're in such an unbelievable time to be building and deploying technology. The fact that anybody, anywhere in the world, could get 90th-percentile expertise on any topic instantly is an unbelievable thing.

If you had contemplated that 3 years ago, based on our understanding of technology at the time, it wouldn't have been conceivable. I just think it's an incredibly exciting moment, and we're super excited to bring it to the enterprise.

Nathan Labenz

Terrific stuff in both respects. Aaron Levie, founder and CEO of Box, thank you for being part of The Cognitive Revolution.

Aaron Levie, CEO of Box, on Box AI, Enterprise Enthusiasm, and the Evolution of SaaS | BidClub