Box CEO Aaron Levie:谈 Box AI、企业热情与 SaaS 的演进
企业对 AI 的需求远超云时代的热情,但生产级应用仍处于“非常早期”。 云计算要求企业不情愿地放弃物理基础设施、信任陌生供应商;AI 则让高管主动提出“尽可能多的用例”,有时甚至超出实际可行范围。Levie 认为,大型科技公司、企业软件厂商和新兴智能体创业公司,正处在一个相对狭窄的需求承接窗口中。
AI 正把 IT 从软件赋能部门改造成数字劳动力的运营部门。 业务部门未来向 IT 提出的,不只是部署 CRM 或 HR 系统,还包括配置能够执行销售活动、审阅合同和发票、完成入职流程的智能体。借用 Jensen 在 NVIDIA 的说法,“IT 部门会变成 AI 的 HR 部门”,这要求 IT 掌握更深的业务知识,并拥有更高的战略地位。
Box 的 RAG 优势部分来自 ChatGPT 之前就已搭建的架构,也有运气成分。 Box Hubs 允许用户整理权威文档,无需复制文件或改变权限,从而降低企业知识库中大量草稿污染检索结果的风险。由于企业缺少类似公共互联网 PageRank 的权威信号,Levie 认为,这种人工策展让 Box 的 RAG 至少比对全公司数据进行无差别搜索好 100 倍。
战略终点是一个新的“智能系统”,把结构化控制与概率性判断结合起来。 Box 当前的智能体包含模型访问、工具、技能、指令和企业数据,但 Levie 也承认,其中很多在两年前还会被称为助手。更大的机会在于让智能体审阅内容、作出依赖上下文的判断、分派工作,并与 Salesforce、ServiceNow、Microsoft 或人类协同。
达到人类水平或便宜 40%,仍不足以克服企业采用阻力。 买方仍要面对 17 个竞争项目、AI 委员会审批、可能耗时 6 个月的测试,以及隐私、安全和员工转型问题;涉及重大后果的工作还可能要求“99.99999% 的可靠性”,而不是 98%。Levie 的商业门槛是成本、质量或能力提升一个数量级,尤其适用于此前从未被自动化的工作。
在自然应用领域,现有厂商应继续守住自己的阵地;创业公司则应争取跨平台或此前无人服务的工作流。 “AI-first CRM”必须假设 Salesforce 也会变成 AI-first,Workday 和 ServiceNow 同样会捍卫各自领域;因此,叠加在现有厂商或 OpenAI 之上的薄层产品都很脆弱。独立的合规、智能体红队测试、跨应用工作流,以及需要大量非 AI 界面的产品,仍是创业公司的可信机会。
AI 可能重新打开 SaaS 定价空间,并在宏观统计明确反映之前先改善生产率。 Levie 预计,按结果收费、按算力消耗收费和订阅制将长期并存,结束持续 20 年的按席位 SaaS 主导格局;Box 自身仍在测试。公司内部,编程工具带来的提升从部分工程师的 5–10% 到新员工可能达到的 50% 不等,但如果从几年后开始计算,Levie 愿意押注全经济生产率最终会超过预测。
1. 企业热情颠覆了云计算的采用路径
Levie 不愿给 AGI 下明确结论,因为它的定义仍然模糊,而且自己“取决于 Ilya、Sam 或 Greg 在谈什么”。如果当前进展速度持续,他认为数学、逻辑和编码方面的推理能力仍会提升;基准测试持续改善,可能为能够自主学习的模型提供基础,让“某种我们过去会定义为 AGI 的东西”在几年内成为可能。
个人采用速度提供了一个小规模需求信号:过去 6 个月,Levie 至少把 5 个 AI 应用加入了手机主屏幕;而在此前 10 年里,他每 1 到 2 年才会有大约 1 次有意义的新增。他经常使用 Gemini Voice、ChatGPT 的语音和视频功能,也用 Perplexity,尝试 xAI 和 Grok,并用 Artifacts 和 Claude 搭建过原型。
云计算最初几年,市场怀疑企业是否应把基础设施移出数据中心、是否应信任 Amazon 等陌生供应商,以及是否应以 API、仪表盘和审计报告取代物理控制。在 AI 处于相同发展阶段时,客户正在创造用例,“可能很多时候已经超过了实际可行的范围”。
但部署仍是关键限制:热情尚未转化为大量规模化的企业使用。尽管如此,Levie 认为,大型科技公司如 Microsoft、Oracle 和 Google,企业软件厂商如 Box、Salesforce 和 ServiceNow,以及全新的智能体公司,都处在一个“相对狭窄的窗口”内,可以解决企业过去无法着手的问题。
2. IT 成为 AI 劳动力的运营部门
传统 IT 负责选择、部署、保护和管理系统,销售、财务、市场和 HR 等业务职能则负责执行。CRM 能提高销售人员的生产率,但销售活动本身并不归 CRM 执行。
AI“把这一切彻底倒转”。销售负责人可能要求 IT 为一场营销活动配置智能体,运营团队则可能要求智能体审阅发票、处理合同或管理客户入职;技术部门开始参与完成工作,而不只是提供支持。
NVIDIA 的 Jensen 给出的表述准确概括了这一组织变化:“IT 部门会变成 AI 的 HR 部门。”IT 必须理解运营流程、模型能力、智能体供应商和更广泛的生态,才能配置数字劳动力;这会让 IT 变得更具战略性,但也迫使其发生剧烈转型。
Levie 表示,Box 服务约 15,000 家客户企业,存储文件超过 1000亿个,正在把企业内容定位为这类数字劳动力的基础。Box AI 在保留权限、隐私、安全和访问控制的同时,把模型连接到文档,提取结构化元数据,并最终让智能体能够处理合同、财务记录、研究资料、媒体文件和产品计划。
3. 权威策展,而非单靠嵌入,让企业 RAG 真正可用
Labenz 的质疑指向人们熟悉的“幻灭低谷”:许多 RAG 系统在生成开始前就已经失败,因为向量搜索召回了错误来源。相似的嵌入无法可靠地区分权威的业绩报告与企业多年积累的“Earnings_Final”“Earnings_Draft_1”和“Earnings_Draft_1_Sally_Edits”等版本。
Box 恰好在 ChatGPT 出现前约 1 年开始建设 Hubs。其多对多架构允许同一份规范文档出现在 20 个主题 Hub 中,无需移动、复制文件或改变权限;源文件更新后,所有 Hub 会自动同步。Levie 坦言:“这次我们完全是走运了。”
公共搜索可以借助类似 PageRank 的证据,判断某篇文章或某个网站是否比其他来源更权威;混乱的企业内容通常没有类似信号。通过决定哪些内容应进入销售、产品、区域或 HR Hub,员工同时确定了可信语料库,以及适合在其中提出的问题。
这种边界明确、经过策展的检索层才是其所谓的突破。用户不再对 1亿份异构文件提出泛化问题,而是在销售 Hub 内查询权威销售材料;Levie 估计,这种服务至少比跨全部数据的广泛 RAG 好 100 倍。
4. 智能体从品牌助手走向概率性工作流
Box 有意采用宽泛的智能体定义,避免让客户面对“17 个不同版本的同一种东西”。一个智能体由 1 个或多个模型、平台工具、底层技能或能力、系统提示词、专有架构,以及对企业数据的权限访问组成。
第一代产品相对有限:用户可以与单份文档对话、查询多个文件、生成内容、提取元数据,或创建一个具备专门语言和指令的销售智能体。Levie 承认,“其中很多智能体,按我们两年前的叫法,其实会被称为助手”。
终点更接近 Labenz 提出的自主性测试:多步骤工作流,并拥有实质性的决策自主权。智能体审阅合同、识别风险条款、判断下一步应触发什么流程,再把结果分派给另一个智能体或人类;在缺少清晰 API 的系统之间,浏览器操作也可以充当桥梁。
Levie 将这一变化划分为 3 个时代:记录系统以确定性方式修改数据库行;互动系统让人类协作变得更流畅、结构更弱;智能系统则把记录系统的控制能力与互动系统的适应性结合起来。由于大多数真实工作“确实需要判断”,概率性工作流将扩大企业软件可以数字化的范围。
5. 可靠性重要,但企业惯性要求 10 倍结果
Levie 同意,企业对已有能力的使用仍然不足。即便自己沉浸在 AI 中,他也必须提醒自己,与其找人处理一项任务,不如“去试试”用 AI 创建它;而行业之外的认知差距可能更大。
但他拒绝接受涉及重大后果的流程只有 98% 可靠:没人能接受抵达机场后才发现已订机票中有 2% 根本不存在。企业可能要求“99.99999% 的可靠性”;当 0.01% 的错误率作用于 10亿笔金融交易时,也足以让产品出局。
成本是另一项约束。客户可能希望部署 10,000 个智能体来解决一个问题,随后才发现自己暂时负担不起所需 AI;隐私、安全、员工转型和再培训又会增加阻力。因此,Levie 预计企业真正变成 AI-first,至少还需要 10 年的变化。
Labenz 提出了更尖锐的挑战:只要存在明确算法,就应使用确定性代码,因为它更快、更便宜、更可靠;智能能力应留给模糊工作,而经过良好上下文处理的模型可能已经能达到人类水平。Levie 同意,另一个商业障碍在于,“略好一点、略便宜一点、略快一点”通常不足以通过企业内部的优先级排序。
6. 创业公司需要正交市场,而不是复制现有厂商
如果把现有流程的成本降低 40%,放进经济模型里看似很有吸引力,但项目可能仍排在 17 个项目中的第 9 位。买方还要获得 AI 委员会批准,甚至进行 6 个月的验证,因此 Levie 认为,供应商需要把成本降到十分之一、把质量提升 10 倍,或实现类似数量级的改善。
全新的增强型产品面临的阻力小于替代型产品。Copilot 和 Cursor 之所以有效,是因为用户继续编码,同时获得即时生产率提升。Levie 还认为,10 倍更好的结果,例如更有效的癌症发现,足以支撑采用;此前从未被自动化的工作,可以在不先拆除现有人员配置的情况下增加新能力。
Levie 所说的“永恒”创业规则,是去做现有厂商无法自然吸收的事情。叠加在 OpenAI 或 Salesforce 之上的薄层产品不是好赌注:“你应该预期 Salesforce 会成为 AI-for-CRM 系统”;Workday 和 ServiceNow 也会为各自领域构建 AI 产品。
机会仍存在于跨应用工作流、需要大量非 AI 界面,或与现有厂商战略正交的领域。Levie 认为,独立的智能体安全、合规和红队测试可以成为可行方向;而覆盖 Salesforce、HR、ERP 和 Box 的供应商,则可能占据单一应用未必拥有的中间层。
7. 定价与生产率会发生变化,但扩散速度决定时间表
对于 Klarna 关闭几个记录系统的报道,Levie 的看法刻意保持混合:这一说法可能被过度渲染,但技术上并非不可能。为了节省几十万美元而自研一个 AI Workday 替代品,“极具挑衅性”,但他怀疑 90% 的企业是否会把这件事列为优先事项。
AI 更接近最终结果,重新打开了定价空间。获客智能体可以按每条线索收费,按生成 10,000 条线索所消耗的算力收费,或通过固定订阅吸收用量波动;在按席位 SaaS 主导超过 20 年后,Box 仍在测试不同模式,没有给出一套普适建议。
在 Box 内部,新员工学习销售产品时,可以查询销售 Hub,“就像在和公司里最顶尖的专家交流”,并且全天候可用。据称,编程生产率提升从 5–10% 到新员工可能达到的 50% 不等;实际 KPI 很简单,就是交付更多软件。
如果从几年后开始计算,Levie 愿意押注 AI 最终会让年度生产率提升超过 0.5 个百分点。即便即时获得某个主题约“第 90 百分位的专业能力”已经让技术突破变得不可忽视,通过“人类介导的环境”完成扩散,仍会比预期更慢。
The rate of change that we're seeing and the rate of just exponential improvement that we're seeing from AI models is incredible. I would assume that if we keep that pace up, these systems will increasingly be able to perform any kind of general task.
Jensen at Nvidia put it the best: effectively, the IT department becomes the HR department of AI. That just opens up so many new questions about what the future of IT looks like. I think we're entering a new era with systems of intelligence that let us combine data, AI, and underlying enterprise software to automate really anything about our business.
There's going to be a tremendous amount of AI startup opportunity, but it will not come from just doing an AI-first CRM system, because you should anticipate that Salesforce is an AI-for-CRM system.
Aaron Levie, founder and CEO of Box, welcome to The Cognitive Revolution.
Thank you. Good to be here. A lot of stuff going on in AI land right now.
How are you doing?
Never a dull moment, that is for sure.
Let's do a quick warm-up, and then we'll get into what's going on with AI in the enterprise and what you guys are bringing to the enterprise with your latest AI features. For starters, what are you doing with AI in your personal life, and what is your AI worldview, particularly as it relates to whether we're going to get AGI soon? What are you expecting over the next couple of years?
It's hard to say, obviously, based on the very amorphous definition of AGI. I'm downstream of whatever Ilya, Sam, or Greg are talking about, so I have no better ability to predict where that's going than anybody else right now.
All I would say is that the rate of change and the rate of exponential improvement we're seeing from AI models is incredible. I would assume that if we keep that pace up, these systems will increasingly be able to perform any kind of general task that you give them. With these reasoning models, we're seeing incredible results in math, very complex logic, and obviously coding.
Once you have those basic foundations, if we can continue to improve more and more on the benchmarks, that gives you the building blocks for models that can continue to learn on their own and get as advanced as we need them to be for basically any task that you'd want to give them. It seems like within the next few years, we could be getting some form of whatever we would have previously defined AGI as.
Yeah, not that much further to go, I would say. Certainly, the checkpoints seem to be coming, if anything, more densely than they were not long ago. How about your day-to-day use? Do you have favorite use cases? Do you feel like you're properly challenging AI Pro, or are you doing more basic stuff?
Partly because I always get enamored with new technology, I try everything out. One thing I was reflecting on recently is that I have more new apps on my home screen in the past 6 months than probably at any other time in the past decade or decade and a half.
My home screen used to be, “Okay, you added Uber, then you added Spotify, maybe one social app, and then WhatsApp.” Only every 1 or 2 years did something get to the home page. Recently, I have at least 5 new apps that I've added into the mix, which to me is a proxy for just how much infusion of AI has already occurred in our personal lives.
I'm talking to Gemini Voice and OpenAI with ChatGPT voice and video pretty regularly, using Perplexity to get different kinds of answers, and playing with xAI and Grok. It's incredibly exciting to see the range of what these AI models can do. I built a prototype with Artifacts and Claude. Everywhere in my personal and professional life, AI is getting added into the mix.
You recently put out a LinkedIn post that I thought was pretty interesting, where you basically said that this is the most energized you've seen enterprise companies about a new technology at any point in your career. Obviously, you made your company on the cloud wave, so that's another big wave that people were pretty enthused about.
For people who aren't in the room and aren't hearing these conversations with senior leadership at enterprise companies, what is the vibe like? What are they doing? Are they hands-on? What are you seeing?
Given that you brought up cloud, it's actually a really interesting comparison between the two. In the early days of the cloud, I would not describe the energy as super excited, super animated, and aggressively pursuing moving to the cloud. Most of the conversations we had with enterprises involved some degree of skepticism, resistance, and friction. There wasn't a lot of pure excitement with no friction around it.
The reasons were that this was a very big shift for enterprises. They had to move their infrastructure from the data centers they managed to the cloud. They had to trust these new vendors that they hadn't worked with before. If you were an enterprise, you'd never worked with Amazon as an enterprise vendor, which was obviously very atypical. You also had to change your whole notion about data privacy and compliance, from a world of physical management of hardware to a world where you get a set of APIs and maybe some dashboards and audit reports, and that's your only way of controlling these systems.
Enterprises were very reluctant early on to move to the cloud. If I compare that to today, and I snap the line at 2 or 3 years into cloud versus 2 or 3 years into AI, the environments couldn't be more different. Enterprises are looking for almost as many use cases as possible in which they can deploy AI, probably in many cases more than is actually practical.
There's a sense of creativity, excitement, and innovation that didn't necessarily exist with the cloud. With the cloud, it was very pragmatic. It was, “I could take this piece of infrastructure and move it into a virtualized environment.” There's nothing particularly exciting about that. It's very utilitarian.
With AI, people are saying, “What if we could actually solve this business problem that we've never been able to attack before?” Or, “What if I could deploy human resources against much more interesting problems than where my talent is currently dedicated?” All of a sudden, the creativity, energy, and opportunity are totally different for these enterprises.
There's a big asterisk, which is that we're still very early. The actual amount of at-scale deployment in the enterprise is still in the very early innings in terms of where this technology has been deployed. But the excitement level is totally different.
To me, as someone in enterprise software, this means there's going to be an insane amount of opportunity. We're going to see opportunity for bigger companies like Microsoft, Oracle, and Google. There's going to be opportunity for the software stack, like Box, Salesforce, and ServiceNow. There's going to be a lot of opportunity for brand-new startups to build AI agents that solve a whole new set of problems for the enterprise.
I think there's a relatively narrow window of opportunity right now. You're going to see a tremendous amount of change and a tremendous number of new startups emerge, and enterprises are very much ready and excited to pursue that.
What do you think that change looks like in practice? One of the comments that you made in that post was that you expect IT departments, in a way, to undergo almost a reversal from the cloud.
In the cloud, it was, “You aren't going to manage this physical stuff anymore. More of that responsibility is going to shift to the vendor in terms of making sure we have uptime, and we'll just consume that and build apps on top of it.” But if I understand you correctly, you're saying almost the opposite when you say IT departments are going to go from supporting the work to actually doing more of the work.
What does that mean, and how many enterprise IT departments are up to that challenge today? How much are they going to have to transform to take on that challenge?
I think enterprise IT departments are going to have to transform pretty dramatically. The history of IT was that, many times in partnership with—and even proactively with—the business, the IT department would work with the business. By “the business,” I mean the marketing team, sales team, finance team, or somebody else in the business who had a particular need.
They might say, “I want to have a CRM tool so I can track all my leads,” or, “I want an HR system so, as we hire more people globally, we can make sure we're managing all the local laws from an HR standpoint.” They would go to the IT team, review a set of vendors, and then basically hand it off to IT to manage the system, deploy the technology, pick the vendor, and enable the business.
That relationship has been well-defined and codified for a few decades as we've had modern IT environments. But it was still the responsibility of the business to make use of the technology, be productive, and drive execution in the company. The sales team still owned all of the execution, and they were using a CRM system from the IT team to be productive.
AI flips that on its head. Now it's not just the people in the business who are using the technology to enable the business. The business is going to go to IT and say, “I need you to deploy AI labor against different kinds of problems that I'm dealing with.”
You might have a future where the head of sales goes to the IT team and says, “I need to spin up a new sales campaign,” if you have AI agents that can help do that. Or a back-office processing team might go to the IT team and say, “Do you have AI agents that can review all my contracts, review all these invoices, or help with all my client-onboarding workflows?”
This is a totally different era for IT: being in the position of actually solving the business problem, not just deploying the software to help the business solve its problem. This means you're going to have to understand the business much more if you're in IT. You're going to have to become even more strategic for the company. You're certainly going to have to understand all the trends happening in AI and the ecosystem, including all the surrounding companies that produce AI agents and the software around them.
Jensen at NVIDIA put it the best: effectively, the IT department becomes the HR department of AI. That opens up so many new questions about what the future of IT looks like, all of which are much more exciting than the past. But we are in for quite a bit of change in this space.
If I torture a metaphor here, if IT departments are the HR departments of AI, then Box is applying for, or looking for, a promotion in a lot of these organizations. You've just been through what you might call an upskilling, with a bunch of new generative AI product and feature releases.
Tell us what's new as we enter 2025 at Box.
We work with about 15,000 customer companies, and what they use Box for is to secure, collaborate on, and automate workflows around their content. That could be their contracts, financial documents, marketing assets, research data, or product plans. We store all those documents, media assets, digital images, and other content; help secure it; automate workflows around it; and enable companies to collaborate on that data.
AI is the next frontier of what we can do with all of this unstructured data in the enterprise. We built a layer called Box AI that connects AI models to enterprise content in a secure way. That lets the enterprise continue to have its security, privacy, access controls, and permissions. We handle all of that, and then connect AI models with this abstraction layer.
What we've been building out is a set of capabilities to let people transform and unleash the power of their data. The first big use case is being able to talk to your data for the first time. We store well over 100 billion files in Box, and you can imagine that every single one has incredible insights and value inside of it. But most of the time, you don't know what's inside your data because until you search for it and look at it, you don't know what's actually out there.
Now, for the first time, you can just talk to your data. I could take 100 sales presentations and ask, “What's the best practice for selling this product to a new customer?” Our system will read through all those documents and produce exactly the right answer for whatever you're doing with that customer.
You could take a bunch of medical research or life sciences documents and research and talk to all that data to find interesting trends or patterns in that information. That first use case is basically retrieval-augmented generation over many documents, with an AI layer that powers it.
The next big use case, which we've just started working on, is the ability to read through any kind of content and extract the most important structured data from that content. Take a contract: you want to pull out the renewal date or the names of the parties, and then store that in a structured database. That way, you can query it later, create dashboards, and automate workflows. That's the next big use case we're working on, and it's rolling out imminently.
Where this is all going is eventually to agentic workflows within our platform or connected to other systems. The very big prize in this space is taking all of the information in your enterprise and having agents operate on that data.
I could have a contract agent, a marketing agent, or a sales agent that performs tasks to make me more productive. It could review a contract, pull out the riskiest clauses, and route it to the right person. I could design an agent to perform those workflows across my entire data set and then connect it to another system.
I might want to connect it to Salesforce, ServiceNow, or Microsoft. Box becomes a layer for managing intelligent content, and then we connect to other technologies that handle the AI in their particular part of the workflow.
This is the future we're building out. I think it's the biggest set of changes and the most we've built in the history of the company. It's unbelievably exciting because on a daily basis, you're seeing new things that were never possible before with our data.
Let me ask 2 questions, digging into the retrieval and then the agents, one by one. Retrieval-augmented generation has obviously been a huge trend, and a lot of people and companies have rushed out and implemented a version of it. I've heard over and over again that it hasn't worked super well for people.
When I get under the hood and explore why it isn't giving people the answer they want, it seems like the most common failure is that the vector-embedding-mediated search isn't retrieving the right content in the first place. If you don't have the right sources, you don't get the right answers. What have you done to deal with that? Perhaps structuring the unstructured data is part of it, but how do you address that problem, which seems to pop up so often?
That's the first trough of disillusionment for so many people. Let me ask you a question: when you see that trough of disillusionment, is it often with data sets that are quite broad or heterogeneous?
We've seen this a lot with the use case of, “I have my email, my data, and my calendar, and I want to search across all of it with some kind of AI system.” Are you seeing it in those kinds of scenarios, or in other ones?
Honestly, I feel like it almost always pops up, even if it isn't a crazy complicated situation.
I'd love to attribute this to pure genius, but we got lucky. We were building a product about a year before ChatGPT launched. It's now called Hubs, and it was the ability to organize content on a topic-by-topic basis so you could share or search that content on a per-topic basis.
The novel idea was that within your Box account, you manage files inside folders, but you could create a hub that points to content within your Box account. The whole power of it was to create a many-to-many relationship. I could have the same documents show up in 20 different hubs without ever moving them or changing their permissions. If I made one update to that document, it would propagate to all the hubs where it was accessed.
We created this architecture because we saw that a lot of people wanted to create a sales hub that pointed to sales data. Then they would say, “I want a sales hub for the Japan region, and I want to create a sales hub for a particular product line. I don't want to change the documents, fork them, and keep track of why the one in the Japan hub is updated but the one in the product hub isn't.”
The idea was to create virtual pointers back to the same source of truth, with virtual hubs that you could create. We were working on that about a year before ChatGPT, unrelated to AI, with search and discovery as common use cases. As soon as ChatGPT launched, we thought, “Obviously, the next big thing we should do is let you talk to all the data in a hub.”
What we discovered was that this was a breakthrough architecture for a RAG use case. Where RAG runs into problems is when you have, for example, 100 million files inside your enterprise and somebody asks, “What was the last revenue figure in our last quarterly earnings?”
The challenge is that you might have a document called “Earnings_Final,” another called “Earnings_Draft_1,” and another called “Earnings_Draft_1_Sally_Edits.” The vector embeddings on all of those documents look very relevant to producing the answer, but the accuracy and authoritativeness of any one of those versions might be totally wrong based on where it was in the editing process.
Now multiply that by years of data and signals from other systems, such as email or other data sources. All of a sudden, it becomes very hard to ask any generic question of your internal enterprise data.
This is less of a problem for public services like Perplexity because you get the benefit of almost a PageRank algorithm for public data. You can look at a curve and say, “This CNN article is more authoritative than this random blog,” or, “This blog is more authoritative than this random website that was created 2 days ago.” You get an authoritative score on the public internet in a way that corporate data doesn't really have.
Corporate data is much messier. It tends not to have a PageRank element, so it's very hard to know the ultimate source of truth. Hubs basically solve this problem because users tell us what the authoritative copy of the data is that they're putting into a hub.
If I created a sales hub with sales presentations and product information, I would put only the authoritative records into that hub. When you're asking a question in that hub, you're asking questions about sales. You aren't going to ask the sales hub an HR question; you're going to go to the HR hub.
We've been able to get the user to ask the right types of questions where the data set is authoritative and isn't as messy as doing a broad-based RAG environment over all your data. Those are 2 or 3 different hacks we've locked into and doubled down on, and they've made our particular RAG service, I think, at least 100 times better than just doing a broad-based deployment across all your data.
That's interesting. I've heard so many stories about how somebody just happened to be building the right thing that was perfectly complemented by AI, and then the whole business shifted as a result of those things coming together.
We pinch ourselves. If we hadn't been building that, I get very scared. It was about a year to a year and a half of deep architecture work. There was no way to build it any faster. You had to create a file system that could have a virtual sharing component, and for the complexity of our platform, that took a year or more.
If we hadn't already been a year or a year and a half into it, I fear we would have felt that it was too daunting based on how fast AI was moving. I'm not sure we would have landed on this exact architecture. We might have done a less optimal version. In this case, we got totally lucky but ultimately built the exact right thing for the exact right moment.
On the agents: 2025 is the year of agents. We thought it might be 2024, but it turns out to be 2025. What does an agent mean to you?
I have my own sense of what the definition is, and I contrast an agent with what I tend to call an intelligent workflow based on how much autonomy or decision-making discretion the AI ultimately has. How do you think about it, and how much autonomy or discretion are you giving the agents in your platform?
I like that definition, and I would fully subscribe to it. In our platform, we've taken probably the broadest definition because we didn't want there to be 17 different versions of a thing, with the user having to understand all the differences. We've defined an agent as a mix of an AI model or many AI models, a set of tool use within the platform, and underlying skills or capabilities.
Those capabilities are some mix of system prompts, a proprietary architectural change, and access to your data. That whole collection of capabilities creates an agent.
Within our platform, you can do very simple things with an agent. You can talk to a single document, and you're using an agent, but it isn't doing anything agentic. You're talking to either the core Box AI agent or a custom agent.
You could create a sales agent with a custom prompt and custom instructions that contain information about how your sales workflows work or the kind of language you should use. You can create custom agents to let you talk to your data, extract metadata from documents, talk to many files at once, or create content.
That's our first era of agents. Many of these agents are what we would have called assistants 2 years ago, but we've created a universal language for this. Where it's really going is much closer to your definition: agentic workflows where the agent performs many sequential tasks and there are some probabilistic elements to those tasks. For example, “I want to review a document, and based on how I review that document, I want to kick off another process, either to another agent or to a human.” Or, “I want to take a lot of data, collate it, and produce something.”
These are much more multistep flows that we expect to make very agentic in the future. Right now, the limitation of software is that it's very good at deterministic workflows, but the vast majority of work is not deterministic. It requires judgment, and you change your answer based on other inputs or insights.
The majority of work in the future will be nondeterministic, judgment-oriented, agentic work. That's an incredible opportunity for enterprise software because that's what we can finally digitize.
The way I've been thinking about this recently is in terms of the eras of enterprise software. Forty years ago, we had the initial wave of systems of record, such as CRM and ERP systems. This was the definition of the most deterministic technology possible. You're effectively changing rows inside databases, and based on that, it kicks off a process. That was 100% deterministic.
Then we moved to systems of engagement. This was the idea of collaborative systems—Slack, Box, and other tools—where things are a lot less structured and deterministic. The workflows are more fluid and can adapt a bit more because they're human-to-human and a little messy.
Now we're in a new era of systems of intelligence. This is the era in which AI automates those workflows. These systems have all the properties and benefits of being structured like systems of record, but they have all the flexibility of a system of engagement. They can adapt and change because that's what AI can do.
I think we're entering a new era with systems of intelligence that let us combine data, AI, and underlying enterprise software to automate really anything about our business.
I often say that intelligence has been debated for a long time, and I don't pretend to be the final word, but my working definition of intelligence is the ability to do useful work when there is no explicit algorithm that tells you what every step ought to be.
I think that's a great, very practical definition, especially for enterprise use cases. You don't want to have a predetermined or prewired workflow because it's going to change. You need it to adapt to new data. Maybe there aren't even APIs available for the thing you're trying to do.
That's where even browser access gets very exciting. There are a lot of future workflows where I want to automate 5 systems and have them talk to each other, but there aren't clean APIs for making them all communicate.
I find myself in my AI-assisted coding workflows, when they're working well, basically copy-pasting things around most of the time. When they're not working well, I start to have to troubleshoot. The smooth thing is that I'm just the copy-and-paste monkey gluing these other things together. It's a weird experience when it's working well.
What would you say are the biggest bottlenecks that enterprise companies face today in realizing value? I have a thesis that the AIs are good enough to do a lot more than people are actually deploying them to do. Tyler Cowen recently said, in front of an audience, “All of you humans, you are the bottlenecks.” What do you see as the practical barriers that people are still struggling to get over?
I 100% agree that AI can do vastly more than most enterprises either think or have the near-term appetite to deploy. To some extent, there's more technology available at the moment than most people realize.
Even I have to remind myself: that task I would normally have asked somebody to work on—I'm reminding myself more and more, “No, go and try to create it in Artifacts or use AI to do this.” Even right in the center of AI, watching everything happen around the world, I have to trigger my brain to remember how much these things can do. I can only imagine that if you're not in it every day, that gap is probably fairly massive.
That being said, I don't want to let the AI model providers off the hook. One thing that prevents AI deployment is that you can't have an enterprise workflow, particularly in a regulated industry, that works 98% of the time.
You wouldn't find it acceptable if 98% of the flights you scheduled were successful, but 2% of the time you showed up at the airport and didn't actually have a ticket. Enterprises need 99.99999% reliability on almost anything that's important.
If it's a creative task—write a marketing campaign, review a blog, write a blog post, or send an email on this topic—you have some room for hallucination, or you can edit it later. But if you're going to have a billion financial transactions submitted every week or month and AI gets 0.01% of them wrong, that's a nonstarter for deploying those systems.
What we need from the model providers—and you can see it in the benchmarks—is for accuracy to keep getting higher. We need to get to the point where all the evaluations have to be completely reset because everything has hit 100%.
I think we're still a little bit technology-dependent and AI-model-dependent. We also need the cost of AI to continue to come down. The excitement from customers is real: they'll say, “I'd love to deploy 10,000 AI agents against this problem,” but then they'll look at the cost and say, “I still can't afford what that would look like in my business.” Maybe we have to take it in a more stepwise fashion until costs come down.
We need AI to be cheaper and the performance of these models to get even higher. From there, you're running into Tyler's point: all of the classic, human-based change-management difficulties.
Those range from privacy and security concerns, in some cases, to the very real issue of, “I still have humans doing that thing, so until I can transition them to a different role or teach them a new skill, we're not going to be able to automate that particular workflow.”
All of that is still what we're in for in the enterprise, and it's going to take years. This is going to be, at a minimum, a decade-long change in how enterprises become more AI-first.
But I think we have a roadmap because we did it with cloud. We also have an increasingly clear vision because we're starting to understand what this world could look like: What would an AI-first enterprise look like? How would agents be part of the workforce? How do we get better insights from all our data? How do we automate almost any workflow in our enterprise?
I think you see an increasing understanding of what an AI-first future could look like in most organizations.
To push a little harder on this, I feel like a corollary to my definition of intelligence is that you should only use intelligence where there is no explicit algorithm. If you have a way to do something with traditional code, you probably should. It will be faster, cheaper, and more reliable.
If we bring AI back into the domain of all these fuzzy things for which we don't have algorithms, and ask how good AI is at doing those things compared with a human, my belief is that with some elbow grease—making sure you have the right context, putting a few examples together, and applying the best practices—you can most of the time get to the point where AI can do as well as a human at a much lower cost and much faster.
We've seen that in medical diagnosis recently. If you buy that, doesn't it suggest that something else is going on? Maybe it's a fallacy—or not necessarily a fallacy, but an attitude—that we need AIs to be not just on par but to have a 10-times-lower defect rate or something. That seems to be the case with self-driving.
I think you just nailed it. You can't go to a company and say, “I can do exactly what you're doing today, and you'll save 40%.” An economist would say, “Everybody would do that deal all day long,” but once that meets real life, the person has 17 other projects.
There is an incredible number of people, priorities, and demands all competing for their time. If you could wave a magic wand and make something 40% cheaper, you would totally do it, but of all the things you have to do, that might be number 9 on the list.
To your point, there's something else going on. That person now has to go to the AI council and get approval. They have to run a full 6-month test to make sure that it actually is 40% cheaper. They have to weigh that against everything else they're doing in their business.
I think we need AI to produce multiples-better improvements over the status quo. That's how you compel motivation. You can't be incrementally better, cheaper, or faster. You have to be an order of magnitude better on one of those dimensions.
If you could go to a company and say, “I can be literally one-tenth the cost of what you do today to review your contracts, review your invoices, or automate this back-end supply-chain process,” then you're talking. You're saying, “I could save you millions of dollars.”
Or you could say, “I can do a 10-times-better job than your human-based workflow today. We'll discover cancer more effectively, or we'll target even more automation across your enterprise.”
This often explains AI's biggest opportunities. I think the flavor of what you're saying is that you go after work that isn't automated today. You're not even replacing something that already exists; you're layering onto an existing workflow and making it better.
I would say this largely explains the breakthrough in Copilot or Cursor. I get to do exactly the same thing I'm doing now, but I see incremental productivity gains immediately without really changing any behavior.
The more AI can solve those problems—net-new use cases that add incremental productivity—the easier it is on the change-management front. Anything that replaces an existing process and saves only a little money is a much harder problem to go after in the enterprise.
What do you make of the debates around the future of enterprise software? We've heard conflicting narratives because everything is happening so fast.
On the one hand, it was, “It's never been a better time to start a startup.” Then it was, “Actually, this technology is so easy to implement that incumbents will probably capture most of the value because they'll be able to roll it out to their existing customers.” They might have a little bit of a go-to-market lead, but incumbents already have the customers, so they'll beat you if you're trying to start up rather than having them figure it out and deploy it to their current customer base.
Then we have the Klarna narrative, where they've allegedly or reportedly shut down a couple of systems of record. There's also the pricing debate: do we still charge by seats, or do we have to move to per-outcome-based pricing? There's a lot there that could probably consume the rest of our time. What do you make of all that?
Exactly. If I had heard that when I was just starting, I would have been way too stressed to start a company.
There are definitely a lot of variables in flux. I compare that with our early days, when we had plenty of other problems but not the fundamental variables of the company's business model.
I think you're right. You have to ask which spaces give incumbents a natural advantage, whether the underlying billing model is seat-based or outcome-based, and whether there's a future of some kind of AGI-light that makes some software irrelevant because you don't need the software in the first place.
All of those things will happen. I would lean more toward timeless lessons of competitive strategy. If you're a brand-new startup, go after things that aren't easy for the incumbent to pursue.
If all you're doing is building a thin layer on top of OpenAI, that's a bad idea. If you're building a thin layer on top of Salesforce with AI, that's a bad idea. Salesforce is very competent; it will build the CRM AI product. Workday will build the HR AI product. ServiceNow will build the ServiceNow AI product.
Equally, if you're just finding a slight gap in what OpenAI does today, you have some risk of them moving up the stack, or the model getting better and eventually bringing that capability into the model layer.
That being said, I can think of a number of things that would be unnatural for OpenAI to do because they might involve a lot of non-AI interface and workflow work to solve a particular problem. You could be doing things inside the sales, HR, or IT service-management worlds that the incumbents equally aren't going to do.
Maybe it's cross-platform AI workflows that aren't natural for any one of those players to pursue. Maybe it's building AI agents that are so orthogonal to the normal strategy of Salesforce, Workday, or ServiceNow that those companies wouldn't think to pursue them. You can get enough traction fast enough to create some degree of a moat.
I think there's going to be a tremendous amount of AI startup opportunity, but it won't come from just doing an AI-first CRM system. You should anticipate that Salesforce is an AI-for-CRM system. We're seeing startups all the time that are finding those windows of opportunity right now.
What do you make of Klarna? I've heard every take on it, from “It's all hype; they're not really doing it,” to “Maybe, but it's the exception that proves the rule.”
Right now, it's the exception. As I've seen the reports, I'm more in the camp that maybe it's overplayed a little bit, but nothing about it is impossible.
Given that it's not impossible to do what they've said, maybe they've chosen this as something that will differentiate them as a company. It's still different from every company on the planet to do it in a homegrown way.
My understanding was that they were going to build their own Workday system with AI, and that's just not a priority. Again, where are you in the prioritization stack? Most companies simply aren't focused on building their own HR system to save a couple hundred thousand dollars.
I think what they're doing is super provocative and super interesting, but not translatable to the broader economy. It's fun to watch. I invite as many companies as possible to try that experience and share their lessons along the way. It makes the conversation in the ecosystem much more interesting and dynamic.
But I'm not convinced that 90% of corporations would ever do what they're doing.
Any pricing guidance?
I don't really have any guidance because we're testing all the business models ourselves. In general, one of the biggest benefits of AI is that you can get much closer to software solving an outcome.
The conclusion of that theory would be that your pricing model should be closer to that outcome. That could be the outcome very literally: you pay an AI agent to generate leads, and therefore you pay per lead. Or it could be the consumption that goes into that outcome: “I want the AI to generate 10,000 leads, and that takes a certain amount of compute capacity,” so I'm paying for the consumption of that compute capacity.
Then there are traditional subscription models: “I want an ongoing license that will roughly do this much volume for me. Sometimes it's a little more expensive, sometimes it's a little more volume, sometimes it's a little less, but I pay the same fixed rate the entire time.”
I think we'll see every version of these business models pursued. I put it in the category of something that's intellectually interesting because we're in such a dynamic period. We haven't had open questions about business models in software for 2 decades.
Salesforce, perhaps with a couple of other companies, essentially invented the idea that you pay per seat on a subscription basis. That's been the business model of SaaS for more than 20 years. Now we have a chance to say, “There are other business models that will begin to emerge.”
I find that incredibly fascinating, but each company has to develop its own understanding of what its customers are looking to pay for. That determines which business model makes the most sense.
One niche that's really interesting to me, in terms of things people might want to buy separately, is safety or compliance. Maybe they're going to continue being a Salesforce customer and get all the agents or whatever, but perhaps they might want to buy something separately from somebody who red-teams the agents. Do you see that as viable?
Absolutely.
That's interesting, because the other option would be for Salesforce to deliver that as another feature.
It's really important to understand the layers of the stack. If you think about an IT stack—Salesforce, HR, an ERP system, Box—which things cut across all of those? That's where you need a new vendor independent of any one of those players.
Some things make sense to exist in one of those systems. Other things are more likely to be technology that works across all the apps in your enterprise. That determines whether you could be a startup or whether it's really an incumbent game in that market.
That's a good perspective. One last question: it's been famously said that we see the impact of computers everywhere except in the productivity statistics. It seems like AI may still be in that zone.
What are you seeing internally at Box when it comes to AI-enabled productivity boosts? Is that something you can measure, or is it something you believe in and encourage on faith right now? What do you have for other leaders who want to make sure they're getting the productivity that's promised?
We're 100% committed to being an AI-first enterprise and company. That's particularly important because we sell AI technology to enterprises. We need to be the first to understand where this is all going, but I also think it's going to be a way to run a better company in the future.
It's showing up in a handful of ways already. With Box AI, this is how we work with our unstructured data. If you're a new employee and want to learn how to sell our product, you go to our sales hub and ask it any kind of question. It gives you an answer back, and it's basically like talking to a top expert in the company.
You're getting all the value of talking to the smartest existing employee, but now you can do it 24/7. You don't have to wait for somebody to respond to your Slack message. That's probably less measurable because it permeates everything we do and simply improves productivity.
In other areas, it's anecdotal, but we've deployed AI coding tools. I'll get ranges of feedback: somebody will say they were 5% or 10% more productive, while a new hire might be 50% more productive because they can ramp up so much faster.
I think the biggest way this will show up is that we'll ship more software. That will be the measure of productivity we care about internally. I think AI will certainly be the first distinct technology category to show up in the long-term GDP graph that people talk about as driving productivity gains from technology.
So, more than 0.5% a year, which is what Tyler Cowen recently said? You'd take the over?
If you let us start the clock in a few years, I think so. You still have diffusion of the technology across the economy, and that takes much longer than I ever think it should. But that's just life in human-based, human-mediated environments.
I'm with you on that. I always underestimate the timelines. Anything else you want to touch on or leave with the audience before we break?
We're in such an unbelievable time to be building and deploying technology. The fact that anybody, anywhere in the world, could get 90th-percentile expertise on any topic instantly is an unbelievable thing.
If you had contemplated that 3 years ago, based on our understanding of technology at the time, it wouldn't have been conceivable. I just think it's an incredibly exciting moment, and we're super excited to bring it to the enterprise.
Terrific stuff in both respects. Aaron Levie, founder and CEO of Box, thank you for being part of The Cognitive Revolution.