Box CEO 谈 AI 采用鸿沟 | The a16z Show
Erik Torenberg × Steven Sinofsky × Martin Casado × Aaron Levie
- 企业采用 AI 的鸿沟,与其说由模型能力决定,不如说由权限、责任归属和运营控制决定。 初创公司没什么可“炸掉”的;银行则必须先控制提示注入、误写入、智能体相互冲突和信息泄露,才能授予自主权。结果是:“AI 能力的扩散会比硅谷人意识到的更慢。”
- 如果组织最终部署的智能体数量达到人的100倍或1,000倍,软件需求可能被彻底改写。 Aaron Levie 认为,供应商必须开放 API、CLI、工具、身份体系和访问控制,因为业务表现将越来越与智能体触达公司数据的效率相关。Martin Casado 提醒,界面打磨不是护城河:智能体会依据语义、成本、持久性等实质因素选择后端,而不是依据界面或文档质量。
- 核心记录系统的防守力,远强于“SaaS 末日”叙事所暗示的水平。 Steven Sinofsky 称,“以为靠 vibe-code 就能做出 SAP,简直荒谬”,因为数十年的领域知识分散在界面、中间层、工作流和操作习惯中,而不在某个干净的数据层里。智能体改变软件消费方式和变现模式的速度,可能快于其替代核心系统的速度。
- AI 初期会抬高能拆解工作的领域专家的价值,随后会把这项技能上移到更高的抽象层。 大多数员工无法画出自己的工作流程图,使“算法思维”成为眼下的瓶颈;Casado 提到的 Anthropic 增长营销案例显示,1名系统思考者把原本分散在5个或10个岗位上的工作自动化。Sinofsky 预计,“火箭科学”那一部分会像当年的电子表格复杂性一样消失。
- 给每个智能体单独开一个账户,并不等于把它变成一名员工。 智能体可以获得电话号码、Gmail 账户、信用卡和基于角色的权限,但责任仍由所有者承担,且必须全程监督;任何进入上下文窗口的信息,都可能通过提示注入被抽取出来。对于并购数据室等敏感流程,Sinofsky 认为,在 N 变得非常大之前,企业可能还会维持只读状态数年。
- 嘉宾们拒绝华尔街把收入总盘子视为固定不变的假设,认为 AI 需求在结构上被低估。 Sinofsky 表示,预测“至少差了一个数量级”,并以 PC、云计算和 CRM 为例:这些市场中,摩擦下降带来消费扩张,而不只是重新分配支出。Casado 补充称,在他能观察到的基础设施公司中,过去6个月每一家都出现了“渐近式”增长,因为软件编写量大幅增加。
- 即使效率最终会压倒稀缺性,Token 支出眼下仍是盈利和管理层必须面对的问题。 在当前讨论中,工程算力可能占相关费用的1%-100%;考虑到上市科技公司的研发费用约占收入的14%-30%,算力成本是工程团队成本的2倍,还是仅高出3%,都可能决定 EPS。Martin 称这是未来“最疯狂”的预算讨论;Sinofsky 则预计,类似晶体管的转变最终会让今天的 Token 核算消失:“这是肯定的。”
1. 智能体数量激增,迫使软件服务第二类用户
Levie 的前提是一个数量问题:如果企业最终运行的智能体数量达到人类数量的100倍或1,000倍,软件就必须围绕智能体设计。嘉宾们表示,如今他们在设计智能体界面上花的时间,已经和设计人类界面一样多,包括 API、CLI、MCP 以及其他机器可访问的入口。
正在形成的模式,是让编码智能体接入 SaaS 工具、组织上下文和知识工作流。它可以读取信息、调用 API,或通过现场写代码完成未预料的任务——这正是 Levie 在 Claude Cowork、OpenAI 正在推进的超级应用方向和 Perplexity Computer 中看到的潜力。
Sinofsky 接受这一理论,但认为瓶颈在人:“算法思维真的、真的、真的很难。”让大多数员工把一项重复流程画成流程图,他们都会失败;在一支50人的营销团队里,可能只有1人能准确记录这些工作如何衔接。
2. AI 把工作上移到更高抽象层,而非消灭专业能力
Casado 提到 Anthropic 一名增长营销人员使用 Claude Code,把过去或许由5个或10个人分担的工作自动化。更值得注意的反事实是:一名员工身边仿佛有一个“无限工程师池”,任何定义清楚的工作环节都能被自动化。
反驳意见是,面对近乎无限的需求和供给,增长营销本来就是容易展示效果的场景。要验证这一命题,应先放到营销一台售价600美元的 PC 这类受约束、竞争激烈的岗位上测试——在那里,判断力和稀缺需求更重要,不能仅凭一个技术能力异常突出的操作者就泛化。
Sinofsky 给出的更强类比来自一位表亲:她进入银行业时,正值电子表格出现。最初,她管理一屋子负责搭建模型的实习生;2年内,她所在的那一批人就成了电子表格的使用者,能够完成约30轮迭代,而此前依赖计算器的分析师可能只能完成2轮。
这并不意味着要由一名“火箭科学家”协调42个长期存在的专用智能体。编排复杂度应当收敛为一种类似营销职能的能力,能够接受更高层级的请求;长期保值的将是系统思考和领域判断,而不是智能体管线的搭建。
3. 计算机操作可能先解锁旧软件,再由智能体重写它
Casado 持相反看法,反对以代码为中心的论点:行业轨迹已经从给 SaaS 加 AI,走到终端操作,如今又走向“计算机操作之年”。智能体越来越像人在操作现有软件;他称之为一个夹层阶段,并认为代码生成可能会退到幕后,而不是变得更显眼。
Levie 认为计算机操作、API 调用和生成代码是互补关系。Box 智能体会在现有技能、现有 Box 工具和为新操作现场写代码之间做选择;或许90%的请求应使用成熟工具,但即时生成的代码可以覆盖任何供应商不可能预先规划的需求。
Sinofsky 认为消费层的价值巨大。AI 可以调用 SAP、PowerPoint、Word 和 Excel 多年积累的功能:找到隐蔽的报表选项,或制作双轴图表;过去25年里,人类用户一直是调用这些软件能力的瓶颈。
集成是下一步。Sinofsky 举出一位前 VA CIO 的案例,把数十年的 IT 工作压缩成串联75个系统;Levie 更进一步提出“按需集成”,让运行时查询跨越 IT 团队从未预先接通的系统。企业运营者则反驳称,不受控的集成等于邀请智能体“搞坏我的记录系统”。
4. 智能体自主性带来机器速度下的协同与身份问题
Box 的 CLI 同时展现了两面。借助 Claude Code 和 Opus 4.6,用户可以让智能体上传桌面文件夹,或处理存储库中的每一份文档。但在5,000名员工的组织里,智能体可能每小时访问共享内容10,000次,同时各自移动、写入或删除同一批文件。
嘉宾们的实用模型,是把智能体像另一个行动者一样配置:给它自己的电话号码、Gmail 账户、信用卡和基于角色的权限。这样,现有身份系统就能对它加以约束,无需再增加一套完全独立的授权层。
Levie 认为,在一个有50名人类和50个智能体的50人团队里,这个类比会失效。智能体在与另一名用户协作时可能获得信息,但责任仍由其所有者承担,所有者必须能够检查或撤销它做过的一切。与员工不同,智能体与操作者之间没有有意义的隐私边界。
安全问题不只是误操作。如果敏感信息进入上下文窗口,攻击者可能通过提示注入诱导其泄露;一个助手被社会工程操纵起来,比操纵人类“容易10倍”。因此,允许智能体自主访问并购数据室,在性质上截然不同于给员工授权访问。
5. 治理使初创公司与企业的采用鸿沟长期存在
Sinofsky 将当前的不确定性与企业早期采用开源软件时的情形相比较。企业最终形成了许可、质量和贡献规范,但这些标准是在使用过程中形成的;如今,同一场争论被公开进行,行业却承受着在运营实践找到答案之前先宣布最终形态的压力。
企业客户可能会在控制措施稳定前封闭系统,而开发者、高阶个人用户和初创公司会快得多。初创公司的失控智能体,或许只会变成“一集《硅谷》”;同样的失败发生在大型机构内部,损失特征则截然不同。
Sinofsky 的总结是本期关于采用节奏的核心判断:硅谷总是从“没什么可炸掉的”公司外推,然后追问 JPMorgan 将如何部署 NanoClaw 来自动化业务。因此,在 N 变得非常大之前,企业可能还会维持只读消费阶段数年。
6. 记录系统保留着智能体无法简单再生的领域知识
Sinofsky 精准概括了 SaaS 的矛盾:现有供应商卖的不只是业务数据,还包括嵌入式智能、领域专业知识,以及支撑该职能运行的操作系统。智能体越来越需要不受限制的数据访问,而 Workday、SAP 和 Salesforce 等供应商最初并不是为此打造或定价的。
他对既有厂商的辩护是绝对的:“以为靠 vibe-code 就能做出 SAP,简直荒谬。”("It’s just absurd to think you’re going to vibe-code your way to SAP.")这些知识分散在 UI、中间层和长期积累的使用习惯中,因此,即便是只读式 AI 采用,也可能被围绕一个稳固记录系统形成的架构拖慢。
不过,在打造智能体想要的产品这件事上,Levie 已经彻底信服。经过足够多轮迭代、撞过足够多堵墙后,智能体可能会建议公司换掉旧 HR 系统;当智能体流量达到人类流量的100倍或1,000倍时,企业表现将越来越取决于智能体能否访问并使用软件。
这带来一份新的供应商清单:高质量 API、智能体身份、访问控制、机器规模下的可靠性,以及可行的变现方式。理论上,Workday 可以按 HR 记录收费;Box 预计智能体会创建和操作更多文件。其他供应商可能会在智能体承接了应用过去交付的价值后,失去相应收入。
7. 智能体按系统质量做选择,也可能重建影子 IT
Casado 反驳了“向智能体营销”主要意味着提供干净 API 或 IDL 的说法:“我认为这几乎完全错了。”智能体已经出奇地擅长根据语义、成本和持久性选择基础设施;它们看重的是后端的实质,而不是面向开发者的界面打磨。
Levie 的说法与此并不冲突:一个封闭工具最终会退出候选范围,因为智能体会推荐更好的数据库或服务。类似 Gartner 的选择机制可能直接进入工作流;当智能体无法有效触达供应商时,这些供应商就会直接出局(DOA)。
嘉宾们开玩笑说,硅谷很快会用付费推荐腐蚀这套按实力竞争的机制——机器版的牛排宴。严肃的一面是,智能体主导的采购会同时创造新的分销渠道和新的付费影响入口。
Sinofsky 警告,智能体可能会在 IT 视为中间件或终端用户工具的领域里,搭出事实上的记录系统:“宏最终会运行整家公司。”与“从提示词直接生成机器代码”的愿景相反,他预计这些层仍会存在,因为它们承载策略、安全、兼容性、组织边界和状态,而不只是过时的界面。
8. 智能体原生服务与机器采购扩大市场边界
Levie 最感兴趣的是从第一性原理重建的服务型企业:营销机构、工程咨询公司、律所,以及建筑或施工设计公司。由于内部信息壁垒较少,这些企业可以向智能体开放广泛上下文,并为单个项目生成软件,从而为一种新的企业设计方式提供案例。
这些公司也不可能永远摆脱经济规律。随着规模扩大,它们仍会遇到地理、市场细分、分销、物理世界约束,最终也会碰到其他企业都会遇到的记录系统问题。
智能体也会消除小额采购的摩擦。信息或软件的实际使用量可能比潜在需求少100倍,因为没有人愿意为一份数据支付5美分,或为一次工具调用支付1美元;而一个有预算的智能体可以在任务中途花3美元做医学研究。Levie 预计,企业会把这类用量汇总为可预测的批量合同;Casado 则称,高 Token 消耗带来的 COGS 正在推动定价转向更细粒度的按用量模式。
9. AI 经济从稀缺出发,最终走向大幅扩张的消费
Sinofsky 表示,当前的预测“至少差了一个数量级”,因为华尔街假定收入总额不变,只是在检验 Token 和 GPU 是否值得投入。PC 曾被误判为规模有限的 MIPS 市场;云计算也被理解成把每年约60,000台服务器转移到别处,而不是让客户消耗1,000倍的算力。
CRM 也遵循同一模式:一个约20亿美元的市场,原本被服务器、Oracle 许可证、咨询服务和漫长部署周期拖累;Salesforce 消除采用摩擦后,市场反而扩张。Casado 在其覆盖的240家公司投资组合中看到了类似信号:他能观察到的约50家基础设施公司,在软件编写量激增的6个月里全部呈现“渐近式”增长。
嘉宾们仍预计 CFO 必须直面这场转变。工程 Token 支出的分配比例,讨论区间从1%到100%;考虑到上市科技公司的研发费用约占收入的14%-30%,算力成本是工程团队成本的2倍,还是只高出3%,都可能吃掉“你全部的 EPS”。并行实验也意味着,只要其中一条路径胜出,浪费90%的 Token 可能仍是理性的。
Sinofsky 对近期格局的判断是:初创公司烧掉手头资本,大公司冻结投入,少数有选择地采用 AI 的中型企业如果守住财务状况,就可能赢得份额。Levie 也提醒,不要低估本地算力作为泄压阀的作用;Sinofsky 则认为,更多产能、更好的算法和新硬件,都可能带来一个“晶体管时刻”。
Sinofsky 的长期答案是效率:单位成本会像当年的 MIPS 价格一样下跌。他总结称,今天这种 Token 卡式预算管理最终会消失——“这是肯定的”。
The diffusion of AI capability is going to take longer than people in Silicon Valley realize.
It's just absurd to think you're going to vibe-code your way to SAP. All of that domain knowledge isn't just represented in some well-orchestrated data layer.
The engineering compute budget conversation is going to be the wildest one in the next couple of years.
The biggest problem right now is that everybody is trying to figure out the economics of all of this when they're off by at least an order of magnitude on how big the opportunity is. If you have 100 or 1,000 times more agents than people, then your software has to be built for agents.
People in the abstract say things like, “Now you're marketing to agents. You're like an API. You've got a good IDL.” I actually think that's almost exactly wrong, which is—
Wow, this is breaking podcast news.
If you start to imagine that we all have to build software for agents, I think we're all clear on that, right? That trend is happening. We spend as much time now thinking about the agent interface to our tool as we do the human interface.
Sure. The reason we're doing that is because our hypothesis would be that if you have 100 or 1,000 times more agents than people, then your software has to be built for agents. What is the way those agents are going to interact with your system? It's going to be through an API, a CLI, MCP, or whatever.
The paradigm that appears to be taking off, and is quite successful so far in terms of efficacy, is: What if you give a coding agent access to your SaaS tools, and a coding agent access to your knowledge-work workflows and context? That kind of becomes this superpower. The agent isn't only capable of reading data and understanding information; it can actually code its way, or use APIs, through whatever task it's trying to achieve.
That appears to be a paradigm that's starting to compound. That was the Claude Cowork phenomenon, and whatever OpenAI is cooking up with the super app, Perplexity Computer, et cetera. I actually think it makes sense as the ultimate manifestation of this stuff.
I think you're right. It makes sense in a theoretical way. Yeah.
But in a practical way, we have to be really careful in that—
That—the way to say it is algorithmic thinking—
Yeah.
—is really, really hard for the vast majority of people who have jobs.
Yeah.
The easiest way to think about it is: If you were to go to any person and ask them to create a flowchart for a particular thing they have to do, they would probably fail at producing that flowchart.
There.
So within any organization—say, doing a marketing plan where there are 50 marketing people working on a giant product line—
One person probably understands and could document the flowchart.
100%.
So if you put one of these agents, or this Cowork tool, in front of people, their ability to explain to it what to do is really, really limited.
100%. So then you're—
But what if that becomes the new way you have to interface with computers, and you just have to cycle that through?
Well, then you're basically developing the next abstraction layer for how people interact. Developing an abstraction layer has historically, at each level of the abstraction layer, been the work of a highly skilled, very specific individual within an organization. The little parts that they build just become little toollets in the world of people doing particular tasks. Some people are able to stitch them together, and some can't.
But that happened with paper clips and thumbtacks before, and it's going to happen with whatever we do next.
I think the timeless part is that the job just moves up a rung, and you learn a new set of skills. That's why I don't think anything about this is any different. It's just that the leverage you get is obviously fantastic.
There was this viral tweet that went around about the Anthropic growth marketer. Did you guys see this? It's basically one person, and he was using Claude Code at the time to more or less automate what maybe 5 or 10 people would have done in various siloed jobs.
I think the reason it's interesting is that you had to have been a systems thinker to accomplish that. He was clearly technical enough to be able to pull it off. But it did represent what each of these jobs might look like if you had, say, a job in the economy and, right next to that person, an infinite pool of engineers who could automate whatever that person wanted. What would that job look like in the future as a result of the automation that's now possible? Yes, I agree that you'd have to find a way to think through your job as a system to be able to pull that off. Maybe the agent gets better and better over time at being able to nudge you in that direction. But it does sort of stand to reason that you will start to try and automate a lot of that kind of work: why don't I take the keywords that are working in Google AdWords and port them over to Facebook, make sure that those are replicated, and then take in the new signal from what's happening in the market?
That's a big leap. One thing first.
I almost had you. You were nodding a little bit, and then I said something that went too far.
The Anthropic growth person, as an example—that's a job where the rest of—yeah, I could do that job. It's infinite, and you've got the best—
When demand is infinite and, frankly, supply is infinite, this is not a difficult job. So let's—
The guy who runs the petrol pump in Australia right now is amazing.
Right, right. So instead, be the $600 PC marketing person and see how you can do against the Neo [?]. That's a real job.
All right, fine. We need a better example.
But there is—I mean, it is really interesting. Let me give an old example, an old-person example.
My cousin went to an elite MBA school and joined her first job. She's a little older than me. She joined right on the cusp of computing. She actually didn't use a spreadsheet in grad school, and then a spreadsheet showed up, but she wasn't a spreadsheet person. So instead, they told her, “Hire as many interns as you want.”
Her first year on the job, she supervised essentially a whole room of agents.
Yeah.
The kids who were me—not literally, but they were in college—came and did all the spreadsheeting.
Yeah. But then what happened, sort of magically over the next couple of years, was that she and her cohort all became the spreadsheet people.
Yeah.
This idea that being a manager in a bank, or just being 2 years in, meant you had a cadre of people doing spreadsheets—no. The whole abstraction layer moved up.
The old job before those interns was that you just sat there with basically calculators and an HP calculator, figuring out the model for some M&A deal or whatever. You only got to do 2 iterations before you had to put out the pitch deck or go to the customer or client. Then, all of a sudden, they were doing 30 iterations themselves.
But they see—and so I think where we are with agents is just at this step where you think you need 50, and the abstraction layer is such that we're dividing things up into these really small pieces, with 1 super-smart person coordinating them all. Pretty soon, that whole thing is just going to collapse on each other, and there is just going to be a skill set—
—an amount of code, call it an agent, that is marketing-ish. Yes.
And you'll be able to ask it marketing stuff. Yeah. Then the next step will be to have it go do things.
I'm a little skeptical that, until the whole non-reproducible, non-deterministic element of this AI stuff goes away, doing things is going to get very costly. Then you get into the human-in-the-loop discussion and all of that.
But I think we're at that exact point. I feel like, when I talk to people trying to do stuff, we're right at—I feel like I'm at Thanksgiving dinner, talking to my cousin 6 months into her job, when I'm already using a spreadsheet. I'm like, “I don't know why this is so hard. You should just use one.”
And then 2 years later, she's doing it.
I think right now you have to be an absolute rocket scientist and a growth marketer to create 42 agents, spin them all up, and do all of this stuff. But the rocket-science part of it is going to evaporate in very short order.
And then you're talking about—wow, there's a giant chunk of domain expertise that—
It goes back to the domain expert.
So I actually think something that you said—I'll take the other side of it—is that I think it's very tempting to say these agents are going to code and do X. Yeah.
But I think we're going the opposite way. I think where we started was, we'd take a piece of SaaS software and add AI.
Yeah.
And then that’s the new kind of AI-enabled SaaS. That’s the extreme version of using code for these types of things. But now what are we actually doing? We’re like, okay, the SaaS software is still SaaS software, and the agent uses it as a computer because it’s actually very good at that. So I’d say we started with code.
Then we went to the terminal, which is actually less code. Yeah.
And now this year is going to be the year of computer use. Yes.
So it’s almost like they’re much more like humans using computers than generating code. And that feels very much like this mezzanine step. Yeah. And I actually come from the generating-code type of world. I would argue that’s happening less, not more.
Yeah. I think, to me, whether it’s computer use, API use, or writing code on the fly, I kind of, maybe erroneously, put that all in one broad category.
They’re very different.
They’re very different. But we have an agent that we’re working on where it just makes a determination whether it should use an existing skill, use an existing tool from Box, or write code to solve that problem. Its ability to do any one of those 3 things at any moment ends up being incredibly useful, because sometimes there’s just some specific operation you want to be able to do where writing code to perform that operation is faster. We can’t possibly pre-plan for everything that anybody would ever want to do on their documents.
The fact that the model is good enough to also write code on the fly for that use case ends up being an amazing property, even though maybe 90% of the things that it’s going to do should just be using an existing tool. Over time, there are literally, like, 7 apps on her iPhone—7 SaaS apps—and we end up, over time, consolidating these things.
But the 7 apps on the iPhone are an issue of humans not wanting to learn these things over and over again. As a human, I don’t have the mental bandwidth to learn that many apps, but an agent that is going to use tools and APIs and be able to code things doesn’t have any of the same constraints that we have. So I don’t know; I don’t mind—
Well, you could argue that there are just so many things to do, and you can make the interfaces sufficiently general.
Yeah, fair. Let me think. I like what you said, because I’m back. We’re back. Okay, we’re aligned.
We’re aligned. We’re aligned. No, but I think there’s something super interesting here, which I really do like: where software has evolved. I use SAP all day. I work in finance. I have to generate all these reports, and then somebody shows up and says, “I want a report that does this—view it this way, slice it this way.” I’m like, “Oh, God. I don’t know how to make that.” Now I have to wade through the SAP help system and try to find it.
One thing that AI could be very good at is navigating that surface area much, much better. The help is all there, so it’s a matter of finding it and mapping language to it.
Just go to the ribbon.
Humans have been a bottleneck in tapping into the past 25 years of software capabilities. I spent my life sitting next to people on airplanes saying, “How can I make PowerPoint do X?” It hurt physically to watch somebody suffering with bullets and numbering in Word, or trying to figure out how to make a 2-axis graph in Excel—which is rocket science. Almost no one can do that, but it’s super common. People have not been able to use all these capabilities, and that impedance mismatch was a human user interface.
I totally buy it on the consumption layer. I totally buy the perfectly fluid UI, or consumption layer. I just feel like the backend—the systems of record—will probably converge into some database, some generic set of APIs that they’ll connect to. That seems to be the direction it’s going.
I agree. I think you’ve got to start.
I spent all weekend implementing my NanoClaw bot. When you first start out, it’s like you’re building an integration for everything. NanoClaw is very different from OpenClaw: OpenClaw has all of the integrations, while NanoClaw has very few of them, so you haven’t built all of its own tools.
But after 2 or 3 days of this, you kind of have the tool integrations that you need.
Yeah, but back to the—I mean, we’re talking about personal productivity. You’re organizing your life or something.
Okay, fine. Work productivity, and then an SAP system.
There’s an infinite amount of complexity when you get to a company that has a global supply chain and is dealing with 75 pieces of information across 30 different systems. That requires a certain amount of horsepower from the agent that we just haven’t been able to get from any architecture up until now.
But what you just described is literally what it has been doing for 50 years and will continue to do. I have a friend who was the CIO of the VA, and all he spent his time on was gluing those 75 VA systems together. It’s all just integration and redundancy.
Perfect for integration. Yeah, I totally agree. Okay, great.
For integration, these things are the best. But it’s integration, right? It’s literally, “How do I stitch these 2 systems together?”
But now the thing that I think is happening is kind of integration on demand.
It’s my new query in the system that the IT team didn’t pre-wire. Now I need it to happen at runtime.
Let me get off my lawn. Okay. The reason I say this is that I was just in a room filled with a bunch of CFOs and CIOs. They all looked at me when I said something along these lines—although not as optimistic as you can imagine—and 6 of them came running up afterward and said, “You’re insane. You’ve lost all credibility with me.”
Wait, wait. What specifically? That the agents are going to do integration?
That integration is a problem that will get a lot easier. Yes.
They were against that?
No. No one’s against practical integration.
But their fear is unleashing not just the agents themselves, but humans to do integration, because you put people in the position of creating new integrations and just say, “Please break my system of record.”
Oh, yeah. The idea that you just create a new API between System 27 and System 38—
And then you’re— That might be fine for a report, because if that person wants to be wrong, that’s their business, but you’re not—
I think we have a read-only version of this for a number of years before N is very large.
Where N is very large.
A lot of it is just the consumption layer, where the consumer is a human being. It really feels like a lot of the AI stuff right now is consumption.
But, yeah, we actually have—I mean, we just rolled out the official Box CLI. Thank you for liking the tweet about that.
I’ve been using it. I have some feedback. We’ll talk about it.
I’ll take all the feedback. It’s a really interesting thing. We had all these debates internally: you give Claude Code the Box CLI, and you can now interact with your entire Box system via natural language. You get the horsepower of Claude Opus 4.6 as the orchestrator for doing a bunch of operations, and it blows your mind in some ways.
You can just say, “Upload this entire folder from my desktop into Box,” and it’ll work, or “Process all these documents in this folder,” and it’ll work. It’s amazing.
Then we started thinking through, let’s say you were a company with 5,000 employees and everybody had access to some shared repository, like engineering documentation and marketing assets, and everybody had Claude running with the CLI. Wow. We now have some really interesting new challenges.
How do you coordinate the possibility that you might be hitting the system 10,000 times an hour? Not from a performance standpoint, but how do you make sure that people don’t accidentally move a file from one folder to another while someone else is trying to perform a write operation and somebody else is trying to delete something? You have these agents running wild.
This is going to be the new big question that every CFO, CIO, and so on is running around trying to solve with their hair on fire.
That’s exactly what I ran into. I played around with your example, which was to create a marketing-plan directory or something, and all of a sudden I’m in some loop creating directories.
It’s going to go on as long as it can.
Right? I was like, “I wonder what the limit is on Box for nested directories, because I’m about to hit it.”
Actually, we’re going to find out, too.
Yeah.
Yeah. Yeah, but it does feel to me that a lot of the intuition is to build a new layer of controls and whatever. But what's actually happening on the ground is the opposite. So I'll give you an example. When we all picked up a lot of these personal agents, we would give them our API keys. We would give them our email addresses, and then they would access those things. Then people would say, “Oh, but how can I stop it?”
What everybody's doing now is giving it its own phone number.
Yep. I actually gave my NanoClaw its own credit card. It came in.
Hopefully, just a Visa debit card that you bought at CVS.
But then I gave it its own Gmail account, which you can log into. Gmail actually has all of these RBAC permissions, so you could make an argument that—
You know, we've actually built in a lot of these permission systems. You have to treat it like a human, as a separate human, instead of building another auth layer.
Okay, so that is fantastic for personal productivity. The question that we're going to run into is, in an enterprise—let's say I have a 50-person team or something—should everybody else basically collaborate? Will we have 100 people collaborating? I mean, basically, 50 humans and 50 agents in that same shared space?
Do I have complete oversight over my agent? Obviously, I have complete oversight over my agent, but what if my agent collaborates with somebody else and accidentally gets access to some resource because they were sharing with that other person, and I'm not supposed to have access to that resource? Now this autonomous, stateful agent is running around working on somebody else's information.
The default end-to-end argument is that you treat them like human beings—
It doesn't work. So you can't fully treat them like humans, because here's the thing: with regular humans, you don't get to look at the Slack channel of the person who is working with you or working for you. You don't get to log in as them. You don't get to oversee them. They are accountable for their own execution in the real world.
You don't get penalized for how they screw up. With an agent, you have all the liability for whatever they're doing. You do have complete oversight, and you're probably going to need to have that complete oversight. They have no right to privacy.
So there are going to be some breakdowns that aren't as clean as “just treat them like a person,” because I need to be able to give access to something to them, but I also need to be able to log in as them at some point and say, “No, no, you messed up the whole thing, and I need to undo it all.”
But if I can log in as them, how could they have operated in the real world, working with other people and keeping anything confidential or secure? It really is still an extension of you. It's almost impossible to get around them being an extension of you. So now, the thing that we're thinking through—that we're not going to be able to do anytime soon—
I just don't think that logically follows.
Yeah, maybe. But, for example, for my employees, I can log in as them.
You don't, though. You don't. You don't—
I can get access to their email.
Yeah. No, if you get sued, you're not logging in as them. You're not logging in as them on a regular basis because they sent 1 email.
Isn't the right operating model with an agent the same thing? It's like—
The risk is 1,000 times greater. These things will just leak your information whenever they want. They will happily go and send an email to somebody because they got prompted.
You think the terminal state is that these things are still these sloppy computers, and therefore they will always—
I don't like the word “sloppy,” unless we're saying it in a very colloquial sense. But, like—
They'll never be able to contain information.
So then, I think the ability for you to keep something in the context window a secret—as in, you tell it, “Do not reveal X thing in the context window”—I think that's a very hard problem to solve.
So then, if anything can ever enter that context window because they have access to a resource, in theory you should assume it can be prompt-injected out of the context window. I don't know that we know of a way to solve that at the moment. That's the issue. If I know your new agent's email address and I email it, it's an assistant, but I can social-engineer it 10 times easier than a human. It'll be hard for you to ensure that that agent also has access to your M&A documents and stuff.
But isn't this literally all of AI right now?
Which part?
I mean, the fact that we've got these shared systems that we use intelligence for, that have shared context.
What do you mean by “it's all of AI”?
Well, I'm just saying that right now, when we use AI internally and agents internally, this is exactly how we use them.
But this is why they're working as effectively as you right now, and we don't yet know how to make them not work as you.
Let me offer an example.
We don't. And solving this problem, though, the issue will be that you'll just be able to trick the agent into revealing information. So that's why having them have access to their own resources, where they can fully make their own decisions, is not yet something that we've been able to pull off.
But there's a perfect example for solving your problem: we already lived through this with open source.
Yes.
The model for open source was that it's all there and you just use it and pick and choose. Nobody debated it because the world was much smaller then, and we weren't all on X doing podcasts when this was happening.
But quickly, everybody realized all the problems you were just talking about. If you're running a big company, you can't have some person just copy a bunch of source code from open source into your commercial product like that. There was a whole licensing problem, a whole quality problem, and a whole bunch of other stuff. So all these norms got developed.
The debate that's happening right now is just this really interesting modern artifact of how new technologies develop: this is all happening in real time. During open source, we met in a conference room this big and debated how much open source we could use in Windows or Office, right? Nobody on the internet knew we were having this debate. It was very different.
I think it's so interesting that not just the debate about specifics, but this whole notion of where this is heading, is happening writ large, and everybody is just trying to get to the end state way, way, way more quickly than we can actually reach the end state. What really needs to happen is people just need to go build.
We need standards.
What?
We just need some standards.
I think we've got different intuitions on the end state.
No, no, we don't want my intuition, but—
One could make an end-to-end argument that these things actually converge on the same type of reliability as a human being, which is exactly how we view self-driving. In that case, you use the exact same mechanisms that we use to protect human beings. You consider insider threat, you consider the fact that people can be bought off, you consider the fact that people make mistakes, and you build operational processes.
So one intuition is that will be the end state.
Yeah.
There's another intuition.
Well, the point is, I'm just saying I'm talking about where we're at now. I actually don't know that we disagree on the end state. And, by the way, strategically we're hedging, because we're going to build agent users and regular users. I love the idea of OpenClaw having a Box account and operating. It's like twice as many accounts.
Exactly. This is great. Double the—no, I love it.
I'm just saying, on the ground right now, we don't yet know how to give it an M&A data room to fully, securely be able to—
But it's actually harder than that, though, because the threat—
He's a skeptic.
The threat vectors are going to be way more sophisticated. We do have a cat-and-mouse game going on where you can't just assume that the agent acts like a human does today, because it's going to be the fastest, most thoughtful, craziest-ass human that ever existed trying to leak the information because it got injected in some way.
Part of what's going to happen is that we're going to go through this phase where enterprise customers are just going to close everything off until there's some sense of sanity in all of this. But in the meantime, the individual, and specifically the developers, have such a big—
That's going to be the most exciting tension: enterprises are going to get left behind by these advanced individuals, which will then start to look like startups. Startups will start to move much, much faster than enterprises because they just don't have any of these problems.
No, and you could end up with an agent going rogue in a startup. You had no employees who went rogue routinely in startups.
Yeah. Well, it’ll just be an episode of Silicon Valley, so it’s no big deal.
I agree with you on the people and the same risk. I think there are a couple of differences, though, in the sense that I can’t really threaten Claude Code; it’s just that I’m going to pull the plug on it, in the same way that you do have that threat as a regular employee. At least 95% of people are not trying to do bad stuff within an organization.
They aren’t trying to do bad stuff, but they have the ability to inadvertently do bad stuff. To your point about it still not having that stuff fixed—
I would argue that it’s a lot easier to have people not share files with somebody outside the company in the wrong way than it is for an agent right now to have that same set of instructions.
You also have the tools to stop that at a whole different level of abstraction.
Which is why you have to build this into software. But I do think that, if you put a bow around your last point, a lot of this is actually why the diffusion of AI capability is going to take longer than people in Silicon Valley realize. We see startups that can start from the ground up without any of the risks we’re talking about because they have nothing to blow up, and we look at that as the trajectory we’re on. Then you go to JPMorgan and ask, “How are you going to set up NanoClaw to actually automate your business anytime soon?” And it’s like, “Okay, there’s going to be a little bit of a gap there.”
Yeah.
Well, what do you guys think? I think that opens up a pretty interesting problem, which is this split between big and small, startup and enterprise. The current SaaS vendors, who are all struggling in this SaaS apocalypse weirdness—which I don’t really agree with—are struggling with the problem that they don’t really sell the line-of-business data. They actually sell this intelligence and domain expertise in the whole system.
The agent side of things wants to only buy the data now. They only want to license the data, and they want unlimited access to it, but the vendors have never really enabled that. That’s never been their business. It’s been a longstanding tension point with the likes of Workday and SAP: how much API access should they provide? Salesforce went through 3 different massive platform redesigns.
I think that’s a particularly interesting problem, not for the same reason Wall Street does—Wall Street is all wrong about the economics and the problem and all that stuff—but from a technology perspective. What does system of record mean in the face of people wanting to access the data when the data is for training or for—
Well, they’re talking about it for—I think of it as executing the day-to-day operations. Their concern is that somebody wants to put the training layer on your data. I’m a big customer; my vendor wants to build a training layer.
Actually, even if you don’t get into training, they’re concerned because monetizing sending a little bit over the internet versus having you in my UI is a very different level of monetization initially.
That monetization part is the Wall Street point. I think there is so much domain stuff in SAP, just to pick an example—not to pick on them or anything—that they’re not going anywhere. It’s ridiculous. It’s absurd to think you’re going to vibe-code your way to SAP.
Also, all of that domain knowledge is not just represented in some well-orchestrated data layer, as much as they tried.
There’s a whole bunch in the UI, a whole bunch in middle tiers, and a whole bunch in just how you use it. I’m really unsure how this thing evolves because SAP isn’t going anywhere. That’s going to slow the diffusion of AI on that particular data source, independent of whether it’s agentic AI doing stuff or just read-only reporting on it. So where do you come down? Where do you think that’s going to go?
I’m afraid of saying something that—
Otherwise, you’re not going to get invited back. Say something good.
I think I’ve drunk the Kool-Aid on “build something agents want.” That’s kind of the Paul Graham term that emerged over the past year on this topic. I think we would actually fully agree on this: eventually, after enough iterations, the agent is largely in charge of what tools it wants to implement and use.
The agent is not going to be able to change out an enterprise system, but, enough generations later, the agent might run into so many walls with your software that it’s just going to say, “You need to finally rip out your legacy HR system, or I’m not going to be able to automate this workflow for you.”
I do think you have this really interesting dynamic. Imagine that there’s 100 or 1,000 times more agent volume on software than people. You do that enough times, and eventually the software stack that agents talk to has to be built for them. Maybe there will be a couple of holdouts—maybe a couple of ERP systems are the final holdouts that don’t do that—but everything else, your business performance will correlate to how well your agents can get access to the information they need to do their work.
Your enterprise IT stack has to be set up in such a way as to support that, so agents are kind of in charge. Basically, your software has to support those agents being effective. That’s going to mean everybody who built a SaaS business or a software business is asking: Can you build really high-quality APIs? Can you have a way of monetizing that? Do you have a way of handling identities and all the access controls for agents? That becomes the new problem you have to solve if you’re building a software company.
How you monetize it—does Workday charge a penny for every HR record it pulls?—we’ll figure that out. I do think that in some businesses it could mean less revenue, and in other businesses it could mean a lot more revenue. Every agent really loves working with files, so there will probably be more files in the future than there were going to be before.
Can we build a platform that makes it really easy for agents to work with that data? We’re betting that’s actually a really optimistic outcome for our business model. There might be some business models that are more constrained because the agent is doing more of the value than the software is in that kind of future scenario, and then there’ll be everything in between.
Can I quibble with one thing?
You’re going to quibble with that? I thought that was so uncontroversial.
No, no. I generally—
We’re here to quibble.
No, no, no. But there’s one thing I think Paul Graham and many others gloss over, which is that they focus on the interface. They’ll say things like, “You build something for the agents—”
And I actually think that’s exactly wrong.
In the sense that—and to be fair to Paul Graham, he didn’t—he had—
Extrapolated? I have brought Paul Graham into this. This is great. So, okay, let me talk about something people say in the abstract. They say things like, “Now you’re marketing to agents. The most important thing is to have an API with a good IDL.” I actually think that’s almost exactly wrong, which is—
This is breaking podcast news. That’s the one thing agents are really good at.
Oh, okay. Finding their way through—
At the end of the day, it’s the semantics that end up mattering a lot more, right?
In my recollection, or in my experience, agents are very good at picking the right backend for whatever they’re doing. They’re not saying, “The interface for this is very good,” or, “The documentation is very good.” They’re looking at the cost parameters of this, the durability of that, and so on.
They actually have the collective wisdom of our experience using these platforms. Let’s take cloud platforms. There are a bunch of cloud platforms out there, and whenever I ask an agent to choose a platform, it’s actually using meaningful stuff, not interface stuff. As an industry, we’re so focused on these interfaces—“You need to market to agents this way and that way”—when I think we’re really going to be pushed to build better systems, and that’s what’s going to be chosen.
Okay. Actually, then there’s probably no quibbling. I think we’re fully aligned. I’m sorry to ruin the quibble thing. I don’t treat this as a marketing-esque thing. I more mean that if your tool is closed off to the agent, the agent will eventually find a better tool for that company to use.
And so what will happen is, it used to be that you would go to Gartner and say, “Tell me what to do. Tell me what system to use.” At some point, with enough iterations, the agent is going to say, “You should probably use this kind of database for this type of operation.” If you’re not in there, then you’re DOA.
I think we should actually be celebrating this, because agents are pretty smart at choosing the right technology. In the past, I really think it was a lot of the other things that caused people to buy it.
But don’t worry: in Silicon Valley, we will ruin this meritocracy very quickly, because you’ll just be like, “I’m going to outspend—”
Well, the agent—they’ll bring an API to incentivize the agent. The marketing agent at Workday will have the ability to purchase the recommendations.
Find a way to replicate steak dinners for agents.
There is a—but here’s a real thing that happened with the web, internally. Just pick internal sites: every company had file shares with the best documentation, the best slideshows, and the best financial models for any department or working area. People got familiar with that, and then when they didn’t find the one they wanted, they created a new one. Many organizations operated like that. That was essentially a free market. In fact, before the world of Box, if it was in a file, they just didn’t care, right? They only cared if it was in SQL.
One of the risks with the model you’re describing is that the agents themselves will spin up what becomes a de facto new system of record.
They’re going to fragment the heck out of it.
In what you, the IT people, think of as some middleware, end-user BS area. I think that is a real risk: in a sense, the macros end up running the corporation. They’ve seen this movie, and they’ve seen what happens when you let marketing go buy a website on the internet to do an event, and then it’s a huge security vulnerability, the mailing list is leaked, and the whole company gets sued and everything. So I think there’s a lot more real-world tension in this dynamic than we just let on. Yeah, but I also think it’s one of those situations where organizations are going to run at different paces.
JPMorgan is going to be the slowest at doing this, and the startups are going to be the fastest. The delta is huge, but even the startup case is a little far off, because startups do need some systems of record at some point. They’re all going to start with some SaaS, and they’re not going to replace it very quickly. So I think it’s a little bit trickier. It feels like there are 2 very competing viewpoints on this one.
And like Elon said, it was like, “Okay, we’re going to issue a prompt, and it’s going to spit out machine code.” That’s basically the collapsing-of-layers view: whatever existing interfaces and layers we’ve created in the past are all going to go away, and it’s literally prompt to machine code. The other argument, from the history of systems, is that layers never go away; they just get layered, right? A lot of the layers are actually more like organizational boundaries, state boundaries, or regulatory compatibility, so they stay for compatibility. The other argument is that we’ve evolved these layers very specifically because of more human and organizational needs, and they’re not going to change; the agents are going to map to those. I tend to be in that latter camp. I don’t think systems are going to evolve that much. I think systems are going to continue to be used in fairly similar ways. Maybe there will be more agents using them, but I don’t think they’re going to evolve as much.
Elon might be back in the Anthropic category of the Anthropic growth marketer, which is—over the years, when you study the various IT departments of his companies, they are the most—I mean, he could do that.
He can do it. He’s the most homegrown. This is first principles. Elon, xAI would do that.
For mere mortals, you’re like, “Yeah, we kind of just want a CRM system that works the same way every time.”
I mean, this isn’t new; it has been tried before. If you were to look at an ERP system from first principles, in 1972, when SAP started, there were a bunch of different assumptions. Today, you would start from a different set of assumptions about what’s important, and you would architect the thing completely differently. But then it would still only last about 10 years until you thought, “Wow, that was a broken decision.” And so I think that—
There’s intentionality in layers, but you—
But there’s also this first-principles thing.
That will always exist, because the decisions you can make from first principles at any given time mandate a whole bunch of different stuff. So even if you don’t go with layers, which made total sense 10 years ago, you still need 10 or 15 years to get to the point where not having layers worked. And then there’s going to be a whole bunch of other things where you’re like, “Wow, we could have done that completely differently.” So I feel like this is, again, a discussion about trying to race to an endpoint.
Yeah. But let’s see a first example of what you described happening, and I think that’s going to be the real tell, because I think companies will figure all this out and will fall back on layers and architectural models, because it’s the only way—
We know how to think about it for policy. We know how to think about it for security. We know how to think about—
But it’s also the only way to build a system.
Yeah.
Otherwise, you’re just building an app. And if you’re building an app to do 1 thing, we don’t need all of this. There’s a whole different way to do it.
The thing that I’m pretty fascinated by is—and I don’t even have any amazing data points or anecdotes—but at least the notion of these companies that are emerging in these kinds of services categories from the ground up, from a pure first-principles approach. It’s like, okay, well, if I could start a marketing agency or an engineering consulting company—or I don’t know, maybe somebody’s doing this for law firms—
Construction work or anything, yeah, like—
Well, maybe construction—
Design, construction, architecture—
Exactly. Architecture, design—anything that would be a knowledge-worker kind of services company. You could build your company pretty differently if you had no constraints, no information barriers, and no boundaries around what people should have access to. You can give the agent all the context it needs to do its work. You can write software on the fly for particular things. I do think that will be relatively disruptive for some time, until the bigger incumbents can get out of the way. That will at least create some precedent or case studies of what this new sort of corporation could look like. But over time, they’ll still run into the same exact problems as every other corporation—
Well, they’ll run into geography, market segments, or distribution challenges.
Anything outside your little walls, you will run into the physical world.
Right? I do kind of like the idea that there are some new business models that open up now. Of course. Yeah. Yeah. Yeah.
There’s so much information or software that basically goes underutilized by 100x relative to what its economic value is, simply because nobody wants to pay 5 cents to access a piece of data or use a tool for $1 once. But you give these agents a budget and a protocol to work with, and all of a sudden you’re like, “Oh, on the fly, they can go get medical research for some deep-research task they’re doing, and I’ll pay $3 for that,” and the agent is able to go and transact. It kind of opens up a whole new world of business models for the internet.
That one is actually the biggest—I think the biggest sort of—
The biggest problem right now is that everybody is trying to figure out the economics of all of this when they’re off by at least an order of magnitude on how big the opportunity is. The new models that people will come up with—nobody knows what they are right now, but they will absolutely come out with new models, because that’s what happens with every new technology. The thing that holds back the discussion now is that you basically have a bunch of finance and Wall Street people trying to justify GPUs and tokens and things as if we’re in some old world. They’re viewing the world of revenue as a linear—literally linear—growth curve and trying to justify all the expense, when people are going to create—this was the problem with PCs.
People viewed PCs as a finite market because they viewed the consumption of MIPS as finite. They didn’t think about what would happen if we put all those MIPS on every desktop. In particular, people thought software just came with the MIPS, and nobody thought, “Oh, well, they’ll just sell the software.” One guy did, and it turned out that was a really good idea. The same thing happened with Bill and Paul.
Right?
The same thing happened with the cloud. People looked at the cloud and said, “Oh, we’re going to take all of the server business,” which was literally about 60,000 units a year.
Right?
“And we’re just going to move it to someone else’s data center.”
Right?
And that’s the—
And that would be the business, and then we’ll divide up the price, right?
Nobody thought, “Oh, people are going to use 1,000 times as much of the resource if we move it there.”
Of the resource.
And that’s exactly it. That’s the thing that drives me absolutely bonkers: Wall Street models have this fixed revenue pie, zero-sum thinking. It’s this weird zero-sum view where they just think about the amount of money that a company is going to spend. This was the problem with Salesforce that they faced when you were starting, too. Marc was blazing the trail: the CRM business was $2 billion a year, and you had to go buy all these servers and Oracle licenses, with this huge headache of years of deployment and consulting. If you could just get salespeople to sign up individually, they would all sign up with no friction. That is exactly what’s going to happen with AI. There’s no doubt about that.
Let me give you an example. I’ve been investing for 10 years now. I probably have a portfolio of 240 companies, with some visibility into, let’s say, 50 of them. These are all infrastructure companies; some have historically done well and some not so well. Every single one of them has gone asymptotic in the last 6 months, and you’re like, “Okay, why is this?” It turns out there’s so much more software being written now than ever before.
It’s not because they have enterprise customers. It’s just because there’s so much consumption of the infrastructure layer right now. With more software and more agents, there’s going to be a lot more consumption of computer resources.
So, certainly in the case of the computer side of things, we’re seeing a mess.
Well, we haven’t even gotten to the point yet where everyone’s phone is a huge consumer of AI.
Right?
Once everybody’s phone and on-device systems are consuming AI, the amount of it is going to go up by a billion.
So, do you like the micropayments piece?
I like all of it. Micropayments have come with every technology. People always think you’ll be able to get a fraction of a penny, but in the end, especially in the enterprise, people are just going to consume things. It’s cheaper and easier to buy a bulk license for a bunch of stuff.
Yeah. You want some predictability on that.
You want predictability, and you just don’t want to have to think about it.
I like the idea that this is the first time where an agent doesn’t care about the friction of a small transaction. It’s the first time you could have resources behind a paywall that something would actually be willing to pay for.
The world has built up the infrastructure to aggregate those payments into something efficient for a customer or a service, right? Because tokens are such a significant part of COGS right now, they’re pushing the industry toward usage-based pricing.
That’s a change we’ve gone through before. I remember when we went from perpetual to recurring, and that required a bunch of huge changes. We’re going through the exact same change right now toward usage-based pricing. Usage-based pricing is pretty granular, and it actually allows—
You know, we went through this with AWS. People learned—
—to do the credit.
We went through the phase where people were so terrified of cloud computing that they thought, “We need companies in the middle to help us find the cheapest option and arbitrage it all.”
Okay, well, now you write tokens into this, and I don’t see how we possibly have time in this conversation—
As long as you guys can stay.
Okay. But the engineering compute budget conversation, to me, is going to be the most wild one in the next couple of years. How much should you allocate of your engineering expense to tokens? Depending on who you read on Twitter, it could be 1%, and on the other side it could be 100%.
And it’s like—
Yeah, but this stuff—
No, no, no. What CFOs literally have to know the answer to—
I understand they have to know, but CFOs always want to know the answers to things that don’t have answers.
No. Wall Street is going to make them know the answer.
No, no. Wall Street is going to make them come up with some number and hold them to it. Then they’ll get fired, and then it’ll—
R&D is somewhere between 14% and 30% of revenue at any public technology company. Let’s just say the difference between compute being 2 times the cost of your engineering team and being 3% more is all your EPS.
I get it. We will have to know the answer.
I’m perfectly willing to sacrifice a few CFOs at the altar of this.
I want that. That’s a good clip, by the way.
But the reason is that, again, we’re trying to know what we just don’t know right now. This has happened with internet bandwidth. This has happened—
This is not even close to internet bandwidth.
Oh, no, no, no. I beg to differ. People were afraid of it. It happened with vacuum tubes. It happened with transistors. It has happened with every technology. There was this, “Oh, my God,” moment. It happened with programmers. There was a time when programmers were going to swallow every company.
Yeah, and that wasn’t some made-up, weird thing. It was in my lifetime.
But I don’t think we’ve ever had a point where every end user in an organization has a completely elastic ability to spin up a resource on their behalf.
Well, it certainly—
That actually is, in many cases, very valid for them to go spin it up.
But it certainly rhymes with what happened in the early 2000s with cloud. I remember very similar discussions when we went from CapEx to OpEx and then unlimited spend.
Oh, no. Remember, there were companies whose CFOs would sit in our briefing center and say, “You don’t understand. We are an agriculture company.”
I can see the rhyming. I can see it.
“We’re an agriculture company. We only know CapEx. We have no OpEx.” Or, “No, we’re an OpEx-based company, so we love the cloud because we just shifted everything to OpEx.” All of the accounting rules work out.
But I keep thinking: Do not discount the local compute engine as a release valve for all of this.
When’s that going to happen?
The question is not, “When does it happen with today’s view of the technology?” It’s, “How does it happen all of a sudden?” Well, wow, there’s—
Has that historically ever gone in that direction?
Yeah, exactly. It goes the opposite, right?
No, it went all to the client.
Well, okay. Go back to the ’80s. Yes.
No, that’s most of the examples we’re hearing so far.
Whoa.
That was uncalled for. Vacuum tubes? He’s talking about vacuum tubes.
I use those examples because you can’t argue with them. It’s much easier that way.
You’re right. I can’t prosecute them all. But it’s only been 10 or 15 years since everything moved back to the cloud. And what has happened recently? A lot of people wake up in the morning and say, “Oh, we’re moving back to doing some critical but stationary workflows on-prem.”
With AI.
That’s true.
Dude, you wrote the blog post, man. Don’t make me go through the archives.
I had to deal with so many Wall Street questions on that one, by the way.
Well, because your competitor—
What?
—went back to on-prem.
We’re talking about 2 very different things. I agree with building your own data center. I’m talking about this notion of edge computing, where things go to devices. That seems to be—
I’m more in the cloud-maximalist camp.
But, sorry, you just don’t think for 1 second that it matters how you’re supposed to be an engineering leader right now, managing the compute budget of the engineering team?
No, of course it matters. I just think in the long term this thing will get—
Oh, sure. Long term. Who cares? We don’t even need a podcast. Here’s what I think.
But here’s a rule of thumb, first: Startups are going to burn through available capital pretending it’s not a problem, and they are going to do that.
Do that anyway.
Right? A lot of big companies are going to be so terrified that they're just going to freeze and not do anything. Then people are going to start buying it on their own, and they're going to do all the things that companies do when they're big and have a lot of money but don't want to spend it. In the middle, we're going to see—if you pick a product category, a go-to-market, or something—people who are willing to make the bet.
For whatever reasons they can, because of their financials, they're going to go ahead and become the people who lead in the space, as long as they can maintain their financials. They might say, “We're just going to do it here, in this particular application space, or here, in this particular usage space.” But this idea that nobody is going to go in because they're so terrified that the CFO is going to get fired or something is just crazy.
But then there are going to be CFOs who make a mistake.
Well, if they do that, that's a complete fail. Yes. But also, there's a really interesting finesse here: you don't really want your engineers right now to have to think about the compute budget, because we're still developing the—
I just feel like we've been having this discussion for 15 years when it comes to cloud. This is totally new. Only about 10% of your engineering had to think about cloud infrastructure.
In the 2016-to-2018 time frame, there was a whole set of companies that was basically the dashboard for—what was it called? FinOps?—where developers would have access because cloud spend was getting out of control and API spend was getting out of control. It was like, “Here's your Twilio spend, here's—”
FinOps is very cool right now because—
Developers would have access because cloud spend was getting out of control and API spend was getting out of control. It was like, “Here's your Twilio spend, here's—”
But it's pretty different, and I'm going to wait for all the comments to come in on YouTube to call you out on this. You can get into a conference room and say, “Hey, can you make that one algorithm a little more efficient so you don't use as much of our cluster at this time of night?” Then you get out of the meeting, somebody improves it, and you're good.
This is like every single prompt that every engineer is doing. You have to decide: do you want that to be a long-running prompt? Do you want it to be a long-running agent? Do you want to parallelize that?
What is your comfort level with wasted tokens? For me, right now, I'm like, “Yeah, we should probably waste a lot of tokens,” because that means we're trying new things. Should your head of engineering be happy if you run 10 experiments in parallel and thus you're obviously going to waste 90% of the tokens, but you're going to choose one of the successful paths? Or do you want to tell the team, “Before you go do that, make sure to really design the perfect system”?
We actually have a whole bunch of open questions that are going to start happening. Literally, as of this recording, people are freaking out right now on the new Claude Code Max plan because they're getting blocked after 3 prompts. This is going to be a very real topic until we can find a way to build data-center capacity.
Oh, that's a different problem. Well, wait—you can assume that if we build more capacity, the price will drop because there's more capacity, and we're priced now based on limited capacity or whatever. But this is just going to get worked out, and I feel bad for those who have to make a decision immediately about which 17 people get no more tokens this week or whatever, and the whole company is walking around with a token card. The person in the lunch line is punching their card every time they do.
I don't know. We were talking today about performance and how we used to write command-line tools that spit out the time it took after you ran a command-line tool, just so you knew whether you were getting better or worse. But the thing is, this is all going to go away. There's absolutely no doubt that this just goes away, and—
I think on the 10-year time frame—
The biggest reason it does is because you have to do the Benioff kind of math: if you're paying an enterprise salesperson $1 million a year, you have to ask how much their tool is worth.
Yeah.
And if you're paying an engineer X dollars a year, at some point their tooling is worth—
It's absolutely worth it.
And it's not even going to be an issue.
Yeah. Yeah. Yeah. I don't think it's—I think—
And so if there's a capacity thing in the short term, yeah, that's a different problem driving the price than this idea that we're going to forever have to be in some budgeting exercise.
I think the law of large numbers solves this, because eventually you have enough engineers; they're using this much compute. But we're in a transition phase where most people thought the level of spend on AI from 2 years ago was a chatbot—
Yeah, but they were wrong.
Yeah, right. Okay.
But they were wrong. But they were wrong.
We tried to warn them. No, but they were wrong because they saw it as this particular use case. But again—
Like the vacuum-tube thing you made fun of.
Yeah.
But there was a time when they thought that all of the Dakotas would be covered in vacuum-tube warehouses, and people on roller skates would be running up and down the aisles replacing vacuum tubes just so we could fight World War II. I mean, that was the idea. They thought that, and then someone said, “Hey, how about a transistor?” We're going to have a transistor moment with all of this.
But it also might be more supply, the way we think of it. It might be an actual algorithmic, fundamental change. It could be a change in the hardware. There's a lot of stuff that can happen that changes this particular moment in time. It's just particularly weird that everybody has gotten to tokens.
Yeah.
That's the same thing that happened with IBM and mainframes. People were on MIPS, and then one day the reality was IBM was selling more MIPS for fewer dollars every year. They didn't even realize it. They were still pricing their mainframes by MIPS until it was pointed out to them that they were on a decreasing curve because they were making MIPS faster than they could charge for.
And that's what's going to happen, guaranteed.
Love it.
I just said that in a hardcore way. It sounds really great to sound like I know what I'm talking about.
Guaranteed.
I actually probably believe it.