为 AI 智能体构建云基础设施 | AWS CEO Matt Garman
AWS押注AI需求足够持久,足以支撑其2026年2200亿美元的资本开支,且预计不会放缓。 Matt Garman认为,AWS客户基础足够分散,单一客户占比最高也只是个位数百分比;已有生产环境工作负载持续产生正回报。“不存在一种泡沫情形,会让他们停止在这上面的投入。”
AWS正在整个生态中分配稀缺的加速器,而不是把芯片全部卖给前沿实验室。 尽管AWS可以把所有可用芯片卖给几家大型实验室,但它会为初创公司保留产能,并最终以某种形式满足收到请求的约60%;有时是在其他区域或采用不同配置。AWS计划在未来几年采购200万块 NVIDIA GPU,但Garman也承认:“谁知道这够不够?”
智能体工作负载正推动云架构转向高速、可能随时弃用且权限严格限定的资源,并与长期运行的系统并存。 智能体关注p99.9延迟、3秒创建数据库、计算沙箱、网关和限时权限,而不是直接继承某个人的访问权限。Garman的核心判断是明确的:“智能体工作流在AWS上的表现往往好于其他任何地方。”
AWS正在消除传统的上手摩擦,但不会把初创公司锁死在简化版终点上。 新流程铺开后,用户可以用Gmail开户、无需信用卡,并在30秒内开始使用,VPC和IAM默认配置则在后台处理。关键在于,这仍然是“一个真正的AWS账户”,客户之后可以逐步启用完整的安全和组织控制,无需迁移。
AWS的定制芯片路径从Nitro、Graviton延伸至Trainium。 Graviton的成本低20%、性能高20%,AWS前100大客户中约90%多在使用;Trainium 3的产能可能已经售罄至明年年底,而Trainium 4已经宣布但尚未发布。Garman称,按绝对性能和成本表现计算,Trainium可能是“目前市场上最好的推理芯片”。
企业采用智能体的瓶颈,与其说是模型访问,不如说是工作流重构、评估体系和信任。 Garman反对简单复制人类的五步流程:智能体可以并行尝试50种方法,但企业在赋予自主权前,需要权限体系、护栏、标注数据、生产环境度量和漂移测试。AWS的前置部署模式旨在用45天教会客户这些能力,然后退出,而不是制造持续数年的咨询依赖。
AWS围绕企业数据托管来定位Bedrock,同时保留专有模型和开放权重模型两条路径。 Garman表示,提示词和数据都留在客户的VPC内,永远不会回到模型提供商手中;拥有真正专有数据的客户可以对开放权重模型进行后训练或微调,甚至蒸馏,以更低成本获得更好表现。前提是必须建立评估体系,证明这种优势,而不是想当然。
在Amazon内部,智能体已经压缩软件、产品和业务流程,并开始重塑团队设计。 Garman称,采用智能体优先模式的“frontier teams”由员工管理能够写出全部代码的智能体;HR和财务员工也在为规划、税务和合规工作构建自己的智能体。过去需要10人完成的能力,如今可能只需3或4人,由此产生了新的问题:在小团队快速转向其他项目时,谁来持续维护已经上线的产品。
1. AWS仍认为自己处于早期阶段
Raghu Raghuram先给出规模参照:AWS从赚到第一美元收入,发展到如今约1690亿-1700亿美元的规模,目前增速为37%。但Garman仍称,这只是“业务可能达到的规模的早期阶段”,因为大量工作负载仍在本地运行,而AI正在扩大每天完成的总算力。
Garman在AWS的第一份工作,是2005年商学院实习期间分析谁最看重这项拟议中的服务。答案是初创公司;它们后来成为“我们核心业务的命脉”,因为AWS的价值主张对初创公司尤其有吸引力,也帮助它们搭建能够逐步扩展的架构。
这种关系已经形成商业复利:AWS估计,目前30%-40%的收入来自AWS存续期间曾经处于初创阶段的公司。那家只有2人的公司,其价值不仅在于未来可能成为大型企业客户,也在于它可能提前发出能力需求信号——5年后,银行、医疗企业和政府可能就会需要同样的能力。
初创公司的基准线已经彻底改变。过去,一家公司可能融资1000万美元来反复迭代一款应用;如今Garman看到的团队,可能从第一天起就有2亿美元融资和10亿美元估值,相应的目标、模型训练成本和基础设施扩张规模也都更大。
2. 智能体暴露出新的云性能边界
一些初创公司的需求并没有改变:创始人仍然需要可扩展架构、安全性、性能以及能够支撑公司从3名员工继续增长的IAM。Garman称,正是这层完整能力,让许多公司最终更偏好AWS而非新型云厂商,即便它们眼下的直接需求只是GPU。
改变的是使用者。AWS如今不仅为人操作的云设计,也为智能体操作的云设计,重点包括大规模API访问、快速创建资源和可预测的性能。Garman举了一个具体例子:“如何在3秒内启动一个数据库?”
AgentCore和Bedrock是明确面向智能体构建的服务;处于预览或测试阶段的“AWS Context”,则旨在为S3、Aurora及其他AWS存储中的数据建立统一的上下文层。智能体可以穿越这些彼此分离的数据资产,而人通常做不到。
智能体还会放大尾部延迟。人可能根本察觉不到S3的p99.9性能,但智能体工作流可能因此被阻塞;延迟、吞吐量和底层引擎的响应速度,因而从抽象的基础设施指标变成了编排约束。
3. AWS简化上手路径,但没有移除控制能力
Raghuram提出的反问是:如今编程智能体会自行选择数据库、邮件系统和部署目标,可能让运行了10年的服务显得“过时”。Garman回应称,核心构件仍然稳健;真正的缺口,是让没有现成云账户的人,或某个智能体,能够更容易地开始使用云。
在典型的成熟工作流中,客户告诉Kiro、Claude或Codex在AWS上构建,提供凭证和部署指令,之后由智能体处理其余工作。但如果一个尚未上云的实验只说一句“部署”,它可能会选择上层体验更简单的合作方;Garman称,AWS欢迎这种结果,但也希望直接解决这个问题。
AWS正在逐步推出新的开户流程:用户可以用Gmail注册、无需信用卡,也不必手动定义VPC或IAM角色。默认配置由后台处理,账户可在30秒内投入使用。
Garman最强调的设计选择是连续性:这不是一个之后必须迁移的玩具账户。当客户需要组织、定制VPC或细粒度IAM时,“你已经处在一个真正的AWS账户里”,可以逐步开放这些控制能力。
4. 可弃用的智能体基础设施需要生产级退路
智能体经常创建一个数据库、完成一项小任务,然后将其丢弃,由此产生一个真实的设计问题:这个数据库是否需要5个9级别的持久性?AWS传统的Aurora定位,是把数据库视为需要持久性和可用性的生产资产;对一次性智能体任务而言,这可能属于“过度设计”。
Garman不愿意用一个随意降低持久性的选项来解决,因为AWS并不总能知道临时资源最终会不会变成永久资源。因此,工程目标是打造双用途基础设施:创建速度足够快、资源足够轻,可以随时丢弃;同时又能够在不更换系统的情况下扩展成持久化生产数据库。
一些智能体需求是真正新增的基础构件,而不是旧系统的新用法,包括计算沙箱、网关,以及区别于人类或服务角色权限的权限体系。“你不能只是把Raghu的权限给它”;智能体可能只需要完成一项任务所需的狭窄、限时权限,甚至不应访问某个工具的全部能力。
Firecracker微型虚拟机提供了现成底座。它大约在10年前开发,如今已被许多沙箱初创公司采用;其启动速度快、传统虚拟机开销低,同时保有强安全边界,非常适合执行智能体任务。
5. 稀缺性使GPU分配成为战略组合决策
Raghuram直接点出矛盾:前沿实验室可以“吞掉所有可用GPU”,而规模较小的公司没有相当的信用和融资能力,却可能成为明日的企业客户。Garman承认,AWS完全可以把所有加速器分配给少数实验室,但公司有意为初创企业、企业客户和更广泛的生态保留供给。
扩建规模极其庞大:Garman称,AWS计划在2026年投入2200亿美元资本开支,并预计不会放缓,因为需求依然“巨大”。电力、数据中心、资本、内存、芯片,甚至建设这些设施所需的施工劳动力,都可能轮流成为约束。
AWS称,最终收到的请求中,约60%会以某种形式得到肯定答复,但交付可能要延后,也可能安排在其他区域,或采用不同配置。每家初创公司仍然想要更多;AWS宣布将在未来几年采购200万块 NVIDIA GPU,但这个数字本身也可能不够。
Garman将AWS与另一类供应商区分开来:后者最大的1到2个客户可能占到30%-60%的产能。AWS的单一客户集中度最高也只是个位数百分比,而大部分使用量来自与应用绑定的核心算力、存储和推理;他询问的几乎每一家企业都表示,当前AI能力已经带来正回报。
6. 限制因素持续向供应链下游移动
当被问及2027-2028年的瓶颈时,Garman引用了《目标》中的一句话:“从来不会只有一个约束,永远只是最新的约束。”电力缓解后,限制因素可能转向内存、TSMC产能、HBM、网络设备、连接器、硬盘或SSD。
地理位置会进一步放大问题,因为产能并非完全可互换:印度尼西亚有充足电力,并不能解决德国的短缺。AWS跟踪的零部件达到数万乃至数十万种,并沿供应链向上追溯4到5层,寻找任何可能阻断部署的部件。
规划周期已经从向公用事业公司申请再增加几兆瓦,扩展到为太阳能、核能和其他电力项目提供融资;有时这些项目位于表后,有时则向电网供电。电力和输电决策如今延伸至20年,服务器、内存和芯片需求的规划则横跨2026年、2027年和2028年。
数据中心引发的反对也带来沟通挑战。Garman称,AWS需要更清楚地说明可再生能源、用水和就业方面的收益。他举例称,某个县的居民据报道每年少缴5000美元税款,原因就是AWS带来的税收;但这一看不见的收益此前没有被告知居民。
7. Nitro带领AWS从虚拟化卸载走向定制AI芯片
AWS的芯片路径始于“虚拟化税”。公司最初将网络虚拟化迁移到一块卸载卡上,随后与一个小团队合作,由其搭载Arm核心的卡片吸收存储虚拟化和其他功能;收购该团队后,AWS最终借助Nitro通过API提供接近裸机的资源。
收益不仅体现在利用率和性能上,也体现在隔离能力上:Garman称,AWS可以有底气告诉客户,自己无法访问客户正在运行的虚拟机。由于这套架构并不是其他人可以直接购买的通用组件,他认为AWS因此获得了长达10年的领先优势。
将这些Arm核心做成服务器后,AWS最初得到的是性能不足的Graviton;随着Arm性能提升,Graviton随后成为“势不可挡的爆款”。Garman称,过去5到6年里,Graviton一直保持成本低约20%、性能高约20%;AWS前100大客户中约90%多在使用,一些客户迁移整个服务器集群后,服务器数量减少了一半。
AWS在5到6年前启动Trainium项目,如今已经开始交付Trainium 3,产能可能已经排到明年年底。Bedrock的大部分流量运行在Trainium上,此外还有Anthropic和OpenAI的合作,以及大约6到12家在其上构建产品的较小型初创公司。
8. 企业自主性取决于重构、评估和数据信任
如今大多数企业智能体仍然相对简单、缺乏自主性,但客户已经报告了实际价值。Garman的第一条建议,是停止复刻“Bob完成第1、2、3、4、5步”;智能体可以并行化处理、尝试50种方法,以不同于人类工作流的方式解决目标问题。
第二个障碍,是企业对自主性的担忧有充分理由。企业需要护栏、沙箱、数据权限,以及明确哪些环节必须保留人在回路中;当智能体可能删除生产数据库时,“放手去做”不能算一条可接受的指令。
Raghuram列出了评估体系的要求:标注数据、持续测试循环、目标达成标准、生产环境度量、回测和漂移检测。Garman称,企业目前并不知道如何解决这些问题。他怀疑真正擅长解决这些问题的人并不存在,因此AWS才推动前置部署工程服务,目标是在45天内教会客户这项能力,然后让客户能够自行运转。
在模型选择上,Garman认同企业数据是“它们最有价值的资产”。Bedrock保证数据留在客户的VPC内,模型提供商永远看不到提示词;这是AWS优先建设的底层原则,即便3年前批评者还认为AWS推进得太慢。
9. 开放模型与机器速度安全扩大AWS的服务边界
拥有重要专有数据的客户,可以对开放权重模型进行后训练或微调,再进行蒸馏,最终有可能以更低成本获得更好的性能。Garman始终保留条件:客户需要拥有合适的数据和专业能力,并通过评估证明定制模型确实更优。
Raghuram称,这让SageMaker“重新获得了生命力”;Garman则表示,SageMaker一直就是模型构建平台,如今非常适合企业构建定制模型。他希望AWS逐步简化开放权重模型之间的比较、调优和测试。
在生存风险争论、安全漏洞和Hugging Face遭攻击的背景下,CEO们最关心的问题仍然很实际:如何信任运行在自身环境中的智能体?AWS给出的答案包括权限、沙箱、护栏、明确允许的行为,以及在适当场景下保留人工审核。
Garman将Continuum描述为一套把强大模型用于防御的系统:它扫描环境中的漏洞,再结合权限和补偿性控制措施等上下文,对漏洞进行优先级排序。他希望最终实现“机器速度的安全防护”,替代警报发出后等待人工调查的工作流。
10. Amazon自身的智能体采用开始重塑团队
Amazon正在安全和软件开发领域使用AI,Amazon Q也已经向每一名员工推出。Garman称,HR员工把原本需要数周的团队规划工作压缩到数小时完成,财务团队则使用智能体收集税务规则并确保合规。
最大的提升来自软件和产品开发。AWS的“frontier teams”采用智能体优先模式,而不是把AI当作代码补全工具:智能体负责编写代码,员工则管理智能体团队,由此带来Garman所说的客户能力发布“涡轮增压”。
组织设计仍未解决。Garman预计,在很长一段时间内,人仍然不可或缺,但一项过去需要10名员工负责的产品能力,如今可能只需要3或4人;而且它可能建得足够快,以至于这些人应该转去解决另一个问题。
AWS正在试验pod和更灵活的人员配置,同时面对维护问题:小团队构建的产品,必须有人继续运营。Garman目前没有提出万能的组织结构,只是观察到员工喜欢更快地构建、完成更多工作,而且“那里确实有真正的工作”。
完整逐字稿
Agentic workflows tend to perform better on AWS than anywhere else. Compute sandboxes, gateways, agent permissions versus people—a lot of those are things that we have built and are building and thinking actively about.
The top frontier labs gobble up all the available GPUs. At the same time, you want to promote newer companies that are going to become the enterprises of tomorrow. How are you thinking about balancing that?
From the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do. These are the innovators that are at the edge of technology, understanding what’s possible. We’re very intentional about keeping capacity available for the startups. We recently announced that we’re going to be buying 2 million NVIDIA GPUs over the next couple of years.
Your capex is what, $200 billion or—
$220 billion for 2026. We don’t anticipate slowing down anytime soon because the demand is just massive.
With all the debate around AI existential risk and the Hugging Face attack, what are CEOs asking you about all these things?
1. 169B, growing 37
Welcome to the pod, Matt. What a time we’re living in. I have a lot of topics to talk to you about.
Awesome. Thanks for having me. I’m excited.
Yeah, absolutely. So let’s start right from the beginning. You were the first GM for EC2.
Mm-hmm.
And that was 2006, right? Today, you guys are at $160 billion or $170 billion in revenue.
Yeah, about $169 billion or $170 billion.
Growing 30—
37%. 37%.
That’s 37% at $169 billion.
Yeah, that’s crazy.
There’s a ton of opportunity, and it’s interesting to think about it. We were there on day one, when we had the first dollar of revenue.
2. The GPU allocation problem
It’s still early stages of what the business can be and what the opportunity is for customers. Most workloads still live on-premises today, and the amount of compute that people are doing every single day is more than it was the day before. You see the tailwind from AI, you see the tailwind from migration into the cloud, and the business has grown really fast. It’s been a super-fun thing to be a part of.
It is. They’ll be writing history books and business books about this for a long time to come. I want to touch on the on-premises market, which sort of boggles my mind. I obviously did my best to keep them there for a long time.
You built a lot of stuff on-premises back in the day, trying to now get all of that to move into AWS.
3. Rethinking cloud for agents
Yeah, we can talk about that later. But if you think about EC2 in the early days, you got your start with obviously selling to startups. Today, as you reflect on the evolution, what stands out?
Yeah.
Part 1, and then part 2, we’ll talk about how the nature of how you serve startups has changed.
4. Why startups are the lifeblood
Sure. Well, like you said, it’s funny. I actually interned for AWS in 2005, when it was first an internal project. It was my business-school internship, and my project was to come up with an analysis of who we thought AWS would be most interesting to. The answer was startups, probably not surprisingly.
From the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do for a number of reasons. One is that the value proposition is just so attractive—what AWS provides to startups. We spend a lot of time and effort making sure that we’re great partners to the startups, helping them not just provide infrastructure but also providing advice on how to get their company up and running, how to think about their architecture so it’ll scale eventually, and a whole bunch of things that we do for startups.
We also think that, for us, it’s just good business, because the startups today are the enterprises of tomorrow. It’s an imprecise number, but today we estimate that maybe 30% to 40% of AWS revenue comes from companies that were at one time startups in AWS’s lifetime. It’s fun for us to see the companies grow over time, get bigger, and become enterprises effectively.
That’s why we invest so much in startups and pay so much attention to the brand-new, two-people-in-a-garage-type startups. It’s not just because of the business outcomes, but also because that’s who we learn from. These are the innovators that are at the edge of technology, understanding what’s possible and pushing our services to say what they’d want more of, what could help them go faster, and what could help them achieve their outcomes.
A lot of times, they’re pushing more than the banks, the healthcare companies, or the governments. The startups are the ones pushing that envelope. It really helps us to be better and make sure that we’re ahead of that wave, where the banks are going to want some technology or capability 5 years down the road that startups want today.
Yeah. Compared with then and now, how have the startups changed in what they want from you? Obviously—
They all want a lot of GPUs, and we can talk about that, but besides that, I’ll say there are a couple of things that have changed.
One is that startups started at a much smaller size than when we first started. They might have gotten $10 million of funding and had an app idea that they were slowly iterating on. Now, from day one, they’re valued at $1 billion and have $200 million of funding. It’s a team and an idea, and all of a sudden they’re worth $1 billion.
They just go from our offices to your offices.
That’s right. The size, scale, and ambition of the ideas require more capital. They’re bigger to start with, and they’re obviously more expensive to pursue, whether it’s training a model or doing something that a lot of the folks are doing today. So that’s number 1: the size at which they start is really big.
Number 2 is that some things haven’t changed. They’re still thinking about how to build an architecture that will work once they scale. How do they think about security? How do they think about performance? How do they think about having all the capabilities they need? How do they think about setting up their IAM so that when they have more than 3 employees, this thing is going to work and scale?
That’s a lot of times why they like AWS, as opposed to just going to a neocloud or something like that. They need all of those other security capabilities that come along with AWS, and so I think that’s something that hasn’t changed and is exactly the same. The scale at which they ramp up is definitely different today.
I also think, increasingly, we’re seeing teams that want a cloud that is great to work with agents and not just with people. We’ve spent a lot of time thinking about exactly how you think about broad scale, performance, and an interface that is a well-defined API interface that agents can easily traverse and work across.
That’s something we were naturally set up to do well, but we’ve also doubled down on it to ensure things like how you can start a database in 3 seconds and really get to capabilities that agents are excited about.
I should have looked, but I haven’t lately. Have you introduced any specific new services that are explicitly targeted at agents or people building agents?
I’d say what we’ve done is optimize some existing services so that they can work for both people and agents. Sometimes the answer is that there are definitely some services designed for building agents, such as AgentCore and Bedrock.
But when we look at taking an underlying component like S3, where the vast majority of companies store their data and have their data lakes, it turns out that a lot of the use cases for people and agents are similar. You want to have—and one of the things that we have in preview right now, or in beta, is called AWS Context. It allows you to build a context layer so that agents can more easily find all of the data they want across your various data lakes.
Whether you have your data stored in Aurora, in S3, or somewhere else in AWS, you can build this context layer. People aren’t necessarily going to access data in that way, but agents are happy to go across lots of those different things. There are some services that we’re building like that, but for the most part, you also care about what latency and throughput look like. You want to make sure that the underlying engine is fast and scalable.
Agents actually care a lot about tail latencies, which is interesting. People don’t always care about the p99.9 S3 latency, but agents do care and get blocked by that. That’s something we’ve cared about for a long time, and from a performance perspective, it’s one of the things that really popped.
That’s why agentic workflows tend to perform better on AWS than anywhere else.
5. Is the AI CapEx a bubble?
Yeah, obviously, we see a lot of companies starting out, and the very common refrain, of course, is that agents are writing all the code for them, right? And agents are selecting the databases, the email servers, and everything that you can name, right? So has there been a lot of thinking—and, obviously, they read the documentation—about how to rearchitect your, it feels funny to say, quote-unquote, legacy services that have been around for a decade?
I think there are some things that we’ve thought about. Actually, the core underlying building blocks, I think, are in really good shape. I think there’s a usability layer that we’re thinking about: how we make that easier to use for, we’ll call it, the very simplest use case you call out, where somebody’s just coding up an app really quickly and they ask it to deploy.
We find the vast majority of our customers will go and tell their coding agent—whether it’s Kiro, whether it’s Claude, whether it’s Codex—“I want to build on AWS. Here are my credentials, here’s my stuff. This is how I want to deploy it.” And then the agents are great and they go do it. There are use cases, though, where you’re not already on a cloud, you’re just trying something. You say, “Deploy.” And frankly, a lot of times that’ll go to some of our partners that have an easier-to-use layer on top, which we love, by the way. And we love our partners on those fronts, too.
But I do think that there are some things that we’re doing where, if you’re brand new, you don’t already have an AWS account and you haven’t already set up your IAM, today—or we’ll say a couple of months ago—it was much harder to start up an account, right? You had to—it’s been like this for 20 years—you had to define your VPCs, you had to define your IAM roles, all these kinds of things, which are actually super important once you get to be large. What we hear from customers is that it was a hard trade-off because they know that they’re going to want those things months or years down the line, but for now they just kind of want to use the services and not worry about that.
What we’ve actually launched is, if you go create a new AWS account today—and we’re slowly rolling this out; I don’t know if it’s actually fully rolled out yet—you don’t have to do those things. You don’t have to give it a credit card. You can sign up with your Gmail account. All of those things are handled by default behind the scenes, and within less than 30 seconds, you’re up and running and can be operating in a full AWS account, which is much more what the agents want to be able to do for those types of systems, because they don’t want to have to go through all that setup of your VPCs, your IAM, and pieces like that. So there are some things where we’re adding some ease of use to that, which we’re quite excited about, and I’ve seen some really positive feedback from customers on. We’ll keep doing more things like that over time.
One more thing, which is the great part about that, though, is it’s not like a simplistic account and then you have to migrate. If you basically say, “Okay, now I actually do want to develop an organization. I want to go and kind of fine-tune some of the things,” you can easily come—you’re already in a real AWS account—and you can actually then go and do all of those things later when you need them. So there’s no migration or move later, and that’s part of the hard work that we really think about: how do we make it not a choice for the customer, but an easy on-ramp into the depth of features that we know customers and startups are going to want when they start scaling?
Yeah. Got it. What have been some of the hardest things to accommodate as agents have taken over? I mean, you made a name—this is the way you guys became the phenomenon that you became—by serving developers, right? And then infrastructure teams. Now developers are being substituted by agents, and pretty soon infrastructure teams are being substituted by agents. What have been some of the hardest things for you, as the largest service provider in the world, to handle in that transition?
Well, I think, like I said, much of our infrastructure was pretty well set up to handle the scale, which is good. I think there have been a couple of things that are interesting—interesting paradigm shifts—where you could argue that some of our systems are, I wouldn’t say overengineered, but, as an example, many agents want to create a database, do a little bit of work, and then have the database go away. You really need that database to have five nines of durability? Exactly, right. And so there are some things that we rethink there, where, when you’re creating an Aurora database as your production database, you do want five nines of data durability—you need durability, you want availability, you want all of those things. For the agent use case, that’s arguably—or maybe not arguably—overengineered for what we need.
We don’t really want to have a nondurable option that’s going to cause problems either, because you never quite know if the database that’s created wants to stay around for a long time or a short amount of time. So we’re trying to do the hard work to think about how you accomplish both of those things, where you can create it quickly and throw it away and you don’t really waste a lot of resources, but if you do want it, it can be durable and stay for a long time. It can actually grow into a big production database.
Those are some trade-offs that we think about actively as we think about how the more traditional, “This is going to be my production system” capabilities match with some of the more transient nature of the infrastructure that agents want to use. That’s one, I think, but there’s a bunch. I think the scale, speed, and latency of creation of durable resources is another one that’s interesting.
I think the other one that we actively think about, too, is whether there are new building blocks that agents are going to want that we just didn’t really need before. Exactly. And so compute sandboxes, gateways, agent permissions versus people or service-role permissions—a lot of those are things that we have built and are building and thinking actively about, because they are just brand-new building blocks.
It’s not using the existing building blocks differently, but brand-new ones, where pretty clearly you want different permissions for agents. You don’t just want to give it Raghu’s permissions and let it go do whatever you can do. You actually want very time-boxed, short-term permissions to just go do a task. You may not want to give permissions to a whole tool at all. You may want to get very fine-grained permissions for what an agent is able to do from a sandbox, right? You actually want a sandbox that—
Virtualization is having its act together.
That’s right. Exactly. And they’re different. It’s similar—it’s the same idea—but how do you have that lightweight? Unfortunately, we’ve done a lot of work with Firecracker, our kind of microVMs. In fact, a lot of the sandbox company startups use Firecracker. They all use Firecracker, which we invented 10 years ago, maybe, or something like that.
It wasn’t purpose-built for agents, but it’s actually quite good for agents because you can spin them up really rapidly. They have a great security boundary, and you don’t have a lot of that virtualization overhead that you have from a traditional large VM. So I think those are some new building blocks that we’re thinking about, and more emerge every single day.
Yeah. So I’ve rattled on enough about agents. I’ll come back to it later, but it’s a fascinating topic. Matt Garman
Super cool.
But let me switch gears a little bit and ask about something that every one of our startups faces and that we get asked about most frequently, which is: how can we get GPUs?
I mean, obviously, you guys have massive GPU farms, and you’re increasing them every day, but the top frontier labs gobble up every available GPU. More power to them. How do you deal with that? At the same time, you want to promote newer companies that are going to become the enterprises of tomorrow, like you said. How are you thinking about balancing your capacity needs and these small companies that don’t have a lot of credit and whatnot?
Well, yeah, it’s a great question. A couple of these things make it more challenging. One is that the capex expense needed to deploy all of the compute that everyone needs right now is massive, right? And so—
Your capex is the same as what, $200 billion or something?
$220 billion for 2026. At one point I saw that it’s a pretty—I mean, that is a larger expense than we’ve ever had, maybe any company has ever had, in a single year. We don’t anticipate slowing down anytime soon because the demand is just massive. At some point, you’re limited by how fast you can build data centers, how fast you can deploy capital, and how fast you can get memory and chips and all of those kinds of things.
All of those are, at different points, constraints that we're dealing with—whether it's power, data centers, capital, memory, or chips. They're all constraints at various times. Construction people to build buildings are at a premium today, so we work really hard to make sure all of those things happen.
Then, with the capacity that we're able to deploy—which is still a huge amount, and not enough—we think very intentionally about our allocation strategy. We're great partners with the large frontier labs—Anthropic, OpenAI, Meta, and other large customers—and those folks are really good customers of ours.
Yep.
We want to make sure that we invest in them. We also have large enterprise customers, whether it's Salesforce, JPMorgan Chase, or other large companies, that have demand for fewer GPUs or accelerators. Sometimes they want Trainium; sometimes they want NVIDIA GPUs. We want to make sure that we can support them as well, and we're very intentional about making sure that we have capacity for startups.
What we do is allocate capacity. We basically say, “Okay, we could—you’re right—we could sell every single GPU or AI accelerator we had to probably just the big frontier labs and call it a day.”
We choose not to do that because we actually want to keep growing the full ecosystem. They get a large number, but we want to keep supporting a broader set of customers because we actually think that having the whole ecosystem will be healthier for us. There's some diversification, but it's also that we know these are going to be big companies over time, so we try to support them.
I saw recently that we say yes, in some way, shape, or form, to something like 60% of the requests we eventually get. Sometimes it's a little bit later; sometimes it's in a different region; sometimes it's a slightly different configuration than the customers are looking for. But we really try to lean in and allocate as much as we can.
Every single startup wants more, so we're working hard to make sure we have capacity for all of them. But it's hard. As a lot of people say, it's a good problem to have, but it's a problem nonetheless.
We continue to look at it. We recently announced that we're going to be buying 2 million NVIDIA GPUs over the next couple of years. We're landing a massive amount of capacity.
And 2 million?
2 million. So it's a lot, and it's over the next couple of years. Who knows if that's enough? At some point, again, we're limited by other components as well. But we're very intentional about keeping capacity available for startups. I know it's painful not to have enough, but we keep pushing.
When did your capex cross—I mean, pure AWS? Amazon was a bigger company.
Yeah.
Cross, let's call it, even $1 billion or $10 billion a year?
Oh, I don't know. I'd have to go back and look. I'm not sure about that, but it definitely scaled up over the last couple of years in a pretty meaningful way.
The AI buildout has definitely ramped our capex spending. Over the last 3 or 4 years, our capex has accelerated pretty meaningfully. I don't know when we crossed $1 billion, but given that we're at $220 billion now, we've been spending capex for a while. The company has been really good about funding that, and obviously the AWS business is a good one that we like to invest in.
Yeah, yeah, of course. For the longest time, we were investing ahead of where the demand was. Matt Garman
One of the most painful things is that, with the real ramp-up of GPUs, a lot of the elasticity has unfortunately gone away. Hopefully, in our core compute and storage, that elasticity is still there.
We've been spending for quite a bit of time, and we feel really good about the spend that we're making now. I get lots of questions about how we feel about that spend and whether we're nervous about a bubble or other things like that.
Because of the position we have, we take a diversified approach. Not all of our capacity is bundled up in one customer. You go to some of these neoclouds or some of the other providers out there, and you'll sometimes see concentrations of 30%, 40%, 50%, or 60% with 1 or 2 customers. We're nowhere near that. Obviously, we're in the single-digit percentages at the highest, and usually it's less than that.
For one, I think we have a lot less risk with any particular customer. But also, because we have that rich set of services, AWS is where people are really coming to launch their production workloads. The majority of our usage today is actually either core compute and storage or inference, which is part of that application. Those are the workloads that I think just aren't going to go away.
We're seeing enterprises get positive ROI. You go talk to the customers and say, “At the capability today and the cost today, are you seeing positive returns for your business?” Almost to a person, they'll say, “Oh, yeah.”
You're like, well, that's not going to go away. There's no bubble in which they stop spending on that. The VC model is that every billion-dollar startup isn't going to make it. No, they won't. But that's kind of the game, and it's been true for 50 years. They haven't always been billion-dollar startups; the startup numbers have changed, but the principles are the same, right? You bet on 10, and 1 makes it and pays for the other 10, or whatever the percentage is. Hopefully it's higher than 1.
You saw that with the internet, where there was a bubble and a bunch of internet companies didn't make it, and the internet is still a thing. A lot of the companies that had durable businesses—Google, Amazon, and others like that—did pretty well. So we feel really good about that investment and the continued investment going forward.
You guys have a view of the demand that's unparalleled, right? You're seeing across the globe and across every segment: enterprise, the big labs, AI-native companies, and so on and so forth. If anybody should be able to call it, you should be able to call it first.
I would hope so. Yeah, that's our plan. Honestly, we spend a lot of time thinking about it. We're very intentional about how we spend our shareholder capital, and we think that we're making great investments. We have a lot of good protections around how we intentionally invest that money.
We're very bullish on it. Andy has been public about saying this: The potential for AWS is really, really large. Over the next decade, the potential is there, and anytime you have an opportunity that's that big, you want to invest to go after it.
As the scale of these numbers gets larger and larger, right?
Yeah.
Has your planning and process dramatically changed? You're not writing billion-dollar checks; you're writing $50 billion checks or $20 billion checks, or whatever it is.
Yeah. Yes and no. A lot of times, we're still very bullish about the investments and lean forward, but the process has changed. There are a lot of things that have completely changed.
How do you even estimate 2028 demand? There are things that we have to think about now that we just never had to think about. If you go back 15 years, if we needed more power, we asked the power company to give us another couple of megawatts, whatever it was—tens of megawatts—and they would just give it to us because 10 megawatts wasn't that much. That was plenty for us to keep growing.
Now we have to bring our own power. We pay for power projects and renewable projects. We're one of the biggest renewable power purchasers each year, and we've been doing that for the last 10 years. We're regularly bringing on new solar projects and new nuclear projects.
So you're talking about behind the meter, or are you talking about working with the operator?
We'll do both. Often, with these power projects that we bring on, we'll pay the capital and pay for the project, and then it will go into the grid and we'll get credit for that. We'll bring those on.
Sometimes we'll do behind the meter, too. It's a mix. At the scale that we're doing, you have to think about all of those things. That's planning where you're thinking 20 years out about how you're going to think about power, how you're going to think about transmission, and how you're going to think about that capital. That's stuff we'd never had to think about before.
That’s planning we just never had to do. The other thing is, when we used to think about the server demand we needed, we’d have multiple-quarter demand forecasts and talk with our suppliers. Now we work multiple years out, just because the scale is so large that we have to think about what we need for 2026, 2027, and 2028.
But it’s also one of the values that we bring to customers, right? That’s something that customers legitimately can’t do themselves. They’re not going to do power, and they’re not going to plan their memory footprint in 2028. They just can’t do that. That’s one of the real values that we bring to our broad set of customers: it’s a whole set of things they don’t have to worry about, and that we spend a huge amount of time thinking about.
Yeah. You guys are generally, I think for the record, the largest buyers of practically every component of a server. Correct?
I don’t know that. I’m sure we’re one of the bigger purchasers of components out there, for sure. Who knows about everyone? It depends on how you think about them and how you measure it.
So where I was leading to: where do you see the constraints being most severe, let’s call it 2027 and 2028, and where do you see the constraints easing up?
It’s funny. I’ll answer this in a roundabout way. I remember that in undergrad—we actually read this a long time ago, and I never thought it would be a useful book—but we read “The Goal.” Have you read “The Goal”? It turns out there’s never one constraint; there’s always just the latest constraint. You have to think about all of them, and as soon as you hit one, there’s another one.
So what is the constraint that’s going to happen in 2028? I actually don’t know. I think there will be one. Every month for us, it’s: Do you have enough power? As soon as power is no longer the constraint, it might be memory, TSMC capacity, HBM, networking components, or a blip somewhere in the supply chain, like connectors. At some point, you have to think about all of those pieces.
It’s not just where they happen; that matters, too. We may have a ton of power in Indonesia, but not enough in Germany. You think about where in the world you want that capacity, too, because it turns out that not everything is totally fungible. Some is, and some isn’t.
It might be disk drives, it might be SSDs, or it might be something else. We think about every single component, and we have tens or hundreds of thousands of components that we track and think about. Some we rely on our suppliers to manage, and some we directly manage. We saw this problem coming probably a decade ago and really started not just thinking about how many servers we needed to track, but thinking all the way through the supply chain—four or five tiers down—to identify the component that could cause an issue for us and make sure that we had guaranteed supply of it.
If you remember—
I don’t remember when this was—over a decade ago, when there were floods in Thailand and no one had disk drives anymore.
There were disk-drive crises, and then there was a memory crisis.
Exactly. I think we think through all of those things. We also think about where there’s diversification in manufacturing—all of those pieces we try to work through. We’re never going to be perfect at it, but there’s always a different supply constraint.
Yeah. Obviously, there’s a lot of wide-ranging debate about data centers.
Mhm.
It’s clear where folks like us stand. But do you think, as an industry, we haven’t done a good job of explaining why data centers are good for America and generally the world? What’s the internal discussion amongst Andy’s team about how to deal with this?
Well, look, I think you’ll hear more from us on this, and I agree. I think we need to be more vocal and more upfront. We actually do a ton that’s really beneficial, both for the communities we operate in and more broadly. We think a ton about how to bring renewable energy to these data centers and how to be water-positive.
Actually, our data centers use a really, really small amount of water. We mostly use free-air cooling. We think about how to be great participants in the communities where we are and how to bring high-paying jobs to the communities we operate in.
Not all data center operators do that. There are some well-chronicled examples of others out there that aren’t great at it. They don’t really pay attention to regulations, they think the rules don’t apply, or they launch really quickly without thinking about those things. I think that causes problems for the whole industry because everybody gets lumped into that.
I think you’re right. We’re vocally self-critical. We need to be more vocal about the benefits that we do bring and think about additional ways that we can help communities understand those benefits, both for the services they use and more broadly. If you go to a community and say, “Do you not want to use Netflix?” they’ll say, “No, no. I still want Netflix.” It’s important for us to highlight the benefits that we bring.
I recently saw a report where, in one of the communities that we operate in, everybody in that county pays $5,000 a year less in taxes because of the taxes that we bring. We don’t tell them. They don’t even know it; it’s just invisible to them. I think we need to be more clear about those benefits, because if you told the communities, “By the way, your tax bill is $5,000 less than it would otherwise be if we weren’t here,” they might have a slightly different thought about the building that’s over there.
Not everyone does that.
Presumably, there’s more transparency needed.
I think much of the data center community, not just us, is actually made up of pretty good actors. There are just a few that aren’t, and I think they’ve caused some of the angst recently. We need to do a better job of highlighting who’s being a good citizen and who’s not.
Before we leave the hardware topic, I want to touch upon training and your whole history with building your own chips. We were one of your first partners using Nitro a long time ago.
Mhm.
Since then, with Graviton, you’ve made tremendous progress. What was the thinking that led to saying, “Look, we’re going to do our own thing”? How has that progress been, and where are you?
It’s actually a fascinating story, and I think it’s a great example of how AWS will innovate and iterate over time and continue to think bigger about what we can do, but prove our way there.
Take, as an example, this was probably about 13 or 14 years ago. We were seeing that there was a pretty significant virtualization tax on the overall number of resources. Our customers were telling us, “I want bare-metal performance,” and they were comparing it to having all the resources of a server.
The first thing we did was take a network offload card and move all of our network virtualization onto it, so that network virtualization got closer to bare-metal performance. Back then, it wasn’t quite bare-metal, but it was closer.
Then we got really excited about that and said, “What if we could move storage virtualization off as well?” None of the network offload cards could do that. Then we found this one company that had some Arm cores on an offload card. They were doing it for other reasons—I can’t remember their original purpose—but we thought, “Could you use those to do storage virtualization and some of these other functions?” They said, “Maybe.”
We really iterated with them. This was the product, and this was the team, and I just loved that team: really innovative, really mission-driven, and really wanting to solve problems.
We acquired them and said, “Look, could you build us a slightly bigger card that could take all the network virtualization off and basically give us a bare-metal server that had no virtualization on it—no VM virtualization? Everything was through APIs on the card.” We had the view that performance would be much better and resource utilization would be better.
Security isolation would be much, much better, and we could then legitimately tell people, “We have no access to any of your VMs that are running there.” This has been a huge benefit for us for the last decade where, frankly, we’ve been leading and others have been kind of slow to do this, because this is not a generalized thing that people can do.
We got to that, and we basically said, “Look, we’re making a lot of progress here. What if we take this?” There were a bunch of Arm cores on this offload card, and we said, “What if we turn that into a server?” We did that first with Graviton. It was a very underpowered, very small server that we launched.
Customers were excited. They said, “I’d love to have an Arm server. This is super interesting.” So we went down that path. With Graviton, part of what we did was look and see that the slope of Arm cores getting faster was increasing, and we saw the power utilization. You could see where the intercept was going to happen—where this architecture was going to be really good for workloads—and they just needed somebody to drive the ecosystem and get some of the pieces in place. So we did that with Graviton, and Graviton has been a runaway hit at this point.
Have you been public about what percent of your fleet is Graviton?
We ship more Graviton chips every year than any other type, so it’s very popular. They’re 20% cheaper with 20% better performance, and have been like that for the last 5 or 6 years. That’s an easy value proposition.
I think the vast majority—something like 90-plus percent—of our top 100 customers all use Graviton in some way, shape, or form across their fleets. That’s been a huge win for us and for customers. It’s the single biggest, easiest way that customers lower their bill: They move to Graviton.
We’ve had examples where people have moved their whole fleets and cut the number of servers they had in half.
The performance is so much better.
It’s amazing. Half the number of servers, and each server costs less. It’s a big win.
About 5 or 6 years ago, we saw the rise of AI compute happening. We didn’t nearly expect what it is today, but we still saw it was going to be a big, material mover. So we went in and built our first chips, Inferentia and Trainium, and we’re now in market with the 3rd generation, Trainium 3.
We’ve seen fantastic results. We’re sold out for capacity through probably toward the end of next year or something like that. We’re trying to figure out how we can save some capacity and get startups to be able to use some of it, because we see great results.
The majority of traffic on Bedrock runs on Trainium. We have great deals with both Anthropic and OpenAI to build on top of Trainium, as well as a set of smaller startups. I think we have half a dozen to a dozen startups that are building on top of Trainium now, too.
The name suggests it’s a training chip.
Yeah, we’re bad.
But everybody’s using it for inference, too. So where is the advantage architecturally, and where is it going?
It’s a good point. Look, we’re vocally self-critical: We’re terrible at naming. It’s not our strength. Originally, we had a chip called Inferentia for inference and Trainium for training. As the models got bigger and bigger, it turns out you actually want to run the inference on these really large systems. It turns out that Trainium is maybe the best inference chip on the market right now, from a pure performance and cost-performance point of view.
Is it better memory bandwidth? Where is the advantage?
The architecture is just a little bit different from others, and it’s much cheaper. From a cost-performance perspective and an absolute-performance perspective, Trainium is great.
We use it a ton for inference, and it drives much of the Bedrock inference today. It’s also a good training chip. A lot of our broad set of customers’ usage is not in training models, but in using models, so that’s where a lot of people get to use it under the covers. That’s where we’re excited about it.
A lot of the big customers are interested in it for training clusters as well, particularly as you get to Trainium 3 and 4. We’ve announced Trainium 4, but we haven’t launched it yet. Folks have looked at that architecture and said, “That’s the future of where I want my training clusters to be as well.” We’re quite excited about where that goes for these really broad-scale training clusters.
6. Where enterprises are stuck on agents
But, yeah, it’s built.
Yeah. Now let’s get back to talking about agents, but from the perspective of large enterprises or large and medium-sized enterprises. Where are they in their adoption? What sort of benefits are you seeing them reap already, and what is the roadmap for them, as far as you can tell from your vantage point?
It’s a really good question, and it’s one that we’ve spent a lot of time thinking about. When I talk to customers out there today, they view themselves as getting a lot of value out of what they’ve done. I would say the agents that most enterprises have built are relatively simple and straightforward, and they’re starting to think about how to make them autonomous in a safe way.
Mostly nonautonomous, right?
They’re still people in the loop, if you will. Customers are still getting lots of value out of that today. They’re really thinking about how to make these systems autonomous in a safe way.
I think there are 2 things that hold customers back today from continuing to scale. It’s already a pretty big business today, but I think it has a massive opportunity to really change every single customer and every single workflow.
Number 1 is just how to think about it. What we originally saw was that enterprises had a workflow and were saying, “Great, I would have an agent go do the same workflow.” What we encourage them to do is not just replicate it: Bob does steps 1, 2, 3, 4, and 5, so the agent is going to do steps 1, 2, 3, 4, and 5, and then Bob’s going to check it at the end. That’s not really the model you want.
You want to step back and say, “If I want to accomplish something, how can an agent do it differently?” It can do it in a massively parallelized way. It can try 50 different things and get to the right outcome. How do you help it get to that right outcome? You rethink how a computer would solve a problem versus how a human would solve a problem.
One of the things is helping customers understand how to think about that and really have that blank slate, because that’s where you really get value. It’s not just replicating what you’re doing today, but thinking from a greenfield approach about how you solve a problem completely differently. I’m sure that’s how many of your startups are thinking about this, too: How do you help customers greenfield-solve a problem, not replicate the thing that happens today?
So that’s number 1.
People running fleets of agents, swarms, whatever you want to call them. Yeah—
Very common these days.
And you just want to think about it. Enterprises are not as forward-leaning. Again, this is one where you learn from the startups and try to apply that to an enterprise world. An insurance company is not necessarily as forward-leaning, but they would love to figure out how they can have a better approval workflow or something like that.
That’s number 1. The second one is how you turn those into fully autonomous workflows and how you actually trust the agents. We’re spending a lot of time thinking about how to build services to help enterprises feel like their systems are secure and that they can trust an agent to make a decision, that they can have the right guardrails, that it can have the right permissions on their data, that it’s not going to delete production systems, and that it’s not going to make tragic mistakes.
Right now, I think that nervousness is holding people back—maybe appropriately, by the way—from just saying, “Okay, go nuts.” You don’t actually want an agent to go crazy and accidentally delete a production database. That’s going to be pretty bad.
We’re actively working through this with customers. How do we both help them architect and, frankly, invent new technologies and capabilities that are going to help them solve that problem? That’s one of the areas where I think we’ll continue to innovate, and we’ll get there. I think we have some really good ideas and some good technologies brewing that can really help.
So, are enterprises learning how to do eval systems and so on and so forth to keep the agents? They need help, honestly. Both with evals—how do you have a constant loop of testing? How do you think about goal-seeking in a reasonable way? How do you have your data labeled in such a way that it actually makes sense, so the eval can approximate what you're going to be doing in production? How do you measure in production and back-test it so you're not seeing drift?
All of those things are problems that enterprises don't know how to solve today. I don't know if anyone really is great at solving these today. It's why you've seen so many FDE teams spin up, and AWS and our partners are really leaning into the FDE motion to go and help. This is the single biggest area where customers need help.
And when we think about how to do that, my view is that we want to train our customers to be able to do this themselves. This is not the traditional notion where I want to have a people-driven business that goes on forever, where you just keep paying consultants over and over and over again. Our view on how FDEs should work is that we want to go into a customer who's ready to accept owning this when we're done, and in 45 days do work where we can teach them how to make an eval and teach them how to get their data in a labeled way.
We do the work alongside them, and then, at the end of 45 days, we leave and that customer is good and ready to go and trained up. That's what our customers tell us they want. They don't want to be beholden to an external workforce for the next 5 years.
But they need help today.
Yeah.
You made a massive investment in FDEs.
FDEs.
7. AI risk, Hugging Face, and Continuum
So, taking that even one step further, some of your industry peers have said, “Look, you can't have all of your data going into a big frontier model. What enterprises should really do is take an open-source model and then post-train on your own data, workflows, preferences, and whatnot.”
Where do you stand on that? Are you seeing customers actually trying to do that, or how do you think about that?
It's a great—the first point, I wholeheartedly agree on that first point: enterprise data is their most valuable asset. From the very beginning, that's why we built Bedrock like we did. We have a guarantee that your data never leaves your VPC. If you're running inside of Bedrock, your data doesn't go back to the model provider. They never see your prompts. That stays inside of your own trusted environment.
That is why enterprises prefer to run on top of Bedrock, and it's why you see that business growing massively. Every month, we see that business exploding, and it's why you see OpenAI workloads migrating to Bedrock. It's why you see Anthropic really growing rapidly.
Whether you're using open models or closed frontier models, I think Bedrock is a great solution. Our customers told us this, by the way. If you remember, 3 years ago I got a lot of heat.
Speaking of bad names, that's a good name, though.
8. The Graviton and Trainium bet
Yeah, Bedrock is a good name. That's good. But we got a lot of heat, actually, for being slow to the AI world because we built the foundations of this. We said, “Look, we're not just going to rush out a service. We really want to think about how we make sure that we protect our customers' data and build a service that we think is going to be durable for the use cases that we knew about.”
And if you remember, we got a lot of heat, and we said, “Look, we're going to go build the right thing.” Now, as people move from proof of concepts to production, the vast majority of them are landing in AWS on Bedrock. One of the reasons is because of this. It's also because of the set of services that we have.
We offer open models. We offer proprietary models. We offer a whole set of capabilities around those—AgentCore. We build these building blocks so it's easier to build agents with any of the models that you want, whether they're in Bedrock or out of Bedrock, for that matter. You can use Gemini or other things for it.
I think that's a differentiating piece for us, and it's a super important thing to think about because having that data go back into the model provider is a dangerous thing. You talk about open-weights models, though. I do think that there's an area that I'm excited about. We're really ramping up our support of open-weights models and trying to build a good environment.
Frankly, this is where today I think—and I think it's true of a lot of customers—people believe that they have meaningful proprietary data that, if they could mix it in and do some post-training or fine-tuning to an open-weights model, they could distill down and actually get a better-performing model at a lower price.
There are a bunch of pieces in here that have to work out well. They actually have to have a good eval to prove that that's true. Most people are doing that in SageMaker today. I think there's more that we can do to make that easier. But if you go look at where people are doing that, they actually do it in SageMaker on AWS today, and they host the inference via SageMaker.
SageMaker is getting a new lease on life as a—
I mean, it is. It was always kind of a model-building platform, right? And now, if you think about what enterprises are doing in this world, that's what they're doing: they're basically building their own custom models. SageMaker is a great place for doing that.
I think there are some things that we need to keep building on to make that easier and easier to do and to test across different open-weights models and things like that. But that's an evolving space. I think it's a super interesting one, and it's one we want to make sure that we have all the right things for customers to be able to do if they have the right data and expertise to actually go down that path.
And with all the debate around AI existential risk, this, that, and the other, and the security vulnerabilities and so on and so forth, and the Hugging Face attack, what do CEOs ask you about all these things?
Yeah, there's a bunch. They mostly want to say—and it goes back to this—“When I launch agents, how can I trust that they're going to do what I want them to do?”
We're heavily investing in building capabilities that allow people to deploy agents safely into their environment and think about those controls. Some of those are: How do you make sure the agents have the right permissions? How do they have the right sandboxing? How do you make sure that you have the right set of guardrails? How do you intentionally think about what you want the agents to do and not do? Is there a human in the loop or not?
We spend a lot of time with our customers thinking about safe agent deployment and how to get better over time, and what other capabilities we need to build to help people deploy agents safely into their environment. So, there's a lot there.
The other angle that a lot of people are worried about is whether some of these really powerful models are going to be attack surfaces and be misused. I have a view on that: yes, that is a real risk to customer environments, but it's also a real opportunity.
We recently launched a service called Continuum that uses these powerful models to help customers secure their environment. We'll look across the environment and look for vulnerabilities. We'll use some of these powerful models to help customers find vulnerabilities they haven't found before and, most importantly, prioritize which ones, because we know context about their environment: how it's set up, where their permissions are, and where they may have compensating controls that make it harder or easier to exploit them.
Continuum is incredibly popular with customers. We're really bullish about what's possible from AI to help with AI-powered security, because at some point customers are going to need security at machine speed, not at human speed—not an alert that someone goes in and looks at.
We're running fast to help build that for customers and help them protect their environments, and I'm very excited about what the Continuum team is building on that front, too.
So, within AWS itself—
Mm-hmm.
What's the state of adoption and usage of agents broadly?
Well, Continuum is basically us trying to expose what we do internally. We use AI extensively for our own security. We use AI extensively for our own software development. We use agents—actually, one of the things that's really cool is that we use agents across our entire business.
We rolled out Amazon Q to every single Amazon employee. Now I see HR teams building agents to help drive what used to take teams of people weeks to do, which a single person can now do in a couple of hours, such as team planning and resource management.
I have finance teams that are building agents to pull tax rules from everywhere and ensure compliance on a bunch of different pieces. It’s super cool to see that things that used to be blocked by software developers—actually, the line-of-business folks are able to go and unblock themselves and innovate more quickly.
Amazon Q has been an enormous boon, and that has grown like wildfire. We see customers like small startups using it all the way to the largest enterprises in the world rolling it out to their entire customer base to get the benefit of being able to access all of their enterprise data and easily apply agents and capabilities to help accelerate their jobs.
And so, as I said, we use it across everything from software development to security to driving HR policies. You guys are notorious for measuring everything about your operation. Where have you seen the biggest gains?
Yeah, I mean, obviously, you know the answer: software development. That’s the real answer.
The derivatives of that.
But yeah, I mean, the speed of software development—and really, product development as a whole, not just coding—is absolutely the case. The pace at which we’re deploying new products is massively different than it has been in the past.
I think you can see this. AWS has always been known for rolling out features really quickly, and we’ve seen a turbo boost on that in the last year or so as we’ve developed what we call frontier teams, as they think about agentic development as opposed to traditional development. It’s not code completion; it really is agent-first. The agents write all of the code. You’re just managing a team of agents and driving that.
It’s been fun to see the pace at which innovations for customers have been happening. It has to be, because our customers out there have an almost insatiable appetite for new capabilities, and that’s what we’ve got to do.
Yeah. So, organizationally, do you have any insights on how organizations should change?
I don’t know. I’ll say that—
Agents manage people, people manage agents?
There are going to be lots of people for a long period of time. I do think organizations will change. I don’t know the magic answer yet, but we’re actively thinking about it.
Inside, you’re just trying various experiments, or what?
We’re trying experiments. We’re thinking about pods. As you think about it, here’s one example: In a product organization, you used to have a team that would own a particular capability for a long time, and you might have 10 people working on that thing. Well, today, you can innovate so rapidly that it doesn’t have to be 10 people. It can be 3 to 4 people, and they build something so fast that you actually want to move them to different projects and problems.
Thinking about how you both operate and maintain the things that you’ve built while being agile and flexible enough to move around in an organization as big as AWS is—that’s an active area that we’re experimenting with and playing with. It’s fun, and it’s enabling for our employees. They actually love it because they can build faster and do more. But there’s real work there.
Yeah, it’s a fascinating time. Thank you very much for your time. We could be talking about this for hours together, but thanks for all your insights.
Yeah, thank you for having me. We love having all your companies as customers, and we love learning from them. I appreciate you having me here.
Yeah, we’ll keep sending them your way.
Excellent. Thank you.
Thanks.