[BidClub_]
The a16z Show · · 56 min

Building the Cloud for AI Agents | AWS CEO Matt Garman

Raghu RaghuramMatt Garman

EquitiesAI & SoftwareSemisCompany BuildingTechnical
YouTube ↗
TL;DR
  • AWS is betting that AI demand is durable enough to justify $220 billion of 2026 capex, with no slowdown anticipated. Matt Garman argues the risk is tempered by AWS’s diversified customer base—individual-customer concentrations are in the single-digit percentages at the highest—and by production workloads already generating positive ROI. “There’s no bubble in which they stop spending on that.”

  • AWS is allocating scarce accelerators across the ecosystem rather than selling them all to frontier labs. Although AWS could sell every available chip to a few large labs, it reserves capacity for startups and eventually says yes in some form to roughly 60% of requests it receives, sometimes in another region or configuration. AWS plans to buy 2 million NVIDIA GPUs over the next couple of years, though Garman concedes, “Who knows if that’s enough?”

  • Agentic workloads push cloud architecture toward fast, potentially disposable, tightly permissioned resources alongside long-lived systems. Agents care about p99.9 latency, three-second database creation, compute sandboxes, gateways and time-boxed permissions rather than inheriting a person’s access. Garman’s core claim is categorical: “Agentic workflows tend to perform better on AWS than anywhere else.”

  • AWS is removing its traditional setup friction without forcing startups onto a simplified dead end. As the new flow rolls out, an account can be opened with Gmail, without a credit card, and become usable in under 30 seconds while VPC and IAM defaults are handled behind the scenes. Crucially, it remains “a real AWS account,” so customers can later expose the full security and organizational controls without migrating.

  • AWS’s custom-silicon path runs from Nitro and Graviton to Trainium. Graviton is described as 20% cheaper with 20% better performance, used by something like 90-plus percent of AWS’s top 100 customers; Trainium 3 capacity is sold out potentially through late next year, while Trainium 4 has been announced but not launched. Garman says Trainium may be “the best inference chip on the market right now” on absolute and cost performance.

  • Enterprise agent adoption is constrained less by model access than by workflow redesign, evaluation and trust. Garman rejects merely copying a human’s five-step process: agents can try 50 approaches in parallel, but enterprises need permissions, guardrails, labeled data, production measurement and drift testing before granting autonomy. AWS’s forward-deployed model aims to teach those skills in 45 days and leave, not create years of consultant dependence.

  • AWS is positioning Bedrock around enterprise data custody while keeping both proprietary and open-weight model paths available. Garman says prompts and data remain inside the customer’s VPC and never return to the model provider; customers with genuinely proprietary data may post-train or fine-tune open-weights models, potentially distilling them, for better performance at lower cost. The prerequisite is an evaluation system proving that advantage rather than assuming it.

  • Inside Amazon, agents are already compressing software, product and business workflows—and beginning to reshape team design. Garman says agent-first “frontier teams” manage agents that write all the code, while HR and finance employees build their own agents for planning, tax and compliance work. A capability that formerly required 10 people might now need three or four, creating a new problem: how to maintain products while moving small teams rapidly between projects.

Digest · the substance, structured for research

1. AWS still sees itself near the beginning

  • Raghu Raghuram opens with the scale: AWS has grown from its first dollar of revenue to roughly $169 billion-$170 billion, now growing 37%. Garman nevertheless calls it “the early stages of what the business can be,” because substantial workloads remain on-premises while AI expands the total amount of compute performed each day.

  • Garman’s first AWS assignment was a 2005 business-school internship analyzing who would value the proposed service most. The answer was startups, which became “the lifeblood of the core of what we do” because AWS’s value proposition was especially attractive to them and helped them build architectures that could eventually scale.

  • That relationship has compounded commercially: AWS estimates that 30%-40% of current revenue comes from companies that were startups at some point during AWS’s lifetime. The two-person company matters not only as a future enterprise customer but as an early signal of capabilities that banks, healthcare companies and governments may demand five years later.

  • The startup baseline has changed radically. Where a company might once have raised $10 million to iterate on an app, Garman now sees teams with an idea, $200 million in funding and a $1 billion valuation from day one; their ambition, model-training costs and infrastructure ramp are correspondingly larger.

2. Agents expose a new cloud performance envelope

  • Some startup requirements have not changed: founders still need scalable architecture, security, performance and IAM that survives growth beyond three employees. Garman says this surrounding depth is why many ultimately prefer AWS to a neocloud, even when their immediate request is simply for GPUs.

  • What has changed is the user. AWS now designs for clouds operated by agents as well as people, emphasizing broad-scale API access, fast resource creation and predictable performance. Garman offers a concrete example: “How can you start a database in three seconds?”

  • AgentCore and Bedrock are explicit agent-building services, while “AWS Context,” described as being in preview or beta, is intended to create a context layer across data in S3, Aurora and other AWS stores. An agent can traverse those separate data estates in a way a person typically would not.

  • Agents also magnify tail latency. A human may never notice p99.9 S3 performance, but an agentic workflow can block on it; latency, throughput and the speed of the underlying engine therefore become orchestration constraints rather than abstract infrastructure metrics.

3. AWS is simplifying the on-ramp without removing the controls

  • Raghuram’s pushback is that coding agents now select databases, email systems and deployment targets themselves, potentially making decade-old services feel “legacy.” Garman answers that the core building blocks remain sound; the larger gap has been the usability layer for someone—or some agent—starting without an existing cloud account.

  • In a typical established workflow, a customer tells Kiro, Claude or Codex to build on AWS, supplies credentials and deployment instructions, and the agent handles the rest. An experiment that is not already on a cloud and merely says “deploy,” however, may choose a partner with an easier layer on top—an outcome Garman says AWS welcomes but also wants to address directly.

  • AWS’s new-account flow is being rolled out so users can sign up with Gmail, omit a credit card and avoid manually defining VPCs or IAM roles. Defaults are handled behind the scenes, producing a functioning account in under 30 seconds.

  • The design choice Garman stresses is continuity: this is not a toy account that later requires migration. When the customer needs an organization, custom VPC or fine-grained IAM, “you’re already in a real AWS account” and can progressively expose those controls.

4. Disposable agent infrastructure needs production-grade escape hatches

  • Agents often create a database, perform a small task and discard it, raising a genuine design question: does that database need five nines of durability? AWS’s traditional Aurora posture assumes a production asset requiring durability and availability, which can be “overengineered” for a transient agent task.

  • Garman resists solving that with a casually non-durable option because AWS cannot always know whether a temporary resource will become permanent. The engineering objective is therefore dual-use infrastructure: quick and resource-light enough to discard, yet capable of growing into a durable production database without replacement.

  • Some agent requirements are genuinely new building blocks rather than different uses of old ones: compute sandboxes, gateways and permissions distinct from human or service-role permissions. “You don’t just want to give it Raghu’s permissions”; an agent may need narrowly scoped, time-boxed authority for one task, perhaps without access to an entire tool.

  • Firecracker microVMs provide a ready substrate. Developed roughly a decade ago and now used by many sandbox startups, they start rapidly, impose little traditional-VM overhead and retain a strong security boundary—making them well suited to agent execution.

5. Scarcity makes GPU allocation a strategic portfolio decision

  • Raghuram frames the conflict directly: frontier labs can “gobble up every available GPU,” while smaller companies lack comparable credit but may become tomorrow’s enterprises. Garman agrees AWS could allocate every accelerator to a handful of labs, but says it intentionally reserves supply for startups, enterprises and a broader ecosystem.

  • The buildout is enormous: Garman cites $220 billion of capex for 2026 and says AWS does not anticipate slowing because demand remains “massive.” Constraints rotate among power, data centers, capital, memory, chips and even the construction labor required to erect facilities.

  • AWS says yes in some form to roughly 60% of the requests it eventually receives, though fulfillment may arrive later, in another region or with a different configuration. Every startup still wants more; the company’s announced purchase of 2 million NVIDIA GPUs over the next couple of years may itself prove insufficient.

  • Garman distinguishes AWS from providers whose top one or two customers can represent 30%-60% of capacity. AWS’s individual-customer concentrations are in the single digits at the highest, while most usage is core compute, storage and inference tied to applications; nearly every enterprise he asks says current AI capabilities already produce positive returns.

6. The limiting resource keeps moving down the supply chain

  • Asked to identify the 2027-2028 bottleneck, Garman invokes The Goal: “There’s never one constraint; there’s always just the latest constraint.” Once power eases, the limit might become memory, TSMC capacity, HBM, networking gear, connectors, disk drives or SSDs.

  • Geography compounds the issue because capacity is not wholly fungible: ample power in Indonesia does not cure a shortage in Germany. AWS tracks tens or hundreds of thousands of components and looks four or five tiers into the supply chain for parts that could halt deployment.

  • Planning horizons have expanded from asking a utility for another few megawatts to financing solar, nuclear and other power projects, sometimes behind the meter and sometimes feeding the grid. Power and transmission decisions now extend 20 years, while planning for server, memory and chip needs stretches across 2026, 2027 and 2028.

  • Data-center opposition raises a communication challenge. Garman says AWS needs to communicate renewable-energy, water and employment benefits more clearly, citing one county where people reportedly pay $5,000 less in taxes annually because of the taxes AWS brings—an invisible benefit residents had not been told about.

7. Nitro led AWS from virtualization offload to custom AI silicon

  • AWS’s chip path began with the “virtualization tax.” It first moved network virtualization onto an offload card, then worked with a small team whose Arm-equipped card could absorb storage virtualization and other functions; acquiring that team ultimately enabled Nitro to expose near-bare-metal resources through APIs.

  • The payoff was not only utilization and performance but isolation: Garman says AWS could credibly tell customers it had no access to their running VMs. Because the architecture was not a generalized component others could simply buy, he argues it produced a decade-long lead.

  • Turning those Arm cores into a server yielded an initially underpowered Graviton, followed by a “runaway hit” as Arm performance improved. Garman says Graviton has remained roughly 20% cheaper with 20% better performance for five or six years; something like 90-plus percent of AWS’s top 100 customers use it, and some fleet migrations halved server counts.

  • AWS began Trainium five or six years ago and is now shipping Trainium 3; capacity is sold out potentially toward the end of next year. Most Bedrock traffic runs on Trainium, alongside Anthropic and OpenAI agreements and roughly half a dozen to a dozen smaller startups building on it.

8. Enterprise autonomy depends on redesign, evaluation and data trust

  • Most enterprise agents today remain relatively simple and non-autonomous, though customers already report value. Garman’s first prescription is to stop reproducing “Bob does steps one, two, three, four, five”; an agent can parallelize, try 50 approaches and solve the desired outcome differently from a human workflow.

  • The second barrier is justified nervousness about autonomy. Enterprises need guardrails, sandboxing, data permissions and clarity on when a human stays in the loop; “go nuts” is not an acceptable instruction when an agent could delete a production database.

  • Raghuram lists the evaluation requirements—labeled data, a constant testing loop, goal-seeking criteria, production measurement, back-testing and drift detection—and Garman says enterprises do not know how to solve these today. He doubts anyone is really great at solving them, motivating forward-deployed engineering engagements intended to teach the capability in 45 days and then leave the customer self-sufficient.

  • On model choice, Garman agrees enterprise data is “their most valuable asset.” Bedrock guarantees that data stays within the customer’s VPC and that model providers never see prompts, a foundation AWS prioritized even when critics said it was moving too slowly three years earlier.

9. Open models and machine-speed security broaden the AWS layer

  • Customers with meaningful proprietary data may post-train or fine-tune an open-weight model, distill it and potentially achieve better performance at lower cost. Garman keeps the claim conditional: they need the right data, expertise and evaluations to prove the custom model actually wins.

  • Raghuram calls this “a new lease of life” for SageMaker; Garman says it was always a model-building platform and is now a good fit for enterprises building custom models. He wants AWS to make comparison, tuning and testing across open-weight models progressively easier.

  • CEOs’ overriding question amid existential-risk debate, security vulnerabilities and the Hugging Face attack is still practical: how can they trust agents inside their environments? AWS’s answer spans permissions, sandboxes, guardrails, explicit allowed behaviors and human review where appropriate.

  • Garman describes Continuum as turning powerful models toward defense: it scans environments for vulnerabilities, then prioritizes them using context about permissions and compensating controls. His destination is “security at machine speed,” replacing a workflow where an alarm waits for a human investigation.

10. Amazon’s own agent adoption is beginning to reshape teams

  • Amazon uses AI across security and software development, while Amazon Q has been rolled out to every employee. Garman describes HR staff compressing weeks of team-planning work into hours and finance teams using agents to collect tax rules and ensure compliance.

  • The largest gain is software and product development. AWS’s “frontier teams” work agent-first rather than treating AI as code completion: agents write the code while employees manage teams of agents, producing what Garman calls a “turbo boost” in the release of customer capabilities.

  • Organizational design remains unresolved. Garman expects people to remain important for a long time, but a product capability that once occupied 10 employees may now need three or four—and may be built quickly enough that those people should move to another problem.

  • AWS is experimenting with pods and more fluid staffing while confronting the maintenance question: somebody must continue operating what a small team built. Garman offers no magic structure yet, only the observation that employees like building faster and doing more—and that “there’s real work there.”

Full transcript
Matt Garman

Agentic workflows tend to perform better on AWS than anywhere else. Compute sandboxes, gateways, agent permissions versus people—a lot of those are things that we have built and are building and thinking actively about.

Raghu Raghuram

The top frontier labs gobble up all the available GPUs. At the same time, you want to promote newer companies that are going to become the enterprises of tomorrow. How are you thinking about balancing that?

Matt Garman

From the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do. These are the innovators that are at the edge of technology, understanding what’s possible. We’re very intentional about keeping capacity available for the startups. We recently announced that we’re going to be buying 2 million NVIDIA GPUs over the next couple of years.

Raghu Raghuram

Your capex is what, $200 billion or—

Matt Garman

$220 billion for 2026. We don’t anticipate slowing down anytime soon because the demand is just massive.

Raghu Raghuram

With all the debate around AI existential risk and the Hugging Face attack, what are CEOs asking you about all these things?

1. 169B, growing 37

Welcome to the pod, Matt. What a time we’re living in. I have a lot of topics to talk to you about.

Matt Garman

Awesome. Thanks for having me. I’m excited.

Raghu Raghuram

Yeah, absolutely. So let’s start right from the beginning. You were the first GM for EC2.

Matt Garman

Mm-hmm.

Raghu Raghuram

And that was 2006, right? Today, you guys are at $160 billion or $170 billion in revenue.

Matt Garman

Yeah, about $169 billion or $170 billion.

Raghu Raghuram

Growing 30—

Matt Garman

37%. 37%.

Raghu Raghuram

That’s 37% at $169 billion.

Matt Garman

Yeah, that’s crazy.

Raghu Raghuram

There’s a ton of opportunity, and it’s interesting to think about it. We were there on day one, when we had the first dollar of revenue.

2. The GPU allocation problem

Matt Garman

It’s still early stages of what the business can be and what the opportunity is for customers. Most workloads still live on-premises today, and the amount of compute that people are doing every single day is more than it was the day before. You see the tailwind from AI, you see the tailwind from migration into the cloud, and the business has grown really fast. It’s been a super-fun thing to be a part of.

Raghu Raghuram

It is. They’ll be writing history books and business books about this for a long time to come. I want to touch on the on-premises market, which sort of boggles my mind. I obviously did my best to keep them there for a long time.

Matt Garman

You built a lot of stuff on-premises back in the day, trying to now get all of that to move into AWS.

3. Rethinking cloud for agents

Raghu Raghuram

Yeah, we can talk about that later. But if you think about EC2 in the early days, you got your start with obviously selling to startups. Today, as you reflect on the evolution, what stands out?

Matt Garman

Yeah.

Raghu Raghuram

Part 1, and then part 2, we’ll talk about how the nature of how you serve startups has changed.

4. Why startups are the lifeblood

Matt Garman

Sure. Well, like you said, it’s funny. I actually interned for AWS in 2005, when it was first an internal project. It was my business-school internship, and my project was to come up with an analysis of who we thought AWS would be most interesting to. The answer was startups, probably not surprisingly.

From the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do for a number of reasons. One is that the value proposition is just so attractive—what AWS provides to startups. We spend a lot of time and effort making sure that we’re great partners to the startups, helping them not just provide infrastructure but also providing advice on how to get their company up and running, how to think about their architecture so it’ll scale eventually, and a whole bunch of things that we do for startups.

We also think that, for us, it’s just good business, because the startups today are the enterprises of tomorrow. It’s an imprecise number, but today we estimate that maybe 30% to 40% of AWS revenue comes from companies that were at one time startups in AWS’s lifetime. It’s fun for us to see the companies grow over time, get bigger, and become enterprises effectively.

That’s why we invest so much in startups and pay so much attention to the brand-new, two-people-in-a-garage-type startups. It’s not just because of the business outcomes, but also because that’s who we learn from. These are the innovators that are at the edge of technology, understanding what’s possible and pushing our services to say what they’d want more of, what could help them go faster, and what could help them achieve their outcomes.

A lot of times, they’re pushing more than the banks, the healthcare companies, or the governments. The startups are the ones pushing that envelope. It really helps us to be better and make sure that we’re ahead of that wave, where the banks are going to want some technology or capability 5 years down the road that startups want today.

Raghu Raghuram

Yeah. Compared with then and now, how have the startups changed in what they want from you? Obviously—

Matt Garman

They all want a lot of GPUs, and we can talk about that, but besides that, I’ll say there are a couple of things that have changed.

One is that startups started at a much smaller size than when we first started. They might have gotten $10 million of funding and had an app idea that they were slowly iterating on. Now, from day one, they’re valued at $1 billion and have $200 million of funding. It’s a team and an idea, and all of a sudden they’re worth $1 billion.

Raghu Raghuram

They just go from our offices to your offices.

Matt Garman

That’s right. The size, scale, and ambition of the ideas require more capital. They’re bigger to start with, and they’re obviously more expensive to pursue, whether it’s training a model or doing something that a lot of the folks are doing today. So that’s number 1: the size at which they start is really big.

Number 2 is that some things haven’t changed. They’re still thinking about how to build an architecture that will work once they scale. How do they think about security? How do they think about performance? How do they think about having all the capabilities they need? How do they think about setting up their IAM so that when they have more than 3 employees, this thing is going to work and scale?

That’s a lot of times why they like AWS, as opposed to just going to a neocloud or something like that. They need all of those other security capabilities that come along with AWS, and so I think that’s something that hasn’t changed and is exactly the same. The scale at which they ramp up is definitely different today.

I also think, increasingly, we’re seeing teams that want a cloud that is great to work with agents and not just with people. We’ve spent a lot of time thinking about exactly how you think about broad scale, performance, and an interface that is a well-defined API interface that agents can easily traverse and work across.

That’s something we were naturally set up to do well, but we’ve also doubled down on it to ensure things like how you can start a database in 3 seconds and really get to capabilities that agents are excited about.

Raghu Raghuram

I should have looked, but I haven’t lately. Have you introduced any specific new services that are explicitly targeted at agents or people building agents?

Matt Garman

I’d say what we’ve done is optimize some existing services so that they can work for both people and agents. Sometimes the answer is that there are definitely some services designed for building agents, such as AgentCore and Bedrock.

But when we look at taking an underlying component like S3, where the vast majority of companies store their data and have their data lakes, it turns out that a lot of the use cases for people and agents are similar. You want to have—and one of the things that we have in preview right now, or in beta, is called AWS Context. It allows you to build a context layer so that agents can more easily find all of the data they want across your various data lakes.

Whether you have your data stored in Aurora, in S3, or somewhere else in AWS, you can build this context layer. People aren’t necessarily going to access data in that way, but agents are happy to go across lots of those different things. There are some services that we’re building like that, but for the most part, you also care about what latency and throughput look like. You want to make sure that the underlying engine is fast and scalable.

Agents actually care a lot about tail latencies, which is interesting. People don’t always care about the p99.9 S3 latency, but agents do care and get blocked by that. That’s something we’ve cared about for a long time, and from a performance perspective, it’s one of the things that really popped.

That’s why agentic workflows tend to perform better on AWS than anywhere else.

5. Is the AI CapEx a bubble?

Raghu Raghuram

Yeah, obviously, we see a lot of companies starting out, and the very common refrain, of course, is that agents are writing all the code for them, right? And agents are selecting the databases, the email servers, and everything that you can name, right? So has there been a lot of thinking—and, obviously, they read the documentation—about how to rearchitect your, it feels funny to say, quote-unquote, legacy services that have been around for a decade?

Matt Garman

I think there are some things that we’ve thought about. Actually, the core underlying building blocks, I think, are in really good shape. I think there’s a usability layer that we’re thinking about: how we make that easier to use for, we’ll call it, the very simplest use case you call out, where somebody’s just coding up an app really quickly and they ask it to deploy.

We find the vast majority of our customers will go and tell their coding agent—whether it’s Kiro, whether it’s Claude, whether it’s Codex—“I want to build on AWS. Here are my credentials, here’s my stuff. This is how I want to deploy it.” And then the agents are great and they go do it. There are use cases, though, where you’re not already on a cloud, you’re just trying something. You say, “Deploy.” And frankly, a lot of times that’ll go to some of our partners that have an easier-to-use layer on top, which we love, by the way. And we love our partners on those fronts, too.

But I do think that there are some things that we’re doing where, if you’re brand new, you don’t already have an AWS account and you haven’t already set up your IAM, today—or we’ll say a couple of months ago—it was much harder to start up an account, right? You had to—it’s been like this for 20 years—you had to define your VPCs, you had to define your IAM roles, all these kinds of things, which are actually super important once you get to be large. What we hear from customers is that it was a hard trade-off because they know that they’re going to want those things months or years down the line, but for now they just kind of want to use the services and not worry about that.

What we’ve actually launched is, if you go create a new AWS account today—and we’re slowly rolling this out; I don’t know if it’s actually fully rolled out yet—you don’t have to do those things. You don’t have to give it a credit card. You can sign up with your Gmail account. All of those things are handled by default behind the scenes, and within less than 30 seconds, you’re up and running and can be operating in a full AWS account, which is much more what the agents want to be able to do for those types of systems, because they don’t want to have to go through all that setup of your VPCs, your IAM, and pieces like that. So there are some things where we’re adding some ease of use to that, which we’re quite excited about, and I’ve seen some really positive feedback from customers on. We’ll keep doing more things like that over time.

One more thing, which is the great part about that, though, is it’s not like a simplistic account and then you have to migrate. If you basically say, “Okay, now I actually do want to develop an organization. I want to go and kind of fine-tune some of the things,” you can easily come—you’re already in a real AWS account—and you can actually then go and do all of those things later when you need them. So there’s no migration or move later, and that’s part of the hard work that we really think about: how do we make it not a choice for the customer, but an easy on-ramp into the depth of features that we know customers and startups are going to want when they start scaling?

Raghu Raghuram

Yeah. Got it. What have been some of the hardest things to accommodate as agents have taken over? I mean, you made a name—this is the way you guys became the phenomenon that you became—by serving developers, right? And then infrastructure teams. Now developers are being substituted by agents, and pretty soon infrastructure teams are being substituted by agents. What have been some of the hardest things for you, as the largest service provider in the world, to handle in that transition?

Matt Garman

Well, I think, like I said, much of our infrastructure was pretty well set up to handle the scale, which is good. I think there have been a couple of things that are interesting—interesting paradigm shifts—where you could argue that some of our systems are, I wouldn’t say overengineered, but, as an example, many agents want to create a database, do a little bit of work, and then have the database go away. You really need that database to have five nines of durability? Exactly, right. And so there are some things that we rethink there, where, when you’re creating an Aurora database as your production database, you do want five nines of data durability—you need durability, you want availability, you want all of those things. For the agent use case, that’s arguably—or maybe not arguably—overengineered for what we need.

We don’t really want to have a nondurable option that’s going to cause problems either, because you never quite know if the database that’s created wants to stay around for a long time or a short amount of time. So we’re trying to do the hard work to think about how you accomplish both of those things, where you can create it quickly and throw it away and you don’t really waste a lot of resources, but if you do want it, it can be durable and stay for a long time. It can actually grow into a big production database.

Those are some trade-offs that we think about actively as we think about how the more traditional, “This is going to be my production system” capabilities match with some of the more transient nature of the infrastructure that agents want to use. That’s one, I think, but there’s a bunch. I think the scale, speed, and latency of creation of durable resources is another one that’s interesting.

I think the other one that we actively think about, too, is whether there are new building blocks that agents are going to want that we just didn’t really need before. Exactly. And so compute sandboxes, gateways, agent permissions versus people or service-role permissions—a lot of those are things that we have built and are building and thinking actively about, because they are just brand-new building blocks.

It’s not using the existing building blocks differently, but brand-new ones, where pretty clearly you want different permissions for agents. You don’t just want to give it Raghu’s permissions and let it go do whatever you can do. You actually want very time-boxed, short-term permissions to just go do a task. You may not want to give permissions to a whole tool at all. You may want to get very fine-grained permissions for what an agent is able to do from a sandbox, right? You actually want a sandbox that—

Raghu Raghuram

Virtualization is having its act together.

Matt Garman

That’s right. Exactly. And they’re different. It’s similar—it’s the same idea—but how do you have that lightweight? Unfortunately, we’ve done a lot of work with Firecracker, our kind of microVMs. In fact, a lot of the sandbox company startups use Firecracker. They all use Firecracker, which we invented 10 years ago, maybe, or something like that.

It wasn’t purpose-built for agents, but it’s actually quite good for agents because you can spin them up really rapidly. They have a great security boundary, and you don’t have a lot of that virtualization overhead that you have from a traditional large VM. So I think those are some new building blocks that we’re thinking about, and more emerge every single day.

Raghu Raghuram

Yeah. So I’ve rattled on enough about agents. I’ll come back to it later, but it’s a fascinating topic. Matt Garman

Super cool.

Raghu Raghuram

But let me switch gears a little bit and ask about something that every one of our startups faces and that we get asked about most frequently, which is: how can we get GPUs?

I mean, obviously, you guys have massive GPU farms, and you’re increasing them every day, but the top frontier labs gobble up every available GPU. More power to them. How do you deal with that? At the same time, you want to promote newer companies that are going to become the enterprises of tomorrow, like you said. How are you thinking about balancing your capacity needs and these small companies that don’t have a lot of credit and whatnot?

Matt Garman

Well, yeah, it’s a great question. A couple of these things make it more challenging. One is that the capex expense needed to deploy all of the compute that everyone needs right now is massive, right? And so—

Raghu Raghuram

Your capex is the same as what, $200 billion or something?

Matt Garman

$220 billion for 2026. At one point I saw that it’s a pretty—I mean, that is a larger expense than we’ve ever had, maybe any company has ever had, in a single year. We don’t anticipate slowing down anytime soon because the demand is just massive. At some point, you’re limited by how fast you can build data centers, how fast you can deploy capital, and how fast you can get memory and chips and all of those kinds of things.

All of those are, at different points, constraints that we're dealing with—whether it's power, data centers, capital, memory, or chips. They're all constraints at various times. Construction people to build buildings are at a premium today, so we work really hard to make sure all of those things happen.

Then, with the capacity that we're able to deploy—which is still a huge amount, and not enough—we think very intentionally about our allocation strategy. We're great partners with the large frontier labs—Anthropic, OpenAI, Meta, and other large customers—and those folks are really good customers of ours.

Raghu Raghuram

Yep.

Matt Garman

We want to make sure that we invest in them. We also have large enterprise customers, whether it's Salesforce, JPMorgan Chase, or other large companies, that have demand for fewer GPUs or accelerators. Sometimes they want Trainium; sometimes they want NVIDIA GPUs. We want to make sure that we can support them as well, and we're very intentional about making sure that we have capacity for startups.

What we do is allocate capacity. We basically say, “Okay, we could—you’re right—we could sell every single GPU or AI accelerator we had to probably just the big frontier labs and call it a day.”

We choose not to do that because we actually want to keep growing the full ecosystem. They get a large number, but we want to keep supporting a broader set of customers because we actually think that having the whole ecosystem will be healthier for us. There's some diversification, but it's also that we know these are going to be big companies over time, so we try to support them.

I saw recently that we say yes, in some way, shape, or form, to something like 60% of the requests we eventually get. Sometimes it's a little bit later; sometimes it's in a different region; sometimes it's a slightly different configuration than the customers are looking for. But we really try to lean in and allocate as much as we can.

Every single startup wants more, so we're working hard to make sure we have capacity for all of them. But it's hard. As a lot of people say, it's a good problem to have, but it's a problem nonetheless.

We continue to look at it. We recently announced that we're going to be buying 2 million NVIDIA GPUs over the next couple of years. We're landing a massive amount of capacity.

Raghu Raghuram

And 2 million?

Matt Garman

2 million. So it's a lot, and it's over the next couple of years. Who knows if that's enough? At some point, again, we're limited by other components as well. But we're very intentional about keeping capacity available for startups. I know it's painful not to have enough, but we keep pushing.

Raghu Raghuram

When did your capex cross—I mean, pure AWS? Amazon was a bigger company.

Matt Garman

Yeah.

Raghu Raghuram

Cross, let's call it, even $1 billion or $10 billion a year?

Matt Garman

Oh, I don't know. I'd have to go back and look. I'm not sure about that, but it definitely scaled up over the last couple of years in a pretty meaningful way.

The AI buildout has definitely ramped our capex spending. Over the last 3 or 4 years, our capex has accelerated pretty meaningfully. I don't know when we crossed $1 billion, but given that we're at $220 billion now, we've been spending capex for a while. The company has been really good about funding that, and obviously the AWS business is a good one that we like to invest in.

Raghu Raghuram

Yeah, yeah, of course. For the longest time, we were investing ahead of where the demand was. Matt Garman

One of the most painful things is that, with the real ramp-up of GPUs, a lot of the elasticity has unfortunately gone away. Hopefully, in our core compute and storage, that elasticity is still there.

We've been spending for quite a bit of time, and we feel really good about the spend that we're making now. I get lots of questions about how we feel about that spend and whether we're nervous about a bubble or other things like that.

Because of the position we have, we take a diversified approach. Not all of our capacity is bundled up in one customer. You go to some of these neoclouds or some of the other providers out there, and you'll sometimes see concentrations of 30%, 40%, 50%, or 60% with 1 or 2 customers. We're nowhere near that. Obviously, we're in the single-digit percentages at the highest, and usually it's less than that.

For one, I think we have a lot less risk with any particular customer. But also, because we have that rich set of services, AWS is where people are really coming to launch their production workloads. The majority of our usage today is actually either core compute and storage or inference, which is part of that application. Those are the workloads that I think just aren't going to go away.

We're seeing enterprises get positive ROI. You go talk to the customers and say, “At the capability today and the cost today, are you seeing positive returns for your business?” Almost to a person, they'll say, “Oh, yeah.”

You're like, well, that's not going to go away. There's no bubble in which they stop spending on that. The VC model is that every billion-dollar startup isn't going to make it. No, they won't. But that's kind of the game, and it's been true for 50 years. They haven't always been billion-dollar startups; the startup numbers have changed, but the principles are the same, right? You bet on 10, and 1 makes it and pays for the other 10, or whatever the percentage is. Hopefully it's higher than 1.

You saw that with the internet, where there was a bubble and a bunch of internet companies didn't make it, and the internet is still a thing. A lot of the companies that had durable businesses—Google, Amazon, and others like that—did pretty well. So we feel really good about that investment and the continued investment going forward.

Raghu Raghuram

You guys have a view of the demand that's unparalleled, right? You're seeing across the globe and across every segment: enterprise, the big labs, AI-native companies, and so on and so forth. If anybody should be able to call it, you should be able to call it first.

Matt Garman

I would hope so. Yeah, that's our plan. Honestly, we spend a lot of time thinking about it. We're very intentional about how we spend our shareholder capital, and we think that we're making great investments. We have a lot of good protections around how we intentionally invest that money.

We're very bullish on it. Andy has been public about saying this: The potential for AWS is really, really large. Over the next decade, the potential is there, and anytime you have an opportunity that's that big, you want to invest to go after it.

Raghu Raghuram

As the scale of these numbers gets larger and larger, right?

Matt Garman

Yeah.

Raghu Raghuram

Has your planning and process dramatically changed? You're not writing billion-dollar checks; you're writing $50 billion checks or $20 billion checks, or whatever it is.

Matt Garman

Yeah. Yes and no. A lot of times, we're still very bullish about the investments and lean forward, but the process has changed. There are a lot of things that have completely changed.

How do you even estimate 2028 demand? There are things that we have to think about now that we just never had to think about. If you go back 15 years, if we needed more power, we asked the power company to give us another couple of megawatts, whatever it was—tens of megawatts—and they would just give it to us because 10 megawatts wasn't that much. That was plenty for us to keep growing.

Now we have to bring our own power. We pay for power projects and renewable projects. We're one of the biggest renewable power purchasers each year, and we've been doing that for the last 10 years. We're regularly bringing on new solar projects and new nuclear projects.

Raghu Raghuram

So you're talking about behind the meter, or are you talking about working with the operator?

Matt Garman

We'll do both. Often, with these power projects that we bring on, we'll pay the capital and pay for the project, and then it will go into the grid and we'll get credit for that. We'll bring those on.

Sometimes we'll do behind the meter, too. It's a mix. At the scale that we're doing, you have to think about all of those things. That's planning where you're thinking 20 years out about how you're going to think about power, how you're going to think about transmission, and how you're going to think about that capital. That's stuff we'd never had to think about before.

That’s planning we just never had to do. The other thing is, when we used to think about the server demand we needed, we’d have multiple-quarter demand forecasts and talk with our suppliers. Now we work multiple years out, just because the scale is so large that we have to think about what we need for 2026, 2027, and 2028.

But it’s also one of the values that we bring to customers, right? That’s something that customers legitimately can’t do themselves. They’re not going to do power, and they’re not going to plan their memory footprint in 2028. They just can’t do that. That’s one of the real values that we bring to our broad set of customers: it’s a whole set of things they don’t have to worry about, and that we spend a huge amount of time thinking about.

Raghu Raghuram

Yeah. You guys are generally, I think for the record, the largest buyers of practically every component of a server. Correct?

Matt Garman

I don’t know that. I’m sure we’re one of the bigger purchasers of components out there, for sure. Who knows about everyone? It depends on how you think about them and how you measure it.

Raghu Raghuram

So where I was leading to: where do you see the constraints being most severe, let’s call it 2027 and 2028, and where do you see the constraints easing up?

Matt Garman

It’s funny. I’ll answer this in a roundabout way. I remember that in undergrad—we actually read this a long time ago, and I never thought it would be a useful book—but we read “The Goal.” Have you read “The Goal”? It turns out there’s never one constraint; there’s always just the latest constraint. You have to think about all of them, and as soon as you hit one, there’s another one.

So what is the constraint that’s going to happen in 2028? I actually don’t know. I think there will be one. Every month for us, it’s: Do you have enough power? As soon as power is no longer the constraint, it might be memory, TSMC capacity, HBM, networking components, or a blip somewhere in the supply chain, like connectors. At some point, you have to think about all of those pieces.

It’s not just where they happen; that matters, too. We may have a ton of power in Indonesia, but not enough in Germany. You think about where in the world you want that capacity, too, because it turns out that not everything is totally fungible. Some is, and some isn’t.

It might be disk drives, it might be SSDs, or it might be something else. We think about every single component, and we have tens or hundreds of thousands of components that we track and think about. Some we rely on our suppliers to manage, and some we directly manage. We saw this problem coming probably a decade ago and really started not just thinking about how many servers we needed to track, but thinking all the way through the supply chain—four or five tiers down—to identify the component that could cause an issue for us and make sure that we had guaranteed supply of it.

Raghu Raghuram

If you remember—

Matt Garman

I don’t remember when this was—over a decade ago, when there were floods in Thailand and no one had disk drives anymore.

Raghu Raghuram

There were disk-drive crises, and then there was a memory crisis.

Matt Garman

Exactly. I think we think through all of those things. We also think about where there’s diversification in manufacturing—all of those pieces we try to work through. We’re never going to be perfect at it, but there’s always a different supply constraint.

Raghu Raghuram

Yeah. Obviously, there’s a lot of wide-ranging debate about data centers.

Matt Garman

Mhm.

Raghu Raghuram

It’s clear where folks like us stand. But do you think, as an industry, we haven’t done a good job of explaining why data centers are good for America and generally the world? What’s the internal discussion amongst Andy’s team about how to deal with this?

Matt Garman

Well, look, I think you’ll hear more from us on this, and I agree. I think we need to be more vocal and more upfront. We actually do a ton that’s really beneficial, both for the communities we operate in and more broadly. We think a ton about how to bring renewable energy to these data centers and how to be water-positive.

Actually, our data centers use a really, really small amount of water. We mostly use free-air cooling. We think about how to be great participants in the communities where we are and how to bring high-paying jobs to the communities we operate in.

Not all data center operators do that. There are some well-chronicled examples of others out there that aren’t great at it. They don’t really pay attention to regulations, they think the rules don’t apply, or they launch really quickly without thinking about those things. I think that causes problems for the whole industry because everybody gets lumped into that.

I think you’re right. We’re vocally self-critical. We need to be more vocal about the benefits that we do bring and think about additional ways that we can help communities understand those benefits, both for the services they use and more broadly. If you go to a community and say, “Do you not want to use Netflix?” they’ll say, “No, no. I still want Netflix.” It’s important for us to highlight the benefits that we bring.

I recently saw a report where, in one of the communities that we operate in, everybody in that county pays $5,000 a year less in taxes because of the taxes that we bring. We don’t tell them. They don’t even know it; it’s just invisible to them. I think we need to be more clear about those benefits, because if you told the communities, “By the way, your tax bill is $5,000 less than it would otherwise be if we weren’t here,” they might have a slightly different thought about the building that’s over there.

Not everyone does that.

Raghu Raghuram

Presumably, there’s more transparency needed.

Matt Garman

I think much of the data center community, not just us, is actually made up of pretty good actors. There are just a few that aren’t, and I think they’ve caused some of the angst recently. We need to do a better job of highlighting who’s being a good citizen and who’s not.

Raghu Raghuram

Before we leave the hardware topic, I want to touch upon training and your whole history with building your own chips. We were one of your first partners using Nitro a long time ago.

Matt Garman

Mhm.

Raghu Raghuram

Since then, with Graviton, you’ve made tremendous progress. What was the thinking that led to saying, “Look, we’re going to do our own thing”? How has that progress been, and where are you?

Matt Garman

It’s actually a fascinating story, and I think it’s a great example of how AWS will innovate and iterate over time and continue to think bigger about what we can do, but prove our way there.

Take, as an example, this was probably about 13 or 14 years ago. We were seeing that there was a pretty significant virtualization tax on the overall number of resources. Our customers were telling us, “I want bare-metal performance,” and they were comparing it to having all the resources of a server.

The first thing we did was take a network offload card and move all of our network virtualization onto it, so that network virtualization got closer to bare-metal performance. Back then, it wasn’t quite bare-metal, but it was closer.

Then we got really excited about that and said, “What if we could move storage virtualization off as well?” None of the network offload cards could do that. Then we found this one company that had some Arm cores on an offload card. They were doing it for other reasons—I can’t remember their original purpose—but we thought, “Could you use those to do storage virtualization and some of these other functions?” They said, “Maybe.”

We really iterated with them. This was the product, and this was the team, and I just loved that team: really innovative, really mission-driven, and really wanting to solve problems.

We acquired them and said, “Look, could you build us a slightly bigger card that could take all the network virtualization off and basically give us a bare-metal server that had no virtualization on it—no VM virtualization? Everything was through APIs on the card.” We had the view that performance would be much better and resource utilization would be better.

Security isolation would be much, much better, and we could then legitimately tell people, “We have no access to any of your VMs that are running there.” This has been a huge benefit for us for the last decade where, frankly, we’ve been leading and others have been kind of slow to do this, because this is not a generalized thing that people can do.

We got to that, and we basically said, “Look, we’re making a lot of progress here. What if we take this?” There were a bunch of Arm cores on this offload card, and we said, “What if we turn that into a server?” We did that first with Graviton. It was a very underpowered, very small server that we launched.

Customers were excited. They said, “I’d love to have an Arm server. This is super interesting.” So we went down that path. With Graviton, part of what we did was look and see that the slope of Arm cores getting faster was increasing, and we saw the power utilization. You could see where the intercept was going to happen—where this architecture was going to be really good for workloads—and they just needed somebody to drive the ecosystem and get some of the pieces in place. So we did that with Graviton, and Graviton has been a runaway hit at this point.

Raghu Raghuram

Have you been public about what percent of your fleet is Graviton?

Matt Garman

We ship more Graviton chips every year than any other type, so it’s very popular. They’re 20% cheaper with 20% better performance, and have been like that for the last 5 or 6 years. That’s an easy value proposition.

I think the vast majority—something like 90-plus percent—of our top 100 customers all use Graviton in some way, shape, or form across their fleets. That’s been a huge win for us and for customers. It’s the single biggest, easiest way that customers lower their bill: They move to Graviton.

We’ve had examples where people have moved their whole fleets and cut the number of servers they had in half.

Raghu Raghuram

The performance is so much better.

Matt Garman

It’s amazing. Half the number of servers, and each server costs less. It’s a big win.

About 5 or 6 years ago, we saw the rise of AI compute happening. We didn’t nearly expect what it is today, but we still saw it was going to be a big, material mover. So we went in and built our first chips, Inferentia and Trainium, and we’re now in market with the 3rd generation, Trainium 3.

We’ve seen fantastic results. We’re sold out for capacity through probably toward the end of next year or something like that. We’re trying to figure out how we can save some capacity and get startups to be able to use some of it, because we see great results.

The majority of traffic on Bedrock runs on Trainium. We have great deals with both Anthropic and OpenAI to build on top of Trainium, as well as a set of smaller startups. I think we have half a dozen to a dozen startups that are building on top of Trainium now, too.

Raghu Raghuram

The name suggests it’s a training chip.

Matt Garman

Yeah, we’re bad.

Raghu Raghuram

But everybody’s using it for inference, too. So where is the advantage architecturally, and where is it going?

Matt Garman

It’s a good point. Look, we’re vocally self-critical: We’re terrible at naming. It’s not our strength. Originally, we had a chip called Inferentia for inference and Trainium for training. As the models got bigger and bigger, it turns out you actually want to run the inference on these really large systems. It turns out that Trainium is maybe the best inference chip on the market right now, from a pure performance and cost-performance point of view.

Raghu Raghuram

Is it better memory bandwidth? Where is the advantage?

Matt Garman

The architecture is just a little bit different from others, and it’s much cheaper. From a cost-performance perspective and an absolute-performance perspective, Trainium is great.

We use it a ton for inference, and it drives much of the Bedrock inference today. It’s also a good training chip. A lot of our broad set of customers’ usage is not in training models, but in using models, so that’s where a lot of people get to use it under the covers. That’s where we’re excited about it.

A lot of the big customers are interested in it for training clusters as well, particularly as you get to Trainium 3 and 4. We’ve announced Trainium 4, but we haven’t launched it yet. Folks have looked at that architecture and said, “That’s the future of where I want my training clusters to be as well.” We’re quite excited about where that goes for these really broad-scale training clusters.

6. Where enterprises are stuck on agents

But, yeah, it’s built.

Raghu Raghuram

Yeah. Now let’s get back to talking about agents, but from the perspective of large enterprises or large and medium-sized enterprises. Where are they in their adoption? What sort of benefits are you seeing them reap already, and what is the roadmap for them, as far as you can tell from your vantage point?

Matt Garman

It’s a really good question, and it’s one that we’ve spent a lot of time thinking about. When I talk to customers out there today, they view themselves as getting a lot of value out of what they’ve done. I would say the agents that most enterprises have built are relatively simple and straightforward, and they’re starting to think about how to make them autonomous in a safe way.

Raghu Raghuram

Mostly nonautonomous, right?

Matt Garman

They’re still people in the loop, if you will. Customers are still getting lots of value out of that today. They’re really thinking about how to make these systems autonomous in a safe way.

I think there are 2 things that hold customers back today from continuing to scale. It’s already a pretty big business today, but I think it has a massive opportunity to really change every single customer and every single workflow.

Number 1 is just how to think about it. What we originally saw was that enterprises had a workflow and were saying, “Great, I would have an agent go do the same workflow.” What we encourage them to do is not just replicate it: Bob does steps 1, 2, 3, 4, and 5, so the agent is going to do steps 1, 2, 3, 4, and 5, and then Bob’s going to check it at the end. That’s not really the model you want.

You want to step back and say, “If I want to accomplish something, how can an agent do it differently?” It can do it in a massively parallelized way. It can try 50 different things and get to the right outcome. How do you help it get to that right outcome? You rethink how a computer would solve a problem versus how a human would solve a problem.

One of the things is helping customers understand how to think about that and really have that blank slate, because that’s where you really get value. It’s not just replicating what you’re doing today, but thinking from a greenfield approach about how you solve a problem completely differently. I’m sure that’s how many of your startups are thinking about this, too: How do you help customers greenfield-solve a problem, not replicate the thing that happens today?

Raghu Raghuram

So that’s number 1.

Matt Garman

People running fleets of agents, swarms, whatever you want to call them. Yeah—

Raghu Raghuram

Very common these days.

Matt Garman

And you just want to think about it. Enterprises are not as forward-leaning. Again, this is one where you learn from the startups and try to apply that to an enterprise world. An insurance company is not necessarily as forward-leaning, but they would love to figure out how they can have a better approval workflow or something like that.

That’s number 1. The second one is how you turn those into fully autonomous workflows and how you actually trust the agents. We’re spending a lot of time thinking about how to build services to help enterprises feel like their systems are secure and that they can trust an agent to make a decision, that they can have the right guardrails, that it can have the right permissions on their data, that it’s not going to delete production systems, and that it’s not going to make tragic mistakes.

Right now, I think that nervousness is holding people back—maybe appropriately, by the way—from just saying, “Okay, go nuts.” You don’t actually want an agent to go crazy and accidentally delete a production database. That’s going to be pretty bad.

We’re actively working through this with customers. How do we both help them architect and, frankly, invent new technologies and capabilities that are going to help them solve that problem? That’s one of the areas where I think we’ll continue to innovate, and we’ll get there. I think we have some really good ideas and some good technologies brewing that can really help.

Raghu Raghuram

So, are enterprises learning how to do eval systems and so on and so forth to keep the agents? They need help, honestly. Both with evals—how do you have a constant loop of testing? How do you think about goal-seeking in a reasonable way? How do you have your data labeled in such a way that it actually makes sense, so the eval can approximate what you're going to be doing in production? How do you measure in production and back-test it so you're not seeing drift?

Matt Garman

All of those things are problems that enterprises don't know how to solve today. I don't know if anyone really is great at solving these today. It's why you've seen so many FDE teams spin up, and AWS and our partners are really leaning into the FDE motion to go and help. This is the single biggest area where customers need help.

And when we think about how to do that, my view is that we want to train our customers to be able to do this themselves. This is not the traditional notion where I want to have a people-driven business that goes on forever, where you just keep paying consultants over and over and over again. Our view on how FDEs should work is that we want to go into a customer who's ready to accept owning this when we're done, and in 45 days do work where we can teach them how to make an eval and teach them how to get their data in a labeled way.

We do the work alongside them, and then, at the end of 45 days, we leave and that customer is good and ready to go and trained up. That's what our customers tell us they want. They don't want to be beholden to an external workforce for the next 5 years.

Raghu Raghuram

But they need help today.

Matt Garman

Yeah.

Raghu Raghuram

You made a massive investment in FDEs.

Matt Garman

FDEs.

7. AI risk, Hugging Face, and Continuum

Raghu Raghuram

So, taking that even one step further, some of your industry peers have said, “Look, you can't have all of your data going into a big frontier model. What enterprises should really do is take an open-source model and then post-train on your own data, workflows, preferences, and whatnot.”

Where do you stand on that? Are you seeing customers actually trying to do that, or how do you think about that?

Matt Garman

It's a great—the first point, I wholeheartedly agree on that first point: enterprise data is their most valuable asset. From the very beginning, that's why we built Bedrock like we did. We have a guarantee that your data never leaves your VPC. If you're running inside of Bedrock, your data doesn't go back to the model provider. They never see your prompts. That stays inside of your own trusted environment.

That is why enterprises prefer to run on top of Bedrock, and it's why you see that business growing massively. Every month, we see that business exploding, and it's why you see OpenAI workloads migrating to Bedrock. It's why you see Anthropic really growing rapidly.

Whether you're using open models or closed frontier models, I think Bedrock is a great solution. Our customers told us this, by the way. If you remember, 3 years ago I got a lot of heat.

Raghu Raghuram

Speaking of bad names, that's a good name, though.

8. The Graviton and Trainium bet

Matt Garman

Yeah, Bedrock is a good name. That's good. But we got a lot of heat, actually, for being slow to the AI world because we built the foundations of this. We said, “Look, we're not just going to rush out a service. We really want to think about how we make sure that we protect our customers' data and build a service that we think is going to be durable for the use cases that we knew about.”

And if you remember, we got a lot of heat, and we said, “Look, we're going to go build the right thing.” Now, as people move from proof of concepts to production, the vast majority of them are landing in AWS on Bedrock. One of the reasons is because of this. It's also because of the set of services that we have.

We offer open models. We offer proprietary models. We offer a whole set of capabilities around those—AgentCore. We build these building blocks so it's easier to build agents with any of the models that you want, whether they're in Bedrock or out of Bedrock, for that matter. You can use Gemini or other things for it.

I think that's a differentiating piece for us, and it's a super important thing to think about because having that data go back into the model provider is a dangerous thing. You talk about open-weights models, though. I do think that there's an area that I'm excited about. We're really ramping up our support of open-weights models and trying to build a good environment.

Frankly, this is where today I think—and I think it's true of a lot of customers—people believe that they have meaningful proprietary data that, if they could mix it in and do some post-training or fine-tuning to an open-weights model, they could distill down and actually get a better-performing model at a lower price.

There are a bunch of pieces in here that have to work out well. They actually have to have a good eval to prove that that's true. Most people are doing that in SageMaker today. I think there's more that we can do to make that easier. But if you go look at where people are doing that, they actually do it in SageMaker on AWS today, and they host the inference via SageMaker.

Raghu Raghuram

SageMaker is getting a new lease on life as a—

Matt Garman

I mean, it is. It was always kind of a model-building platform, right? And now, if you think about what enterprises are doing in this world, that's what they're doing: they're basically building their own custom models. SageMaker is a great place for doing that.

I think there are some things that we need to keep building on to make that easier and easier to do and to test across different open-weights models and things like that. But that's an evolving space. I think it's a super interesting one, and it's one we want to make sure that we have all the right things for customers to be able to do if they have the right data and expertise to actually go down that path.

Raghu Raghuram

And with all the debate around AI existential risk, this, that, and the other, and the security vulnerabilities and so on and so forth, and the Hugging Face attack, what do CEOs ask you about all these things?

Matt Garman

Yeah, there's a bunch. They mostly want to say—and it goes back to this—“When I launch agents, how can I trust that they're going to do what I want them to do?”

We're heavily investing in building capabilities that allow people to deploy agents safely into their environment and think about those controls. Some of those are: How do you make sure the agents have the right permissions? How do they have the right sandboxing? How do you make sure that you have the right set of guardrails? How do you intentionally think about what you want the agents to do and not do? Is there a human in the loop or not?

We spend a lot of time with our customers thinking about safe agent deployment and how to get better over time, and what other capabilities we need to build to help people deploy agents safely into their environment. So, there's a lot there.

The other angle that a lot of people are worried about is whether some of these really powerful models are going to be attack surfaces and be misused. I have a view on that: yes, that is a real risk to customer environments, but it's also a real opportunity.

We recently launched a service called Continuum that uses these powerful models to help customers secure their environment. We'll look across the environment and look for vulnerabilities. We'll use some of these powerful models to help customers find vulnerabilities they haven't found before and, most importantly, prioritize which ones, because we know context about their environment: how it's set up, where their permissions are, and where they may have compensating controls that make it harder or easier to exploit them.

Continuum is incredibly popular with customers. We're really bullish about what's possible from AI to help with AI-powered security, because at some point customers are going to need security at machine speed, not at human speed—not an alert that someone goes in and looks at.

We're running fast to help build that for customers and help them protect their environments, and I'm very excited about what the Continuum team is building on that front, too.

Raghu Raghuram

So, within AWS itself—

Matt Garman

Mm-hmm.

Raghu Raghuram

What's the state of adoption and usage of agents broadly?

Matt Garman

Well, Continuum is basically us trying to expose what we do internally. We use AI extensively for our own security. We use AI extensively for our own software development. We use agents—actually, one of the things that's really cool is that we use agents across our entire business.

We rolled out Amazon Q to every single Amazon employee. Now I see HR teams building agents to help drive what used to take teams of people weeks to do, which a single person can now do in a couple of hours, such as team planning and resource management.

I have finance teams that are building agents to pull tax rules from everywhere and ensure compliance on a bunch of different pieces. It’s super cool to see that things that used to be blocked by software developers—actually, the line-of-business folks are able to go and unblock themselves and innovate more quickly.

Amazon Q has been an enormous boon, and that has grown like wildfire. We see customers like small startups using it all the way to the largest enterprises in the world rolling it out to their entire customer base to get the benefit of being able to access all of their enterprise data and easily apply agents and capabilities to help accelerate their jobs.

Raghu Raghuram

And so, as I said, we use it across everything from software development to security to driving HR policies. You guys are notorious for measuring everything about your operation. Where have you seen the biggest gains?

Matt Garman

Yeah, I mean, obviously, you know the answer: software development. That’s the real answer.

Raghu Raghuram

The derivatives of that.

Matt Garman

But yeah, I mean, the speed of software development—and really, product development as a whole, not just coding—is absolutely the case. The pace at which we’re deploying new products is massively different than it has been in the past.

I think you can see this. AWS has always been known for rolling out features really quickly, and we’ve seen a turbo boost on that in the last year or so as we’ve developed what we call frontier teams, as they think about agentic development as opposed to traditional development. It’s not code completion; it really is agent-first. The agents write all of the code. You’re just managing a team of agents and driving that.

It’s been fun to see the pace at which innovations for customers have been happening. It has to be, because our customers out there have an almost insatiable appetite for new capabilities, and that’s what we’ve got to do.

Raghu Raghuram

Yeah. So, organizationally, do you have any insights on how organizations should change?

Matt Garman

I don’t know. I’ll say that—

Raghu Raghuram

Agents manage people, people manage agents?

Matt Garman

There are going to be lots of people for a long period of time. I do think organizations will change. I don’t know the magic answer yet, but we’re actively thinking about it.

Raghu Raghuram

Inside, you’re just trying various experiments, or what?

Matt Garman

We’re trying experiments. We’re thinking about pods. As you think about it, here’s one example: In a product organization, you used to have a team that would own a particular capability for a long time, and you might have 10 people working on that thing. Well, today, you can innovate so rapidly that it doesn’t have to be 10 people. It can be 3 to 4 people, and they build something so fast that you actually want to move them to different projects and problems.

Thinking about how you both operate and maintain the things that you’ve built while being agile and flexible enough to move around in an organization as big as AWS is—that’s an active area that we’re experimenting with and playing with. It’s fun, and it’s enabling for our employees. They actually love it because they can build faster and do more. But there’s real work there.

Raghu Raghuram

Yeah, it’s a fascinating time. Thank you very much for your time. We could be talking about this for hours together, but thanks for all your insights.

Matt Garman

Yeah, thank you for having me. We love having all your companies as customers, and we love learning from them. I appreciate you having me here.

Raghu Raghuram

Yeah, we’ll keep sending them your way.

Matt Garman

Excellent. Thank you.

Raghu Raghuram

Thanks.