[BidClub_]
SemiAnalysis · · 47 min

AWS Trainium: How Amazon Built Their Own AI Chips | Researcher Conversations at GTC

Jordan NanosDan NishballKang Wen CheangZane Fong

YouTube
TL;DR
  • AWS says its cloud fleet contains about 2 million GPUs and plans to add another 1 million during 2026. That increment equals “50% of the entire footprint we’ve built in the last 15 years,” which AWS presents as evidence of customer demand and its ability to scale quickly.
  • AWS says it looks forward to bringing next-generation NVIDIA Rubin GPUs or systems this year, extending a partnership approaching 15 years. Yet GPU instances alone do not deliver production AI: customers “hit a wall” around data pipelines, cost controls, security and compliance.
  • Trainium is AWS’s cost-and-capacity option amid constrained chips and power. AWS claims 30–40% better price-performance than alternative accelerators available in the cloud, arguing that broader hardware choice helps customers access capacity and scale.
  • Trainium 3 promises two to three times better performance than Trainium 2, with AWS planning more than 1 million chips in 2026 and more next year. The speaker cites Anthropic in connection with heavy training and inference, but the wording does not clearly establish Trainium use; he says OpenAI labs probably are and will be using Trainium for large-scale training and industry inference, while startups are showing interest.
  • The Cerebras partnership targets inference economics through disaggregated prefill-and-decode. AWS says its primary goal is lowering customer costs, especially “dollar per token,” which customers identify as necessary for broader GenAI deployment.
Digest · the substance, structured for research

1. AWS plans to add half its historic GPU footprint in one year

  • The AWS executive says the NVIDIA partnership dates back almost 15 years, when AWS was among the first cloud providers to offer GPUs, before GenAI was really a thing.
  • AWS reports about 2 million GPUs today and another 1 million planned for 2026—“50% of the entire footprint we’ve built in the last 15 years.”
  • AWS says it looks forward to bringing next-generation Rubin GPUs or Rubin systems this year (2026); no more precise availability date is offered.

2. Production AI requires far more than GPU instances

  • AWS’s product-marketing lead says demos take “days or even hours,” but customers “hit a wall” when moving AI into production.
  • It is not just about deploying applications on GPU instances: the surrounding stack includes data pipelines, cost control, security and compliance.
  • AWS points to 20 years of operating infrastructure at massive scale and says it offers the largest, broadest and deepest set of services for taking AI applications from proof of concept to production.

3. Trainium addresses both cost and constrained capacity

  • Custom silicon follows AWS’s strategic aim of lowering infrastructure costs and widening hardware choice; Graviton began that effort eight or nine years ago.
  • AWS says the Trainium line offers 30–40% better price-performance than alternative accelerators available in the cloud.
  • AWS frames that choice as a way to free customers to access alternative infrastructure amid constraints on chips, power and “everything in between,” so capacity need not prevent scaling.

4. Trainium 3 is planned to exceed one million chips in 2026

  • Trainium 3 offers two to three times better performance than Trainium 2; AWS plans to scale it to more than 1 million chips in 2026 and more the following year.
  • The speaker says Anthropic has been “asked heavily for training and inference,” but the wording does not clearly establish that Anthropic is using Trainium. He adds that OpenAI labs probably are—and will be—using it for large-scale training and industry inference workloads.
  • AWS also reports interest from a broader customer base, especially startups, looking to use Trainium.

5. Disaggregated inference is ultimately a token-cost strategy

  • The Cerebras-and-Trainium product supports disaggregated inference using prefill and decode.
  • AWS says the partnership’s primary goal is bringing down customer costs, especially inference costs and “dollar per token.”
  • Customers tell AWS that deploying more generative-AI applications depends on costs coming down.
Speaker 1

We have 2 partners from AWS with us today. We're doing a discussion on your partnership with NVIDIA, your statement on the 1 million-GPU deployment, and a discussion on training, the progress there, and the partnership with Cerebras. Could you elaborate more on that?

Speaker 2

Well, thanks. It's great to be here at GDC 2026, and thanks for having us here. Our partnership with NVIDIA goes back almost 15 years. We were one of the first cloud providers to start offering GPUs, before generative AI was really a thing. It's a long, deep partnership.

1. AWS Expands Its GPU Footprint

If you look at our infrastructure today, with the new GPUs, we offer about 2 million GPUs in the cloud—the largest missile provider to offer 2 million GPUs via cloud. We recently launched a blog saying that we're going to add another 1 million GPUs just this calendar year. Basically, that's 50% of the entire footprint we've built in the last 15 years.

I think that says 2 things. 1 is the huge customer demand we're seeing on AWS for this infrastructure, and 2 is a testament to our ability to scale so fast and in such a short period of time. We deeply value the NVIDIA partnership and look forward to bringing the next generation of Rubin GPUs, or Rubin systems, this year. Our customers are excited to have them.

Speaker 1

What are the challenges to scaling, and how is AWS getting better at that? What are your advantages there? Also, when do you expect Rubin to come online?

2. Scaling AI Into Production

Speaker 3

I can start with that. I lead product marketing for AI infrastructure at AWS. A common problem we're seeing with customers is that it's easy to build a very cool generative AI demo in days or even hours, as you can see at all these trade shows and industry events. But when customers really hit a wall is when they deploy these AI solutions in production.

It's not just about deploying applications on GPU instances; it's also everything around them, like the data pipeline, cost control, security, and compliance. That's really where AWS can help customers, with our 20 years of operating infrastructure at massive scale. We offer the largest, broadest, and deepest set of services that customers need to bring these applications to production, versus doing a proof of concept at a much smaller scale.

Speaker 1

But I think, back to what you were saying, let's go more into the training side. You had the Trainium announcement at re:Invent last year.

That was a Trainium 3 announcement, with a brand-new Scalapack architecture, right? You recently announced a partnership with Cerebras to do this disaggregated prefill-and-decode architecture. Could you elaborate more on that? How are your customers viewing these options, and what applications do you see Trainium bringing to your customers?

3. Why AWS Built Trainium

Speaker 2

One of the core strategic principles for AWS when it started 20 years ago was to bring down costs for customers accessing infrastructure and also give them a wider selection of different hardware platforms. That's why we've been investing in custom silicon, going back to our Graviton generation of CPUs some 8 or 9 years ago. The same thing is why we've invested in Trainium as our AI custom silicon chips.

4. Trainium Wins On Cost And Capacity

We want to bring down costs for our customers and give them a broad selection where capacity is still available, so that it's not an issue for them to scale up. From a cost perspective, I think the Trainium line of products offers 30% to 40% better price performance compared to the alternative accelerators available in the cloud.

From a capacity-availability perspective, it frees customers up to access alternative infrastructure, especially in this environment where there are supply constraints on the chip side, the power side, and everything in between. Customers appreciate and value that, and I think that's why we've been investing heavily in Trainium.

5. Trainium Three Scales Training

Trainium 3 is our latest-generation offering, with 2 to 3 times better performance compared to the previous edition, Trainium 2. We plan to scale it to more than 1 million chips this year, and more next year. Some of the recent customers are opening eyes. Obviously, Anthropic has been asked heavily for training and inference.

We're excited about partnering with AI labs, but at the same time, we're also seeing a lot of interest from a broader set of customers, especially in the startup arena, looking at using Trainium for their workloads.

Speaker 1

Is that for their chip-development workloads?

Speaker 2

Yeah, it's from AI labs.

Speaker 1

AI labs, yeah.

Speaker 2

The OpenAI labs, probably—I mean, they are and will be using it for large-scale training and industry inference workloads. Recently, we announced the Cerebras and Trainium product supporting disaggregated inference. And the team is also similar there: bringing down costs for our customers, especially for inference, and bringing down the dollar per token.

That's something we hear from customers: for us to deploy more generative AI applications, we want to see costs come down. The partnership's primary goal is to try to bring down costs for our customers.

Speaker 1

Thank you. I think—thank you very much for going through the efforts you're making to help the ecosystem build, and for the efforts you're putting into the infrastructure and offering the widest and broadest range of products for our customers.

Speaker 2

Great. Thanks for having us.

Speaker 3

Thanks for having us.