Speaker 1
We have 2 partners from AWS with us today. We're doing a discussion on your partnership with NVIDIA, your statement on the 1 million-GPU deployment, and a discussion on training, the progress there, and the partnership with Cerebras. Could you elaborate more on that?
Speaker 2
Well, thanks. It's great to be here at GDC 2026, and thanks for having us here. Our partnership with NVIDIA goes back almost 15 years. We were one of the first cloud providers to start offering GPUs, before generative AI was really a thing. It's a long, deep partnership.
1. AWS Expands Its GPU Footprint
If you look at our infrastructure today, with the new GPUs, we offer about 2 million GPUs in the cloud—the largest missile provider to offer 2 million GPUs via cloud. We recently launched a blog saying that we're going to add another 1 million GPUs just this calendar year. Basically, that's 50% of the entire footprint we've built in the last 15 years.
I think that says 2 things. 1 is the huge customer demand we're seeing on AWS for this infrastructure, and 2 is a testament to our ability to scale so fast and in such a short period of time. We deeply value the NVIDIA partnership and look forward to bringing the next generation of Rubin GPUs, or Rubin systems, this year. Our customers are excited to have them.
Speaker 1
What are the challenges to scaling, and how is AWS getting better at that? What are your advantages there? Also, when do you expect Rubin to come online?
2. Scaling AI Into Production
Speaker 3
I can start with that. I lead product marketing for AI infrastructure at AWS. A common problem we're seeing with customers is that it's easy to build a very cool generative AI demo in days or even hours, as you can see at all these trade shows and industry events. But when customers really hit a wall is when they deploy these AI solutions in production.
It's not just about deploying applications on GPU instances; it's also everything around them, like the data pipeline, cost control, security, and compliance. That's really where AWS can help customers, with our 20 years of operating infrastructure at massive scale. We offer the largest, broadest, and deepest set of services that customers need to bring these applications to production, versus doing a proof of concept at a much smaller scale.
Speaker 1
But I think, back to what you were saying, let's go more into the training side. You had the Trainium announcement at re:Invent last year.
That was a Trainium 3 announcement, with a brand-new Scalapack architecture, right? You recently announced a partnership with Cerebras to do this disaggregated prefill-and-decode architecture. Could you elaborate more on that? How are your customers viewing these options, and what applications do you see Trainium bringing to your customers?
3. Why AWS Built Trainium
Speaker 2
One of the core strategic principles for AWS when it started 20 years ago was to bring down costs for customers accessing infrastructure and also give them a wider selection of different hardware platforms. That's why we've been investing in custom silicon, going back to our Graviton generation of CPUs some 8 or 9 years ago. The same thing is why we've invested in Trainium as our AI custom silicon chips.
4. Trainium Wins On Cost And Capacity
We want to bring down costs for our customers and give them a broad selection where capacity is still available, so that it's not an issue for them to scale up. From a cost perspective, I think the Trainium line of products offers 30% to 40% better price performance compared to the alternative accelerators available in the cloud.
From a capacity-availability perspective, it frees customers up to access alternative infrastructure, especially in this environment where there are supply constraints on the chip side, the power side, and everything in between. Customers appreciate and value that, and I think that's why we've been investing heavily in Trainium.
5. Trainium Three Scales Training
Trainium 3 is our latest-generation offering, with 2 to 3 times better performance compared to the previous edition, Trainium 2. We plan to scale it to more than 1 million chips this year, and more next year. Some of the recent customers are opening eyes. Obviously, Anthropic has been asked heavily for training and inference.
We're excited about partnering with AI labs, but at the same time, we're also seeing a lot of interest from a broader set of customers, especially in the startup arena, looking at using Trainium for their workloads.
Speaker 1
Is that for their chip-development workloads?
Speaker 2
Yeah, it's from AI labs.
Speaker 1
AI labs, yeah.
Speaker 2
The OpenAI labs, probably—I mean, they are and will be using it for large-scale training and industry inference workloads. Recently, we announced the Cerebras and Trainium product supporting disaggregated inference. And the team is also similar there: bringing down costs for our customers, especially for inference, and bringing down the dollar per token.
That's something we hear from customers: for us to deploy more generative AI applications, we want to see costs come down. The partnership's primary goal is to try to bring down costs for our customers.
Speaker 1
Thank you. I think—thank you very much for going through the efforts you're making to help the ecosystem build, and for the efforts you're putting into the infrastructure and offering the widest and broadest range of products for our customers.
Speaker 2
Great. Thanks for having us.
Speaker 3
Thanks for having us.