[BidClub_]
SemiAnalysis · · 47 分钟

AWS Trainium:亚马逊如何打造自有AI芯片|GTC研究员对谈

Jordan NanosDan NishballKang Wen CheangZane Fong

YouTube
TL;DR
  • AWS称,其云端集群目前约有200万颗GPU,并计划在2026年再增加100万颗。 这部分增量相当于“过去15年搭建的全部规模的50%”,AWS据此将其视为客户需求旺盛、且自身能够快速扩容的证据。
  • AWS称,期待在今年引入下一代 NVIDIA Rubin GPU或系统,延续一段接近15年的合作关系。 但仅有GPU实例并不能交付生产级AI:客户在数据管道、成本控制、安全与合规等环节推进到生产环境时会“撞墙”。
  • 在芯片和电力供给受限之际,Trainium是AWS兼顾成本与算力供给的选项。 AWS称,与云端可用的替代加速器相比,Trainium的性价比高出30–40%;AWS据此认为,更丰富的硬件选择有助于客户获得算力供给并实现扩容。
  • Trainium 3的性能目标是Trainium 2的2至3倍,AWS计划在2026年部署超过100万颗芯片,明年还会继续增加。 发言者提到Anthropic与大规模训练和推理有关,但措辞并未明确证明Anthropic正在使用Trainium;他表示,OpenAI实验室可能已经、且未来也会使用Trainium进行大规模训练和行业推理,初创公司也在表现出兴趣。
  • Cerebras合作瞄准的是通过prefill与decode解耦来改善推理经济性。 AWS称,合作的首要目标是降低客户成本,尤其是“每token成本”(“dollar per token”);客户认为,这是生成式AI进一步普及部署的必要条件。
摘要 · 为研究而整理的核心内容

1. AWS计划一年内新增相当于历史GPU规模一半的部署

  • AWS高管称,AWS与NVIDIA的合作始于近15年前,当时AWS是最早提供GPU的云服务商之一,而生成式AI尚未真正形成气候。
  • AWS称,目前拥有约200万颗GPU,并计划在2026年再增加100万颗——相当于“过去15年搭建的全部规模的50%”。
  • AWS称,期待在今年(2026年)引入下一代Rubin GPU或Rubin系统,但未给出更具体的上线时间。

2. 生产级AI远不止GPU实例

  • AWS负责产品营销的负责人称,演示通常只需“几天甚至几小时”,但客户将AI投入生产环境时会“撞墙”。
  • 问题不只是把应用部署在GPU实例上,外围技术栈还包括数据管道、成本控制、安全与合规。
  • AWS强调自身拥有20年大规模运营基础设施的经验,并称其提供规模最大、覆盖最广且能力最深的一套服务,帮助客户将AI应用从概念验证推进到生产环境。

3. Trainium同时应对成本与供给受限

  • 定制芯片延续了AWS降低基础设施成本、扩大硬件选择的战略目标;Graviton在8或9年前开启了这一努力。
  • AWS称,Trainium产品线相较于云端可用的替代加速器,性价比高出30–40%。
  • AWS将这种选择定位为:在芯片、电力以及“介于两者之间的一切”都受到约束时,帮助客户获得替代基础设施,从而避免算力供给限制扩容。

4. Trainium 3计划在2026年部署超过100万颗芯片

  • Trainium 3的性能是Trainium 2的2至3倍;AWS计划在2026年将其部署规模扩大到超过100万颗芯片,次年还会继续增加。
  • 发言者称,Anthropic对训练和推理的需求一直“非常旺盛”,但这段表述并未明确证明Anthropic正在使用Trainium。他补充称,OpenAI实验室可能已经、且未来也会使用Trainium,承担大规模训练和行业推理工作负载。
  • AWS还称,更广泛的客户群体,尤其是初创公司,也在寻求使用Trainium。

5. 推理解耦最终服务于降低token成本

  • Cerebras与Trainium的产品支持通过prefill与decode解耦实现推理。
  • AWS称,合作的首要目标是降低客户成本,尤其是推理成本和“每token成本”。
  • 客户告诉AWS,要部署更多生成式AI应用,关键取决于成本下降。
Speaker 1

We have 2 partners from AWS with us today. We're doing a discussion on your partnership with NVIDIA, your statement on the 1 million-GPU deployment, and a discussion on training, the progress there, and the partnership with Cerebras. Could you elaborate more on that?

Speaker 2

Well, thanks. It's great to be here at GDC 2026, and thanks for having us here. Our partnership with NVIDIA goes back almost 15 years. We were one of the first cloud providers to start offering GPUs, before generative AI was really a thing. It's a long, deep partnership.

1. AWS Expands Its GPU Footprint

If you look at our infrastructure today, with the new GPUs, we offer about 2 million GPUs in the cloud—the largest missile provider to offer 2 million GPUs via cloud. We recently launched a blog saying that we're going to add another 1 million GPUs just this calendar year. Basically, that's 50% of the entire footprint we've built in the last 15 years.

I think that says 2 things. 1 is the huge customer demand we're seeing on AWS for this infrastructure, and 2 is a testament to our ability to scale so fast and in such a short period of time. We deeply value the NVIDIA partnership and look forward to bringing the next generation of Rubin GPUs, or Rubin systems, this year. Our customers are excited to have them.

Speaker 1

What are the challenges to scaling, and how is AWS getting better at that? What are your advantages there? Also, when do you expect Rubin to come online?

2. Scaling AI Into Production

Speaker 3

I can start with that. I lead product marketing for AI infrastructure at AWS. A common problem we're seeing with customers is that it's easy to build a very cool generative AI demo in days or even hours, as you can see at all these trade shows and industry events. But when customers really hit a wall is when they deploy these AI solutions in production.

It's not just about deploying applications on GPU instances; it's also everything around them, like the data pipeline, cost control, security, and compliance. That's really where AWS can help customers, with our 20 years of operating infrastructure at massive scale. We offer the largest, broadest, and deepest set of services that customers need to bring these applications to production, versus doing a proof of concept at a much smaller scale.

Speaker 1

But I think, back to what you were saying, let's go more into the training side. You had the Trainium announcement at re:Invent last year.

That was a Trainium 3 announcement, with a brand-new Scalapack architecture, right? You recently announced a partnership with Cerebras to do this disaggregated prefill-and-decode architecture. Could you elaborate more on that? How are your customers viewing these options, and what applications do you see Trainium bringing to your customers?

3. Why AWS Built Trainium

Speaker 2

One of the core strategic principles for AWS when it started 20 years ago was to bring down costs for customers accessing infrastructure and also give them a wider selection of different hardware platforms. That's why we've been investing in custom silicon, going back to our Graviton generation of CPUs some 8 or 9 years ago. The same thing is why we've invested in Trainium as our AI custom silicon chips.

4. Trainium Wins On Cost And Capacity

We want to bring down costs for our customers and give them a broad selection where capacity is still available, so that it's not an issue for them to scale up. From a cost perspective, I think the Trainium line of products offers 30% to 40% better price performance compared to the alternative accelerators available in the cloud.

From a capacity-availability perspective, it frees customers up to access alternative infrastructure, especially in this environment where there are supply constraints on the chip side, the power side, and everything in between. Customers appreciate and value that, and I think that's why we've been investing heavily in Trainium.

5. Trainium Three Scales Training

Trainium 3 is our latest-generation offering, with 2 to 3 times better performance compared to the previous edition, Trainium 2. We plan to scale it to more than 1 million chips this year, and more next year. Some of the recent customers are opening eyes. Obviously, Anthropic has been asked heavily for training and inference.

We're excited about partnering with AI labs, but at the same time, we're also seeing a lot of interest from a broader set of customers, especially in the startup arena, looking at using Trainium for their workloads.

Speaker 1

Is that for their chip-development workloads?

Speaker 2

Yeah, it's from AI labs.

Speaker 1

AI labs, yeah.

Speaker 2

The OpenAI labs, probably—I mean, they are and will be using it for large-scale training and industry inference workloads. Recently, we announced the Cerebras and Trainium product supporting disaggregated inference. And the team is also similar there: bringing down costs for our customers, especially for inference, and bringing down the dollar per token.

That's something we hear from customers: for us to deploy more generative AI applications, we want to see costs come down. The partnership's primary goal is to try to bring down costs for our customers.

Speaker 1

Thank you. I think—thank you very much for going through the efforts you're making to help the ecosystem build, and for the efforts you're putting into the infrastructure and offering the widest and broadest range of products for our customers.

Speaker 2

Great. Thanks for having us.

Speaker 3

Thanks for having us.