SemiAnalysis · · 47 min
AWS Trainium: How Amazon Built Their Own AI Chips | Researcher Conversations at GTC
Jordan NanosDan NishballKang Wen CheangZane Fong
TL;DR
- AWS says its cloud fleet contains about 2 million GPUs and plans to add another 1 million during 2026. That increment equals “50% of the entire footprint we’ve built in the last 15 years,” which AWS presents as evidence of customer demand and its ability to scale quickly.
- AWS says it looks forward to bringing next-generation NVIDIA Rubin GPUs or systems this year, extending a partnership approaching 15 years. Yet GPU instances alone do not deliver production AI: customers “hit a wall” around data pipelines, cost controls, security and compliance.
- Trainium is AWS’s cost-and-capacity option amid constrained chips and power. AWS claims 30–40% better price-performance than alternative accelerators available in the cloud, arguing that broader hardware choice helps customers access capacity and scale.
- Trainium 3 promises two to three times better performance than Trainium 2, with AWS planning more than 1 million chips in 2026 and more next year. The speaker cites Anthropic in connection with heavy training and inference, but the wording does not clearly establish Trainium use; he says OpenAI labs probably are and will be using Trainium for large-scale training and industry inference, while startups are showing interest.
- The Cerebras partnership targets inference economics through disaggregated prefill-and-decode. AWS says its primary goal is lowering customer costs, especially “dollar per token,” which customers identify as necessary for broader GenAI deployment.
Digest · the substance, structured for research
1. AWS plans to add half its historic GPU footprint in one year
- The AWS executive says the NVIDIA partnership dates back almost 15 years, when AWS was among the first cloud providers to offer GPUs, before GenAI was really a thing.
- AWS reports about 2 million GPUs today and another 1 million planned for 2026—“50% of the entire footprint we’ve built in the last 15 years.”
- AWS says it looks forward to bringing next-generation Rubin GPUs or Rubin systems this year (2026); no more precise availability date is offered.
2. Production AI requires far more than GPU instances
- AWS’s product-marketing lead says demos take “days or even hours,” but customers “hit a wall” when moving AI into production.
- It is not just about deploying applications on GPU instances: the surrounding stack includes data pipelines, cost control, security and compliance.
- AWS points to 20 years of operating infrastructure at massive scale and says it offers the largest, broadest and deepest set of services for taking AI applications from proof of concept to production.
3. Trainium addresses both cost and constrained capacity
- Custom silicon follows AWS’s strategic aim of lowering infrastructure costs and widening hardware choice; Graviton began that effort eight or nine years ago.
- AWS says the Trainium line offers 30–40% better price-performance than alternative accelerators available in the cloud.
- AWS frames that choice as a way to free customers to access alternative infrastructure amid constraints on chips, power and “everything in between,” so capacity need not prevent scaling.
4. Trainium 3 is planned to exceed one million chips in 2026
- Trainium 3 offers two to three times better performance than Trainium 2; AWS plans to scale it to more than 1 million chips in 2026 and more the following year.
- The speaker says Anthropic has been “asked heavily for training and inference,” but the wording does not clearly establish that Anthropic is using Trainium. He adds that OpenAI labs probably are—and will be—using it for large-scale training and industry inference workloads.
- AWS also reports interest from a broader customer base, especially startups, looking to use Trainium.
5. Disaggregated inference is ultimately a token-cost strategy
- The Cerebras-and-Trainium product supports disaggregated inference using prefill and decode.
- AWS says the partnership’s primary goal is bringing down customer costs, especially inference costs and “dollar per token.”
- Customers tell AWS that deploying more generative-AI applications depends on costs coming down.