[BidClub_]
The a16z Show · · 66 min

Dylan Patel on GPT-5’s Router Moment, GPUs vs TPUs, Monetization

Dylan PatelErin Price-WrightGuido AppenzellerErik Torenberg

YouTube
TL;DR
  • GPT-5 is less a frontier-compute leap than an “economic release” built around routing. Dylan Patel argues that power users lost access to GPT-4.5 and o3—with o3 thinking roughly 30 seconds on average versus GPT-5’s 5-10 seconds—while free users sometimes receive reasoning they never had before. The router lets OpenAI choose regular, mini, or thinking models and “gracefully degrade” service, trading maximum capability for dramatically greater token capacity.

  • The router’s larger prize is matching inference spend to each query’s monetizable value. A “why is the sky blue?” request can go to mini, while a search for the best DUI lawyer, flight, or product can receive “ungodly amounts of compute” because OpenAI could complete the transaction and take a cut. With an estimated 10% of Etsy traffic already coming from ChatGPT, Dylan sees agentic commerce—not result-degrading ads—as the route to monetizing free users.

  • Flat-rate AI subscriptions are colliding with 20x differences in consumer usage. Anthropic and coding products have tightened rate limits after heavy users exploited “negative gross margin” plans; one developer reportedly rearranged sleep into sailors’ power naps, while a Reddit leaderboard included usage worth roughly $30,000 a month. Enterprises may support commitments or averaged flat fees, but consumers increasingly point toward usage pricing—even as products use subscriptions and superior review interfaces to create stickiness.

  • AI may already create more value than its infrastructure costs, but the labs capture only a fraction of it. Dylan Patel’s coding thought experiment—30 million developers, productivity doubled, $100,000 of value each—produces $3 trillion of potential GDP value from one use case, while Dylan believes OpenAI captures “not even 10%” of the value ChatGPT has created. That mismatch need not stop capex: hyperscalers could grow spending another 20-30%, with CoreWeave, Oracle, infrastructure funds, and sovereign capital adding less immediately economic capacity.

  • Custom silicon is Nvidia’s largest threat only if AI demand remains concentrated among a few giant buyers. Google is making millions of highly utilized TPUs, Amazon millions of Trainium chips, and Meta is sharply increasing internal-silicon orders; Dylan thinks Google should physically sell TPUs, not merely rent them. If open models and cheap deployment disperse demand, however, Nvidia’s universal ecosystem strengthens—and independent challengers must be “like 5x better” before supply-chain, software, and margin disadvantages erase the lead.

  • American AI deployment is constrained less by electricity’s cost than by powered sites, grid equipment, and construction speed. Dylan puts roughly 80% of a Blackwell data center’s cost in GPUs, networking, buildings, and power-conversion capital, leaving only 20% for land, electricity, cooling, backup power, and related items. That makes paying extra to launch three months earlier rational; meanwhile, China’s current constraint is capital and chip quality rather than power, despite its ability to scale generation faster.

  • Intel needs immediate operational surgery and capital, while several platform incumbents need product urgency. Dylan says Intel’s five-to-six-year design cycles and as many as 14 silicon revisions must fall toward two-to-three years and one-to-three revisions; without a major cash infusion or severe cost cuts, it could “literally” go bankrupt before a formal separation is completed. His broader calls: Nvidia should reinvest its projected $100 billion-plus cash pile into infrastructure, Google should open TPUs, Apple should spend perhaps $50 billion on AI infrastructure, and Erik says Microsoft must “shake the crap out of the company” despite its extraordinary starting position.

Digest · the substance, structured for research

1. GPT-5’s real breakthrough is control over inference economics

  • Dylan’s disappointment is user-tier specific: paid power users lost GPT-4.5, which he still considers the better pretrained model for some work, and o3, which thought for roughly 30 seconds on average. GPT-5 thinking typically runs only 5-10 seconds, so less compute reaches his average query.

  • That does not mean no progress. GPT-5 is roughly the same model size and materially improves on the vanilla predecessor, while avoiding pathological reasoning such as o3 spending 48 seconds deciding whether pork is red or white meat. Anthropic had already shown that comparable or better answers could require far less thinking.

  • The router now chooses among the regular model, mini after rate limits, and thinking—with control over how long reasoning runs. Free users occasionally receive a much stronger answer; OpenAI can also “gracefully degrade them” when capacity tightens. Guido floated a meme—explicitly saying it was not true—that OpenAI had put o3 and smaller models behind a router at a lower blended price; Dylan said there was “a little bit of that.”

2. Query value can determine how much intelligence OpenAI spends

  • Dylan’s business framing: “the router points to the future of OpenAI.” Traditional ads conflict with a helpful assistant—injecting promotions can worsen answers, while banner ads fit poorly—so the free user needs a transaction-native monetization path.

  • His contrast makes the allocation logic concrete. “Why is the sky blue?” deserves mini; finding the best DUI lawyer might justify searching court filings, contacting local firms, comparing results, and spending “ungodly amounts of compute” because OpenAI can take a cut from a high-stakes transaction.

  • Shopping and travel are the obvious wedge: Dylan cites 10% of Etsy traffic coming from ChatGPT while OpenAI earns nothing, partly because Amazon blocks ChatGPT. His advice is to add a credit card, calendar, and preferences such as aisle versus window, then let the agent book while charging an agreed take rate.

3. AI pricing is exposing the extremes hidden inside subscriptions

  • OpenAI said it doubled rate limits for large numbers of users and dramatically increased served tokens, making GPT-5 “an economic release.” The implication is cheaper blended inference, not simply a new winner on MMLU or another intelligence benchmark.

  • Heavy coders expose flat-rate plans’ fragility. After Anthropic imposed hour-based as well as weekly limits, one user reportedly adopted fragmented sleep like a solo sailor; a Reddit leaderboard featured someone consuming roughly $30,000 monthly through a subscription. “People are taking advantage of the negative gross margin.”

  • Consumers can vary by about 20x, pushing model vendors toward usage pricing; enterprises can average predictable full-time developer behavior and may pay substantial commitments to avoid open-ended bills. One source of product stickiness is the verification loop: visualizing changed files, consequences, diagrams, and fast versus complex feedback better than competitors.

4. Nvidia’s runway rests on accelerating demand and abundant speculative capital

  • Dylan divides current chip demand into rough thirds. OpenAI and Anthropic alone account for about 30%: Anthropic’s compute comes from Google and Amazon, while OpenAI’s comes from Microsoft, CoreWeave, and Oracle. Advertising buyers such as Meta and ByteDance take another third; the remaining, less clearly economic providers may struggle to keep raising ever-larger rounds.

  • The hosts challenge the ceiling through coding alone: even a conventional GitHub Copilot rollout may yield 15%, while Dylan insists better products can do much more. At 30 million developers, doubled productivity, and $100,000 of value each, the hypothetical reaches $3 trillion—before counting other AI applications.

  • Dylan’s correction to the “$600 billion problem” is that infrastructure purchased now accounts for perhaps five years of rising revenue rather than one year of flat revenue. AI already creates more value than the spend, he argues, but “value capture is broken”: his four-person team uses inexpensive Gemini APIs to analyze permits, regulatory filings, satellite imagery, generators, cooling towers, substations, and construction progress, then captures far more downstream value than the model provider.

  • Economically justified capex has limits, yet available funding does not. Hyperscalers could grow capex another 20-30% next year; CoreWeave and Oracle can raise substantially more through capital markets; Brookfield, Blackstone, G42, GIC, and other sovereign or infrastructure pools have “barely started touching AI.”

5. Market concentration determines whether TPUs or Nvidia win

  • Custom silicon is the central Nvidia threat. Guido says Google and Amazon are making millions of chips and that Meta is sharply increasing orders; Dylan says Google’s TPUs are essentially 100% utilized, while Amazon’s Trainium still lags in utilization. Guido characterizes Microsoft’s custom silicon as “kind of sucks”; Dylan later calls Microsoft’s internal chip effort the worst of any hyperscaler.

  • The governing variable is concentration. A few enormous AI buyers can amortize their own chips and compress Nvidia’s margin; dispersed demand driven by Chinese open models and cheap inference libraries favors Nvidia’s broadly supported platform. Dylan therefore allows that Nvidia could remain the world’s most valuable company for a long time.

  • Dylan asks whether Google should start selling chips to everyone, and Guido agrees that it should sell TPUs externally, not merely rent them. Doing so would require a cultural reorganization across Cloud, TPU, JAX, and XLA. Dylan argues that a larger TPU business could support a higher Google valuation, while Guido notes that Google leadership would still believe Gemini will ultimately be worth much more.

6. A startup must beat Nvidia by 5x before reality cuts the lead down

  • Capital is funding Etched, Rivos, MatX, and others before public chips exist, alongside established challengers such as Groq, Cerebras, SambaNova, Tenstorrent, and SoftBank-owned Graphcore. The hyperscalers possess an advantage none of them shares: a captive customer willing to accept a specialized chip as a margin-compression exercise.

  • Independent vendors must design silicon and software, assemble IP, manage chips, racks, networking, memory, and customers, then earn a margin. AMD illustrates the difficulty: despite strong engineering, it uses more silicon area and memory for comparable performance and sells near 50% gross margin against Nvidia’s roughly 75%.

  • Hardware-model co-design becomes a trap when research moves. Cerebras, Groq, and SambaNova emphasized more on-chip SRAM and less DRAM for then-leading workloads; larger models and vision transformers changed the economics. Newer startups optimized giant systolic arrays for dense transformers, only to encounter DeepSeek-like workloads with much smaller shapes and many small matrix multiplications.

  • Nvidia brings better networking, HBM, process access, ramp speed, and bargaining power with TSMC, SK hynix, rack suppliers, and cable vendors. A challenger’s “5x” architectural edge can become 2.5x after supply-chain penalties, then roughly 50% after Nvidia compresses margin and deploys its software defense. Meanwhile, “you have to advance in the tech tree” without knowing where models will branch.

7. China has power, but capital efficiency and chip access still bind

  • Dylan notes that Chinese provinces have rules saying the H20 is insufficiently efficient, even though he considers it China’s best available AI chip and says Huawei remains behind. In power-constrained America, even a free H20 can be unattractive: consuming scarce megawatts with weaker silicon means lower compute capacity.

  • Nvidia’s export argument is ecosystem control. Chinese developers contribute valuable Nvidia-compatible software—including Triton extensions—so selling GPUs can prevent Huawei from establishing a rival stack. Dylan believes models deliver more economic value to society than hardware; supplying H20s and a cut-down Blackwell could therefore transfer more economic capability than chip sales capture.

  • China’s immediate bottleneck is “always capital,” not electricity. Its AI capex is growing faster in percentage terms than America’s, but from lower absolute dollars and with worse output per dollar; it already subsidizes semiconductors by roughly $150-$200 billion annually and could fund a Meta- or Google-scale national effort if it chose.

8. Powered land—not cheap cooling—is the scarce AI asset

  • Chinese firms bypass domestic limitations by renting superior GPUs abroad or building through Singapore-linked entities. ByteDance is one of Google Cloud’s largest customers and also rents from Oracle and Microsoft because overseas Blackwell capacity can beat self-built domestic infrastructure on dollars per unit of output.

  • America’s physical bottleneck leaves purchased chips idle. Google has TPUs waiting for powered facilities, while Meta and others have GPUs in the same state; chips alone represent roughly 60-80% of cluster cost. Interconnections, transmission, substations, electrical contractors, and travel electricians have become schedule-critical inputs.

  • CoreWeave’s value is substantially its willingness to move fast—converting crypto sites, maintaining bare-metal clusters, and buying powered assets. Google took an 8% position in TeraWulf for its power, while hyperscalers have effectively said, “Screw it,” to their sustainability pledges as speed overtook prior commitments.

  • Guido says cooling is secondary: alfalfa uses about 100x as much water as AI data centers today, while underwater facilities save perhaps 5-10% but become unserviceable. For a Blackwell facility, about 80% of cost is capital and 20% covers land, power, cooling, backup systems, and generators; Elon’s expensive temporary generators and chillers were rational because they enabled the data center to come online three months faster.

9. Intel’s survival matters more than an elegant corporate separation

  • Guido says the world needs Intel because Samsung appears further behind on 2-nanometer-class process development, while TSMC holds a practical leading-edge monopoly. TSMC is raising some prices only 3-10% next year despite its pricing power; if something happened to Taiwan, Intel could possess the world’s most advanced process technology, albeit uneconomically.

  • Intel’s fab and design businesses should eventually separate, but executing the split could consume the management time the company does not have. The urgent defects are operational: five-to-six-year design-to-launch cycles and sometimes 14 tape-out revisions, versus roughly three years and one-to-three revisions for strong competitors.

  • Dylan’s prescription for CEO Lip-Bu Tan is to run the cultures separately, remove layers and weak managers, retain the engineers who led process technology for two decades, and reduce design-to-launch toward two-to-three years. The x86 and PC franchises can remain highly profitable without AI-leading growth, potentially with one-third or half as many people.

  • The fab requires much more capital for subsequent generations. Without a major infusion or drastic cuts, Dylan warns Intel could “literally” go bankrupt before restructuring finishes; Erik’s hoped-for backstop is each major hyperscaler contributing perhaps $5 billion before TSMC’s margin potentially approaches 75%.

10. Nvidia, Google, and Meta should turn compute ownership into distribution

  • Dylan would tell Jensen Huang to invest Nvidia’s war chest into the infrastructure layer. Year-one depreciation of GPU clusters under the new Trump tax bill has major tax implications for customers, while Nvidia itself could hold north of $100 billion in cash by year-end; buybacks and dividends would squander the chance to accelerate powered capacity and control more of the stack.

  • Dylan says Google should open up more of the ecosystem around TPUs and XLA, sell hardware, build data centers faster, and recover its former compute lead. Sergey Brin works closely with DeepMind, but physical infrastructure and product shipping remain too slow as purchasing agents threaten to disintermediate monetizable search queries.

  • Mark Zuckerberg already recognizes urgency—building temporary “tents,” hiring aggressively after unsuccessful efforts to buy Thinking Machines Lab or SSI for $30 billion, and linking superintelligence to wearables and assistants. Dylan’s complaint is execution outside Meta’s gardens: products are often “kind of mid,” so it should ship explicit ChatGPT and Claude Code competitors faster.

11. Apple and Microsoft risk losing the interface; Elon risks losing focus

  • Apple’s hardware and form-factor work remains strong, but Dylan thinks it could “lose the boat” without perhaps $50 billion of AI infrastructure. As agents integrate calendars, messages, preferences, and transactions, AI becomes the computing interface and weakens Apple’s ability to control experience through touch, keyboards, and its walled garden.

  • Microsoft was aggressive in 2023 and 2024, then pulled back on data centers while OpenAI began slipping from its grasp. Dylan calls its internal chip program the hyperscalers’ worst, MAI failing, and Azure vulnerable to Oracle, CoreWeave, and Google. Erik says GitHub Copilot and Microsoft Copilot are weak despite Microsoft starting with GitHub, the best source-code repository, enterprise distribution, first-mover status, and a strong relationship with a model company.

  • Dylan’s Elon Musk assessment stays hedged: porn models could accelerate xAI revenue, robotaxi is “starting to look good,” and Musk remains a magnet for exceptional builders. But talent losses, killed projects, and snap decisions are now damaging alongside their historic upside; the closing suggestion is simply to “focus on the products again.”

Dylan Patel

Nvidia is going to have better networking than you. They’re going to have better HBM, a better process node, and they’re going to come to market faster. They’re going to be able to ramp faster and have better negotiations with whether it’s TSMC or SK hynix on the memory and silicon side, or with all the rack people, copper cables—everything.

They’re going to have better cost efficiency. You can’t just do the same thing as Nvidia. You have to really leap forward in some other way. You have to be 5× better.

Erik Torenberg

Dylan, welcome to the podcast.

Dylan Patel

Thank you for having me.

Erik Torenberg

We’ve been trying to get you for a while. You’re a busy man, but it worked out. Guido, do you want to introduce why we’re so excited to have Dylan on the podcast and what we’re excited to discuss?

Guido Appenzeller

Look, I think Dylan, you’ve done an exceptional job covering what’s happening in the AI hardware space, the AI semiconductor space, and now more and more in the data center space as well. Currently, the most valuable company on the planet is an AI semiconductor company, right? The biggest IPO so far in AI was an AI cloud company.

This is currently where it’s happening. In any gold rush, in the early days, it’s the picks and shovels that make money, and I think this is the stage that we’re in. So, we’re super excited to have you here today.

Dylan Patel

Awesome. Thank you. Happy to talk about my favorite topics.

Erik Torenberg

Amazing. Well, maybe let’s start with GPT-5. We just had some of the researchers from OpenAI—Christina and Isabella—on here last week. You said it was disappointing. Why don’t you share your reactions, what capabilities you were hoping to see, or your overall reaction?

Dylan Patel

I think it depends on what tier of user you are, right? If you’re just using GPT-5 and before you were a $20 or $200-a-month subscriber, you no longer have access to GPT-4.5, which, in my opinion, is still a better pretrained model for certain things. Or you no longer have access to o3, which would think for 30 seconds on average, maybe, right?

Whereas GPT-5, even when you’re using thinking, only thinks for 5 to 10 seconds on average, which is an interesting phenomenon. Basically, GPT-5 is not spending more compute, per se. The model did get a little bit better on a vanilla basis. GPT-4.5 is actually quite a bit better.

But when you think about what this curve of intelligence looks like, the more compute you spend, the better the model gets. That’s whether it’s a bigger model—which GPT-5 isn’t; you can see it’s not a bigger model, it’s roughly the same size—or whether you think more. OpenAI’s first thinking models, the first few generations of o1 and o3, would think for a long time and waste a lot of tokens, if you will.

When you look at, for example, Anthropic’s thinking models, they actually think a lot less when you put them in thinking mode, while getting to the same results or better results than OpenAI was. OpenAI optimized a lot of that. The silliest example I had was when I asked o3, “Is pork red meat or white meat?” It thought for 48 seconds. I was like, “What are you doing? This should just tell me the answer.”

The nice thing is that GPT-5 will think a lot less, even if you select thinking manually. More importantly, they have the auto functionality—the router—which lets them decide whether to route you to the regular model, perhaps to a mini model if you’re out of rate limits, or to the thinking model, and how much to think.

In general, the thinking model will think less. So, there’s less compute going into a power user’s average query than before.

Guido Appenzeller

But isn’t it even more interesting? OpenAI can now control how much compute it wants to allocate to you, right? If we’re in a high-load situation, maybe it tunes the router a little bit so it uses less. I have no idea what they’re doing behind the curtain, but there’s this meme out there at the moment that all they did—which is a meme, right? It’s not true—is take o3, plus a couple of smaller models, put a router in front, and offer that at a lower blended price, essentially.

Dylan Patel

I think there’s a little bit of that. Cost suddenly matters, and they figured out a way to steer that. They talked about how they’ve been able to dramatically increase their infrastructure capacity. I myself was regularly using o3 or GPT-4.5, and now I’m forced to use auto, which sometimes gives me the o3-equivalent thinking model but sometimes gives me just the regular base model, which sucks.

For the free user, though, it’s actually quite interesting. The free user was not getting thinking models pretty much ever—not using them, or in many cases just opening the website and asking a query. Now, sometimes the query gets routed there, so sometimes they get a way better model. But now OpenAI can gracefully degrade them if they need to.

I think the router points to the future of OpenAI as a business. You can look at the model companies: Anthropic is fully focused on B2B—APIs, coding, Claude Code, whatever it is. OpenAI has that business, Codex, and its API business, but the majority of its revenue is really from consumers. It’s consumer subscriptions, but it has no way to upsell or make money off all the free users.

In any other consumer app, the free user still pays via ads. But that’s not compatible with AI. It’s a helpful assistant; you can’t just make the user’s result worse by injecting ads. Banner ads don’t really work in AI, either. So how do you monetize them now?

I think with the router, they’re getting really close to figuring out how to monetize that user. If you saw the new CEO of applications’ product that she launched at Shopify—I think it was Shopify—it was an agent for shopping. Now this immediately clicks: if the user asks a low-value query like, “Why is the sky blue?” just route them to mini. The model can answer perfectly fine, and that’s a large chunk of queries.

But if they ask, “What’s the best DUI lawyer near me?” all of a sudden, you’re in jail and you have one shot. You’re like, “Screw it. Let me ask ChatGPT what the best DUI lawyer is.” Soon enough, the model will be able to contact all the lawyers in the area, figure out what their results are, search their court filings, and book the best lawyer—or an airplane.

Guido Appenzeller

Negotiate a cut as part of that, you know.

Dylan Patel

Yeah, of course they’re going to take a cut. But this is a much better way of monetizing the free user. Etsy gets 10% of its traffic now from ChatGPT, and OpenAI makes nothing off of that, but they really will soon. Partially, that’s because Amazon blocks ChatGPT, but there’s a way to make money from shopping decisions, whether it’s booking flights or looking for items.

You can say, “Free user, I don’t care. I’m going to send you to my best model. I’m going to send you to agents. I’m going to spend ungodly amounts of compute on you because I can make money off of this.” But if it’s a query like “Help me with my homework,” I’ll send you to a decent model. I don’t need to spend money on you.

That’s how I think OpenAI can finally make money off the free user. I think that’s the biggest thing about the router.

Guido Appenzeller

This is super interesting. I think this is the first time we’ve seen a launch of a new model where, to some degree, cost is the headline item. So far, it was always, “Who has the smartest model? Who has the highest MMLU score?” Now we suddenly have people who use models for coding for 8 hours a day and are surprised that, if you take a large context window and the best model, it creates thousands of dollars of cost a month.

Cost matters. To some degree, the shift on the Pareto frontier between cost and performance is the new benchmark for model competitiveness. Is that what we’re seeing here, or is cost alone…?

Dylan Patel

I mean, definitely. OpenAI said it doubled its rate limits for large numbers of users. It has dramatically increased the number of tokens it’s serving from this launch, which effectively says this is an economic release.

Guido Appenzeller

And it probably also means the tokens are all cheaper, right? Otherwise—

Dylan Patel

Yeah, for sure. I think the funniest thing about this whole cost issue is that we’ve seen it in the coding space. Cursor had to pull away the unlimited cloud code. Initially, it had a super-expensive plan with unlimited usage, and then it had only a weekly rate limit.

Now they have hour-based rate limits. I saw the craziest thread on Twitter where this guy said he changed his sleep schedule, modeled after how sailors operate when they're solo sailing. You can't sleep uninterrupted, so they'll take power naps when they get to the right spots so they can still be safe.

Erin Price-Wright

They do that in the morning, when it's not very windy. But they can't sleep uninterrupted, right? Because Anthropic had to put rate limits in place that are based not just on weeks, but on a number of hours, he basically sleeps multiple times a day in small chunks just so he can maximize his usage. There's also a leaderboard on Reddit where people are competing to see how many tokens they're using through their subscription. There's a guy spending $30,000 a month.

Erik Torenberg

So I'm going to find a developer in India to do pair programming with. I can get the day cycle, he can get the night cycle, and we can both maximize the account's quota. Is that the future, then?

Dylan Patel

It's clear people are taking advantage of the negative gross margin in the subscriptions that are offered. I think Anthropic probably makes a positive gross margin off my subscription. I don't code enough, but plenty of people are definitely losing money. As you said, it's an economic issue.

Guido Appenzeller

It'll push more and more, I think, toward usage-based pricing. If you have an underlying commodity that you're reselling to some degree, and that commodity is that large a part of your cost of goods, you need to go to usage-based pricing.

Erik Torenberg

How much do you think the customer capture and stickiness for these code products is? I'm curious what you think. Once you use an IDE, once you integrate one of the CLI products, how sticky is it, or is it?

Guido Appenzeller

A billion-dollar question. That's a very conservative estimate. Look, Andrej Karpathy has this great slide where he basically says that if you're building an agentic system today, fundamentally what it is is this loop. Half of the loop is the model thinking and trying to do something. The other half is the user verifying what the agent did, deciding whether it's the right thing, providing feedback, and trying to steer it in the right direction, because we can't run forever. Eventually, you need to steer it back.

One half of that is the model provider. They're trying to build the best models. The other half is really about designing the best possible UI to enable the user to give feedback. I think there's value in that, so I think there's a certain amount of stickiness there. Take code editing: What are all the different tools for visualizing the code changes? How can I most easily visualize what they impact and which files? How can I get very quick feedback for small changes, versus complex feedback for complex ones? There are some tools that actually draw diagrams for you of what they do.

I think this will be the battle. I think there's stickiness in that. How much exactly? This was a great question.

Erik Torenberg

So in that sense, people should be doing subscriptions to get people locked in, instead of moving to usage-based pricing?

Guido Appenzeller

I think it's the customers that don't want to do usage-based pricing, because it's so hard to guarantee. It's hard for usage to get away from them. You actually want guarantees, and you're willing to commit to pretty high spend in order to avoid usage-based pricing. I think it's the model companies that want usage-based pricing.

Erin Price-Wright

I think with consumers, it's frankly very hard not to have usage-based pricing, just because the variability is so massive. Us coding versus somebody who does this as their full-time job—you just have a factor of 20 or so difference in usage. If that costs a lot of money, I think for enterprises we could see more flat-fee pricing, because you can average it out more.

Guido Appenzeller

You have a developer who's using it all day. You kind of know, in a general sense, how many hours a day they're programming and what that sort of looks like.

Erin Price-Wright

The vibe quotas are harder.

Erik Torenberg

Yeah.

Erik Torenberg

Before we leave OpenAI, I want to ask a broad question. If someone was sitting here and saying, “Hey, Dylan, I'll listen to anything you tell me to do. Any advice you have, as long as it makes OpenAI more valuable,” what would you tell them?

Dylan Patel

I would say immediately launch a method for you to input your credit card into ChatGPT and agree that, for anything it agentically does for you, it'll take a cut. Then launch that product, because shopping is where we know Anthropic, OpenAI, and all the other labs are buying RL environments from Amazon, Shopify, Etsy, and all the different ways to shop on the internet, as well as airline websites.

Just integrate my calendar. I want to fly there on Thursday. Make sure I don't miss a meeting. Great—book it. Do that integration extremely well. Know my preferences, whether I like an aisle or window seat, and just take a take rate. I think this will make them so much money the moment they launch it.

I think they're working on it already, but I'd like to hear how Sam Altman thinks about it, because he's shifted his tone massively on ads over the last 6 months. He used to be like, “No way,” and now he's like, “Maybe there's a way to do it without harming the user.” I think this is how you monetize the free user. That's probably what I'd tell him, or ask him about—a whole line of questions around this.

Erik Torenberg

I want to shift to Nvidia. Nvidia is having a monster year. They're up almost 70%. What are the possible paths from here? How do you see it playing out?

Guido Appenzeller

It depends on how bullish you are on the continued growth, but I think you guys have a good vantage point. We have a good vantage point of how fast revenue is growing for a lot of these companies, especially the code companies, but even many other applications. I think we can clearly see the demand side is accelerating.

On the training side, I think the race is on. Meta is upping hugely. Google is upping hugely. If you just look at OpenAI and Anthropic and the compute that they have and are getting this year—from Google and Amazon for Anthropic, and from Microsoft, CoreWeave, and Oracle for OpenAI—30% of the chips are going to those 2 companies alone.

Well, 1/3 of it is ads, whether it be ByteDance, Meta, or many of the other people who are doing ads. So then it's still, okay, where are the rest of these 1/3 of the chips coming from? They're mostly uneconomic providers, and I don't think it's an obvious bet that they're going to keep growing and raising bigger and bigger rounds. So what happens there?

We talked about coding earlier. Qwen3-Coder is actually super cheap if you're running it on-prem or in the cloud with all these inference libraries, and there's stuff like that as well. I think the question is how much does it keep growing? Clearly, the first 1/3—OpenAI and Anthropic lab spend—is definitely skyrocketing. The second 1/3, ads, is going to grow. It's not going to grow like crazy, but I think there's definitely an inflection point that could be hit with generative AI ads.

I know Meta has been experimenting with it a lot, but I could totally be convinced that there's going to be a huge inflection in the take rate there, where you start showing me personalized ads. Every ad looks like me, and I'll be like, “Okay, yes.” Except it's slightly better, so I feel better, and I want to buy it.

It's interesting. I have no idea how this is going to scale, but if you ask how much it could scale, how much value we're creating here, can we create enough value to actually keep growing for a long time? If you just take AI software development as an example, right?

Erin Price-Wright

We know we can easily get about 15% more productivity out of—

Dylan Patel

I don't think that's right. I think it's way higher.

Erin Price-Wright

No, no. With a straight—I've talked to a lot of enterprises—a classical enterprise, straight-up GitHub Copilot deployment gives you about 15%. We can do much more than that.

Dylan Patel

But, bro, you know how bad GitHub Copilot is? Look at the revenue ARR chart. It's so funny: in 3 months, Claude Code has surpassed them, and Cursor easily surpassed them. Even companies like Replit and Windsurf/Cognition are going to pass them. It's like you're preaching to the choir. So look, let's assume we can get this to 100%.

So we can double the productivity of a developer, right? About 30 million developers worldwide, give or take. Let’s say $100,000 in value added per developer. This might be a little high worldwide; the U.S. is low, but worldwide is high. So it’s $3 trillion.

Erin Price-Wright

Yeah, yeah.

Guido Appenzeller

Right. So we’re probably building technology here that adds $3 trillion of GDP value.

Erik Torenberg

Just from a coding model?

Dylan Patel

Just from this one use case. So at least in theory, the value generation is here to keep growing, right?

Now, how that translates to the industry is much more complicated. There’s the whole famous $300 billion problem—or $200 billion problem. Now it’s the $600 billion problem. I’m sure someone is going to put out the $1 trillion problem soon enough. But there is some reality in that, of course. It ignores that infrastructure spend today is accounting for 5 years of revenue, not 1, and the revenue looks like this, not a flat line.

But I think the main thing is that AI is already generating more value than the spend. The value capture is broken, right? I legitimately believe OpenAI is not even capturing 10% of the value they’ve created in the world already, just through usage of ChatGPT. I think the same applies to Anthropic, Cursor, and whoever else you’re looking at. I think the value capture is really broken.

Even internally, I think what we’ve been able to do with 4 developers in terms of automation is remarkable. Our spend on the Gemini API is absurdly low, and yet we go through every single permit and regulatory filing around every single data center with AI. We take satellite photos of every data center, and we’re able to label our data set and recognize the generators people are using, the cooling towers, the construction progress, and the substations.

All this stuff is automated, and it’s only possible because of generative AI. We do it with very few developers, and the value capture that I’m able to generate by selling this data and consulting with it is so high. But the companies making it get nothing out of it, right? I think there’s a value-capture challenge here that far exceeds the creation. As you get models like GPT-5 or open-source models continuing to drive costs down, the value capture is just harder and harder and harder for these companies because they’re making 50% gross margin on inference, if they’re making that, or less in many cases.

Erik Torenberg

In so many words, you’re saying we’re getting commoditized, and therefore you can’t capture the value. Thus, you should temper your expectations of how much you can spend on GPUs.

Dylan Patel

Well, no. I think there are still ways to inflect hugely on value capture, right? I mentioned that ads are a huge value capture.

Erik Torenberg

That needs to happen before we see a massive increase.

Dylan Patel

No, I think the other thing is that there’s a lot of capital that hasn’t been spent, right? The hyperscalers can still grow CapEx 20% to 30% next year from what they’re doing this year. In addition, companies like CoreWeave and Oracle, because they’re tapping capital markets, can raise way more than 20% to 30% in CapEx.

And you go down the list further, and it’s like, “Oh, the largest infrastructure funds in the world, like Brookfield and Blackstone.” Actually, they’re turning all of their attention to investing even more in AI infrastructure. Then you have the sovereign wealth funds of the world, like G42, Norway’s sovereign wealth fund, or GIC in Singapore. These people have barely started touching AI.

So I think there’s a whole lot more CapEx that can come without it necessarily being economically motivated on day 1. I’m more so saying economically motivated CapEx can only grow so much, but there’s so much other spending where it’s not clear from a spreadsheet, if you’re basing it on a real business, that you should actually spend this much. But people will because they believe—I believe, I think you believe, and I think people in infrastructure believe—that you’ll get profit out of it. But there’s no 100% certain way to argue it.

Guido Appenzeller

How threatened, if at all, is Nvidia by custom silicon? I think that’s the biggest thing, right? When we look at orders from Google and Amazon, especially, and Meta—not Microsoft—their custom silicon kind of sucks. But the other 3 are really upping their orders massively over the last year.

Dylan Patel

Amazon is making millions of Trainium chips, and Google is making millions of TPUs. TPUs are clearly 100% utilized, right? Trainium’s not there, but I think Amazon will figure out how to do that. And Anthropic will.

I think that’s the biggest threat to Nvidia: people figuring out how to use custom silicon more broadly. If AI is concentrated, then custom silicon will do better. And that’s not even talking about OpenAI’s silicon team and stuff, right?

If AI is really concentrated, then custom silicon will do better. But if it gets dispersed broadly because there are all these open-source models from China, along with open-source software libraries from Nvidia and China, and that makes the deployment costs rock-bottom, then potentially—hear me out here—if Google’s TPU is able to compete with Nvidia, in theory, Google could do it on the open market. The TPU business is worth more than Google these days. Shouldn’t Google start selling its chips to everyone? In theory, it should be able to achieve a higher market cap.

Guido Appenzeller

I absolutely think so. I think Google’s even discussing it internally. It would require a big reorganization of the culture, how Google Cloud works, how the TPU team works, and how the JAX and XLA software teams work.

I totally think they could do it. It would just take them shaking themselves pretty hard to be able to do it. I totally think Google should sell TPUs externally—not just renting them, but selling them physically.

Erik Torenberg

It’s kind of funny if a side hobby, in theory, has higher company-value potential than your—

Erin Price-Wright

—entire business, especially as you think about the degradation of the search business.

Guido Appenzeller

Yeah. But I think if you were to ask Sergey, “Do you think selling chips and racks is more valuable than cloud or Gemini?” he’d be like, “No, no, no, no, no. Gemini is going to be worth way, way, way more. It’s just not there yet today, right?”

So I think today you say Nvidia is the most valuable company in the world. Again, it’s a whole concentration thing, right? If the world is super concentrated in terms of customers, then Nvidia will not be the most valuable company in the world. But if it gets dispersed more and more—which arguably we’re starting to see, with a lot of these open-source models getting better and better and better, and with the ease of deploying them improving—then you could argue Nvidia will remain the most valuable company in the world for a long period of time.

Historically, no pun intended, software has eaten the world in most markets. If you look at the early networking days, Cisco was the most valuable company on the planet for a while. It’s no longer. The companies that built services on top, like Google, Amazon, or Meta, eventually eclipsed them.

Dylan Patel

Well, which is why Nvidia is making all these software libraries, right? That’s why they’re trying to commoditize inference. I don’t think you guys even have an inference API provider investment, do you?

I think I talked to one of the team members—maybe Rajko or someone—about why you guys didn’t invest in Together or Fireworks. The argument was, “We think just serving models alone, without making them, will be commoditized,” right?

Erin Price-Wright

We have all kinds of model providers.

Dylan Patel

Model providers, but I’m talking about a pure API provider investment. I think that’s right, isn’t it? I think I talked to one of the team members about why you guys didn’t invest in Together or Fireworks. The argument was that just serving models alone, without making them, would be commoditized, right?

Guido Appenzeller

We have some in the Stable Diffusion ecosystem, like fal. It’s a little bit different dynamically there, I think.

Dylan Patel

They tend to make much more compound models than the LLM folks. I think that’s a little bit different. But you guys don’t have one of these, like Baseten, or any of these API investments, because you think—this is from someone on the infrastructure team—that it’ll get commoditized.

The software Nvidia is making, along with vLLM and SGLang, which are open-source software projects coming out of Berkeley and now sort of have their own environments and are supported by many, means that API providers aren’t necessarily worth a ton, right? That’s sort of your argument, maybe. I think that’s relevant to this whole thing, which is: Why would you do this?

Erik Torenberg

Shifting gears, what about the silicon startups? What’s your take on those? There’s a ton of capital flowing into that, right? We’ve seen—I don’t have the numbers—but probably billions being invested in chip startups.

Dylan Patel

Yeah, for sure.

Whether you're looking at companies like Etched and Rivos, and a number of other companies like MatX and others, I think it's pretty impressive that they've gotten the amount of funding they've had without even launching a chip. In the past, silicon companies would make money or raise money, but they would at least launch a chip before they got a big round. But Etched and Rivos have raised a lot of money without ever launching a chip publicly, which I think speaks to—

Guido Appenzeller

Yes, silicon is super capital-intensive if you're building a chip, especially an accelerator, which has so many moving pieces. There are 10 different AI accelerator companies out there that are newish in the last few years.

Erin Price-Wright

I think there are a lot more.

Guido Appenzeller

Yeah. Yeah. Yeah. That's fair. And then there's the old guard, which continues to raise money, right? Groq, Cerebras, SambaNova, Tenstorrent, and so on and so forth. Or Graphcore getting bought out by SoftBank, with SoftBank dumping money into this effort as well, right? There's a lot of capital being invested to displace Nvidia's top position.

Erin Price-Wright

They're a captive customer, which is themselves, right? It's huge—

Guido Appenzeller

And they can just win on supply chain, right, by using cheaper providers.

Erin Price-Wright

It's a margin-compression exercise, essentially.

Guido Appenzeller

Yeah. And maybe for certain workloads, like Meta's recommendation systems, they'll have a better—you know, they can specialize more. But for the most part, it's like, no, we're targeting the same workloads. We can just simplify the supply chain or bring a lot of it in-house, compress margin, and it'll be fine.

But in the case of these other companies, it's like, well, they don't have a captive customer, so now you have to contend with the fact that you're using the same ecosystem. Either I can use some custom silicon provider who's going to take a margin anyway on top, and that's going to compress what I can sell for, or I can try to bring everything in-house. But then it's like, this is really hard, right? I'm going to do all the software design, all the silicon design, build all this different IP, and manage the supply-chain pain on chips, on racks, on everything. It ends up being a huge effort in terms of team size.

At the end of it all, it's like, hey, I make a 75% gross margin as Nvidia. AMD sells its GPUs for a 50% gross margin, and they have a hard time out-engineering Nvidia—and they're great at engineering, right? But they still take more silicon area and more memory to achieve the same performance, and they have to sell for less, so their margin gets compressed.

Erik Torenberg

That makes sense. But look, historically, if you look at it, typically, if new entrants in markets didn't win by marginally improving on something existing—that happens sometimes—but more likely, they jumped up some kind of disruptive technology leap, where it's like, we have a different approach, we have different technology. Is that possible here?

I mean, to some degree? Maybe this is simplifying it a little bit, but I think part of the reason why the transformer model won was because it runs so incredibly well on GPUs. A recurrent neural network is similarly performant, but it runs terribly on a GPU. So, did we sort of pick the model for an architecture? And now it's hard to come up with an architecture that—

Guido Appenzeller

Well, it's hardware-software co-design, right? There's all this hype about neuromorphic computing, right? Theoretically, it's amazing and super-efficient. It's like, okay, great, but there's no ecosystem of hardware and there's no ecosystem of software. It would take tens of thousands of people who are the best at AI today focusing on that to even prove out whether it's worthwhile or not, right? On the hardware side, on the software side, and on the model side.

And so you look at Groq, Cerebras, and SambaNova. They all sort of over-indexed on the models that were leading at the time when they designed their chips, and so they made certain trade-offs, right? They put a lot more memory on-chip, and Nvidia was like, well, we're not going to do that.

Erik Torenberg

A lot faster, at least, right?

Dylan Patel

Well, more like—if you compare the amount of SRAM on NVIDIA's chips, it's much, much lower.

Erik Torenberg

Yes, correct. They went SRAM instead of DRAM, but then they usually have less DRAM. So there's a trade-off there as well.

Dylan Patel

Right? There's less DRAM, there's more SRAM, and because there's more SRAM on the chip, you have to have less compute on the chip. And so they ended up losing, right? Because the model sizes got too big and all this, right?

Guido Appenzeller

And so you have this super-weird dynamic where they bet on something that was actually better, right? I have no doubt that Cerebras would run certain types of models better than NVIDIA or Groq or, hey, Dojo, right? Dojo runs certain types of models way better than NVIDIA's chips because they're optimized for that.

But then it's like, oh, well, actually, even in vision tasks, people use vision transformers now. So it's like, okay, cool. Model sizes grew and all these things. It ends up being a catch-22 in that you optimize for something, and so now today you have this new age of AI accelerator companies that are like, okay, we're going to optimize for transformers.

But by the time they started designing, they're like, okay, transformers are dense models that are this big. What's the best—you know, the hidden dimension is 8K, and your batch sizes are this big and your sequence lengths are this big, so let's just make a super-large systolic array so you can create the maximum efficiency.

And then it turns out—oh, look at DeepSeek, or go look at what the labs are doing. Actually, their shapes are much smaller. You need to do a bunch of small matrix multiplies, not massive, massive, massive single matrix multiplies per layer. And then it ends up being, oh, well, that chip you're designing for that is actually not super-effective for that.

Erik Torenberg

There's still the software angle, right? NVIDIA has fantastic—

Dylan Patel

Yeah, and then there's software as well, right? But NVIDIA's going to have better networking than you. They're going to have better HBM, they're going to have a better process node, they're going to come to market faster, and they're going to be able to ramp faster. They're going to have better negotiations with whether it's TSMC or SK hynix on the memory and silicon side, or all the rack people, or copper cables—everything. They're going to have better cost efficiency, so you have to be 5× better.

Erin Price-Wright

But to be fair, if somebody had a viable competitor that was even marginally cost-competitive, my guess is many of the big consumers of GPUs would immediately shift some revenue there just to have a number 2, right? Just to—

Erik Torenberg

There's still pretty limited traction, though, right?

Dylan Patel

Sure. Meta continues to buy from them, and Microsoft did buy a bunch, and then they stopped because it's like, well, yes, AMD is giving you all these advantages, but it still ends up not being better on a performance-per-watt basis. They have a way bigger software team, and they're somewhat competitive on all these dynamics that I mentioned, right?

So you can't just do the same thing as NVIDIA and do it better, or try to execute better. Like AMD, you have to really leap forward in some other way. But the design cycle takes so long that models will shift, right? Because they're like, okay, what does the next generation of TPU and GPU look like? Okay, let's optimize for that. And the research path is, great, yes, neuromorphic computing could be the most optimal thing for us to do.

But no one's working on that because you have to advance in the tech tree. You've chosen it, right? If you restart the tech tree, you're going to be like, "Well, this sucks." And so, if it branches this way and you're over here, you're screwed, because you have to be 5x. There's a moat, because the supply chain stuff means that 5x actually turns into a 2.5x, and then NVIDIA can compress its margin a little bit if you're actually competitive. Then that 2.5x becomes like 50% better, and so it ends up being way too difficult.

Guido Appenzeller

And defense supply chain, for sure.

Dylan Patel

Yeah, defense supply chain. And then they get that right. Lutnick himself said, "We had to do this for rare earth minerals," and it's interesting. There are provinces in China with rules saying that the H20 is not efficient enough to be deployed, which is super bizarre because it's clearly the best AI chip China has. Huawei is still a little bit behind.

Guido Appenzeller

Well, what's interesting is that efficiency is so much less of an issue in China than here because they have the power infrastructure to support it. So even if they're running less powerful chips, you would imagine that it doesn't really matter, because China has such an infinite supply of power that they'd sort of be okay with it. So it's interesting.

Dylan Patel

Which is a big challenge in America, right? There have been companies that said—Jensen keeps saying he couldn't give away H20 in America for free, but I've literally heard companies say, "No, I mean, I wouldn't, because I only have this much power. How am I going to power data centers ready to go over the next year if I bought an H20? I'd literally have less compute capacity, and then I'd lose." Even if it was free, it doesn't make sense. Whereas China doesn't care. They can build these data centers; they have the muscle.

I'm curious how this all shakes out. China's posturing really hard. They even put out something that was like, "We're investigating to see if there are back doors in the H20." It's like, "There's no back door in the H20. Chill." GPUs are usually firewalled from the public internet anyway. You step through stuff before you get to the GPU clusters, so a backdoor wouldn't even matter. I don't know. I think it'll be interesting to see, because China can definitely deploy way, way, way more power to AI the moment they decide to, but there are competing interests, right?

Guido Appenzeller

Because they want Huawei to be better than NVIDIA.

Dylan Patel

Yeah. And this is how NVIDIA argued to the administration: if we don't do this, I think it's a very powerful argument. For example, within Triton, which is a common ML library, ByteDance has open-sourced some stuff that plugs into it that's super awesome. There are all these other libraries—not just models—that Chinese companies open-source for NVIDIA.

In a sense, NVIDIA's argument was that by selling GPUs, they were able to stop Huawei from building up a software ecosystem, and the Western ecosystem is better. But on the flip side, if you believe the models deliver more economic value to society than the hardware—which I actually think they do; it's just that there's a value-capture problem today—then you're giving China way more by giving them H20s and soon a cut-down version of Blackwell, as Trump said, versus selling them the chips. The economic value derived from selling them the chips is not as large as being able to somehow sell them AI services.

Guido Appenzeller

So, is China gatekeeping power for AI?

Dylan Patel

I don't think so. Again, what we see is that even with H20s being sold into China, as well as future versions of the chip—the H20E and other chips—we still see Chinese companies like Alibaba renting GPUs outside of China because the GPUs they can get outside of China are so much better on a dollar-spend performance basis. They're renting them or even going through a Singaporean company that's effectively a Chinese company, building data centers, and putting chips in them.

I don't think China is limiting the power per se. It's that Chinese companies are growing their capex way more than US companies on a percentage basis next year. The absolute dollar number is obviously still higher for US companies spending on AI. But on a percentage basis, Chinese companies are growing more next year, and you still have the problem that dollar spend to AI output—in tokens or whatever—is going to be lower because these chips are worse.

Power is not the gating factor. It's always capital, at least today. China can spend a lot more capital if it wanted to. It's subsidizing the semiconductor industry to the tune of $150 billion to $200 billion a year through SOEs and through capex that's not generating revenue, et cetera. So it's not like China couldn't do this to the AI ecosystem, right?

Given that Meta's capex is like $60 billion and Google's capex is like $80 billion, China could totally spend way more than that on a single effort. They just haven't decided to. And I think for the US, our build-outs are constrained by power, right? Google has a ton of TPUs sitting and waiting for data centers to be powered and ready, as does Meta with GPUs. We posted about how Meta is now building these effectively as tents.

Guido Appenzeller

Isn't this to some degree also coupled to their unwillingness to sell them to a broader ecosystem? I mean, if they want to be confined to their own data centers and they didn't rent data-center build-out for their own hyperscale use cases quickly enough, then yes, that constrains them, right? If they were on the open market, would we still be constrained?

Dylan Patel

Yeah, yeah, for sure. Because companies like CoreWeave—why is CoreWeave valuable? It's really because they build infrastructure really fast. Their software is nice, I think, but a lot of their customers are bare metal: just replace the GPUs whenever they're broken and network aggressively, and I think they'll go anywhere.

Guido Appenzeller

Jensen likes them as well.

Dylan Patel

Yeah, they'll go anywhere. Yeah, that's very important as well. But they'll go because it deconcentrates the ecosystem, which is better for NVIDIA. Having worked at Intel, I know exactly what's going through his mind.

Guido Appenzeller

Yeah. So, I think what's really important is that CoreWeave doesn't care, right? They're like, "Oh, crypto data center. I will convert it to an AI data center." They bought a company for like $10 billion that was doing crypto mining, which was worth like $2 billion a couple of years ago. And it's not because their Bitcoin-mining business is growing. It's because they have powered data centers, right?

Anywhere and everywhere, people are trying to build powered data centers. Companies like CoreWeave and Oracle are moving to that. Actually, today Google just bought 8% of a crypto-mining company called TeraWulf, right?

Dylan Patel

Not because they're getting into crypto mining.

Guido Appenzeller

No, because they need the data. They want the power.

Dylan Patel

They need the power, right? And all the hyperscalers have said, "Screw it," to their sustainability pledges because they need power as fast as possible. They're doing things that take a little bit longer to move the ship, but even if you didn't do it in your own self-built data centers, there are still a lot of challenges in the open market.

There's a deficit, right? And that's constraining American AI build-outs heavily. Others could maybe do it a little bit faster, like CoreWeave or others. Oracle has an open mind as well, but it's still constraining US build-outs heavily. Even though the capital's been spent, the chips are 60% to 80% of the cost of the cluster, depending on what chips you're getting. They've already bought the chips; they just can't put them anywhere because the data centers aren't ready. That applies to Google, Microsoft, Meta, and a lot of folks.

Guido Appenzeller

I mean, it's really hard to build infrastructure in the US. Power-grid interconnections, transmission, substations—all of this stuff, including electrical contractors and electricians. In Texas, if you're willing to be a traveling electrician, it's oil pay, right? It used to be that if you're physically adept, you could go make $100,000 in West Texas, but who the fuck wants to do that?

Now it's like, well, you could go 200 miles away from Dallas, to what's still a reasonable town, and build a data center and work on the wiring within the data center and all this other stuff—the transmission stuff—and your pay is up 2x now versus what it was just a few years ago.

This labor problem is a challenge too. I think in China they don’t have any of these problems, but they just haven’t spent the capital yet. Capital is an issue as well, because of the scale of what’s being spent, right? NVIDIA’s revenue this year is going to be over $200 billion, and next year it expects over $300 billion. Plus, Google is going to spend around $50 billion on TPU data centers, and Amazon is going to spend tons and tons on Trainium data centers.

The scale of dollars is quickly growing to nation-state-level stuff. What’s more important is being able to decide to spend the dollars and what’s cost-effective. To some extent, China is still constrained by that, but it can smuggle chips in. It can build data centers outside China and rent data centers outside China, and have the most cost-effective Blackwell chips or whatever, right?

ByteDance is either the biggest or the second-biggest customer of Google Cloud for a reason, right? They’re getting many, many Blackwell chips from them, right? The same is true with Oracle and Microsoft. All these other companies are renting tons of chips to China anyway, because it’s more cost-effective to do that than build it yourself.

It’s not like China has this mentality where it only has to—well, the government does, but the infrastructure companies don’t. Alibaba, Tencent, ByteDance, and so on don’t.

Erik Torenberg

So what’s the end game for data centers? We need more power; we need more cooling. Will, at the end, all data need to be next to a nuclear reactor or lots of solar, next to deep-sea water that we use for cooling or something like that? What’s the end game?

Guido Appenzeller

I think that cooling—the physical cooling of a data center—is not as significant as people think. There’s this whole narrative that AI uses so much power, and it’s not really true. Farming alfalfa uses 100 times the water of AI data centers; even by the end of the decade, it’ll be the same. Alfalfa is worth very little.

People have experimented with undersea data centers to reduce cooling costs, but that doesn’t make sense. It’s like 5% to 10% savings, but if you want to get the water out of the ocean, then put the data center into the ocean. If you want to service it, you’re screwed, right?

The same goes for power. We talk a lot about power, but it’s not actually that expensive. It’s just hard to build, right? It’s about getting to the right place, getting to the right space, and converting it down to the voltages and all the stuff that chips need.

Erik Torenberg

So it’s less the magnitude of power and more where it is and how it moves.

Guido Appenzeller

The magnitude too, right? In terms of total worldwide energy consumption, AI data centers are still a fraction of a percent. Even by the end of the decade, the US will have around 10% of its electricity going to data centers, which is still a small fraction of total energy consumption. In terms of energy, that’s an even smaller fraction, right?

Erik Torenberg

Oh, yeah, because shifting to electric vehicles could probably make a bigger swing than all the AI data centers we can build.

Guido Appenzeller

But outside the US—in Europe, for example—that number is not moving up that fast. In all these other countries, I think we need to build a lot more power, but it’s not some crazy amount. Doing it properly is the hard thing.

Again, the cost of power—if you look at the deals people are signing, even though the price has skyrocketed from a few cents a kilowatt-hour for these massive, massive purchases to around 10 cents, it’s still, when you think about the full TCO of the cluster, the GPU cost, the networking, and all of that stuff, far outstrips the power. The same goes for cooling.

Erik Torenberg

What percentage is power? If I do a 4-year amortized GPU data center, what percentage would be power?

Guido Appenzeller

About 80% of the cost of a GPU data center, if you’re building Blackwell, is capital, right? It’s the GPU purchases, the networking, the physical data center, the power-conversion equipment. All of this stuff is 80% of the cost.

Then 20% is your land, power, cooling, cooling towers, backup power, generators, and all this stuff. It’s like nothing, which is why it doesn’t matter if you spend 10% or 50% more on that.

At the end of the day, the expensive thing is the infrastructure itself. This is why what Elon did would seem silly: They spent a lot more money on generators outside the data center and mobile chillers to cool the water for their liquid cooling instead of the more cost-effective option, because it got the data center up 3 months faster.

That 3 months of additional training time is worth way, way more on a TCO basis, right? The performance you got out of the chips, the time to market, and all of this is way, way faster. Therefore, it was the right decision, even though this part of the data center ballooned in cost.

Everything else is still there, and you’re still paying for the chips. If they were sitting idle, it’s not worth it, right?

Erik Torenberg

Just by bypassing the grid, bypassing anything to do with interconnect, anything to do with public utilities.

Guido Appenzeller

Exactly. Exactly.

Erik Torenberg

What’s your take on Intel? Where’s Intel going?

Guido Appenzeller

I think the world—well, the US needs Intel. The world needs Intel because Samsung is doing worse than Intel on leading-edge process development, in my opinion. Based on various customers in the industry having done test chips at Intel versus Samsung, I think the industry generally agrees that Intel is further along in 2-nanometer-class process technology than Samsung, but both are way behind TSMC. TSMC is a monopoly, to some extent.

The number-one question people ask is, why is TSMC not making more money? Why are they only raising prices 3% to 10% next year, depending on what it is? TSMC is a monopoly. They could raise prices a lot more, but they’re good Taiwanese people rather than dirty American capitalists.

If TSMC were owned or managed by Americans—I think most of the ownership is actually American, in terms of the stock market, on the New York Stock Exchange and all this—they would have raised prices a lot more. There’s this difficult thing to be done: There’s 1 island that controls all leading-edge semiconductors, and not just all leading-edge—the majority of trailing-edge production as well. Something needs to be done. Intel is behind, but not absurdly so. If something were to happen to Taiwan, Intel would have the most advanced technology in the world. It's just not economic.

Erik Torenberg

Can you keep Intel as 1 company if you want it to be competitive?

Guido Appenzeller

I think the process of splitting it would take so much executive time and effort that Intel would have gone bankrupt by then, right? That’s the big challenge. I think Intel should be separate, right? But properly splitting the company, with all the management time that’s needed, is absurd.

Instead, what you need is Lip-Bu Tan, the CEO of Intel. There’s a lot of drama going around about him because he’s one of the greatest semiconductor investors ever, right? He’s invested in so many different companies. First, he was on the board of SMIC, which is effectively China’s TSMC, and that’s a big source of drama. He was the first investor in some of the biggest tool companies in China because it was a multipolar world there and he was making good investments.

Now people are getting mad about that, but he recognizes the companies. He understands the supply chain. He needs not to spend his time splitting the company, because then he never actually fixes the company, right?

Intel’s problem is that it takes 5 to 6 years to go from design to shipping the product. In some cases, more. When they tape out a chip, you send the design to the fab, the fab brings back the chip, and they go through 14 revisions in some cases. The rest of the industry goes through 1 to 3 revisions if they’re good: Send the design in, get the chip back, test it, send the design in again. For a public launch, they’ll launch a chip in 3 years.

Erik Torenberg

But if you look at Intel today, they still don’t have a competitive entry on the AI side, and they will, right? So what does that mean for their offering? They’re still doing great on CPUs. They don’t have a good AI chip product. Is that long-term sustainable positioning as a standalone chip company?

IBM still makes more money off every mainframe launch. So it’s not like x86 is dead. You don’t get the growth rates, but you could totally run this as a very profitable enterprise.

The same goes for PCs. There’s some turmoil, some ARM entry, and some AMD competition, but I think it can be a very profitable business if it had 1/3 the people or half the people working on it.

Dylan Patel

And so, Lip-Bu Tan, to fix Intel, needs to go into the design side of the company and lay off a shitload of people, but keep all the good people. He needs to make sure they’re designing fast and launching fast—that from design conception to launch is 2 to 3 years, not 5 to 6. That’s on the design side, and he needs to make that profitable.

On the fabs, you have to do the same thing. There are all these people. One of the heads of fab automation at Intel—I explicitly told Lip-Bu Tan, because we have a couple of ex-Intel people who are actually good in the company who worked on the fab side. We were like, “Who’s the worst?” and they were like, “Oh, this guy sucks.”

I explicitly told Lip-Bu Tan. He had never talked to the guy because he was 4 layers down. The company has absurd amounts of hierarchy. He goes and talks to the guy, and the guy’s out, right? He figures out who’s bad and who’s good. The vast majority of the team at Intel is the team that led the world in production and process technology for 20 years, right? But there’s a lot of built-up crap.

Erin Price-Wright

So, he has to go figure this out, right? He can’t waste his time on, “Oh, all this structuring to split.” I think it would be better if the company split; I just don’t think he can spend the time to do that.

If the design side of the company isn’t really going to get into AI, you have to make some money there. But the fabs, I think, could truly become a competitor, but they’re going to go bankrupt by the time anything can happen. So, they have to figure out how to get capital.

Guido Appenzeller

I think the goals are completely correct. I think the big challenge, just reflecting back on my time there, is that right now, if you look at Intel, they essentially have software, chip design, and then the core manufacturing part. They have 3 very different cultures, and it’s very hard to get everything under 1 umbrella. I think that is the big challenge.

Dylan Patel

I think you could—you should—run the company separately, but you can’t physically separate them entity-wise because it’s going to take so long to sever all these things, and he doesn’t have time. Intel is literally going to go bankrupt if they don’t have a big cash infusion or lay off half the company. Some could argue you need to lay off 30% of the company anyway, but a lot of bad things happen if that happens.

They need to spend a lot more on building the next-generation fab. Even if they fix the fab, they don’t have money for that. There are a lot more important problems than physically separating the company, even though I think long term, yes, the fab has to be separate from the chip design and software part of the company. That’s going to make each company much more accountable and able to service its customers better. It’s just that’s going to take too long, and they’re going to go bankrupt by then.

Erik Torenberg

Awesome. I hope someone does something right. I hope, I pray someone does something right: You get a big capital infusion. No, no, no—the big hyperscalers have muscled in like, “Oh, okay, wait. If TSMC eventually grows its margin to 75% because of the monopoly, plus it takes in all this stuff like co-packaged optics and power delivery and all this, all of a sudden the cost is going to spike. So, we should actually just throw $5 billion at Intel each, right? Screw it.”

That could actually give Intel enough of a lifeline to potentially get to something and maybe be competitive. That’s the hope. Can we finish by finishing this game that we started when we gave Sam Altman advice? If Jensen was here, what advice would you have for him?

Dylan Patel

I think he has a massive, massive balance sheet. Jensen does. His free cash flow is ridiculous. The new Trump tax bill institutes something really incredible, which is that you can depreciate all of the GPU cluster costs in year 1. We put out a note about how the tax implications for Meta are $10 billion a year, and across each of the major hyperscalers, it’s massive.

NVIDIA is going to spend tons and tons of cash. They’re going to spend, like, tens of billions of dollars on taxes. Why don’t they get into the infrastructure game somehow? Now, this is obviously crazy because they’re buying GPUs—their own GPUs—and putting them in data centers, and they’re competing with their own customers. But they’re already doing that anyway because their customers are trying to make chips.

They should accelerate the data center ecosystem with investments, because we think we can have a very high degree of accuracy on what they’re going to do next year in terms of revenue. It’s just the number of data center watts that are being built. This is a harder thing to shift up and down. Right now, there’s a little bit of share difference between how much is TPU versus GPU, but you have to accelerate the infrastructure, and you need to spend all of this capital that you’re building, right?

Do you want to go the route of doing buybacks and dividends? Great—you’re a loser if you do that, right? You can make more money by reinvesting in building a bigger company that’s not just chips into the ecosystem or servers into the ecosystem, but actually controls the infrastructure end to end somehow. I think there’s something he could do there with this massive war chest and this massive balance sheet.

There’s a reason NVIDIA has done some buybacks and increased dividends, but the cash on its balance sheet keeps growing. They’re going to have north of $100 billion of cash on their balance sheet by the end of this year, I think. What are you going to do with that? I think there’s something in moving much more into the infrastructure layer that they could do if he really wants to be the king of the world, which I think he does.

Erik Torenberg

Sergey and Sundar.

Dylan Patel

I think they should open up the ecosystem around TPUs, right? Start selling them. Open up the software—open-source a lot more of the XLA software, because there’s OpenXLA and there’s XLA, but the vast majority is closed source. Really, really open up the ecosystem around that, and be a lot more aggressive.

They’re still not very aggressive on data centers. They’re not very aggressive on a lot of elements of the company. The TPU team’s next-generation designs are not very aggressive, partially because a lot of the TPU team has left to go to OpenAI—the best people that I knew. It was really annoying. I knew 4 or 5 people, and they all went to OpenAI, so now I don’t get as much in. I met some other people, right? But I think they could be a lot more aggressive in many ways across the company.

They don’t have to be, but they could. Because AI—take ChatGPT, for example—the shift of search queries, especially the monetizable ones, from Google to purchasing agents is going to really screw Google long term if they don’t get their act together. I think they’ve gotten their act together on DeepMind. There are still some inefficiencies, but Sergey works within DeepMind a lot, and they’re driving hard.

They’re still a little bit behind, but I think physical infrastructure, TPUs, how much money they could make, and how much they could take the wind out of everyone else’s sails if they start selling TPUs externally all matter. They should reorganize around building data centers much faster so that they once again have the most compute in the world, because they did. There are certain companies that are going to surpass them over the next few years if they don’t really get their act together. So, that’s what I would say for them. Also, learn how to ship products better, right?

Erik Torenberg

Zuck.

Dylan Patel

I think Zuck—it remains to be seen what goes on with superintelligence, but they’re trying to move super fast with the data centers. Screw it, they’ll build tents instead of physical data centers because they only need these for 5 years anyway. The superintelligence moves, you could say whatever you want, but trying to buy Thinking Machines Lab for $30 billion or SSI for $30 billion didn’t work out. So then they spent not even that much on hiring—not $30 billion on hiring all these people.

I think that he recognizes the urgency with the models and the infrastructure. I really think he needs to move faster. If you read his website post about AI, I think he sees the vision: There are the wearables, there’s integrating AI into those, and there’s being your AI assistant to do all this purchasing and stuff. I think he sees the vision, but I think he also needs to focus on actually releasing that faster.

The products that they do outside of their core IP—every time they launch something, it’s kind of mid, right? Meta’s Reality Labs is doing well, but I think they should go more explicit...

Have a ChatGPT competitor, have a Claude Code competitor—just start releasing way more products. They're really focused on their individual gardens rather than branching outside of them.

Erik Torenberg

Do you think Apple should have that same sense of urgency? If Tim Cook were here, what would you tell him?

Dylan Patel

The funny thing is, some of Apple's best AI people are now at Meta Superintelligence Labs. They're building an AI accelerator. They have AI models, but they're just way slower. They did mention on the last earnings call that they're going to allocate more capital to this, but it's like, guys, Apple, you're going to lose the boat if you don't spend like $50 billion on infrastructure.

Erik Torenberg

You don't think the current Siri will cut it?

Dylan Patel

I think more and more you'll see people say, “Great, Apple has this walled garden,” but they can only do so much to protect it, right? IDFA—they shut down ads, or data sharing, to Meta, but Meta made better models, and now they have way more data and way more power over the user than they ever did before. It was good that Meta kicked the crutch off of them—or Apple did.

The same applies to AI. Yes, they have access to the text and they have access to this, but I think other people are going to be able to integrate user data. Agents will be able to integrate all this user data, and they'll start to lose control of what the user experience is as more and more gets disintermediated by AI being the interface, rather than touch, rather than a touchpad and keyboard.

I don't think they've truly realized what happens when the interface to computing is AI. They market it, but that's going to shift computing really heavily. They have great hardware, and their hardware teams are working on awesome stuff and form factors, but I just don't know if they get what is actually going to happen to the world in the next 5 years truly well enough, and they're not building fast enough for it.

Erik Torenberg

What about Microsoft, to that end?

Dylan Patel

Microsoft has the same problem, I think. They were super aggressive in 2023 and 2024, and then they pulled back heavily, right? Now OpenAI is slipping through their grasp. There's that whole thing there. They cut back on data center investments heavily. They were going to be the largest infrastructure company in the world by a factor of 2x, which—you could argue maybe that was too much, and maybe it wouldn't have been economical—but they're losing their grasp on OpenAI.

Their internal model efforts are failing spectacularly. They're on LLM Arena right now, and they're pretty decent there, but that's just a synthetic model. It's a code name, but whatever. MAI is failing. Azure is losing a lot of share to Oracle, CoreWeave, and Google, and so on and so forth. Their internal chip effort is by far the worst of any hyperscaler. They're just mis-executed.

GitHub—how is GitHub not the highest-ARR software code model? They had the best IDE, the best source-code repository, the best enterprise sales force, the best relationship with a model company, and they were first to market. They had everything going for them.

Erik Torenberg

And there's just nothing, right? GitHub Copilot is failing, and Microsoft Copilot is still crap, right?

Dylan Patel

Yeah, it's unusable.

Erik Torenberg

What is going on? You need to shake the crap out of the company. I think they win a lot because they have the best business-to-business relationship with so many enterprises.

Dylan Patel

Best sales force on the planet.

Erik Torenberg

Yeah, but they end up not having the actual product to sell them, which is really scary. They need to really work on product.

Dylan Patel

Satya has done great on sales and stuff, but—yeah.

Erik Torenberg

If Elon were here, what advice would you give him?

Dylan Patel

A lot of people at xAI are mad about the porn models, the porn stuff. It's fine. You're going to make a ton of money off this. This is how you accelerate the revenue of that company.

But he's losing a lot of talent and axing a lot of good projects. Elon is a magnet for amazing talent and building stuff, so I won't bet against him. Since he left the administration and focused on stuff again, I don't know. I think he's focused on a lot of things.

I think the robotaxi is starting to look good again, actually. I haven't ridden one yet, but I have some friends who have ridden one. It looks pretty decent. He could not make these snap decisions, which often are the reason why he's amazing, but some of these snap decisions are hurting him. I'm not sure I can give Elon that much great advice.

Speaker 2

I think maybe it's just: focus on the products again, more. But he is working on that stuff a lot. Yeah.

Erik Torenberg

I think that might be a good place to wrap. Awesome. It was a great discussion, Dylan. Thanks so much for joining us.

Dylan Patel

Thank you for having me.

Erik Torenberg

Thank you.

Dylan Patel on GPT-5’s Router Moment, GPUs vs TPUs, Monetization | BidClub