[BidClub_]
All-In · · 98 min

Four CEOs on the Future of AI: CoreWeave, Perplexity, Mistral, and IREN

Michael IntratorAravind SrinivasArthur MenschDaniel Roberts

YouTube
TL;DR
  • CoreWeave’s central claim is that AI compute decommoditizes at cluster scale, while contractual cash flow lasts far longer than the market’s feared chip cycle. Michael Intrator called 16-month-or-so GPU obsolescence “nonsense”: the average contract is five years, CoreWeave depreciates equipment over six, expects useful life beyond that, and says A100 pricing appreciated during the year. Hardware becomes obsolete when data-center power can earn a higher margin elsewhere.

  • CoreWeave’s growth engine is structured finance, not uncontracted GPU speculation. Each discrete “box” contains a customer contract, GPUs and a data-center agreement; customer payments cover facilities, power, interest and principal before residual cash returns to CoreWeave. That structure supported $35 billion of financing in 18 months, pays for everything within 2½ years of a five-year deal and helped lower its cost of capital by 600 basis points.

  • Infrastructure demand remains above physical supply, but the bottleneck now spans memory, power, optics, networking and construction—not merely NVIDIA allocations. Intrator said four years of demand have “overwhelmed the global capacity of the world,” while IREN’s Daniel Roberts said there are “no idle GPUs in the world sitting in a data center.” Long-term contracts protect CoreWeave against an air pocket; IREN’s position is 4.5 GW of power and an eight-year head start tying up land and power.

  • Perplexity is betting that value migrates from any single frontier model to the neutral conductor that selects and coordinates them. Aravind Srinivas called the company “Switzerland”: GPT, Gemini, Claude, Kimi, Nemotron and Qwen can specialize while Perplexity auto-routes work, or its Model Council explains where multiple models agree and disagree. The company serves several tens of millions of consumers, says enterprise revenue is growing faster, and reports positive gross margins on “every single penny” of revenue despite not yet being profitable overall.

  • Perplexity’s road map turns AI from an answer box into the computer and eventually the operating system. Personal Computer will use a Mac mini as a private local runtime while delegating permitted frontier-model calls or long tasks to a user-isolated server; Srinivas described it as “Open Claw for dummies.” The interface starts with objectives rather than programs, while Linux, files, models, connectors and sandboxes sit beneath the orchestration layer.

  • The near-term disruption is bespoke software and autonomous back-office work, not a single all-purpose chatbot. Perplexity Computer already produced the company’s board memo, partnership deck and press briefing; Srinivas’s longer-term target is a small business whose AI runs ads, support, billing and feature development while its owner is “sipping wine in Napa.” He stressed that this future “is not there yet,” and conceded temporary job displacement even as he argued it could restore agency to people who dislike conventional jobs.

  • Mistral’s enterprise thesis combines open models, human expert signal and strict execution controls. Arthur Mensch said synthetic data can warm up and compress a smaller model, “but eventually you do need to have human signal”; Mistral therefore deploys portable training tools and PhD-level engineers inside customer infrastructure so data need not flow back. For production agents, OpenClaw-style autonomy is insufficient without deterministic gates, sandboxes, role-based controls and a “context engine” that prevents compensation or other restricted data from leaking across an organization.

  • IREN’s legacy power portfolio has become its scarce AI asset, with a $9.7 billion Microsoft contract consuming only 5% of capacity. Its 750 MW Texas flagship sits inside a 4.5 GW portfolio, while its operations have used 100% renewable energy since inception—British Columbia hydro and West Texas wind and solar. Roberts’s demand model is Jevons paradox: if 10 times more compute cuts image generation from minutes to 5-10 seconds, users will generate many more images, not bank the efficiency savings.

Digest · the substance, structured for research

1. CoreWeave bought compute optionality before it sold AI cloud

  • Intrator traced CoreWeave to a natural-gas algorithmic hedge fund whose team became interested in crypto during downtime. They rejected Bitcoin ASICs because the chip designer likely had the operating advantage, but GPUs could mine Ethereum and serve other workloads—so from the beginning they treated “the compute as an option.”

  • Founded in 2017, the company spent roughly three years mining while its hedge-fund risk discipline helped it survive multiple crypto winters. It then climbed the workload stack from CGI rendering to batch computing and medical research, looking for steadier uses of inherently volatile crypto demand.

  • The decisive tuition payment came in 2020-21: CoreWeave bought A100s and donated capacity to EleutherAI while learning neural-network infrastructure. Free compute meant the researchers “can’t really get pissed at us if we’re not very good at it initially”; when the volunteers returned to their day jobs, they asked for the same infrastructure, creating CoreWeave’s commercial opening.

2. At cluster scale, compute stops being a commodity

  • The scaling laws changed CoreWeave’s framing before the ChatGPT moment. Intrator’s distinction was blunt: “Anybody can run a GPU, but can you run a cluster that’s large enough to train a model that can change the world?” In his formulation, “compute decommoditizes at scale.”

  • CoreWeave deliberately occupies the layer “above the NVIDIA GPUs but below the models,” integrating software, operations and observability for one purpose-built workload. Intrator contrasted that focus with AWS: conventional cloud solved one problem well, while accelerated AI required a new solution.

  • EleutherAI was followed by Inflection, working with Mustafa, and then hyperscalers, OpenAI and other foundation-model customers. Training created the initial business; productized AI is now pushing inference from the organizational fringe into core workflows.

3. Inference monetizes AI—and extends every GPU generation

  • Intrator called inference “the monetization of the investment in artificial intelligence”: asking for an answer or asking the model to act is where model capability crosses into real-world value. He therefore watches inference volume on CoreWeave’s infrastructure as a direct gauge of the ecosystem’s health.

  • CoreWeave says it led scaled commercial deployment of H100s, H200s and GB200s and identifies GB300s as the next architecture. The lifecycle is not train once and discard: bleeding-edge chips train models, move into experiments, then supply inference “for a very, very long time.”

  • The host’s pushback surfaced the public bear case that rapid architectural progress must shorten asset life. Intrator rejected the premise: the market’s growing demand for the newest chips does not eliminate the long tail of experiments, rendering, smaller models and inference that older architectures can serve.

4. Five-year contracts puncture the short GPU-life thesis

  • Intrator called the depreciation debate “nonsense” promoted by traders with short positions. CoreWeave customers contract for five or six years, with an average of five; claims that the equipment becomes obsolete in 16 months simply “don’t in any way match up with the facts on the ground.”

  • CoreWeave uses six-year depreciation and expects GPUs to operate beyond six years. More strikingly, Intrator said A100/Ampere pricing appreciated during the year as new companies and workflows absorbed capacity that could not justify or obtain the latest H100-class systems.

  • The host’s iPhone analogy carried the resale logic: an iPhone 12 may look dated to its original owner yet retain substantial value in another market. Likewise, an older GPU can migrate to a less bleeding-edge customer, a different data center or another geography rather than disappearing economically.

  • Intrator’s actual obsolescence test is opportunity cost: the moment arrives when a data center’s constrained power can generate a higher margin with replacement infrastructure. Until then, older infrastructure remains “extraordinarily profitable,” and CoreWeave was still selling available older capacity at a higher price than the prior year.

5. The “box” turns customer paper into cheaper infrastructure debt

  • CoreWeave begins with a creditworthy customer contract—Intrator used Microsoft as the example—then places that contract, the purchased GPUs and the data-center agreement into a financing structure called “the box.” The customer pays the box rather than CoreWeave directly.

  • Its cash waterfall first pays the data center and power bill, then interest and principal; only the residual returns to CoreWeave. Lenders therefore underwrite contracted customer paper and physical collateral, not a generalized promise that future spot demand will remain high.

  • Intrator said this structure let a relatively little-known company raise $35 billion in 18 months. Within 2½ years of a five-year deal, the box has paid for everything, including principal and interest, while already generating a return for CoreWeave.

  • Each box is discrete, limiting contagion if one transaction fails. Repeated performance lowered CoreWeave’s cost of capital by 600 basis points in two years, but requires discipline: Intrator will reject a one-year GPU request because it is too short to amortize the build.

6. Scarcity has spread from GPUs into the entire supply chain

  • Intrator said demand has been “relentless” for four years and continues to overwhelm global compute capacity. Yet CoreWeave still models an air pocket caused by technology, war or another shock; the reason is irrelevant to risk management, so five-year contracts and strong counterparties are the protection.

  • GPUs are only one throttle. Power shelves, memory, storage, networking and optics can each cap deployment, and Intrator identified memory as the current bottleneck—the collision of surging AI demand with fab investment that probably needed to begin in 2023.

  • The host framed the resulting boom-bust cycle as both destructive and generative: it “clears out the underbrush,” rewards survivors and leaves infrastructure behind. He linked fiber overbuild to YouTube, whose free uploads benefited from storage and bandwidth becoming radically cheaper.

  • The host also cited OpenAI CFO Sarah Friar’s comparison: a million tokens cost $32 and change when GPT-3 arrived versus 9 cents now. Intrator agreed that falling unit costs show how capital markets, capitalism and engineering fuel competition, lowering the barrier that kept human creativity “contained” and giving potentially eight billion people a tool to build from an idea.

7. Perplexity is turning answer accuracy into a computer

  • Srinivas presented Perplexity Ask, Comet and Computer as one trust progression. Internet access improved answer accuracy; full browser access improved task accuracy; full computer access now lets the AI perform the same digital work the user would perform manually.

  • His signature metaphor was an orchestra: hundreds of specialized models are instruments, sub-agents are musicians and Perplexity is the conductor. Coding, writing, image, video and audio capabilities matter only insofar as they combine into “the music you play”—the completed work.

  • Personal Computer will synchronize Perplexity with a Mac mini acting as a local server. Private-data orchestration can remain on that hardware, while permitted frontier-model calls or complicated long-running jobs can move to a server-side computer accessible only to that user.

  • Srinivas called the packaged experience “Open Claw for dummies”: one executable, no API-key management, no separate billing across roughly 100 services and no manual permission maze. The strategic choice is end-to-end integration rather than exposing every component.

8. Local AI becomes an appliance, while AI becomes the operating system

  • The host reported running Kimi 2.5 on a Mac Studio and getting perhaps 80% of frontier capability locally for free; he also pointed to a Dell-NVIDIA workstation with 750 GB of RAM. His economic question was whether a $10,000 desktop could replace part of a $500 monthly cloud bill while improving privacy.

  • Srinivas expects local models to begin as privacy-sensitive sub-agents for tax returns, photos, emails, calendars and notes. He rejected a fully local-versus-cloud dichotomy: Google Workspace data already resides on servers, while phone users mostly care that the task completes, not which permitted machine ran.

  • Over time, he expects the home AI server to feel like “buying a refrigerator” or an internet modem and to orchestrate household sensors. Its operating system begins with objectives rather than instructions; Linux, files, models, connectors and code sandboxes sit underneath an AI abstraction that decides how to execute.

9. Model specialization is Perplexity’s Switzerland moat

  • Perplexity has several tens of millions of monthly users, while enterprise revenue is growing faster than consumer revenue. Enterprise Pro costs $40 per user monthly; Enterprise Max costs $400, with usage charged after included Computer credits, and Srinivas said Max customers have collectively saved more than $100 million.

  • His economic defense was unusually specific: every source of revenue has positive gross margin because Perplexity combines recurring subscriptions, model routing, retrieval and search rather than reselling maximum-context tokens indiscriminately. The company itself, however, is “still not yet profitable.”

  • Against acquisition speculation and rivals with far larger capital budgets, the roughly 400-person company’s answer is neutrality: “We’re like Switzerland.” Kimi, Nemotron and Qwen can join GPT, Gemini or Claude under the hood, and a winner in any model family does not strand Perplexity’s application layer.

  • Specialization strengthens that position. Srinivas said Perplexity’s iOS engineers love using Codex while back-end engineers love Claude Code; the product auto-routes each prompt but preserves manual selection. Model Council goes further, comparing several answers and identifying exact agreements, disagreements and nuances.

10. Faster shipping points toward software made on demand

  • Srinivas called speed Perplexity’s mode, provided it can preserve quality and trust. Inside Perplexity, even non-engineers now ask the Computer Slack bot to fix bugs.

  • Computer prepared Perplexity’s board memo, one-shot a partnership deck and replaced a communications briefing that previously required staff work. The product’s improving memory compounds that leverage because it can draw on prior meetings, decks and relationships rather than rebuilding context each time.

  • Srinivas credited sandboxes, terminals, files, sub-agents, skills and command-line tools with making long orchestration reliable. The model can load only the context it needs and discard it later, so context-window limits stop being the defining constraint.

  • Srinivas said he had an agent download every All-In episode, extract public-company mentions, chart frequency over time, analyze stock-price impact and sentiment, and provide timestamps linking back to the relevant audio. The host separately described using an agent—after becoming “Claude-pilled”—that offered to build a bespoke CRM rather than merely export a spreadsheet.

11. Autonomous businesses are plausible, but not yet turnkey

  • Srinivas distinguished the slogan of a one-person billion-dollar company from actual new GDP: replacing researchers with AI is not the same as creating a billion dollars of value. His target is smaller and more concrete—helping someone buy a Mac mini and operate a business producing hundreds of thousands or millions annually.

  • The envisioned Computer would run Instagram and Google ads, integrate SEM and SEO tools, find users, charge them through Stripe, ship features and handle support through Intercom while the owner is “sipping wine in Napa.” His caveat was explicit: “Everybody thinks AI is already there. It’s not there yet.”

  • Srinivas put his own outlook at 70-80% highly positive and 20% worried about rapid job losses. He accepted temporary displacement but argued that many people dislike conventional jobs; the desired end state is agency, ownership and small-business creation, provided individuals remain active, resilient and resourceful.

12. Paid browser access is the missing agency layer

  • Computer is already available inside the Perplexity app, while Comet for iOS brings the browser experience to mobile. The browser remains strategically distinct: until websites are organized around APIs and command-line tools, agents must still open tabs, complete forms, click controls and upload files. Native browser control lets Perplexity automate that residual web work.

  • The host pressed on websites such as Reddit and LinkedIn becoming resistant to automated activity and proposed paid, authenticated access. His model was bounded behavior—summarize subscribed content or locate seven employees, with explicit quotas and no posting or voting—rather than unlimited scraping.

  • Srinivas would not discuss individual negotiations but endorsed official APIs and a “win-win” based on user choice. The host extended the proposal to paid news subscriptions: agent access could become incremental revenue for publishers and make subscriptions more useful and sticky.

13. Mistral uses open models to bring enterprise IP inside the model

  • Mensch announced that Mistral would train its next frontier generation with NVIDIA, extending work begun with Mistral NeMo roughly 18 months earlier. The goal is to produce leading open-source models that Forge can specialize for engineering, physics, science, financial services and government languages.

  • Mistral is European but not confined to Europe: Mensch said 25% of its business and 25% of its researchers are in the US, alongside teams in France, the UK and Singapore. Europe offers more manufacturing exposure and companies seeking to leap forward after lagging in adoption.

  • General-purpose models remain necessary for orchestration, but enterprises possess decades of IP and physical-system signals that closed models cannot absorb deeply or transparently. Open models permit new parameters, bespoke harnesses and deployment on any cloud, customer hardware or the edge.

  • Forge combines model customization with forward-deployed engineering, while Studio builds end-to-end agents. Mensch’s value proposition is lower cost and greater control, particularly for companies such as ASML whose proprietary data can make a model uniquely capable in a critical industrial process.

14. Human signal and deterministic controls separate enterprise agents from demos

  • Mistral’s platform is portable into a customer’s own infrastructure, so “there’s no data flow coming back to Mistral.” PhD-level engineers work with subject-matter experts on tasks such as image scanning and defect detection, then transfer enough knowledge for the customer to retrain without permanent Mistral involvement.

  • Synthetic data is useful for warming up and compressing models: a large teacher can efficiently generate training material for a smaller model. Mensch’s limit was firm—“eventually you do need to have human signal,” even though expert feedback is expensive to acquire.

  • OpenClaw demonstrated individual autonomy, but Mensch said it lacks enterprise primitives. An HSBC KYC process needs deterministic gates that always execute, remain observable and can be guaranteed to management; production therefore requires a control plane, sandboxes and rigorously enforced access controls.

  • The host’s compensation-data example exposed the risk of giving an agent root access across Gmail and Slack. Mensch’s answer was a “context engine” mapping data and metadata to roles, denying an engineer’s compensation query when unauthorized; once those information flows change, management layers and customer-service organizations may also need redesign.

15. IREN used Bitcoin to finance its pivot into AI

  • Roberts said IREN’s founding thesis was that exponential digital demand would eventually collide with the physical world. Bitcoin mining was the first cash-generating workload used to bootstrap land, power and data centers until “higher and better use cases” emerged.

  • A 2020 Dell memorandum of understanding was a false dawn: it was too early to bring on customers and compute, so IREN returned to Bitcoin. Demand became tangible roughly two years before the interview and has escalated monthly; the company is now “swapping out all the Bitcoin for AI chips.”

  • Microsoft signed a $9.7 billion contract late in the prior year, yet Roberts said it consumes only 5% of IREN’s capacity. The company develops sites itself—land, permits, grid connections and construction—with a 750 MW flagship in Texas and 4.5 GW overall, nearly the Bay Area’s annual power use by his comparison.

  • Because IREN began tying up land and power eight years ago, Roberts said its constraint is now “time to compute.” Construction labor, cooling, memory and supply chains create “permanent whack-a-mole” as thousands of people try to convert remote physical infrastructure into live clusters.

16. Power geography and Jevons paradox define IREN’s next build cycle

  • IREN says it has used 100% renewable energy since inception: hydro in British Columbia and wind and solar in West Texas. The West Texas opportunity is geographic—45-50 GW of renewable generation against only 12 GW of transmission capacity to Dallas and Houston—so compute goes to excess power and leaves as a digital commodity.

  • The host expected batteries to smooth intermittent supply, but Roberts said IREN does not need them because the utility guarantees 24/7 reliable power once a scarce grid connection is secured. Locally, IREN hires outward in 20-mile increments, has awarded $1 million in cumulative community grants and is partnering with universities and trade colleges; Roberts said the lower end of the host’s $150,000-$300,000 trades estimate was directionally right.

  • Roberts called demand “gangbusters” and said, categorically, “There are no idle GPUs in the world sitting in a data center.” Efficiency reinforces consumption: if 10 times the compute reduces image generation from minutes to 5-10 seconds, users make more images—Jevons paradox, not demand destruction.

  • NVIDIA remains the safest near-term roadmap because of its ecosystem and standards, though custom silicon is beginning to seek data-center capacity. Nuclear could expand supply but likely needs a decade or longer; networking matters immediately, with 6 ms round-trip latency from IREN’s West Texas site to Dallas, while space-based data centers remain constrained by launch cost, radiation and difficult engineering.

Speaker 1

One of the great companies of the AI era is, of course, CoreWeave. They’re building massive infrastructure for these hyperscalers, and in some ways, Michael Intrator, welcome to the program. You’re the original hyperscaler. You guys got in very early and secured your—I don’t know which GPUs you wound up getting—but you were very early to this trend. How did you get to it so early, and how did you build out this first, I guess at the time, neocloud?

Michael Intrator

We didn’t really start it as a neocloud. I was running an algorithmic hedge fund focused on natural gas, and when you build an algorithmic hedge fund, once the algorithms are built, you’re really just monitoring them, testing different pieces, and doing all that. But there’s also a lot of downtime, and we got super interested in crypto.

We’re pretty nerdy. We dig under the hood, and we started to get interested in the security layer. We looked at Bitcoin and Bitcoin mining, and we didn’t like it. We thought there was some brilliant engineer who built the ASIC, and they were probably going to be better at running it than we were.

We really began to focus on GPUs, mostly because you could mine Ethereum with them, but you could also do all these other things. Right from the start, we looked at compute as an option to deploy our computing power to different use cases.

We began the company in 2017 and spent the first 3 years mining crypto. We went through a couple of crypto winters, and because we had come from a hedge fund, we had real chops in risk management and how we thought about capital, risk exposure, allocation, and all of that. We were really careful around that right from the start.

We weathered crypto winter really well and began to scale the company. We immediately started to look for other use cases for this compute because crypto was pretty volatile.

Speaker 1

Yeah, and crypto was a question mark at that time.

Michael Intrator

Absolutely. Bitcoin was speculative, and there were many other speculative projects. The only other people using this type of hardware were quants and medical researchers.

A good way to think about it is the progression of products that we started to work on. First was crypto, but we immediately moved from crypto to CGI rendering. We built projects that allowed folks who were trying to animate and render images—the things that make movies cool—to do that.

Then we moved to batch computing and started to look at medical research and different ways of using compute to drive science. We just kept moving up the stack in terms of the complexity of how GPUs could be used.

Ultimately, around 2020 or 2021, we started to figure out how GPUs could be used for neural networks. That wasn’t something we knew how to do, so we went out and bought a bunch of A100s and donated them to a group that was working on Luther AI. They were working on an open-source project, and the thought was that since these guys were taking the GPU compute as a donation, they couldn’t really get pissed at us if we weren’t very good at it initially.

Speaker 1

They couldn’t complain about the SLA.

Michael Intrator

They kept telling us, “We need more of this. You’ve got to work on this.” That began to give us an understanding of what was necessary to run scalable, parallelized computing.

I feel like buying those initial GPUs was the tuition we paid to learn how to run this business. One of the interesting things is that all of those guys went back to their day jobs because they were all volunteers working on this. They were like-minded scientists.

When they got to their day jobs, they were all like, “I want that infrastructure. It’s built the right way. That’s the way research is going to want to use it.” That launched our business.

Speaker 1

It was an amazing story. You went from crypto to these researchers, into academia and deep research. What’s the next card to turn over in the poker game?

Michael Intrator

What became very clear to us very early on was that the scaling laws were going to drive this. Remember, this was back in 2020 and 2021, before the ChatGPT moment occurred.

We began to understand that compute decommoditizes at scale. Anybody can run a GPU, but can you run a cluster that’s large enough to train a model that can change the world? That’s a different question.

We began to think about how to scale up our delivery of this computing to larger and larger clients. That was the next card to turn: thinking about how there was a component of this that would lean into our ability to access the capital necessary to deliver our solution to the broadest possible audience, to the most sophisticated consumers of this compute.

The next card was thinking about it as a business rather than as an engineering project: how to deliver the infrastructure and the software, and really everything in between. When you’re thinking about what we do, we live above the NVIDIA GPUs but below the models.

Everything in there—the software, the integration of software and operations, observability, and all the things you need to build a cloud that’s purpose-built for this one specific use case—is what we focus on. We don’t do everything. We really focus on one use case.

Speaker 1

Web servers are different. You’ve got AWS. They do a great job. It’s a great solution. It was a brilliant solution to solve a problem.

Michael Intrator

We just looked at it and said, “There’s a new problem. Let’s look at this problem and try to come up with a solution to deliver compute that solves it.”

Speaker 1

When did the language models start dialing and calling you for capacity?

Michael Intrator

Our first language model was really EleutherAI. Our first large commercial customer was Inflection. We worked with Mustafa and Inflection, and then we diversified from there into the hyperscalers and OpenAI, across the foundation-model landscape.

We just kept scaling, with the belief that the decommoditization of compute and the ability to deliver a solution were going to matter. The solution is building supercomputers that can change the world. That’s really what we began to focus on.

That led into training, and now the world has gone through this moment where we’ve moved from research into the productization of this. It’s beginning to work its way in from the fringe of organizations into the core of what they do.

You can see that every day in the amount of inference compute being driven through our infrastructure layer, which is massive.

Speaker 1

They’re consuming it, not just building models, but deploying and utilizing them.

I always think of inference as the monetization of the investment in artificial intelligence.

Michael Intrator

When we see our compute being used to stand up the massive scale of inference hitting our compute every day, inference is when people ask the model a question and it comes back with an answer. That’s an inference. Or when you ask the model a question and then ask it to go do something, that’s inference.

That’s where you have the opportunity to really drive value outside of the model itself and into the real world. That’s exciting for us. That’s what we like to watch, and what I like to watch in terms of gauging the health of the business.

Speaker 1

What chips are those?

Michael Intrator

We are the tip of the spear in bringing the new architectures out of NVIDIA into commercial production at scale. We were the first ones to bring the H100s at scale. We were the first ones to bring the H200s at scale, the first ones with the GB200s, and now you’ve got the GB300s.

One of the things that’s amazing and really fascinating for us is that people are using the bleeding-edge GPUs to train models as the new architectures come out. Then they take those GPUs and move them into different experiments. Over time, they move them into inference, and they continue to use them in inference for a very, very long time.

Speaker 1

What is the shelf life of an H100 right now? That’s been a big debate, I think, for your company and for Microsoft. I guess Michael Berry—you must have known him when you were a quant—has been saying, “Oh, my God, the whole industry is falling. The sky is falling.”

We all know in the industry that people don’t just throw this hardware away. They find uses for it. The market finds its own use for technology. So, what’s the reality of the lifespan of these things?

Michael Intrator

My take on the GPU depreciation debate is that it’s nonsense. It’s a debate being brought to the forefront by some traders who have a short position in the stock, and they’re trying to talk it down.

Look, here's what we know. When we buy infrastructure, we're a success-based company, right? We're a small company on a relative basis compared to the enormous companies that we're competing with. Our clients come to us and buy compute for 5 years, for 6 years. Our average contract is 5 years. So any commentary by anyone, either inside or outside of the industry, that this stuff becomes obsolete in 16 months or whatever nonsense they're spewing, it doesn't in any way match up with the facts on the ground. The fact on the ground is they're buying it for 5 years, right? My approach to this has always been: if people are willing to pay me for it, it still has value.

Speaker 1

Correct.

Michael Intrator

Pretty simple way of approaching it. We use a 6-year depreciation. We believe that the GPUs will last in excess of 6 years, but we felt like that was a fair and reasonable approach to a technology cycle that's moving at this velocity. The A100s, the Ampere GPUs—this year, the price has appreciated through the year. Now, why is that? I think it's because, as more installed capacity becomes available, you have new companies that come into existence with new use cases and different-size models. They're trying to build new commercial ventures that maybe have been blocked out of the H100s and never had an opportunity to run on them.

Daniel Roberts

To make a very simple example for the audience, when you trade in your iPhone after 3 or 4 years, you're like, “Who's going to use an iPhone 12?” And it's like, “Have you been to South America or Africa, where you go to the store and buy an iPhone 12 or the Pixel 7, and it costs $50? That's still got great life left in it.”

Michael Intrator

Absolutely. Yeah. Well, and so, look, we find these amazing use cases: new companies that have come into existence or existing companies that have integrated new models into their workflow and are able to use the Ampere GPUs. And so they keep buying any GPUs that we have available. Once again, the concept that a GPU is no longer relevant or commercially viable after 16 or 18 months or 2 years—

Daniel Roberts

Yeah, as far as it goes, it just doesn't make any sense.

Michael Intrator

It goes as far as it goes. I think sometimes people get caught up in Moore's law or in just how fast our industry is growing, and that there's so much at stake that big companies are demanding the most recent products. That doesn't mean that the lifespan has gotten shorter; it means the opportunity and the surface area of the opportunity have gotten much larger.

The industry has gotten so much attention for the unprecedented scale of capital that's coming to bear on this. Because of that, there tends to be an incredible focus on the companies that are building on the most advanced chipsets. The truth of the matter is that even within those companies, they have a long tail of useful life to provide inference horsepower, work on other experiments, and do less bleeding-edge activity that still needs to be done. Rendering comes to mind as well. We're making images on Nano Banana. There will be a use for it.

There is a moment in time where maybe the compute-to-power ratio doesn't make sense. My expectation is that obsolescence will be defined by the moment in time when the power in the data center, for me, will be able to be repurposed for a higher margin than the existing infrastructure provides. As I said, I fully expect this infrastructure to last in excess of 6 years, but the standard in the space has really been 5 years, with 1 exception, which is Amazon, at 6 years. That seems like the right schedule. I'm not making it up. That's what everybody's using.

Speaker 1

The energy cost is the opportunity cost because we need that space. There's a better reward here, and that hardware might get resold to somebody else who wants it—a hobbyist or something.

Michael Intrator

It could be sent someplace else where they have more capacity and can repurpose it there. I kind of feel like we'll deal with that part of the business when we get there.

What I know right now is that it is extraordinarily profitable. It's very accretive to my company to continue to keep the infrastructure that's been up and running and that's been on these long-term contracts. As it rolls off, as it's been in use for 5 years and becomes available, I'm still able to sell it at a higher price than it was at a year ago.

Speaker 1

There's competition now. When you were buying these from Jensen back in the day, you could buy them and have them shipped, I would assume, within 30 days or less. Nowadays, what's the wait like, even for you, a loyal old customer? Is there a bit of a battle? Is there politics to who gets the servers? You see some very big names talking about how they have to get an allocation. Is it still a little bit crazy? What's it like to be in that category, having to buy something everybody wants?

Michael Intrator

Look, I think of it as an affirmation of the business that we're in, right? The fact that we are attracting competitors means that the business is healthy and that there's a lot of people trying to deliver this service. The need for this infrastructure, and the need to integrate the infrastructure into the software layers to deliver it to artificial intelligence—whether at the model level, the inference level, the application level, or whatever level of the 5-layer cake that Jensen's focused on—is growing. The fact that there are more people coming into this doesn't discourage me.

As far as getting access to the GPUs, we show up like everybody else with a PO: we'd like to buy, and we're ready to pay.

Speaker 1

What's the wait time like? Is it just really competitive or not? Because I talked to Jensen about it. I said, “How do you manage all these big egos, names, and companies trying to buy stuff?” And he said, “Well, they order it, and we give it to them in the order in which they order it.” Is it really like that?

Michael Intrator

It really is, right? He doesn't want to be in the position of playing favorites. That just seems like a bad place to be with your clients.

Daniel Roberts

Or auctioning them off.

Michael Intrator

Yeah, that would be crazy. I'm not sure that would be good for the long-term business.

Daniel Roberts

No.

Michael Intrator

Yeah, so our approach is—

Daniel Roberts

Get some sovereigns coming in and saying, “I'll pay double.” They do that with Ferraris, too, sometimes.

[Laughter]

Because these are the Ferraris of computing, right?

Michael Intrator

By the way, they are. Yeah, they're the Bugattis. Our approach is to work with clients across the entire space to find opportunities with really interesting companies that can fit into our contracting requirements, where we're going to be able to go out and structure the debt that we require in order to build infrastructure at this scale.

Speaker 1

How does all that debt work? That is something that you guys specialize in. Corporate debt—I'm in the venture business. People are like, “Why should I be in venture when corporate debt pays so well?” Corporate paper is so huge. I'm curious how this fits in and what interest rate people are paying on $1 billion in infrastructure. What do they pay on that?

Michael Intrator

Yeah, so CoreWeave has really been the innovator around a lot of the financing engines that have come to bear on this. We did the first GPU-backed loans. I think it's important—or I'm going to try to explain this in a way people can understand.

What we do is go out and find a client. Let's use Microsoft—you brought them up before, right? Microsoft comes to us and says, “We'd like to buy something.” We say, “Okay, great. We're going to sign a contract.”

Once I have a contract in hand, what I do is create something. It's not a particularly creative name. It's called “the box.” What I do with the box is take my contract with Microsoft and put it in the box. I go to Jensen and buy the GPUs, and put them in the box. I take my data center contract and put it in the box. Now the box governs cash flow. It has a waterfall of cash flow that comes into it and goes out of it.

The way it works is that I build the compute and deliver it to Microsoft, and they pay the box. They don't pay me. It goes into the box, and the first thing it does is pay the data center. It pays the power bill. It pays the interest and the principal. Then whatever's left flows back to us, right?

It is an incredibly well-structured, time-tested, pressure-tested vehicle to borrow money against client paper and all of the other collateral around the deal. That's why CoreWeave, a company that many people haven't ever heard of, was able to go out and raise $35 billion in 18 months to build infrastructure at scale.

What's important to understand is that the economics in this box are such that within 2 and 1/2 years of a 5-year deal, we've paid for everything. The principal has been paid off, and the interest has been paid off. The return into the box is such that we're able to generate returns to our company at the box level, right?

That gives the most sophisticated lenders in the world—whether it's banks, private equity funds, or whoever—confidence that they're going to be able to achieve the 1 rule of lending: “Give me my money back.” Yes, and so it's better when that happens.

Speaker 1

So they look at this box and they're like, “Wow, we're really confident we're going to get our money back.”

Michael Intrator

And maybe they want 10 boxes.

That's correct.

Speaker 1

And if any one box goes upside down, you can deal with it, and it's not as acute.

Michael Intrator

That's correct. They don't cross-pollinate; they don't cause a contagion across the boxes. One, and number two, as you do this and as you show the lenders how this financing tool and how this financing mechanism works, what they do is they continue to lend you money at progressively lower rates.

When you think about our cost of capital over the last 2 years, we have dropped our cost of capital by 600 basis points. Wow. It is enormous, right? You're seeing a company that is driving its cost of capital down toward where the hyperscalers borrow, which will enable us to be competitive with them over time. We have been extremely militant and diligent about feeding, watering, and caring for those boxes so that we continue to have access to the capital markets in a way that allows us to build and drive our business.

Speaker 1

That means you have to say no. You have to say no to maybe some people who want to be in the box?

Michael Intrator

Yeah, we look at some deals and we're just like—they want to buy GPUs for a year, and I look at it and say, "That's not a deal that I can do because it's too short for me to amortize the expenses." And so I won't do that.

Speaker 1

They can go to another provider who maybe wants to take that risk on, who has extra capacity.

Michael Intrator

Absolutely, but our business is really built around the risk management of being able to get to scale because, in my mind, during this period of disequilibrium—during this period where there were not enough GPUs in the world to provide the compute for all of the different use cases in artificial intelligence—the part that's important for me and for my company is to get enormously large so we can drive down our cost of capital, so that we have information flow coming in from all different parts of the market.

The large language models, high-speed trading, search—all of these things are feeding information back in to us that is letting us know what the next product we need to build is, where they need help scaling, or what type of compute they need. All of that information flow is incredibly valuable to us.

Speaker 1

What can you tell us about demand? There have been reports of, "Hey, maybe the Oracle star base thing with OpenAI has been downsized, or maybe not." And then other folks—Microsoft is going big, Google's going big, Meta is going big—and those people obviously have massive cash flow. Apple seems to be MIA. They don't seem to want to play.

Michael Intrator

You've named a lot of really big companies with really big balance sheets that have the capacity to drive a lot of demand. Look, I have been truly steadfast in this for years now. For 4 years, the depth of the demand for the service we provide has been relentless and overwhelms the global capacity of the world to deliver enough compute to enable all of the demand for artificial intelligence to be satisfied. We have been relentless about that.

Speaker 1

Sounds like Knicks tickets during the Patrick Ewing era. They got up to 50,000 people on the waitlist. So if magically the waitlist went away, if the constraint went away, and we just had a large amount of GPUs available, a lot of energy available, and a lot of data center available, how much capacity would just all of a sudden come out of the system? So, what would be deployed, I should say?

Michael Intrator

Remember how we build our business through this box. It's a 5-year box. If we had an air pocket—if demand were suddenly to disappear because of a technology breakthrough, because of a war, anything—the why, from a risk-management perspective, does not matter. You have to prepare your company for what happens if it happens.

Speaker 1

Yeah. And so, by entering into these long-term contracts, by entering into contracts with counterparties that have large balance sheets, you are—or we are—protecting ourselves and our lenders.

Michael Intrator

Yeah, so that we are confident and they are confident, because you can see how confident they are by the rate that they're charging us continuing to decline, that they're ultimately going to get their money back, and that is the one rule of lending.

Daniel Roberts

Yeah. And so, in terms of the capacity, if you were unconstrained and NVIDIA's Jensen says, "Order as many as you want," what would happen?

Michael Intrator

It is also important to understand the constraints aren't just GPUs.

Speaker 1

Right. Right. Electricity.

Michael Intrator

It's power shelves, it's memory, it's storage, it's networking, it's optics—all of the things. There are various throttles.

Speaker 1

Memory is the throttle right now, right?

Michael Intrator

Oh, yeah, it is. Oh, yeah, it is.

Speaker 1

Why? How did memory become the throttle? Memory has historically been a cyclical business, right? We have seen these waves of demand driving up the cost for memory, and then it collapses, and then it drives it up. It's a very boom-and-bust business. It's cyclical in its nature because the fabs are so capital-intensive that people invest in the fabs, build a ton of capacity, and then overbuild if there's any type of downturn. We've seen that cycle again and again.

Michael Intrator

What's happening right now is the confluence of 2 things, right? One is, with all the demand for artificial intelligence and the corresponding demand for compute and the ancillary services around the GPU, the demand is through the roof. That's number one.

Number two is that there was probably an investment cycle that needed to happen back in 2023 that would have brought on the necessary fab capacity to be able to serve the demand. It's impossible to predict what just happened. Just with energy, it's impossible to predict what just happened, and now people are chasing energy. The data centers are going where the energy is. It's not based on real estate; it's based on where there's some wind.

Many times, when you have a capital-intensive business like building fabs, you will get this boom-and-bust cycle. Just like in energy, they overbuild. And then, you know, fiber. Yeah, I mean, there are a lot of examples of that.

Speaker 1

In some ways, when you look at that, it's a beautiful aspect of capitalism: we're able to have a boom-and-bust cycle, and we're able to weather it, right? If you think about capitalism from first principles, something like that happens and we have too much fiber, it creates an opportunity for Google to buy it all up or the next person.

Listen, the boom-and-bust cycle does a lot of things. It clears out the underbrush. The strongest companies will be able to survive and take advantage of that, and it sows the seeds of future business. The other thing that it does is put that infrastructure into the ground. You put the fiber into the ground, which became the backbone of how we watch movies every day, how we communicate, and how we hop on a Zoom. COVID and all of these things were based on that infrastructure that was available to be consumed.

People don't recognize this fact. The premise of YouTube, from the founders whom I knew, Chad Hurley and his other partner, was that they basically had the realization that storage was coming down so quickly that they could offer free, unlimited uploads, and bandwidth was coming down, so they didn't have to charge people for sharing a video online.

Before that, if your video went viral, people were going to have their minds blown, but your server would turn off and it would say, "This person needs to pay their bill." They were getting charged for carriage by the megabit going out.

The business models change and evolve, and, like you said, Moore's law—and certainly NVIDIA's Jensen will talk about the fact that what is going on within accelerated compute dwarfs Moore's law. All of that is going to lead to more opportunity to build more companies that are going to do things like YouTube did, which has really changed the world.

The concept—I don't know if it was a million hours being uploaded every hour or minute—but at some point, Susan Wojcicki, rest in peace, said to me, "How much was being uploaded every minute?" It made no logical sense until she realized, "Well, there are 2 or 3 billion people on the service, and 10 bips upload." It's like, okay, 1 in 1,000 people upload. It's a big denominator.

I was sitting on a panel with Sarah Friar, CFO of OpenAI, and she every once in a while really puts out interesting information. She was talking about the cost of a million tokens when GPT-3 came out, and it was $32 and change. Now a million tokens costs 9 cents.

Michael Intrator

Right, and so you just see the incredible power of how the capital markets, how capitalism, are fueling engineering and fueling competition.

Speaker 1

It becomes recursive now, too. I mean, these models—if you say to the model, "Hey, make yourself more efficient, spend less money, and lower the cost of tokens," it'll be like, "Okay, captain." I don't know if you saw Karpathy's recursive thing last weekend, but now civilians who've never worked in a language model or done computer science are like, "I'm going to try to do something recursive this weekend."

Michael Intrator

You know, it's one of the things that I talked to the other founders about. When you think about some of the things that AI does, right, it's lowering the barrier to operations.

So if you have a good idea or a great idea, you can open up your model and tell your model—you can vibe-code it, you can do all kinds of different things—and create things that never existed before. That's amazing, right? That's bringing down this incredible barrier that kept human creativity contained, and now, all of a sudden, there's this whole new vector of medical research or different approaches to baseball cards or whatever you want.

If you've got a great idea, if you've got a new creative idea, that's the valuable kernel right now that allows you to build new things and create new things. I just think that's incredibly exciting. You're bringing the minds of 8 billion people a tool that allows them to overcome what was insurmountable forever for humanity.

Speaker 1

Yeah. It's a bright new future, Michael. I appreciate you sharing the information with us and the vision. I am really delighted to have Aravind Srinivas on the program.

Aravind Srinivas

Thank you for having me. It's so great.

Speaker 1

I want to go through 3 stages in which I fell in love with your product. The first phase was that I could go in and pick my language model if I wanted to use OpenAI, Claude, or whatever it was. That was a real unlock for me.

On the sidebar, I noticed you had done essentially what Yahoo did in the early days: finance, sports. When I pulled my Knicks game up, it gave me a live version of that. When I pulled my stocks up, it summarized the news in real time, and I was like, "Wow, this execution's great."

I made you my front door to different models, and it made it easier for me to check them. Then you came out with the Comet browser, and I was like, "Holy cow, I can give this a series of instructions: Go to my LinkedIn, find everybody from this company, and put them into a Google Sheet." Boom, you were the first out of the gate with that.

Then, just in the last couple of weeks, I'd been Claude-pilled and using Open Claude, but you came out with Computer. I started using Computer, and boy, it's good. It's a really strong start, allowing me to do repetitive tasks very similar, in some ways, to Claude Cowork, or basically an engineer or developer using it.

So, are these the evolution of the company, and should I think about it that way? How do you look at Perplexity now? You have a very loyal fan base. You're making a lot of money. I don't know if you disclose it, but I think it's hundreds of millions to billions. You can tell us.

What is Perplexity in the face of Claude having a great run, OpenAI still doing strong, Grok doing very well, and Gemini coming on strong? There's like 6 or 7 of you, and you just happen to be one of my top 2 right now. Thank you. So, tell me.

Aravind Srinivas

First of all, thank you. Thank you so much. Perplexity has always been built for people who are always looking for the extra edge—the curious people. So, it's very natural that you are one of our power users.

One common theme for us for the last 3 and a half years is accuracy. Perplexity wants to be the company that's building the most accurate AI. When you want to give somebody answers, accuracy is very essential for building trust, because only then is the user going to ask the next set of questions.

It turns out it was a great idea to give AI access to the internet to be accurate. So that's the Perplexity Ask product. It turns out it's a great idea for AI to have full access to a browser so that it can be accurate when you task it to go do something that you would do yourself on a browser. Agentic browsing: Comet.

Now, the last phase is that it turns out it's a great idea for AI to be given full access to a computer, so that it can do whatever you do on a computer on its own, essentially becoming the computer itself.

It's an orchestra of everything AI can do today—every single capability each individual AI model has, be it GPT, Claude, Gemini, or anything else. An orchestra of all those capabilities. That's what Perplexity Computer is.

All these sub-agents that are running inside Computer are the musicians. The models are essentially the instruments. There are hundreds of models out there, each having its own specialization. Some are good at coding, some are good at writing, some are good at multimodal visual synthesis, image generation, video generation, or audio.

But what matters is the end output, the music you play. That's the work AI gets done for you, and that's what Perplexity Computer is. The AI itself is the computer now.

Speaker 1

It still lives inside of a browser. Have you considered giving it desktop root access? That feels like the next place this is going, but that comes with a lot of security issues and a lot of trust issues.

As you mentioned, trust is paramount. Getting the right answer is what builds it, but also not getting hacked and not having it delete your files. So, how do you think about root access to my Windows machine? Obviously, iOS won't let you, but with an Android phone, it would let you. Do you have that in the works?

Aravind Srinivas

Yes. We announced something called Personal Computer—Perplexity Personal Computer. That's essentially going to take all the trust and reliability and the server-side execution of Perplexity Computer, but synchronize it with your local computer so that you can use it from your phone.

We're going to do this with the Mac mini, where you synchronize your computer with the Mac mini so that it becomes your local server. All the agent orchestration that has to do with your local private data will run on that local orchestration loop—that runtime—with the Mac mini.

Daniel Roberts

Not on your servers, not on Anthropic's.

Aravind Srinivas

Exactly. It could still ping frontier models if it needs to, with your permission, but it will be orchestrating everything on your local hardware.

If it needs to run on the server-side hardware, if you don't want very complicated, long-running tasks to be running on your local hardware, you can delegate it to run on your server-side computer, which is again only accessible to you and you alone.

That way, we're going to bring this perfect, trustworthy hybrid between local and server-side.

Speaker 1

And you'll make it easy to do. It'll just be abstracted. You install 1 executable, and boom, it's done.

Aravind Srinivas

It's like OpenClaw for dummies. Nobody needs to learn how to use it. Nobody needs to manage API keys. Nobody needs to manage separate billing across 100 different services, or figure out what you can give access to and what you can't access. We take care of that.

So it's the Steve Jobs way of doing it: end-to-end integration.

Speaker 1

And how do you think about local models? I've started running Kimmy 2.5 on a Mac Studio. It's not as good as Claude, Gemini, or Grok, but you can probably do about 80% there for free.

Aravind Srinivas

Yeah, essentially.

Speaker 1

Do you have one of those? Have you started testing on your local Mac Studio? I assume you have a Mac Studio and you're doing this yourself?

Aravind Srinivas

Yeah. I don't know if you saw Dell and NVIDIA announce a giant workstation. Is it a 3800? Something like that, with 750 GB of RAM.

Speaker 1

Something like that, with 750 GB of RAM. So, what do you think about the desktop going back to workstation-server status?

Aravind Srinivas

I think it's very promising. My prediction is it'll initially start off as a sub-agent. Whatever you need to go—your tax returns, your personal photos, your emails, your calendar, all that stuff, those local apps, your personal notes, very personal notes—you can make sure that the models that access those tokens will be running on your local hardware if you want to, if you're that privacy-conscious.

More complicated stuff that accesses your data is already on the server side. For example, your Google Calendar—

Speaker 1

Your Gmail.

Aravind Srinivas

This is personal data still, but an AI runtime can access that through your connector—your Google Calendar connector, your Google Workspace connector. That could run on the server side because, anyway, the data's on the servers. It's not even lying on your device.

So that sort of hybrid orchestration is where we are headed. I don't think it's a dichotomy between fully local versus fully server-side. It's all about choice.

Anyway, when you're on your phone, you don't actually care which server that workload's running from, because it's not going to be able to run on your phone anyway. The chips need to exist on a Mac Studio or a Mac mini, or on the new Dell that's coming out.

Speaker 1

I really think the idea of spending $10,000 on a powerful desktop will appeal to people if it lowers their $500-a-month cloud bill.

Aravind Srinivas

Yes. This is an incredible savings, plus you get the benefit—

Daniel Roberts

Yes, of privacy and not educating the language models on your personal data.

Aravind Srinivas

Yes. And it's going to be like you're buying a refrigerator, your internet modem. The cost for these will eventually go down, but it's not going to feel like you're wasting your money.

Every home has a lot of other sensors that run your home. They'll also be part of this orchestration loop. That's where it gets exciting, because now you can just dictate something to your phone, and that can control your entire home.

That's the dream that everybody has, and that entire orchestration loop can run on your local hardware, no problem.

Speaker 1

And I'm curious what you think of the operating system. What's eventually going to be the operating system of this workstation?

Aravind Srinivas

AI is the operating system. Earlier, in the traditional operating system, you executed programmatically. Now you start with objectives, not specific instructions. You come up with a high-level objective: "Go build this website for me that takes all the transcripts of the All-In podcast and tracks the stock price just before the podcast and after."

Daniel Roberts

Yeah, and charted for the max 7. Yeah, and charted over time. You can—so that's the objective.

Aravind Srinivas

But individually, it's running a file system, a code sandbox, and access to the internet. It's got its own HTML tools. So I think that's basically where models, systems, files, and connectors are all coming together. You would think of that as an OS.

Daniel Roberts

Mhm. Except you're operating at an abstraction above that, where you're thinking in terms of objectives. Does it need to eventually become its own operating system in your mind?

Aravind Srinivas

It could be. People could think about it as, “Yeah, I have my Perplexity Computer running all the time.” Essentially, it runs on Linux machines right now. Every server-side computer is a Linux machine. Mark Zuckerberg recently tweeted, right after our release, “Turns out Linux computers were the right idea. Desktop Linux computers are finally going to work.”

Linux machines are stable and customizable, and you're not at the mercy of Apple's desire to contain the experience or Microsoft's attack surface for hackers. You build something rock-solid, and it does feel like Linux might actually become the eventual winner. It may not need to have a front end. You could access the Linux machine on your phone, running iOS or Android. It doesn't matter. The actual valuable runtime is running on Linux on the server.

Speaker 1

You've done great as a consumer company. A lot of love there. Now I'm starting to see corporations engage with it. In fact, you'll be happy to know this: Last week, I took 2 people in my back office and said, “Stop working on OpenClaw. Your job is to do the back-office automation at our venture firm, using only Perplexity.” They said, “Perplexity Computer?” I said, “It will.” I'm going to see Aravind, so I'll talk to you about that. We need a really strong Slack connector.

Aravind Srinivas

It's already out.

Daniel Roberts

It is? Okay, great. At first, we were sending reports in, but it wasn't interactive. That's perfect. So now you've got your company going in 2 different directions: this incredible consumer run you have. How many people are using the product every month?

Aravind Srinivas

Several tens of millions. Our computer exists as a Slack bot right now that you can add to your Slack workspace on the enterprise plan. Our entire company works like that. People are talking more to the computer on Slack than to other people.

Daniel Roberts

That's the first volley. We were sending reports in, but it wasn't interactive. So now you've got your company going in 2 different directions. This incredible consumer run you have—how many people are using the product every month?

Speaker 1

So tens of millions of people. That's very similar to the trajectory of Google's and Yahoo's consumer businesses. Now you've got corporate. How are you doing on the corporate side? Thousands of companies?

Aravind Srinivas

It's a growing business for us. It's growing faster than the consumer business in revenue. Things like Computer unlock entirely new possibilities. For example, we've saved more than $100 million for our Enterprise Max customers, who are on the highest tier of enterprise.

Speaker 1

Explain that. What does it cost? $200 a month per person?

Aravind Srinivas

There are 2 tiers. One is Enterprise Pro, which is $40 a month, and there's Enterprise Max, which is $400 a month. On Computer, after you run out of your credits, you pay for the tokens. You pay for the usage.

Speaker 1

Are you making money on the $400-a-month, $5,000-a-year one, or at this point are people going so crazy that—

Aravind Srinivas

One thing Perplexity has, unlike certain other wrapper companies, is that every dollar of revenue we make has positive gross margins. We’re not just selling tokens. Most of our revenue is recurring because people are paying a subscription fee. Because we route through multiple different models, we're very efficient in terms of how we spend on tokens. Because we have all this advantage with RAG, orchestration, and search, we don't actually need to blow up the context window of the models.

As a result, we have positive gross margins on all the revenue we make. Every single penny we make, we make profits on. Overall, the company is still not yet profitable, but we're working toward that.

Speaker 1

You've had the opportunity to exit. There were a lot of rumors that Apple and other people were saying, “Hey, this is a great team.” How many people are on the team now?

Aravind Srinivas

About 400.

Daniel Roberts

You've got a very coveted team. You obviously understand consumer, you obviously understand business, and it's a product-driven organization. Reports are that you declined the offers. But the world's getting hypercompetitive here. How do you keep up as a 400-person organization when you've got Sam Altman over here raising $100 billion, Elon putting data centers in space and merging with SpaceX and Twitter, Google with unlimited resources, and Amazon getting into the game? Gemini is a very strong product, and Google is really good at consumer.

I think we'd all agree that Facebook and Meta haven't figured it out yet, except maybe for serving us better ads. They haven't figured out the consumer case yet, but they'll copy it. They always do. How do you look at the playing field? The degree of difficulty—this isn't playing checkers. This is like playing against the 10 best chess players in the world. That's what you have to do every day. How do you think about it long-term? An independent company? Do you think you'll need to join forces at some point? And why didn't you take the deal? The deals that you were offered were incredible.

Aravind Srinivas

One advantage we have that all these companies you mentioned don't have is multi-model orchestration. We're like Switzerland. We don't have to have 1 horse in the race. If GPT wins, Gemini wins, Claude wins, or Llama wins, it doesn't matter to us. Even open-source models can win. No problem.

Speaker 1

And you have them on the service? You have DeepSeek and Kimi?

Aravind Srinivas

We have Kimi, we have Nemotron, and we have a lot of usage of Qwen.

Speaker 1

Alibaba Qwen?

Aravind Srinivas

Yes, silently under the hood. For us, the advantage of being able to take the best in each model and give the user the orchestration of everything they can do is something I don't think any of the companies you mentioned can do.

Daniel Roberts

Right. Nor would they.

Aravind Srinivas

Nor would they. It makes no sense for them. It would be an admission that all the data centers and capex they've built out mean they still couldn't produce the best model themselves.

Dario Amodei, the CEO of Anthropic, said recently in an interview that models are specializing. Toward the beginning of last year, people thought models were going to commoditize, but toward the end of last year, models started specializing. Even within coding, Claude Code and Codex have very different capabilities. Our iOS engineers love using Codex. Our back-end engineers love using Claude Code.

Even within a specialization like coding, models have their own unique specialties. There are many other use cases outside coding where different models are good at different things. That means the orchestra conductor, which has no single horse in the race, can win by providing a very unique value and service to the customer that each of these amazing companies cannot.

Speaker 1

So you're buying tokens wholesale from them and then charging customers for it?

Aravind Srinivas

We're going to take care of all that orchestration, so you don't have to manage tokens across different models.

Speaker 1

I authenticate a couple of my different accounts—my Pro accounts—into Perplexity, but I don't have enough knowledge to know if you're abstracting that and people can just search across them as part of their Perplexity subscription.

Aravind Srinivas

No, we're not bundling subscriptions from other AIs. We just ping the models directly. What you get from us is Perplexity orchestration. When models are specializing, there's a bigger value in the one who knows how to build a great harness that can take the best from each model.

Speaker 1

Does it auto-route today, or do you still have the dropdown? Somebody has to pick?

Aravind Srinivas

It definitely auto-routes to the best model for each prompt, but we also give users the flexibility to pick whatever model they want.

Speaker 1

I've seen a bunch of startups hack this together. What do you think of doing the same query across multiple models?

Aravind Srinivas

We built a thing called Model Council.

Speaker 1

Model Council, yeah. I saw Jensen Huang say in one of his interviews that he puts the same prompt into 5 different AIs and sees what each of them says. Everybody does that.

Aravind Srinivas

But then you still have to apply your biological compute to read every answer and figure out where they differ. It's like talking to 5 different doctors and trying to figure it out.

Daniel Roberts

Exactly. It's dumb.

Aravind Srinivas

Model Council is a feature we built where it will not just give you the answers from each model. It will tell you exactly where they agree, where they disagree, and where the nuances are.

Speaker 1

And that's in the interface? Model Council? I didn't know it was there. You release products at a pretty great cadence. Where did you learn that, and what's your philosophy of shipping product?

Aravind Srinivas

Our philosophy is that speed is our mode. One of the things that big companies cannot do is move at the speed we do and serve customers at the speed and quality we do. It's very hard to maintain quality, speed, and trust at the same time. Apple takes a long time to ship anything.

Speaker 1

Right.

Aravind Srinivas

Because they're very worried about people not trusting them. And so some companies are bureaucratic, and they just take forever to ship something. They don't maintain what they ship. They may make a big deal about an event, but nobody even knows how to go and use that feature. They get abandoned.

Speaker 1

Exactly. So, Perplexity has those advantages of being very small.

Aravind Srinivas

Mhm. Which is honestly one of the reasons why we built Computer, because now even non-engineers are shipping code here by just pinging a Slack bot and asking it to fix bugs.

Speaker 1

So, this iteration has just been exponential. The moment I became Claude-pilled was when I was working with it and I was like, “Hey, I want to build my network. I know these 20 people in Japan. I had dinner with them during my recent trip. I want to know who they know, so check out LinkedIn and other things to see who they're associated with, and make me a mind map of it.”

“On the next trip, I want to meet with the next circle of those connections.” So I started asking. It said, “Okay, I got the results.” I was like, “Great.” It said, “Where do you want me to put them?” And I was like, “Well, where can you put them?”

It said, “I could put it in a Google Sheet. I could put it in a Notion table. I could put it here. I could give you a PDF. I could give you a CSV file, or I could write you a CRM.” And I was like, “Yeah, sure. Make me a CRM system.” And it made a CRM system.

Aravind Srinivas

Yeah. And I think maybe 1 out of 1,000 people working with AI have had that experience. Maybe it's 1 in 10,000, where your agent says, “I'll make you bespoke software.” Have you had that yet? Do you see that as a part of Computer, where when a person needs a spreadsheet, you don't launch Excel or Google Sheets—you just pop up a spreadsheet?

Yeah. Well, we have a board meeting tomorrow.

Speaker 1

Okay, I'll come. [Laughter] So, pitch it to the board.

Aravind Srinivas

Sure. Computer made the memo. And we had a partner meeting to pitch a partnership idea. Earlier, we would have a design team do the whole deck. Computer did it in one shot.

I had a press briefing with a bunch of journalists. My comms person would usually give me a memo about what to say. Computer did it in one shot.

Speaker 1

That's brutal.

Aravind Srinivas

It's crazy, and the context is so good because the memory is getting better. So, it knows that journalist from the last time. It knows the board meeting. It has all the previous decks.

Speaker 1

When did that happen?

Aravind Srinivas

I think it happened with Opus or Phi.

Speaker 1

Uh-huh.

Aravind Srinivas

Anthropic's Opus or Phi—that was the inflection point when models started being amazingly good at orchestration, reasoning, and tool calls. Claude Code brought in this new idea in AI that everything can happen inside a sandbox, a console, or a terminal with access to tools, where tools are just command-line tools. They don't even need to have a graphical user interface.

So, when you did that, and when you organized around files, subagents, skills, and CLIs, the model became very good at handling the context. The context window no longer became a problem. It just put whatever was necessary into the context whenever it wanted to and dumped it away when it wanted to. That made it suddenly so good at doing very long orchestration tasks.

Speaker 1

Yeah, it's pretty crazy. I have every episode of This Week in Startups, all the transcripts, and then all of All-In.

Aravind Srinivas

One of the tasks I did, by the way—I can send it to you. I asked it, “I want you to download every All-In podcast.”

Speaker 1

Yeah, since the beginning.

Aravind Srinivas

Yes. I want you to take a mention of all the public companies they mentioned during each episode. I want you to have a histogram of the counts, and I also want you to chart it across time. Then I want you to analyze the impact on the stock price and the sentiment of what we said.

Speaker 1

Exactly. And it did. It clearly said—

Moving stocks?

Aravind Srinivas

It was about Google's stock going up. Yes, prior to that, you guys were talking a lot about Google.

Speaker 1

And I said, “I made a bet publicly on the thing. I said, ‘I am buying a bunch of Google because I believe, even though they're behind, it's because they're too precious.’” You were mentioning a company that might be too precious at times and doesn't release. I was like, “That's that company. They need to release more.” And I told Sergey, “Give us the good stuff.”

Aravind Srinivas

Yeah. He started giving us the good stuff.

Speaker 1

It literally gives you the timestamps of every single mention, and then I can click on it and actually hear exactly—

Aravind Srinivas

Yeah. Sweet. That's when I was like, “Damn. This would have been a week-long project.”

Speaker 1

It would have been 10 hours a week of a researcher. I'm experiencing the same thing. When I do research notes, I've created my own mega-prompt, and it will tell me where you worked before, who's in your circle, who your competitors are, who your friends are, and so on.

Then it will find old podcasts. That's one of my secrets if you're an interviewer watching. I try to find what the person was talking about 5 years ago, 10 years ago, and then over 10 years ago. I've gone into interviews now with Michael Dell and talked about things he was talking about in the '90s.

It finds me some ancient stuff. You would pay a researcher or a producer $70,000 or $80,000 a year to do this, and they would have done a third of the job in 10 times longer.

Aravind Srinivas

Yeah. It's really gotten weird just in the last 6 months.

Speaker 1

What do you think the next 6 months looks like?

Aravind Srinivas

I think the dream is to help businesses run as autonomously as possible. Everybody talks about how AI is going to create this 1-person, $1 billion company. Some people say it's already happened because people pay researchers like $1 billion. But it's not truly moving the GDP by $1 billion. It's not truly creating new value.

So, the best way to do that is to actually help a small business—people who would otherwise drive Ubers for extra passive income—buy a Mac mini, set up Perplexity Computer, and run their business on that, or run it on the server. It doesn't matter. They can actually make real money—hundreds of thousands or even millions a year—and grow it.

Have Computer run your ad campaigns on Instagram or Google. Integrate with SEM and SEO tools, find new users, integrate with Stripe, charge them, ship new features, have your own Intercom integration for customer support, and have all of this working while you can be sipping wine in Napa.

That's the dream. It feels awesome to say. Everybody thinks AI is already there. It's not there yet. Someone has to do that hard work. That's what we want to do.

Speaker 1

It's a great vision, because when I watched startups 20 years ago, there were so many checkboxes they had to do. I have to find an office space. I have to put up a bunch of servers. I have to hire an HR firm. I have to hire a PR person. All this stuff.

Now I talk to young founders who have a 3-person team. They've come out of a16z, my Launch Accelerator program, or Y Combinator. And I'm like, “Okay, you raised $500,000. You raised $1 million. Who are you hiring?”

They're like, “I don't know if we need to hire anybody.” I'm like, “If you could hire somebody, who would you hire?” They're like, “Well, I do my own HR. I have this partner.” I'm like, “How are you doing hiring, anyway?”

They're like, “Well, I put out an ad, and then it sorts and ranks the candidates. It emails the top 10, asks them a bunch of questions, and then I meet with the last 2.” I'm like, “That's what a recruiter did.”

Aravind Srinivas

The entire recruiting job has been abstracted, and a tool like Computer is going to make that even faster.

Speaker 1

There's work to do. A lot of connectors. A lot of specific workflows. People don't want to learn how to write essay-long prompts. It needs to be so quick, fast, and autonomous. You just set it up and it's done. You have an idea, you can turn it into a business, and start making money.

It's an incredible future, and it feels like it's right here. How do you think about job displacement? You're actually making the tool that enables people to be solo entrepreneurs and get to $1 million in revenue, but it's also the same tool that doesn't require them to hire. We've had this debate a million times on the podcast. Do you have moments where you're like, “Oh my God, this is really terrifying?”

Aravind Srinivas

Yeah. A lot of people are going to lose their jobs really fast.

Speaker 1

Yeah. And then, oh my God, you can learn any skill you want, and all the things that were hard are now easy.

Aravind Srinivas

Yeah. I go back and forth. I'm 70–80% super positive about this, but 20% of the time I'm a little worried. Where do you sit?

Speaker 1

I mean, America has always been about entrepreneurship, right? We've been about trying to build new things, discover new things, and go explore.

I think Henry Ford came and built factories and brought in jobs and things like that, and put people into a box. But the reality is that most people don't enjoy their jobs. They're doing it for—

Aravind Srinivas

They hate them. Exactly. So, there is suddenly a new possibility, a new opportunity to use these tools, learn them, and start your own mini-business.

If it pays for your needs for a year or multiple years, lets you have a high-quality life and good work-life balance, and gives you a true feeling of agency, ownership, and passion to get your ideas out there, I think that—even if there is temporary job displacement to deal with—that sort of glorious future is what we should look forward to.

I think you're exactly right. If there will be some displacement, then there's also going to be so many opportunities opening up, and it requires the individual to not be passive.

Daniel Roberts

They have to be rugged individualists. They have to be resilient. Yeah, and they have to be resourceful. I think once you start playing with these tools, that's what happens.

Aravind Srinivas

Exactly. You all of a sudden feel—it brings out the best in you if you truly are in a good space. Yeah. And then today, Comet for iOS is out.

Daniel Roberts

Yeah. I'm a Comet super fan. I required everybody to put this on. You were nice enough; I emailed you. I was like, “Can you send me some licenses?” You sent me a bunch of licenses, and I said, “Everybody, put this on.”

Because it was $300 a month when you first came out with the Comet browser. Now it's free, I think, for all users. Highly recommend it. Highly recommend getting a Pro account. It's only $20 a month to get into Perplexity, which is a joke. You can get on board for nothing—less than a dollar a day.

But what does iOS allow me to do? And how does it connect to Computer? That's another thing I'm having trouble with. Claude Code and Computer—there's not a good enough integration with this mobile device yet.

Aravind Srinivas

Yeah. Computer is already on the Perplexity app, so you can just toggle to Computer and start using it. Comet's uniqueness in Perplexity, for the company and the strategy, is the fact that you can control the browser.

The browser also becomes a tool for Computer, just like your Google Workspace and all these other things. Until the whole world is organized around CLIs and tools, there's still a lot of tasks we have to do manually on the web, on the browser: open tabs full of forms, click on things, upload stuff.

All that stuff, if you want to automate, you need a browser. You need an AI that can natively control the browser. So, that is Comet. That's why, no matter how many other tools in the market exist, like OpenClaw or Claude Cowork, executing tasks on a browser on the server side, along with all the other things, is something uniquely Perplexity can do.

Daniel Roberts

Yeah, my dream is that you'll create an Android app that roots my Android phone. You just take over and see everything. One of the blockers I have now is that some of the websites have gotten a little persnickety.

I don't want to mention too many, but Reddit and LinkedIn. I'm a great Reddit user. I'm a great LinkedIn supporter. But sometimes I need to get my InMail from LinkedIn, and I just need to find 7 people at a company.

Is there going to be a solution between the LinkedIns and Reddits of the world and Claude and Perplexity? How is that negotiation going? You don't have to speak about any specific ones unless you want to. It feels like there's got to be a solution, and I'm willing to pay for it as a user. I'm willing to pay Reddit to allow my bot to show up and behave properly.

Aravind Srinivas

Yeah. Well, I cannot speak about any particular company, but we are happy to work with anyone. With Comet, our idea is to give people the flexibility to set things up on their own. Any official APIs that anyone's willing to offer, we're always happy to put that as part of Computer.

Daniel Roberts

Here's what I think should happen. Let me see if you agree. This is for Steve Huffman at Reddit. I go on Reddit and get a Pro account for $20 a month. When I do that, I can authenticate whatever tool I want to do a series of well-behaved things a certain number of times a day.

It's not unlimited. I'm not going to scrape the whole site, but I would like to let Perplexity or Computer go and tell me, “Hey, what are people saying on the This Week in Startups and All-In subreddits? Summarize it for me so I get the customer feedback.”

I would literally name my agent and say, “It won't post on my behalf. It won't vote on my behalf. I just need it to do a couple of little read-only things.” This would be an easy solution.

Or with LinkedIn, I already pay them like $50 a month. They should just let the $50-a-month account work with Computer.

Aravind Srinivas

Yeah, absolutely. This is for Satya Nadella: Let LinkedIn work with Perplexity and the other players, and we'll pay you extra. Perfect.

Daniel Roberts

It's a revenue stream. Don't you think API access for our customers is a revenue stream?

Aravind Srinivas

I think so. Fundamentally, giving users a choice and setting it up as a win-win for both the business and the user is where the world should head. I would say the same thing applies to any website in the world. If you want an AI to use it on your behalf, it should be okay, because that's what the user wants.

I have a paid New York Times subscription. Let me go in there and do 100 searches a day, a week, a month—whatever they choose. That would make the subscription that much stickier.

Daniel Roberts

Exactly. All right, Aravind, love the product. Anybody at home, it's just tremendous. Go learn Computer and get the Comet browser. It has changed my business for the last 2 years. Love the product, and we'll have you back soon when you launch your operating system and come up with your own server and desktop server, but business is the focus. Yes. Great seeing you.

We have an amazing guest. Arthur Mensch is here, the CEO of Mistral AI. How are you doing, sir?

Arthur Mensch

Great. Thank you for having me.

Speaker 1

You're here at NVIDIA's big conference, and there's a big announcement. You're going to be working with NVIDIA to build models and open-source them. What is the big announcement here?

Arthur Mensch

We're announcing that we're going to be training the next generation of frontier models with NVIDIA. It's something that we've done before with NVIDIA with Mistral NeMo, something we did around 18 months ago.

The point for us is really to be able to produce the best open-source models out there so that we can use those assets to specialize them through products that we do for our customers, like Forge, which helps us customize the models for the enterprises we work with in engineering, physics, and science, and make them better at certain languages when we work with governments, et cetera.

Speaker 1

Mistral is obviously based in France. You're the leading AI company there. What's it like running the company and building a large language model in Europe? Obviously, there's regulation and all kinds of considerations around privacy. The French are known for protecting privacy. In the United States, we're known for taking it away.

How is the landscape there, and what do you have to deal with there that maybe you wouldn't have to deal with in America? What are the pros and the cons?

Arthur Mensch

Let's say, first, 25% of our business is in the US, and 25% of our researchers are actually here. I spend a lot of time here, as well as in France, the UK, and Singapore, where we are.

Of course, they're different markets—markets where language is a topic, where manufacturing is a bigger piece of the pie than it is here. I'd say our strength has also been to work with European companies that are a bit lagging behind and want to adopt the technology to leap forward.

We've been able to do that through forward-deployed engineering engagements, through our Forge product, or through our Studio product, which allows you to deploy agents that do end-to-end automation.

On top of that, the thing that we announced today, like Forge, is something that's actually being used today with customers in the US because they come to us with needs for post-training, for making models more capable, and specifically good at financial services. What's happening is that we have this product, and we can bring the models to specialize them as well.

Daniel Roberts

Your belief is that specialized, verticalized models—health care, finance, engineering, and different verticals—will win the day, or that a global model will win the day that does everything?

Arthur Mensch

You need general-purpose models to do the orchestration part, et cetera. But at some point, enterprises sit on a lot of intellectual property and a lot of signals coming from physical systems, factories, and tools. It's actually not trivial to connect those systems, to connect that data to models that are closed-source.

If you have open models, you can add new parameters. You can make a lot of deeper things that you cannot do with closed models. You can also—and that's something that we do—we not only work on the model side, but also on the orchestration side.

We sit with subject-matter experts to understand their needs, and we build business applications that are fully bespoke to their needs by modifying the models, but also modifying the harness on top.

So, we believe that eventually, building an open-source technology is a way to save costs and have better control, because you can see the thing on every cloud that you want, on your hardware if you want, and deploy it on the edge if you want.

Eventually, from a customization perspective and from leveraging your decades of IP that you've been accruing in financial services and heavy manufacturing, companies like ASML, for instance, benefit from working with us because we take their data and build models that are specifically good for them and their purposes.

Speaker 1

This training data uses experts to come in and refine a model. Most people don't know this business that well, but this has become a very large part of the industry. Obviously, Scale AI was doing it. They went to Facebook and lost a lot of the customer base who didn't want to send their data, I guess, over to Meta.

Speaker 1

We're investors in a company called micro1 that's doing pretty well in this space. There are other folks doing it. Explain to the audience what you're doing specifically for companies, how this training works in a verticalized way, and how you silo that data. If you're working with one customer in aerospace or fintech, they might have a need set, but they may not want that training to go to a competitor.

Arthur Mensch

I can give you a few examples. I think overall, the data segregation is super important. The way we have solved that is through a portable platform. Our technology is a set of services, a set of training tools, and a set of data-processing tools that I can take and put on the infrastructure of my customers.

Suddenly, from an IT perspective, when we talk to the CIOs, they realize that, from a security perspective, the flow of data doesn't go anywhere. There's no data flow coming back to Mistral because everything stays there. The way we then use that technology that has been deployed is that we're going to be working with the teams that are doing image scanning and defect detection with ASML, for instance.

We're going to send forward-deployed engineers and scientists. They're all PhDs, and they know how to train models. They spend some time with the subject-matter experts who can explain how an image is being detected, how you detect defects, and so on. Based on that, we're going to work out what kind of data needs to be used to train the models that are going to solve the task itself.

We send the technology and, typically, a few scientists because you do need that expertise transfer and knowledge transfer between our teams and the vertical experts. Then we make sure that eventually our team no longer needs to be there to retrain the models, get more data access, and so on. That combination of data segregation, expertise transfer, and knowledge transfer is the one thing that makes us quite unique and allows us to serve the most critical use cases and processes in industries that need to take their data and put it into models for them to work.

Speaker 1

It seems that once we've exhausted the entire open web—what was available legally, on the gray market, and so on—I wouldn't have you comment on that controversy, but we've kind of exhausted what's in the open crawl, yeah?

Arthur Mensch

We have.

Speaker 1

And it's time to either make synthetic data or use experts. Do you believe in synthetic data, and where does that work and where does it fail?

Arthur Mensch

We use synthetic data as a way to warm up the models. It's a way to be quite efficient at the beginning. If you have a large model and you want to train a small model, you will use your large model to process and produce a lot of synthetic data at the beginning.

Eventually, though, you do need to have human signal. Human signal is always a bit costly to acquire because you need to talk to the experts, and they need to give feedback to the machines. At the beginning, synthetic data allows you to do the compression, to further compress the models.

At the end, you do need to go and get data that is produced by humans. So, it's mostly an efficient way of training models, with bigger models used as teachers for smaller models, but it's not enough. You also need human signal.

Speaker 1

Arthur, we've seen an incredible explosion. We're sitting here 52 days after OpenClaw, the year of our Lord, 52 days in. When you first saw OpenClaw and saw the reaction of hackers, founders, and startup CEOs—the amount of energy, and seeing it race to the top of GitHub with the most stars and likes, along with all these contributors—what did that say to you as an executive in the space who's been grinding on this for many years? What did that OpenClaw moment mean?

Arthur Mensch

It resonates a lot with what we are doing with our customers because, pretty quickly, enterprises realized that if they wanted to make some gains with artificial intelligence and generative AI, they would need to automate full processes. To automate a full process as an enterprise, you can use OpenClaw, but it's actually not really enough because you have data problems and governance problems. You can't observe the process that's running, and in many cases you can't control it.

When you run a KYC process, if you're HSBC, for instance, one of our customers, you will want to have deterministic gates that are always going to do the same thing, in a way that is observable and allows you to guarantee to the CEO that it's always going to go through these gates. That's not something OpenClaw is providing because it doesn't have the kind of primitives that you need to work on collective productivity, observable productivity, and mission-critical systems.

On the other hand, the autonomy it gives and brings to people who are just individuals hacking things together is also a way to show enterprises that if you set up the right control plane and the right sandboxes, connect to the right data sources, and make sure that your access controls are well respected, then you can actually unleash the power of agents doing things for your employees. That can work on the platform; otherwise, you will not be at ease when you're sleeping.

Speaker 1

It is definitely something you have to be thoughtful about. When I installed it, I gave my agent root access to my Google Docs, my G Suite, my Notion, my Zoom, and my Calendar—everything.

Then I realized, "Wow, with my enterprise edition of Gmail, I can essentially summarize every conversation going on in Gmail for my entire 21-person investment company, and then correlate it with every conversation in Slack." Then I realized, "Oh my gosh, there are compensation discussions going on. There is a person on a preferred performance improvement plan, or something like that."

I have to make sure nobody else can access this because the power comes from giving it access to data, but with great power comes great responsibility. I think people are learning that in real time.

Arthur Mensch

It's a big problem because enterprise data is not a single thing that you want to put into a single system that's going to be accessible by everyone. You need to have this layer that actually understands what is in the data. You need to have a semantic of what can actually be exposed to HR or what can be exposed to engineering.

Typically, compensation is one of these things. You want to make sure that compensation data does not flow back to all of the enterprise because you're going to have a lot of problems if that's the case.

What you actually need, and which is hard to do, is what we call a context engine: a mapping of where the data sits that comes with a certain amount of metadata telling you that this data is not accessible to this part of the company. If someone in engineering is asking for something related to compensation, the system is going to tell you, "Look, you actually can't access that data."

That's hard. It's actually hard. You need to rethink entirely the way your IT systems are being connected. At some point, you also need to think about your management because your information flow is completely different today if you're connecting agents together with your data sources than it used to be.

Suddenly, maybe you don't need that manager whose only purpose was to take information from the bottom and put the information at the top. There are so many problems to solve. You need the right primitives, you need sandboxes, and you need role-based access control and these kinds of things. You have changes to make.

You need to rethink your entire customer-service department because suddenly you don't need that much transfer of information operated by humans.

Speaker 1

All right, you have to go—you've got a flight to catch. It is so great to see you, Arthur. Continued success with Mistral.

I'm really lucky to have Daniel Roberts here. He's the co-CEO and co-founder, along with his brother, of IREN. They're a publicly traded company. They started in BTC.

Daniel Roberts

Thanks. Pleasure to be here.

Speaker 1

You and your brother started in Sydney 7 or 8 years ago, and you got in early on Bitcoin. All these Bitcoin miners wanted to have data centers, huh?

Daniel Roberts

That's directionally right. The thesis we saw was this explosion of the digital world and the growth in the online world, and at some point the real world was going to struggle. So, we set about to build out large-scale data centers.

Yes, the first use case was Bitcoin mining, but as we said to our seed investors, use that to bootstrap the platform, generate cash flow, and layer in higher and better use cases over time as they emerge. Here we are today with AI. We are swapping out all the Bitcoin for AI chips.

Speaker 1

When did you first start seeing the demand in the company shift from, "Hey, Bitcoin miners, we need some H100s, whatever it is," to, "Hey, we're this nonprofit OpenAI. Hey, we're this research lab. We need some AI compute"? When did that start hitting?

Daniel Roberts

Look, we had a bit of a false dawn, I would say, back in 2020. We signed an MOU with Dell to start bringing on customers and compute, but in hindsight it was too early. So, we went back to Bitcoin and kept bootstrapping the platform.

I would say about 2 years ago, and month by month, the demand just continues to escalate.

Speaker 1

And you were in so early that when you were looking at data-center space in the United States, you were one of 1 or 2 or 3 people looking at the space.

They were trying to sell you on space, yeah?

Daniel Roberts

Yeah, so we actually developed the data centers ourselves. We go and find the land, get the permits, and apply for grid connections, and we were doing it at a scale that just amazed people at the time. Our flagship Texas site is 750 MW. Four years ago, that was unheard of. In the middle of the desert, we're building these big data centers, and the traditional data center industry was going, “What are you guys doing?” We said, “We believe in the future of digitization, high-performance computing,” and obviously, today, it's paying dividends.

Speaker 1

I don't think anybody could have predicted, when ChatGPT came out, OpenClaw recently as a turning point, and then Microsoft, Google, and everybody embracing this. That's your big partner, Microsoft.

Daniel Roberts

Yes, Microsoft is one of our early partners. We signed a $9.7 billion contract with them late last year. But, as I was explaining to you before the show, that's 5% of our capacity. So, things are busy at the moment.

Speaker 1

And when you do these build-outs, the big conversation today is no longer the number of GPUs we're putting in; it's just power. Power is the constraint today, yeah?

Daniel Roberts

For many in the industry, it is. But for us, because we started eight years ago tying up all this land and power, it isn't. We've got 4.5 GW. For context, that's almost as much power annually as the Bay Area uses in its entirety each year. It's huge. So, for us, the hurdle, or the constraint, is really time to compute, and that's emerging across the industry as well.

Speaker 1

And time to compute means tradespeople coming to West Texas, living in a trailer that you set up, and then breaking ground on a data center, building foundations, and building water-cooling systems. This is hard manual labor going on, yeah?

Daniel Roberts

Exactly. And this is the whole real-world challenge of responding to these digital exponential demand curves. They're unconstrained by the real world in terms of their appetite, and it just compounds. You need thousands of people out in these locations that haven't supported it. You put stress on supply chains. We're seeing what's happening with memory—every aspect of it. So, it's permanent whack-a-mole, permanently putting out fires to try and bring this compute online.

Speaker 1

And you get to spend time there. What's it like when you set up a town or bring 1,000 or 2,000 people to what's pretty much a remote small town? I'm assuming that when you bring 1,000 people, there might only be 500 living there right now. What are those towns like? It sounds to me like something out of the gold-mining era, when people first went and were prospectors. It's a prospecting town?

Daniel Roberts

Pretty much. I mean, the barbecue's great. That was the drawcard. Then, apart from that, we've always had a policy of hiring local and supporting the local community. This year, we're hitting $1 million in community grants cumulatively. That's things like local playgrounds and supporting the fire departments. But we will hire locally. Once we can't find that trade locally, we'll expand the radius by 20 miles and hire out of that, and so on and so on.

Speaker 1

That's very thoughtful. These folks are coming—say, an electrician or a construction worker—and they've built houses or maybe corporate offices. Now they come for a tour of duty here, and the salaries go up massively, but they have to leave their family for a 3-month tour or something?

Daniel Roberts

Yes and no, because typically, where we locate is where there's heavy electrical infrastructure. Where there's heavy electrical infrastructure is typically where old manufacturing and industry have closed down. So, we go in, leverage that sunk CapEx, rehire and retrain local workforces, and bring a new industry to town in these data centers.

Speaker 1

Has that workforce now been completely depleted? Do we need to train another generation, a younger generation, to really embrace the trades?

Daniel Roberts

100%. We're partnering with universities and trade colleges, absolutely.

Speaker 1

And you go to a trade school, you go to a college, and people are getting degrees in philosophy and English literature. They're going $50K a year into debt, $200K a year into debt. What's the starting salary for a tradesperson working on a data center, doing electrical or construction?

Daniel Roberts

Exactly.

Speaker 1

What's the ballpark range?

Daniel Roberts

Look, I won't talk specifics, but they are going up. The price is going up. It depends on the level, but yes, there is a rush for good tradespeople.

Speaker 1

I'm hearing $150K to $300K. Am I in the ballpark?

Daniel Roberts

At the lower end, directionally, you're right.

Speaker 1

It's incredible when you think about it. There's concern about, “Hey, AI is taking jobs,” and then, on this other side of the ledger, we can't find enough talent to service it. Talk to me about energy sources and how you think about that. President Trump, Chris Wright, and the administration started with, “Hey, clean, beautiful coal.” Year 2, they're like, “All sources matter. Nuclear.” Obviously, natural gas is plentiful in that area, and we've got a lot of oil. People don't know this about Texas: in the United States, it's the number 1 source of solar installations. Talk to us about energy.

Daniel Roberts

Our philosophy has been sustainability from day 1. We've used 100% renewable energy since inception.

Speaker 1

What? 100%? Wait, how is that possible?

Daniel Roberts

We use hydro in British Columbia, and we use wind and solar in West Texas. In West Texas, there's around 45 to 50 GW of wind and solar. The transmission line to export that down to the load centers in Dallas and Houston is 12 GW. So, you locate to the closest source of low-cost excess renewable energy, monetize it into this digital commodity, and export it at the speed of light as a token.

Speaker 1

Great arbitrage. The wind is producing a lot, but it's harder to get the power from those areas where people are willing to put it up. People don't understand how big West Texas is. It's an incredible amount of land. And you're coming from Australia, where, on the west side, people also don't understand exactly how much pure undeveloped land there is.

Daniel Roberts

So much land. And the issue is distance. You've got to spend billions of dollars on this transmission-connection infrastructure to move that power to where people actually want it. You can build wind farms and solar farms, but if you build them in the desert and no one can use them, then what's the point? So, the whole opportunity for our industry is to go to the source of that power and monetize it. The data centers follow the wind turbines and the solar installations.

Speaker 1

How do you think about batteries? Are you able to put those online? Obviously, you're going to have periods where it's not a windy day. In Texas, we have very few days when it's overcast, so that problem's pretty much solved, but you're going to have 50 days where the sun's not beating down. How do you deal with the demand and soften that duck curve?

Daniel Roberts

We don't need to.

Speaker 1

Huh.

Daniel Roberts

The utility does that on our behalf. This is why these grid connections are so scarce, so hard to get, and so highly valued, because once you get that grid connection, the utility underwrites all of that variability. They guarantee you 24/7 reliable power.

Speaker 1

Got it. So, on their side, they're figuring it out. If something goes down, they could fall back, even though you're 100% committed to renewables. If they needed to fall back to gas or whatever, they have that ability out there, so you have that as a backup. There's a lot of talk, or a debate, about whether we're getting ahead of our skis and whether people are slowing down. There was some talk about the OpenAI project maybe downscaling a little bit. Is OpenAI a partner as well, or—

Daniel Roberts

I can't comment.

Speaker 1

Can't comment. Okay, so we'll read into that whatever we want. Are there pockets where people are saying, “Hey, let's slow down,” or is it still gangbusters?

Daniel Roberts

It's at the upper end of the spectrum. It's gangbusters. We cannot meet demand. That's why the whole industry now is around time to compute. There are no idle GPUs in the world sitting in a data center.

Speaker 1

What's your take on what happens when software makes things more efficient? This was a big discussion from Jensen himself during his 2.5-hour keynote yesterday. We're sitting here Wednesday; I think he did his keynote on Tuesday. He was talking about, “Hey, software is going to make it 50 times more efficient and lower the cost of tokens 50x.” Then you have transport also contributing to that. When do you think the curve goes from parabolic to simply growing at a ridiculous level? Is there a slowdown coming, or how are you planning for the future?

Daniel Roberts

Look, I think it's actually the opposite. I think it feeds on itself. I'll give you 1 example. You go into ChatGPT today and generate an image. You hit enter on the prompt, and it's like the dial-up internet days. It takes minutes, and you're like, “I better get this prompt right.” Finally, 2 minutes later, it comes. Now, I'll give you an example: if we 10x the amount of compute available—which is an enormous task from where we are today—and those images take 5 to 10 seconds, are we going to generate more or fewer images? Many more. This is Jevons paradox. This is the theory of induced traffic. You build a couple more lanes, and people start to think, “Well, maybe the distance from Bondi Beach to the central business district in Sydney would be an acceptable commute.”

Speaker 1

Love the analogy. Yeah. What do you think about, or what are you seeing? We're here at NVIDIA. Obviously, they make the leading-edge chips. They just bought Rocks, so now you've got 2 of the leading-edge chips coming out of the same company. But custom silicon is becoming a big discussion. Has that started to land in the data centers yet? Obviously, Google—I don't know if they're a customer you can tell us about—but they're making custom silicon.

Amazon is making custom silicon. Meta is making custom silicon. Talk to me about that revolution. Is it actually making it to the data centers yet?

Daniel Roberts

Look, to various degrees, it is. They're promoting their products. They're trying to tie up data center capacity. So, yes, there's multiple silicon looking for homes. I think it's fair to say NVIDIA has a massive head start.

The ecosystem they've incubated and the standards that they're setting mean I would say the safest pathway to build out at scale early is to follow the NVIDIA roadmaps. But absolutely, over time, we are seeing these chips emerge.

Speaker 1

And in terms of desktop computing, I know there was a survey announcement that Dell and NVIDIA are making a really powerful desktop—750 GB of RAM and a lot of power. You're going to be able to run some local models, open source, with OpenClaw and open-source models coming from Kimi and a bunch of the models out of China. And in the hacker group—which I think you started in, like I did, probably around a similar time—people are starting to get really obsessed with having a $10,000 or $20,000 desktop setup and running this locally. What do you think of that trend? I'm curious.

Daniel Roberts

Yeah, the breakthroughs we're seeing in software—the way it's distributing power to every man and every woman in every house, and their ability to code and use products like OpenClaw—the generation of demand and appetite for compute at the local level, all the way through to these mega data centers, it's absolutely real. And as we see the emergence of agents using more and more compute, and as we see autonomous vehicles and other automation and robotics, it's absolutely going to compound.

Speaker 1

And what about nuclear? The Trump administration really seemed to flip the switch on a growing belief that, hey, wait, nuclear's pretty great. It's clean. It's the original renewable, in a way. These new modular reactors have nothing to do with Chernobyl, Fukushima, or Three Mile Island. They're much safer. They're a completely different architecture.

Have those started to land yet? And since you correctly followed that trend in the great state of Texas, where I'm from, are you following nuclear?

Daniel Roberts

I think you have to. I think the reality is it's going to take a decade, a bit longer, by the time big projects can come into commissioning. But now is the time to start that conversation, put in place policies, mobilize capital, and start that ball rolling.

Speaker 1

Do you have a data center going up near nuclear?

Daniel Roberts

No, not at the moment.

Speaker 1

But you're actively tracking that activity?

Daniel Roberts

Yes. Yeah, this seems pretty inevitable.

Speaker 1

Feels like it. And if that happens, what impact does it have on your industry? Obviously, it's happening in China, and people always point to the Bitcoin miners—they were like the canary in the coal mine—near the hydro dams and near the nuclear plants where there was excess capacity. What impact do you think this has if you could actually have small modular reactors next to data centers?

Daniel Roberts

Well, I think it just opens up the market and enhances the US's competitive advantage in this space. AI is inevitable. Robotics is inevitable. The reality is the correlation between human progress and energy consumption is really, really high over a very long time period.

So, if we can find a way to unlock new generation—clean generation, as nuclear—and locate that more at the source and enable more compute on a distributed basis, all those use cases we just discussed become easier, more fluid, faster. Then you get that positive flywheel around Jevons paradox and demand.

Speaker 1

Talk to me about the architecture today of Ethernet and data moving between data centers and within data centers. That backbone is going through a paradigm shift as well, yeah?

Daniel Roberts

Yeah, it is. Jensen coins the term, “The data center is the new computer.” You need to step back and say, “Right, this big building is essentially the old desktop PC we had under our desk at home.” You go, “Right, how does that work?”

All the cabling, the latency, the number of hops between each GPU, how they talk to each other, and the fabric around InfiniBand and Ethernet—it’s absolutely critical because every millisecond matters in terms of the performance of that cluster.

Speaker 1

And where do you think—or what do you think of Elon’s vision? It’s obviously a longer-term vision of putting data centers in space, and there are a couple of other people working on it as well.

Daniel Roberts

I mean, it’s very hard to argue with Elon. He’s been very right on a number of things for a very long time. I think sitting here today, it feels exceptionally difficult, given the cost of moving things to space and the challenges around radiation. There’s a huge amount of engineering challenges, but that’s never scared Elon before.

He’s uniquely qualified, and he’s inevitably right, but sometimes he’s late. He might be late to the party, he might show up at dessert, but generally he nails it.

Speaker 1

How much of an issue is getting the data out of the data center to consumers today? Is that not something people are worried about when you’re building something out in West Texas? All that data and fiber—all that’s been taken care of, or does that become a gating issue at some point?

Daniel Roberts

So, this was one of the big myths that we had to bust when we started this business, because everyone said, “Data centers must be located close to population centers, metropolitan areas. Latency is really important.” And we said, “Yeah, that’s right. Latency is important.”

But the reality is, in the US, Texas especially, there is fiber everywhere underneath the ground. Lots and lots and lots of it. And when you look at latency from our site, in the middle of the desert in West Texas, down to Dallas, the big carrier hotel—

Speaker 1

Yeah. 6 ms round-trip latency. What’s 6 ms? There are 1,000 milliseconds in a second.

Daniel Roberts

Yeah, we’re talking six. It’s effectively adjacent. It’s not even—yeah, it’s definitely not material.

Speaker 1

Listen, continued success, and you’re hiring a lot of people.

Daniel Roberts

Yeah, I think we’ve got 129 job advertisements up at the moment.

Speaker 1

The company’s doing fantastic. Thanks for spending some time with us here at GTC.

Daniel Roberts

Thanks, Jason. Appreciate it.

Four CEOs on the Future of AI: CoreWeave, Perplexity, Mistral, and IREN | BidClub