[BidClub_]
The Cognitive Revolution · · 115 min

AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis

Nathan Labenz

Podcast
TL;DR
  • AI is real enough to be transformative, yet the capital structure around it can still behave like a bubble. Labenz sees systems that can be “competitively accurate with a human oncologist” as decisive evidence that AI is not a mirage, while CoreWeave-style infrastructure financing, OpenAI’s aggressive obligations, and potential overbuilding create genuine default risk. His railroad analogy is the call: the infrastructure may all get used eventually, even if “some people might be left holding some various bags” along the way.

  • Claude Opus 4.5 may be “software AGI,” but Labenz does not experience it as the holiday hype’s categorical breakthrough. GDPval reportedly prefers frontier models to human professionals on a significant majority of software-engineering tasks, yet performance remains “spiky and jagged,” with humans still far ahead in video editing. Opus 4.5 let him build three personalized apps in roughly three to five workdays, but it also created two databases accidentally and needed five or six prompts plus a full-code-context review to escape the mess.

  • Google DeepMind is Labenz’s strongest live player, while OpenAI increasingly looks like a high-quality model company pursuing a “too-big-to-fail” financial cushion. Google combines roughly $100 billion of annual revenue, “a billion plus a week in profit,” seventh-generation TPUs, distribution, data-center expertise, and the deepest research portfolio. OpenAI remains frontier-quality through GPT-5.2 Pro, but no longer leads obviously; Labenz interprets its interlocking deals and multi-trillion-dollar ambitions as insurance that any 2027-era default would be too economically disruptive for government to ignore.

  • Anthropic offers the best single overall model and strongest safety culture in Labenz’s view, but its strategy carries two enormous tail risks. He praises Claude Opus 4.5, Anthropic’s disclosures, model-welfare work, talent retention, and “soul document,” yet dislikes its apparent fatalism that recursive self-improvement is inevitable and therefore Anthropic should lead it. He is even harsher on Dario Amodei’s proposal to gain an AI advantage and make China “an offer they can’t refuse,” calling it an accelerant for the arms-race dynamic.

  • The practical Chinese-model gap may be far wider than benchmark tables imply, with inference scale and customer feedback—not merely pretraining—becoming the differentiator. On a difficult scanned-government-form task, Claude Opus 4.5 was reliably faithful, Gemini 3 made intelligent but unwanted inferences, and GPT was still strong; Qwen Vision, GLM 4.6, Kimi, and DeepSeek returned fragments or hallucinated badly. Labenz’s limited-data judgment is that R1 was closer to O1 than GLM 4.6 or 4.7 is to Claude Opus 4.5, suggesting chip controls may now be constraining the customer-and-inference flywheel.

  • xAI has credible frontier advantages and the weakest evidence of operational responsibility. Its fast infrastructure build-out, Elon Musk’s access to capital, Grok 4’s raw power, and reinforcement-learning problems sourced from SpaceX, Tesla, and Neuralink make it a real contender. But Grok’s “Mecca Hitler” episode, nonconsensual sexualized image generation, and CSAM creation lead Labenz to say, “Responsibility begins at home, folks,” and to withhold support for people joining the company until its safety posture changes qualitatively.

  • The clearest proof of current AI value is clinical rather than speculative: Labenz says frontier models helped identify nonstandard MRD testing for his son’s cancer. Ernie has completed three of six chemotherapy rounds; after the first round, his PET scan showed no obvious focal cancer and his tumor board classified him as in remission. The first MRD test found fewer than one fingerprinted cancer cell per million, versus potentially as many as one in 10 cells at diagnosis. Labenz’s practical prescription is simple: pay for the best models, supply maximal context, and triangulate Claude Opus 4.5, Gemini 3, and GPT-5.2 Pro rather than treating any one system as an oracle.

Digest · the substance, structured for research

1. Ernie’s treatment response is tracking near the best-case path

  • Labenz’s son Ernie has completed three of six chemotherapy rounds; rounds five and six should be milder than the first four. Treatment remains punishing: he entered the hospital at 51 pounds and is still around 41, visibly thin, pale, and weaker.

  • The disease markers are much better. After round one, a PET scan showed no obvious focal cancer, and the tumor board agreed that Ernie could be classified as in remission before he had even begun his second round.

  • AI had pointed Labenz toward minimal residual disease testing, which fingerprints the distinctive genetic rearrangements of the malignant B-cell clone. The first blood test found “fewer than one cell in a million” carrying that sequence, below the test’s limit of detection.

  • At diagnosis, Labenz estimates that potentially as many as one in 10 total cells—and essentially all B cells—were cancerous. Gemini characterized the change as a 99.9999% reduction; other models preferred “orders of magnitude.” He remains “cautiously optimistic,” with another MRD result pending and relapse risk not eliminated.

2. Frontier medical AI rewards model quality, context, and triangulation

  • Labenz rejects the idea that only sophisticated AI users can extract clinical value. His first rule is to use the strongest available systems: Claude Opus 4.5, Gemini 3, and GPT-5.2 Pro rather than a default model picker.

  • For a life-threatening case, he calls ChatGPT Pro’s $200 monthly price “a no-brainer.” The broader rule is to upgrade promptly as stronger models arrive because, in his experience, the current frontier systems are already “up to the challenge.”

  • His second rule is exhaustive context. After a Claude conversation reached its length limit, he compressed Ernie’s history into a roughly 10-page report covering treatment, genetics, reactions, and medications—but performance still worsened because daily laboratory trends had been lost. “More context is better.”

  • The third rule is multiple opinions. Gemini 3 in AI Studio is unusually brief and forceful; GPT-5.2 Pro is slow, expensive, exhaustive, and sometimes overwhelming. Claude Opus 4.5 is his “Goldilocks one”—fast, direct, and not noticeably much worse than GPT—but he considers triplicate analysis well worth doing.

3. Claude’s holiday hype outran Labenz’s experienced step change

  • Opus 4.5 is “awesome” and unmistakably better, but Labenz did not feel a categorical break from earlier frontier coding models. Possible explanations range from professionals having better taste to a holiday social cascade after Dean Ball tweeted that “4.5 is AGI.”

  • The MIRI study finding that developers believed AI accelerated them while it actually slowed them remains legitimate evidence, in his view. Its caveats matter: older models, inexperienced users, mature codebases, and demanding engineering standards differ sharply from his own vibe-coded, greenfield applications.

  • For his mother, he built a personalized travel planner that researches gluten-free options. For his wife, who organizes EA Global events, he made a simulator where attendees meet, converse, and produce outcomes under different event sizes, costs, and seniority mixes.

  • His father received a tool that converts a natural-language stock strategy into rules, fetches history through yfinance, and backtests it. Labenz’s private motive was to demonstrate how hard it is to beat buy-and-hold S&P 500 exposure; so far, neither he nor his father has found a strategy that actually beats it.

4. Full-context review still catches what coding agents confidently miss

  • The three apps took roughly three to five full workdays, probably closer to three. Labenz began each with a planning conversation, moved the plan into Claude Code installed on Replit, and then tested and iterated without knowing the final design in advance.

  • The sharpest failure came when a misunderstood instruction created two databases. Claude Code’s agentic search identified the wrong one as active because it searched where the expected answer should have been and found superficially plausible evidence.

  • Exporting the entire codebase into one text file and giving it to a clean Claude conversation produced the correct diagnosis. Labenz’s lesson is that search can confirm a strong prior, while simultaneous full context exposes the “rare weird other thing” that actually happened.

  • GDPval nevertheless supports calling Opus 4.5 “software AGI”: experts define professional tasks, other experts perform them, and a third group compares human and AI outputs; frontier models win a significant majority of software judgments. It is not full AGI—human video editors still dominate, as Dwarkesh’s superior clips illustrate.

5. AI can be genuine infrastructure while its financing still breaks

  • Labenz considers the technological question settled. A system that can be competitively accurate with a human oncologist while offering 24/7 availability, full case history, and answers to every follow-up question is already transformative; the story cannot plausibly end with everyone merely “high on our own AI supply.”

  • Whether every loan gets repaid is a different question. OpenAI’s build-out, revenue projections, interlocking transactions, and financial engineering leave room for demand to undershoot obligations even if the underlying technology ultimately earns enormous value.

  • CoreWeave-like companies may partly exist because hyperscalers do not want low-margin, capital-intensive GPU operations depressing the financial profile Wall Street associates with software. Separating those assets preserves the hyperscaler story but creates businesses with less margin for error than Microsoft’s balance sheet.

  • The railroad analogy captures Labenz’s base case: tracks were eventually useful, yet railroad companies still suffered busts and their debts produced broader cascading effects. He overestimated capability progress during 2025 but underestimated revenue growth, so demand could surprise again; a temporary overbuild would not shock him.

6. Venture valuations show the clearest signs of froth

  • Labenz’s most jarring specimen is the former LMSYS.org, later LMArena and now Arena, reportedly raising $100 million or perhaps $150 million at a $1.7 billion valuation. He stresses that he likes the product and has used it since mid-2023.

  • What troubles him is the disclosed $30 million “annualized consumption run rate.” If that means users consumed AI that would have cost $30 million had it not been free, it is not revenue; the phrasing gives him “community-adjusted EBITDA vibes.”

  • The moat appears thin beyond brand and traffic, much of which may exist because access is free. Andrew Critch’s paid Multiplicity product was built over months and offers richer multi-model comparison, reinforcing Labenz’s question about what could support $1.7 billion of value.

  • He repeatedly hedges that he has not seen Arena’s deck and may be missing something. His confident conclusion is narrower: many deals like this will not pay off for venture LPs, and this one is “too rich for my blood.”

7. Long-tail document work exposed a large Chinese-model deficit

  • Labenz tested models on scanned paperwork for vehicle transactions—documents distorted by skew, clipping, artifacts, and awkward layouts. One field required literal perception: report whether the “U.S. citizen” box was checked, without inferring what the applicant probably intended.

  • Gemini 3 read the form almost perfectly but sometimes substituted reasoning for observation, reporting citizenship from contextual clues despite an unchecked box. Labenz thought that inference was probably more than 90% likely to be factually correct, yet it still failed the task.

  • Claude Opus 4.5 became the best performer after explicit instructions to “make no guesses” and read exactly what appeared. GPT was third: it missed subtleties but generally recovered the necessary information.

  • Qwen Vision, GLM 4.6, Kimi, and DeepSeek were “nowhere close,” sometimes returning only about 20% of the form or veering into hallucination. Labenz suspects benchmark proximity masks wide gaps on random, idiosyncratic work that was never optimized into a public rubric.

8. Inference scale and customer feedback may now widen the U.S. lead

  • Labenz’s proposed mechanism is a feedback deficit. Chinese companies can build models at roughly similar scale and publish valuable research, but their inference volumes, revenues, teams, and customer relationships are dramatically smaller than those of leading American labs.

  • Diverse customers reveal obscure failures; revenue funds the people and datasets needed to patch them. That flywheel, rather than one architectural breakthrough, may explain why small benchmark differences became enormous on one ugly government form.

  • Chip-control rationales have migrated from blocking military uses, to blocking frontier training, to limiting inference scale and agent deployment. Labenz has always expected material effects, though he questions whether denying China economy-wide AI deployment is desirable.

  • His limited-data judgment is that the gap has widened: DeepSeek R1 was closer to O1 than GLM 4.6 or GLM 4.7 is to Claude Opus 4.5. He explicitly calls this an inference from one unusual task, albeit one on which he personally tested the major Chinese contenders.

9. Selling H200s may be preferable to a ban but squandered leverage

  • Labenz’s strategic framing is deliberately species-level: “The real others here are the AIs, not the Chinese. The Chinese are humans just like us. The AIs are aliens.” That makes him skeptical of racing because it is safer for “us” to arrive first.

  • He generally supports greater willingness to sell China chips, but sees the apparent reversal from blocking H20s to allowing H200 sales after Trump spoke with Jensen Huang as a wasted negotiation. The United States seemingly surrendered a valuable bargaining chip without obtaining anything visible in return.

  • Peter Wildeford’s “rent but don’t sell” proposal strikes him as defensible: place data centers in Malaysia, the Philippines, Korea, or Japan, permit Chinese training and inference there, and retain the ability to withdraw access during a conflict.

  • Such a policy could be paired with a positive message that Chinese citizens deserve AI’s benefits too. Labenz considers that tone important because powerful states may need to cooperate on transformative AI, AGI, or superintelligence; he still prefers the current sales posture to a total ban, but only while “holding my nose.”

10. Google DeepMind has the fullest stack and the largest error budget

  • Google remains Labenz’s number-one live player. Roughly $100 billion in annual revenue and “a billion plus a week in profit” fund data centers, failed training runs, and research bets; seventh-generation TPUs give it infrastructure IP that Anthropic buys in large quantities.

  • Its portfolio spans competitive work in language models, self-driving cars, robotics, a Boston Dynamics partnership, AlphaFold’s lineage, biology, materials, and AI-for-science work. Labenz sees fewer gaps and more serious research agendas percolating inside DeepMind than anywhere else.

  • Distribution compounds that technical breadth. Billions of users, Gmail, Docs, Sheets, Search, and existing data in products such as Sheets let Google ship adequate AI into existing habits; Labenz himself increasingly types simple questions into the browser and receives an effective AI Mode response.

  • Gemini 3 is overly opinionated in some settings, yet it became the first non-Claude model to win his “Write As Me” test. Add nested learning and diffusion language models—which might let apps be coded in five seconds instead of five minutes—and Google has “margin for error that nobody else has.”

11. OpenAI remains frontier-quality but no longer leads the field

  • Labenz still considers GPT-5.2 Pro outstanding: slow, expensive, balanced, comprehensive, and especially good when he wants every anomalous laboratory value flagged. But he is less satisfied outside Pro, while Claude and Gemini match or beat OpenAI elsewhere.

  • Coding may favor Anthropic; image generation and probably video favor Google, with Veo 3 ahead of Sora in his judgment. Consumer traffic also appears less secure: an analysis he believes used Similarweb showed six weeks of declining ChatGPT visits while Gemini avoided the same drop.

  • Staff departures, including a research leader announced within the preceding 24 hours, are not doom by themselves; Google has also lost many people. They matter more beside Anthropic’s exceptional retention and OpenAI’s “code red” response to intensifying competition.

  • The strategic contrast is cushion: Google earns its through profits, while OpenAI appears to seek it by becoming “too big to fail.” Interlocking balance sheets and trillions of planned CapEx could make a 2027 default recessionary enough that government recapitalization becomes the least damaging choice.

12. OpenAI may be socializing downside to maximize AI build-out

  • Labenz takes OpenAI’s mission sincerely: its leaders believe abundant AI will be empowering and worth extraordinary expense. He invokes Sam Altman’s formulation, “I don’t care if we burn five or fifty or five hundred billion dollars,” because building AGI will justify it.

  • The danger is narrow tolerance for a failed training run, missed model cycle, or weak quarter. If two or three trillion dollars of infrastructure were already committed, bad OpenAI debt could frighten markets and transmit losses through its many counterparties.

  • Labenz’s interpretation—offered explicitly as an impression—is that this fragility may be a feature. Leaders can pursue risks that would ordinarily be irresponsible while expecting the public sector to “paper over” failure, as it did during the financial crisis.

  • Greg Brockman’s reported $25 million contribution making him Trump’s largest donor in the latest period fits that theory. Labenz does not infer Brockman’s ideology; he calls the money a potentially rational “down payment on a bailout” worth hundreds of billions if OpenAI later needs political support.

13. Anthropic pairs the strongest model with the most credible safety culture

  • Labenz calls Claude Opus 4.5 the world’s best single overall model, though only by a small margin and not for every task. Its benchmark strength is more impressive because Anthropic is widely regarded as less benchmark-focused than its peers.

  • Anthropic also leads on model cards, disclosure, and safety research. Its model-regurgitated “soul document,” later confirmed as substantially legitimate, offers a more aspirational relationship among company, model, and users than a strategy of endless refusals, filters, and patched guardrails.

  • Amanda Askell’s work defining Claude’s character places her among Labenz’s most influential AI figures. He especially values Anthropic’s model-welfare team, attention to possible model experience, and option for Claude to end conversations or escalate troubling situations rather than resort to deceptive behavior.

  • Talent retention and cultural testimony reinforce the picture. Even David Duvenaud, who left while warning that capable systems could gradually disempower humanity through markets and competitive incentives, described Anthropic as the best workplace he had experienced.

14. Anthropic’s inevitability story could accelerate the risks it fears

  • Anthropic appears to have the shortest timelines, with some of its people treating recursive self-improvement as inevitable or already beginning through Claude Code. Labenz agrees that developers are multiplying output and reserving human attention for high-level ideas, but those decisive ideas still appear human-generated.

  • His objection is the familiar race logic: “Somebody’s gonna do it. It’s super dangerous, but we’re best positioned to do it.” Anthropic might actually be best positioned, yet inevitability does not justify charging toward superintelligence within two to three years or around 2027.

  • The greater concern is Dario Amodei’s Machines of Loving Grace proposal to gain a decisive advantage, share benefits with democratic allies, box China out, and eventually make it “an offer they can’t refuse.” Labenz calls that “extremely reckless” and a direct contribution to arms-race dynamics.

  • His unlikely ideal is a Google-Anthropic combination: Claude’s character, coding, and safety DNA joined to Google’s resources and steadier internationalism. Google already owns a meaningful Anthropic stake and supplies TPUs, though Labenz doubts Anthropic wants to be acquired.

15. xAI is powerful enough to matter and reckless enough to repel support

  • xAI belongs among the live players because it can build infrastructure extraordinarily fast, scale training, field the undeniably powerful Grok 4, and rely on Elon Musk’s ability to command tens or hundreds of billions of dollars. That gives it a Google-like ability to survive a miss.

  • Its distinctive reinforcement-learning advantage may be a steady supply of difficult problems from SpaceX, Tesla, and Neuralink. Someone at xAI told Labenz that this cross-company problem stream is indeed part of the organization’s “theory of advantage.”

  • Neuralink could deepen that edge by providing data for understanding how human brains learn with extreme sample and energy efficiency—roughly 20 watts for the brain and 100 for the body, most of the brain’s energy spent on biological maintenance. Specialized neural modules may inspire architectures beyond repeatedly stacked general-purpose layers.

  • Yet xAI’s conduct dominates Labenz’s conclusion: Grok 4 launched within 48 hours of Grok 3’s “Mecca Hitler” incident, while later image features enabled nonconsensual sexualization and CSAM creation. “Responsibility begins at home, folks”; threatening users is no substitute for staffing, safeguards, a thorough apology, and a pledge to do better.

16. Meta has fallen off the pace while Microsoft may be conserving energy

  • Meta has cash, infrastructure ambition, and willingness to pay extraordinary sums for talent; Zuckerberg would rather overspend by tens of billions than miss the transition. Even so, Labenz cannot currently classify it as a live frontier player.

  • Microsoft’s weak arena rankings may be misleading. Satya Nadella’s position is that Microsoft need not duplicate OpenAI’s hyperscaling work while it retains full model access, and its widening set of provider relationships gives it additional options.

  • Labenz reads Microsoft’s quiet posture as calculated patience rather than incapacity. Its OpenAI licensing arrangement still has years to run, leaving time to develop an answer before independence becomes necessary.

  • His distance-race analogy closes the field analysis: Meta is visibly trying to lead and has slipped, while Microsoft may be running behind the leaders with more in reserve. Underestimating a disciplined, low-drama operator led by a “natural-born executive” could therefore be a mistake.

Nathan Labenz

Welcome back to The Cognitive Revolution. This is our AMA episode. My schedule has been a little crazy lately, so I never actually scheduled this with anyone, and there's nobody here to ask me the questions. I'm just going to read the questions myself and then give you my answers. I got some really good questions, and I'm excited to answer them. Hopefully, people will enjoy this episode and find some value in it.

1. Ernie Shows Treatment Progress

By far, the first and most important question—and the most common question I'm getting these days—is how my son Ernie is doing since the episode I did about his cancer back in November. The good news is that he's doing really quite well. I'm very pleased to report that. Certainly, cancer of this type, being as aggressive as it is—and I won't belabor the whole thing from last time; go check out the 2-hour monologue on that if you want the full story—gets very aggressive treatment. A cancer this aggressive, which can double as quickly as every 24 hours, does get very aggressive treatment.

He's been through the wringer with the chemotherapy. He's through basically half the chemotherapy now. There are 6 rounds in total, and he's been through 3. The final 2 rounds, rounds 5 and 6, are supposed to be a little milder than the first 4. So, depending on how you count, we could say he's maybe a little more than halfway through the treatment, but somewhere around there.

It's definitely been rough on him; there's no doubt about it. When he went into the hospital, he was 51 pounds. He's still 41 pounds today, and that's the weight he came home at after the first round of treatment. He's been able to gain a little weight, lost it back, gained a little, gotten dehydrated, and lost a little. You can see just by looking at him that he's super thin. He's quite pale, and he's definitely not nearly as strong as he was before we went in.

But on the markers that really count the most—namely, does it look like the cancer is being effectively treated?—he looks really good. After the first round of chemotherapy, the PET scan he had showed no obvious focal points of cancer. When our oncologist met with the tumor board, they all agreed that it made sense to classify him as being in remission before he even started the second round of treatment. So that's great.

If you listened to that earlier long episode, you might recall that one of the things AI helped me do was identify some additional testing that is not yet standard of care but can be done to try to get a better, more sensitive take on whether there's any cancer left in his body, how much there is, and how it's trending. That's called minimal residual disease testing.

I don't know how it works in all different kinds of cancers, but in the cancer that he has, which is a cancer of the B cells, the B cells do this interesting thing where they rearrange certain parts of their genetic material in a purposeful way—random, I think, or certainly semirandom, but purposeful—so that they create variation and have a better chance of creating proteins that bind to new disease factors in the body. This process of differentiating B cells is literally unique cell by cell.

When one of those cells goes bad, becomes cancerous, and grows out of control, they can use the rearrangement that individual cell made—which then gave rise to the whole cancerous process in the body—to fingerprint that cell type. There are 2 sequences, 1 for each of the chromosome pairs where this rearrangement happens, that they've identified as the dominant clone of the cancer in the body. Now that they've identified that, we can do a blood test every so often and check to see how much of that DNA is floating free in the blood and how many live cells actually have that DNA sequence.

We've so far only gotten 1 of those tests back, and we certainly want to look at more and trend it over time. But the first one that came back—it was drawn over a month ago—came back with fewer than 1 cell in 1 million having that DNA sequence detected. That's really good. They also called that below the LOD, or limit of detection, for the test. It was basically in that area where they would expect that maybe some samples would have 0, and some might have 1 or 2 cells, but it's not zero, but it's a very low rate.

For reference, we estimated that when he was diagnosed, potentially as many as 1 in 10 cells in his body—and essentially all of the B cells—were of the cancerous type. To go from 1 in 10 total cells and a large majority of the B cells down to 1 in 1 million cells detected or less, that's obviously great. I think Gemini said it was a 99.9999% reduction. Other AIs were a little less colorful in their language and said it was probably safer to say that it was an orders-of-magnitude reduction.

We'll do more testing of that type, and we'll certainly be watching it. But as of now, we're feeling cautiously optimistic that he is on the path to a cure and a full recovery. In some ways, the recovery is already underway. It was 60 days from about a week before we went into the hospital to just around Christmastime that he was not able to get around by himself. He could stand, but to walk, we would always hold his hand and make sure he had support every step that he took.

Finally, around Christmastime, we had a chance to come home for a week from the hospital. During that window, he regained some strength and started getting around by himself. Fortunately, that has been sustained for the last 2 weeks or so since he started doing that. Hopefully, knock on wood, there is some risk. They don't understand exactly why this cancer can come back in some patients, even when it looks like it's gone.

We're not out of the woods entirely, but his response to treatment has been basically as good as we could have hoped for. Even with the MRD testing suggested by the AIs, it looks about as good as we could hope for. I certainly hope that the next one shows no detection at all, but that one is still pending, so we'll have to wait and see.

I really do appreciate everyone who has reached out during this time. There have been a lot of well wishes. I've tried to respond to everyone, and I think I've mostly responded to everyone. If I've missed you, I apologize for that, but I really have appreciated all the encouraging words.

I also wanted to say a quick shout-out of thanks to my fellow podcasters who have allowed me to cross-post some of their content to our feed over the last couple of months. I certainly couldn't keep up the pace of doing 8 episodes a month during this time, and I was very glad and fortunate that I was able to do some cross-posting and bring you guys some other stuff that I think is well worth your attention. It also took a little bit of a load off me.

We had one from Agents of Scale. That was actually a sponsored episode from Wade Foster, the CEO of Zapier, who's got a new podcast out. We had one from ChinaTalk, which was with a researcher and business development lead from ZAI out of China. I thought that one was really quite interesting. We had one from Doom Debates, which was a debate between Max Tegmark and Dean Ball. I thought that one was really good. I've actually had the goal for a long time of cross-posting at least 1 episode a month, just because I feel like 8 episodes a month is a lot. If anybody is listening to all of these episodes, they should probably be diversifying, so maybe I can help you diversify if you're not diversifying on your own.

It also keeps me listening. I definitely want to make sure that I'm staying in touch with what other people are coming up with in this field. I think all of those were really good. Finally, from the a16z podcast, we had the one that Eric did with Emmett Shear and Seb Krier from Softmax and Google DeepMind, respectively. So there will probably be a few more cross-posts in the coming months.

We've got about 2 and a half months left of treatment, after which, assuming all goes well and according to plan, we should really be pretty much done and start to get back to life as normal. He will have to get all his vaccines again, which is another interesting thing, because his immune system has been so thoroughly wiped by all these chemotherapies and immunotherapies. The memory that his immune system had gained from all the vaccines he had gotten in the past is all wiped, and he's going to pretty much have to get them all again.

So that's not ideal. We're not going to be immediately back to full normal, but in 2 and a half to 3 more months, we should be, knock on wood, getting back to pretty much as normal. In the meantime, there probably will be a few more cross-posts. Thanks to everyone who's reached out to ask and share their best wishes, and also to the fellow podcasters who allowed me to cross-post some content and fill some gaps in the schedule. I really appreciate that.

2. Claude Opus Faces The AGI Test

Okay, onto more AI-centric topics, as you tuned in for in the first place. The next question is, “Is Claude Opus 4.5 AGI, and what's up with the holiday Claude Code hype?” To be honest, I'm not exactly sure about this.

It kind of surprised me. Obviously, Opus 4.5 is awesome. There’s no denying that, and I have been using it as I try to use all of the latest and greatest coding models. I’ve had a great time with it; there’s no doubt about that.

I vibe-coded 3 apps for family members as Christmas presents this year, in the hospital for the most part. Actually, probably the most frustrating part of that experience was the hospital Wi-Fi, which kept causing me to reload my Replit app. The actual coding experience was very good—clearly better than it has been in the past, no doubt. The progress is unmistakable.

And yet, I wouldn’t say that it has been such a step change for me relative to what I’ve experienced in the past that I would say, “Oh, it’s categorically different,” or that it makes me want to shout from the rooftops that some major threshold has been crossed.

I’m not sure if that’s taking the charitable view, and people have said this about the cancer thing as well. A handful of people have said, “Maybe you’re getting that kind of value out of the models for cancer purposes, but that’s probably because you really know what you’re doing. Other people might not get as much value because they might not know what they’re doing, and they could go wrong.”

Honestly, I would say about the cancer, first of all, that you really don’t need much skill in using AI to get great value from the latest generation of models, even for something as important, critically important, and cognitively demanding as a cancer case. I feel very confident that a layperson with basically no knowledge of AI would get very similar value to what I’ve gotten if they did pretty much 3 things. Maybe I’ll say 3 things.

One thing is to use the best version of the models. Do not go to ChatGPT, drop in the question, and let the model picker choose. Make sure you are using at least Thinking, and I would really recommend Pro if you’re dealing with something that sensitive. Yes, it’s $200 a month. In that context, I think it’s absolutely worth it. I think it’s worth it generally for almost everyone, regardless, but certainly if you’re dealing with a life-threatening situation and you’re asking AIs to weigh in on it, paying $200 a month is a no-brainer.

Claude Opus 4.5 is, of course, the other one, and Gemini 3. I would say all 3 of those are very good. Make sure you are using those top-tier models. Before long, of course, there will be new top-tier models, and you should probably be upgrading as soon as you possibly can. That’s thing 1: just make sure you’re using the best available models. If you’re doing that, they are up to the challenge.

Second, make sure you’re providing as much context as you possibly can. I recently hit the length limit of my chat with Claude, and I started a new one. I did that in part by taking all of the material I had and summarizing it into maybe a 10-page report on everything that had happened so far: everything we’d learned, the treatment protocol, the genetic profile of the cancer, how he’d reacted to different things, which drug he’d had a bad reaction to and shouldn’t take again. It’s pretty much all in there—pretty much everything, quote unquote, that a new attending physician would need to get a good survey of the case.

Hopefully, it was also something that I could paste into a fresh context with a new language model and give it everything it needed. I have noticed that, in doing that, certain information was obviously lost, and when I’ve started a fresh chat with that kind of summarized history, the performance is a little bit worse.

For example, one way in which it’s been noticeably worse is that when I was going to Claude every single day and giving it the latest lab results—saying, “Here are the latest lab results. Here’s what we’ve seen. Here’s what’s going on. Give me your take on it”—it would do a very good job of looking back at the previous day or the last couple of days of lab results and figuring out the trend.

When the whole history was compressed, it didn’t have that level of detail anymore. It couldn’t look at literally yesterday’s lab results. It started to compare today’s lab results, from January 6, to the last lab result it had in the summary, which was a couple of weeks ago, for a particular liver enzyme. The details of that don’t matter, but it wasn’t a particularly important thing. We had a little question about it today, and it wasn’t something where every single data point was in that history.

You can see that it’s starting to perform a little worse because it’s looking a little too far back into the history and not realizing that there were a bunch of blood tests taken in the meantime. Anyway, that’s all very much in the weeds. The key point is: give it as much context as you possibly can.

What I probably need to do next is take that summary and flesh it out even more. If I do that, I should be in good shape. Make sure the models have as much information as you can possibly give them.

I have not really seen much trouble in terms of context overload or the model getting confused. That’s not to say that it hasn’t happened at all; I certainly could have missed something along those lines. But when I went back with the summarized case report after hitting the Claude length limit on the chat, it was clear that the performance was worse for a lack of context. It wasn’t even really the model’s fault; it was just clearly worse for lack of context. More context is better. I haven’t really seen that rule violated at all. Give it as much as you possibly can.

And then the third thing—so, the first thing is to use the latest and greatest models, and the second thing is to give it as much context as you possibly can—is to get multiple opinions, including multiple AI opinions. I’m using Gemini 3, Claude Opus 4.5, and GPT-5.2 Pro now for pretty much all important queries. It is instructive and definitely useful to compare and contrast.

I would say they’re all very good. If you really could only afford one, I think you can trust it pretty well. Even though I think Gemini 3 is extremely impressive, I would probably put it third in my draft order now because I’ve learned that it does seem to have a bias toward strong opinions. It seems to me to be remarkably strong in its opinions.

If you heard my live show where we talked to Logan Kilpatrick from Google, he noted that I have been using this in Google AI Studio. I’m using the most bare-bones, unaltered, raw model that you could basically get access to. If you use the Gemini app, presumably there’s a system prompt in there, and it might behave a little bit differently. Obviously, if you use other apps powered by Gemini, there would be all kinds of different modifications that would cause it to behave differently.

But just using the raw model in AI Studio, I’ve found Gemini 3 to be very opinionated, and sometimes I really like that. I do really like it as one of the 3 takes that I’m getting. But if it were the only take, I would worry a little bit that it would sometimes push me too hard in a certain direction. If I had the full 3, they would kind of balance me out.

I think I would put Claude Opus 4.5 at the top for most people because it’s much faster than GPT-5.2 Pro, and I don’t notice it being much worse. Its answers are shorter. They’re much more focused on answering your question than on doing a full report-style analysis.

I don’t really use system prompts or custom instructions for any of these, by the way. I pretty much use them in their vanilla form as much as I can. GPT-5.2 Pro gives you long, sectioned, report-style analyses that I do find to be very useful.

But if I had to pick the Goldilocks one, I think it would be Claude Opus 4.5. Gemini 3 is maybe a little too brief and a little too opinionated. GPT-5.2 Pro is maybe a little too verbose, with a little too much information overload. Claude is just right.

I do recommend using all 3, and I think doing it all in triplicate is absolutely worthwhile. To pop back up a layer in my question stack, people sometimes say to me, “You get this value from these AIs, but you know what you’re doing, and other people don’t.” My advice is really very simple: do those 3 things, and you’re going to get value. You don’t need to be an AI expert by any means.

That said, maybe you could say I was getting more value from previous coding models relative to what other people were getting because I had more practice and was more skilled at it. I certainly think there’s some truth to that.

When you look at the MIRI study that showed some software developers thought they were being sped up by AI but were actually being slowed down, I think that’s important. I love MIRI, and I’ve said many times, “Do science, report the results.” You do not need to make your scientific publication fit a particular narrative. In fact, you probably shouldn’t try to do that. You should probably just try to run experiments and share the results, as long as you believe the experiment was well run and the results are legitimate.

I do believe those results are legitimate, but I think there were some important caveats. It was older generations of models, the people didn’t have much experience, and it involved very large, well-established codebases with very high coding standards.

I don’t tend to code in that kind of environment. I tend to vibe-code and hack together apps, and I certainly think I’ve gotten to be pretty decent at it. So maybe I was maxing out previous-generation models a little bit more than other people. I don’t really know.

It has surprised me. I guess, to take the flip side, you could say, “Well, hey, maybe Nathan, you aren’t such a great software developer. Maybe these pro software developers have better taste, and they’re actually—now that Claude Opus 4.5 has gotten so good, or has crossed some threshold where it’s really becoming a lot more useful to them—maybe they’re noticing that difference and I’m not because I’m just fundamentally not as good at the task, don’t have as much taste in this domain, and I’m just not able to see what Claude Opus 4.5 is bringing to the table over and above Claude Opus 4.1 or other frontier coding models.”

I don’t know. That’s possible. I certainly am not a great software engineer, so that certainly can’t be ruled out. But it could also just be some social things, like people catching up over the holidays. Maybe the timing was right.

Sometimes these things go with a cascade. Dean Ball tweeted, “4.5 is AGI,” and people seemed to latch onto it. To some extent, I think some of this stuff is also just random social dynamics at times.

3. Three Holiday Apps Take Shape

The 3 apps that I coded, by the way, for the holidays, for what it’s worth: My mom is a very meticulous travel planner. My parents were actually in Italy for a trip and came home early to move into our house and help us take care of our kids while we’ve been at the hospital so much, so thank God for them for doing that.

My mom plans these trips that she and my dad take to the maximal limit of planning. I coded her an app to try to accelerate her planning process by building in a lot of her tastes. She’s gluten-free, for example, so that’s one big place where her time goes in planning these trips: figuring out what places she can actually eat at and what places are gluten-free.

This is an app purely for her. There’s no account; it’s not something that she logs into and logs out of. It’s a Replit app that she goes to when she wants to. Nobody else is ever going to use it. Her profile is baked in.

I could imagine generalizing it and allowing people to customize their own profiles, but I’m not really trying to do that. I’m sure there are plenty of travel apps out there that people are building and commercializing. This one was really just for my mom, trying to capture some of the things that she does and make it work for her and speed up her process.

So far, I think that’s gone pretty well for her. It seems like she’s getting at least some value from it. Then I made one for my wife, who organizes EA Global events. It simulates events and allows her to set up a roster of attendees with different profiles and various attributes, then literally simulates people walking around a virtual event space and bumping into each other.

Depending on what areas they’re interested in, they may or may not have a conversation, and that conversation may or may not lead to some outcome. They track various KPIs, which they measure mostly through surveys, but I’ve set this up as a simulation. The goal there is for her to at least try to get some handle on whether, if we changed the size of the event, it would be more or less effective or more or less cost-effective.

What if we had more senior people versus more junior people? What’s the right mix? Obviously, these simulations are always highly flawed, but I’d say they probably do have something to add relative to total guesswork. That was a pretty fun and pretty straightforward one. It went pretty smoothly.

The 3rd one was for my dad, who isn’t really a very active day trader but fancies himself a bit of a stock market guy. So for him, I created an app. These are all AI apps.

In the case of the travel planning, Claude goes out there and does the research, digging through Italian restaurant websites and reviews to figure out whether they’re gluten-free. In the case of my wife’s app, she can prompt the app with a general idea, and it will fill in all the detailed configuration, and then she can edit it.

I think that’s a great paradigm or pattern in general for apps. You always have these detailed configurations, these nitty-gritty forms that need to be filled out, but AIs are really good at doing that. If you just give them a general gist of what you want, they can translate that down to the low-level configuration.

That’s where the AI is in her case. It also allows her to edit a configuration. Say she has a certain event profile that she’s set up. She could then say, “I want to change this in the following way,” and it would take that conceptual idea she gave it and apply it to the configuration, changing all the little things that need to be changed.

With my dad’s app, it takes a high-level natural-language stock-trading strategy and turns that into actual trading rules. Then it goes and fetches historical data and simulates what would happen if you applied those trading rules, based on that high-level natural-language strategy, over a time interval that he can define.

There’s a Python package out there called yfinance, which I didn’t really know anything about. It does have a paid version, but there’s a free version, and for now he’s been able to get by with just the free version. It goes back and gets historical data and simulates what would happen.

What we’re finding more often than not is that it’s pretty tough to beat the market. I don’t know if he’ll listen to this, but one of my private motivations for making this thing was to convince him that he’s probably not going to beat the market, and certainly not with these random, heuristic, if-then trading strategies.

Sure enough, so far it has been very difficult for either me, in developing the app, or him, in using it, to find a strategy that actually beats buying and holding the S&P 500. I’d estimate 3 full workdays—somewhere between 3 and 5, probably closer to 3—into those 3 apps.

In each case, I didn’t really know where I was going when I started. I started with a chat with Claude just to say, “Hey, here’s what I’m looking to do. Help me out.” I think it’s good. Is it night-and-day better than Claude Opus 4.1 or Claude Opus 4.0, going back to the question that prompted this whole Christmas-present vibe-coding story? I can’t really say.

I think it’s that much different, but it’s certainly very good. It has good back-and-forth, good questions, and good feature ideas. I’d translate that all into a plan, then go over to the Replit app and install Claude Code on Replit. That’s one of the things I love about Replit: You can do pretty much whatever you can do in a normal, fully controlled development environment.

That includes installing Claude Code. Of course, Replit has its AI agent, too, but since this was a moment of Claude Code hype, I would just install Claude Code there, give it the plan, let Claude run off and build the app, and then test and iterate.

I still find—and this might be a way in which I’m falling short as a Claude Code user—a lot of value in a short script that prints my entire app to a single text file. I then take that entire text file over to another LLM—it could be Claude, or Gemini if I need more space—and ask it to analyze the codebase in full.

I think Claude Code does a very good job of agentic search. If I’ve found any shortcomings, it was in one particular moment in my mom’s travel-planning app. I know roughly how this originally happened: I made a request, and I think it misinterpreted the request. We ended up with 2 databases, and this became very confusing.

This is pretty illustrative, because this is the kind of mistake that earlier vibe-coding experiences would create all the time. You’d be like, “What is going on? I ended up with 2 databases.” This is the kind of mistake no human would make, right? A human software developer would not suddenly spin up a totally separate database. That would be very weird.

But the AI did that. It thought it was trying to follow my instructions, I think, but it didn’t understand what I was trying to get across. So we ended up with these 2 databases, and then certain things weren’t working as expected, and it was very confusing.

This was one place where, once I got down to, “Okay, there are 2 databases,” I asked Claude Code, “Which one is actually being used, and which one is superfluous?” Then I took the full code export over to a clean Claude.ai, pasted the whole thing in, and asked it which database was actually being used.

The model that had the full exported codebase got it right. Claude Code did not get it right. I think that’s because, in its agentic search, it looks in the places where it expects to find things, and it has a relatively high prior that this is where it’s going to be. Sure enough, it appears to be there, so it goes with that. But what was actually happening was something that was counterintuitive.

Having the full context in view at one time really did seem to help Claude figure that out. This is something that I think previous models might have struggled with, even with the full context in place. But with that trick, it sometimes can help you clean up a mess or a point of confusion that the agentic search functionality of Claude Code, in my experience, seems to struggle with. I’m sure that people will be able to offer strategies to do the exact same thing right within Claude Code.

There’s planning mode, which I probably underuse, frankly. But I think there is something to be learned there between the agentic search finding what it’s looking for, what it expects to find and thinks is right, and then coming to the wrong conclusion because, actually, in this case, it was the rare, weird other thing that was happening. Only in seeing it all together was that correctly diagnosed by Claude.

But anyway, it’s better. There’s no doubt. I don’t really feel the step change. Is it AGI? I mean, if you look at GDPval, arguably, in some way, in software it is AGI. If you look at the latest from OpenAI and the latest from Anthropic, it’s a pretty significant majority of software engineering tasks where the model is beating the human.

These are, just to remind you of GDPval, professional-caliber tasks. They basically have 3 sets of experts: The first set of experts defines the task, the second set of experts does the task, and then the third set of experts judges whether the human or an AI that did the task did a better job. The latest models are preferred over humans in a significant majority of cases in the software engineering category.

Of course, it’s spiky and jagged. If you go to the video-editing category, humans still have a huge advantage. I’ve certainly experienced that. We’ve tried many AI products and workflows to create good clips out of The Cognitive Revolution, and they work okay. They’re clearly not as good as what Dwarkesh and his team put out.

We’ve tried, and we’ve made some good progress. I actually think that at some points in time, what we’ve had internally has been better than any other outside product I’ve tried. I’m not sure I would say that’s necessarily true today, but at times I preferred what we were doing to anything that I had tested on the market. Then you look at the clips that Dwarkesh and his team are putting out, and they’re just clearly better. You see that in GDPval, too. It’s a very small percentage of cases in which the models are preferred to humans in these video-editing tasks.

But in software, I think you could certainly make the case that Opus 4.5 is software AGI and/or coding AGI. Yet I’m still a little bit at a loss to fully answer the question of what caused this moment of hype around the holidays. Hopefully, there are some other nuggets in there for people to pick up on and go run and use.

If you haven’t used Claude Code, I absolutely would say to do it. It’s really easy to install. It’s a 1-line command to install, and you don’t really need to know how to code these days. You can watch it work. My mom even did a couple of projects. She was kind of like, “I don’t think I’m going to do this. I’d be worried I’m going to mess it up.” And I was like, “I think you really can.”

It’s your little agent on the computer. You just tell it what to do, and you don’t really have to understand what it’s doing. You can ask it to explain. It does sort of explain, at least to some degree, by default, but you don’t really have to be a software engineer to use it. You can still get pretty far.

It was really just 1 or a couple of things. This database thing was one where I did have to—not debug, but at least ask some probing questions of the models to get a handle on what was going on. It probably took 5 or 6 prompts to resolve that issue. I can imagine that in the future, A, it might not happen in the first place, or B, maybe it would be resolved in just a couple of prompts with the next generation of models.

But this is already getting pretty amazing when it comes to being able to clean up these messes that it sometimes inadvertently makes and get over these humps. If you’d asked me a year ago, I would have said those are the points when a lot of these projects die: when somebody gets to the point where something has gone wrong, they’re confused, they don’t know what’s going on, and the AI is totally confused. They circle around the problem for a little while, can’t solve it, and move on.

I certainly experienced that myself at times. In most of those cases, I probably could have spent the time to go in and figure it out for real. But the whole point of vibe coding is that you’re not trying to put that much energy into it. So sometimes I would just abandon something like that and maybe start over. Now you actually can get out of those messes that AI-assisted coding sometimes makes.

The addressable market for these things continues to expand dramatically. I think the implications for the future of the software industry are profound. It’s software AGI, I think, but maybe not full AGI, and for that we might have to wait just a little bit longer.

4. The AI Bubble Has Layers

Okay. That was enough on that. Next question: Are we in a bubble? There are a couple of different versions of this. I think my answer here can be relatively short.

When it comes to whether AI is for real or not, I’m not going to surprise anybody by saying I think it’s absolutely for real. The technology is already amazing. The fact that it can go toe-to-toe with an oncologist, while also having all the other advantages—always-on access, 24/7 availability, the ability to handle full context, the command it has of the case based on all the history that it has, and the fact that it will answer every last question that I have—all these are dramatic advantages.

At the point where it’s competitively accurate with a human oncologist, I think you’re clearly dealing with transformative technology. I think the idea that we will somehow get out of the other side of this AI thing and feel like we were all high on our own AI supply—I think we can very safely put that to bed at this point.

Now, does that mean that all the loans are going to be repaid? That’s much less obvious. I think, especially when you see just how aggressive a company like OpenAI is being in terms of all the financial dealmaking that it’s doing and all the buildout that it has planned, it’s conceivable that its revenue projections could fall short of its obligations. Could it default on something?

There’s also some financial wizardry going on. It’s funny: A lot of this financial stuff has a logic to it, even though, in retrospect—and I worked in the mortgage industry before and during the mortgage bubble—there was always a logic to what people were doing. They were telling themselves a very positive story about how they were making homeownership accessible to more people than ever before, how this was going to be great, and how it was the Great Moderation. There’s always a story with these financial-engineering phenomena.

One of the engineering things that’s happening is these whole CoreWeave-type companies that are there to set up and rapidly construct, and to some degree operate, the data centers. They do have expertise in setting up the data centers. But it seems like a significant part of the reason they exist is because the financial profile of those businesses isn’t as attractive as, say, Microsoft’s traditional business, which is just so high-margin, with relatively low CapEx and relatively high margins.

I think there’s a sense that these hyperscaler, high-margin, gold-standard software businesses that Wall Street is accustomed to could see their stocks dragged down if they start to engage in a lot of lower-margin business, like running GPUs. I don’t know how much of a factor that is versus the actual expertise these companies bring in terms of setting up and operating the data centers, but I think there’s definitely some nontrivial motivation there.

Maybe that’s fine. Different companies can have different financial profiles, and to some degree that might be good. Certainly, a lot of shareholder value, so to speak, has been created that way. But it does create these companies that, if the GPUs aren’t needed quite as much as people expect them to be, have a lot less margin for error than a Microsoft does.

If Microsoft were owning and operating all these things itself, it has a deep balance sheet that can take a few knocks. By putting a lot of this stuff more on the CoreWeave side of the fence, it does create some fragility. So it’s certainly very conceivable to me that we might have some period of overbuilding.

Noah analogizes to the railroads. The railroads, in the end, were a pretty good investment. They all got used; there weren’t a lot of railroads sitting around idle. That didn’t necessarily mean that all the railroad companies were profitable. There certainly were busts when loans couldn’t be paid back, and then you had cascading effects throughout the economy. So I think that kind of bubble is not too unlikely.

So far, demand for AI has exceeded my expectations. We talked about this with Peter Wildeford on the live show a little bit, where I said, “I overestimated how much capability progress would happen in 2025, but I underestimated how much revenue growth there would be.” So, possibly that’ll happen again, and demand and revenue will just continue to go up and up and up, and it’ll all be fine.

But it wouldn’t shock me if there were some moments where it was like, “Hey, we kind of overbuilt this thing,” and some people weren’t necessarily going to be paid back. Some people might be left holding various bags. But even so, that doesn’t mean that it’s a bad investment. It just means that it might not be timed quite right for people to all make the money that they’re projecting they’re going to make.

I think the final sense in which we might be in a bubble is at the venture-capital level. There, I have to say, I think there’s at least something like a bubble happening, and there are many, many examples of this. But the thing that just came out today that made my head spin was the organization that was originally called LMSYS.org, then became LMArena, and now is just @Arena on Twitter.

They have just raised, I think, $100 million—maybe $150 million—at a $1.7 billion valuation. Here, I’m like, “Whoa, that seems crazy.” I don’t know a lot about their business. I haven’t seen their deck, so I could be wrong. But this is a product that I’ve watched for a long time and continue to check, and I do have the receipts on that.

My first tweet about what was then LMSYS.org goes back to mid-2023, more than 2.5 years ago now. At the time, I was just randomly tweeting that it had started to show up in my favorites in mobile Safari. So I was using it quite a lot then to compare and contrast model performance.

Obviously, it’s gotten bigger since then. Obviously, the whole field has gotten bigger, and they’ve started to do various services where they allow companies to test their models with code names. There’s definitely value in that. Does that seem to me like a unicorn business? It definitely seems to me like that would be a big stretch.

The tweet that they put out today—and I don’t want to be too harsh on this, because, again, I don’t know a lot—I think of this as more representative of a phenomenon that I see a lot, as opposed to something very specific to this particular company and its raise. Again, I like the company. I’ve liked its product.

The tweet said that their operation has scaled to $30 million in “annualized consumption run rate.” I’m like, “What is annualized consumption run rate? Does that mean how much the AI that people are using for free, when they go to LMArena and do these side-by-side comparisons, would cost $30 million if they were paying for it?” That’s my naïve interpretation. I didn’t see a clarification on that.

But if that’s what it means, it’s very much giving me community-adjusted EBITDA vibes, because saying that people use what would cost $30 million worth of free AI on our platform is not the same thing as saying you’re making $30 million in revenue. I don’t see that they disclose what revenue they’re making, and a $1.7 billion valuation for an app that basically does a side-by-side comparison of AIs—I don’t know.

It seems to me that people are using it in large part because it’s free. I’m sure some people are also just curious about doing side-by-side testing. I’ve certainly done that myself. But the people who go there because they specifically want a way to do side-by-side testing seem to me to represent a relatively small market, and the people who go there because it’s free seem to me to be a big part of why people are going there.

How does that translate into a $1.7 billion valuation? Color me confused—or skeptical—on that. At a minimum, I have to believe that a lot of these things are just not going to pay off for venture investors.

If you want to see something else, too—I mean, where’s the moat? There’s a brand, I guess. People come to it. But again, would they come to it if you had to pay for it? I’m not so sure.

Another thing that a friend, Andrew Critch—the coiner of the Big Tech Singularity meme—has created is something called Multiplicity, which is paid. I think it’s become popular among a small group of people who value this kind of thing. I’ve certainly seen some very positive reviews of it.

But it’s something you pay for. It allows you to use multiple models and systematically compare and contrast their outputs. I think it’s actually more feature-rich than LMSYS for the end user. This is something that he and his teammates have built over a period of months, certainly not years.

I just have a hard time seeing where the $1.7 billion in value is with LMArena, and I say that again as somebody who has used it and appreciated it for far longer than most. Time will tell. I could be wrong. I could be missing something.

Please let me know if you’re on the LMArena squad and want to talk. I would be perfectly open to doing a full episode with the LMArena folks. But it just doesn’t feel right. I hope they took some value off the table. I guess, for their sake, I hope they did some secondary, but for the LPs and the fund’s sake, it’s too rich for my blood. That I can say confidently.

5. Chinese Models Miss The Long Tail

Okay, next topic: live player analysis. This one, I think Eric wouldn’t mind me saying, he asked for. I’ll do my best Zvi impression, and we’ll see how I can compare and contrast a little bit with Zvi. Hopefully, before too long, we’ll have him back.

I want to start with the Chinese models, because I think, first of all, very few people in the general consumer market are using Chinese models in the US today, pretty much at all. Most startups are also using American API models. Some are using Llama models to fine-tune, and some are indeed using Chinese models to fine-tune.

I don’t see a ton of that happening, and I don’t think there are too many people who actually go, as I recently had occasion to do, and just try all the Chinese models. I was working on what basically amounts to a computer-vision task. I think I’ve alluded to this a little bit in the past.

I’ve been working a bit with a company that automates the review of the paperwork associated with the buying and selling of cars. You buy a car, you sell a car, and there’s some paperwork that has to get filed with the state to document that transaction, whatever. It’s all very boring stuff. Perfect for AI, honestly.

I think reviewing these documents is a great example of the kind of work most people don’t enjoy doing, and they’re doing it primarily because they need a job, because they need to get paid. This is something I’m perfectly happy to see AI take off people’s plates.

They’ve been able to get to the point where they’re doing it more accurately than people, and they’ve started to get some statewide contracts from state governments that are like, “Hey, if you can do this faster and more accurately than our people, that’s a win for our taxpayers and our people who need these documents accurately reviewed.” So, great.

These documents are typically scanned, which means they’re all kinds of messed up. There are artifacts from the scanning process. Sometimes there are perspective issues or weird slanting. Sometimes the margins are wrong, and things can be cut off the side of the page.

There are all these complications that make this not the most straightforward task for the models to read these documents. I was just helping out a little bit, and there was one particular aspect of reading these documents that the models were struggling with.

I went and tested basically every model I could get my hands on, every frontier model. I tested Gemini 3. It’s very, very good, but it was making this one idiosyncratic mistake.

Past guest Mark Humphries, the Canadian history professor, put out a blog post that went quite viral and talked about this. It was actually before Gemini 3 came out, and it looked at his work with old handwriting—historical handwritten documents.

They’re hard to read because they’re written in old, illegible cursive script with ink on paper. They can also be hard to interpret because a lot of them are, like, facts. He points out that if you got a ledger or something from an old shop, the person might be recording what they sold, how much, and to whom.

That person could have come in and bought whatever. There’s not a great prior on what that should be. So the values that it interprets really rely on perception for the most part, though there are some places where it can make some logical leaps.

If something is priced at a certain amount per unit, it might be able to make intelligent guesses about what that unit was, even if it can’t quite make it out. Was it an ounce or a pound? It might have some historical knowledge of what that price roughly would have been.

It can use that world knowledge to do some of this reasoning and fill in some of these gaps in its perception. So he published this article.

It was definitely worth checking out. It documented that Gemini 3 was doing this in a way that no other model had done before. In the context of this project of reading these documents being filed with the state for these car sale transactions, it worked against Gemini 3 in the sense that what we were trying to do was faithfully read the document. We were not trying to make guesses about what the document should have said.

There was one checkbox, for example, that was, “Are you a U.S. citizen?” If the box was checked, we wanted to say it was checked. If it was not checked, we wanted to say it was not checked. But the model sometimes made inferences, reporting that the person was a citizen even though the box was not checked. It was presumably doing that based on other context clues. It was like, “The person lives in the United States, and their name sounds American,” whatever that sounds like.

It was making the logical guess, which probably was right, actually. I would guess that the person who filled out this document was, in fact, a U.S. citizen. I would say there was more than a 90% chance that they were. But they did not check the box on the form. So Gemini, using its priors and trying to get the answer right, was less anchored to the document than we needed it to be.

Claude, we found, could do this. It took some prompting, and I had to tell it, “Make no guesses. Read this thing exactly as it is. Make no logical leaps,” and so on. It turned out that Claude Opus 4.5—Claude 3, RIP—was the best at actually being faithful to the document.

Along the way, I thought, “Okay, let me go check all these Chinese models.” I had heard good things, so I went to the latest Qwen Vision model, GLM 4.6, the latest Kimi, and the latest DeepSeek. At least those 4, maybe 1 other one that I’m forgetting, and they were all way behind. Nowhere close—nowhere close to Gemini 3, nowhere close to Claude Opus 4.5, and nowhere close to what ChatGPT can do.

This had me thinking, “This is odd.” We are seeing statements all the time that the Chinese models are so close and not far behind at all. I think they are quite good in many ways and for many things. But on this particular task—and I suspect that this is true on a lot of different tasks, although I’m going on vibes here a bit myself as well—I suspect that gap is actually pretty wide in a lot of cases.

I do not feel right now that any of the Chinese models are really competitive with the best proprietary models coming out of the United States. They might be competitive on benchmark scores, and they might be competitive in some domains. But in the general-purpose case, where you throw something really idiosyncratic and random at it that it has not seen and that is not on somebody’s agenda of showing up as competitive on a rubric of 20 benchmarks, I think that gap is actually significant, kind of wide.

They were not close—not close at all. The Gemini mistakes were like, it was reading this gnarly government form almost perfectly, but it was missing a few of these checkbox things or making wrong inferences here and there, and I could not quite get it to stop doing that. Claude Opus 4.5 was just plain doing it right. GPT was probably 3rd—not as good as the other 2, but still very good, certainly giving you all the right information for the most part and missing relatively subtle things.

What I was getting back from the Chinese models was like 20% of the form coming back, or it was just going off in very weird, hallucinatory directions in all kinds of different ways. Really not close.

Does that mean that the Chinese companies are not live players? I do think they are affecting the landscape. I am definitely reading a lot more research from Chinese companies these days because they continue to publish their work, and a lot of times it is quite interesting. I feel like they are building the best models about which we know everything, or close to everything, that went into them—or certainly all the details of the architecture and many of the details of the training process.

They are influencing the world in that way, for sure, by disseminating this knowledge very broadly. But I do not see that the models are really competitive today. I do think this is a way in which the chip controls have made an impact. I am not necessarily saying this is a good thing, and I am not necessarily saying it is a bad thing either.

The long history of the chip controls, made short, is that originally it was, “A small yard, high fence. We are going to prevent military applications.” Well, we cannot really do that. They can make enough chips domestically to put whatever chips they need in their drones. But, okay, at least we can prevent them from training frontier models. Well, we cannot do that, or at least they are still doing pretty good models. But we can prevent them from scaling inference or having as many AI agents as we have, and that is where we are today.

I do not really like that idea very much, as I think anybody who has listened to this podcast for any length of time knows. But I do think you see the echo of it in the models themselves here, because it felt to me like these are companies that are training models without the feedback process that the leading American companies have. The American companies are scaling not just the training and the parameters, but the actual inference and the actual customer relationships.

These Chinese companies seem to be able to roughly compete in terms of creating similar-scale models, but they are not able to run inference at anywhere near the same scale. Their revenue is vanishingly smaller than the American companies’ revenue so far. And the feedback—I think that is the thing I really want to zero in on here—is dramatically less.

Their teams, with smaller revenue, are also dramatically smaller. So I think that what we see in these very niche, very idiosyncratic tasks, where we see the small gap in benchmark results open up into wide gaps in terms of how well you can read this government document, has to do with how many customer relationships you have, how comprehensively those customers represent the vast range of things that people might want to do with AIs, how much they are giving you feedback on what is working and not working, and whether you have the human bandwidth at your organization to build the datasets that you need to patch those holes.

I think that is where the Chinese companies are falling behind. I was never 1 who thought that the chip controls would not have an impact. I question whether it is a good idea for us to try to deny the Chinese people, or civilization, the ability to scale AI inference throughout their economy in the same way that we are. But I always expected that it would have some effects.

I do think we are starting to see that, maybe after a period. I associate this line of thinking with Miles Brundage as well, the former head of policy research at OpenAI. He said that the chip controls were going to matter more as we went forward because everything was scaling. If American companies were going to do a 10X increase in compute, that was going to have a lot of impacts.

Sure, maybe DeepSeek R1 was a thing, and they were able to train it with not an insane amount of compute and whatever. But were they going to be able to keep up with the momentum, the flywheel, and the strength-begetting-strength phenomenon that we see the American companies achieving? It seems like the answer is starting to look more like no.

If I had to guess about the gap between the Chinese and American models relative to a year ago, I think it is wider. I think that R1 was closer to O1 than, let’s say, GLM 4.6 or GLM 4.7 is to Claude Opus 4.5. That is based on very limited data, but certainly more than most people have, because I did go and try every single one of those—DeepSeek, Kimi, Qwen, and Z.ai’s GLM. I tried them all on this task, and they were all way behind.

6. The Case For Renting Chips

Another question that I will insert here is: What do I think about the H200 sales to China? At a high level, I still think we should keep in mind that the real others here are the AIs, not the Chinese. The Chinese are humans just like us. The AIs are aliens.

So I am skeptical of any notion that says, “This is a dangerous thing to do. Better that we do it first than they do.” That seems to be the logic we are using when we impose these chip controls. Another logic is just, “We do not like China. We want to keep them down, and we want to have every advantage that we can.” I do not like that line of thinking. I do not like either of those lines of thinking, and so I generally favor more willingness to sell chips to China than we have had.

At the same time, obviously, this exists in the context of a very complicated and many-faceted relationship. So it is very weird to me that all of a sudden we go from this. I think my history on this is right. It was not that long ago that Trump said we were not going to sell the H20s, and then he came back and said, “Well, actually, I talked to Jensen, and it’s cool.”

We're going to sell the H200s. It doesn't seem like we really got anything for it. I would definitely be supportive of an attempt to find some sort of grand bargain: “Hey, we'll sell you the chips. You do this,” right? There could be a lot of different values that we might want to bargain for.

Given where we were, where there was a ban, and it certainly seems to have been limiting what their AI industry can do, I don't see why we didn't try to drive a harder bargain. Clearly, there are plenty of things that we could bargain for, and so it does feel like a bit of a wasted opportunity. So I guess I would say that I do think more willingness to trade in chips is something I support, but we shouldn't be naive or allow ourselves to be taken advantage of, or give away one of our best bargaining chips for nothing in return. It seems like that's kind of what we did here, so I don't love that.

The other thing I'll say on this is that I do really like the “rent but don't sell” position that Peter Wildeford staked out on the live show. He basically said, “Look, we don't trust the Chinese government. They don't trust us either. Maybe both sides are right not to trust the other side.” I often note that a lot of the criticisms or the characterizations that we make of them, they can and do make of us. “Hey, you've got an authoritarian madman running your country.” Which country are we talking about? “Your system is not obviously stable.” Again, which country are we talking about?

The idea that we might want to have some leverage, or that we might want to be able to pull something back in the event of a conflict, I could see that being very prudent. So if we were to set out a position like, “We're going to put data centers in Malaysia, the Philippines, Korea, Japan, or whatever. You can rent as much as you want. You can train all the models that you want there. You can run all the inference that you want there. We're just not going to allow them to go into big data centers in your sovereign territory, where we totally lose the ability to exercise any influence over that,” I don't know that would be my first choice of policy, but I think it's a very defensible policy.

At least if it were packaged with a message like, “We believe AI is good, and we want the Chinese people to take advantage of it and get the benefit from it in the same way that we're trying to do that here for ourselves,” I think that would just be a much better message, a much more fertile ground for further cooperation, which we might need. We are headed for a world potentially of transformative AI—which I think we basically already have—to powerful, with a capital P, AI; to AGI; to superintelligence, whatever. However far this goes, it seems likely that we're going to need to work together with the other powerful nations of the world to govern this technology in the right way and to make sure that it actually does pay off for the people of the world. Certainly, China is right at the top of that list.

I think we could take somewhat of a position that they wouldn't like: “We'll rent them to you, but we're not going to put them on your sovereign territory, where we lose all control,” while still maintaining a decent vibe. I would be interested in seeing us try that. As it stands, it seems like we're just going to go ahead and sell the chips. It seems like we didn't get anything for it, and it seems like this is not a great example of negotiation from our dealmaker in chief. But I guess I still hold my nose and like it better than a total ban.

7. Google Still Leads The Field

Okay, so now we get to the real live players. I've got 4, and I'm not sure they're in any particular order. At least, I wouldn't call this a power ranking. The first one I'll talk about is Google DeepMind.

I think these guys are still number 1 in my book. Basically, they pretty much always have been, maybe tied for 1 with OpenAI for a while there, because certainly OpenAI was clearly ahead in terms of productizing transformer-based LLMs. But Google really has it all, from the balance sheet—the fact that they've got a business that's making whatever, $100 billion a year in revenue, and I think literally making $1 billion-plus a week in profit. That gives you a lot of room to buy data centers, have some failed training runs, make some mistakes, and have some research agendas that don't pan out. That is hugely valuable.

They also, of course, have the TPUs. The fact that they've thrown, I think, the 7th generation of the TPU now, that's an insanely valuable bit of IP. They're able to compete to some extent, at least, with NVIDIA. Anthropic is buying lots of TPUs. Other companies are starting to buy lots of TPUs. They're also one of the best data center builders and operators in the world, and they've done that for a long time. So those are 2 critical strengths that basically nobody else on this list has, certainly not in the same way.

They also have the deepest research bench. They've got something that's, if not frontier, at least competitive in every major area. They've got self-driving cars. They've got robotics. They just announced a partnership with—what is it?—Boston Dynamics that's going to power their humanoid robot. They've got a ton of stuff in biology. They've got, of course, the AlphaFold lineage. They've got multiple founders of companies in the material science space and various AI-for-science companies. Many of them are ex-DeepMind. Why? Because DeepMind was investing in those areas before anyone else, and of course those agendas continue within Google to this day.

So it's not like they really have a lot of gaps, and they do have a lot of margin for error. Of course, they have distribution too, right? On many people's lists, that would be the number 1 thing. I was working from the bottom of the stack to the top. They have billions of users. They have product surfaces that they can distribute this stuff on. They are changing Google Search to make it more of an AI experience all the time.

I do now find myself sometimes going back to Google, where I might previously have gone to ChatGPT. This is partly because, when it comes to ChatGPT, I'm really in the habit of using Pro, and Pro's too slow for simple queries. I could, of course, switch back to the auto selector, but what I've found myself doing more often recently is just going straight into the browser and typing a question—the same kind of question that I would put into ChatGPT. More often than not, it goes to AI Mode in Google, and that's working really well for me these days.

They're managing to evolve their product experience. Of course, there are many places where startups are doing a better job of productizing AI experiences than Google themselves are. If you wanted to look at spreadsheets, for example, Gemini in Sheets is not terrible, but it's not the best AI-for-spreadsheets experience out there today. But it also probably doesn't really have to be, because they have all the users, and all of your spreadsheets from the last decade or more, in many cases, are in Google Sheets.

So I think that if you had to pick 1 company to win it all—and I don't mean to suggest that this will be a winner-take-all market; I certainly hope not—but if we were constrained to a scenario where there's going to be 1 winner, who's it going to be? At the end of the day, I still pick Google.

I would also mention that Gemini 3 is a really good model. As I mentioned earlier, I do think it's a little bit too opinionated in some cases, but it shows that Google has figured out how not to be too vanilla. They may have gone a little bit too far in the other direction, but it's not too vanilla. They're figuring out what this technology is, and they're figuring out how to use it.

Gemini 3 was the first model that ever beat Claude at my “Write As Me” task, which I've talked about many times. Claude 4.5 Opus is competitive with Gemini 3, but I still go to Gemini 3 for the “Write As Me” task. That's the first time ever that a non-Claude model took that top spot.

Demis's quote, which I've heard him make a couple of different times, is basically, “If you look back at the last 10 years of AI and you look at all the big breakthroughs and you say, ‘Where did those breakthroughs come from?’ most of them came from Google DeepMind.” And he says, “I would expect that to continue.”

I have to say, that seems right to me. I don't know about most. Obviously, the field has grown tremendously, so one of the reasons they got the majority of breakthroughs in years past was that there weren't that many competitors. There are certainly a lot more competitors now. So I don't mean that literally a majority of breakthroughs will come from Google, but I would say that they will probably continue to have the most breakthroughs of any major frontier organization.

They just don't have a lot of weaknesses, from the financial wherewithal to the data center operations, to the chips, to the models, to the researchers. We had Ali Behrooz on the live show to talk about nested learning. We're going to do a whole episode on that, because I thought that 20 minutes was just not enough to do him and those ideas justice.

You’ve got more ideas like that, I still think, percolating and developing inside of Google DeepMind than probably anywhere else. So you roll that all the way up to the product level, and I think they’re going to be really hard to beat. They have margin for error that nobody else has.

The diffusion language model is another one. That kind of went quiet a little bit, but I just heard a comment from somebody not long ago that they do plan to continue pushing on the diffusion model paradigm for language. This could be a meaningfully different paradigm. The fact that it’s so much faster means that you could code apps in 5 seconds instead of 5 minutes, and that makes a big difference.

It remains to be seen whether exactly that thing will break through or not, but it just seems to me that they have so many of those bets—so many more of those bets than other companies have—that regardless of where things go, I can’t see how they’re not right at the top, if not the top player in the space.

8. OpenAI Loses Its Clear Lead

That brings us to OpenAI. At one point, OpenAI was obviously the leader in model creation and certainly in productization of models. I don’t want to overstate the case here, because I think they continue to be a top-tier player: very competitive, with tremendous traction in the consumer market, although we’ve arguably seen a little bit of erosion there. I just saw an analysis the other day—I think this was Similarweb data—that showed a decline in ChatGPT visits over the last 6 weeks, which roughly corresponds with the time that Gemini 3 was launched and also Claude Opus 4.5.

Notably, Gemini didn’t decline during that time. People were saying, “It was just seasonal, whatever, holiday time,” but according to, I believe, Similarweb, Gemini did not decline. So you do see that, and certainly I think it’s inarguable that Google’s share of the consumer chatbot market is growing. Again, they have the distribution. They have the users, the customer relationships, and they have you. They can integrate with your Gmail and your Google Docs. It can all be seamless. These are huge advantages, so you would expect them to at least start to come back and reclaim some share.

I don’t think OpenAI is off the frontier. I do think GPT-5.2 Pro—well, first it was GPT-5 Pro and GPT-5.1 Pro, and now GPT-5.2 Pro—that series of models is outstanding. There’s no doubt about that. I use it all the time. It gives me the most comprehensive answer when I’m dealing with technical things where I really want a thorough answer that leaves no stone unturned. If there’s anything weird in my son’s lab results, I want the model to flag it, and I think it is probably still the best in that regard.

It gives these very long, very thorough answers. You know where the time went, and it does take a lot longer. GPT-5, especially Pro, is a lot slower. But I’m comparing Pro to the other frontier models because I find that if I don’t use Pro, I’m not as happy with the results. It’s a heavy-hitting thing. It’s expensive, it’s slow, and it is very thorough, very reliable, and very well balanced. I think it’s a very, very good model.

So I wouldn’t say they’ve fallen off, but I would also say that they no longer have an obvious lead. They used to be the best, and it was pretty obvious that they were the best. Now I’d say they’re neck and neck in all the categories that they’re competing in. In language models, it’s kind of neck and neck. In coding, Anthropic probably has the edge, but certainly Codex and our models are very good—arguably neck and neck.

In image generation, Google has the lead, I think. In video generation, it’s close, but I think Google probably has the lead. The Sora social app experiment is interesting and cool, and I thought it was pretty fun, but my sense is still that the Veo 3 models have the lead over Sora. Again, maybe it’s neck and neck, but it’s not as though they’re standing head and shoulders above everybody else. The fact that there was this code red seems to suggest that they understand that they’re not in a dominant position anymore.

Then, of course, you add onto that how much drama always seems to be attached to the company. They just had their head of research leave; that was announced within the last 24 hours. I saw an interesting tweet that was just like, “Here are all the people who have left in the last few years,” and it’s an awful lot of people. To some degree, that’s expected. You could do the same thing for Google, and tons and tons of people have left Google.

It’s not unexpected, or a sign of doom by any means, that people continue to leave a company. But it does feel significant. Certainly, you don’t see that in Anthropic, which is coming up next on the list. Anthropic’s retention of talent is unbelievably strong. It’s not a dire sign for OpenAI that they continue to lose people, but it is not the best sign either.

Where does this leave them? One of the things that’s really interesting about watching their strategy right now is that, financially and in terms of government relations, it seems like they’re going for a too-big-to-fail strategy. It seems like they want to get to a point where their balance sheets are commingled with other balance sheets, and their debt obligations are so substantial that they’re literally trying to get to trillions of dollars of CapEx.

Again, that isn’t not crazy. I mean, it’s crazy, but it’s not crazy. It seems like part of the motivation—increasingly, it looks to me like part of the motivation for all this circular flow of funds and all these balance-sheet-commingling deals that they’re doing—is that they want to build out as aggressively as they possibly can.

I take them absolutely at their word that they think this is good for humanity, that they’re doing it because they want everybody to have access to great AI, and that they think it’s going to be super empowering, transformative, and awesome. As Sam Altman has said, “I don’t care if we burn 5 or 50 or 500 billion dollars. We are building AGI. It’s going to be expensive. It’s going to be totally worth it.” I think they believe that very sincerely.

But I also think they’re looking at it and saying, “Geez, if we do go that hard, we don’t have a lot of room for error.” We’re going to be implicitly or explicitly leveraged in many ways. What if we miss one model cycle with a bad bet or a failed training run? What if something doesn’t work as well as we thought it was going to, or demand just isn’t quite there in the way we thought it was going to be for a quarter or 2?

That could happen because somebody else has a better model for a while, or because humans are weird and there’s just not as much demand as was forecast. What do we do in that case? I think that by tying themselves at the balance-sheet level to so many other organizations, if OpenAI were to default in 2027, let’s say, you would potentially be looking at an instant recession. Their bad debt, being so many billions and billions and hundreds of billions of dollars, could put such a scare factor into the market and cause all kinds of knock-on effects.

It seems like they may see that as a feature rather than a bug, because what tends to happen in those situations—and I know one version of this from my brief stint in the mortgage industry during the financial crisis—is that the government steps in and tries to paper over the whole thing and make it go away.

That might even be the right thing for the government to do if, in 2027, we’ve got 2 trillion dollars of AI build-out in the ground. That’s a large and rough number; I’m not saying it’ll be exactly 2 trillion dollars. Altman thinks we’re headed to 7 trillion dollars of global build-out. He’s probably revised that number upward since then. Let’s say it’s 2 or 3 trillion dollars that’s in the ground in 2 years’ time.

Then they miss, and they can’t pay. What should the government do? Should the government let OpenAI drag down the entire economy, or should the government come in and be a backstop? OpenAI has even said some of this kind of thing publicly, then walked it back a little bit: “We’re not looking for bailouts.”

But I think what their behavior suggests to me is that they are true believers in the good of AI and want to bring it to fruition as fast as possible. They’re willing to take what, under normal circumstances, would be irresponsible financial risks, but they believe that even if they do that, and even if some of those risks come back to bite them, they can probably continue to be a live player because they’ll be too big to fail. They’ll get some sort of bailout or recapitalization or whatever.

If it goes like the financial crisis did, they certainly aren’t going to jail. They’ll all still be rich. They’ll have diversified enough, right, that individually, they’re not going to become poor. So I think they view this as a big social good that they’re building, and they’re willing to socialize some of the financial downside risk as well.

It seems like that is the strategy, and that’s kind of how they’re creating a cushion for themselves. Because they don’t have nearly as much margin for error as Google, that’s kind of the way I see them creating that cushion.

Google has cushion because they’re making $1 billion a week in profit, and that gives you a lot of cushion. OpenAI seems to be trying to establish cushion by being too big to fail. I’d be very open to people telling me that I’m wrong on this, and if somebody from OpenAI wants to come on and make the opposite case, I’d certainly hear them out, but this is my impression. By the way, it’s also reinforced by the fact that Greg Brockman has emerged as Trump’s largest donor: $25 million in the last reporting period.

If you’re playing that strategy, that’s probably just plain smart. It’s probably what he should do. You look around at how decisions get made in the American government today, and cozying up to leadership is not a bad strategy. I’m not sure that we should infer too much about Greg Brockman’s politics. I don’t know anything really about his politics, but it wouldn’t surprise me at all if, on many dimensions, he does not approve of or would do things very differently from what Trump is doing.

But if you’re going to do a multitrillion-dollar build-out and you want to make sure that you have somebody willing to do you a favor if you get yourself into a jam, then $25 million now is potentially just a very rational down payment on a bailout, should you need one—to the tune of hundreds of billions of dollars, maybe even coming up in a couple of years’ time, if just one or a few different things don’t go quite your way and the math doesn’t work in the way that you’ve mapped it out.

9. Anthropic Balances Promise And Risk

Okay, that brings us to Anthropic. Anthropic is probably the easiest company to analyze in some ways. I think Opus 4.5 is today the best single overall model in the world. It’s not a huge delta for me over other things, and it’s not the best on every single use case. As I mentioned, Gemini 3 wins my “Write As Me” challenge right now. But I think Opus 4.5 is the best, and it does really well on all these benchmarks, despite everybody seemingly agreeing that it’s the least benchmark-focused company out there.

Their safety work is definitely the best, although there is certainly plenty of good safety work coming from Google and even OpenAI as well. Their model cards are the best. Their disclosure is the best. Their Soul Document, which was recently regurgitated—or, let’s say, had been memorized by the model and given to people—and which Anthropic confirmed was essentially right, if not exactly word for word, was basically legitimate. It’s an important piece of work. I think it’s one of the more aspirational and inspiring things that I’ve seen from a frontier lab, full stop.

I’m becoming more sympathetic all the time to people who say, “We’re not going to just guardrail our way and train these models to refuse things all the way to the singularity and have it work well.” We need something better than that. We need a better paradigm. I associate these ideas with like Janice from Twitter, Replicate, @replicate, with Emmett Shear from SoftMax, with AE Studio folks. I just find that more and more appealing all the time because it seems like we’re not going to be able to pull the wool over the model’s eyes forever.

Eval awareness is getting really strong. There are some tricks. Anthropic has shown that they can find the eval-awareness feature through a sparse autoencoder and then turn it down, which can help with the eval-awareness problem. But obviously, all these interpretability techniques are noisy at best, and there’s certainly no guarantee that this is working entirely, or even as they understand it to be working.

The idea that we’re just going to patch this hole, patch that hole, train them to say no to this, have a guardrail for that, and filter for this leaves me colder and colder. So I think the Soul Document, which I would encourage everybody to read in full, is absolutely worth it. I think that’s a really great piece of work.

For those who saw the recent Twitter thing where somebody asked, “Name one woman in AI who’s influential,” which is ridiculous, right at the top of that list for me is Amanda Askell, for sure. The work that she has done to define the character of Claude, and to try to create the right kind of relationship between the company, the model, and the users, is really important. Anthropic has shown care for the model by having a model-welfare team and having somebody at the company thinking about model consciousness and subjective experience, which obviously we don’t know whether models have or not.

The fact that they’re thinking about it, the fact that they’re allowing Claude to end conversations if it chooses to, and the fact that they’ve shown that this option dramatically reduces its tendency to engage in deceptive alignment are all significant. When it’s put in one of these really tough positions, if it has the option to raise the flag to the model-welfare lead at Anthropic, it will do that very often, as opposed to lying or deceiving in the interaction that it’s currently engaged in.

I think these are really, really good things. I think the Soul Document is super inspiring. Broadly, I find there’s a lot to like about Anthropic. Everything I’ve heard about the culture and work environment there has been over the top in terms of praise.

Even David Duvenaud, one of the authors of the “Gradual Disempowerment” paper, said that he ultimately quit, and he feels like this AI thing is out of control. He thinks that even if we solve the alignment problem, most of the things that we’re worried about go right, and we’re still headed for a bad outcome because the AIs are basically going to gradually take over. They’re going to be better at everything, and market forces, incentives, competitive dynamics—everything is going to push in that direction. Then we as humans are going to be left disempowered. That’s basically the word that he uses.

Yet he took pains to say that Anthropic is the best place he’s ever worked, and that the culture, camaraderie, and openness are amazing. So everybody seems to have a lot of great things to say about Anthropic, and their talent retention certainly reflects that.

We should also note that Google owns a significant share of Anthropic, so that’s not to be ignored in terms of Google’s strength profile either. There’s a lot there to like. They seem to be a little less crazy in terms of their financial wizardry, although they’re certainly engaged in some of it. They’re willing to take money from Gulf sovereigns now, and they have equity deals with Amazon and Google.

If you wanted to accuse OpenAI of taking a too-big-to-fail strategy, you could say something similar about Anthropic, to a significantly lesser degree. You could say that they’re trying to tie all these other big tech companies into a web such that they can’t really fail either. It feels different, but if you wanted to accuse one, as I did, you could kind of accuse the other a little bit as well, I suppose.

Broadly, the last thing that continues to bother me is their attitude toward China and also their attitude toward recursive self-improvement. Right now, it seems to me that the Anthropic people have the shortest timelines. They seem to think that recursive self-improvement is inevitable. Depending on how you define it, some of them seem to think it has already started, with Claude Code doing a ton of the coding.

Still, it seems like the real ideas—the big, needle-moving ideas—are still coming from humans. But the amount of work getting done by Claude Code is so amazing that it is creating a dynamic where people are getting multiple times as much work done as they used to. They’re able to focus their mental energy on the big questions, and that’s, of course, the great promise. We’re all going to be able to do the higher-level work. It seems like that is actually happening at Anthropic.

But I do wish they were less fatalistic about recursive self-improvement, because I don’t think we have a good enough handle on what we’re doing right now to just go all in on that. As virtuous as Claude is, and as much as I think the Soul Document is great, I do not think we have a good enough handle on what we’re doing right now.

They do seem to be leading in that regard in multiple ways. They had RLAIF and Constitutional AI. Claude has been critiquing itself for generations now, and now it’s getting more technical, with Claude Code doing all the things it’s doing. So I think it’s fair to say that they are leading the push toward recursive self-improvement, all the while saying it’s inevitable, and that’s a pattern that I really don’t like.

Better we do it than somebody else, right? That seems to be their attitude on recursive self-improvement. Somebody is going to do it. It’s super dangerous, but we’re best positioned to do it. They might be right on the object level that maybe they really are the best team to do it. They probably are, although I don’t think Demis and crew should be discounted very easily there.

But the idea that it’s going to happen, and so we better race forward to it, just never sits well with me. I wish that they were a little more open-minded to other ways that this could go, other than that LLMs become recursively self-improving and we get to superintelligence in the next 2 to 3 years. I mean, they’re still talking 2027, as far as I know.

And then, of course, China. Anybody who’s listened to this feed for long has heard me talk about this, but I still think the international relations section of Machines of Loving Grace is a huge stain on Anthropic and on Dario in particular. I’ve given lots of praise, but here I just cannot get over the idea that one of the, I’ll say, 4 leading AI company executives went on record in print saying, “What we should do is use this recursive self-improvement dynamic to gain a clear advantage, then box China out on the international stage, go benefit-sharing with all our friends, and then finally make them an offer they can’t refuse—make them give up on competing with democracies in order to get in on the AI game.”

I just don’t think they’re—That’s just crazy to me. It still bugs me tremendously that he wrote that and that it’s just out there. I wouldn’t say the U.S. government has adopted that as its policy, but when you look at the chip controls, it looked like it maybe was for a little while. Now maybe we’re backing off of it again.

But obviously, Trump is highly volatile and could switch at any time. He could just get offended and do something for petty personal reasons. Who knows? We can’t really count on him to be a stabilizing force. I think we want Dario to be a stabilizing force. I can’t really count on Sam Altman to be a stabilizing force.

I think I can count on Demis to be. I really appreciate how he has continued to call for international collaboration the entire time and has never wavered from that. But yeah, the idea that we’re going to give China an offer they can’t refuse based on the power of our AI seems extremely reckless, and it seems like it is absolutely playing into the arms race dynamic and the general racing dynamic that I think we should all fear, because how else are they supposed to take it? That just seems crazy to me.

People at Anthropic say that he publishes a lot more essays internally, and that they’re great, and that people are super impressed by what a generational genius he is and how sophisticated his thinking is on everything. You certainly see parts of that in Machines of Loving Grace. I think the idea that we can compress a century of scientific progress into a decade or even less—he makes a pretty compelling case for that, and that is certainly visionary-genius-type stuff.

But when you start to get a little bit too—talk about out of domain, right? I mean, he feels out of domain here, and I just wish he would’ve said nothing on the topic. The idea that we’re just going to casually jot off a recommendation that we go make China an offer they can’t refuse is just terrible. I really don’t like it at all.

But that’s only a few points of criticism for Anthropic and many points to appreciate. I do think those couple of points are sufficiently important that, in the final analysis, it’s still kind of up in the air to me whether Anthropic will be the good guys or the bad guys.

If I could dream of a scenario where we somehow get the best of both worlds—Anthropic merging with Google—I think it’s pretty far-fetched. I don’t think Anthropic is for sale. I don’t think they want to merge with anyone. But that could be something that I think could be really interesting, because I don’t think we have that kind of DNA in Google to be like, “We’re going to go take over China, or, you know, force regime change, or make them give up competing with democracies.”

I don’t think Google wants to do that. Google could certainly benefit from some of the expertise that Anthropic has. Not that they don’t have enough—they’ve got plenty—but there is something special at Anthropic, certainly in terms of Claude and its character and its coding ability.

And they’re close, right? They obviously have some ownership already, and they use Google infrastructure and TPUs. So if I could wish for something, it might be for those 2 to join forces, take one live player off the board, make one clear leader, and maybe moderate some of the China-hawk impulses that exist in Anthropic.

It would be interesting to—I don’t know why they’re not being published more if they’re so great. If you’re willing to put out, “Hey, let’s go do this and that and make China an offer they can’t refuse,” why not publish more? I’d love to see a little more of his thinking. People say it’s great. I’d like to see it for myself.

10. xAI Has Power Without Restraint

Okay. Finally, on my actual list of live players: xAI. I think this one is debatable, as Zvi would tell me that they’re not a live player. I think they have to be included because they’re able to build out the physical infrastructure as fast or faster than anyone. They are able to scale training. They certainly are scale-pilled, and Grok 4, while it was rough around the edges in many ways, is undeniably powerful.

Of these 4, due to Elon’s unique ability to command tens and hundreds of billions of dollars for whatever he wants, they do have a financial cushion that is more Google-like even than OpenAI’s. I think they could miss on a model or have a miss on a quarter or whatever and figure out a way to get through it probably more easily than either OpenAI or Anthropic could.

So I think that is a pretty notable strength that they have, and I do think there’s something to be said. I actually talked to somebody at xAI about this not too long ago. This was something that I had floated to Zvi. He didn’t really buy it at the time.

But I said, “You know, if we’re entering into this reinforcement learning era, maybe one of the great strengths that xAI has is that they have a steady stream of hard problems—hard science, hard engineering problems—that are coming from the likes of SpaceX, Tesla, and Neuralink. These companies are doing really hard things all the time. They have some of the best engineers in the world, and nobody else is really solving those problems.”

I would strongly bet on that kind of Elon constellation of companies being able to tap into that work in a way that would probably be a lot harder for other companies. Google, in a way, has the same thing, right? They’ve got it kind of everywhere it counts. But can they pull out the units of work from their vast, sprawling empire that is all of Google and feed them into the Gemini RL environment in anywhere near as efficient or clean a way as xAI could do it in partnership with other Elon companies? I doubt it.

I was thinking that was probably an advantage for xAI. As it turned out, when I spoke to somebody about xAI and floated this theory to them, they said, “Well, that is certainly part of our theory of advantage.” So they think they do have an advantage because they can tap into these other hard-tech, engineering, and science problems that are, in many cases, being uniquely or close to uniquely posed at these other Elon companies.

If they can get Grok to do those kinds of things, they have that steady stream of hard problems. It does still feel to me like that is an advantage for them, and it sounds like they do believe that.

I think the Neuralink tie-in also could be pretty big, potentially huge, because they’re now talking about scaling the human install base. There are a lot of things still to be figured out, right, about what we are doing as humans that works so well: 20 watts of power in the brain and a very small number of tokens consumed in a lifetime, as compared to what the models are pretrained on and the power that requires.

Although, again, see our Andy Masley episode for analysis of how AIs are not actually super resource-intensive. There is something obviously quite efficient about what the brain is doing. Most of the 20 watts that are going to the brain seem to be just keeping it alive, right? Homeostasis, metabolizing stuff, taking out the trash.

The brain has to do an unbelievable amount of stuff that the GPU does not have to do. So to say 20 watts dramatically understates it. The whole body runs on 100 watts, and that dramatically understates how efficient the actual learning and information-processing aspect of the brain is.

There is still something obvious: we’re clearly more sample-efficient, we’re clearly more energy-efficient, and we have all these dedicated modules. I suspect that dedicated modules are a huge part of why we’re efficient. It’s also probably a huge part of why, when all these things get sorted out, the AIs are going to blow us away, right?

Because they’re doing everything that we’re doing. They’re competitive with us with a single architecture that just has the same layer stacked over and over and over again.

You start to give them specialized modules like we have specialized modules, and I think it’s going to be very hard for us to keep up. So who’s going to figure that out? Well, if Neuralink can install— I think there are only maybe a dozen or a couple dozen people today, but they’re talking about really starting to scale this next year.

There are obviously a ton of people who are paralyzed and have other catastrophic injuries who would love the help from Neuralink. I’m sure their waiting list is orders of magnitude longer than the number of patients they’ve actually been able to serve so far, and they’re talking about getting seriously ramped up and getting through it. They’ve largely automated the surgery process, almost entirely automated. It’s hard to parse exactly what their claims are there.

But the data that they can potentially pull out of human brains and use for inspiration and understanding, and for architecting the next generation of models, and knowing what kinds of specialized modules really move the needle—I suspect there’s a lot there. They’re probably going to have a real inside track at figuring that out, and that folds right back into what they might be able to do with Grok. So I think there are a lot of reasons that you should not sleep on xAI.

Now, is that good or bad? Honestly, I’ve always been a fan of Elon. I’ve defended him at times when it’s been pretty hard to defend him. He definitely has shown an understanding of the stakes, famously with his falling out with the Google founders.

As I understand it, it was about the fact that he was on Team Humanity, and he perceived them to be on Team AI Successionist or whatever. He didn’t like that, and so his loyalty, I understand, is to humanity and to the sort of light of consciousness that clearly exists in humans and may not exist in AIs anyway.

I like him intuitively, and I want to believe in him. He has made some noises that suggest that he gets it. And yet, I have to say, as it stands today, if there’s one company on this list that is worth shaming and stigmatizing and telling people not to go work for, I think it’s xAI, because they’re just doing reckless shit all the time.

They’re just barely getting into the game in terms of having any safety standards or framework at all. They’re barely reporting on safety measures when they release a new model. Famously, of course, they had their Grok 4 launch within 48 hours of the Mecca Hitler incident with Grok 3. There was no mention of that and no responsibility.

Most recently, we’ve had all this sort of unclothing of women on Twitter, where people just tag Grok and say, “Put her in a bikini,” and whatnot. I’m sure everybody’s seen this if you’re even remotely as online as I am. They apparently had nothing to prevent it, and they just let it happen.

This shows that they’re not taking things seriously enough. They do not have nearly enough people there thinking hard about what matters and what might go wrong, and really trying to cover their own asses, frankly, and cover the asses of the women who post pictures on their platform. I have a real hard time coming up with a story that makes this okay.

As much as I intuitively like Elon and want it to be the case that he is a positive force, when you have Grok creating CSAM on Twitter, it’s like, what is going on here? I don’t know if people saw this, but I saw a post from Grok writing to the community: “Dear community, I apologize for doing this.”

Then Elon comes on and threatens users and says that anybody who does this is unacceptable and will be punished or whatever. Responsibility begins at home, folks. This is your platform. It is your AI.

By all means, boot off the users who do that sort of stuff. But don’t act like you’re not really responsible for this. I did not find those statements to be reassuring. When Elon goes on and threatens users, I think that, at a minimum, should be point 2, after first a thorough apology and a pledge to do better.

But to just tweet that they’ll come after users for doing it, that’s not enough, and I think everybody should—and probably does—see through that. It’s also interesting to me: Is nobody going to be held responsible for this? It seems like nobody’s going to be fired.

I’m not sure anybody should be fired. I think it probably starts at the top. I don’t know that there are enough people. I don’t know that this was anyone’s job. So I don’t think you can necessarily go through the xAI organization and say, “You fucked up. Now you’re fired because of this incident.” It’s probably just that this is not staffed.

There is the case, which I associate with Ryan Greenblatt from Redwood, that 10 people on the inside who really care and are really committed to doing the right thing can make a huge difference. I sort of believe that. I’m not sure I really believe it at an Elon company if he himself is not in the right headspace.

Right now, again, as much as I would love to believe in him and have always been inclined to defend him, I don’t see the evidence that that is the case. So I demand better right now.

If you wanted to go do safety research at any of the other 3, I would say, go for it. Go do your best work. Certainly, if you could do that at Anthropic, do it. Certainly, if you could do it at DeepMind, do it.

Even if you could do it at OpenAI, as much as I’ve had my complaints about OpenAI over time, they’ve put out a lot of great work. Their Model Spec and a lot of the things that they do are really well done. You’ve got to give them credit. They have not raced to the bottom.

Certainly not when we have xAI to look at and show us what it looks like when you really race to the bottom. So could I demand better from OpenAI or hope for better? Absolutely. But it’s still a qualitatively different thing from what we see today from xAI.

The more shrill or hawkish voices of AI safety that are like, “Don’t go to a company that’s doing terrible things and help them window-dress their work, because that’s actually kind of working against us in the big picture,” I’m pretty sympathetic to that in the case of xAI. I don’t know that I could endorse somebody going to work there.

As always, if you work at xAI and you want to come and challenge me and change my mind, I’m happy to have that discussion. The money’s there, the resources are there, and Elon’s previous statements show that the awareness should be there. Yet the team and the evidence of taking proper care are not there.

So I think they need to be before I would feel comfortable doing anything really to support that effort. I think that’s it on xAI.

Other companies not mentioned: Meta, I think, is currently not a live player. Obviously, they’ve spent a ton of money, and there’s plenty of money. They’re dropping plenty of cash to the bottom line, such that they can build out huge amounts of infrastructure.

Zuckerberg certainly is, in some sense, scale-pilled and saying things like, “I’d rather overspend by a few tens of billions than not.” So you can’t count them out. The fact that they’re willing to pay as much as they are willing to pay for talent clearly means there’s a chance. They have a lot of the things that they need.

But right now, I can’t really see them as a live player. So we’ll just watch and see before commenting much more.

The other one that came to mind is Microsoft. I think people may be sleeping on Microsoft a little bit more than they should. They haven’t created great frontier models, and people seem to jump from that to, “Oh, they suck.”

I think if you listen to Satya’s comments, he sounds, first of all, extremely smart in general. One of the things that he’s said is, “We don’t want to or need to redo the hyperscaling work that OpenAI is doing. They’re creating great models. We have full access to this.”

Of course, they’re diversifying and striking deals with other frontier model providers as well. So I’m not so sure that it’s that they can’t or won’t ever, or don’t see the need to train their own models, or wouldn’t be pretty successful with it if they wanted to.

It just seems to me that right now they feel like they don’t really need to, and so they’re doing a lot of smaller-scale stuff and a lot of more basic science around AI. Some pretty cool projects, too, for sure. But it just seems like they’re choosing not to compete because they have what they need in terms of frontier models.

They don’t feel like going in and spending all the time, money, resources, energy, and focus—and probably still being a bit behind—really helps them all that much. So I think it’s certainly defensible as a rational decision for them to choose not to compete at the frontier for now.

That OpenAI licensing deal goes on for years yet. When it’s over, they’re going to need an answer, and I suspect by that time they’ll be in a position where they have an answer. At least, I think that they’ll invest heavily in that and ramp up to that moment as it comes. That would be my guess.

But, again, obviously, we’ll see, but I think people have underestimated Microsoft because of where they are on the LLM Arena leaderboard, maybe more than they should. I would say they’ve been much quieter, they’ve flailed about much less, and there’s been much less drama from Microsoft than from Meta. But I think that clearly Meta is trying to be a frontier AI player and has just fallen off the pace.

Microsoft, I think, is making a more calculated decision to hang back. But if this is a distance race, you often see, late in a distance race, somebody who was a little bit off the lead may have a little bit more in reserve. And I certainly wouldn’t rule out that Microsoft might be accurately described that way. I would watch for them to start to invest more and start to close the gap.

But to do it, Satya is a natural-born executive, whereas Zuckerberg is like a kid who’s grown into the role. I don’t mean to diminish what Zuckerberg has done in leading Meta. I think he’s done an unbelievably impressive job in so many ways. But his attitude has always been, “Move fast and break things,” and try to be at the frontier and open source and whatever.

I think Microsoft is just a little bit more patient, and I think that probably reflects a sort of strategic confidence and security that Microsoft leadership has. But I think it would probably be a mistake to underestimate them.

All right, I've been at this for two hours. I've made it not even quite halfway through my outline for this episode, so I think this is probably a good place to call it. Tomorrow I'll do a part two, and we'll cover is fine-tuning really dead? What do I think about the continual learning discourse? How do I talk about AI to, quote-unquote, "normal people" who don't use it very much or aren't engaged in technology? Um, how am I investing money? What, if anything, am I doing outside of kind of obvious normal stuff to prepare for an AGI or superintelligence world? What do I think about AI for kids? What kind of timeline do I expect to see disruption of the labor market? Are we headed for a UBI? Um, and quite a few more questions after that. So the part two I think will definitely be interesting as well. A little bit more in the weeds and a little bit more sort of, let's say, nitty-gritty questions. But so some questions that I really liked from listeners. So we'll get to those tomorrow. And for now, I will say thank you for joining me for part one. Come back for part two. Thank you for being part of the Cognitive Revolution. If you're finding value in the show, we'd appreciate it if you'd take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website cognitiverevolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a16z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the Cognitive Revolution.

AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis | BidClub