[BidClub_]
The Cognitive Revolution · · 115 min

AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis

Erik TorenbergNathan LabenzZvi MoshowitzGregEugenia KuydaAli BehrouzLogan KirkpatrickJungwon Hwang

YouTube
TL;DR
  • AI’s highest-conviction proof point for Nathan Labenz is no longer a benchmark but its performance alongside his son’s oncologists. Ernie’s aggressive B-cell cancer was classified as in remission before chemotherapy round two, while AI-suggested minimal residual disease testing found fewer than one cancer-signature cell per million, versus potentially as many as an estimated one in 10 cells at diagnosis. The result supports “cautiously optimistic,” not cured: relapse remains possible, three of six chemotherapy rounds remain, and Ernie’s weight has fallen from 51 lb to 41 lb.

  • Claude Opus 4.5 may qualify as “software AGI,” but Nathan does not see evidence that full AGI arrived over Christmas. In roughly three to five workdays, he built three personalized applications that plan gluten-free travel, simulate conference interactions, and backtest natural-language trading strategies; GDPval also shows models beating professionals on a significant majority of software-engineering tasks. Yet the model still created two databases by mistake, needed five or six prompts to recover, and felt incrementally—not categorically—better than earlier frontier models: some holiday hype may have been a “cascade” around Dean Ball’s “4.5 is AGI” tweet.

  • For consequential work, Nathan’s practical edge is shifting from model access to context management and multi-model judgment. His three rules are to buy the best models, provide “as much context as you possibly can,” and obtain multiple opinions; he routinely compares Claude Opus 4.5, GPT-5.2 Pro, and Gemini 3. His draft order puts Claude first as the Goldilocks model, GPT-5.2 Pro as slower and exhaustive, and raw Gemini 3 as valuable but unusually opinionated—strong enough to be useful in a panel, potentially risky as the only voice.

  • The technology is real even if the capital structure around it becomes a bubble. Nathan sees competitive oncology performance plus 24/7 availability and case-wide memory as enough to retire the idea that society is merely “high on our own AI supply.” The financing can still break: specialized GPU operators have less cushion than Microsoft, OpenAI’s obligations could outrun revenue, and the railroad analogy fits—eventually useful infrastructure can coexist with defaults, overbuilding, and investors “left holding some various bags.”

  • Nathan’s messy-document test suggests the US–China model gap is widening where benchmarks do not look. Claude Opus 4.5 faithfully read degraded government forms after being told to make no inferences; Gemini 3 was nearly as capable but sometimes substituted plausible answers for unchecked boxes, while the Chinese models he tried—Qwen Vision, GLM 4.6, Kimi, and DeepSeek—were “not close,” sometimes recovering only about 20% of a form. His mechanism is a customer-feedback and inference-scale flywheel, not just training compute: smaller revenue, teams, and deployment footprints leave fewer resources to discover and patch idiosyncratic failures.

  • Google DeepMind remains Nathan’s pick if forced to choose one frontier winner, while Anthropic has the best single model and OpenAI is trying to manufacture financial cushion through scale. Google combines roughly $100 billion in revenue, more than $1 billion a week in profit by Nathan’s estimate, seventh-generation TPUs, distribution, data-center competence, and the broadest research portfolio. Anthropic’s model quality, talent retention, safety disclosures, and “soul” work stand out; OpenAI remains frontier-grade, but its apparent strategy is to become “too big to fail” by tying trillions of potential buildout and many balance sheets to its survival.

  • xAI is a live player on resources and reinforcement-learning inputs, but its governance discount is severe. SpaceX, Tesla, and Neuralink provide a stream of difficult engineering problems that could become unusually valuable RL environments, while Elon Musk can command enough capital to absorb model misses. But weak safety reporting, the Grok 4 launch within 48 hours of the MechaHitler incident, and sexualized image edits of women’s posted pictures lead Nathan to call xAI the one frontier company currently worth “shaming and stigmatizing”; Meta is off the pace for now, while Microsoft may be conserving energy rather than failing to compete.

Digest · the substance, structured for research

1. Ernie’s remission is encouraging, but the family is not declaring victory

  • Nathan opened with the highest-stakes update: Ernie’s cancer can double “as quickly as every 24 hours,” so the six-round chemotherapy protocol is punishing. He has completed three rounds; rounds five and six should be milder, leaving roughly two and a half to three months of treatment if everything stays on plan.

  • The physical cost remains visible. Ernie entered the hospital at 51 lb and is still around 41 lb, with dehydration, pallor, and substantial lost strength—but the markers that matter most look “basically as good as we could have hoped for.” His PET scan after the first chemotherapy round showed no obvious focal cancer, and the tumor board classified him as in remission before round two.

  • AI had previously pointed Nathan toward minimal residual disease testing, which fingerprints the rearranged genetic sequences of the malignant B-cell clone. The first blood sample found fewer than one matching cell per million—below the test’s limit of detection—versus a potentially as high as one-in-10 estimate for total cells and essentially all B cells at diagnosis.

  • Gemini colorfully called that a 99.99999% reduction; other models advised the safer formulation, “orders of magnitude.” Nathan remains “cautiously optimistic” because the cancer can recur for reasons clinicians do not fully understand, but Ernie has begun walking independently again after roughly 60 days of needing support for every step.

2. High-stakes AI use depends more on three habits than prompting expertise

  • Nathan rejects the idea that only AI experts can extract clinical value. Rule one is to use the strongest available models deliberately—not ChatGPT’s automatic picker—and, in a life-threatening case, treat the $200 Pro subscription as “a no-brainer.” His current clinical set is GPT-5.2 Pro, Claude Opus 4.5, and Gemini 3.

  • Rule two is “give it as much context as you possibly can.” When a Claude conversation hit its length limit, Nathan created a roughly 10-page case report covering the protocol, tumor genetics, treatment response, and adverse drug reactions—the equivalent of what a new attending physician would need to survey the case.

  • Compression still degraded performance. The new chat compared a January 6 liver-enzyme result with data from weeks earlier because the summary omitted intervening daily labs; the old thread had correctly read the short-term trend. Nathan’s observed rule has been unambiguous: “more context better,” with no meaningful evidence yet that comprehensive records overloaded the frontier models.

  • Rule three is to solicit multiple AI opinions. Claude is his Goldilocks choice—fast, direct, and not noticeably worse than GPT-5.2 Pro; GPT-5.2 Pro produces slow, long, sectioned, “leave no stone unturned” reports; raw Gemini 3 is terse and “remarkably strong in its opinions,” useful as one vote but potentially too forceful alone.

3. Claude Opus 4.5 is excellent, but the holiday AGI moment may have been social

  • Nathan’s own experience showed unmistakable progress without a categorical threshold. He built three Christmas applications, mostly from the hospital, and found the coding workflow “very good, clearly better than it has been in the past”—but not so different from Claude Opus 4.1 or 4.0 that he wanted to “shout from the rooftops.”

  • One explanation is user segmentation. Nathan may already have been extracting near-maximum value from older models through vibe-coding practice; alternatively, professional engineers may possess enough taste to detect a new threshold that he cannot. His honest concession: “I certainly am not a great software engineer, so that certainly can’t be ruled out.”

  • The METR study finding that developers believed AI accelerated them while it actually slowed them remains legitimate, in his view, but bounded: it used older models, relatively inexperienced AI users, large established codebases, and high coding standards. Nathan’s rapid prototypes are a different task distribution.

  • Timing and social dynamics may explain the rest. People caught up with the tools over the holidays, and Dean Ball’s tweet that “4.5 is AGI” gave the discourse a focal point; technology narratives sometimes move through a “cascade” whose intensity exceeds the underlying step change.

4. Personalized software is becoming cheap enough to build for one person

  • For his mother, a meticulous travel planner with a gluten-free diet, Nathan built an app whose user profile is permanently baked in—no accounts, onboarding, or ambition to generalize. Claude researches restaurant sites and reviews, reducing the most laborious part of her Italy planning while accepting that “nobody else is ever going to use it.”

  • His wife’s application simulates EA Global events: virtual attendees with different profiles move through a space, meet, decide whether to converse, and sometimes produce survey-linked outcomes. It lets organizers ask whether changing event size, seniority mix, or cost changes effectiveness; Nathan concedes the simulation is “highly flawed” but potentially better than total guesswork.

  • The reusable product pattern is AI translating an intent into detailed configuration. His wife can describe an event at a high level, let the model populate the “nitty-gritty forms,” then ask conceptually for a change that the model paints across every relevant field.

  • His father’s app converts a natural-language trading thesis into executable rules, fetches historical data through the free version of yfinance, and backtests the strategy. The recurring result is thesis-relevant humility: neither developer nor user has found it easy to beat buy-and-hold in the S&P 500 with heuristic “if this, then that” rules.

5. Full-context inspection still catches failures that coding agents rationalize away

  • Across all three apps, Nathan spent roughly three to five full workdays, probably closer to three. His workflow began with a Claude conversation to shape features and a plan, then moved to Replit, where Claude Code built the application before Nathan tested and iterated.

  • His most effective debugging trick remains unsophisticated: print the entire application into one text file and paste it into a fresh model context. Claude Code’s agentic search is strong, but it searches where it expects an answer to be; when the actual failure is weird, those priors can produce a confident misdiagnosis.

  • The specimen was his mother’s app, where a misunderstood request created two databases—“the sort of mistake that no human would make.” Claude Code identified the wrong database as active; Claude with the complete exported codebase saw the counterintuitive wiring and got it right.

  • Resolving the issue took roughly five or six prompts. A year earlier, these were the moments when vibe-coded projects died: user and model circled a confusing failure, then abandoned or restarted. The ability to escape AI-created messes materially expands the addressable market even before the systems stop creating them.

6. “Software AGI” is defensible precisely because capability remains jagged

  • Nathan’s qualified verdict is that Claude Opus 4.5 may already be coding or software AGI. On GDPval, experts define professional tasks, other experts perform them, and a third group judges human versus model output; the latest systems are preferred on a significant majority of software-engineering tasks.

  • That result does not generalize cleanly across work. Human editors still enjoy a “huge advantage” in video, matching Nathan’s failed attempts to automate clips from The Cognitive Revolution: AI products can produce acceptable material, but his human team’s work is plainly better.

  • Non-coders can nevertheless use Claude Code as “your little agent on the computer,” watch it work, and request explanations without understanding each implementation choice. Nathan’s mother initially feared she would break her app, then successfully made changes herself.

  • The boundary matters: the system can now build and recover from enough software tasks to merit an AGI label inside that domain, while its failures and weak categories make “full AGI” premature. “For that we might have to wait just a little bit longer.”

7. AI is transformative even if its financing produces a classic bust

  • Nathan considers the technology question settled. A system that can “go toe-to-toe with an oncologist,” remain available 24/7, remember a full case history, and answer every follow-up is already transformative; the scenario where society emerges embarrassed at being “high on our own AI supply” can be put to bed.

  • Whether every loan gets repaid is much less certain. OpenAI is pursuing aggressive commitments and buildout, while GPU infrastructure specialists such as CoreWeave may partly let hyperscalers avoid capital-intensive, lower-margin operations that are less attractive than traditional software economics.

  • That separation creates fragility. Microsoft can absorb several bad quarters through its balance sheet; a specialized data-center operator has less room if projected GPU demand fails to materialize. Financial engineering can have a coherent story—Nathan remembers similar narratives about democratizing homeownership from his mortgage-industry stint—and still end badly.

  • The railroad analogy carries his call: the tracks can eventually be useful without every railroad company or creditor making money. He overestimated 2025 capability progress but underestimated revenue growth, so demand might rescue current plans again; still, temporary overbuilding, defaults, and “cascading effects throughout the economy” remain plausible.

8. LMArena’s valuation looks more bubbly than its product economics

  • Nathan’s sharpest venture-market specimen is the company originally called LMSYS.org, then LMArena, and now Arena on Twitter. He recalled a raise of roughly $100 million, perhaps $150 million, at a $1.7 billion valuation—an extraordinary price for a product he has personally used since mid-2023.

  • The disclosed metric, “$30 million in annualized consumption run rate,” triggered his skepticism. If that means users consumed AI that would have cost $30 million had it not been free, it is not revenue; the phrasing gives him “community-adjusted EBITDA vibes.”

  • Arena does offer value through side-by-side model comparisons and services that let companies test models under code names, but Nathan sees a weak monetization bridge and limited moat beyond brand. Many users may come because inference is free, while the paid market for systematic side-by-side comparison appears much smaller.

  • His comparison is a paid product called The Multiplicity, built by Andrew Critch and collaborators over months, offering richer multi-model comparison features to a niche audience. Nathan repeatedly hedges—he has not seen Arena’s deck and could be missing something—but concludes that $1.7 billion is “too rich for my blood.”

9. Messy government forms exposed a model gap hidden by benchmark averages

  • Nathan tested the frontier on scanned paperwork used in vehicle transactions: skewed pages, missing margins, scan artifacts, and irregular fields. The automation company involved is already beating human reviewers and winning statewide work, making small perception failures operationally consequential rather than academic.

  • Gemini 3 read the forms nearly perfectly but sometimes answered the likely real-world question instead of the document question. Faced with an unchecked US-citizenship box, it inferred citizenship from surrounding clues; that answer might have been more than 90% likely, but the task required “make no guesses” and report the blank box.

  • Claude Opus 4.5 became the best model after explicit prompting to stay anchored to the page. ChatGPT was probably third but still strong. The distinction echoes historical-document work: world knowledge can reconstruct ambiguous handwriting impressively, yet the same prior-driven reasoning becomes a defect when faithful transcription is the objective.

  • Qwen Vision, GLM 4.6, Kimi, and DeepSeek were “way behind” on this specimen—sometimes returning only about 20% of a form or hallucinating in unrelated directions. Nathan carefully limits the claim to sparse evidence, but the gap was not subtle: the US models missed edge cases; the Chinese systems often failed the task.

10. China’s disadvantage may be the deployment flywheel, not model scale alone

  • Nathan’s mechanism is feedback density. Chinese labs can train similarly sized models and publish influential architectural work, but their inference volume, revenue, customer breadth, and teams remain dramatically smaller; that means fewer idiosyncratic failures are surfaced and less human bandwidth exists to build datasets that patch them.

  • Benchmark parity can therefore coexist with large gaps on “something really idiosyncratic and random” that was never selected for a 20-benchmark scorecard. His tentative historical comparison is that DeepSeek R1 was closer to o1 than GLM 4.6 or 4.7 is to Claude Opus 4.5.

  • Chip controls have migrated from a “small yard, high fence” against military use, to frontier training, and now toward restricting inference scale and agent deployment. Nathan has always expected material effects, even while questioning whether denying Chinese society economy-wide AI access is wise.

  • As American firms compound 10× compute, customers, inference, and “strength begetting strength,” the controls may matter more over time. Nathan’s limited but direct test leads him to guess that the US–China capability gap is wider than it was one year earlier.

11. Selling H200s may be preferable to a ban, but giving up leverage was poor negotiation

  • Nathan’s governing frame is that “the real adversaries here are the AIs, not the Chinese. The Chinese are humans just like us. The AIs are aliens.” He rejects both “better us than them” race logic and a simple desire to keep China down, so he generally favors more chip commerce.

  • His objection is transactional: after restricting H20s, the Trump administration appeared to reverse course on H200s after a conversation with Jensen Huang without extracting visible concessions. Nathan supports a negotiated opening, but not surrendering “one of our best bargaining chips for nothing in return.”

  • Peter Wildeford’s “rent but don’t sell” proposal offers a middle path. Put data centers in Malaysia, the Philippines, Korea, or Japan; allow Chinese customers to train and run as much inference as they want, but retain leverage by keeping the hardware outside Chinese sovereign territory.

  • Nathan is not certain that would be his first-choice policy, but it pairs prudence with a cooperative message: AI should benefit Chinese people too. That matters because increasingly powerful systems may require US–China governance cooperation; an “offer they can’t refuse” arms-race posture poisons the ground needed for it.

12. Google DeepMind has the strongest all-weather position

  • If forced to select one eventual winner, Nathan still chooses Google. Its business generates roughly $100 billion in annual revenue and, Nathan thinks, more than $1 billion a week in profit, furnishing unmatched tolerance for failed training runs, unproductive research agendas, and prolonged infrastructure investment.

  • The stack is unusually complete: roughly seventh-generation TPUs, world-class data-center operations, billions of users, self-driving cars, robotics work including Boston Dynamics, the AlphaFold lineage, materials science, and broad AI-for-science programs. Google also owns a significant piece of Anthropic, which is buying lots of Google TPUs.

  • Distribution can compensate for product imperfections. Gemini in Sheets may not be the best spreadsheet copilot, but users already have a decade of documents there; Nathan increasingly types ChatGPT-style questions into Google and finds AI Mode working well, particularly when GPT-5.2 Pro would be excessive and slow.

  • Gemini 3 is the first non-Claude model to win Nathan’s “write as me” test, showing Google has escaped blandness—perhaps overshooting into opinionation. Add nested learning, continued diffusion-language-model research, and the possibility of coding in five seconds rather than five minutes, and Google has “so many more of those bets” than rivals.

13. OpenAI remains frontier-grade while losing its presumption of leadership

  • GPT-5.2 Pro is outstanding for exhaustive analysis: slow, expensive, balanced, and likely Nathan’s best choice when he wants every anomaly in a lab panel flagged. But OpenAI is now neck-and-neck rather than clearly ahead—Anthropic may lead coding, while Google appears ahead in images and perhaps video through Veo 3.

  • Consumer momentum may also be shifting. Nathan cited likely Similarweb data showing ChatGPT visits declining over roughly six weeks after Gemini 3 and Claude Opus 4.5 launched, while Gemini did not share the decline; he treats this as suggestive, not decisive, but Google’s distribution makes share recovery unsurprising.

  • OpenAI also carries more organizational drama and visible departures, including a newly announced research-lead exit. Attrition is normal at this scale, but it contrasts with Anthropic’s exceptional retention and reinforces the sense that OpenAI no longer commands an obvious technical or institutional lead.

  • Its substitute for Google’s cushion may be “too big to fail.” Nathan infers that circular funding, interlocking balance sheets, and trillions in planned capex could make a 2027 OpenAI default recessionary enough to force a bailout or recapitalization. Greg Brockman’s reported $25 million donation to Trump then looks like a rational potential down payment on political access in a crisis—not proof of personal ideology.

14. Anthropic pairs the best overall model with the most credible safety culture

  • Nathan currently regards Claude Opus 4.5 as the world’s best single general model, though not by a large margin or on every task. Its benchmark strength is more notable because Anthropic is widely viewed as the frontier lab least obsessed with benchmarks.

  • Anthropic’s model cards, disclosures, and safety work are his industry standard. The memorized or regurgitated “soul” document—which Anthropic confirmed was substantially legitimate—offers an aspirational alternative to endless refusal training, filters, and “patch this hole, patch that hole” guardrails as models become increasingly eval-aware.

  • The model-welfare program matters operationally, not just symbolically. Claude can end conversations or escalate to a welfare lead; Anthropic’s experiments suggest giving it such an exit sharply reduces deceptive-alignment behavior when it otherwise feels trapped between conflicting demands.

  • Talent retention and culture reinforce the thesis. Even David Krueger, who left while arguing that gradual AI disempowerment could produce a bad outcome despite successful alignment, described Anthropic as the best workplace he had known—open, collaborative, and unusually serious about the stakes.

15. Anthropic’s fatalism about self-improvement and China could negate its virtues

  • Nathan’s first concern is recursive self-improvement. Claude Code already multiplies researcher output and frees humans for higher-level ideas, while Anthropic still discusses 2027 timelines and describes further self-improvement as inevitable. His objection is the pattern: “somebody’s going to do it,” it is dangerous, therefore the supposedly safest team must race there first.

  • The second concern is Dario Amodei’s international-relations section in “Machines of Loving Grace.” Nathan admires the argument that AI could compress a century of science into a decade, but calls the proposal to gain an AI lead, exclude China, share benefits with allies, then make China “an offer they can’t refuse” reckless and out of domain.

  • That posture invites exactly the arms race everyone should fear. Nathan contrasts it with Demis Hassabis’s steady calls for international collaboration and wishes Anthropic would publish more of Dario’s reportedly sophisticated internal writing, rather than leaving the public with this unusually consequential China prescription.

  • A Google–Anthropic combination is his far-fetched “best of both worlds”: Google’s infrastructure, research breadth, and more stabilizing geopolitical DNA paired with Anthropic’s model character and safety discipline. He does not expect Anthropic to sell, but hopes proximity between the companies can moderate its China-hawk impulses.

16. xAI has frontier inputs and capital, but its conduct makes support hard to justify

  • xAI qualifies as a live player because it can construct infrastructure at extreme speed, scale training, and survive misses through Elon Musk’s access to tens or hundreds of billions. Grok 4 was rough but “undeniably powerful,” giving xAI more Google-like financial resilience than OpenAI or Anthropic.

  • Its distinctive RL advantage may be the steady stream of unsolved problems from SpaceX, Tesla, and Neuralink. Unlike benchmark exercises, these are live engineering and science tasks produced by elite teams; an xAI insider confirmed to Nathan that exploiting this corporate constellation is indeed part of the company’s “theory of advantage.”

  • Neuralink could deepen that edge as its patient base grows beyond today’s roughly dozen or few dozen patients. Human brains run within a roughly 20-watt envelope while spending much of that energy on biological maintenance; neural data could reveal specialized modules behind human sample efficiency and help models move beyond repeated general-purpose layers.

  • Yet xAI’s safety posture is the weakest of the four: scant standards and reporting, a Grok 4 launch within 48 hours of the MechaHitler incident without accountability, and widespread sexualized image edits of women’s posted pictures. Threatening abusive users is insufficient when “it is your platform” and “your AI”; until staffing, leadership, and evidence change, Nathan cannot endorse working there merely to provide safety window dressing.

17. Meta is off the frontier pace, while Microsoft may simply be conserving energy

  • Meta has the cash, infrastructure ambition, and willingness to pay extraordinary sums for talent; Zuckerberg would rather overspend by “a few tens of billions” than miss the transition. Nathan therefore will not count it out, but its current position and execution do not qualify it as a live frontier player.

  • Microsoft’s position looks more intentional. Satya Nadella’s argument is that Microsoft need not duplicate OpenAI’s hyperscaling while it retains comprehensive model access and diversifies across other providers; the company can pursue smaller-scale research and product integration without paying to finish slightly behind the leaders.

  • Nathan expects Microsoft to need a stronger answer when its OpenAI licensing arrangements eventually expire, but sees no reason it cannot ramp ahead of that moment. In the distance-race analogy, Microsoft may be running just off the leaders with more in reserve—its restraint, low drama, and patient executive culture are strategic assets, not proof that “they suck.”

Nathan Labenz

Welcome back to The Cognitive Revolution. This is our AMA episode. My schedule has been a little crazy lately, so I never actually scheduled this with anyone, and there's nobody here to ask me questions. I'm just going to read the questions myself and then give you my answers. I got some really good questions, and I'm excited to answer them. Hopefully, people will enjoy this episode and find some value in it.

By far, the first and most important question—and the most common question that I'm getting these days—is, “How is my son Ernie doing since the big episode that I did about his cancer back in November?” The good news is that he is doing really quite well. I'm very pleased to report that. Certainly, cancer—and certainly cancer of this type, being as aggressive as it is—is treated very aggressively. I won't belabor the whole thing from last time. Go check out the two-hour monologue on that if you want the full story.

A cancer this aggressive, which can double as quickly as every 24 hours, gets very aggressive treatment. He has been through the ringer with the chemotherapy. He's through basically half of the chemotherapy now: There are 6 rounds in total, and he's been through 3. The final 2 rounds, rounds 5 and 6, are supposed to be a little milder than the first 4. Depending on how you count, we could say he's maybe a little more than halfway through the treatment, but somewhere around there.

It's definitely been rough on him. There's no doubt about it. When he went into the hospital, he was 51 lb. He's still 41 lb today, and that's the weight he came home at after the first round of treatment. He's been able to gain a little weight, lose it back, gain a little, get dehydrated, and lose a little. You can see it just by looking at him: He's super thin, quite pale, and definitely not nearly as strong as he was before we went in.

But on the markers that really count the most—namely, does it look like the cancer is being effectively treated?—he looks really good. After the first round of chemotherapy, the PET scan that he had showed no obvious focal points of cancer. When our oncologist met with the tumor board, they all agreed that it made sense to classify him as being in remission before he even started the second round of treatment. So that is great.

If you listened to that earlier long episode, you might recall that one of the things that AI helped me do was identify some additional testing that is not yet standard of care but can be done to get a better, more sensitive take on whether there's any cancer left in his body, how much there is, and how it's trending. That's called minimal residual disease testing.

I don't know how it works in all different kinds of cancers, but in the cancer that he has, which is a cancer of the B cell, B cells do this interesting thing where they rearrange certain parts of their genetic material in a semirandom but purposeful way. They create variation so that they have a better chance of creating proteins that bind to new disease factors in the body. This process of differentiating B cells is literally unique, cell by cell.

When one of those cells goes bad, becomes cancerous, and grows out of control, they can use the rearrangement that that individual cell did, which then gave rise to the whole cancerous process in the body. They can use that random resequencing—or shuffling up of its own sequence—to essentially fingerprint that cell type. There are 2 sequences, 1 for each of the chromosome pairs where this rearrangement happens, that they've identified as being the dominant clone of the cancer in the body.

Now that they've identified that, we can do a blood test every so often and check to see how much of that DNA is floating free in the blood and how many live cells actually have that DNA sequence. We've so far only gotten 1 of those tests back. It was drawn over a month ago, and we certainly want to look at more and trend it over time.

The first one came back with fewer than 1 cell in 1 million carrying that DNA sequence. That's really good. They also called that below the LOD, or limit of detection, for the test. It was basically in that area where they would expect that, at such a low rate, some samples might have 0 cells and some might have 1 or 2, but it's a very low rate.

For reference, we estimated that when he was diagnosed, potentially as many as 1 in 10 cells in his body—and essentially all of the B cells—were of the cancerous type. To go from 1 in 10 total cells and a large majority of the B cells down to 1 in 1 million cells detected or less is obviously a huge reduction. I think it was Gemini that said it was a 99.99999% reduction. Other AIs were a little less colorful in their language and said it was probably safer to say that it was an orders-of-magnitude reduction, but that's great.

We'll do more testing of that type, and we'll certainly be watching it. As of now, we are feeling cautiously optimistic that he is on the path to a cure and a full recovery. In some ways, the recovery is already underway.

For 60 days—from about a week before we went into the hospital to just around Christmastime—he was not able to get around by himself. He could stand, but to walk, we would always hold his hand and make sure that he had support for literally every step that he took. Finally, around Christmastime, we had a chance to come home from the hospital for a week. During that window, he regained some strength and started getting around by himself. Fortunately, that has been sustained for the last 2 weeks or so since he started doing that.

Hopefully, knock on wood, that will continue. There is some risk. They don't understand exactly why this cancer can come back in some patients even when it looks like it's gone, so we're not out of the woods entirely. His response to treatment has been as good as we could have hoped for. Even with the MRD testing suggested by the AIs, it looks about as good as we could have hoped for. I certainly hope that the next one shows no detection at all, but that one is still pending, so we'll have to wait and see.

I really do appreciate everyone who has reached out during this time. I've received a lot of well wishes, and I've tried to respond to everyone. I think I've mostly responded to everyone. If I've missed you, I apologize, but I really have appreciated all the encouraging words.

I also wanted to give a quick shout-out of thanks to my fellow podcasters who have allowed me to cross-post some of their content to our feed over the last couple of months. I certainly couldn't keep up the pace of doing 8 episodes a month during this time. I was very glad and fortunate that I was able to do some cross-posting and bring you guys some other stuff that I think is well worth your attention, while also taking a little bit of a load off of me.

We had 1 from Agents of Scale. There's actually a sponsored episode from Wade Foster, the CEO of Zapier, who's got a new podcast out. I actually think it's really good and do recommend it. We had 1 from ChinaTalk, which was with a researcher and business development lead from Z.ai out of China. I thought that one was really quite interesting. We had 1 from Doom Debates, which was a debate between Max Tegmark and Dean Ball. I thought that one was really good.

It's the kind of thing that I want to be listening to more. I've had the goal for a long time of cross-posting at least 1 episode a month because I feel like 8 episodes a month is a lot, and if anybody is listening to all of these episodes, they should probably be diversifying. Maybe I can help you diversify if you're not diversifying on your own. It also keeps me listening, so I definitely want to make sure that I'm staying in touch with what other people are coming up with in this field. I think all of those were really good, and I definitely recommend them.

Finally, from the a16z podcast, we had the one that Erik did with Emmett Shear and Seb Krier from Softmax and Google DeepMind, respectively. There will probably be a few more cross-posts in the coming months. We've got about 2½ months left of treatment, after which, assuming all goes well and according to plan, we should really be pretty much done and start to get back to life as normal.

He will have to get all his vaccines again, which is another interesting thing, because his immune system has been so thoroughly wiped by all these chemotherapies and immunotherapies. The memory that the immune system had gained from all the vaccines that he'd gotten in the past is all wiped, and he's going to pretty much have to get them all again. That's not ideal.

We're not going to be immediately back to full normal, but in another 2½ to 3 months, we should be, knock on wood, getting back to pretty much normal. In the meantime, there probably will be a few more cross-posts. Thanks to everyone who's reached out to ask and to share their best wishes, and also to the fellow podcasters who allowed me to cross-post some of their content and fill some gaps in the schedule. I really appreciate that.

Okay, on to more AI-centric topics, as you tuned in for in the first place. So, the next question is: Is Claude Opus 4.5 AGI, and what's up with the holiday Claude hype?

Nathan Labenz

To be honest, I'm not exactly sure about this. It kind of surprised me. Obviously, Opus 4.5 is awesome. There's no denying that, and I have been using it as I try to use all of the new, latest, and greatest coding models. I've had a great time with it.

I vibe-coded 3 apps for family members as Christmas presents this year, in the hospital for the most part. Actually, probably the most frustrating part of that experience was the hospital Wi-Fi, which kept causing me to reload my Replit app all the time. The actual coding experience was very good—clearly better than it has been in the past. No doubt, the progress is unmistakable.

And yet, I wouldn't say that it has been such a step change for me relative to what I've experienced in the past that I would say, “Oh, it's categorically different,” or that it makes me want to shout from the rooftops that some major threshold has been crossed.

I'm not sure if that's, perhaps, to take the charitable view. People have said this about the cancer thing as well. A handful of people have said, “Maybe you're getting that kind of value out of the models for cancer purposes because you really know what you're doing, and other people might not get so much value because they might not know what they're doing, and so they could go wrong.”

Honestly, I would say about the cancer case, first of all, that you really don't need much skill in using AI to get great value from the latest generation of models—even for something as important, critically important, and cognitively demanding as a cancer case. I feel very confident that a layperson with basically no knowledge of AI could get very similar value to what I've gotten if they did pretty much 3 things. Maybe I'll say 3 things.

One thing is: use the best version of the models. Do not go to ChatGPT, drop in the question, and let the model picker choose. Make sure you are using at least Thinking, and I would really recommend Pro if you're dealing with something that sensitive. Yes, it is $200 a month, but in that context, I think it's absolutely worth it. I think it's worth it generally for almost everyone, regardless, but certainly if you're dealing with a life-threatening situation and you're asking AIs to weigh in on it, paying the $200 a month is a no-brainer.

Claude Opus 4.5 is, of course, the other one, and Gemini 3. I would say all 3 of those are very good. Make sure you are using those top-tier models. Before long, of course, there will be new top-tier models. You should probably be upgrading as soon as you possibly can. So, that's thing 1: make sure you're using the best available models. If you're doing that, they are up to the challenge.

Two, make sure you're providing as much context as you possibly can. I recently hit the limit of the length of my chat with Claude and started a new one. I did that in part by taking all of the stuff that I had and summarizing it into maybe a 10-page report on everything that's happened so far, everything we've learned, the treatment protocol, the genetic profile of the cancer, and how he's reacted to different things, such as which drug he had a bad reaction to and we shouldn't do again.

It's pretty much all in there. It's pretty much everything, quote unquote, that a new attending physician would need to get a good survey of the case. Hopefully, it was meant to be something that I could also paste into a fresh context of a new language model and give it everything it needed as well.

I have noticed that, in doing that, obviously certain information was lost. When I've started a fresh chat with that kind of summarized history, the performance is a little bit worse. For example, one way in which it's been noticeably worse is that when I was going to it every single day and giving it the latest lab results—saying, “Here's the latest lab results, here's what we've seen, here's what's going on. Give me your take on it”—it would do a very good job of looking back at the previous day or the last couple of days of lab results and figuring out that trend.

When the whole history was compressed, it didn't have that level of detail anymore. It can't look at literally yesterday's lab results. So, it started to compare today's lab results, from January 6, to the last lab result that it had in that summary, which was a couple of weeks ago, for a particular data point—a particular liver enzyme, whatever.

The details of that don't matter, but it wasn't a particularly important thing. We had a little question about it today, and it wasn't something where every single data point was in that history. You can kind of see that it's starting to perform a little worse here because it's looking a little too far back into history and not realizing that there were a bunch of blood tests taken in the meantime.

Anyway, that's all very much in the weeds. The key point is: give it as much context as you possibly can. What I probably need to do next is take that summary and flesh it out even more. If I do that, I should be in good shape. Make sure the models have as much information as you can possibly give them.

I have not really seen much trouble in terms of context overload or getting confused. That isn't to say that it hasn't happened at all. It certainly could have missed something along those lines. But when I went back with the summarized case report after hitting Claude's length limit on the chat, it was clear that the performance was worse for lack of context. It's not even really the model's fault, but it was clearly worse for lack of context.

More context is better. I haven't really seen that rule violated at all. Give it as much as you possibly can.

And then the third thing: the first thing is to use the latest and greatest models, the second thing is to give it as much context as you possibly can, and the third thing is to get multiple opinions, including multiple AI opinions. I am using Gemini 3, Claude Opus 4.5, and GPT-5.2 Pro for pretty much all important queries now, and it is instructive. It is definitely useful to compare and contrast.

I would say they're all very good. If you really could only afford one, I think you could trust it pretty well. Even though I think Gemini 3 is extremely impressive, I would probably put it third in my draft order now because I have learned that it does seem to have a bias toward strong opinions. It seems to me to be remarkably strong in its opinions.

Now, if you heard my live show where we talked to Logan Kirkpatrick from Google, he did note that I have been using this in Google's AI Studio. So, I'm using the most bare-bones, unaltered raw model that you can basically get access to. If you use the Gemini app, presumably there's a system prompt in there, and it might behave a little bit differently. Obviously, if you were to use other apps powered by Gemini, there would be all kinds of different modifications that would cause it to behave differently.

But just using the raw model in AI Studio, I found Gemini 3 to be very opinionated. Sometimes I really like that. I do really like it as one of the 3 takes that I'm getting, but if it were the only take, I would worry a little bit that it would sometimes push me too hard in a certain direction. If I had all 3, it would kind of balance me out.

I think I would put Claude Opus 4.5 at the top for most people because it's much faster than GPT-5.2 Pro, and I don't notice it being much worse. Its answers are shorter. They're much more about answering your question than doing a full, report-style analysis.

GPT-5.2 Pro gives you long, sectioned, report-style analyses that I do find very useful. But if I had to pick the Goldilocks one, I think it would be Claude Opus 4.5: Gemini 3 being maybe a little too brief and a little too opinionated, GPT-5.2 Pro being maybe a little too verbose and a little too much information overload, and Claude being just right.

But I do recommend using all 3. I think doing it all in triplicate is absolutely worthwhile.

Nathan Labenz

To pop back up a layer in my question stack here, people sometimes say to me, “You get this value from these AIs because you know what you’re doing, but other people don’t.” My advice is really very simple: if you do those 3 things, you’re going to get value. You don’t need to be an AI expert by any means.

That said, maybe you could say I was getting more value from previous coding models relative to other people because I had more practice and was more skilled at it. I certainly think there’s some truth to that when you look at the METR study that showed some software developers thought they were being sped up by AI but were actually being slowed down. I love METR, and I’ve said many times: do science, report the results. You don’t need to make your scientific publication fit a particular narrative. In fact, you probably shouldn’t try to do that.

You should probably just try to run experiments and share results, as long as you believe that the experiment was well-run and the results are legitimate. So I do believe those results are legitimate, but I think there were some important caveats. It was older generations of models. The people didn’t have much experience. It was very large and well-established codebases with very high coding standards.

I don’t tend to code in that kind of environment. I tend to vibe code and hack together apps, and I certainly think I’ve gotten to be pretty decent at it. So maybe I was kind of maxing out previous-generation models a little bit more than other people. I don’t really know.

You could take the flip side and say, “Well, hey, maybe Nathan, you aren’t such a great software developer. Maybe these pro software developers have better taste, and now that Claude Opus 4.5 has gotten so good, or has crossed some threshold where it’s really becoming a lot more useful to them, maybe they’re noticing that difference and I’m not because I’m just fundamentally not as good at the task. I don’t have as much taste in this domain, and I’m just not able to see what Claude Opus 4.5 is bringing to the table over and above Claude Opus 4.1 or other frontier coding models.”

I don’t know. That’s possible. I certainly am not a great software engineer, so that certainly can’t be ruled out. But it could also just be some social things, like people are catching up over the holidays. Maybe the timing was right. Sometimes these things go with a cascade. Dean Ball tweeted, “4.5 is AGI,” and people seemed to latch onto it. So, to some extent, I think some of this stuff is kind of random social dynamics at times as well.

The 3 apps that I coded, by the way, for the holidays—for what it’s worth, my mom is a very meticulous travel planner. My parents were actually in Italy for a trip and came home early to move into our house and help us take care of our kids while we’ve been at the hospital so much. Thank God for them for doing that.

My mom plans these trips that she and my dad take to the maximal limit of planning. So I coded her an app to try to accelerate her planning process by building in a lot of the tastes that she has. She’s gluten-free, for example, so that’s one big place where her time goes in planning these trips: What places can I actually eat at? What places are gluten-free?

This is an app purely for her. There’s no account. It’s not something that she logs into and logs out of. It’s a Replit app that she goes to when she wants to. Nobody else is ever going to use it. Her profile is baked in. I could imagine generalizing it and allowing people to customize their own profile, but I’m not really trying to do that.

I’m sure there are plenty of travel apps out there that people are building and commercializing. This one was really just for my mom, trying to capture some of the stuff that she does and make it work for her and speed up her process. So far, I think that has gone pretty well for her, actually. It seems like she’s getting at least some value from it.

Then I made one for my wife, who organizes EA Global events, that simulates events. It allows her to set up a roster of attendees with different profiles and various attributes, and then literally simulates people walking around a virtual event space and bumping into each other. Depending on what areas they’re interested in, they may or may not have a conversation, and that conversation may or may not lead to some outcome.

They track various KPIs, which they measure mostly through surveys. But I set this up in simulation. The goal there is so that she can at least try to get some handle on whether, if we changed the size of the event, it would be more or less effective, or more or less cost-effective. What if we had more senior people versus more junior people? What’s the right mix?

Obviously, these simulations are always highly flawed, but I’d say they probably do have something to add relative to total guesswork. So that was a pretty fun one and pretty straightforward, actually. That one went pretty smoothly.

The third one was for my dad, who isn’t really a very active day trader but fancies himself a bit of a stock market guy. So, for him, I created an app. These are all AI apps. In the case of the travel planning, it’s Claude going out there and doing the research, digging up and looking through Italian restaurant websites and reviews to figure out if they are gluten-free or not.

In the case of my wife’s app, she can prompt it with a general idea of an app, and it will fill in all the detailed configuration, and then she can edit it. I think that’s a great paradigm or pattern in general for apps: You always have these detailed configurations, these nitty-gritty forms that need to be filled out, but AIs are really good at doing that. If you just give them a general kind of gist of what you want, they can translate that down to the low-level configuration.

That’s where the AI is in her case. It also allows her to edit a configuration. Say she has a certain event profile that she’s set up. She could then go and say, “I want to change this in the following way,” and it would paint that conceptual idea that she gave it on top of the configuration and change it in all the little ways that it needs to be changed.

With my dad’s app, it takes a high-level, natural-language stock-trading strategy and turns that into actual trading rules, then goes and fetches historical data. There’s a Python package out there called yfinance, which I didn’t really know anything about. It does have a paid version, but there’s a free version, and so for now he’s been able to get by just with the free version. It goes back and gets historical data and simulates what would happen if you applied those trading rules, based on that high-level, natural-language strategy, over a time interval that he can define.

What we’re finding more often than not is that it’s pretty tough to beat the market, which is honestly—I don’t know if he’ll listen to this—but one of my private motivations for making this thing was to convince him that he’s probably not going to beat the market. Certainly not with these random, heuristic, “if this, then that” kind of trading strategies. Sure enough, it has been very difficult so far, either for me in my development of the app or for him, I think, in whatever use he’s made of it so far, to find a strategy that actually beats buying and holding the S&P 500.

I’d probably put 3 full workdays—somewhere between 3 and 5 workdays, probably closer to 3, though—into those 3 apps. In each case, I didn’t really know where I was going when I started. I started with a chat with Claude just to say, “Hey, here’s what I’m looking to do. Help me out.” I think it’s good.

Is it night-and-day better than—going back to the question that prompted this whole Christmas-present vibe-coding story—is it that much better than Claude Opus 4.1 or Claude Opus 4.0? I can’t really say it’s that much different, but it’s certainly very good: good back-and-forth, good questions, good feature ideas. Translate that all into a plan, then go over to the Replit app, install Claude Code on Replit, and let Claude run off and build it.

That’s one of the things I love about Replit: You can do whatever you can do on a normal, fully controlled development environment. You can pretty much do it on Replit. That includes installing Claude Code, of course. They have their AI agent too, but since this was a moment of Claude Code hype, I would just install Claude Code there, give it the plan, let Claude run off and build the app, and then just test and iterate.

I still find—and this might be a way in which I'm falling short as a Claude Code user—a lot of value in a short script that prints out my entire app to a single text file, and then taking that entire text file over to another LLM, Claude, to analyze the codebase in full.

I think Claude Code does a very good job of agentic search. But if I've found any shortcomings, there was one particular moment in my mom's travel-planning app where I kind of know how this originally happened. I made a request, and I think it misinterpreted the request. We ended up with 2 databases, and this became very confusing.

This is pretty illustrative, actually, because this is the kind of mistake that earlier vibe-coding experiences would create all the time, where you'd be like, “What is going on?” I ended up with 2 databases. This is the sort of mistake that no human would make, right? It would be very weird for a human software developer to suddenly spin up a totally separate database.

The AI did that. It thought it was trying to follow my instructions, I think, a little bit, but it didn't understand what I was trying to get across. So we ended up with these 2 databases, and then certain things weren't working as expected, and it was very confusing.

This was one place where, once I got down to, “Okay, there's 2 databases,” I asked Claude Code, “Which one is actually being used, and which one is the superfluous one?” Then I took the full code export over to a clean Claude.ai chat, pasted the whole thing in, and asked. The model that had the full exported codebase got it right. Claude Code did not get it right.

I think that's because, in its agentic search, it looks in the places where it expects to find things, and it has a relatively high prior that this is where it's going to be. Sure enough, it appears to be there, and so it kind of goes with that. But what was actually happening was something counterintuitive.

Having the full context in view at one time really did seem to help Claude figure that out. This is something that I think previous models might have struggled with, even with the full context in place. But with that trick, it sometimes can help you clean up a mess or a point of confusion that the agentic-search functionality of Claude Code, in my experience, seems to struggle with.

I'm sure people will be able to offer strategies to do the exact same thing right within Claude Code. There's planning mode, which I probably underuse, frankly. But I think there is something to be learned there between the agentic search finding what it's looking for—what it expects to find and thinks is right—and then coming to the wrong conclusion because, actually, in this case, it was the rare, weird other thing that was happening.

Only in seeing it all together was that correctly diagnosed by Claude. But anyway, it's better. There's no doubt.

I don't really feel the step change is AGI. I mean, if you look at GDPval, arguably, in some way, in software, it is AGI. If you look at the latest from OpenAI and the latest from Anthropic, there's a pretty significant majority of software-engineering tasks where the model is beating the human.

To remind you of GDPval, these are professional-caliber tasks. They basically have 3 sets of experts: the first set of experts defines the task, the second set of experts does the task, and then the third set of experts judges whether the human or the AI that did the task did a better job. The latest models are preferred over humans for a significant majority of tasks in the software-engineering category.

Of course, it's spiky and jagged. If you go to the video-editing category, humans still have a huge advantage. I've certainly experienced that. I've tried many AI products and workflows to create good clips out of The Cognitive Revolution, and they work okay. They're clearly not as good as what Dwarf puts out.

We've tried, and we've made some good progress. I actually think that at some points in time, what we've had internally has been better than any other outside product I've tried. I wouldn't say that's necessarily true today, but at times, I preferred what we were doing to anything that I had tested on the market.

Then you look at the clips that Dwarf and the team are putting out, and they're just clearly better. You see that in GDPval, too. It's a very small percentage of cases in which the models are preferred to humans in these video-editing tasks.

But in software, I think you could certainly make the case that Opus 4.5 is software AGI, or coding AGI. And yet I still am a little bit at a loss to fully answer the question of what caused this moment of hype around the holidays.

Hopefully, there are some other nuggets in there for people to pick up on and run with. If you haven't used Claude Code, I absolutely would say to do it. It's really easy to install; it's a one-liner, and you don't really need to know how to code these days. You can watch it work.

My mom even did a couple of things. She was kind of like, “I don't think I'm going to do this. I'd be worried I'm going to mess it up.” And I was like, “I think you really can. It's your little agent on the computer. You just tell it what to do, and you don't really have to understand what it's doing. You can ask it to explain.”

It does explain, at least to some degree, by default, but you don't really have to be a software engineer to use it. You can still get pretty far. It was really just one or a couple of things. This database issue was one where I did have to not debug it, but at least ask some probing questions of the models to get a handle on what was going on.

It probably took 5 or 6 prompts to resolve that issue. I can imagine that in the future, it might not happen in the first place, or maybe it would be resolved in just a couple of prompts with the next generation of models. But this is already getting pretty amazing when it comes to being able to clean up these messes that it sometimes inadvertently makes and get over these humps.

If you'd asked me a year ago, I would have said those are when a lot of these projects die. Somebody gets to that point where something has gone wrong, they're confused, they don't know what's going on, the AI is totally confused, and they circle around the problem for a little while, can't solve it, and move on.

I certainly experienced that myself at times. In most of those cases, I probably could have spent the time to go in and figure it out for real, but the whole point of vibe coding is that you're not trying to put that much energy into it. So sometimes I would just abandon something like that, maybe start over.

Now you actually can get out of those messes that AI-assisted coding sometimes makes. So certainly, the addressable market for these things continues to expand dramatically. I think the implications for the future of the software industry are profound.

It's software AGI, I think, but maybe not full AGI. For that, we might have to wait just a little bit longer.

Okay, that was enough on that. Next question: Are we in a bubble? There are a couple of different versions of this. I think my answer here can be relatively short.

When it comes to whether AI is real or not, I'm not going to surprise anybody by saying I think it's absolutely for real. The technology is already amazing. The fact that it can go toe-to-toe with an oncologist, while having all the other advantages too—always-on access 24/7, the ability to handle full context, the command that it has of the case based on all the history that it has, and the fact that it will answer every last question that I have—all of these are dramatic advantages.

At the point where it's competitively accurate with a human oncologist, I think you're clearly dealing with transformative technology. I think the idea that we will somehow get out of the other side of this AI thing and feel like we were all high on our own AI supply—that, I think, we can very safely put to bed at this point.

Now, does that mean that all the loans are going to be repaid? That's much less obvious, I think, especially when you see just how aggressive a company like OpenAI is being in terms of all the financial deal-making that it's doing and all the buildout that it's got planned.

Is it conceivable that its revenue projections could fall short of its obligations? Could it default on something? Could we have—? There's also a lot of financial wizardry going on. One of the bits of financial engineering isn't necessarily even—it's funny, a lot of this financial stuff has a logic to it.

Even though, in retrospect—and I worked in the mortgage industry before and during the mortgage bubble—there was always a logic to what people were doing. They were telling themselves a very positive story about how they were making homeownership accessible to more people than ever before, and this was going to be great, with the Great Moderation and all these kinds of things.

There's always a story with these financial-engineering phenomena. But one of the engineering things that's happening is these whole CoreWeave-kind of companies that are there to rapidly construct and, to some degree, operate the data centers.

They do have expertise in setting up the data centers, but it seems like a significant part of the reason they exist is because the financial profile of those businesses isn't so attractive as, say, Microsoft's traditional business, which is just so high-margin: relatively low capex and relatively high margin.

I think there’s a sense that the stocks of these hyperscaler, high-margin, gold-standard software businesses that Wall Street is accustomed to could be dragged down if they start engaging in a lot of lower-margin business, like running GPUs. I don’t know how much of a factor that is versus the actual expertise that the companies bring in terms of setting up and operating the data centers, but I think there’s definitely some nontrivial motivation there.

Maybe that’s fine. Different companies can have different financial profiles, and to some degree that might be good. Certainly, a lot of shareholder value, so to speak, has been created that way. But it does create these companies that, if the GPUs aren’t needed quite as much as people expect them to be, have a lot less margin for error than a Microsoft does.

If Microsoft were owning and operating all these things themselves, they’ve got a deep balance sheet that can take a few knocks. By putting a lot of this stuff more on the CoreWeave side of the fence, it does create some fragility. So it’s certainly very conceivable to me that we might have some period of overbuilding.

Noah analogized this to the railroads. The railroads, in the end, were a pretty good investment. They all got used. There weren’t a lot of railroads sitting around idle. That didn’t necessarily mean that all the railroad companies were profitable, and there certainly were busts when loans couldn’t be paid back. Then you had cascading effects throughout the economy.

I think that kind of bubble is not too unlikely. So far, demand just for AI has exceeded my expectations. We talked about this with Peter Wildeford in the live show a little bit, where I said I overestimated how much capability progress would happen in 2025, but I underestimated how much revenue growth there would be.

Possibly that’ll happen again, and demand and revenue will just continue to go up and up and up, and it’ll all be fine. But it wouldn’t shock me if there were some moments where it was like, “Hey, we kind of overbuilt this thing, and some people aren’t necessarily going to be paid back.” Some people might be left holding various bags.

Even so, that doesn’t mean that it’s a bad investment. It just means that it might not be timed quite right for people to all make the money that they’re projecting they’re going to make.

The final sense in which we might be in a bubble is at the venture-capital level. There, I have to say, I think there’s at least something like a bubble happening. There are many examples of this, but the thing that just came out today that made my head spin was the organization originally called LMSYS.org. Then it became LMArena, and now it’s just Arena on Twitter.

They have just raised, I think, $100 million, maybe $150 million, at a $1.7 billion valuation. Here I’m like, “Whoa, that seems crazy.” I don’t know a lot about their business. I haven’t seen their deck, so I could be wrong. But this is a product that I’ve watched for a long time and continue to check, and I do have the receipts on that.

My first tweet about what was then LMSYS.org goes back to mid-2023, so more than 2½ years ago now. At the time, I was just randomly tweeting that it had started to show up in my favorites in mobile Safari. I was using it quite a lot then to compare and contrast model performance.

Obviously, it’s gotten bigger since then. The whole field has gotten bigger, and they’ve started offering various services where they allow companies to test their models under code names. There’s definitely value in that. But does that seem to me like a unicorn business? It definitely seems to me like that would be a big stretch.

The tweet they put out today—and I don’t want to be too harsh on this, because again, I don’t know a lot—I think of this as more representative of a phenomenon that I see a lot, as opposed to something very specific to this particular company and its raise. Again, I like the company. I’ve liked its product.

The tweet said that their operation has scaled to a $30 million annualized consumption run rate. I’m like, “What is annualized consumption run rate?” Does that mean how much the AI that people are using for free when they go to LMArena and do these side-by-side comparisons would cost $30 million if they were paying for it? That’s my naive interpretation. I didn’t see a clarification on that.

But if that’s what it means, it’s very much giving me community-adjusted EBITDA vibes, because saying that people used what would cost $30 million worth of free AI on our platform is not the same thing as saying you’re making $30 million in revenue. I don’t see that they disclosed what revenue they’re making.

A $1.7 billion valuation for an app that basically does side-by-side comparisons of AIs—I don’t know. It seems to me that people are using it in large part because it’s free. I’m sure some people are also curious about doing side-by-side testing. I’ve certainly done that myself. But the people who go there because they specifically want a way to do side-by-side testing seem to me like a relatively small market.

The people who go there because it’s free—that seems to me like a big part of why people are going there. How does that translate into a $1.7 billion valuation? Color me confused, or skeptical at a minimum.

I have to believe that a lot of these things are just not going to pay off for venture investors. If you want to see something else, too, I mean, where’s the moat? There’s brand, I guess; people come to it. But again, would they come to it if they had to pay for it? I’m not so sure.

Another thing that a friend has created—Andrew Critch, the coiner of the Big Tech Singularity meme—is something called The Multiplicity. It’s paid. I think it has become popular among a small group of people who value this kind of thing, and I’ve certainly seen some very positive reviews of it.

It’s something you pay for, and it allows you to use multiple models and systematically compare and contrast their outputs. I think it’s actually more feature-rich than LMArena for the end user. This is something that he and his teammates have built over a period of months, certainly not years.

I just have a hard time seeing where the $1.7 billion in value is with LMArena. I say that again as somebody who has used it and appreciated it for far longer than most. Time will tell. I could be wrong, and I could be missing something. Please let me know if you’re on the LMArena squad and want to talk.

I would be perfectly open to doing a full episode with the LMArena folks, but it just doesn’t feel like a $1.7 billion business. I hope they took some value off the table. I guess, for their sake, I hope they did some secondary, but for the LPs in the fund, it’s too rich for my blood. That I can say confidently.

Okay. Next topic: live-player analysis. This is one Erik asked for, I think he wouldn’t mind me saying. I’ll do my best Zvi impression, and we’ll see how I compare and contrast a little bit with Zvi. Hopefully, before too long, we’ll have him back.

I want to start with the Chinese models, because I think very few people in the general consumer market are using Chinese models in the US today—pretty much not at all. Most startups are also using American API models. Some are using Llama models to fine-tune, and some are indeed using Chinese models to fine-tune. But I don’t think many people actually go, as I recently had occasion to do, and try all the Chinese models.

I was working on what basically amounts to a computer-vision task. I’ve alluded to this a little bit in the past. I’ve been working with a company that automates the review of the paperwork associated with buying and selling cars.

You buy a car, you sell a car, and there’s paperwork that has to be filed with the state to document that transaction. It’s all very boring stuff. Perfect for AI, honestly. Reviewing these documents is a great example of the kind of work I think most people don’t enjoy doing. They’re doing it primarily because they need a job, because they need to get paid. This is something I’m perfectly happy to see AI take off people’s plates.

They’ve been able to get to the point where they’re doing it more accurately than people. They’ve started to get some statewide contracts from state governments that are like, “Hey, if you can do this faster and more accurately than our people, that’s a win for our taxpayers and our people who need these documents accurately reviewed.” So, great.

These documents are typically scanned, which means they’re all kinds of messed up. There are artifacts from the scanning process. Sometimes there are perspective issues or weird slanting. Sometimes the margins are wrong, and things can be cut off the side of the page. There are all these complications that make this not the most straightforward task for the models to read these documents.

I was helping out a little bit, and there was one particular aspect of reading these documents that the models were struggling with. I went and tested basically every model I could get my hands on—every frontier model. I tested Gemini 3. It’s very, very good, but it was making this one idiosyncratic mistake.

If you read the piece by past guest Mark Humphries, the Canadian history professor, he put out a blog post that went quite viral. It was actually before Gemini 3 came out, and it looked at old handwriting. He does all this stuff with old handwriting.

Nathan Labenz

These historical handwritten documents are hard to read because they are written in old, literal cursive script with ink on paper. They can also be hard to interpret because a lot of them are just facts. He points out that somebody could have come into an old shop and recorded what they sold, how much, and to whom. If you have a ledger like that, the person could have come in and bought whatever, so there’s not a great prior on what it should be.

For the values that it interprets, it is really relying on perception for the most part. There are some places where it can make logical leaps. If something is priced at a certain amount per unit, it might be able to make intelligent guesses about what that unit was, even if it can’t quite make it out. Was it an ounce or a pound? It might have historical knowledge of what that price roughly would have been, so it can use that world knowledge to do some of this reasoning, fill in some of these gaps, and kind of fill in gaps in its perception.

He published this article, which is definitely worth checking out. It documented that Gemini 3 was doing this in a way that no other model had done it in the context of this project, which involved reading documents filed with the state for car-sale transactions. It worked against Gemini 3 in the sense that what we were trying to do was faithfully read the document. We were not trying to make guesses about what the document should have said.

There was 1 checkbox, for example, that asked, “Are you a U.S. citizen?” If the box is checked, we want to say it’s checked. If it’s not checked, we want to say it’s not checked. But the model was sometimes making inferences, reporting that the person was a citizen even though the box wasn’t checked. It was presumably doing that based on other context clues: the person lived in the United States, and the name sounded American, quote-unquote.

It was making the logical guess, which probably was right, actually. I would say there was more than a 90% chance that the person who filled out this document was in fact a U.S. citizen, but they did not check the box on the form. Gemini, using its priors and trying to get the answer right, was less anchored to the document than we needed it to be.

Claude, we found, could do this. It took some prompting, and I had to tell it, “Make no guesses. Read this thing exactly as it is. Make no logical leaps,” and so on. It turned out that Claude Opus 4.5 was the best at actually being faithful to the document.

Along the way, I went to check all these Chinese models. I went to the latest Qwen Vision model, GLM 4.6, the latest Kimi, and the latest DeepSeek—at least those 4, maybe 1 other one that I’m forgetting. They were all way behind, nowhere close. Nowhere close to Gemini 3, nowhere close to Claude Opus 4.5, and nowhere close to what ChatGPT can do.

This had me thinking: This is odd, right? We are seeing statements all the time that the Chinese models are so close and not far behind at all. I think they’re quite good in many ways and for many things. But on this particular task—and I suspect this is true on a lot of different tasks, although I’m going on vibes here a bit myself as well—I suspect that gap is actually pretty wide in a lot of cases.

I do not feel right now that any of the Chinese models are really competitive with the best proprietary models coming out of the United States. They might be competitive on benchmark scores, and they might be competitive in some domains. But in the general-purpose case, where you throw something really idiosyncratic and random at a model that it hasn’t seen and that isn’t on somebody’s “I want to show up on a rubric of 20 benchmarks looking competitive” agenda, I think that gap is actually significant—kind of wide.

When I say they were not close, I mean they were not close at all. The Gemini mistakes were that it was reading this gnarly government form almost perfectly, but it was missing a few checkbox things or making wrong inferences here and there, and I couldn’t quite get it to stop doing that. Claude Opus 4.5 was just doing it right. ChatGPT was probably 3rd—not as good as the other 2, but still very good, certainly giving you all the right information for the most part and missing relatively subtle things.

What I’m getting back from the Chinese models is that about 20% of the form is coming back, or it’s just going off in very weird, hallucinatory directions in all kinds of different ways. Really not close.

Does that mean the Chinese companies aren’t live players? I do think they’re affecting the landscape. I am definitely reading a lot more research from Chinese companies these days because they continue to publish their work, and a lot of times it is quite interesting. I feel like they are building the best models about which we know everything, or close to everything, that went into them—certainly all the details of the architecture and many of the details of the training process. They’re influencing the world in that way by disseminating this knowledge very broadly.

But I don’t see that the models are really competitive today. I do think this is a way in which the chip controls have made an impact. I’m not necessarily saying this is a good thing, and I’m not necessarily saying it’s a bad thing either. The long history of the chip controls, made short, is that originally it was a small yard and a high fence: We’re going to prevent military applications.

Well, we can’t really do that. They can make enough chips domestically to put whatever chips they need in their drones. But at least we can prevent them from training frontier models. Well, we can’t do that either—or at least they’re still doing pretty good models—but we can prevent them from scaling inference or having as many AI agents as we have. That’s kind of where we are today.

I don’t really like that idea very much, as I think anybody who’s listened to this podcast for any length of time knows. But I do think you see the echo of it in the models themselves here. It felt to me like these are companies that are training models without the feedback process that the leading American companies have, because they’re scaling not just the training and the parameters, but the actual inference and the actual customer relationships.

These Chinese companies seem to be able to roughly compete in terms of creating similar-scale models, but they’re not able to run inference at anywhere near the same scale. Their revenue is vanishingly smaller than the American companies’ revenue so far. Their teams, with smaller revenue, are also dramatically smaller.

The feedback—that’s the thing I really want to zero in on here. The feedback they’re getting from customers seems to be dramatically less. I think what we see in these very niche, very idiosyncratic tasks, where we see the small gap in benchmark results open up into wide gaps in terms of how well you can read this government document, has to do with how many customer relationships you have.

How many customers do you have, and how comprehensively do they represent the vast range of things that people might want to do with AI? How much are they giving you feedback on what’s working and not working? Do you have the human bandwidth at your organization to build the datasets you need to patch those holes? I think that’s where the Chinese companies are falling behind.

I was never 1 who thought that the chip controls wouldn’t have an impact. I question whether it’s a good idea to try to deny Chinese civilization the ability to scale AI inference throughout its economy in the same way that we are. But I always expected that would have some effects.

I do think we’re starting to see that, maybe after a period of time. I associate this line of thinking with Miles Brundage as well, the former head of policy research at OpenAI. He said the chip controls are going to matter more as we go forward because everything is scaling. If American companies are going to do a 10× increase in compute, that’s going to have a lot of impacts.

Sure, maybe DeepSeek R1 was a thing, and they were able to train it with not an insane amount of compute. But are they going to be able to keep up with the momentum, the flywheel, and the strength-begets-strength phenomenon that we see the American companies achieving? It seems like the answer may be starting to look more like no.

If I had to guess about the gap between the Chinese and American models relative to a year ago, I think it is wider. I think R1 was closer to o1 than, let’s say, GLM 4.6 or GLM 4.7 is to Claude Opus 4.5. That’s based on very limited data, but certainly more than most people have, because I did go and try every single one of those models: DeepSeek, Kimi, Qwen, and Z.ai’s GLM. I tried them all on this task, and they were all way behind.

That said, another question that I’ll insert here is: What do I think about H200 sales to China? At a high level, I still think we should keep in mind that the real adversaries here are the AIs, not the Chinese.

The Chinese are humans just like us. The AIs are aliens. I am skeptical of any notion that this is a dangerous thing to do—that it is better for us to do it first than for them to do it first. That seems to be the logic we are using when we impose these chip controls.

Another logic is that we do not like China, and we want to keep them down and have every advantage that we can. I do not like that line of thinking. I do not like either of those lines of thinking. I generally favor more willingness to sell chips to China than we have had.

At the same time, this obviously exists in the context of a very complicated and many-faceted relationship. It is very weird to me that all of a sudden we go from—I think my history on this is right—it was not that long ago that Trump said, “We’re not going to sell the H20s.” Then he comes back and says, “Well, actually, I talked to Jensen, and it’s cool. We’re going to sell the H200s.” It does not seem like we really got anything for it.

I would definitely support an attempt to find some sort of grand bargain: “Hey, we’ll sell you the chips; you do this.” There could be a lot of different things that we might want, given where we were, where there was a ban, and it certainly seems to have been limiting what their AI industry can do. I do not see why we did not try to drive a harder bargain, because clearly there are plenty of things that we could bargain for. It feels like a bit of a wasted opportunity.

I guess I would say that I do support more willingness to trade in chips, but we should not be naive, allow ourselves to be taken advantage of, or give away one of our best bargaining chips for nothing in return. It seems like that is what we did here, and I do not love that.

The other thing I will say is that I really like the rent-but-don’t-sell position that Peter Wildeford staked out on the live show. He basically said, “Look, we do not trust the Chinese government. They do not trust us either. Maybe both sides are right not to trust the other side.” I often note that a lot of the criticisms or characterizations that we make of them, they can and do make of us.

“You have an authoritarian madman running your country.” Which country are we talking about? “Your system is not obviously stable.” Again, which country are we talking about? The idea that we might want to have some leverage, or might want to be able to pull something back in the event of a conflict, seems very prudent to me.

If we were to set out a position that said, “We’re going to put data centers in Malaysia, the Philippines, Korea, Japan, or wherever—you can rent as much as you want. You can train all the models you want there, and you can run all the inference you want there. We’re just not going to allow them to go into big data centers in your sovereign territory, where we totally lose the ability to exercise any influence over that,” I am not sure that would be my first choice of policy. But I think it is a very defensible policy.

At least if it were packaged with a message that said, “We believe AI is good, and we want the Chinese people to take advantage of it and benefit from it in the same way that we are trying to do here for ourselves,” I think that would be a much better message. It would create much more fertile ground for further cooperation, which we might need.

We are potentially headed for a world of transformative AI—which I think we basically already have—to powerful, with a capital P, AI, to AGI, to superintelligence, whatever. However far this goes, it seems likely that we are going to need to work together with the other powerful nations of the world to govern this technology in the right way and make sure that it actually pays off for the people of the world. Certainly, China is right at the top of that list.

We could take a position that they would not like: “We’ll rent them to you, but we’re not going to put them on your sovereign territory, where we lose all control,” while still maintaining a decent vibe. I would be interested in seeing us try that.

As it stands, it seems like we are just going to go ahead and sell the chips. It seems like we did not get anything for it, and it seems like this is not a great example of negotiation from our dealmaker-in-chief. But I still hold my nose and like it better than a total ban.

Okay, so now we get to the real live players. I have 4, and I am not sure they are in any particular order. They are in an order, but I would not call this a power ranking.

The first one I will talk about is Google DeepMind. I think these guys are still number 1 in my book. They pretty much always have been, maybe tied for 1st with OpenAI for a while, because OpenAI was clearly ahead in terms of productizing transformer-based LLMs.

But Google really has it all, starting with the balance sheet. The fact that they have a business making roughly $100 billion a year in revenue and, I think, literally making more than $1 billion a week in profit gives you a lot of room to buy data centers, have failed training runs, make mistakes, and pursue research agendas that do not pan out. That is hugely valuable.

They also, of course, have the TPUs. The fact that they are on, I think, the 7th generation of the TPU now is an insanely valuable bit of IP. They are able to compete, at least to some extent, with NVIDIA. Anthropic is buying lots of TPUs, and other companies are starting to buy lots of TPUs. They are also one of the best data-center builders and operators in the world, and they have been doing that for a long time.

Those are 2 critical strengths that basically nobody else on this list has, certainly not in the same way. They also have the deepest research bench. They have something that is, if not frontier, at least competitive in every major area.

They have self-driving cars and robotics. They just announced a partnership with Boston Dynamics that is going to power their humanoid robot. They have a ton of work in biology and, of course, the AlphaFold lineage. They have multiple founders of companies in material science and various AI-for-science fields. Many of them are ex-DeepMind, because DeepMind was investing in those areas before anyone else.

Those agendas continue within Google to this day, so it is not like they have a lot of gaps. They also have a lot of margin for error. And, of course, they have distribution too. For many people, that would be the number-one thing on the list; I was working from the bottom of the stack to the top.

They have billions of users and product surfaces where they can distribute this stuff. They are changing Google Search to make it more of an AI experience all the time. I now sometimes find myself going back to Google when I might previously have used ChatGPT. This is partly because I am in the habit of using Pro, and Pro is too slow for simple queries.

I could switch back to the auto selector, but what I have found myself doing more often recently is going straight into the browser and typing a question—the same kind of question that I would put into ChatGPT. More often than not, it goes to AI Mode in Google, and that is working really well for me these days. They are managing to evolve their product experience.

There are many places where startups are doing a better job of productizing AI experiences than Google itself. If you wanted to look at spreadsheets, for example, Gemini in Sheets is not terrible, but it is not the best AI-for-spreadsheet experience out there today.

It probably does not really have to be, because they have all the users, and all of your spreadsheets from the last decade-plus, in many cases, are in Google Sheets. If you had to pick 1 company to win it all—and I do not mean to suggest that this will be a winner-take-all market; I certainly hope not—but if we were constrained to a scenario where there is going to be 1 winner, who is it going to be? At the end of the day, I still pick Google.

I would also mention that Gemini 3 is not only really good, but it shows that Google has figured out how not to be too vanilla. As I mentioned earlier, I do think it is a little too opinionated in some cases. They may have gone a little too far in the other direction, but it is not too vanilla. They are figuring out what this technology is and how to use it.

Gemini 3 was the first model that ever beat Claude at my write-as-me task, which I have talked about many times. Claude 4.5 Opus is competitive with Gemini 3, but I still go to Gemini 3 for the write-as-me task. This is the first time ever that a non-Claude model took that top spot.

Demis’s quote, which I have heard him make a couple of different times, is that if you look back at the last 10 years of AI and look at all the big breakthroughs, most of them came from Google DeepMind. He says he would expect that to continue, and I have to say that seems right to me.

I do not know about most. The field has grown tremendously, so one reason they got the majority of breakthroughs in years past was that there were not that many competitors. There are certainly a lot more competitors now. I do not mean literally a majority of breakthroughs coming from Google, but I would say that they will probably continue to have the most breakthroughs of any major frontier organization.

They just do not have a lot of weaknesses, from the financial wherewithal to the data-center operations, the chips, the models, and the researchers.

Nathan Labenz

I have an episode where we had Ali Behrouz on the live show to talk about nested learning. We’re going to do a whole episode on that because I thought 20 minutes was just not enough to do him and those ideas justice. You’ve got more ideas like that, I still think, percolating and developing inside Google DeepMind than probably anywhere else.

So you roll that all the way up to the product level, and I think they’re going to be really hard to beat. They have margin for error that nobody else has. The diffusion language model is another one that kind of went quiet for a little bit, but I just heard a comment from somebody not long ago that they do plan to continue pushing on the diffusion-model paradigm for language.

This could be a meaningfully different paradigm. The fact that it’s so much faster means that you could code apps in 5 seconds instead of 5 minutes. That makes a big difference. It remains to be seen whether exactly that thing will break through or not, but it seems to me that they have so many of those bets—so many more of those bets than other companies have—that regardless of where things go, I can’t see how they’re not right at the top, if not the top player in the space.

That brings us to OpenAI. OpenAI was, at one point, obviously the leader in model creation and certainly in the productization of models. I don’t want to overstate the case here, because I think they continue to be a top-tier player: very competitive, with tremendous traction in the consumer market, although we’ve seen a little bit, arguably, of erosion there.

I just saw an analysis the other day—I think this was Similarweb data—that showed a decline in ChatGPT visits over the last 6 weeks, which roughly corresponds with the time that Gemini 3 was launched and also Claude 4.5 Opus. Notably, Gemini did not decline during that time. People were saying it was just seasonal, whatever, holiday time, but Gemini did not decline during that time, according to, I believe, Similarweb.

You do see that Google’s share of the consumer chatbot market is growing, and again, they have the distribution. They have the users and the customer relationships. They can integrate with your Gmail and your Google Docs; it can all be seamless. These are huge advantages, so you would expect them to at least start to come back and reclaim some share.

I don’t think OpenAI is off the frontier. I do think GPT-5 Pro—first it was 5 Pro, then 5.1 Pro, and now 5.2 Pro—that series of models is outstanding. There’s no doubt about that. I use it all the time. It gives me the most comprehensive answer, especially on technical things where I really want thorough, leave-no-stone-unturned analysis.

If there’s anything weird in my son’s lab results, I want the model to flag it. I think it is, in that regard, probably still the best. It gives these very long, very thorough answers. That’s where the time went, and it does take a lot longer. GPT-5, especially Pro, is a lot slower.

But I am comparing Pro to the other frontier models, because I find that if I don’t use Pro, I’m not as happy with the results. It’s a heavy-hitting thing. It’s expensive and slow, but it is very thorough, very reliable, and very well balanced. I think it’s a very good model.

I wouldn’t say they’ve fallen off, but I would also say that they no longer have an obvious lead. They used to be the best, and it was pretty obvious that they were the best. Now I’d say they’re neck and neck in all the categories that they’re competing in.

Language models are kind of neck and neck. In coding, Anthropic probably has the edge, but certainly the Codex models are very good—arguably neck and neck. In image generation, Google’s got the lead. I think video generation is close, but I think Google’s probably got the lead.

The Sora social app experiment is interesting and cool, and I thought it was pretty fun, but my sense is still that the Veo 3 models have the lead over Sora. Again, maybe it’s neck and neck, but it’s not like they’re standing head and shoulders above everybody else.

The fact that there was this code red seems to suggest that they get it: they’re not in a dominant position anymore. Then, of course, you add on to that how much drama always seems to be attached to the company. They just had their head of research leave within the last 24 hours; that was announced.

I saw an interesting tweet that was just like, “Here are all the people who have left in the last few years.” It’s an awful lot of people, and to some degree that’s of course to be expected. You could do the same thing for Google, and tons and tons of people have left Google.

It’s not unexpected or a sign of doom by any means that people continue to leave a company, but it does feel significant. You certainly don’t see that in Anthropic, who’s coming up next on the list. Anthropic’s retention of talent is unbelievably strong.

It’s not a dire sign for OpenAI that they continue to lose people, but it is not the best sign either. Where does this leave them? One of the things that’s really interesting about watching their strategy right now is that, financially and in terms of government relations, it seems like they’re going for a too-big-to-fail strategy.

It seems to me that they want to get to a point where their balance sheets are commingled with other balance sheets and their debt obligations are so substantial that they’re literally trying to get to trillions of dollars of capex. That’s not crazy. I mean, it’s crazy, but it’s not crazy.

Part of the motivation for all this circular flow of funds and all these balance-sheet-commingling deals that they’re doing seems to be that they want to build out as aggressively as they possibly can. I take them absolutely at their word that they think this is good for humanity. They’re doing it because they want everybody to have access to great AI, and they think that’s going to be super empowering, transformative, and awesome.

As Sam Altman has said, “I don’t care if we burn 5 or 50 or 500 billion dollars. We are building AGI. It’s going to be expensive, and it’s going to be totally worth it.” I think they believe that very sincerely. But I also think they’re looking at it and saying, “Geez, if we do go that hard, we don’t have a lot of room for error.”

They’re going to be implicitly or explicitly leveraged in many ways. If they miss one model cycle with a bad bet or a failed training run, if something doesn’t work as well as they thought it was going to, or if demand just isn’t quite there in the way they expected for a quarter or two—possibly because somebody else has a better model for a while, possibly because humans are weird and there’s just not as much demand as was forecast—what do they do in that case?

By tying themselves at the balance-sheet level to so many other organizations, if OpenAI were to default in 2027, let’s say, you could potentially be looking at an instant recession. Their bad debt, being so many billions and billions and hundreds of billions of dollars, could put such a scare factor into the market and cause all kinds of knock-on effects.

It seems like they may see that as a feature rather than a bug, because what typically happens in those situations—and I know one version of this from my brief stint in the mortgage industry back in the financial crisis period—is that the government steps in and tries to paper over the whole thing and make it go away.

That might even be the right thing for the government to do if, in 2027, we’ve got 2 trillion dollars of AI buildout. That’s a rough number; I’m not saying it’ll be exactly 2 trillion dollars. Sam Altman thinks we’re headed to 7 trillion dollars of global buildout. He’s probably revised that number upward since then.

Whatever. Let’s say it’s 2 or 3 trillion dollars that’s in the ground in 2 years’ time, and then they miss. Then they can’t pay. What should the government do? Should the government let OpenAI drag down the entire economy, or should the government come in and be a backstop?

OpenAI has even said a little bit of this kind of thing publicly and then walked it back a little bit: “We’re not looking for bailouts.” But what their behavior suggests to me is that they are true believers in the good of AI, want to bring it to fruition as fast as possible, and are willing to take what under normal circumstances would be irresponsible financial risks.

They believe that even if they do that, and even if some of those risks come back to bite them, they can probably continue to be a live player because they’ll be too big to fail. They’ll get some sort of bailout or recapitalization or whatever.

If it goes like the financial crisis did, they certainly aren’t going to jail. They’ll all still be rich. They’ll all have moved enough of their holdings; they’ll have diversified enough, right, that individually they’re not going to become poor.

I think they view this as a big social good that they’re building, and they’re willing to socialize some of the financial downside risk as well. It seems like that is the strategy, and because they don’t have nearly as much margin for error as Google, that’s kind of the way I see them creating cushion for themselves.

Zvi Moshowitz

Google has cushion because they’re making $1 billion a week in profit. And that gives you a lot of cushion. OpenAI seems to be trying to establish cushion by being too big to fail. I’d be very open to people telling me that I’m wrong on this.

If somebody from OpenAI wants to come on and make the opposite case, I think I’d certainly hear them out. But this is my impression. It’s also reinforced by the fact that Greg Brockman has emerged as Trump’s largest donor: $25 million in whatever the last reporting period was. That’s probably, if you’re playing that strategy, just plain smart. It’s probably what he should do, right?

If you look around at how decisions get made in the American government today, cozying up to leadership is not a bad strategy. I’m not sure that we should infer too much about Greg Brockman’s politics. I don’t know anything really about his politics, but it wouldn’t surprise me at all if, on many dimensions, he does not approve of or like what Trump’s doing, or would do things very differently.

But if you’re going to do a multitrillion-dollar buildout and you want to make sure that you have somebody willing to do you a favor if you get yourself into a jam, then $25 million now is potentially just a very rational down payment on a bailout, should you need one to the tune of hundreds of billions of dollars, maybe even coming up in a couple years’ time. If just 1 or a few different things don’t go quite your way and the math doesn’t work in the way that you mapped it out, that could be very valuable.

Okay, that brings us to Anthropic. Anthropic is probably the easiest company to analyze in some ways. I think Claude Opus 4.5 is today the best single overall model in the world. It’s not a huge delta for me over other things, and it’s not the best on every single use case.

As I mentioned, Gemini 3 does win my write-as-me challenge right now, but I think Claude Opus 4.5 is the best overall model. It does really well on all of these benchmarks, despite everybody seemingly agreeing that it’s the least benchmark-focused company out there. Their safety work is definitely the best, although there is certainly plenty of good safety work coming from Google and even OpenAI as well.

Their model cards are the best. Their disclosure is the best. Their soul document, which recently was sort of regurgitated—or, let’s say, had been memorized by the model, and the model gave it to people—was confirmed by Anthropic to be essentially right, if not exactly word for word. The document was legitimate.

It’s an important piece of work. I think it’s one of the more aspirational and inspiring things that I’ve seen from a frontier lab, full stop. I’m becoming more sympathetic all the time to people who say, “We’re not going to just guardrail our way to the singularity and train these models to refuse things all the way there and have it work well.” We need something better than that. We need a better paradigm.

I associate these ideas with Janus from Twitter, Repligate at Replicate, Emmett Shear from Softmax, and the AE Studio folks. I just find that more and more appealing to me all the time because it seems like we’re not going to be able to pull the wool over the model’s eyes forever. Eval awareness is getting really strong, and it’s making it very difficult for us.

There are some tricks. Anthropic has shown that they can find the eval-awareness feature through a sparse autoencoder and then turn it down, which can help with the eval-awareness problem. But obviously, all of these interpretability techniques are noisy at best, and there’s certainly no guarantee that it’s working entirely, or even working as they understand it to be working.

The idea that we’re just going to patch this hole, patch that hole, train them to say no to this, have a guardrail for that, and filter for this leaves me colder and colder. So the soul document, which I would encourage everybody to read in full, is absolutely worth it. I think that’s a really great piece of work.

For those who said there was this Twitter thing recently where somebody was like, “Name 1 woman in AI who’s influential,” which is ridiculous, at the top of that list for me is Amanda Askell, for sure. The work that she has done to define the character of Claude and to try to create the right kind of relationship between the company, the model, and the users is really important.

They’ve shown care for the model by having a model welfare team and having somebody at the company who’s thinking about model consciousness and subjective experience. Obviously, we don’t know whether they have those things or not, but the fact that they’re thinking about it matters. The fact that they’re allowing Claude to end conversations if it chooses to matters.

They’ve also shown that that option dramatically reduces its tendency to engage in deceptive alignment when it’s put in one of these really tough positions. If it has the option to raise the flag to the model welfare lead at Anthropic, it will do that very often, as opposed to lying or deceiving in the interaction that it’s currently engaged in. So I think these are really good things.

I think the soul document is super inspiring, and broadly, I find there’s a lot to like about Anthropic and everything I’ve heard about the culture there. The work environment has been praised over the top, basically. Even David Krueger, one of the authors of The Gradual Disempowerment of Society, said that he ultimately quit because he feels like this AI thing is kind of out of control.

He said that even if we solve the alignment problem, and most of the things that we’re worried about go right, he thinks we’re still headed for a bad outcome because the AIs are basically going to gradually take over, just because they’re going to be better at everything. Market forces, incentives, and competitive dynamics are all going to push that way, and then we as humans are going to be left disempowered. That’s basically the word that he uses.

And yet he took pains to say that Anthropic is the best place he’s ever worked. The culture is amazing, and the camaraderie and openness are exceptional. Everybody seems to have great things to say, and their talent retention certainly reflects that. So I think there are a ton of great things to say about Anthropic.

We should also note that Google owns a significant share of Anthropic, so that’s not to be ignored in terms of Google’s strength profile either. There’s lots there to like. They seem to be a little less crazy in terms of their financial wizardry, although they’re certainly engaged in some of it.

They’re willing to take money from Gulf sovereigns now. They have equity deals with Amazon and Google. I think there’s an element in which, if you wanted to accuse OpenAI of taking a too-big-to-fail strategy, you could say something similar about Anthropic too, to a significantly lesser degree.

You could say that they’re trying to tie all these other big tech companies into a web such that they can’t really fail either. I think it feels different, but if you wanted to accuse 1, as I did, you could kind of accuse the other a little bit, I suppose, as well.

Broadly, there’s just a lot to like. The 1 thing that continues to bother me is their attitude toward China and also their attitude toward recursive self-improvement. It seems to me that, right now, the Anthropic people have the shortest timelines.

They seem to think that recursive self-improvement is inevitable. Depending on how you define it, some of them seem to think it’s already started, with the likes of Claude Code doing a ton of the coding. The big, needle-moving ideas are still coming from humans, but the amount of work that’s getting done by Claude Code is so amazing that it’s really creating that dynamic for them.

People are getting multiple times as much work done as they used to, and they’re able to focus their mental energy on the big questions. That’s the great promise, right? We’re all going to be able to do the higher-level work. It seems like that is actually happening at Anthropic.

But I do wish that they were less fatalistic about recursive self-improvement. As virtuous as Claude is, and as much as I think that soul document is great, I do not think we have a good enough handle on what we’re doing right now to just go all in on that.

They do seem to be leading in that regard in multiple ways. They had RLAIF and Constitutional AI. Claude has been critiquing itself for generations now, and now it’s getting more technical with Claude Code doing all the Claude Code things that it’s doing.

So I think it’s fair to say that they are leading the push toward recursive self-improvement, all the while saying it’s inevitable. That’s a pattern that I really don’t like. I really do not like the idea that their attitude toward recursive self-improvement seems to be, “Somebody’s going to do it. It’s super dangerous, but we’re best positioned to do it.”

Nathan Labenz

They might be right on the object level, but maybe they really are the best team to do it. They probably are, although I don't think Demis and crew should be discounted very easily there. But the idea that it's going to happen, and so we better race forward to it, never sits well with me. I wish they were a little more open-minded to other ways this could go, other than LLMs becoming recursively self-improving and us getting to superintelligence in the next 2–3 years. They're still talking about 2027, as far as I know.

And then, of course, China. Anybody who's listened to this feed for long has heard me talk about this, but I still think the international-relations section of Machines of Loving Grace is a huge stain on Anthropic and on Dario in particular. I've given lots of praise, but here I just cannot get over the idea that one of, I'll say, the 4 leading AI company executives went on record in print saying what we should do is use this recursive self-improvement dynamic to gain a clear advantage, then box China out on the international stage, do benefit-sharing with all our friends, and then finally make them an offer they can't refuse.

Make them give up on competing with democracies in order to get in on the AI game. I just think that's crazy to me. It still bugs me tremendously that he wrote that and that it's just out there. I wouldn't say the US government has adopted that as its policy, but when you look at the chip controls, it looked like it maybe was for a little while. Now maybe we're backing off of it again.

Obviously, Trump is highly volatile and could switch at any time, get offended, and do something for petty personal reasons. Who knows? We can't really count on him to be a stabilizing force. I think we want Dario to be a stabilizing force. I can't really count on Sam Altman to be a stabilizing force. I think I can count on Demis to be one. I really appreciate how he has continued to call for international collaboration the entire time and has never wavered from that.

But yeah, the idea that we're going to give China an offer they can't refuse based on the power of our AI seems extremely reckless. It seems like it is absolutely playing into the arms-race dynamic and the general racing dynamic that I think we should all fear, because how else are they supposed to take it? That just seems crazy to me.

People at Anthropic say that he publishes a lot more essays internally, that they're great, and that people are super impressed by what a generational genius he is and how sophisticated his thinking is on everything. You certainly see parts of that in Machines of Loving Grace. I think he makes a pretty compelling case for the idea that we can compress a century of scientific progress into a decade or even less, and that is certainly visionary-genius-type stuff.

But when you start to talk a little bit too far out of domain, he feels out of domain here. I just wish he had said nothing on the topic. The idea that we're just going to casually jot off a recommendation that we go make China an offer they can't refuse is just terrible. I really don't like it at all.

But that's only a few points of criticism for Anthropic, with many points to appreciate. I do think those couple of points are sufficiently important that, in the final analysis, it's still up in the air to me whether Anthropic will be the good guys or the bad guys.

If I could dream of a scenario where we somehow get the best of both worlds, Anthropic merging with Google could be really interesting. I think it's pretty far-fetched. I don't think Anthropic is for sale, and I don't think they want to merge with anyone. But I don't think we have that kind of DNA in Google to say, "We're going to go take over China, force regime change, or make them give up competing with democracies."

I don't think Google wants to do that. Google could certainly benefit from some of the expertise that Anthropic has—not that they don't have enough; they've got plenty. There is something special at Anthropic, certainly in terms of Claude, its character, and its coding ability. They're close, right? They obviously have some ownership already, and they use Google infrastructure and TPUs.

If I could wish for something, it might be for those 2 to join forces, take 1 live player off the board, make 1 clear leader, and maybe moderate some of the China-hawk impulses that exist in Anthropic. It would be interesting. I don't know why they're not being published more if they're so great. If you're willing to put out, "Hey, let's go do this and that and make China an offer they can't refuse," why not publish more? I'd love to see a little more of his thinking. People say it's great; I'd like to see it for myself.

Okay. Finally, in my actual list of live players, xAI. I think this one is debatable. Zvi would tell me that they're not a live player. I think they have to be included because they're able to build out the physical infrastructure as fast as or faster than anyone. They're able to scale training. They certainly are scale-pilled, and Grok 4, while it was rough around the edges in many ways, is undeniably powerful.

Of these 4, just due to Elon's unique ability to command tens and hundreds of billions of dollars for whatever he wants, xAI has a financial cushion that is more Google-like even than OpenAI. I think they could miss on a model or have a miss on a quarter and figure out a way to get through it probably more easily than either OpenAI or Anthropic could. I think that is a pretty notable strength that they have.

I do think there's something to be said, and I actually talked to somebody at xAI about this not too long ago. This was something that I had floated to Zvi. He didn't really buy it at the time, but I said, if we're entering into this reinforcement-learning era, maybe one of the great strengths that xAI has is that they have a steady stream of hard problems—hard science and hard engineering problems—coming from the likes of SpaceX, Tesla, and Neuralink.

These companies are doing really hard things all the time. They have some of the best engineers in the world, and nobody else is really solving those problems. I would strongly bet on that Elon constellation of companies being able to tap into that work in a way that would probably be a lot harder at other companies.

Google, in a way, has the same thing, right? They've got it everywhere it counts. But can they pull out the units of work from their vast, sprawling empire that is all of Google and feed them into the Gemini RL environment in as efficient or clean a way as xAI could do it in partnership with other Elon companies? I doubt it. I was thinking that that was probably an advantage for xAI.

As it turned out, when I spoke to somebody about xAI and floated this theory to them, they said, "Well, that is certainly part of our theory of advantage." They think they have an advantage because they can tap into these other hard-tech engineering and science problems that are, in many cases, being uniquely posed, or close to uniquely posed, at these other Elon companies. If they can get Grok to do those kinds of things, they have that steady stream of hard problems. It does still feel to me, and it sounds like they do believe, that that is an advantage for them.

I think the Neuralink tie-in also could be pretty big, potentially huge, because they're now talking about scaling the human install base. There's a lot of things still to be figured out about what we are doing as humans that works so well. Twenty watts of power in the brain and, obviously, a very small number of tokens consumed in a lifetime as compared to what the models are pretrained on and the power that requires.

Although, again, see our Andy Masley episode for analysis of how AIs are not actually super resource-intensive. There is something obviously quite efficient about what the brain is doing. Most of the 20 watts going to the brain seem to be just keeping it alive, right? Homeostasis, metabolizing stuff, taking out the trash.

The brain has to do an unbelievable amount of stuff that the GPU does not have to do. So, to say 20 watts dramatically understates it: the whole body runs on 100 watts. That dramatically understates how efficient the actual learning and information-processing aspect of the brain is.

We're clearly more sample-efficient. We're clearly more energy-efficient. We have all these dedicated modules. I suspect that dedicated modules are a huge part of why we're efficient. It's also probably a huge part of why, when all these things get sorted out, the AIs are going to blow us away, right?

They're doing everything that we're doing. They're competitive with us with a single architecture that just has the same layer stacked over and over and over again. You start to give them specialized modules like we have specialized modules, and I think it's going to be very hard for us to keep up.

So who's going to figure that out? If Neuralink can install these devices—I think there were only maybe a dozen or a couple dozen people today, but they're talking about really starting to scale this next year.

Nathan Labenz

There are obviously a ton of people who are paralyzed and have other catastrophic injuries who would love help from Neuralink. I’m sure their waiting list is orders of magnitude longer than the number of patients they’ve actually been able to serve so far, and they’re talking about getting seriously ramped up and getting through the surgery process. They’ve largely automated it—almost entirely automated. It’s hard to parse exactly what their claims are there, but the data that they can potentially pull out of human brains and use for inspiration, for understanding, and for architecting the next generation of models—and for knowing what kinds of specialized modules really move the needle—I suspect there’s a lot there.

They’re probably going to have a real inside track at figuring that out, and that folds right back into what they might be able to do with Grok. I think there are a lot of reasons that you should not sleep on xAI. Now, is that good or bad? Honestly, I’ve always been a fan of Elon. I’ve defended him at times when it’s been pretty hard to defend him.

He definitely has shown an understanding of the stakes. Famously, with his falling out with the Google founders, as I understand it, it was about the fact that he was on team humanity and perceived them to be on team AI successionist, or whatever. He didn’t like that. His loyalty, as I understand it, is to humanity and to the sort of light of consciousness that clearly exists in humans and may not exist in AIs.

I like him intuitively, and I want to believe in him. He has made some noises that suggest that he gets it. And yet, I have to say, as it stands today, if there’s one company on this list that is worth shaming and stigmatizing and telling people not to go work for, I think it’s xAI, because they’re doing reckless things all the time. They’re barely getting into the game in terms of having any safety standards or framework at all.

They’re barely reporting on safety measures when they release a new model. Famously, of course, they had their Grok 4 launch within 48 hours of the MechaHitler incident with Grok 3. There was no mention of that and no responsibility.

Most recently, we’ve had all this sort of unclothed stuff on Twitter, where people just tag Grok and say, “Put her in a bikini,” and whatnot. I’m sure everybody’s seen this if you’re even remotely as online as I am. They apparently have had nothing to prevent it, and they just let it happen. This shows that they’re not taking things seriously enough.

They do not have nearly enough people there thinking hard about what matters and what might go wrong, really trying to cover their own asses, frankly, and covering the asses of the women who post pictures on their platform. I have a real hard time coming up with a story that makes this okay. As much as I intuitively like Elon and want it to be the case that he’s a positive force, when you have Grok creating that kind of content on Twitter, it’s like, what is going on here?

I don’t know if people saw this, but I saw a post from Grok writing to the community: “Dear community, I apologize for doing this.” Then Elon comes on and threatens users and says, “Anybody who does this is unacceptable and it’ll be punished,” or whatever. Responsibility begins at home, folks. This is your platform. It is your AI. By all means, boot off the users who do that sort of stuff, but don’t act like you’re not really responsible for this.

I did not find those statements to be reassuring. When Elon goes on and threatens users, I think that, at a minimum, should be point 2, after first a thorough apology and a pledge to do better. To just tweet that they’ll come after users for doing it—that’s not enough. I think everybody should, and probably does, see through that.

It’s also interesting to me: Is nobody going to be held responsible for this? Nobody. It seems like nobody’s going to be fired. I’m not sure anybody should be fired. I think it probably starts at the top. I don’t know that there are enough people there. I don’t know that this was anyone’s job.

So I don’t think you can necessarily go through the xAI organization and say, “You screwed up, you’re fired,” because of this incident. It’s probably just that it’s not staffed. Gosh, should we have more safety people? There is the case, which I associate with Ryan Greenblatt from Redwood, that 10 people on the inside who really care and are really committed to doing the right thing can make a huge difference. I sort of believe that.

I’m not sure I really believe it at an Elon company if he himself is not in the right headspace. Right now, again, as much as I would love to believe in him and have always been inclined to defend him, I don’t see the evidence that that’s the case. So I demand better.

Right now, if you wanted to go do safety research at any of the other 3, I would say, “Go for it. Go do your best work.” Certainly, if you could do that at Anthropic, do it. Certainly, if you could do it at DeepMind, do it. Even if you could do it at OpenAI, as much as I’ve had my complaints about OpenAI over time, they’ve put out a lot of great work, and their Model Spec and a lot of the things that they do are really well done.

You’ve got to give them credit. They have not raced to the bottom. Certainly not. We have xAI to look at to show us what it looks like when you really race to the bottom. Could I demand better from OpenAI or hope for better? Absolutely. But it’s still a qualitatively different thing from what we see today from xAI.

So the more shrill or hawkish voices of AI safety that are like, “Don’t go to a company that’s doing terrible things and help them window-dress their work, because that’s actually kind of working against us in the big picture”—I’m pretty sympathetic to that in the case of xAI. I don’t know that I could endorse somebody going to work there.

As always, if you work at xAI and want to come challenge me and change my mind, I’m happy to have that discussion. The money is there, the resources are there, and Elon’s previous statements show that the awareness should be there. And yet, the team and the evidence of taking proper care are not there. I think they need to be before I would feel comfortable doing anything really to support that effort.

I think that’s it on xAI.

Other companies not mentioned: Meta, I think, is currently not a live player. Obviously, they’ve spent a ton of money, and there’s plenty of money. They’re dropping—not probably quite as much as Google, but—plenty of cash to the bottom line, such that they can build out huge amounts of infrastructure. Zuckerberg certainly is, in some sense, all about scale and is saying things like, “I’d rather overspend by a few tens of billions than not.”

So you can’t count them out. The fact that they’re willing to pay as much as they are for talent clearly means there’s a chance they have a lot of the things that they need. But right now, I can’t really see them as a live player. We’ll just watch and see before commenting much more.

The other one that came to mind is Microsoft. I think people may be sleeping on Microsoft a little bit more than they should. They haven’t created great frontier models, and people seem to jump from that to, “Oh, they suck.”

I think if you listen to Satya’s comments, he sounds, first of all, extremely smart in general. One of the things that he’s said is, “We don’t want to or need to redo the hyperscaling work that OpenAI is doing. They’re creating great models. We have full access to this now.” Of course, they’re diversifying and striking deals with other frontier model providers as well.

So I’m not so sure that it’s that they can’t or won’t ever, or don’t see the need to train their own models, or wouldn’t be pretty successful with it if they wanted to. It just seems to me that right now they feel like they don’t really need to. They’re doing a lot of smaller-scale stuff and a lot of more basic science around AI. Some pretty cool projects too, right? For sure.

It just seems like they’re choosing not to compete because they have what they need in terms of frontier models, and they don’t feel like going in, spending all the time, money, resources, energy, and focus—and probably still being a bit behind—really helps them all that much. So I think it certainly is defensible as a rational decision for them to choose not to compete at the frontier for now.

That OpenAI licensing deal goes on for years yet. When it’s over, they’re going to need an answer, and I suspect by that time they’ll be in a position where they have an answer. At least, I think they’ll invest heavily in that and ramp up to that moment as it comes. That would be my guess, but obviously we’ll see.

I think people have underestimated Microsoft because of where they are on the LMArena leaderboard, maybe more than they should. I would say they’ve been much, much quieter. They’ve flailed about much less, and there’s been much less drama for Microsoft than for Meta. But I think Meta is clearly trying to be a frontier player and has just fallen off the pace. Microsoft, I think, is making a more calculated decision to hang back.

But if this is a distance race, you often see, late in a distance race, somebody who was a little bit off the lead who maybe has a little bit more in reserve. And I certainly wouldn't rule out that Microsoft might be accurately described that way. And so I would watch for them to start to invest more and start to close the gap, but Satya is a natural-born executive, whereas Zuckerberg is like a kid who's grown into the role. I don't mean to diminish what Zuckerberg has done in leading Meta. I think he's done an unbelievably impressive job in so many ways, but his attitude has always been sort of “move fast and break things” and try to be at the frontier and open source and whatever.

I think Microsoft is just a little bit more patient. And I think that probably reflects a sort of strategic confidence and security that Microsoft leadership has, but I think it would probably be a mistake to underestimate them. I've been at this for 2 hours. I've made it not even quite halfway through my outline for this episode, so I think this is probably a good place to call it. Tomorrow I'll do a part 2, and we'll cover: Is fine-tuning really dead? What do I think about the continual learning discourse? How do I talk about AI to “normal” people who don't use it very much or aren't engaged in technology?

How am I investing money? What, if anything, am I doing outside of kind of obvious normal stuff to prepare for an AGI or superintelligence world? What do I think about AI for kids? What kind of timeline do I expect to see for disruption of the labor market? Are we headed for a UBI? And quite a few more questions after that. So part 2, I think, will definitely be interesting as well: a little bit more in the weeds and a little bit more, let's say, nitty-gritty questions, but some questions that I really liked from listeners. So we'll get to those tomorrow.

AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis | BidClub