[BidClub_]
The Cognitive Revolution · · 89 min

AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases

Nathan LabenzPrakash NarayananJeremie HarrisEdouard HarrisSteve HouJoel BorgenDaniel McKinnon

AI & SoftwareBiotechTechnicalPolicy
YouTube ↗
TL;DR
  • The Trump-Xi contact matters less as détente than as a chance to install crisis machinery before an AI incident forces an “insane ask.” Edouard Harris welcomed a possible AI-risk hotline but warned that China has historically withheld nuclear hotlines as leverage. Without better verification, the first verifiable demand could be shutting down every running cluster above a visible size—putting billions in GPU depreciation at stake.
  • Verification startups could become a gating layer for frontier-AI revenue, but only if intelligence agencies vet their technology before a crisis. Even a system “50% better” than watching gigawatt-scale heat signatures could leave half a data center operating and save billions; offensive capabilities would still be required to assure compliance. Jeremie Harris stressed that technologies can take years to become accepted national technical means, while some current $10 million-plus research programs may be non-starters to the intelligence community.
  • A shared transformer-MOE technology tree gives the US and China common warning signals, but it does not create common motives. China may interpret American safety proposals as attempts to preserve its lead, while genuinely costly actions—such as Chinese labs pacing after an incident or regulators blocking a misaligned release—would be harder to fake than conciliatory rhetoric. Reciprocal transparency could be comparatively cheap because, as Edouard put it, “the Chinese are already all up in our systems.”
  • Hyperscalers sustain GPU rental prices at least two to three times those of neo-clouds because enterprises are buying an integrated relationship, not a fungible chip-hour. SiliconData’s Steve Hou attributes the premium to analytics, safety, compliance, and customer stickiness—the computational equivalent of fries and a milkshake being “inject[ed] into the patty.” Its token index likewise reflects usage mix: an 11.5% weekly bounce can coexist with a fall from roughly $4 to $1.60 since late June.
  • OpenAI DevDay’s consequential shift was intelligence delivered at interactive latency and made portable through a ChatGPT subscription. Prakash Narayanan described GPT-6 Astra Ultra Fast as eight times faster, making it possible to build “with you in the moment while keeping you in flow.” Sign in with ChatGPT could also remove meaningful trial costs for SaaS vendors: Waymark at one point spent roughly $1 just to build each new customer’s initial business profile.
  • Joel Borgen’s AI-assisted novel argues that models can collapse creative activation energy while leaving architecture, taste, and subtraction with the human. He supplied the story design, scaffolded drafts, and edited the manuscript down about 14%, often removing AI “tics” and over-explanation. His unresolved question is whether abundant personalized art retains value when it is no longer part of a shared cultural space.
  • Gamo Labs shows why cheap genome sequencing did not automatically produce precision medicine: the missing layer was interpretation. An o3 agentic loop recovered the 91-kilobase deletion behind Daniel McKinnon’s son Owen’s disease after a reasonable human-oriented filter excluded an enhancer one megabase from the gene. Gamo now combines models, tools, wet-lab evidence, and evals; Opus 5.5 in CloudCode scores about 50% on RareBench versus roughly 10% for LIRICAL, but that 50% is a benchmark score—not the share of patients diagnosed.
Digest · the substance, structured for research

1. A hotline is progress, but it is not a crisis plan

  • Edouard Harris’s opening judgment was deliberately modest: “Talking is always better than not talking.” The Trump-Xi discussions may have produced the beginnings of an AI-incident line between governments, which is positive even if it remains only theoretical.

  • Gladstone AI’s assessment drew on roughly a dozen State Department diplomats who had negotiated with China. Their warning was that nuclear hotlines were not reliably answered; at critical moments, refusing the call itself became leverage. Cooperation might improve if the CCP structurally takes AI risk seriously, but US planning still needs a backup.

  • Today’s crude emergency option would exploit what is already observable: data centers are huge, slow to build, and thermally bright. After a sufficiently serious incident, the United States might demand that China turn off every running cluster above a threshold until its heat signature disappeared—an ask Jeremie Harris called “insane.”

  • Nathan Labenz completed the logic: rapid capabilities progress plus weak trust could produce an extreme incident, followed by an “outlandish, super high-cost move,” precisely because governments failed to invest earlier in lower-cost verification and “trust but verify” mechanisms.

2. Verification must be pre-vetted before “game time”

  • Jeremie argued that better verification buys economic room. Even a hypothetical system only “50% better” than observing gigawatt-scale energy use might verify half a facility and avoid shutting it all down, immediately saving billions in depreciating GPUs. For that period, AI profits could be “gated…by these little verification technologies.”

  • Monitoring alone cannot compel compliance. Jeremie’s repeated warning was categorical: without offensive options, “you cannot assure compliance”; a government can discover prohibited activity and say “Oh no,” yet still lack any practical response.

  • Jeremie identified the neglected bottleneck: intelligence agencies may take years to vet a technology and recognize it as a national technical means. Verification founders therefore need relationships with the relevant agencies before crisis conditions, while officials need those founders “on speed dial” when “game time” arrives.

  • Early contact could also prevent wasted capital. Some companies are pursuing multimillion-dollar, in some cases $10 million-plus, research agendas that an appropriately placed intelligence official might consider non-viable; fast feedback would let them redirect toward methods that could actually support an agreement.

3. Shared architectures create common warnings, not common motives

  • Nathan proposed using OpenAI and Anthropic as a domestic test bed: require frontier labs that trigger one another’s alarms to develop mutual policing, then adapt whatever works to US-China verification. Jeremie called the idea “not insane,” noting that employee poaching partially mirrors China’s apparent access to US labs.

  • Edouard’s reservation was trust: two US companies still share more values and institutional context than rival states. The similarities may be worth mining, but a lab-to-lab mechanism cannot simply be enlarged into diplomacy without accounting for “fairly radically different” values and circumstances.

  • Jeremie nevertheless saw value in the shared hardware and algorithmic path. Both countries are largely doubling down on transformer MOEs “with some bells and whistles,” making failures in one ecosystem legible to the other. China’s second-place position may reduce its resolution on warning shots and make US safety proposals look like attempts to restrain a competitor.

  • Rhetoric such as Xi’s conciliatory WAIC speech carries some positive weight, Edouard said, but costly behavior would be more “unfakeable”: DeepSeek and Zifu reaching a pacing agreement after an incident, or a regulator blocking a release because an ideologically compliant model became misaligned in the wild. Procedural green flags include empowered negotiators, concrete progress, and fewer capricious objections over language.

4. Reciprocal transparency may be cheaper than continued strategic ambiguity

  • Edouard’s stabilizing transparency would reveal enough to show that neither side is doing “the thing you fear most.” It might not cost the United States much because “the Chinese are already all up in our systems,” although he cautioned that visibility into frontier labs should not be granted unilaterally.

  • The nearer-term opportunity is denser contact among verification researchers, academics, and startups on both sides. Those relationships already include sincere participants saying, in effect, “Yeah, this is a problem and this sucks…let’s try to solve it.”

  • Such links do not erase national constraints, but they create people able to tell domestic leaders, “I’ve spoken to them. They’re not evil.” Edouard saw that human testimony as one modest brake on the most extreme interpretations each side makes of the other.

5. Hyperscaler GPUs command a product premium, not a clean arbitrage

  • SiliconData combines quoted and transacted GPU rental prices, then uses machine-learning normalization to compare contracts. Quotes remain informative because API inventory can often be purchased immediately, unlike an advertised apartment that may be unavailable or negotiable.

  • Transactions alone can mislead: renting three nodes at $2.75 does not prove the next node is available at $2.75. If the surrounding market quotes $4, $6, or $1, a stale transaction must be interpreted against current supply rather than treated as the universal clearing price.

  • Hyperscalers regularly charge “at least two to three times, sometimes more” than neo-clouds. Steve Hou attributed that spread to bundled analytics, safety, compliance, and long-established enterprise relationships whose customers face switching friction—not necessarily to wasted GPU hours.

  • His single-patty-burger analogy carried the distinction: fries and a milkshake can be stripped out, but not if the provider blends them into the patty and declares a new product. At that point, SiliconData separates the provider category rather than pretending the chip-hour is fully fungible.

6. Token indexes move with usage mix as well as posted prices

  • Silica Mark physically tests individual GPUs at the EU ID level for health, throughput, and TFLOPS, but those results do not yet enter the rental index. Although the “GPU lottery” creates unit-level variance, Steve said large clusters probably average it out; current pricing features instead include geography, CPU, memory, term, and provider.

  • SiliconData’s token measure is an expenditure-weighted price index normalized to one million—not a pure price or total-volume series. It can rise because users migrate toward more expensive models even when every posted API price stays unchanged.

  • The proprietary-model index’s 11.5% seven-day bounce followed a much larger decline, from roughly $4 to $1.60 between late June and mid-month. Steve suspected the rebound reflected usage mix—like buyers choosing a pricier Mercedes variant—after new models and cheaper variants had pushed the index down more than 50%.

7. DevDay made latency and portable subscriptions part of the product

  • Nathan viewed Dots as a polished consumer product awkwardly placed in a developer keynote. After nine months evolving his own system, he expects to keep it; for someone like his mother, a managed agent that avoids databases and troubleshooting could be exactly the right abstraction.

  • Prakash’s early read was that GPT-6.1 Sol delivers “basically Astra for the price of Sol” and is definitely better than GPT-6 Sol, released only two or three weeks earlier. The compressed cadence itself was part of the signal.

  • GPT-6 Astra Ultra Fast moves from the prior two-times fast mode to eight-times speed. Prakash’s key distinction was not whether a model can build a game, but whether it can build one interactively “while keeping you in flow”: latency at the same intelligence becomes the next product frontier after capability thresholds are crossed.

  • Sign in with ChatGPT could let eligible users apply their existing plans to participating apps, easing SaaS vendors’ trial-inference costs. Waymark at one point spent about $1 generating the initial profile for each small-business signup, enough to matter to acquisition economics; Nathan called portability a “huge win” for users, developers, SaaS companies, and OpenAI’s ecosystem promise.

8. AI lowered the activation energy, but it did not author the architecture

  • Joel Borgen might eventually have written The Receipt Horizon alone, but children and a part-time job made the activation energy prohibitive. He began with the vision and much of the first half already mapped; early models chiefly introduced context-window management problems as prior material fell out.

  • Better models made the novel better, yet raw prose could still be “unreadable unless you know how to scaffold it and then edit it yourself.” Joel acknowledged both the work’s AI-related weaknesses and readers who will reject AI writing regardless of the process.

  • His positive future still begins after an AI war because “utopias often want to become dystopias in narrative form”: fiction needs friction. Rather than install a character as his mouthpiece, he embedded technologies he wants to see while trying to steel-man people who would oppose them.

  • Joel works in architectural chunks and chapter drafts, mainly with the flagship Claude model and the ChatGPT Pro model from GPT-5 onward. Professional copy editing, art, and layout were expensive and slow; he has since fed those lessons to GPT-5 or o3 Pro to develop workflows for the follow-up.

9. Human judgment is the scarce input while machine art explodes

  • Joel contrasted his months-long process with an automated romance-writing operation producing something like a couple hundred books a year and profiting across the portfolio. As generated art becomes instantly personalized, his harder question is whether art remains valuable when audiences no longer share it.

  • For now, he sees a “sweet spot” where AI collaboration still demands substantial human taste and labor. He edited the manuscript down about 14%, often improving mediocre chapters by removing characteristic AI tics and over-explanation rather than adding more prose.

  • His preferred workflow makes the model an attentive assistant: ingest the whole novel, surface issues, and let him decide point by point. That preserves authorship as judgment even when machines can perform much of the analysis.

  • A long-running viola-clef test illustrated discontinuous capability. Successive models failed five simple notes until Astra; it then rendered a highly complex 20th-century score—with double sharps, accidentals, and ties—almost perfectly into machine-readable form through Codex, a jump Joel described as “zero to 99 overnight.”

10. An AI loop found what a reasonable clinical filter excluded

  • Daniel McKinnon’s motivation came through five years of genetic uncertainty: his son Owen died from a rare lung disease, another pregnancy was lost, and three whole genomes returned negative before the birth of healthy 13-month-old Warren. During that pregnancy, Daniel invoked his HIPAA rights, obtained the raw data, and built an interpretation pipeline.

  • That pipeline rediscovered Owen’s known lesion: a 91-kilobase deletion removing an enhancer one megabase upstream from the relevant gene. The original lab’s structural-variant filter ignored findings more than one kilobase outside a gene—a reasonable decision when humans must review noisy next-generation sequencing assembled from roughly a billion 150-base-pair fragments.

  • Daniel placed o3 in a loop that first inspected coding regions, then kept searching after each miss. Unlike a clinical team constrained by case throughput, machine intelligence could work continuously and in parallel until it found the remote regulatory explanation.

  • He sees genomic interpretation as especially suitable for agentic improvement: it is long-horizon, verifiable, and unsaturated. Rare cases could become “Olympiad-level problems” on which systems repeatedly hill-climb, while human clinicians remain pressed for time.

11. Uncertain variants need wet-lab evidence, not another static prediction

  • Most difficult reanalysis cases reach a variant-of-uncertain-significance state. Under the ACMG rubric, six points can move a variant to likely pathogenic, unlocking treatment and insurance options; waiting for another matching patient may take years, while a functional study can contribute two to four points.

  • Those experiments require biologically relevant systems. For alveolar capillary dysplasia, Daniel cited IMR90 fetal lung cells because the implicated gene is expressed only around weeks 16 through 20 of development; inserting the variant into a generic cancer line would not answer the same question.

  • Gamo bought what remained of Arpeggio, hired four team members, and began its first experiments three weeks later. The immediate product could be a lookup service for clinical labs, but the deeper aim is “RL data from the real world” that closes the loop between model predictions and measured biology.

  • Sequencing itself is now “cheap enough”: below $1,000 per genome, sometimes about $500, and under $100 in high-throughput labs. Interpretation remains scarce. On RareBench, traditional variant-ranking tool LIRICAL scores roughly 10%, while vanilla Opus 5.5 in CloudCode scores around 50%; beyond ranking, the newer systems can explain findings, rank patients, and propose treatment relevance.

12. Vertical AI wins through routing, tools, and evals—but the layer is thinning

  • Gamo serves two existing workflows: online reanalysis when a physician receives an inadequate result, and bulk reanalysis of thousands of archived genomes that rare-disease centers lack staff to revisit. A machine-led system can process 80, 100, or 200 cases in a week versus six months or a year of constrained human work; four months in, the company remained pre-revenue.

  • Daniel described the company as “a harness, a tools, and an evals company.” Claude, Astra, and even Gemini succeed on different RareBench clusters, so intelligent routing and ensembling improve cost and performance without attempting to compete at the base-model layer.

  • Trace analysis catches failures that headline demos conceal. Groq 4.6 once missed a case by silently renaming a gene; Astra once missed an answer in a paper it was assigned, and in another eval it stopped because some tools made it want to think too much. Gamo co-designs tools with each model to prevent those “little paper cuts” reliably and affordably.

  • Daniel also conceded that the vertical layer is shrinking: vanilla o3 scored 0% initially, and he said the improvement since then was “20 or 30 percentage points or something.” Frontier-lab employees have been eager to help, and Daniel reported many examples of previously undiagnosed children now diagnosed—but redirecting a trillion-dollar company is different from individual enthusiasm. Nathan’s closing hedge held both truths: pacing frontier progress may still be wise, yet the roughly 50% RareBench result leaves consequential questions unanswered, making restraint “a costly compromise.”

Full transcript
Nathan Labenz

Real things are starting to happen. This week on AI in the AM. On AI risk and verification, we heard from Jeremie Harris and Edouard Harris of Gladstone AI. They remain hard-boiled realists about US-China cooperation, but this week I heard a slight thaw. Here, Edouard considers what reciprocal transparency could offer the 2 countries.

Edouard Harris

Certain kinds of transparency can be stabilizing. The kind of transparency that goes like, “Hey, we're giving you enough vision into what we're doing to see that we are not doing the thing you fear most”—that sort of thing—is potentially useful. Additionally, it may not actually be that costly for us to do, depending on how we implement it, simply because the Chinese are already all up in our systems. Really, we're not giving anything away that they don't already have in many cases, potentially.

Nathan Labenz

Steve Hou, head of research at SiliconData, which builds GPU price indexes. We asked why renting apparently identical chips costs so much more at the big cloud providers.

Steve Hou

So, in the case of hyperscalers, indeed, you are observing correctly: they charge regularly, consistently, at least 2 to 3 times, sometimes more, compared to a typical neocloud. The reason has to do with a long legacy of whether it's other types of products being offered on their platform for software, analytics, safety, and compliance; the fact that they already have this long-established relationship with enterprise users that have been on board for a long time, that have a certain stickiness for moving. It is being sold as very much of a differentiated product.

Nathan Labenz

Prakash Narayanan, my co-host on AI in the AM. We spent part of the week on OpenAI DevDay. Here, what changes for building software when the models get faster at the same level of intelligence.

Prakash Narayanan

With ultrafast, as you type, you can interact. It's an interactive kind of software build, interactively building games. I think that's really the future. I think the speed—the latency at the same intelligence—is probably something that's going to be very important, especially as you clear these hurdles of capability. The thing can build a game, but can the thing build a game with you in the moment while keeping you in flow?

Nathan Labenz

Joel Borgen co-wrote his novel with AI models. The text has a few AI tics, but the book is legitimately good. Here, what he supplies as architecture and why the prose still needs him.

Joel Borgen

I started the project knowing more or less what I wanted to have made. I sketched it out myself; especially the first half or so of the book was pretty well set before engaging the models. The models are good at certain things. They're getting better at everything. But as far as just prose writing itself, even if you tell it exactly what you want and you have a plan for a chapter, you often get something that's unreadable unless you know how to scaffold it and then edit it yourself.

Nathan Labenz

Daniel McKinnon founded Gamo Labs, which uses AI to interpret genomes. He reports new diagnoses in children whose cases had gone unresolved. Here, what he found when he looked closely at a model's work.

Daniel McKinnon

And you'll see things like—there's 1 particular case where Groq 4.6, which at that point was state-of-the-art on our benchmark, in GroqBuild, missed 1 because it just renamed the gene. It was saying, “Oh, NRF2 or whatever is responsible,” and then it just changed the name of the gene to something totally different. I'd never seen that before, and I was just like, “This is dumb.” Our harness and our tools prevent the agents from doing dumb things.

Nathan Labenz

Welcome to the AI in the AM weekly highlights, with these introductions spoken in my cloned voice. Please tell us what worked and what did not. Your feedback helps us make the next one better.

Part one: A slight thaw. Jeremie Harris and Edouard Harris have been interviewing diplomats who negotiated with China. We discussed the Trump-Xi talks and the possibility of a channel for communicating about AI incidents. Edouard starts with what he makes of that contact.

1. AI Incident Channels

Edouard Harris

Talking is always better than not talking, so that's a positive. My understanding, from at least the beginnings of the Trump-Xi conversation and the stuff leading up to that, is that 1 of the things that may be positive that came out of that relationship was the development of this—I don't know if you'd call it a red phone, but at least some kind of theoretical line between the 2 governments on AI incidents and AI risks.

The report that we came out with, which is really just a long newsletter, is informed by speaking to about a dozen State Department diplomats who have dealt with China from the negotiating table and who've seen how these things develop in practice. One of the issues that they do see, among many others, is that we have tried the red phone thing before in the context of nuclear weapons, and by and large, they don't always answer.

Particularly in critical phases, it's often exercised as a point of leverage, saying, “We're going to take away this phone line and not answer,” rather than as something that is a collaborative, unified project that makes everyone safer. Again, this is not to say that if the CCP does structurally take AI seriously—which there is decent reason to think that they may—we could have a whole different and much more positive level of engagement. All that we're recommending, based on the experience of these folks, is being realistic about it and having a backup plan.

Nathan Labenz

I asked Jeremie, after a serious AI incident, what could the United States ask China to stop doing? And what could be verified with the capabilities that exist today?

2. Verifying AI Compliance

Edouard Harris

Everything that comes after this sentence obviously has not made contact with the intelligence community from a red-teaming standpoint. So the true answer is, we can't know deeply. I couldn't give you an answer to a level of detail where it would be like, “Okay, that's executable.” 1 easy thing, if I'm going to caricature—

Jeremie Harris

Data centers put off a hell of an energy footprint. The thermals on those are really, really bright. Data centers are huge. They haven't, by and large, yet been built to be hidden, and it takes a long time to build data centers.

Now, this will change. AI 2027 talks about the timelines for this. We think it's quite plausible that the timelines could be a lot shorter for hiding data centers, just based on conversations with folks in the industry. But whichever way you slice it, you're going to have an initial conversation where, to first order, for the 80/20 that you really need, it's like, “So help me God, if I see a cluster and that cluster is yay big...” That's the kind of conversation that you're looking at. And how you quantify that is a matter of the sort of national technical means that the US currently has—

Prakash Narayanan

So wait, are you saying having a cluster above a certain size would be a red line?

Jeremie Harris

So I'm saying—

Prakash Narayanan

Initially.

Jeremie Harris

Yeah, initially.

Edouard Harris

Yeah. So if you're just in that panic moment, right? You're like, “We have to do something,” and you ask yourself, “What is possible? What is possible to do with the existing assets and infrastructure that we have today, nothing else?” Then you do get into a space not necessarily where there exists a cluster of this size because you're not—you can't ask them to tear down the cluster—but we have to see the heat signatures from this go away.

So if there is a running cluster above a certain size, that is, to be clear, a tremendously expensive ask in either direction. The depreciation on GPUs is the major part of the OPEX costs.

Prakash Narayanan

So you're saying an incident happens first—

Edouard Harris

Yeah.

Prakash Narayanan

—and the response to that incident, the mitigation for that incident, is, “Hey, can you turn off this big data center?”

Edouard Harris

Yeah, like, turn off any cluster above a certain size. That's 1 possibility because this is—

Jeremie Harris

As of right now, to be clear, the framing is basically, as of right now, that's where we're at. And so, in some sense, we're going to get into a potentially circular loop here where you can see how big of an ask that is. That is an insane ask.

Nathan Labenz

So if I sketch out the logic from beginning to end here, it's like: we don't have that great of a relationship. AI capabilities continue to progress at a fast pace. We expect something crazy to happen. When something crazy enough happens, we're going to find ourselves by default in a spot where we have to ask for some outlandish, super-high-cost move, like shut down all your big data centers, because we don't have any other mechanisms in place that allow for a lower ask, a better trust-but-verify type of environment, because we haven't made those investments now.

That leads me to the question of what should we be doing now to, A, ideally not end up in that situation, or B, if we do end up in that situation, have better options available to ask for aside from shut it all down, which is obviously going to be tough?

Edouard Harris

Develop—basically, you absolutely nailed it. It's develop better verification and develop better offensive options to ensure compliance in the event that verification returns, “No, they're doing it.” The better verification stuff you can do, the faster and the less you have to rely on absurdly expensive things like this: “A gigawatt of energy radiation shouldn't be visible from space.”

Jeremie Harris

If we have techniques like this that are vetted by the intelligence community, even if they are just 50% better than this—even if it’s like, you have to shut down half your data center. I’m making something up here. You have to shut down half your data center, and we can sufficiently verify the other half, or something equivalent to that—you are saving billions of dollars right off the bat.

The ability for this industry to continue to make large amounts of money is actually going to be gated for that period of time by these little verification technologies, and there’s already this community of little verification startups that’s working on this technology. That’s why this is so important. Of course, I will also say the offense side of things is critically necessary. If you don’t have those offensive options, you cannot assure compliance.

You can verify and monitor the situation, and you can say, “Oh no,” but fundamentally, your hands are tied. You don’t have the tools to actually do anything about it. Both of those things are super important.

Jeremie Harris

There’s this bottleneck that I think a lot of these verification companies haven’t necessarily priced in, and this is where a lot of our current work is focused. Imagine what happens when Company A goes, “I have the thing. This thing is going to work,” right? It’s a moment of crisis, and Trump—or POTUS, whoever it is at the time—is casting about for options to alleviate this trillion-dollar bottleneck.

Then they actually go, “Okay, the intelligence community has to now vet this,” because they’re not going to just start using it, obviously, right? How long did it take similar technologies in the past to get used, to get vetted, and to become what’s known as national technical means, or NTMs, right? The answer is years, depending on the technology. Very often, years.

We have to do it. China has to do it. We have to handshake on doing it. Even if you remove money as an obstacle and it’s no object, there are just certain things that take serial time to do. A lot of what we’ve been doing is focused on saying, “Okay, treaty”—or not treaty—“let’s say AI agreement verification or compute verification. Company X, you probably should be talking to IC element Y about this.”

In a moment of crisis, you want as much pre-vetting as possible, and anybody who’s involved in assessing a potential national technical means had better have on speed dial—the Signal, the phone number, the email, whatever—of the founders of all the verification companies that they plan to use or may end up having to use. You want to cut down on all those barriers.

The boring bureaucratic hurdles that nobody ever thinks about because they’re boring and bureaucratic—these are the things that we’re trying to shatter right now so that when game time happens, things move more quickly. Today, those companies that are potentially pursuing research trajectories or agendas—in some cases, multimillion-dollar, like your $10 million-plus research agendas—that are oriented in a way that an appropriately placed person in the intelligence community would look at and be like, “That’s kind of a nonstarter,” we want them to get that feedback as soon as possible so they can reorient toward things that do have a chance of working.

Nathan Labenz

How important do you think it is that we remain on the same fundamental tech tree or AI paradigm across U.S. and Chinese AI development? My sense is that we’re in some ways in a very fortunate position right now because we’re basically building the same tech in the same way, and we’re sort of encountering the same surprises along the way.

The other question is, I feel like if there’s anything good to be found in the OpenAI-Anthropic adversarial dynamic, it would be that maybe they can be a testbed for techniques that might later scale to a U.S.-China dynamic. If I were the president, I would say, “You two have to figure out a way to police each other.”

Maybe we expand that circle to a few other frontier companies. But you two, you’re the ones setting off all these alarm bells. I need you guys in a room. Whatever technology you need to develop, whatever access you need to give one another, it’s on you to figure out a way that you can trust and verify one another.

Then maybe we can scale that up to a trans-Pacific dynamic that could work similarly.

Jeremie Harris

I think that’s not—

Nathan Labenz

What do you think?

Jeremie Harris

That’s not insane. Obviously, there are big differences between—

Nathan Labenz

Yeah.

Edouard Harris

—what Anthropic and OpenAI respectively have on each other. Although poaching of personnel does a decent job of mirroring the kind of access that China clearly has to the frontier labs anyway. But there—

Edouard Harris

There’s also a basis of trust, I think, between two fundamentally U.S. companies with fairly similar values and blah, blah, blah, versus fairly radically different ones. Not to say that this is totally a miss or whatever, but there are going to be some differences as well as some similarities. I think the similarities might be worth mining. Yep.

Jeremie Harris

Yeah. It’s also the case that you’re talking about the importance of the stacks being aligned. The hardware lottery does a lot of really good things in this space. I think the most crucial thing is that the U.S. and China clearly don’t trust each other in terms of the motives that bring each respective side to the table, right?

When China sees the U.S. come to the table and raise issues like slowing down AI or safety guardrails, whatever, the interpretation that we’ve heard consistently from people involved in Track 2 or Track 1.5 kinds of dialogues—and, Nathan, you’ve been kind of in that ecosystem or touched it as well—is that the Chinese view it as an attempt to curtail their own development because they see themselves as being behind. They’re justified, therefore, in doing things that even wouldn’t be appropriate for America to do in their eyes just to catch up, because they’re in second place.

Being in second place with respect to scale also means that you don’t see the warning shots with the same resolution. However, they seem to also be more public than at least I would have expected. The spillover is something that’s nice, because China can verify themselves directly.

That might be a mitigator if you see the same kinds of failure modes emerging from whole-brain emulation, or if some completely wacky other branch of the tech tree were to become dominant in China. I do think it’s good. I think we kind of get there by default. It’s hard to imagine alternatives that really shake things up at this point.

Jeremie Harris

Quantum machine learning—if you wait long enough, I just don’t think that’s going to be relevant on the timescales that matter and could radically reshape algorithms. But what we’re seeing right now is an industry that’s more or less doubling down on transformer MoEs, with some bells and whistles and some variations here and there, but everything is kind of a transformer, and that’s what seems to ship. So, yeah, I expect that we will have the benefit of that.

It also comes with the benefit of being able to share safety technology, as the U.S. did with Russia during the height of the Cold War at times. In principle, that does mean that we can work on each other’s safety stacks, and that might be the source of some trust-building measures, though that term is also problematic for China, and they’ve pushed back on attempts to do that sort of thing in the past. So, I think it’s a good thing. I don’t know how far it goes.

Nathan Labenz

With the Harris brothers, the discussion returned from verification technology to the diplomatic evidence for cooperation.

Nathan Labenz

I want to not be naive, but I do want to notice and give appropriate weight to positive signals as I see them developing. I think Xi’s speech at the WAIC was pretty friendly and conciliatory. He gave credit to the U.S. for inventing AI. He certainly didn’t call for an international arms race, and he warned against overstretching the national security concept.

We can dismiss that as just nice talk. We probably should have at least some weight on that possibility. But what are the meaningful things that you are watching for—the decision points that will update your thinking on whether they are inclined to at least try to control AI for their own narrow self-interest? Are they inclined to meaningfully cooperate, or are they inclined to seek some sort of domination, as we often project that we are interested in doing?

Edouard Harris

So, in terms of what signs to watch out for, I think anything that looks like a positive sign is at least a positive sign to some degree. A conciliatory speech is good. It’s good. At least it’s not a hostile speech, right? It could be worse. Everything is tempered by the fact that rhetoric is often used for strategic purposes in this way and to shape the battlespace in terms of the narrative and so forth.

The point is not that these are positive signs. It’s just that we have to weigh the evidence in the context of the credibility that has or has not been established by this entity over the past span of time.

In terms of slightly more unfakeable signals that they could give off that would make me go, “Oh, whoa, okay, this looks legit,” you can imagine maybe something like DeepSeek and Zifu coming to some sort of pacing agreement because some crazy thing happened over there—their equivalent to the Hugging Face incident. You can imagine the cyberspace commission that has jurisdiction over whether models can be released and whether they’re properly ideological, actually putting the brakes on something because there was some misalignment thing, and they’re finding that a model that was properly ideological in testing suddenly isn’t being ideological in the wild or something like this.

I would say indications that seem genuine—that they are starting to be on the receiving end of these incidents at the same level of detail as our own frontier labs—I think that begins to make all of us as humans go, “Oh, you know, maybe the thing we should be concerned about is the giant shoggoth thing that we don’t understand and that we’re growing in the labs at an accelerated pace.” That would make me feel a little safer.

Jeremie Harris

Well, and maybe procedurally too. If you take—we have a list of, I think, 5 different historical traps that we’ve seen in U.S.-China diplomacy. These are essentially a brief catalog of the ways in which China behaves when they’re full of shit, at least by the assessment of a lot of the diplomats we spoke to. I should be clear, actually: There was a dissenting diplomat who felt that some of these things were much more sincere, including the use of the language of arms control, the objections over language, this and that. And that itself is the epistemic problem that we talked about earlier.

But basically, I would say, take each of those red flags and flip them over, and you get the corresponding green flag. So, if you don’t see arbitrary concerns raised about language that seems random, that’s a green flag. If you see engagement—and this is actually really important—from empowered people, arguably as we did with Xi, though again, you have to calibrate everything with the fact that we’ve seen this before in other contexts, it’s certainly not a red flag.

But in terms of concrete commitments, that’s the sort of thing you look for: empowered people, the lack of capricious, arbitrary objections. Things that actually look like they’re making qualitative progress are a surprisingly good sign because you can contrast them directly with how things have gone in the past, which is not very good. I mean, the contrast point is actually that low that these can be genuine signals.

Nathan Labenz

What are the things that we can do that are not so costly to us but are still credible signals to them that we are not going to try to use AI to gain a decisive strategic advantage and ultimately make them an offer they can’t refuse? If indeed that is not what we’re going to do, which I’m a little worried we might actually be about to try to do. If we were on the path of trying to seek a Pax Robotica, where we can all benefit from the abundance that AI, especially in its Chinese-manufactured, embodied form, might provide for us, what would be the steps that you would prioritize next on our side?

Edouard Harris

Well, there may be some stuff we can do that’s not functionally really even that costly. Generally, as I think Jared and you guys maybe as well highlighted earlier, certain kinds of transparency can be stabilizing. The kind of transparency that goes like, “Hey, we’re giving you enough vision into what we’re doing to see that we are not doing the thing you fear most”—that sort of thing is potentially useful, and additionally may not actually be that costly for us to do, depending on how we implement it, simply because the Chinese are already all up in our systems.

So really, we’re not giving anything away that they don’t necessarily have already, in many cases potentially. So maybe some kind of visibility into, “Here’s what we’re doing, here’s what the frontier labs are doing,” and so forth. The problem is that’s not necessarily something you want to be doing unilaterally. That would come as part of a trust-building measure, dare I say, between two powers in the wake of a moment like this.

Goodwill-type stuff we can do now would certainly include more interactions between the verification communities in the United States and in China, and this kind of thing is already happening, actually. There are some quite good and positive interactions between those communities. So, yeah, the kind of linkages that you get at the level of academic-to-academic, startup-to-startup, all trying to solve for the same mission—these are very, very positive things.

It’s true that right at the political levels, the 2 countries have started separating out, and even at the level of big companies and stuff like that, you see the classic Chinese spinoff story and all this stuff. But there still are real, genuine linkages between the 2 countries, especially on the academic side and at a number of other levels. There are a bunch of sincere people just talking to a bunch of sincere people about, “Yeah, this is a problem and this sucks. Yeah, I agree. Let’s try to solve it.”

So the more of those linkages there are, the better. The more people there are on both sides of the ocean who have the ability to talk to their own domestic leadership and say, “Look, I’ve spoken to them. They’re not evil. They’re just trying to do this or that,” to whatever extent that’s true, that kind of moderates the more extreme tendencies on both sides. It’s very, very hard to do that completely because the actions of these countries are also constrained in a number of ways, but it really does help, I think.

Nathan Labenz

Part two: The economy of intelligence. Steve Hou leads research at Silicon Data, which builds price indexes for rented GPU capacity and model tokens. Turning compute into a measurable market means deciding which prices can be compared. Steve starts with the inputs to the GPU rental index.

3. GPU Rental Economics

Steve Hou

We use a combination of both quote prices and transaction prices in our calculation, making that distinction clear. This is a large normalization process, using machine learning to help us make the contracts apples-to-apples comparable, to the extent that you're looking at a single chip—let's say an H100—being rented from different parts of the world.

One question that jumps to mind right away is that, for example, in the case of an apartment rental index, you would never consider using just an offer price. Someone who lists a number on the front of a building, saying you can rent an apartment for $2,000 a month—you wouldn't expect to just trust that number. You can walk in. But we do use quote prices. Why is that?

The reason is because, unlike an apartment, you cannot click an API button and just get hold of the apartment. You have to walk in and talk to somebody and negotiate, and that apartment may or may not be available. In this case, very often, with API-executable GPU rental, you can actually get hold of the GPU, the same way you can buy something on Amazon, right?

That being said, we also have a transaction, so we're comparing them. If someone who has 3 nodes of GPU rented out, say, all 3 at 275, I will not presume that I can go back to the same merchant to rent another one at 275. It could very well be the case that the next one is not available anymore. Or if they rent out 2 out of 3 for 275, the next one will not necessarily be 275 again either.

It could be $4 or $6 or $1, depending on what everyone else is quoting. The person quoting that price would be crazy if everybody's quoting at $4 and they continue to rent at $2.75. You would think either they're going to raise the price or there is something wrong with that price, right?

This is the reason why I gave you a long answer again. What we do is that we want to provide as broad a set of coverage of the market as possible and normalize everything so that we're capturing the market as it is, as faithfully as possible.

Nathan Labenz

One big difference I noticed in prices—

Steve Hou

Yeah.

Nathan Labenz

Just browsing the Silicon Data website is—

Steve Hou

Mm-hmm.

Nathan Labenz

Between the neo-clouds and—

Steve Hou

Yes.

Nathan Labenz

The hyperscalers. These are not small differences. These are multiple differences in prices, seemingly 2 to 4×. That's a pretty big delta. Why does that delta exist? Is there no way to arbitrage it? And does that imply that GPU hours are being wasted, or that when I buy from a hyperscaler, I'm sort of preempting their internal work and they're using everything that's not sold at runtime? Give me a little peek behind the curtain there.

Steve Hou

Again, I like to use analogies. I gave you an analogy using a single-bedroom apartment. I'm not going to use an apartment again, although I can. You can, by the way. Let's say we're doing a single-patty burger index, right?

You can buy a burger from a burger stand off the street of New York, or you can walk into a high-end steakhouse and order a burger. Believe you me, those 2 burgers are going to cost very different amounts, right? They're both burgers, and what we try to do is, as much as possible, find the marginal price for a unit of compute—or a burger, in this case—that has the same relative features.

I wouldn't want to compare the price of a 3-patty burger with a single-patty burger, right? But once I make those adjustments, there are some adjustments I cannot reasonably make because it's capturing a different type of premium from product bundling or product differentiation.

In the case of hyperscalers, indeed, you are observing correctly: They charge regularly, consistently, at least 2 to 3 times, sometimes more, compared to a typical neo-cloud. The reason has to do with a long legacy of other types of products being offered on their platform—software, analytics, safety, compliance—the fact that they already have this long-established relationship with enterprise users that have been on board for a long time and that have a certain stickiness for moving. It is being sold as very much a differentiated product.

Going back to that single-patty burger index analogy I gave you, I see people who sell a single-patty burger with fries and a milkshake. I can try to strip out the prices of those 2 elements and isolate what I think would be the price of a single-patty burger from that merchant.

But if they, let's say, blended up the fries and the milkshake and injected it into the patty and said, “This is a brand-new product,” and charged 2 or 3 times the price, I can't very easily strip it out, right? At which point I say, “Okay, I'm raising my hands. You guys are a little bit different of a beast, and let me put you in a different category,” right? And that's how we have so far handled it.

We believe that this way, you are getting to a much purer form of a single unit of compute, right? That is actually getting you closer to this idea of fungibility, to the extent that things are direct substitutes for each other.

Nathan Labenz

The same chip can perform differently in different places. And I think, according to your site, you have some software that runs on the chip—

Steve Hou

Silica Mark, yeah. Mm-hmm.

Nathan Labenz

Yeah, that you—

Steve Hou

Yeah.

Nathan Labenz

That you benchmark the chips with. So I guess every single price in your index has been benchmarked by this system?

Steve Hou

No. We have a physical benchmarking service called Silica Mark that actually visits individual GPU at the EU ID level to try to assay—to assess—the GPU's health, performance, throughput, TFLOPS, and so on. At the moment, that physical benchmarking and physical spec performance does not enter into our pricing, right?

Nathan Labenz

I see.

Steve Hou

We do not see a strong relationship, at least at the moment, given the nature of the market, with how specifically the GPU is performing.

Steve Hou

Even though we do see that in the cross-section, you can actually have a bit of variance from GPU to GPU. This is not surprising or new to anyone who is in this space—the GPU lottery, right? But over time, if you have a cluster, it probably averages out, and the law of large numbers kicks in. So we don't use that. We have 6 features, including geolocation, CPU, memory, various other things, term, and provider. But the physical spec is not one of them, at least at this very moment.

Eventually, when we head toward a scenario where we potentially could have physical delivery—because inference, based on my thesis, could make compute more interchangeable—that could enter into our pricing scheme. But at the moment, it does not.

Nathan Labenz

Another price that I noticed had moved on the website and that caught my attention is the proprietary LLM index.

Steve Hou

Mm-hmm.

Nathan Labenz

There, you break down token pricing into—

Steve Hou

Open and closed, yeah. Overall.

Nathan Labenz

Right.

Steve Hou

Yeah.

Nathan Labenz

And so the proprietary one is up 11.5% over the last 7 days.

Speaker 2

Mm-hmm.

Nathan Labenz

And I guess I'm wondering, how are you measuring that? Because, quality-adjusted, everybody would say prices are coming down, right? Even dramatically so. The retail posted API price hasn't changed, right, except when they introduce new models. So what are you measuring on a day-by-day basis that allows you to say what the proprietary cost is doing at such a fine grain of resolution?

4. Tracking Token Price Changes

Speaker 2

Yeah. So first of all, I want to give a little bit of context for what our token indices mean, because I think they've repeatedly been misunderstood. Partly, I think, it has to do with the unfortunate naming. We called it the expenditure index, and then people maybe thought it either meant price or total volume, when it doesn't actually mean either. It's actually an expenditure-weighted price index that's normalized to 1 million.

That can show trends based on usage mix, which I'll come to in a second. When you point out the proprietary LLM model, which is sort of the frontier model, having recently bounced a little bit higher, we should put it in the context that since maybe late June through basically the middle of this month, it has been on a very sharp downward trend, going from some $4 to $1.60 or something. That's more than a 50% drop, right?

During this time, we've seen a lot of frontier leading labs not just cut prices on their newer variants, but also release cheaper variants of powerful models and new-generation models. We've also seen other proprietary models coming out, like Meta and Groq. Don't forget, these are proprietary AI models as well, right? They've been very aggressive on the price front.

So more recently, I think the bounce could come from a variety of sources. If people decided to say, “Okay, I actually quite like the more expensive, powerful model from Anthropic,” and they used more of it, that could drive up the expenditure-weighted price index.

The analogy I'd like to give people is: Forget that these are all LLMs. Imagine these are cars, right? If you have a Mercedes with cheap and expensive variants, depending on which cars are being sold more and which people like more, that could affect the average price of the cars sold. It's the same thing here, right?

In the market for tokens, there are 2 things happening. Token model prices are changing, but usage behavior is also evolving. To the extent that we observe volume from a handful of these public inference platforms, serving platforms that allow you to look at how people use different types of models, I think this most recent bounce of the frontier proprietary LLM index—I haven't looked into the details, but I suspect it has more to do with usage mix than anything else.

Nathan Labenz

From measuring compute to using it, the next conversation is just the 2 of us discussing OpenAI DevDay. I start with Dots, the personal agents OpenAI presented, and who might want a system they don't have to maintain themselves.

5. OpenAI Expands The Platform

It was funny. When they led off with the whole Dots thing, I thought, “Well, I think I have all of this, and I'm pretty sure I'm going to continue to prefer my version that I've gradually evolved over the last 9 full months now.” So that definitely didn't really feel to me like a developer product. I thought that was a little bit muddled, because it very much felt to me like that's a consumer product, right? That's for ChatGPT users to use, and it wasn't entirely clear how that would be used by developers, if at all.

I haven't been down every last breakout-session video, so there's possibly some more that I missed. But certainly at the keynote level, it felt like that's a product they are offering on a first-party basis to their users, and I do think it will be really useful for people. I guess I would say my guess is that people are really going to love these Dots.

Certainly for my mom, on the other hand, I would say, “Go for it. Just use that.” It's probably pretty easy. You don't have to worry about taking on all this stuff yourself, managing your own database on your computer, or troubleshooting when things go wrong, even though the models are getting so good at that on their own. I do think this higher-level and more polished abstraction will probably be really good for a lot of people who don't care to learn a bunch of new tricks and just want to have this thing that they can delegate to.

Nathan Labenz

Prakash then turned to the models OpenAI presented at DevDay. These are his first impressions of their capability and speed.

Prakash Narayanan

They had GPT-6.1 Sol and GPT-6 Astra Ultra Fast. GPT-6.1 Sol is about the same price, or slightly lower price, than GPT-6 Sol. But it's basically Astra for the price of Sol, which is what they're calling it. It seems to be a very competent model. I've used it. It's better—definitely better—than GPT-6 Sol. GPT-6 Sol was only released 2 weeks ago, or 2 or 3 weeks ago. So the cadence of releases is stepping up.

GPT-6 Astra Ultra Fast—now you have… They used to have a fast mode, which is 2 times the speed. Ultra Fast is 8 times the speed. They demoed how Ultra Fast works. With Ultra Fast, you can interact as you type. It's an interactive kind of build-out, software build-out, interactively building games.

I think that's really the future. I think speed and latency—latency at the same intelligence—is probably going to be very important, especially as you clear these hurdles of capability. The thing can build a game, but can it build a game with you in the moment while keeping you in flow? I think that's what's coming up next.

Speaker 0

Still in our DevDay conversation, I turned to Sign in with ChatGPT, which lets eligible users bring their plan to participating apps. My example comes from Waymark, the video company I co-founded, and concerns the cost of letting a new customer try an AI product.

Nathan Labenz

The classic “try 1,” “try 7 days,” whatever—those are tried-and-true tactics that have become very difficult in the context of, “Oh, but I have to spend a certain amount on tokens for every new user to give them that decent experience,” especially if you have something that involves a decent amount of setup or profile processing.

With my company, Waymark, we're not by any means the most token-hungry business, but the first thing we do when you sign up is make a big profile of your small business so that we can then use that profile as an input to make video content later. That profile-creation process has come down in price. I don't know exactly what it is today. At one point, it was about $1 per customer, and it was like, “Okay, well, this does start to become material when you think about the all-in cost of customer acquisition, whether we can make this flywheel work, how fast we get paid back, and all that sort of stuff.”

But I think this is a great value-add to your ChatGPT subscription that you can now take around. Of course, developers will need to implement this, but it shouldn't take too long to tell your coding agents to implement it. So that's great for ChatGPT, great for users, and great for the SaaS companies.

I think this is a huge win, and I do think it's pretty pro-ecosystem. It seems to me that this is one way in which they can actually follow through on the promise of not trying to eat the world and instead trying to empower people to build cool stuff.

Speaker 0

Part 3: The craft of co-writing. Joel Borgen is the author of The Receipt Horizon, a novel set in a world shaped by advanced AI, which he wrote in collaboration with AI models. We discussed what he supplied as the author and how the collaboration changed the work.

6. Writing Novels With AI

Joel Borgen

I may have written the book eventually, but the activation energy required to get started when you have kids and a part-time job was limiting. I started it well over a year ago, and the tools at the time were not as developed as they are now.

There were a lot of limitations, especially around context length. You have the thread where you're trying to work on a chapter or part of the book, and you're constantly having earlier context drop out and needing to manage that carefully. I think it's getting easier and easier to make use of the tools to do something collaborative while still keeping it kind of your vision.

I started the project knowing more or less what I wanted to have made. I sketched it out myself, and especially the first half or so of the book was pretty well set before engaging the models. The models are good at certain things, and they're getting better at everything. But as far as prose writing itself, even if you tell it exactly what you want and you have a plan for a chapter, you often get something that's unreadable unless you know how to scaffold it and then edit it yourself.

I've been working on a follow-up book, actually, and the process has changed quite a bit as the models have become more powerful. The book is definitely a lot better for having done it alongside the AI, but I'm cognizant of the fact that AI writing will be controversial for a lot of readers, and they don't want to engage with it. I think the process made the novel a lot better, and it also had weaknesses that I didn't entirely account for.

Nathan Labenz

I had asked Joel about building a positive vision of the future into a story that still needs conflict, and about the scaffolding he gives the models.

Joel Borgen

Utopias often want to become dystopias in narrative form because you need to have conflict and friction, and it's kind of dull to just have everything work out and everybody's happy, of course. The fact that this was set after a big AI war, and humans are presumably locked into what they're able to do, means you see both incredible flourishing in terms of the technology that's available and the options people have, and also some really major downsides.

I do think so much of what you've talked about, and I agree, is trying to paint a positive vision for the future, and that is difficult to do in narrative form. I didn't want anything in the book or any character to be a mouthpiece for me. A lot of the stuff that I want to see created in the world is a part of the story, and I try to steel-man the people who would have different opinions about that as well.

As far as the scaffolding, I often write out a large-scale architecture for the story and a lot of beats, and then go back and forth with the major frontier models. I've primarily used whatever the flagship Claude model is, and then the ChatGPT Pro model from GPT-5 on, because that was the workhorse. It did the most detailed work. Claude used to be a lot lazier as far as how much it would follow up on.

I'm doing it mostly through the chat interface, which probably has advantages for me and also limitations. Context length, like I said earlier, was a huge unlock as that got longer. At one point in drafting the first one, you could give them the entire book, and every time a new model came out, you'd use it to try to make it better, up your game, and allow it to do more.

As far as the actual scaffolding for the story, you go in chunks often, and I'll draft chapter by chapter. When I was done with it, I got the advice to hire professionals for copy editing, art, and layout. I did that, and I learned a lot from doing it. It was very expensive and time-consuming and kind of slow, and there were definitely frustrations involved.

I've tried to extract what I've learned from that and what made the book work better. I've fed it to GPT-5 or o3 Pro and developed my own workflows for doing the same process with the second book. I think that'll be an interesting experience, to see how well that works.

Nathan Labenz

With Joel, we moved from drafting to judgment, where he still wants a human author making the choices, and to how he edits the prose.

Joel Borgen

As we move into a world where the AI systems can do a lot of what we do—a lot of what we thought was valuable, what we thought was a human contribution—you think about the abilities they have now. The reason I found the book valuable to write is that there was a lot of me in it. It took a lot of effort on my part still.

I think I've heard a podcast episode where there was somebody writing books on Amazon in the romance area, and they were mostly automated and AI-written. I think there were something like a couple hundred per year, so they were playing a scale game. A few of them would make a few dollars, and overall it was profitable. You can certainly do that, and the models are getting so much better that you can probably have customized artwork of the same quality as, or better than, the book that I created over many months, on demand.

What happens to the art at that point? Thinking of it as an art consumer, if there was a new movie by Ingmar Bergman or Stanley Kubrick or something—movies that I've spent many hours thinking about and consuming, and that are deeply moving to me—what would that world look like where there's just a huge plethora of art on demand and it's pretty much custom just to you?

We already have a shared cultural space that's dropping off. People are more fragmented in what they consume, and there's less unification across the cultural landscape. Is that art still valuable at that point if it's not something that we're sharing with other humans? I don't know the answer to that, but I think we're in kind of a sweet spot now where, in order to get a work that you're proud of while working with AI, it still requires a great deal of human judgment and work. I don't know how long that will be the case, given where things are headed.

I had my book edited down about 14%, mostly manually and with AI helping me decide what to cut. Often, you can take a chapter that's so-so and improve it by taking away some of the stuff that AI does particularly badly: the tics and the over-explanation. I think you could architect something now that could create a pretty decent book with the right kind of feedback loops.

A lot of the stuff that I've developed is based on my own preferences. I want to be able to give the entire novel to a model and have it come back with an analysis and a list of things you should consider, because doing it manually and slowly is extremely time-consuming, as you can imagine. Bringing things to your attention like an assistant and saying, “Here's something you could consider,” and then letting you decide for yourself, point by point, is a fairly satisfying way to work.

Nathan Labenz

Joel is also a musician. To describe a capability he had watched arrive, he turned from prose to musical notation and a viola-clef test he had kept trying on successive models.

Joel Borgen

I've had this held-out test for well over a year now, maybe 2 years even. It's very simple: a high-definition capture of 5 notes on the viola. I give it to the reasoning models and ask, “What is the clef? What are the notes? What are the note values, and so forth? What's the time signature?”

Not a single one of them got it right until Astra. OpenAI presumably put in some actual musical scores. It went from basically not being able to do anything to—now I gave it a complex 20th-century score that's almost painful to look at, given how complex it is. There are double sharps and accidentals and ties everywhere. It's difficult.

It did an almost perfect job turning it into a machine-readable form via Codex. This score that you basically just had as a PDF before can now be turned into something that you can manipulate and evaluate. It went from kind of 0 to 99 overnight. I think that's an interesting way to turn some of the scores that aren't machine-readable into something that you could use to train a system.

I would love to see Suno move in the direction of adding a lot of classical stuff to it, because I think, in addition to creating more interest in classical music, classical music requires a different level of understanding of the form because it's often much larger-scale. You'll have a 3-hour-long opera that has some internal structure, or even a 20- or 30-minute single movement, something that has a lot more going on than a pop song. I would love to see that happen.

It does tie into the same element of, if you're making just private art mainly for yourself to listen to, I still think it's valuable. But again, like with writing a book, I would want to do it in tandem with the hypothetical model in the future and create something where a large part of my effort and taste goes into it as well. I don't know how long that interregnum will last, where you have a role for humans and AI to create something together.

Joel Borgen

We use it as a tool, but it still requires a lot of you and your judgment. But I look forward to experimenting with that once the models improve.

Part 4: The diagnoses still waiting. Daniel McKinnon founded Gamo Labs to use AI agents to interpret genomes and revisit difficult, unresolved cases. His son Owen died from a rare lung disease. A human specialist later found the genetic deletion the original sequencing analysis had missed. Daniel subsequently built an AI pipeline that rediscovered it while his family was seeking answers during another pregnancy. That experience helped lead him into this work.

7. AI Finds Missing Diagnoses

Daniel McKinnon

It really leapt out at me last summer when—we blessedly have a healthy 13-month-old right now, but because of our history, we are being monitored very, very carefully. There was just something a little bit questionable in the 16-week anatomy scan that our own maternal-fetal medicine doctor, who is an absolutely wonderful person, said, “Normally I wouldn’t even flag this, but because it’s you guys, we need to look carefully.”

We did a whole-genome analysis of the fetus at that point, and it came back negative. This is actually the third negative whole genome I’ve seen between losing Owen and having our son Warren. We actually lost a second pregnancy due to genetic reasons very late. We know we’ve had basically 5 years of heartbreak before bringing Warren into this world, and all of it was kind of genetically mysterious.

As a family member of a patient, I started to learn a lot about the failings of this system. It was that moment where I said, “I’m heartbroken, but I’m mad, and I’m going to do something, and I have tools to do something.” Fast-forward to last summer: I got this result back and said, “This is unacceptable. This is the number-one thing that I want in my life.” I really wanted to have a family. I wanted to understand what would happen.

I didn’t believe these labs, and I basically called all the labs and said, “Give me my raw data,” which you can get due to our HIPAA rights here. I vibe-coded my own interpretation pipeline. I mean, now vibe coding is crazy, right? Right now, this would be so easy. You could just say, “Claude Code, make me this thing.”

Back then, it was still like the original Codex autocomplete. It did require quite a bit of work on my end, and I did it to see if there was any comfort we could get around the pregnancy. I was shocked that it also outperformed on these other cases. Most clearly, it diagnosed Owen when one of the best—or some might even say the best—prenatal sequencing lab did not. I knew at that moment that this was something I had to contribute to. I didn’t know it would be a company.

Nathan Labenz

The deletion in Owen was 91 kilobases long and removed an enhancer, a region that helps regulate a gene. The specialist had already identified it before Daniel recovered it with AI. We asked Daniel how the original analysis missed that enhancer.

Daniel McKinnon

How you miss it is that this enhancer is a megabase, or 1 million bases, upstream from the gene. What Rady did, as many other labs do—which is a very reasonable assessment—is use a filter for structural variants, which tend to be quite messy in what’s called next-generation sequencing. That’s how sequencing is done today, where you have roughly 1 billion 150-base-pair fragments and need to piece them into this clinical puzzle.

They said, “We’re going to have a filter, and anything more than 1 kilobase up- or downstream of the gene, we’re not going to consider.” This is a totally reasonable trade-off if you have humans looking at all of this stuff. But my thought was—and this was before Codex, before Claude Code—can you put the o3 model, which was the model I used at that point, into some kind of loop and have it keep looking?

It basically looped. The first loop was, “Is there anything wrong with the coding elements?” Those are kind of obvious things, and it was using these bioinformatics tools. This is very crude compared to what we have today, but then it misses something and goes through another loop and another loop, and it just keeps working.

When you are a clinical lab, whether you’re for-profit, nonprofit, or whatever your structure, ultimately you’ve got to move. You have to spend some amount of time on each case, and if you don’t come to a conclusion, you say, “This is nondiagnostic.” This is very common. Most whole genomes, even from infants suspected to have genetic disorders, come back nondiagnostic.

I really think it’s one of these meat-space problems. If we can export these problems onto a machine intelligence that can work nonstop and in parallel, then we will be able to see many more kids, treat many more kids, and do much more interesting analysis on top of the basic things that humans are just pressed for time to do.

I don’t want to claim I’ve reinvented this. People are trying to build software to accelerate genomic interpretation, and they’ve been doing this for a long time. But I think what I probably identified relatively uniquely, early on, was that this is a really great task for agentic AI.

From an improvement perspective, it’s long-horizon, agentic, verifiable, and unsaturated. I suspect that this will be a task like math, where we can just generate these Olympiad-level problems and keep hill-climbing on that until the problem is basically solved.

Nathan Labenz

Daniel McKinnon then turned to the unresolved cases his team is analyzing now. Sometimes the remaining obstacle is a variant whose effect is unknown. We asked where the new biological evidence would come from.

Daniel McKinnon

On the back end, once we’ve done the interpretation, the majority of clinical cases we see right now—and to be clear, we are only seeing hard cases, so this isn’t like most general labs—end up in what’s called a VUS state. That’s a variant of uncertain significance, and these are scored according to a very standard rubric developed by the American College of Medical Geneticists, or ACMG.

You need to get 6 points or more to be bumped into the likely pathogenic category, and that unlocks a lot of treatment options and insurance options. Let’s just say it’s good to be either benign or likely pathogenic. It’s very bad to be in this kind of intermediate stage.

If you are in this intermediate VUS stage, you can do 2 things to get a diagnosis. One is that you just wait, and you can get more points if more patients emerge who have a similar phenotype or a similar condition as you do, and the same genetic variant.

This is often what happens when you read in the news, “This kid has had epilepsy for 10 years. It finally got a diagnosis.” If they’re lucky, “Oh, and since we know that’s the diagnosis, we worked with a pharmaceutical company, and there’s some off-label use of a drug, and it can actually help their condition.”

That’s typically because they just waited until somebody else had the condition, but this is not scalable and takes a long time. Another fork is that you can convince a university lab to care about this problem. What they’ll do is what’s called a functional study.

They’ll make a cell line that’s emblematic of your particular condition. For example, for alveolar capillary dysplasia, that commonly used cell line is called IMR90s. It’s a fetal lung cell line, and this is very, very important because this particular gene is only expressed from week 16 to week 20 of development. You can’t just put them in a cancer cell line.

Then you can do something like edit the genome, insert a plasmid, or conduct any number of different functional studies to say, “Okay, in this physiologically relevant cell line, if I have this genetic mutation, is this gene broken in some way?”

Nathan Labenz

Daniel’s team had just taken over an existing biology lab to test uncertain variants. He described the experiments starting that day and the feedback loop he hopes to build between real biology and the models making predictions.

8. Building The Biology Feedback Loop

Daniel McKinnon

We basically bought what was remaining of Arpeggio, hired 4 people onto the team, and very quickly pivoted to creating this basically like RL data from the real world with biology experiments. We’re 3 weeks in, we’re running our first experiments today, and I’m really looking forward to being able to close that feedback loop.

I think a lot of clinical genetics labs are interested in this as well, since this service is not commercially available. We can not only sell a lookup-table version of this—“I have a patient with this variant. Can you help me get a couple more ACMG points to get this up to likely pathogenic?”—because you can get 2 to 4 points. Remember, you only need 6 points, so 2 to 4 points is a lot of points for a functional study.

But it also lets the machines learn. Right now, the machines aren’t learning. You get to a VUS, and that’s the end. Now it’s like you get to a VUS, you do the biology study, you say, “Oh, this region of the genome is actually quite important for this particular disease. We’re going to do that, we’re going to check these predictions, and then we can improve these predictions over time.”

Nathan Labenz

We had been discussing earlier genetic testing companies, including Counsel.

9. The Vertical AI Harness

Daniel McKinnon

I was like, “Oh my God, we sequenced the genome 25 years ago, and Bill Clinton and Tony Blair got up onstage and said, ‘This is going to revolutionize human medicine.’” I look back and there’s this graveyard of genomics companies. You mentioned Counsyl, which I actually would not say is a graveyard. I think they were modestly successful.

But why is all of health care not based on precision medicine? Everyone in this space is like, “It should be,” and there are examples. It’s just too challenging. The interpretation is too challenging.

I actually look back and I see—and we talked about Counsel ahead of time—I think what was missing prior to now was really that interpretation layer. It’s very, very complex. Genome sequencing got what I would call cheap enough maybe 5 years ago—I mean, below $1,000 a genome, maybe $500 a genome. There are high-throughput labs doing this for less than $100 a genome. You see announcements on Twitter saying, “Oh, I can do a genome for less than $100.” There are many people doing this right now.

The sequencing is not the problem. Given your problem, what insights can you derive from that? That was only possible as of last summer. I’d say o3 is the first example of that.

It’s not just the variant interpretation, right? It’s explaining to the provider why it’s important. It’s ranking variants in a nice way. It’s explaining how you can treat this person. It’s ranking different patients for ASO eligibility. There are many, many, many other things beyond just scoring the variants.

Although I will say we benchmark variant annotation in RareBench and other benchmarks we have. The best-performing traditional machine-learning-based approach in terms of variant ranking is called LIRICAL. I think it scores something like 10% on our benchmark, and right now Opus 5.5 in CloudCode, just a vanilla thing, scores something like 50%. These traditional tools are also getting blown out of the water by this newer approach.

But then they can do much more. You basically need to make a very, very cheap, easy thing that people who are not surrounded by fancy clinical geneticists, genetic counselors, or top-tier hospitals can use, and that’s the problem we’re trying to solve.

There’s already product-market fit for sequencing. Every baby in top NICUs is getting sequenced. Every baby in lower-tier NICUs would get sequenced if they had the resources. The problem is much better scoped: The baby has pulmonary hypertension; figure out why—not, is there something obscure that could be wrong with this baby now or in the future? That is where we are working with various hospitals and families today.

There are 2 forks to this work. One is what I call online. A case comes in; physicians do not get good results from the traditional labs. The patient consents to a reanalysis. They send the data to us, we reanalyze it, and sometimes, but not always, we find additional things that can help them make decisions around either a pregnancy or a newborn.

The second fork is bulk-scale reanalysis. These are rare-disease centers that have thousands of genomes, and they say, “There are probably kids that we can diagnose or even treat in that database, but we have a handful of genetic counselors, a handful of bioinformaticians, and we just don’t have the resources.” Because a machine does most of the work here, we can get a set of 80 or 100 or 200 or whatever and turn them around in a week. That’s kind of 6 months or a year’s worth of work that they really can’t prioritize.

Those are the 2 forks right now, and we’re 4 months in. We’re pre-revenue. We’re really not thinking about what the exactly appropriate business model is here. We’re just trying to diagnose more sick kids.

Stepping one step up, what do we do? We’re basically a harness, tools, and evals company. I think you’ll see this a lot in vertical AI in general. I’ve worked in these frontier labs. I started working on LLMs—actually, the first LLM project at Meta, which was OPT 175.

Unless you are somebody who is very famous with very deep pockets, I don’t want to bet against the frontier labs by any stretch. I think it’s possible, but you’re not going to compete at the model layer, and intelligence is progressing so fast.

We ensemble models for sure, and that helps with both cost and performance. If you look at the confusion matrix of RareBench—we also publish this—you’ll see that the cases kind of cluster. Claude is good at this, Astra is good at this, and even Gemini, right? Gemini is not on the frontier, but they’re kind of good at this, and you can ensemble these things together and get better performance.

I’d say that’s one thing we do. And with ensembling, there are very dumb ways of doing this, but there are also smart ways of doing this, and I don’t have to go into all the details here. I would say figuring out unique routing and intelligence per model is an important edge that I think we and many others are doing.

Another key thing that people forget—and this is what I worked on at both Meta and Google—is evals. Unless you have very structured ways to measure the performance of the system—it’s not just the model at this point; it’s the whole system—you don’t know what to hill-climb.

We do pretty robust trace analysis of how models perform with and without our harnesses. This sounds stupid, but not that many people actually look that closely at data. This is a meme on Twitter, but it’s very true.

You’ll see one particular case where Grok 4.6, which at that point was state-of-the-art on our benchmark in Grok Build, missed one because it just renamed the gene. It was saying, “Oh, NRF2 or whatever is responsible,” and then it changed the name of the gene to something totally different. I’d never seen that before, and I was just like, “This is dumb.” Our harness and tools prevent the agents from doing dumb things.

I think there’s a world where, if I worked at OpenAI or Anthropic and had access to infinite compute, and I could do millions of rollouts, the agents would, with enough compute, all converge on some answer and maybe mitigate all of this work—especially because they can write their own tools these days and everything. But really it’s about consistently and repeatedly and, honestly, affordably getting to the right answer.

For $10 a case, who cares? Or even $100 a case. But if you need to do a million rollouts and all of a sudden this is $1,000 a case or $10,000 a case, it doesn’t make sense. So we’re doing a lot more research around how those cost curves work and how everything works.

Ultimately, our job is to take the smartest intelligences in the world, mix them together, give them access to the best tools, and get answers for our patients. That’s the core of what we do, along with being able to measure whether it’s working. I’d say that’s the core of what we do.

Prakash Narayanan

When you talk about vertical AI, would you say you’re looking at benchmarks where you combine an intelligence with the tools that you have? How do you evaluate the strength of the tools on their own, excluding the models? You’re going to upgrade model families over time, right? How do you measure the performance of what you’re building—that layer in between, in particular?

Daniel McKinnon

Yeah, that’s a good question, and I honestly don’t have a great answer to that question, in that we are co-designing the tools with the models. We have seen cases where—I mean, this is a very dumb example, and if you’re getting deep into this and have a specific eval, you’ll start to see this.

If you’re not deep, it’s very easy to be like, “Oh, I just had Claude Code one-shot this video game. Was it good? I don’t really know. It seems amazing.” I don’t want to communicate that these models aren’t amazing. They absolutely are.

With one tool for a specific type of literature search, the model—I think this was Astra—said, “Oh, yeah, okay, these are all the papers you tagged for me to read.” The answer was in the paper, and it missed it.

We were looking through the trace and the context window, and we were like, “I don’t think you actually read these papers.” There are lots of little things like that. They’re little paper cuts, and that’s what these tools do. They show you, “Oh, you missed a case here because of this.”

You missed a case here because of this. And you kind of put it on these better rails. Actually, Astra is a good example. The first time we ran the eval for Astra, it just stopped because some of the tools we had made it want to think about too much.

You see this on X, where people are like, “I asked the model to do some research for me, and I came back 30 minutes later and it’s looking at flights from Dubai to Calcutta on Emirates Airline.” Why are you looking at that? They just do stuff like this.

So I think we try to co-evolve the tools with the models. Coming back to something that many people might find boring but I find fascinating, evaluating a model is very hard, and we need to have very robust evals that can catch these things early. In that case, we caught it early, updated how Astra called some of the tools, and it works again.

Nathan Labenz

Prakash asked Daniel what he most wants the frontier labs to improve.

Daniel McKinnon

Well, I want them to hill-climb my task, right? Everyone wins when these models get really, really good at clinical genetics. We probably have less to do at the harness layer, which, to be honest, I’m both bullish and bearish on vertical AI, in that there’s a lot of value in owning a customer relationship. I think OpenEvidence has shown this.

But that layer is also getting thinner. When I first started this, vanilla o3 would not do this task. In fact, o3 scored 0% on my benchmark. I re-benchmarked it just for fun. So it was like, “Oh, I need to do all this work and scaffolding to get these things to work.” And now there are 20 or 30 percentage points or something. I mean, it’s definitely shrinking.

But that said, my goal is to build this great AI-native clinical diagnostics company, and I want the best person to do the work.

Nathan Labenz

Prakash also asked how Daniel hopes to make those biology experiments cheap enough to run at scale.

Daniel McKinnon

It’s really an AI and robotics thing. You have an arm there that is doing 384 experiments at a time in a bigger plate. You have a cartridge over there with thousands of plates lined up. You have AI assisting with primer design and experimental design.

This is one of those things where you say, “Can you use AI for this stuff?” You have to get special permission, but you can. At very high throughput, you can design and order the primers, design the experiments, and run the experiments.

Then it comes to analysis as well. You end up with very large-scale data sets that you need AI to poke through. I don’t have some perfect explanation—it’s not just this one invention, but a confluence of all those things.

Nathan Labenz

Daniel also described the response to his diagnostic work from people at the frontier model companies. Here, he reports new diagnoses in other children whose cases had been unresolved. What was the process of talking to frontier model companies about your bio use-case experience like?

Daniel McKinnon

It was really positive for me because it’s extremely obvious why I’m doing this, and I’ve never met anyone who was like, “Why are you doing this?” Every single person I’ve met at any of these labs has been like, “I want to help.”

Whether wanting to help as a person translates into, “I can pivot this trillion-dollar company to work on your problem,” is a different thing, and I understand that. I’ve worked at these trillion-dollar companies too. But it’s been incredibly supportive, and I would be pretty surprised if at least the top frontier labs did not spend some amount of time on this problem.

Again, it’s meaningful and it’s good. Even selfishly for them, there’s a really negative AI narrative swirling around right now. You hear things like in Dario’s recent tweet, where he’s like, “We’re going to cure cancer in 5 years.”

We’re diagnosing kids today. We have many examples of kids who are undiagnosed whom we have diagnosed at a small company 4 months in, and that’s a great story. It should be told by them, even if it’s just for selfish reasons like, “Hey, people of the United States of America who are not happy with data center buildouts or uncomfortable with AI taking jobs, here’s a very concrete way where we are helping people today.”

I also think it comes from the individuals. People find meaning in their work, and I think there are a lot of people, especially mid-career in tech, who are like, “I’ve been doing some kind of ad ranking, data munging, whatever, for my whole career, and what you’re telling me is that you have a very concrete way where you can help kids, and I want to do that.”

So I think there’s a combination of the business reasons and the personal reasons that are driving people to this task.

Nathan Labenz

After that discussion with Daniel, I came back to what slower AI progress could mean for families still waiting for answers. The 50% figure I refer to is a score on RareBench, the company’s variant prioritization benchmark, rather than the share of patients diagnosed.

Intellectual honesty demands recognition that this is probably still one of the things that we will be trading off when we pace the frontier in the near term, to whatever extent that actually happens. I still think that’s probably a good idea and probably worth it.

I don’t think the frontier model companies are at risk of not being able to grow a business, not being able to grow revenue, or becoming surpassed if they pace their frontier efforts. But it is important to keep in mind that there are very real problems in the world that AI is climbing that hill on right now but has not finished climbing. This is one that everybody would obviously love to see solved sooner rather than later.

It’s important for me to stay honest with myself and everybody else that this is a very real cost when you’re talking about individual people and their families. It’s that 50% number that Claude gets to today. That obviously leaves half of the questions unanswered, and those questions are extremely meaningful to people.

So I do think we shouldn’t take that lightly, even if it is a trade-off that, on balance, I think we should probably be willing to make some compromises on. It’s a costly compromise. Thank you for listening. Please tell us what worked and what did not. See you in the morning.

Speaker 9

I carried a tune for 20 years and never had the hands to get it down. Kids asleep and the snow piled high, mine the only window lit in town. Then somebody took the empty chair, said, “Hum it once and I’ll play along.” I hummed 4 bars, you played it twice, and asked me, grinning, “Is that the song?”

Second fiddle, play along. Play me every note there is. You never tire, you never slow. I keep the ones that sound like me and let the others go.

Second fiddle, play it bright. I’ve hummed this tune for 20 years. I’m first chair now every night, so play along, play along, my dear.

Ah. All day I read the slides through glass. I know what’s living and what should go. You played the sunrise 3 times through. I kept the one with frost—if I would let you.

Six hundred pages asking more. I laughed and said, “Let’s find the ending,” and swept the extras to the floor.

Second fiddle, play along. Play me every note there is. You never tire, you never slow. I keep the ones that sound like me and let the others go.

Five little notes I pinned up on the stand. Two winters running, you just couldn’t play. Then one bright morning, you read the whole hard page, ties and flats and all, and I laughed out loud that day.

Hey, now listen to you read. Pull up close and let it sing. Second fiddle, take a bow. They say a revolution’s on. A cognitive one—well, how about that? I say, “Boys, just play the song.”

A stranger wrote me yesterday. Said, “Doc, that tune of yours is good.” I said, “It’s mine,” and you just grinned like any second fiddle would.