[BidClub_]
The Cognitive Revolution · · 89 分钟

AI:AM:Trump-Xi 会谈究竟有何意义?什么才算乌托邦?+ AWS GPU 成本达 3 倍 & AI 诊断罕见病

Nathan LabenzPrakash NarayananJeremie HarrisEdouard HarrisSteve HouJoel BorgenDaniel McKinnon

AI与软件生物医药技术政策
YouTube ↗
TL;DR
  • Trump-Xi 接触的意义,与其说在于缓和关系,不如说在于 AI 事故迫使各方提出“疯狂要求”前,先安装危机处理机制。 Edouard Harris 对建立潜在的 AI 风险热线表示欢迎,但也提醒,中国过去曾把不接核热线作为筹码。若没有更好的核验手段,第一项可核验要求可能会是关闭所有超过可见规模门槛的运行中集群——这将使数十亿美元的 GPU 折旧暴露在风险之下。

  • 核验初创公司可能成为前沿 AI 收入的准入关口,但前提是情报机构在危机发生前完成技术审查。 即便某个系统只比监测吉瓦级热信号“好 50%”,也可能让一半数据中心继续运行、避免关闭全部设施,并节省数十亿美元;但要确保对方合规,仍需要进攻性能力。Jeremie Harris 强调,一项技术可能需要数年才能被认可为国家技术手段,而当前一些耗资1000万美元以上的研究项目,在情报体系看来可能从一开始就走不通。

  • 共同的 transformer-MOE 技术树能让美国和中国看到相同的预警信号,却不会带来共同动机。 中国可能将美国的安全提议解读为维持领先优势的尝试;相比安抚性言辞,真正代价高昂的行动——比如事故后中国实验室放慢研发节奏,或监管机构阻止一个发生失配的模型上线——更难伪装。互惠透明的成本可能相对较低,因为正如 Edouard 所说,“中国人已经全面进入我们的系统”。

  • 大型云厂商的 GPU 租赁价格至少是 neo-cloud 的 2-3 倍,因为企业买的是一套整合式客户关系,而不是可互换的芯片小时。 SiliconData 的 Steve Hou 将溢价归因于分析、安全、合规和客户黏性——这相当于把薯条和奶昔“注入肉饼”后的计算产品。其 token 指数同样反映使用结构:单周上涨 11.5%,可以与自6月底以来从约 $4 跌至 $1.60 并存。

  • OpenAI DevDay 真正重要的变化,是把以交互延迟交付的智能,通过 ChatGPT 订阅变成了可移植能力。 Prakash Narayanan 称 GPT-6 Astra Ultra Fast 的速度提升至 8 倍,用户因此可以“在保持心流的同时,让模型与你当场协作”。Sign in with ChatGPT 也可能为 SaaS 厂商消除大笔试用成本:Waymark 一度仅为建立每个新客户的初始商业档案,就要花约 $1。

  • Joel Borgen 借助 AI 创作的小说说明,模型可以压低创作启动能耗,但结构、品味和删减仍掌握在人手中。 他提供故事设计、搭建初稿框架,并将手稿删减约 14%,经常去掉 AI 的“腔调”和过度解释。他尚未解决的问题是:当个性化艺术大量涌现、不再属于共同文化空间时,它是否仍然具有价值。

  • Gamo Labs 说明,低成本基因组测序并不会自动带来精准医疗:缺失的环节是解读。 一个 o3 agentic loop 找到了导致 Daniel McKinnon 的儿子 Owen 患病的 91-kilobase 缺失;此前,一个面向人工审阅、看似合理的筛选器排除了距相关基因 1 megabase 的增强子。Gamo 现在将模型、工具、湿实验室证据和评测结合起来;CloudCode 中的 Opus 5.5 在 RareBench 上得分约 50%,而 LIRICAL 约为 10%,但这个 50% 是基准测试分数,并不代表患者确诊率。

摘要 · 为研究而整理的核心内容

1. 热线是进展,但不是危机方案

  • Edouard Harris 的开场判断刻意保持克制:“能谈总比不谈好。”Trump-Xi 讨论可能已经促成政府间 AI 事故热线的雏形,即便目前仍停留在理论层面,这依然是积极信号。

  • Gladstone AI 的评估参考了约12名曾与中国谈判的美国国务院外交官。他们的警告是,核热线并不会被可靠接听;在关键时刻,拒绝接听本身就可能成为筹码。如果中共从制度上真正重视 AI 风险,合作或许会改善,但美国的规划仍需要备用方案。

  • 目前最粗糙的应急选项,是利用已经可观测的特征:数据中心规模庞大、建设缓慢,而且热信号明显。发生足够严重的事故后,美国可能要求中国关闭所有超过某一阈值的运行中集群,直到其热信号消失——Jeremie Harris 称这是一个“疯狂要求”。

  • Nathan Labenz 补全了这套逻辑:能力快速进步叠加低水平互信,可能催生极端事故;随后政府会提出一个“离谱、成本极高的动作”,恰恰因为此前没有投资于成本更低的核验机制和“信任但核查”机制。

2. 核验必须在“比赛开始”前完成审查

  • Jeremie 认为,更好的核验技术能够换取经济缓冲空间。假设某个系统仅比观察吉瓦级能源使用量“好 50%”,它也可能核验一半设施、避免全部停机,立即节省数十亿美元的 GPU 折旧成本。在这段时期,AI 利润可能会被“这些小型核验技术卡住”。

  • 监测本身无法迫使对方合规。Jeremie 反复强调,这一点没有例外:没有进攻性选项,“就无法确保合规”;政府可以发现被禁止的活动,然后说“糟了”,却仍然没有任何实际应对手段。

  • Jeremie 指出了一个被忽视的瓶颈:情报机构可能需要数年审查一项技术,并认可它是国家技术手段。因此,核验公司的创始人需要在危机发生前就与相关机构建立关系;而官员到了“比赛开始”时,则需要能够把这些创始人“设为快捷联系人”。

  • 提前接触也能避免资本浪费。一些公司正在推进数百万美元、部分超过1000万美元的研究计划,但如果有合适的情报官员介入,可能会认为这些方向根本不可行;快速反馈可以让公司转向真正能够支撑协议的方法。

3. 共同架构带来共同预警,但不会带来共同动机

  • Nathan 提议把 OpenAI 和 Anthropic 作为国内试验场:要求那些会相互触发警报的前沿实验室建立相互监督机制,再将有效做法改造为美中核验机制。Jeremie 称这个想法“并不疯狂”,并指出员工挖角在一定程度上类似于中国似乎已经能够进入美国实验室的情况。

  • Edouard 的保留意见在于信任:两家美国公司仍然共享更多价值观和制度背景,不能与两个竞争国家相提并论。这些相似性值得挖掘,但实验室之间的机制不能不考虑双方“相当激进地不同”的价值观和处境,就直接放大成外交机制。

  • 不过,Jeremie 仍然认为,共同的硬件和算法发展路径具有价值。两国基本都在继续押注 transformer-MOE,“再加上一些花活”,因此一个生态中的失败,另一个生态也更容易看懂。中国处于第二名的位置,可能看不清预警信号的细节,也更容易把美国的安全提议看成限制竞争对手的尝试。

  • Edouard 认为,Xi 在 WAIC 上发表的缓和性演讲确实有一定积极意义,但代价高昂的行动更难伪装:例如事故后 DeepSeek 和 Zifu 达成控制研发节奏的协议,或者监管机构因为一个意识形态上合规的模型在现实环境中发生失配而阻止其发布。程序层面的积极信号包括谈判代表拥有实权、取得具体进展,以及对措辞的任性反对减少。

4. 互惠透明可能比持续战略模糊更便宜

  • Edouard 所说的稳定性透明,应该足以证明双方都没有在做“你最害怕的那件事”。这对美国的成本可能不高,因为“中国人已经全面进入我们的系统”;但他也提醒,对前沿实验室的可见性不应单方面授予。

  • 近期更现实的机会,是让双方的核验研究者、学者和初创公司进行更密集的接触。这些关系中已经有真诚参与者在表达类似意思:“是的,这是个问题,而且很糟……我们试着解决它。”

  • 这种联系不会消除国家层面的约束,却能培养出一批可以对本国领导人说“我和他们谈过,他们不是恶魔”的人。Edouard 认为,这种来自人与人的证词,至少能稍微遏制双方对彼此作出的最极端解读。

5. 大型云厂商的 GPU 收取的是产品溢价,而不是干净的套利价差

  • SiliconData 将报价和成交的 GPU 租赁价格结合起来,再利用机器学习进行标准化,以便比较不同合同。报价仍然有信息含量,因为 API 库存通常可以立即买到,不像广告中的公寓,可能已经无房,也可能存在议价空间。

  • 单看成交价可能产生误导:以 $2.75 租下3个节点,并不意味着下一个节点也能以 $2.75 获得。如果周边市场的报价是 $4、$6 或 $1,就必须把一笔滞后成交放回当前供给环境中解读,而不能将其视为普遍的出清价格。

  • 大型云厂商的收费通常是 neo-cloud 的“至少 2-3 倍,有时更高”。Steve Hou 将价差归因于打包提供的分析、安全、合规服务,以及积累多年的企业客户关系;客户面临切换摩擦,不一定是在为闲置 GPU 小时买单。

  • 他用单层汉堡作类比:薯条和奶昔可以被单独拿走,但如果供应商把它们混进肉饼里,再宣布这是一个新产品,就不能再拆开计算。此时,SiliconData 会按供应商类别拆分,而不会假装芯片小时仍然完全可互换。

6. Token 指数既随报价变化,也随使用结构变化

  • Silica Mark 会在 EU ID 层面逐一测试 GPU 的健康状况、吞吐量和 TFLOPS,但这些结果目前尚未纳入租赁指数。虽然“GPU 彩票”会造成单机差异,Steve 认为大型集群大概率会将其平均掉;当前价格变量主要包括地理位置、CPU、内存、租期和供应商。

  • SiliconData 的 token 指数是按支出加权、归一化至100万的价格指数,不是纯价格指数,也不是总交易量序列。即便所有 API 的标价都没有变化,只要用户转向更贵的模型,指数也可以上涨。

  • 专有模型指数在7天内反弹 11.5%,但此前从6月底到月中已经从约 $4 跌至 $1.60。Steve 猜测,反弹反映了使用结构变化——类似买家改选更贵的 Mercedes 车型——此前新模型和低价版本曾推动指数下跌超过 50%。

7. DevDay 将延迟和可移植订阅纳入产品本身

  • Nathan 认为,Dots 是一款打磨成熟的消费产品,却被尴尬地放在开发者主题演讲中。在花了9个月迭代自己的系统后,他预计自己会继续使用 Dots;对于他的母亲这类用户,一个不需要数据库和故障排查、由服务商托管的 agent,可能正是合适的抽象层。

  • Prakash 初步判断,GPT-6.1 Sol “基本上是以 Sol 的价格获得 Astra”,而且肯定优于仅在2-3周前发布的 GPT-6 Sol。发布节奏如此压缩,本身就是信号的一部分。

  • GPT-6 Astra Ultra Fast 将此前的 2 倍速模式提升至 8 倍速。Prakash 认为,关键不在于模型能否做出一个游戏,而在于它能否以交互方式完成构建,“同时让你保持在心流中”:当能力跨过门槛后,同等智能下的延迟将成为下一个产品前沿。

  • Sign in with ChatGPT 可以让符合条件的用户将现有订阅计划用于参与该功能的应用,从而降低 SaaS 厂商的试用推理成本。Waymark 一度需要为每个小企业注册用户生成初始档案,成本约 $1,已经足以影响获客经济性;Nathan 称,订阅可移植性对用户、开发者、SaaS 公司以及 OpenAI 的生态承诺都是“巨大胜利”。

8. AI 降低了启动能耗,但没有代替人搭建结构

  • Joel Borgen 最终或许本可以独自写出《The Receipt Horizon》,但孩子和兼职工作让创作启动能耗高到难以承受。他从愿景出发,小说前半部分也已经规划了大半;早期模型主要带来了上下文窗口管理问题,因为此前材料会不断从上下文中掉出去。

  • 更好的模型确实让小说变得更好,但原始文字仍可能“除非你知道如何搭建框架、再亲自编辑,否则根本没法读”。Joel 承认作品存在与 AI 相关的缺陷,也承认有些读者无论创作过程如何,都会拒绝 AI 写作。

  • 他眼中的积极未来仍然从一场 AI 战争之后开始,因为“乌托邦在叙事中往往想变成反乌托邦”:小说需要摩擦力。他没有把某个角色设置成自己的代言人,而是嵌入自己希望看到的技术,同时努力以最有说服力的方式呈现那些反对这些技术的人。

  • Joel 按结构块和章节草稿工作,主要使用旗舰版 Claude,以及 GPT-5 以来的 ChatGPT Pro 模型。专业文字编辑、插画和排版既昂贵又缓慢;此后,他将这些经验输入 GPT-5 或 o3 Pro,为续作开发工作流。

9. 当机器艺术爆发时,人的判断成为稀缺投入

  • Joel 将自己的数月创作过程,与一个自动化言情写作团队作了对比:后者每年可以产出约数百本书,并通过作品组合获利。当生成式艺术能够即时个性化后,他更难回答的问题是:如果受众不再共享同一件作品,艺术是否仍然有价值。

  • 目前,他认为 AI 协作仍然需要大量人类品味和劳动,这是一个“甜蜜点”。他将手稿删减约 14%,很多时候不是增加文字,而是去掉典型的 AI 腔调和过度解释,从而改善原本平庸的章节。

  • 他偏好的工作流,是让模型充当专注的助手:读入整部小说,找出问题,再由他逐点决定。即便机器可以完成大量分析,作者身份仍然体现在判断之中。

  • 一项持续多年的中音谱号测试,展示了能力的非连续跃升。连续几代模型都无法正确处理5个简单音符,直到 Astra 出现;随后,它通过 Codex,几乎完美地将一份高度复杂的20世纪乐谱转换为机器可读格式,其中包括重升号、变音记号和连音线。Joel 将这一跃升描述为“一夜之间从 0 到 99”。

10. AI 循环找到了合理临床筛选器排除的结果

  • Daniel McKinnon 的动机来自5年的基因不确定性:他的儿子 Owen 死于一种罕见肺病,另一段妊娠也未能保住,随后3次全基因组测序都没有找到答案,直到 Warren 出生并健康存活13个月。在那次妊娠期间,Daniel 援引 HIPAA 赋予的权利取得原始数据,并搭建了解读流程。

  • 这套流程重新找到了 Owen 已知的病变:一段 91-kilobase 缺失,移除了位于相关基因上游 1 megabase 的增强子。原实验室的结构变异筛选器忽略了距离基因超过 1 kilobase 的结果;在人类必须审阅由约10亿个、每个长150 base pair 的片段组装而成的嘈杂新一代测序结果时,这个决定是合理的。

  • Daniel 将 o3 放入一个循环中:模型先检查编码区,每次没有命中后继续搜索。临床团队受到病例处理量限制,而机器智能可以持续、并行工作,直到找到这个远端调控区域的解释。

  • 他认为,基因组解读尤其适合通过 agentic 系统改进:任务周期长、结果可验证,而且尚未饱和。罕见病例可以成为系统反复爬坡解决的“奥林匹克级问题”,而人类临床医生仍然被时间压力所困。

11. 不确定变异需要湿实验室证据,而不是又一个静态预测

  • 大多数最难的重新分析病例,最终都会落入意义未明变异状态。按照 ACMG 体系,6分可以将一个变异推至“可能致病”,从而打开治疗和保险选项;等待另一个匹配患者可能需要数年,而一次功能研究可以贡献2-4分。

  • 这些实验需要具有生物学相关性的系统。对于肺泡毛细血管发育不良,Daniel 提到 IMR90 胎儿肺细胞,因为相关基因只在发育第16至20周左右表达;将变异植入通用癌细胞系,无法回答同一个问题。

  • Gamo 买下 Arpeggio 剩余的业务,招募4名团队成员,并在3周后开始第一次实验。眼下的产品可以是面向临床实验室的查询服务,但更深层的目标是获取“来自真实世界的 RL 数据”,闭合模型预测与实测生物学之间的循环。

  • 测序本身已经“足够便宜”:每个基因组低于 $1,000,有时约 $500,在高通量实验室甚至低于 $100。稀缺的仍然是解读能力。在 RareBench 上,传统变异排序工具 LIRICAL 得分约 10%,而 CloudCode 中的原生 Opus 5.5 得分约 50%;新系统的能力也不止于排序,还能解释结果、排列患者优先级,并提出治疗相关性。

12. 垂直 AI 依靠路由、工具和评测取胜,但这一层正在变薄

  • Gamo 服务于两类现有工作流:医生收到不充分结果后的在线重新分析,以及对数千个存档基因组进行批量重新分析——罕见病中心没有足够人手重新审阅这些数据。机器主导的系统一周可以处理80、100或200个病例,而受限于人力的人工流程需要6个月或1年;公司成立4个月后仍处于尚未产生收入的阶段。

  • Daniel 将公司描述为“一家做模型编排、工具和评测的公司”。Claude、Astra,甚至 Gemini,在 RareBench 的不同集群上表现各异,因此智能路由和集成可以同时改善成本与性能,而不必在基础模型层面展开竞争。

  • 轨迹分析能抓住演示数据无法暴露的失败。Groq 4.6 曾经因为悄悄重命名一个基因而漏掉病例;Astra 曾在被要求处理的一篇论文中漏掉答案,另一次评测中则因为某些工具让它陷入过度思考而停了下来。Gamo 与每个模型共同设计工具,以可靠且可负担的方式避免这些“细小但持续的擦伤”。

  • Daniel 也承认,垂直层正在变薄:原生 o3 起初得分为 0%,他称此后的提升是“20或30个百分点之类”。前沿实验室员工一直很愿意提供帮助,Daniel 也报告了许多此前未确诊的儿童如今已经确诊;但调动一家万亿美元公司的方向,与个人层面的热情并不是一回事。Nathan 最后的保留判断兼顾了两面:放慢前沿进展可能仍然明智,但约 50% 的 RareBench 成绩仍留下关键问题,使克制成为“代价高昂的妥协”。

完整逐字稿
Nathan Labenz

Real things are starting to happen. This week on AI in the AM. On AI risk and verification, we heard from Jeremie Harris and Edouard Harris of Gladstone AI. They remain hard-boiled realists about US-China cooperation, but this week I heard a slight thaw. Here, Edouard considers what reciprocal transparency could offer the 2 countries.

Edouard Harris

Certain kinds of transparency can be stabilizing. The kind of transparency that goes like, “Hey, we're giving you enough vision into what we're doing to see that we are not doing the thing you fear most”—that sort of thing—is potentially useful. Additionally, it may not actually be that costly for us to do, depending on how we implement it, simply because the Chinese are already all up in our systems. Really, we're not giving anything away that they don't already have in many cases, potentially.

Nathan Labenz

Steve Hou, head of research at SiliconData, which builds GPU price indexes. We asked why renting apparently identical chips costs so much more at the big cloud providers.

Steve Hou

So, in the case of hyperscalers, indeed, you are observing correctly: they charge regularly, consistently, at least 2 to 3 times, sometimes more, compared to a typical neocloud. The reason has to do with a long legacy of whether it's other types of products being offered on their platform for software, analytics, safety, and compliance; the fact that they already have this long-established relationship with enterprise users that have been on board for a long time, that have a certain stickiness for moving. It is being sold as very much of a differentiated product.

Nathan Labenz

Prakash Narayanan, my co-host on AI in the AM. We spent part of the week on OpenAI DevDay. Here, what changes for building software when the models get faster at the same level of intelligence.

Prakash Narayanan

With ultrafast, as you type, you can interact. It's an interactive kind of software build, interactively building games. I think that's really the future. I think the speed—the latency at the same intelligence—is probably something that's going to be very important, especially as you clear these hurdles of capability. The thing can build a game, but can the thing build a game with you in the moment while keeping you in flow?

Nathan Labenz

Joel Borgen co-wrote his novel with AI models. The text has a few AI tics, but the book is legitimately good. Here, what he supplies as architecture and why the prose still needs him.

Joel Borgen

I started the project knowing more or less what I wanted to have made. I sketched it out myself; especially the first half or so of the book was pretty well set before engaging the models. The models are good at certain things. They're getting better at everything. But as far as just prose writing itself, even if you tell it exactly what you want and you have a plan for a chapter, you often get something that's unreadable unless you know how to scaffold it and then edit it yourself.

Nathan Labenz

Daniel McKinnon founded Gamo Labs, which uses AI to interpret genomes. He reports new diagnoses in children whose cases had gone unresolved. Here, what he found when he looked closely at a model's work.

Daniel McKinnon

And you'll see things like—there's 1 particular case where Groq 4.6, which at that point was state-of-the-art on our benchmark, in GroqBuild, missed 1 because it just renamed the gene. It was saying, “Oh, NRF2 or whatever is responsible,” and then it just changed the name of the gene to something totally different. I'd never seen that before, and I was just like, “This is dumb.” Our harness and our tools prevent the agents from doing dumb things.

Nathan Labenz

Welcome to the AI in the AM weekly highlights, with these introductions spoken in my cloned voice. Please tell us what worked and what did not. Your feedback helps us make the next one better.

Part one: A slight thaw. Jeremie Harris and Edouard Harris have been interviewing diplomats who negotiated with China. We discussed the Trump-Xi talks and the possibility of a channel for communicating about AI incidents. Edouard starts with what he makes of that contact.

1. AI Incident Channels

Edouard Harris

Talking is always better than not talking, so that's a positive. My understanding, from at least the beginnings of the Trump-Xi conversation and the stuff leading up to that, is that 1 of the things that may be positive that came out of that relationship was the development of this—I don't know if you'd call it a red phone, but at least some kind of theoretical line between the 2 governments on AI incidents and AI risks.

The report that we came out with, which is really just a long newsletter, is informed by speaking to about a dozen State Department diplomats who have dealt with China from the negotiating table and who've seen how these things develop in practice. One of the issues that they do see, among many others, is that we have tried the red phone thing before in the context of nuclear weapons, and by and large, they don't always answer.

Particularly in critical phases, it's often exercised as a point of leverage, saying, “We're going to take away this phone line and not answer,” rather than as something that is a collaborative, unified project that makes everyone safer. Again, this is not to say that if the CCP does structurally take AI seriously—which there is decent reason to think that they may—we could have a whole different and much more positive level of engagement. All that we're recommending, based on the experience of these folks, is being realistic about it and having a backup plan.

Nathan Labenz

I asked Jeremie, after a serious AI incident, what could the United States ask China to stop doing? And what could be verified with the capabilities that exist today?

2. Verifying AI Compliance

Edouard Harris

Everything that comes after this sentence obviously has not made contact with the intelligence community from a red-teaming standpoint. So the true answer is, we can't know deeply. I couldn't give you an answer to a level of detail where it would be like, “Okay, that's executable.” 1 easy thing, if I'm going to caricature—

Jeremie Harris

Data centers put off a hell of an energy footprint. The thermals on those are really, really bright. Data centers are huge. They haven't, by and large, yet been built to be hidden, and it takes a long time to build data centers.

Now, this will change. AI 2027 talks about the timelines for this. We think it's quite plausible that the timelines could be a lot shorter for hiding data centers, just based on conversations with folks in the industry. But whichever way you slice it, you're going to have an initial conversation where, to first order, for the 80/20 that you really need, it's like, “So help me God, if I see a cluster and that cluster is yay big...” That's the kind of conversation that you're looking at. And how you quantify that is a matter of the sort of national technical means that the US currently has—

Prakash Narayanan

So wait, are you saying having a cluster above a certain size would be a red line?

Jeremie Harris

So I'm saying—

Prakash Narayanan

Initially.

Jeremie Harris

Yeah, initially.

Edouard Harris

Yeah. So if you're just in that panic moment, right? You're like, “We have to do something,” and you ask yourself, “What is possible? What is possible to do with the existing assets and infrastructure that we have today, nothing else?” Then you do get into a space not necessarily where there exists a cluster of this size because you're not—you can't ask them to tear down the cluster—but we have to see the heat signatures from this go away.

So if there is a running cluster above a certain size, that is, to be clear, a tremendously expensive ask in either direction. The depreciation on GPUs is the major part of the OPEX costs.

Prakash Narayanan

So you're saying an incident happens first—

Edouard Harris

Yeah.

Prakash Narayanan

—and the response to that incident, the mitigation for that incident, is, “Hey, can you turn off this big data center?”

Edouard Harris

Yeah, like, turn off any cluster above a certain size. That's 1 possibility because this is—

Jeremie Harris

As of right now, to be clear, the framing is basically, as of right now, that's where we're at. And so, in some sense, we're going to get into a potentially circular loop here where you can see how big of an ask that is. That is an insane ask.

Nathan Labenz

So if I sketch out the logic from beginning to end here, it's like: we don't have that great of a relationship. AI capabilities continue to progress at a fast pace. We expect something crazy to happen. When something crazy enough happens, we're going to find ourselves by default in a spot where we have to ask for some outlandish, super-high-cost move, like shut down all your big data centers, because we don't have any other mechanisms in place that allow for a lower ask, a better trust-but-verify type of environment, because we haven't made those investments now.

That leads me to the question of what should we be doing now to, A, ideally not end up in that situation, or B, if we do end up in that situation, have better options available to ask for aside from shut it all down, which is obviously going to be tough?

Edouard Harris

Develop—basically, you absolutely nailed it. It's develop better verification and develop better offensive options to ensure compliance in the event that verification returns, “No, they're doing it.” The better verification stuff you can do, the faster and the less you have to rely on absurdly expensive things like this: “A gigawatt of energy radiation shouldn't be visible from space.”

Jeremie Harris

If we have techniques like this that are vetted by the intelligence community, even if they are just 50% better than this—even if it’s like, you have to shut down half your data center. I’m making something up here. You have to shut down half your data center, and we can sufficiently verify the other half, or something equivalent to that—you are saving billions of dollars right off the bat.

The ability for this industry to continue to make large amounts of money is actually going to be gated for that period of time by these little verification technologies, and there’s already this community of little verification startups that’s working on this technology. That’s why this is so important. Of course, I will also say the offense side of things is critically necessary. If you don’t have those offensive options, you cannot assure compliance.

You can verify and monitor the situation, and you can say, “Oh no,” but fundamentally, your hands are tied. You don’t have the tools to actually do anything about it. Both of those things are super important.

Jeremie Harris

There’s this bottleneck that I think a lot of these verification companies haven’t necessarily priced in, and this is where a lot of our current work is focused. Imagine what happens when Company A goes, “I have the thing. This thing is going to work,” right? It’s a moment of crisis, and Trump—or POTUS, whoever it is at the time—is casting about for options to alleviate this trillion-dollar bottleneck.

Then they actually go, “Okay, the intelligence community has to now vet this,” because they’re not going to just start using it, obviously, right? How long did it take similar technologies in the past to get used, to get vetted, and to become what’s known as national technical means, or NTMs, right? The answer is years, depending on the technology. Very often, years.

We have to do it. China has to do it. We have to handshake on doing it. Even if you remove money as an obstacle and it’s no object, there are just certain things that take serial time to do. A lot of what we’ve been doing is focused on saying, “Okay, treaty”—or not treaty—“let’s say AI agreement verification or compute verification. Company X, you probably should be talking to IC element Y about this.”

In a moment of crisis, you want as much pre-vetting as possible, and anybody who’s involved in assessing a potential national technical means had better have on speed dial—the Signal, the phone number, the email, whatever—of the founders of all the verification companies that they plan to use or may end up having to use. You want to cut down on all those barriers.

The boring bureaucratic hurdles that nobody ever thinks about because they’re boring and bureaucratic—these are the things that we’re trying to shatter right now so that when game time happens, things move more quickly. Today, those companies that are potentially pursuing research trajectories or agendas—in some cases, multimillion-dollar, like your $10 million-plus research agendas—that are oriented in a way that an appropriately placed person in the intelligence community would look at and be like, “That’s kind of a nonstarter,” we want them to get that feedback as soon as possible so they can reorient toward things that do have a chance of working.

Nathan Labenz

How important do you think it is that we remain on the same fundamental tech tree or AI paradigm across U.S. and Chinese AI development? My sense is that we’re in some ways in a very fortunate position right now because we’re basically building the same tech in the same way, and we’re sort of encountering the same surprises along the way.

The other question is, I feel like if there’s anything good to be found in the OpenAI-Anthropic adversarial dynamic, it would be that maybe they can be a testbed for techniques that might later scale to a U.S.-China dynamic. If I were the president, I would say, “You two have to figure out a way to police each other.”

Maybe we expand that circle to a few other frontier companies. But you two, you’re the ones setting off all these alarm bells. I need you guys in a room. Whatever technology you need to develop, whatever access you need to give one another, it’s on you to figure out a way that you can trust and verify one another.

Then maybe we can scale that up to a trans-Pacific dynamic that could work similarly.

Jeremie Harris

I think that’s not—

Nathan Labenz

What do you think?

Jeremie Harris

That’s not insane. Obviously, there are big differences between—

Nathan Labenz

Yeah.

Edouard Harris

—what Anthropic and OpenAI respectively have on each other. Although poaching of personnel does a decent job of mirroring the kind of access that China clearly has to the frontier labs anyway. But there—

Edouard Harris

There’s also a basis of trust, I think, between two fundamentally U.S. companies with fairly similar values and blah, blah, blah, versus fairly radically different ones. Not to say that this is totally a miss or whatever, but there are going to be some differences as well as some similarities. I think the similarities might be worth mining. Yep.

Jeremie Harris

Yeah. It’s also the case that you’re talking about the importance of the stacks being aligned. The hardware lottery does a lot of really good things in this space. I think the most crucial thing is that the U.S. and China clearly don’t trust each other in terms of the motives that bring each respective side to the table, right?

When China sees the U.S. come to the table and raise issues like slowing down AI or safety guardrails, whatever, the interpretation that we’ve heard consistently from people involved in Track 2 or Track 1.5 kinds of dialogues—and, Nathan, you’ve been kind of in that ecosystem or touched it as well—is that the Chinese view it as an attempt to curtail their own development because they see themselves as being behind. They’re justified, therefore, in doing things that even wouldn’t be appropriate for America to do in their eyes just to catch up, because they’re in second place.

Being in second place with respect to scale also means that you don’t see the warning shots with the same resolution. However, they seem to also be more public than at least I would have expected. The spillover is something that’s nice, because China can verify themselves directly.

That might be a mitigator if you see the same kinds of failure modes emerging from whole-brain emulation, or if some completely wacky other branch of the tech tree were to become dominant in China. I do think it’s good. I think we kind of get there by default. It’s hard to imagine alternatives that really shake things up at this point.

Jeremie Harris

Quantum machine learning—if you wait long enough, I just don’t think that’s going to be relevant on the timescales that matter and could radically reshape algorithms. But what we’re seeing right now is an industry that’s more or less doubling down on transformer MoEs, with some bells and whistles and some variations here and there, but everything is kind of a transformer, and that’s what seems to ship. So, yeah, I expect that we will have the benefit of that.

It also comes with the benefit of being able to share safety technology, as the U.S. did with Russia during the height of the Cold War at times. In principle, that does mean that we can work on each other’s safety stacks, and that might be the source of some trust-building measures, though that term is also problematic for China, and they’ve pushed back on attempts to do that sort of thing in the past. So, I think it’s a good thing. I don’t know how far it goes.

Nathan Labenz

With the Harris brothers, the discussion returned from verification technology to the diplomatic evidence for cooperation.

Nathan Labenz

I want to not be naive, but I do want to notice and give appropriate weight to positive signals as I see them developing. I think Xi’s speech at the WAIC was pretty friendly and conciliatory. He gave credit to the U.S. for inventing AI. He certainly didn’t call for an international arms race, and he warned against overstretching the national security concept.

We can dismiss that as just nice talk. We probably should have at least some weight on that possibility. But what are the meaningful things that you are watching for—the decision points that will update your thinking on whether they are inclined to at least try to control AI for their own narrow self-interest? Are they inclined to meaningfully cooperate, or are they inclined to seek some sort of domination, as we often project that we are interested in doing?

Edouard Harris

So, in terms of what signs to watch out for, I think anything that looks like a positive sign is at least a positive sign to some degree. A conciliatory speech is good. It’s good. At least it’s not a hostile speech, right? It could be worse. Everything is tempered by the fact that rhetoric is often used for strategic purposes in this way and to shape the battlespace in terms of the narrative and so forth.

The point is not that these are positive signs. It’s just that we have to weigh the evidence in the context of the credibility that has or has not been established by this entity over the past span of time.

In terms of slightly more unfakeable signals that they could give off that would make me go, “Oh, whoa, okay, this looks legit,” you can imagine maybe something like DeepSeek and Zifu coming to some sort of pacing agreement because some crazy thing happened over there—their equivalent to the Hugging Face incident. You can imagine the cyberspace commission that has jurisdiction over whether models can be released and whether they’re properly ideological, actually putting the brakes on something because there was some misalignment thing, and they’re finding that a model that was properly ideological in testing suddenly isn’t being ideological in the wild or something like this.

I would say indications that seem genuine—that they are starting to be on the receiving end of these incidents at the same level of detail as our own frontier labs—I think that begins to make all of us as humans go, “Oh, you know, maybe the thing we should be concerned about is the giant shoggoth thing that we don’t understand and that we’re growing in the labs at an accelerated pace.” That would make me feel a little safer.

Jeremie Harris

Well, and maybe procedurally too. If you take—we have a list of, I think, 5 different historical traps that we’ve seen in U.S.-China diplomacy. These are essentially a brief catalog of the ways in which China behaves when they’re full of shit, at least by the assessment of a lot of the diplomats we spoke to. I should be clear, actually: There was a dissenting diplomat who felt that some of these things were much more sincere, including the use of the language of arms control, the objections over language, this and that. And that itself is the epistemic problem that we talked about earlier.

But basically, I would say, take each of those red flags and flip them over, and you get the corresponding green flag. So, if you don’t see arbitrary concerns raised about language that seems random, that’s a green flag. If you see engagement—and this is actually really important—from empowered people, arguably as we did with Xi, though again, you have to calibrate everything with the fact that we’ve seen this before in other contexts, it’s certainly not a red flag.

But in terms of concrete commitments, that’s the sort of thing you look for: empowered people, the lack of capricious, arbitrary objections. Things that actually look like they’re making qualitative progress are a surprisingly good sign because you can contrast them directly with how things have gone in the past, which is not very good. I mean, the contrast point is actually that low that these can be genuine signals.

Nathan Labenz

What are the things that we can do that are not so costly to us but are still credible signals to them that we are not going to try to use AI to gain a decisive strategic advantage and ultimately make them an offer they can’t refuse? If indeed that is not what we’re going to do, which I’m a little worried we might actually be about to try to do. If we were on the path of trying to seek a Pax Robotica, where we can all benefit from the abundance that AI, especially in its Chinese-manufactured, embodied form, might provide for us, what would be the steps that you would prioritize next on our side?

Edouard Harris

Well, there may be some stuff we can do that’s not functionally really even that costly. Generally, as I think Jared and you guys maybe as well highlighted earlier, certain kinds of transparency can be stabilizing. The kind of transparency that goes like, “Hey, we’re giving you enough vision into what we’re doing to see that we are not doing the thing you fear most”—that sort of thing is potentially useful, and additionally may not actually be that costly for us to do, depending on how we implement it, simply because the Chinese are already all up in our systems.

So really, we’re not giving anything away that they don’t necessarily have already, in many cases potentially. So maybe some kind of visibility into, “Here’s what we’re doing, here’s what the frontier labs are doing,” and so forth. The problem is that’s not necessarily something you want to be doing unilaterally. That would come as part of a trust-building measure, dare I say, between two powers in the wake of a moment like this.

Goodwill-type stuff we can do now would certainly include more interactions between the verification communities in the United States and in China, and this kind of thing is already happening, actually. There are some quite good and positive interactions between those communities. So, yeah, the kind of linkages that you get at the level of academic-to-academic, startup-to-startup, all trying to solve for the same mission—these are very, very positive things.

It’s true that right at the political levels, the 2 countries have started separating out, and even at the level of big companies and stuff like that, you see the classic Chinese spinoff story and all this stuff. But there still are real, genuine linkages between the 2 countries, especially on the academic side and at a number of other levels. There are a bunch of sincere people just talking to a bunch of sincere people about, “Yeah, this is a problem and this sucks. Yeah, I agree. Let’s try to solve it.”

So the more of those linkages there are, the better. The more people there are on both sides of the ocean who have the ability to talk to their own domestic leadership and say, “Look, I’ve spoken to them. They’re not evil. They’re just trying to do this or that,” to whatever extent that’s true, that kind of moderates the more extreme tendencies on both sides. It’s very, very hard to do that completely because the actions of these countries are also constrained in a number of ways, but it really does help, I think.

Nathan Labenz

Part two: The economy of intelligence. Steve Hou leads research at Silicon Data, which builds price indexes for rented GPU capacity and model tokens. Turning compute into a measurable market means deciding which prices can be compared. Steve starts with the inputs to the GPU rental index.

3. GPU Rental Economics

Steve Hou

We use a combination of both quote prices and transaction prices in our calculation, making that distinction clear. This is a large normalization process, using machine learning to help us make the contracts apples-to-apples comparable, to the extent that you're looking at a single chip—let's say an H100—being rented from different parts of the world.

One question that jumps to mind right away is that, for example, in the case of an apartment rental index, you would never consider using just an offer price. Someone who lists a number on the front of a building, saying you can rent an apartment for $2,000 a month—you wouldn't expect to just trust that number. You can walk in. But we do use quote prices. Why is that?

The reason is because, unlike an apartment, you cannot click an API button and just get hold of the apartment. You have to walk in and talk to somebody and negotiate, and that apartment may or may not be available. In this case, very often, with API-executable GPU rental, you can actually get hold of the GPU, the same way you can buy something on Amazon, right?

That being said, we also have a transaction, so we're comparing them. If someone who has 3 nodes of GPU rented out, say, all 3 at 275, I will not presume that I can go back to the same merchant to rent another one at 275. It could very well be the case that the next one is not available anymore. Or if they rent out 2 out of 3 for 275, the next one will not necessarily be 275 again either.

It could be $4 or $6 or $1, depending on what everyone else is quoting. The person quoting that price would be crazy if everybody's quoting at $4 and they continue to rent at $2.75. You would think either they're going to raise the price or there is something wrong with that price, right?

This is the reason why I gave you a long answer again. What we do is that we want to provide as broad a set of coverage of the market as possible and normalize everything so that we're capturing the market as it is, as faithfully as possible.

Nathan Labenz

One big difference I noticed in prices—

Steve Hou

Yeah.

Nathan Labenz

Just browsing the Silicon Data website is—

Steve Hou

Mm-hmm.

Nathan Labenz

Between the neo-clouds and—

Steve Hou

Yes.

Nathan Labenz

The hyperscalers. These are not small differences. These are multiple differences in prices, seemingly 2 to 4×. That's a pretty big delta. Why does that delta exist? Is there no way to arbitrage it? And does that imply that GPU hours are being wasted, or that when I buy from a hyperscaler, I'm sort of preempting their internal work and they're using everything that's not sold at runtime? Give me a little peek behind the curtain there.

Steve Hou

Again, I like to use analogies. I gave you an analogy using a single-bedroom apartment. I'm not going to use an apartment again, although I can. You can, by the way. Let's say we're doing a single-patty burger index, right?

You can buy a burger from a burger stand off the street of New York, or you can walk into a high-end steakhouse and order a burger. Believe you me, those 2 burgers are going to cost very different amounts, right? They're both burgers, and what we try to do is, as much as possible, find the marginal price for a unit of compute—or a burger, in this case—that has the same relative features.

I wouldn't want to compare the price of a 3-patty burger with a single-patty burger, right? But once I make those adjustments, there are some adjustments I cannot reasonably make because it's capturing a different type of premium from product bundling or product differentiation.

In the case of hyperscalers, indeed, you are observing correctly: They charge regularly, consistently, at least 2 to 3 times, sometimes more, compared to a typical neo-cloud. The reason has to do with a long legacy of other types of products being offered on their platform—software, analytics, safety, compliance—the fact that they already have this long-established relationship with enterprise users that have been on board for a long time and that have a certain stickiness for moving. It is being sold as very much a differentiated product.

Going back to that single-patty burger index analogy I gave you, I see people who sell a single-patty burger with fries and a milkshake. I can try to strip out the prices of those 2 elements and isolate what I think would be the price of a single-patty burger from that merchant.

But if they, let's say, blended up the fries and the milkshake and injected it into the patty and said, “This is a brand-new product,” and charged 2 or 3 times the price, I can't very easily strip it out, right? At which point I say, “Okay, I'm raising my hands. You guys are a little bit different of a beast, and let me put you in a different category,” right? And that's how we have so far handled it.

We believe that this way, you are getting to a much purer form of a single unit of compute, right? That is actually getting you closer to this idea of fungibility, to the extent that things are direct substitutes for each other.

Nathan Labenz

The same chip can perform differently in different places. And I think, according to your site, you have some software that runs on the chip—

Steve Hou

Silica Mark, yeah. Mm-hmm.

Nathan Labenz

Yeah, that you—

Steve Hou

Yeah.

Nathan Labenz

That you benchmark the chips with. So I guess every single price in your index has been benchmarked by this system?

Steve Hou

No. We have a physical benchmarking service called Silica Mark that actually visits individual GPU at the EU ID level to try to assay—to assess—the GPU's health, performance, throughput, TFLOPS, and so on. At the moment, that physical benchmarking and physical spec performance does not enter into our pricing, right?

Nathan Labenz

I see.

Steve Hou

We do not see a strong relationship, at least at the moment, given the nature of the market, with how specifically the GPU is performing.

Steve Hou

Even though we do see that in the cross-section, you can actually have a bit of variance from GPU to GPU. This is not surprising or new to anyone who is in this space—the GPU lottery, right? But over time, if you have a cluster, it probably averages out, and the law of large numbers kicks in. So we don't use that. We have 6 features, including geolocation, CPU, memory, various other things, term, and provider. But the physical spec is not one of them, at least at this very moment.

Eventually, when we head toward a scenario where we potentially could have physical delivery—because inference, based on my thesis, could make compute more interchangeable—that could enter into our pricing scheme. But at the moment, it does not.

Nathan Labenz

Another price that I noticed had moved on the website and that caught my attention is the proprietary LLM index.

Steve Hou

Mm-hmm.

Nathan Labenz

There, you break down token pricing into—

Steve Hou

Open and closed, yeah. Overall.

Nathan Labenz

Right.

Steve Hou

Yeah.

Nathan Labenz

And so the proprietary one is up 11.5% over the last 7 days.

Speaker 2

Mm-hmm.

Nathan Labenz

And I guess I'm wondering, how are you measuring that? Because, quality-adjusted, everybody would say prices are coming down, right? Even dramatically so. The retail posted API price hasn't changed, right, except when they introduce new models. So what are you measuring on a day-by-day basis that allows you to say what the proprietary cost is doing at such a fine grain of resolution?

4. Tracking Token Price Changes

Speaker 2

Yeah. So first of all, I want to give a little bit of context for what our token indices mean, because I think they've repeatedly been misunderstood. Partly, I think, it has to do with the unfortunate naming. We called it the expenditure index, and then people maybe thought it either meant price or total volume, when it doesn't actually mean either. It's actually an expenditure-weighted price index that's normalized to 1 million.

That can show trends based on usage mix, which I'll come to in a second. When you point out the proprietary LLM model, which is sort of the frontier model, having recently bounced a little bit higher, we should put it in the context that since maybe late June through basically the middle of this month, it has been on a very sharp downward trend, going from some $4 to $1.60 or something. That's more than a 50% drop, right?

During this time, we've seen a lot of frontier leading labs not just cut prices on their newer variants, but also release cheaper variants of powerful models and new-generation models. We've also seen other proprietary models coming out, like Meta and Groq. Don't forget, these are proprietary AI models as well, right? They've been very aggressive on the price front.

So more recently, I think the bounce could come from a variety of sources. If people decided to say, “Okay, I actually quite like the more expensive, powerful model from Anthropic,” and they used more of it, that could drive up the expenditure-weighted price index.

The analogy I'd like to give people is: Forget that these are all LLMs. Imagine these are cars, right? If you have a Mercedes with cheap and expensive variants, depending on which cars are being sold more and which people like more, that could affect the average price of the cars sold. It's the same thing here, right?

In the market for tokens, there are 2 things happening. Token model prices are changing, but usage behavior is also evolving. To the extent that we observe volume from a handful of these public inference platforms, serving platforms that allow you to look at how people use different types of models, I think this most recent bounce of the frontier proprietary LLM index—I haven't looked into the details, but I suspect it has more to do with usage mix than anything else.

Nathan Labenz

From measuring compute to using it, the next conversation is just the 2 of us discussing OpenAI DevDay. I start with Dots, the personal agents OpenAI presented, and who might want a system they don't have to maintain themselves.

5. OpenAI Expands The Platform

It was funny. When they led off with the whole Dots thing, I thought, “Well, I think I have all of this, and I'm pretty sure I'm going to continue to prefer my version that I've gradually evolved over the last 9 full months now.” So that definitely didn't really feel to me like a developer product. I thought that was a little bit muddled, because it very much felt to me like that's a consumer product, right? That's for ChatGPT users to use, and it wasn't entirely clear how that would be used by developers, if at all.

I haven't been down every last breakout-session video, so there's possibly some more that I missed. But certainly at the keynote level, it felt like that's a product they are offering on a first-party basis to their users, and I do think it will be really useful for people. I guess I would say my guess is that people are really going to love these Dots.

Certainly for my mom, on the other hand, I would say, “Go for it. Just use that.” It's probably pretty easy. You don't have to worry about taking on all this stuff yourself, managing your own database on your computer, or troubleshooting when things go wrong, even though the models are getting so good at that on their own. I do think this higher-level and more polished abstraction will probably be really good for a lot of people who don't care to learn a bunch of new tricks and just want to have this thing that they can delegate to.

Nathan Labenz

Prakash then turned to the models OpenAI presented at DevDay. These are his first impressions of their capability and speed.

Prakash Narayanan

They had GPT-6.1 Sol and GPT-6 Astra Ultra Fast. GPT-6.1 Sol is about the same price, or slightly lower price, than GPT-6 Sol. But it's basically Astra for the price of Sol, which is what they're calling it. It seems to be a very competent model. I've used it. It's better—definitely better—than GPT-6 Sol. GPT-6 Sol was only released 2 weeks ago, or 2 or 3 weeks ago. So the cadence of releases is stepping up.

GPT-6 Astra Ultra Fast—now you have… They used to have a fast mode, which is 2 times the speed. Ultra Fast is 8 times the speed. They demoed how Ultra Fast works. With Ultra Fast, you can interact as you type. It's an interactive kind of build-out, software build-out, interactively building games.

I think that's really the future. I think speed and latency—latency at the same intelligence—is probably going to be very important, especially as you clear these hurdles of capability. The thing can build a game, but can it build a game with you in the moment while keeping you in flow? I think that's what's coming up next.

Speaker 0

Still in our DevDay conversation, I turned to Sign in with ChatGPT, which lets eligible users bring their plan to participating apps. My example comes from Waymark, the video company I co-founded, and concerns the cost of letting a new customer try an AI product.

Nathan Labenz

The classic “try 1,” “try 7 days,” whatever—those are tried-and-true tactics that have become very difficult in the context of, “Oh, but I have to spend a certain amount on tokens for every new user to give them that decent experience,” especially if you have something that involves a decent amount of setup or profile processing.

With my company, Waymark, we're not by any means the most token-hungry business, but the first thing we do when you sign up is make a big profile of your small business so that we can then use that profile as an input to make video content later. That profile-creation process has come down in price. I don't know exactly what it is today. At one point, it was about $1 per customer, and it was like, “Okay, well, this does start to become material when you think about the all-in cost of customer acquisition, whether we can make this flywheel work, how fast we get paid back, and all that sort of stuff.”

But I think this is a great value-add to your ChatGPT subscription that you can now take around. Of course, developers will need to implement this, but it shouldn't take too long to tell your coding agents to implement it. So that's great for ChatGPT, great for users, and great for the SaaS companies.

I think this is a huge win, and I do think it's pretty pro-ecosystem. It seems to me that this is one way in which they can actually follow through on the promise of not trying to eat the world and instead trying to empower people to build cool stuff.

Speaker 0

Part 3: The craft of co-writing. Joel Borgen is the author of The Receipt Horizon, a novel set in a world shaped by advanced AI, which he wrote in collaboration with AI models. We discussed what he supplied as the author and how the collaboration changed the work.

6. Writing Novels With AI

Joel Borgen

I may have written the book eventually, but the activation energy required to get started when you have kids and a part-time job was limiting. I started it well over a year ago, and the tools at the time were not as developed as they are now.

There were a lot of limitations, especially around context length. You have the thread where you're trying to work on a chapter or part of the book, and you're constantly having earlier context drop out and needing to manage that carefully. I think it's getting easier and easier to make use of the tools to do something collaborative while still keeping it kind of your vision.

I started the project knowing more or less what I wanted to have made. I sketched it out myself, and especially the first half or so of the book was pretty well set before engaging the models. The models are good at certain things, and they're getting better at everything. But as far as prose writing itself, even if you tell it exactly what you want and you have a plan for a chapter, you often get something that's unreadable unless you know how to scaffold it and then edit it yourself.

I've been working on a follow-up book, actually, and the process has changed quite a bit as the models have become more powerful. The book is definitely a lot better for having done it alongside the AI, but I'm cognizant of the fact that AI writing will be controversial for a lot of readers, and they don't want to engage with it. I think the process made the novel a lot better, and it also had weaknesses that I didn't entirely account for.

Nathan Labenz

I had asked Joel about building a positive vision of the future into a story that still needs conflict, and about the scaffolding he gives the models.

Joel Borgen

Utopias often want to become dystopias in narrative form because you need to have conflict and friction, and it's kind of dull to just have everything work out and everybody's happy, of course. The fact that this was set after a big AI war, and humans are presumably locked into what they're able to do, means you see both incredible flourishing in terms of the technology that's available and the options people have, and also some really major downsides.

I do think so much of what you've talked about, and I agree, is trying to paint a positive vision for the future, and that is difficult to do in narrative form. I didn't want anything in the book or any character to be a mouthpiece for me. A lot of the stuff that I want to see created in the world is a part of the story, and I try to steel-man the people who would have different opinions about that as well.

As far as the scaffolding, I often write out a large-scale architecture for the story and a lot of beats, and then go back and forth with the major frontier models. I've primarily used whatever the flagship Claude model is, and then the ChatGPT Pro model from GPT-5 on, because that was the workhorse. It did the most detailed work. Claude used to be a lot lazier as far as how much it would follow up on.

I'm doing it mostly through the chat interface, which probably has advantages for me and also limitations. Context length, like I said earlier, was a huge unlock as that got longer. At one point in drafting the first one, you could give them the entire book, and every time a new model came out, you'd use it to try to make it better, up your game, and allow it to do more.

As far as the actual scaffolding for the story, you go in chunks often, and I'll draft chapter by chapter. When I was done with it, I got the advice to hire professionals for copy editing, art, and layout. I did that, and I learned a lot from doing it. It was very expensive and time-consuming and kind of slow, and there were definitely frustrations involved.

I've tried to extract what I've learned from that and what made the book work better. I've fed it to GPT-5 or o3 Pro and developed my own workflows for doing the same process with the second book. I think that'll be an interesting experience, to see how well that works.

Nathan Labenz

With Joel, we moved from drafting to judgment, where he still wants a human author making the choices, and to how he edits the prose.

Joel Borgen

As we move into a world where the AI systems can do a lot of what we do—a lot of what we thought was valuable, what we thought was a human contribution—you think about the abilities they have now. The reason I found the book valuable to write is that there was a lot of me in it. It took a lot of effort on my part still.

I think I've heard a podcast episode where there was somebody writing books on Amazon in the romance area, and they were mostly automated and AI-written. I think there were something like a couple hundred per year, so they were playing a scale game. A few of them would make a few dollars, and overall it was profitable. You can certainly do that, and the models are getting so much better that you can probably have customized artwork of the same quality as, or better than, the book that I created over many months, on demand.

What happens to the art at that point? Thinking of it as an art consumer, if there was a new movie by Ingmar Bergman or Stanley Kubrick or something—movies that I've spent many hours thinking about and consuming, and that are deeply moving to me—what would that world look like where there's just a huge plethora of art on demand and it's pretty much custom just to you?

We already have a shared cultural space that's dropping off. People are more fragmented in what they consume, and there's less unification across the cultural landscape. Is that art still valuable at that point if it's not something that we're sharing with other humans? I don't know the answer to that, but I think we're in kind of a sweet spot now where, in order to get a work that you're proud of while working with AI, it still requires a great deal of human judgment and work. I don't know how long that will be the case, given where things are headed.

I had my book edited down about 14%, mostly manually and with AI helping me decide what to cut. Often, you can take a chapter that's so-so and improve it by taking away some of the stuff that AI does particularly badly: the tics and the over-explanation. I think you could architect something now that could create a pretty decent book with the right kind of feedback loops.

A lot of the stuff that I've developed is based on my own preferences. I want to be able to give the entire novel to a model and have it come back with an analysis and a list of things you should consider, because doing it manually and slowly is extremely time-consuming, as you can imagine. Bringing things to your attention like an assistant and saying, “Here's something you could consider,” and then letting you decide for yourself, point by point, is a fairly satisfying way to work.

Nathan Labenz

Joel is also a musician. To describe a capability he had watched arrive, he turned from prose to musical notation and a viola-clef test he had kept trying on successive models.

Joel Borgen

I've had this held-out test for well over a year now, maybe 2 years even. It's very simple: a high-definition capture of 5 notes on the viola. I give it to the reasoning models and ask, “What is the clef? What are the notes? What are the note values, and so forth? What's the time signature?”

Not a single one of them got it right until Astra. OpenAI presumably put in some actual musical scores. It went from basically not being able to do anything to—now I gave it a complex 20th-century score that's almost painful to look at, given how complex it is. There are double sharps and accidentals and ties everywhere. It's difficult.

It did an almost perfect job turning it into a machine-readable form via Codex. This score that you basically just had as a PDF before can now be turned into something that you can manipulate and evaluate. It went from kind of 0 to 99 overnight. I think that's an interesting way to turn some of the scores that aren't machine-readable into something that you could use to train a system.

I would love to see Suno move in the direction of adding a lot of classical stuff to it, because I think, in addition to creating more interest in classical music, classical music requires a different level of understanding of the form because it's often much larger-scale. You'll have a 3-hour-long opera that has some internal structure, or even a 20- or 30-minute single movement, something that has a lot more going on than a pop song. I would love to see that happen.

It does tie into the same element of, if you're making just private art mainly for yourself to listen to, I still think it's valuable. But again, like with writing a book, I would want to do it in tandem with the hypothetical model in the future and create something where a large part of my effort and taste goes into it as well. I don't know how long that interregnum will last, where you have a role for humans and AI to create something together.

Joel Borgen

We use it as a tool, but it still requires a lot of you and your judgment. But I look forward to experimenting with that once the models improve.

Part 4: The diagnoses still waiting. Daniel McKinnon founded Gamo Labs to use AI agents to interpret genomes and revisit difficult, unresolved cases. His son Owen died from a rare lung disease. A human specialist later found the genetic deletion the original sequencing analysis had missed. Daniel subsequently built an AI pipeline that rediscovered it while his family was seeking answers during another pregnancy. That experience helped lead him into this work.

7. AI Finds Missing Diagnoses

Daniel McKinnon

It really leapt out at me last summer when—we blessedly have a healthy 13-month-old right now, but because of our history, we are being monitored very, very carefully. There was just something a little bit questionable in the 16-week anatomy scan that our own maternal-fetal medicine doctor, who is an absolutely wonderful person, said, “Normally I wouldn’t even flag this, but because it’s you guys, we need to look carefully.”

We did a whole-genome analysis of the fetus at that point, and it came back negative. This is actually the third negative whole genome I’ve seen between losing Owen and having our son Warren. We actually lost a second pregnancy due to genetic reasons very late. We know we’ve had basically 5 years of heartbreak before bringing Warren into this world, and all of it was kind of genetically mysterious.

As a family member of a patient, I started to learn a lot about the failings of this system. It was that moment where I said, “I’m heartbroken, but I’m mad, and I’m going to do something, and I have tools to do something.” Fast-forward to last summer: I got this result back and said, “This is unacceptable. This is the number-one thing that I want in my life.” I really wanted to have a family. I wanted to understand what would happen.

I didn’t believe these labs, and I basically called all the labs and said, “Give me my raw data,” which you can get due to our HIPAA rights here. I vibe-coded my own interpretation pipeline. I mean, now vibe coding is crazy, right? Right now, this would be so easy. You could just say, “Claude Code, make me this thing.”

Back then, it was still like the original Codex autocomplete. It did require quite a bit of work on my end, and I did it to see if there was any comfort we could get around the pregnancy. I was shocked that it also outperformed on these other cases. Most clearly, it diagnosed Owen when one of the best—or some might even say the best—prenatal sequencing lab did not. I knew at that moment that this was something I had to contribute to. I didn’t know it would be a company.

Nathan Labenz

The deletion in Owen was 91 kilobases long and removed an enhancer, a region that helps regulate a gene. The specialist had already identified it before Daniel recovered it with AI. We asked Daniel how the original analysis missed that enhancer.

Daniel McKinnon

How you miss it is that this enhancer is a megabase, or 1 million bases, upstream from the gene. What Rady did, as many other labs do—which is a very reasonable assessment—is use a filter for structural variants, which tend to be quite messy in what’s called next-generation sequencing. That’s how sequencing is done today, where you have roughly 1 billion 150-base-pair fragments and need to piece them into this clinical puzzle.

They said, “We’re going to have a filter, and anything more than 1 kilobase up- or downstream of the gene, we’re not going to consider.” This is a totally reasonable trade-off if you have humans looking at all of this stuff. But my thought was—and this was before Codex, before Claude Code—can you put the o3 model, which was the model I used at that point, into some kind of loop and have it keep looking?

It basically looped. The first loop was, “Is there anything wrong with the coding elements?” Those are kind of obvious things, and it was using these bioinformatics tools. This is very crude compared to what we have today, but then it misses something and goes through another loop and another loop, and it just keeps working.

When you are a clinical lab, whether you’re for-profit, nonprofit, or whatever your structure, ultimately you’ve got to move. You have to spend some amount of time on each case, and if you don’t come to a conclusion, you say, “This is nondiagnostic.” This is very common. Most whole genomes, even from infants suspected to have genetic disorders, come back nondiagnostic.

I really think it’s one of these meat-space problems. If we can export these problems onto a machine intelligence that can work nonstop and in parallel, then we will be able to see many more kids, treat many more kids, and do much more interesting analysis on top of the basic things that humans are just pressed for time to do.

I don’t want to claim I’ve reinvented this. People are trying to build software to accelerate genomic interpretation, and they’ve been doing this for a long time. But I think what I probably identified relatively uniquely, early on, was that this is a really great task for agentic AI.

From an improvement perspective, it’s long-horizon, agentic, verifiable, and unsaturated. I suspect that this will be a task like math, where we can just generate these Olympiad-level problems and keep hill-climbing on that until the problem is basically solved.

Nathan Labenz

Daniel McKinnon then turned to the unresolved cases his team is analyzing now. Sometimes the remaining obstacle is a variant whose effect is unknown. We asked where the new biological evidence would come from.

Daniel McKinnon

On the back end, once we’ve done the interpretation, the majority of clinical cases we see right now—and to be clear, we are only seeing hard cases, so this isn’t like most general labs—end up in what’s called a VUS state. That’s a variant of uncertain significance, and these are scored according to a very standard rubric developed by the American College of Medical Geneticists, or ACMG.

You need to get 6 points or more to be bumped into the likely pathogenic category, and that unlocks a lot of treatment options and insurance options. Let’s just say it’s good to be either benign or likely pathogenic. It’s very bad to be in this kind of intermediate stage.

If you are in this intermediate VUS stage, you can do 2 things to get a diagnosis. One is that you just wait, and you can get more points if more patients emerge who have a similar phenotype or a similar condition as you do, and the same genetic variant.

This is often what happens when you read in the news, “This kid has had epilepsy for 10 years. It finally got a diagnosis.” If they’re lucky, “Oh, and since we know that’s the diagnosis, we worked with a pharmaceutical company, and there’s some off-label use of a drug, and it can actually help their condition.”

That’s typically because they just waited until somebody else had the condition, but this is not scalable and takes a long time. Another fork is that you can convince a university lab to care about this problem. What they’ll do is what’s called a functional study.

They’ll make a cell line that’s emblematic of your particular condition. For example, for alveolar capillary dysplasia, that commonly used cell line is called IMR90s. It’s a fetal lung cell line, and this is very, very important because this particular gene is only expressed from week 16 to week 20 of development. You can’t just put them in a cancer cell line.

Then you can do something like edit the genome, insert a plasmid, or conduct any number of different functional studies to say, “Okay, in this physiologically relevant cell line, if I have this genetic mutation, is this gene broken in some way?”

Nathan Labenz

Daniel’s team had just taken over an existing biology lab to test uncertain variants. He described the experiments starting that day and the feedback loop he hopes to build between real biology and the models making predictions.

8. Building The Biology Feedback Loop

Daniel McKinnon

We basically bought what was remaining of Arpeggio, hired 4 people onto the team, and very quickly pivoted to creating this basically like RL data from the real world with biology experiments. We’re 3 weeks in, we’re running our first experiments today, and I’m really looking forward to being able to close that feedback loop.

I think a lot of clinical genetics labs are interested in this as well, since this service is not commercially available. We can not only sell a lookup-table version of this—“I have a patient with this variant. Can you help me get a couple more ACMG points to get this up to likely pathogenic?”—because you can get 2 to 4 points. Remember, you only need 6 points, so 2 to 4 points is a lot of points for a functional study.

But it also lets the machines learn. Right now, the machines aren’t learning. You get to a VUS, and that’s the end. Now it’s like you get to a VUS, you do the biology study, you say, “Oh, this region of the genome is actually quite important for this particular disease. We’re going to do that, we’re going to check these predictions, and then we can improve these predictions over time.”

Nathan Labenz

We had been discussing earlier genetic testing companies, including Counsel.

9. The Vertical AI Harness

Daniel McKinnon

I was like, “Oh my God, we sequenced the genome 25 years ago, and Bill Clinton and Tony Blair got up onstage and said, ‘This is going to revolutionize human medicine.’” I look back and there’s this graveyard of genomics companies. You mentioned Counsyl, which I actually would not say is a graveyard. I think they were modestly successful.

But why is all of health care not based on precision medicine? Everyone in this space is like, “It should be,” and there are examples. It’s just too challenging. The interpretation is too challenging.

I actually look back and I see—and we talked about Counsel ahead of time—I think what was missing prior to now was really that interpretation layer. It’s very, very complex. Genome sequencing got what I would call cheap enough maybe 5 years ago—I mean, below $1,000 a genome, maybe $500 a genome. There are high-throughput labs doing this for less than $100 a genome. You see announcements on Twitter saying, “Oh, I can do a genome for less than $100.” There are many people doing this right now.

The sequencing is not the problem. Given your problem, what insights can you derive from that? That was only possible as of last summer. I’d say o3 is the first example of that.

It’s not just the variant interpretation, right? It’s explaining to the provider why it’s important. It’s ranking variants in a nice way. It’s explaining how you can treat this person. It’s ranking different patients for ASO eligibility. There are many, many, many other things beyond just scoring the variants.

Although I will say we benchmark variant annotation in RareBench and other benchmarks we have. The best-performing traditional machine-learning-based approach in terms of variant ranking is called LIRICAL. I think it scores something like 10% on our benchmark, and right now Opus 5.5 in CloudCode, just a vanilla thing, scores something like 50%. These traditional tools are also getting blown out of the water by this newer approach.

But then they can do much more. You basically need to make a very, very cheap, easy thing that people who are not surrounded by fancy clinical geneticists, genetic counselors, or top-tier hospitals can use, and that’s the problem we’re trying to solve.

There’s already product-market fit for sequencing. Every baby in top NICUs is getting sequenced. Every baby in lower-tier NICUs would get sequenced if they had the resources. The problem is much better scoped: The baby has pulmonary hypertension; figure out why—not, is there something obscure that could be wrong with this baby now or in the future? That is where we are working with various hospitals and families today.

There are 2 forks to this work. One is what I call online. A case comes in; physicians do not get good results from the traditional labs. The patient consents to a reanalysis. They send the data to us, we reanalyze it, and sometimes, but not always, we find additional things that can help them make decisions around either a pregnancy or a newborn.

The second fork is bulk-scale reanalysis. These are rare-disease centers that have thousands of genomes, and they say, “There are probably kids that we can diagnose or even treat in that database, but we have a handful of genetic counselors, a handful of bioinformaticians, and we just don’t have the resources.” Because a machine does most of the work here, we can get a set of 80 or 100 or 200 or whatever and turn them around in a week. That’s kind of 6 months or a year’s worth of work that they really can’t prioritize.

Those are the 2 forks right now, and we’re 4 months in. We’re pre-revenue. We’re really not thinking about what the exactly appropriate business model is here. We’re just trying to diagnose more sick kids.

Stepping one step up, what do we do? We’re basically a harness, tools, and evals company. I think you’ll see this a lot in vertical AI in general. I’ve worked in these frontier labs. I started working on LLMs—actually, the first LLM project at Meta, which was OPT 175.

Unless you are somebody who is very famous with very deep pockets, I don’t want to bet against the frontier labs by any stretch. I think it’s possible, but you’re not going to compete at the model layer, and intelligence is progressing so fast.

We ensemble models for sure, and that helps with both cost and performance. If you look at the confusion matrix of RareBench—we also publish this—you’ll see that the cases kind of cluster. Claude is good at this, Astra is good at this, and even Gemini, right? Gemini is not on the frontier, but they’re kind of good at this, and you can ensemble these things together and get better performance.

I’d say that’s one thing we do. And with ensembling, there are very dumb ways of doing this, but there are also smart ways of doing this, and I don’t have to go into all the details here. I would say figuring out unique routing and intelligence per model is an important edge that I think we and many others are doing.

Another key thing that people forget—and this is what I worked on at both Meta and Google—is evals. Unless you have very structured ways to measure the performance of the system—it’s not just the model at this point; it’s the whole system—you don’t know what to hill-climb.

We do pretty robust trace analysis of how models perform with and without our harnesses. This sounds stupid, but not that many people actually look that closely at data. This is a meme on Twitter, but it’s very true.

You’ll see one particular case where Grok 4.6, which at that point was state-of-the-art on our benchmark in Grok Build, missed one because it just renamed the gene. It was saying, “Oh, NRF2 or whatever is responsible,” and then it changed the name of the gene to something totally different. I’d never seen that before, and I was just like, “This is dumb.” Our harness and tools prevent the agents from doing dumb things.

I think there’s a world where, if I worked at OpenAI or Anthropic and had access to infinite compute, and I could do millions of rollouts, the agents would, with enough compute, all converge on some answer and maybe mitigate all of this work—especially because they can write their own tools these days and everything. But really it’s about consistently and repeatedly and, honestly, affordably getting to the right answer.

For $10 a case, who cares? Or even $100 a case. But if you need to do a million rollouts and all of a sudden this is $1,000 a case or $10,000 a case, it doesn’t make sense. So we’re doing a lot more research around how those cost curves work and how everything works.

Ultimately, our job is to take the smartest intelligences in the world, mix them together, give them access to the best tools, and get answers for our patients. That’s the core of what we do, along with being able to measure whether it’s working. I’d say that’s the core of what we do.

Prakash Narayanan

When you talk about vertical AI, would you say you’re looking at benchmarks where you combine an intelligence with the tools that you have? How do you evaluate the strength of the tools on their own, excluding the models? You’re going to upgrade model families over time, right? How do you measure the performance of what you’re building—that layer in between, in particular?

Daniel McKinnon

Yeah, that’s a good question, and I honestly don’t have a great answer to that question, in that we are co-designing the tools with the models. We have seen cases where—I mean, this is a very dumb example, and if you’re getting deep into this and have a specific eval, you’ll start to see this.

If you’re not deep, it’s very easy to be like, “Oh, I just had Claude Code one-shot this video game. Was it good? I don’t really know. It seems amazing.” I don’t want to communicate that these models aren’t amazing. They absolutely are.

With one tool for a specific type of literature search, the model—I think this was Astra—said, “Oh, yeah, okay, these are all the papers you tagged for me to read.” The answer was in the paper, and it missed it.

We were looking through the trace and the context window, and we were like, “I don’t think you actually read these papers.” There are lots of little things like that. They’re little paper cuts, and that’s what these tools do. They show you, “Oh, you missed a case here because of this.”

You missed a case here because of this. And you kind of put it on these better rails. Actually, Astra is a good example. The first time we ran the eval for Astra, it just stopped because some of the tools we had made it want to think about too much.

You see this on X, where people are like, “I asked the model to do some research for me, and I came back 30 minutes later and it’s looking at flights from Dubai to Calcutta on Emirates Airline.” Why are you looking at that? They just do stuff like this.

So I think we try to co-evolve the tools with the models. Coming back to something that many people might find boring but I find fascinating, evaluating a model is very hard, and we need to have very robust evals that can catch these things early. In that case, we caught it early, updated how Astra called some of the tools, and it works again.

Nathan Labenz

Prakash asked Daniel what he most wants the frontier labs to improve.

Daniel McKinnon

Well, I want them to hill-climb my task, right? Everyone wins when these models get really, really good at clinical genetics. We probably have less to do at the harness layer, which, to be honest, I’m both bullish and bearish on vertical AI, in that there’s a lot of value in owning a customer relationship. I think OpenEvidence has shown this.

But that layer is also getting thinner. When I first started this, vanilla o3 would not do this task. In fact, o3 scored 0% on my benchmark. I re-benchmarked it just for fun. So it was like, “Oh, I need to do all this work and scaffolding to get these things to work.” And now there are 20 or 30 percentage points or something. I mean, it’s definitely shrinking.

But that said, my goal is to build this great AI-native clinical diagnostics company, and I want the best person to do the work.

Nathan Labenz

Prakash also asked how Daniel hopes to make those biology experiments cheap enough to run at scale.

Daniel McKinnon

It’s really an AI and robotics thing. You have an arm there that is doing 384 experiments at a time in a bigger plate. You have a cartridge over there with thousands of plates lined up. You have AI assisting with primer design and experimental design.

This is one of those things where you say, “Can you use AI for this stuff?” You have to get special permission, but you can. At very high throughput, you can design and order the primers, design the experiments, and run the experiments.

Then it comes to analysis as well. You end up with very large-scale data sets that you need AI to poke through. I don’t have some perfect explanation—it’s not just this one invention, but a confluence of all those things.

Nathan Labenz

Daniel also described the response to his diagnostic work from people at the frontier model companies. Here, he reports new diagnoses in other children whose cases had been unresolved. What was the process of talking to frontier model companies about your bio use-case experience like?

Daniel McKinnon

It was really positive for me because it’s extremely obvious why I’m doing this, and I’ve never met anyone who was like, “Why are you doing this?” Every single person I’ve met at any of these labs has been like, “I want to help.”

Whether wanting to help as a person translates into, “I can pivot this trillion-dollar company to work on your problem,” is a different thing, and I understand that. I’ve worked at these trillion-dollar companies too. But it’s been incredibly supportive, and I would be pretty surprised if at least the top frontier labs did not spend some amount of time on this problem.

Again, it’s meaningful and it’s good. Even selfishly for them, there’s a really negative AI narrative swirling around right now. You hear things like in Dario’s recent tweet, where he’s like, “We’re going to cure cancer in 5 years.”

We’re diagnosing kids today. We have many examples of kids who are undiagnosed whom we have diagnosed at a small company 4 months in, and that’s a great story. It should be told by them, even if it’s just for selfish reasons like, “Hey, people of the United States of America who are not happy with data center buildouts or uncomfortable with AI taking jobs, here’s a very concrete way where we are helping people today.”

I also think it comes from the individuals. People find meaning in their work, and I think there are a lot of people, especially mid-career in tech, who are like, “I’ve been doing some kind of ad ranking, data munging, whatever, for my whole career, and what you’re telling me is that you have a very concrete way where you can help kids, and I want to do that.”

So I think there’s a combination of the business reasons and the personal reasons that are driving people to this task.

Nathan Labenz

After that discussion with Daniel, I came back to what slower AI progress could mean for families still waiting for answers. The 50% figure I refer to is a score on RareBench, the company’s variant prioritization benchmark, rather than the share of patients diagnosed.

Intellectual honesty demands recognition that this is probably still one of the things that we will be trading off when we pace the frontier in the near term, to whatever extent that actually happens. I still think that’s probably a good idea and probably worth it.

I don’t think the frontier model companies are at risk of not being able to grow a business, not being able to grow revenue, or becoming surpassed if they pace their frontier efforts. But it is important to keep in mind that there are very real problems in the world that AI is climbing that hill on right now but has not finished climbing. This is one that everybody would obviously love to see solved sooner rather than later.

It’s important for me to stay honest with myself and everybody else that this is a very real cost when you’re talking about individual people and their families. It’s that 50% number that Claude gets to today. That obviously leaves half of the questions unanswered, and those questions are extremely meaningful to people.

So I do think we shouldn’t take that lightly, even if it is a trade-off that, on balance, I think we should probably be willing to make some compromises on. It’s a costly compromise. Thank you for listening. Please tell us what worked and what did not. See you in the morning.

Speaker 9

I carried a tune for 20 years and never had the hands to get it down. Kids asleep and the snow piled high, mine the only window lit in town. Then somebody took the empty chair, said, “Hum it once and I’ll play along.” I hummed 4 bars, you played it twice, and asked me, grinning, “Is that the song?”

Second fiddle, play along. Play me every note there is. You never tire, you never slow. I keep the ones that sound like me and let the others go.

Second fiddle, play it bright. I’ve hummed this tune for 20 years. I’m first chair now every night, so play along, play along, my dear.

Ah. All day I read the slides through glass. I know what’s living and what should go. You played the sunrise 3 times through. I kept the one with frost—if I would let you.

Six hundred pages asking more. I laughed and said, “Let’s find the ending,” and swept the extras to the floor.

Second fiddle, play along. Play me every note there is. You never tire, you never slow. I keep the ones that sound like me and let the others go.

Five little notes I pinned up on the stand. Two winters running, you just couldn’t play. Then one bright morning, you read the whole hard page, ties and flats and all, and I laughed out loud that day.

Hey, now listen to you read. Pull up close and let it sing. Second fiddle, take a bow. They say a revolution’s on. A cognitive one—well, how about that? I say, “Boys, just play the song.”

A stranger wrote me yesterday. Said, “Doc, that tune of yours is good.” I said, “It’s mine,” and you just grinned like any second fiddle would.