[BidClub_]
The Cognitive Revolution · · 90 分钟

用 o1-pro 猎基因:ChatGPT Pro 资助项目获得者 Brownstein 博士谈如何推理罕见病

Nathan LabenzDr. Catherine Brownstein

YouTube
TL;DR
  • Labenz 将罕见病定义为信息处理瓶颈,而不只是测序问题。 罕见病影响不到 1/2,000 人,或在美国影响不到 200,000 人,但合计患者人数“比美国的天然金发人群还多”。基因组测序的可负担性已提升近 10,000×,从 2007 年超过 100万美元降至几百美元,稀缺资源也因此转向解读能力。
  • 测序成本下降,反而让 Boston Children’s 接收的转诊病例更难,而不是更容易。 早期经过严格筛选的病例诊断率约为 80%;如今明显病例已在其他地方解决,Brownstein 团队接手的多是此前阴性的病例,诊断率降至约 10%。由于“新的东西一直在被发现”,工作流越来越依赖队列构建、新检测手段和定期重新分析。
  • AI 眼下最直接的回报,是释放专家时间,而不是自主诊断。 Brownstein 说,总结论文、基因、陌生技术和审稿意见“改变了我的人生”,包括在约 90 分钟内把 20 个候选基因缩小到 3 个。Labenz 认为,近期机会在于减少文献检索、命令行操作、工具割裂和重复分诊给稀缺专家带来的摩擦。
  • 更好的模型无法弥补不可访问或缺乏代表性的数据。 一个看似致病的变异,可能在很少被测序的族群中其实很常见;因此 Brownstein 认为,“我们需要测序全世界”,才能把疾病信号与背景变异区分开来。研究者激励机制、主动加入的生物样本库、彼此隔离的队列以及薄弱的数据共享默认设置,仍是主要的非 AI 瓶颈。
  • 当前最强的运行模式,仍是多学科人类专家与 AI 协作。 2014 年的一次竞赛中,23 支团队分析 3 个看似符合孟德尔遗传的家系,其中 2 个家系获得诊断;由临床医生、遗传咨询师、研究人员和生物信息学家组成的混合团队,表现优于单一背景团队。Brownstein 的核心观点是,“这一切都不是在真空中发生的”:AI 正在帮助多学科流程,而不是取代它。
  • 高风险领域的落地,将由验证要求和严于人类的错误标准决定。 Brownstein 仍会检查输出,却无法稳定获得引文,有时还会发现推理模型做出令人印象深刻但缺乏依据的额外跳跃;Labenz 预计,医疗 AI 可能需要达到自动驾驶式的门槛,即安全性约为人类的 10 倍。尽管如此,随着“推理变得越来越好”,她的态度已从怀疑转向确信。
  • OpenAI 资助项目正在测试:从分诊走向解决病例,最低需要多少数据和算力。 Brownstein 目前还没有把完整病历和基因组直接输入模型,原因包括成本、能耗、无关细节和跑题风险;她希望找到最小充分输入和可靠工作流。Labenz 预计,o3 级系统可能在一年内给出有意义的结论;Brownstein 则保留了必要的谨慎:“我当然希望我们能用它解决病例。我不知道到时是否真的能做到,但我希望可以。”
摘要 · 为研究而整理的核心内容

1. 罕见病合计起来并不罕见

  • Brownstein 开场时的纠偏构成了本期基础:“罕见病其实很常见。”一种疾病只要影响不到 1/2,000 人,或在美国影响不到 200,000 人,就可能被归入罕见病;但所有罕见病患者加起来,比美国的天然金发人群还多。

  • 分类取决于观察粒度。自闭症本身很常见,但由 KCNJ8 de novo 变异导致的自闭症很罕见;随着研究推进,人们对定义的理解也在不断变化,因此界定并不简单。

  • 因此,Brownstein 更愿意把自己称为“猎基因的人”,而不是某种疾病的专科专家,目标是诊断那些尚未得到诊断的人。

2. 测序成本下降,把瓶颈向上游推到了解读环节

  • Brownstein 2011 年入行时,外显子组——约占基因组 1% 的编码区域——成本约为 3,800-4,000 美元。她最近询价时,同一项检测已降至 160 美元;Labenz 补充说,基因组测序价格也从 2007 年的超过 100万美元降至几百美元,可负担性提升近 10,000×。

  • 当时的稀缺性迫使医生严格筛选患者。临床医生手里存着等待技术成熟、价格下降的 DNA 样本,测序主要集中在那些最可能通过父母子三联体发现明确 de novo 事件、提前终止密码子或大段缺失的病例。

  • 这种筛选带来了约 80% 的诊断率,几乎是“打桶里的鱼”。随着常规检测吸收了这些病例,Boston Children’s 接收的患者越来越多是基因组此前已显示阴性的病例,Brownstein 目前的转诊诊断率降至约 10%。

  • 核心变化在于经济逻辑倒转:生成一份序列已成为常规工作,而解释一份阴性序列,则需要越来越专业的劳动力、文献访问能力和跨病例推理。

3. 诊断之旅最终落在一支刻意多元的团队上

  • 许多家庭要在本地医疗机构和专科医生之间辗转多年,才能来到 Boston Children’s。Manton Center for Orphan Disease Research 接受患者自荐,收集病历和既往遗传学资料,再进行重新测序,或采用 RNA-seq、长读长测序等新方法。

  • 2014 年,Zack Kohane 提议开展一项计划,将 3 个看似符合孟德尔遗传的家系交给 23 支国际团队分析。其中 2 个家系获得诊断;总体而言,由不同背景成员组成的团队明显优于单一由生物信息学家组成的团队。

  • 成功团队中既有研究助理、遗传咨询师和研究人员,也有多种类型的临床医生。Brownstein 对 AI 部署的长期经验是:“让拥有多种能力的多学科团队共同工作,整体表现会好得多。”

4. 诊断是层层叠加的证据,不是一条神奇流水线

  • 新病例往往带着大量病历资料到来,如今还经常附带一只 U 盘里的基因组数据。Brownstein 会运行多条基因组分析流程,因为综合系统可能制造假阳性,而更简单的黑箱系统则可能在不解释原因的情况下悄悄删掉变异。

  • 表型会通过 HPO、ICD-9/10 或 SNOMED 等本体编码,再与不同层级的遗传数据结合,包括原始 FASTQ 文件、比对后的 BAM 文件以及处理过的 VCF 变异列表。流程会根据表型与基因的匹配度,以及预测的致病性,对变异排序。

  • 如果已知疾病基因无法解释病例,研究人员会继续寻找异常结构变化、易位、缺失、重复和受扰乱的基因表达。他们也可能加入表观遗传检测,或通过全基因组分析和 SCAT 等罕见变异检验,考虑多因素解释。

  • 即使从更宽泛的统计口径看,遗传检测阳性率也只有约 25%,意味着约 66% 到 75% 的结果为阴性。Brownstein 的团队可能耗尽整套分析手段,仍然只能把病例搁置一年,等待新的线索出现。

5. 表型翻译质量和关系生物学决定哪些线索会上升

  • “没有眼泪”这样具体的线索,可能强烈指向某一种疾病,但前提是“alacrima”“no tears”以及受文化影响的日常表述能够被准确对应。Brownstein 特别提到 Monarch Initiative 在连接 HPO 术语、患者语言和动物表型方面的工作。

  • 更复杂的流程会利用生物学关系:如果基因 A 与某种表型相关,并且与基因 B 发生相互作用,那么即使 B 尚未建立明确的疾病关联,B 上的大型变异也可能成为候选线索。

  • Brownstein 的第一个成功病例正是如此。一名患者出现发作性共济失调,随后身体变得僵硬并锁定在某个姿势;测序找到了 KCNA1。它不是预期中的基因,但与预期基因相关,因此“一下子升到了候选列表的绝对顶端”。

6. 重新分析能在医学无法改变过去之后,改变一个家庭的未来

  • 重新分析之所以重要,是因为“你不可能成为每个基因、每种疾病和每类结构变异的专家”。某个变异今年被忽略,明年可能因为另一组研究人员建立了缺失的关联,直接成为浏览器给出的第一答案。

  • Brownstein 回忆起 3 名兄弟姐妹,他们约 20 年前因一种肌病去世。DNA 样本耗尽后,剩余的 RNA 仍可用于 RNA-seq,最终发现了一个“我想是 CFL2”的变异,为尚在世的兄弟姐妹提供了携带者检测和家庭规划的信息。

  • 团队曾讨论,在多人已经去世后才完成诊断,是否还能算成功。Brownstein 的结论克制但实际:它无法修复过去,却能让下一代“睁大眼睛”继续生活和做决定。

  • 另一个家庭获得了一种罕见骨病的诊断,患者横跨 3 代,其中还包括一名 90 多岁的亲属。诊断无法撤销此前不必要的手术,却把一段 3 分钟的症状叙述,变成了一个 10 秒就能说清楚的解释。

7. 候选发现如今需要队列、复现和群体背景支持

  • 2011 年,一个惊人的蛋白质预测可能立刻激发论文发表的兴奋感。到了 2024 年,“水位线正在上升”:单个异常变异可能只是随机事件,因此研究人员会寻找同一基因出现变异、且表型相近的其他家系。

  • Matchmaker Exchange 和 Beacon 能帮助研究人员找到这些家系。单个病例可以加入一个已有 19 名患者的病例系列,从而形成更有说服力的基因—疾病关联,也让患者家庭与相关专科医生建立联系。

  • 预测工具只是证据,不是圣旨。Brownstein 会使用 CADD、SIFT、PolyPhen 和蛋白质影响模型,但提醒说,一些已知的疾病关系可能根本过不了这些过滤器,因为在某些基因中,看似轻微的扰动也可能产生重要影响。

  • 一个令人兴奋、且高度保守的变异,也可能在被证明于某个测序不足的族群中很常见后失去意义。因此她提出了一个绝对的数据要求:“我们需要测序全世界”,才能把因果关系与孤立的背景变异区分开来。

8. 生物学缺失的接口层正在浪费专家产能

  • AlphaFold 和 STRING-DB 等蛋白质折叠与相互作用工具“非常惊人”,但生物学相互作用图谱仍有大量区域未被照亮,许多工具也依然难以使用。Brownstein 说,前沿蛋白质分析工具对用户如此不友好,仍让她感到意外。

  • 她最尖锐的例子来自实际操作:一家测序公司只通过命令行指令交付数据,也不提供任何帮助。她不得不临时学习 Harvard/Boston Children’s 的超级计算机,只为传输自己的文件——“这是巨大的时间浪费”。

  • Labenz 认为,机会在于为不会编程的生物学家和医生提供一层 UI 与编排能力。Brownstein 补充说,如果工具只能由熟悉 Unix 的人使用,也会压制其创造者从未设想过的应用场景。

9. 共享激励对发现的约束,超过了原始数据稀缺本身

  • Brownstein 理解年轻研究人员为什么可能不愿共享罕见病例:他们希望靠一篇足以改变职业生涯的论文成为主导作者,而不是在别人的联盟项目中沦为共同作者。但这种激励会让病例停滞,也把研究者利益置于患者利益之前。

  • Boston Children’s 的 CRDC 队列委员会和全院 GeneDx 浏览器提供了更好的模式:研究人员可以查询去标识化的遗传变异,不接触姓名、表型或其他可识别信息;如果发现匹配,再联系负责医生。

  • 一次候选基因查询找到了另外 4 名患者,随后一名医生把 Brownstein 引向荷兰主导的病例系列。结果是更完整的疾病描述、论文发表,以及患者与专家之间的直接连接。

  • 她粗略估计,理想的数据共享会产生巨大影响。许多队列仍“躺在冰柜后面”,而加入多个项目和登记系统,则能提高未来发现真正抵达某个家庭的概率。

10. 行政默认设置可能让有效的科学结果归于无效

  • Brownstein 看到生物样本库政策多年来仍坚持主动加入,即使人们一直在推动更容易地共享废弃组织、尿液和其他样本。她认为,相关改革存在巨大的制度惯性,而受影响家庭真正看重的往往是速度,而不只是隐私。

  • 她还提到 Mew 研发 Milusen 的 N-of-1 药物故事:一次机会,加上一户极其积极的家庭,最终克服了重重障碍。这类鼓舞人心的案例,可能只是大量无法跨越这些障碍的家庭所构成的“暗物质”。

  • 研究发现还需要 Clea 认证实验室确认,并由医生或遗传咨询师向患者反馈。有时医生“就是不配合”——他们认为这件事价值不大,不愿回复邮件,或者单纯拒绝投入确认诊断所需的工作。

  • Brownstein 自己的家庭也曾无法获得一项研究发现的临床确认,因为一名亲属的医生不愿配合。再叠加不足的遗传咨询和错失的研究入组机会,这些看似微小的失败最终变成系统性障碍。

  • Brownstein 希望 AI 能把一部分自主权还给家庭:帮助他们更早识别正确的专科医生、检测项目或登记系统,减少对那些“运行得不够好”的基础设施的依赖。

11. AI 的第一个突破,是消除科学工作的琐碎劳动

  • 在当前病例分诊工作之前,Brownstein 曾通过 Picory 资助项目使用交互式模型,把患者表型映射为 HPO 编码,整个过程由 7 个问题组成;这项工作当时仍在分析中。

  • Brownstein 已经获得的最大收益,恰恰是看似平淡的论文和基因摘要。过去她要找论文、遇到付费墙、登录 Harvard 图书馆,最后才发现摘要与问题无关;现在这套流程“每天还给了我几个小时”。

  • 作为核心设施负责人,她用 ChatGPT 学习新推出的测序方法,为研究人员咨询做准备,并在近 20,000 个基因中总结陌生疾病。它还帮助她破解审稿人 2 号含糊的说法——自己遗漏了“一整套文献”。

  • 共同方向都是从乏味检索走向专家判断:“它正在削减工作中那部分琐碎、耗时、非常无聊的内容”,然后把她送回基因发现本身。

12. 一个 500 人膀胱队列显示,AI 正作为研究过滤器运行

  • Brownstein 研究严重的间质性膀胱炎/膀胱疼痛综合征,这种疾病有时严重到患者无法离开家。在约 500 人的队列中,她和合作者寻找携带变异数量高于预期的基因,以及富集的多基因通路。

  • 一条通路可能包含 12 个基因,统称为“小分子转运”。ChatGPT 可以快速完成第一轮筛查:逐一判断每个基因是否与膀胱疼痛、膀胱生物学或膀胱癌有直接关联;Brownstein 再对回答为“是”的结果进行验证和深入研究。

  • 其中一条通路约 60% 的基因都与膀胱癌有明确关联,提示它们可能涉及膀胱表达和已知扰动。另一个基因与尿路上皮问题的关联,几乎让她尖叫,因为这提供了一个可信机制,足以在第二天早上的会议上讨论。

  • 在另一个病例中,同样的方法帮助她在约 1 个半小时内,把 20 个合理候选基因缩小到 3 个,用于汇报。Brownstein 的判断带有代际意味:未来的遗传学家可能根本不知道,在这些工具出现之前,人们是怎么完成这项工作的。

13. 推理模型的越界,可能是错误,也可能是发现

  • Brownstein 仍在试用 o1 和 GPT-4o,还没有形成稳定的模型分类体系。她没有使用网页搜索,因为“网上有很多垃圾”,并表示即使幻觉和推理能力都在改善,检查输出仍然是最重要的环节。

  • Labenz 会尽量中性地设计提示词,因为模型可能迎合用户偏好的理论。Brownstein 面对的却是相反的问题:她只是问某个基因是否与某种表型有关,模型却可能给出一套聪明的多步机制论证;而她真正想要的只是直接把两者联系起来的论文——“不对,太远了,太远了。”

  • 这段额外推理究竟是幻觉,还是类似 AlphaGo“第 37 手”的发现,目前仍无定论。Brownstein 的诚实回答是:“也许是我不够聪明,没看懂,但它是对的。我不知道。”她提出的办法也很直接:继续测试。

14. 资助项目瞄准最低充分数据、算力与整合能力

  • Brownstein 的 OpenAI 项目要回答的是:如何更快完成诊断、生成新假设,以及一个答案最低需要多少数据和算力。她目前只是总结病例,而不是上传完整病历;即使在 Boston Children’s 受保护的 ChatGPT 部署环境中,她仍然无法让引文稳定工作。

  • 完整病历和基因组会带来跑题、无关相关性和高算力消耗:“不是每个抽烟的人都会得癌症。”每个问题都要消耗能源,因此研究重点不只是最大化上下文,而是筛选真正有意义的证据。

  • 一年后,Brownstein 想象中的系统会是经过封装、带有引导的界面,帮助患者提出正确问题,也让研究人员更清楚地表达使用场景。她希望病例求解能够实现,但拒绝做出承诺:“我不知道到时是否真的能做到。”

  • Labenz 的判断更激进:o3 在极高算力下的 FrontierMath 表现接近 25%,低算力下为 10%,而此前的上限约为 2%。即使受到上下文限制,他仍预计总结和过滤能力足以让部分罕见病病例变得可处理;Brownstein 则提议一年后再聚,看看实际进展如何。

Nathan Labenz

Today, I’m speaking with Dr. Catherine Brownstein, MPH, PhD, an assistant professor at Boston Children’s Hospital and Harvard Medical School whose research focuses on identifying the genetic causes of previously unexplained rare and orphan diseases, and who was recently awarded a ChatGPT Pro Grant from OpenAI.

You might be surprised to learn, as I was, that so-called rare diseases are not necessarily all that rare. Any disease affecting fewer than 1 in 2,000 people, or fewer than 200,000 people in the United States, is classified as a rare disease. Often, families spend painfully frustrating years bouncing around the medical system in search of an accurate diagnosis before ultimately reaching Dr. Brownstein’s elite team at Boston Children’s.

Of course, considering the radical cost reduction we’ve seen in genetic sequencing in recent years—with nearly a 10,000× improvement in affordability, one of the very few cost curves ever to rival that of large language models—there’s been an ongoing revolution in this space even before the current AI moment. In 2007, a genome sequence cost upwards of $1 million. In that era, it was used only in the most challenging cases and was often a difference-maker. Today, it’s just a couple hundred dollars and has become commonplace for individual patients.

But that creates new challenges for specialists like Catherine, who now have to comb through a vast and still exponentially growing literature to find candidate diagnoses for their most challenging cases. This new wealth of information, which, as you’ll hear, could be growing even faster with improved regulations and incentives, makes information-processing capacity relatively scarce and valuable. And you can probably guess where this is going: a great target for the latest generation of reasoning models.

This conversation is, above all, a window into how frontier large language models are starting to become useful in highly specialized fields. Dr. Brownstein is pioneering the application of AI to rare-disease research in real time. She’s using AI to triage potentially relevant research and, in some cases, to connect the dots between subtle clues. She’s working directly with OpenAI to develop use cases and provide feedback.

Considering that every case represents a real person with a life-altering or even life-threatening condition, she’s constantly working to find the right balance between enthusiasm for AI’s capabilities and a healthy skepticism about any specific AI output. As you’ll hear, she’s still figuring out where AIs can be the most valuable, how best to use them, and how much to trust them. That such an established expert is bringing what amounts to a beginner’s mindset to such high-stakes cases may be surprising to some, but I really don’t think it should be.

Even the most AI-obsessed folks like me have only managed to log a few thousand hours with large language models, and nearly all of that was with earlier and less powerful models. So, for the current frontier, we’re all still figuring this out together, and there’s currently an unprecedented opportunity for people with deep experience in specific niche domains to become the leaders in applying AI to their particular fields.

Catherine Brownstein, MPH, PhD, an assistant professor at Boston Children’s Hospital and Harvard Medical School, you specialize in the discovery of new genes for rare and orphan diseases, and you’ve recently been awarded a ChatGPT Pro Grant. Welcome.

Dr. Catherine Brownstein

Thank you so much for having me.

Nathan Labenz

Yeah, I think this is going to be really exciting. As regular listeners know, I have a growing obsession with the intersection of AI and biology. When I saw your name on the o1 Pro blog post announcement, I was excited to reach out and learn more about how you’re applying the latest AI tools to some of these very challenging and pressing problems.

Maybe we could start with a zoomed-out overview of your work, because our listeners are definitely following AI developments. They know about o1, they know about o1 Pro, and probably quite a few have subscribed, even at the $200-a-month level. But they probably don’t know a lot about rare diseases, what the state of knowledge is, or what sorts of techniques people use to try to figure these things out.

I’d love to get a layman’s introduction to your advanced work. This may be a very tough question, maybe the toughest question, but what’s the layman’s introduction to your work?

Dr. Catherine Brownstein

When I’m asked a question, I usually answer it by saying that rare diseases are quite common, actually. There are more people with rare diseases in the United States than there are natural blondes, so it’s actually quite common to have a rare disease.

A lot of it is how you define disease. Autism is really common, but autism due to a de novo variant in KCNJ8 is quite rare. It’s a tricky definition, and it’s always evolving as we learn more, but basically, I consider myself a gene hunter trying to diagnose the undiagnosed.

Nathan Labenz

I read in preparing for this—I think it was Perplexity that gave me this answer—that the definition of a rare disease is one that affects fewer than 200,000 people in the United States. That was a surprisingly large number, wasn’t it?

Dr. Catherine Brownstein

It’s always wild to me because when you think of a city that has 200,000 people, that doesn’t seem like a small town, or at least it doesn’t to me. But that’s the definition of rare in comparison to common disease, which can affect millions. Epilepsy, for example, affects 1% of the population.

Nathan Labenz

Yeah, that’s really interesting. Maybe a little bit more background on the patient’s journey through the medical system to get to you, and then your experience of encountering new patients. I know it’s impossible to give just one story because I’m sure they’re extremely varied, but how do you end up coming into contact with patients? What have they gone through to get to you? And what do you do once you get a new case?

Dr. Catherine Brownstein

I’m really lucky to be at Boston Children’s, which is an internationally known tertiary hospital, so we get really interesting cases from all over the globe. Usually, a patient or family starts out by going to their local medical provider. They can’t figure out what’s wrong with the child or person, so they get referred from specialist to specialist, and they still can’t figure out what’s wrong. Eventually, they get to us.

A lot of times, patients and families have been bounced around for years, trying to figure out what’s going on, what’s next, what they can expect, and just looking for answers. Sometimes that’s not the case. We have a lot of really medically savvy families who know their child, know something’s wrong, and need the best right away. They search on the web, find the person who works on that phenotype, and call every day until they get an appointment.

A lot of times, though, it’s a more circuitous route, going from doctor to doctor to doctor and then finally somehow ending up at Boston Children’s. If they see a clinician who doesn’t know what’s going on, they often refer the case to the organization I work with, the Manton Center for Orphan Disease Research.

We get a lot of the negative cases throughout the hospital where they think it’s genetic in origin. Then we’re able to get the medical records. We’re a philanthropically supported center, and patients can self-refer. We get all the medical records and all the genetics that have been done before. Then we have a huge multidisciplinary team, and we review the case, go through it, and do a reanalysis.

Sometimes we resequence or use a new technology if one is available, like RNA-seq or long-read sequencing. Then we work together to try to figure out what’s going on. When I first started in 2011, genome sequencing and exome sequencing were quite rare. If patients were able to get it, a lot of times it was like shooting fish in a barrel. We would have something like an 80% diagnosis rate.

But now genome sequencing and next-generation sequencing are so common that we only see the families if they've already had a negative sequencing test. So we go from diagnosing roughly 80% of cases to roughly 10%, just because we're getting the most difficult of the difficult cases. They've already been reviewed by really good geneticists and are getting to us because they just can't figure it out.

But that's one thing that I think AI can really address: shortening this diagnostic odyssey for patients who have just been jerked around—not through anyone's fault, but just by the nature of how these things go. Maybe AI can help in analyzing symptoms: maybe you should see this doctor right away, maybe you need this test, or maybe you need to go to this specialist, and just make things happen a lot faster.

Nathan Labenz

That callback to 10 years ago, I think, is quite interesting. Maybe you could give us a little bit of a sense of the relative pass-through rates at these different levels of the filter. People initially go to their local doctor, the local doctor doesn't know what's going on, they get referred, and eventually they get to your hospital, where you've got the best of the best.

There's another related but distinct line of research that has recently been comparing AI's ability to diagnose through a natural-language conversation with patients against doctors. It seems like, against at least the average doctor, the latest models are now very much holding their own. I don't know if that would be true if we were looking at Boston Children's elite clinicians and their ability to diagnose, but they still don't know what's wrong.

In the past, if I understand things correctly, because sequencing was rare, you could often just do a full genome sequence and then be like, “Oh, okay, well, there's your problem.” The literature has characterized this: now that we have this additional information, there's a pretty clear match. Today, that low-hanging fruit is getting absorbed somewhere else in the system before it gets to you, and you're now seeing things that are basically not characterized in the literature at all, or maybe just a little bit.

I'm not sure why the connection wouldn't be one that others could make, but use that prompt and fill in a little more detail, if you would.

Dr. Catherine Brownstein

When I started, I was actually hired as a project manager at Boston Children's to help clinicians get their patients sequenced. Clinicians, even though they weren't geneticists, were really good at identifying cases that were probably genetic in origin. They had freezers full of this DNA, just waiting for the technology to come online so they could analyze it and figure out if there was something genetic that could be discovered.

When I started, an exome, which is just 1% of the genome—it's just the coding region, just the genes—was $3,800, or close to $4,000. It's a good place to start if you're being economical, because a lot of the variants are within the coding region. Now, I just priced out an exome, and it's $160 for that exact same test.

It was so expensive that they went through a rigorous selection process if you were going to get an exome done. If you were going to bet money, you were going to bet money that it was genetic and that you were going to be able to figure it out by doing a trio—that is, the patient and the parents—and that you were going to see something that was de novo, which is basically not in the parents but is in the child. It's like lightning striking, an error happening during development, which causes disease.

The first cases of that were the 80% that I was talking about, and it was because these patients had been collected, in some cases, 20 years ago. They had the DNA there, and sure enough, there was a premature stop codon or a huge deletion of 1 gene that was already hypothesized to be related to this condition or a similar condition. You could point at it and be like, “Yep, that's it.” You would also have multiple cases of the same type, where you'd see 4 families with the same gene missing and the same phenotype, and then you're really confident that that gene is causative of the condition.

As the price dropped, it became less of a thing that happened. It's not because you couldn't get an exome done anywhere; there are a lot of geneticists, a lot of really savvy clinicians, and a lot of for-profit companies that you could send it off to, get a report back, get diagnosed, have more precision-medicine treatment, and go on your way and do very well.

So it was the negative cases that were getting referred up the chain to Boston Children's, because they had already had a genome and it came back negative. That is, there was no obvious variant in a known gene that could explain what was going on. Then it becomes a little trickier.

We start forming cohorts. At the Manton Center, we work with clinicians. We have clinicians in every department of the hospital who are able to refer patients to us. We consent them to our protocol, and then we collect samples and medical records. Sometimes we wait, and we reanalyze.

When we have 4 to 10 patients with the same thing, we're able to look at them together as a whole group and be like, “All right, are there things in the same gene, the same family of genes? What can we come up with as a hypothesis here?”

In 2014, I think Zack Kohane, who had previously been a guest on your podcast, had the idea of having an international competition to solve undiagnosed families. We got 3 families with seemingly Mendelian disorders. That is, we thought they were genetic, and we thought there was something going on with a clear relationship between gene and condition.

We released their data all over the globe to 23 different teams, and we had them compete and each submit a report on what they thought the cause of each family's condition was. It was really interesting. A lot were actually diagnosed from this—I think 2 out of 3 walked away with diagnoses from the process.

We were also able to show that diverse teams did much better. You can't just have a bunch of bioinformaticians in a room together looking at cases and expect them to come up with the right answer. It was teams that had a mix of research assistants, genetic counselors, researchers, clinicians, research clinicians, and clinical geneticists working together. All those diverse perspectives, on the whole, were able to solve more cases.

That was really interesting, and I think that's a recurring theme here when we're talking about LLMs, large language models, and AI. None of this exists in a vacuum. It's helping us along, and maybe it will be enough, but right now, having multidisciplinary teams with multiple strengths all working together means we do much better as a whole.

The other thing I wanted to add is that we still see those slam dunks. We just had a case a little while ago where it was 1 family and 3 generations, all with a rare bone disorder. The matriarch or patriarch was in their 90s, and we were able to give a diagnosis to this person in their 90s, which I thought was really, really cool and shows the power of just having an answer.

They had already gone through surgeries they didn't need to go through and had their whole life with this condition. But something as simple as being able to explain what's going on in 10 seconds, as opposed to 3 minutes of describing symptoms, means a lot to the family.

Nathan Labenz

Yeah, I imagine, especially if you've been dealing with something like that for 90-plus years. That's crazy to think about.

So, a lot of different questions I have about all this, but in these cases where you're getting all the way through the entire medical system, basically, and finally getting to one of these cross-functional teams—

Dr. Catherine Brownstein

Mm-hm.

Nathan Labenz

Can you tell us a little bit more about what the process looks like when that team gets to work? In AI prompting, we talk about thinking step by step and breaking problems down. Maybe one way to frame it would be: What is the sort of collective chain of thought that the group goes through to start with inputs?

Inputs would at least be symptom descriptions and results of genetic testing sequences. I don't know if there's any other inputs that you get at that level. I guess you have the whole scientific literature also as sort of an input. Then you do some thinking and reasoning, maybe some additional testing, and finally you get to a result. What are you doing when you're doing that?

Dr. Catherine Brownstein

Okay. So, when a case comes across my desk, usually there's a medical record that comes along with it because, again, they've been bounced around for a long while. Usually, at this point, they've had some genetic testing that gets transferred to us.

More and more patients are coming with it on a thumb drive, like, “Here's my genome,” which I think is really cool and didn't even happen a few years ago. We run it through our genomic pipelines, and usually we run it through more than 1 because they all have their strengths and weaknesses. Some are more comprehensive but harder to use, and you'll get more false positives because they rule fewer things out. Then you have others that are really easy to use—my kids can use them and understand intuitively what it means—but sometimes they're black boxes, and you don't know the reasoning behind why a variant was eliminated or not.

I'm a PhD, non-MD, so I usually like things to stay anonymous. I don't want to be a walking HIPAA violation, so I kind of don't want to know the names or meet the families, but sometimes I do. I know who they are.

We go through everything case by case and line by line. There are certain phenotypes where I think more information is better. You'll get the occasional phenotype that's only linked to 1 condition, like lack of tears in 1 condition, and that's a really important clue. Then we'll look at that gene.

For what the patient is experiencing overall, generally there are gene lists of what's already been discovered, and you can look at the genomic information for any variation that could be causing disease. We'll call it, for simplicity's sake, pathogenic variation, though suspected pathogenic variation is probably more accurate to say in those genes.

You get the new analysis done, and then you're looking at what's known. If you don't see anything, then you start looking at your special sauce. How am I going to approach this? Where else in the genome is notable? Is there a huge structural change that hasn't been linked to disease, or a translocation where chromosomes break and reattach in the wrong spots? Is there some other deletion or duplication? What's rare? What's unique to this patient?

Now there are also all these new technologies, like looking at epigenetics, where you can kind of predict which genes are turned on and off. Even if you can't see a mutation, is the gene of interest's expression perturbed somehow, or is it constitutively on even though it's not supposed to be? Can you take a look at that?

Sometimes, in the back of your mind, you're thinking, “Is it multifactorial?” It's not just 1 gene impacting it. It's not some big error in 1 gene; it's a bunch of tiny little things scattered throughout the genome. Then there are different types of tests, like looking at a GWAS, or genome-wide association study, or SCAT, where you can look at rare variation weighted by how rare the variation is and how damaging it's predicted to be to a protein. You look at that and see, “Okay, is there some reason why you think that this is going on?”

A lot of the time still—let's say 25% of genetic testing comes back positive—what does that mean? Sixty-six to 75% are negative. Then you go through this whole process, and still most are negative. You put it on the shelf, wait a little bit, and analyze it again a year later.

Reanalysis is actually really, really important because things get discovered all the time. You can't be an expert in every gene, every condition, every structural variation, and other people are actively working on it. A lot of times, you'll take something off the shelf and look at it again, and it rises right to the top. The number-one thing in the genome browser is the answer, and you stared at it a year ago and didn't make that connection. Now, all of a sudden, there is an answer.

Actually, I was asking Alan Beggs and Monica Wojcik, who are the director and medical director of the Manton Center, for success stories—if they had any that stuck out. One was from 20 years ago. There were 3 siblings who all passed away from a type of myopathy, and they couldn't figure it out.

They kept testing and testing and testing, and eventually ran out of DNA. Then we had a pilot grant at the hospital to do RNA-Seq, and Alan and Monica submitted this family because we had some RNA left. We found a variant in CFL2, I think that's the gene name.

Even though it was 20 years ago, the surviving siblings were now planning families, and they had an answer. They could do genetic testing to make sure that there weren't 2 variants and that they were each carriers of 1 variant. They hadn't passed away, so they only had 1 variant, not 2. They could also make sure that their partners didn't have a variant in the same gene and ensure that the next generation wasn't going to have this horrible, fatal myopathy.

In some ways, we had an interesting discussion: “Okay, is that really a success story?” Whenever there are multiple deceased people, is that really a success? Yes, you diagnosed it, but it's not changing anything. But it is changing the future. They're going forward with their eyes wide open and are able to plan as a result.

Nathan Labenz

Yeah, that sounds like certainly some form of success to me. I have 3 young kids, and fortunately, no crazy medical conditions in my family, but we still did a little bit of genetic testing. I would say I was probably never more nervous than when opening that report, just to make sure that I wouldn't have to see something really weird or strange, or that it would change the course of my life.

To be on a potentially negative course and get the assurance that you could confidently get on a path where you'd be able to have healthy children, I think sounds like, to put it mildly, a life-changing development for those folks.

Dr. Catherine Brownstein

That definitely resonates with me.

Nathan Labenz

Okay, let me dig back in at a few points along the way. I'll try to summarize a little bit and interject a couple of questions.

The pipelines that you're describing—I guess those are maybe a mix of commercial options or things that other academic groups have put out. The inputs to those, are they highly structured data?

I mean, I'm thinking here: My sequences are, of course, structured; my symptoms are not, right? I describe myself in words, and the doctor I'm talking to notes that in words. Is there a way that gets translated into specific coded sets of symptoms, or what is the intake of these pipelines?

And then are they basically doing deterministic work, where they're essentially running down a long checklist and saying, “If you have this, we check this. You don't have that, so that's out,” and working down a long set of known possible conditions? Or how would you characterize what those pipelines are doing internally?

Dr. Catherine Brownstein

So, I think you're exactly right. A lot of them let you input the phenotype, and it's coded to ontologies—sometimes HPO codes, sometimes ICD-9 or ICD-10, sometimes SNOMED. There are a bunch of different ontologies. I like HPO the best.

Nathan Labenz

Mm-hm.

Dr. Catherine Brownstein

Let me be real clear: In those sorts of ontologies, something like “no tears” would be a single alacrima HPO item. My condition might be summarized by a set of those. If I had no tears, hair falling out, and loose teeth, that would be 3 things. It would be, “Okay, the patient presents with this bundle of things.”

It's also a huge field of research. My friend Melissa Haendel works with HPO and her site, Monarch Initiative, mapping that onto animal phenotypes and making sure it's one-to-one. Humans don't have paws, but the phenotype that's closest to that can be translated.

Then there's layperson HPO, where we're not saying “alacrima,” but we say “no tears,” or “lazy eye” and “strabismus.” There's a whole mess of work that goes into that, making sure that it's accurate and also culturally sensitive, like “fit” for epilepsy. It's all this stuff that you never think of, and if you don't make those translations, then all of a sudden your phenotype is way less accurate than it could be.

That gets incorporated into the model. Then, when you input that with the genetics, you can have raw data, which is FASTQs—the zeros and ones that come off the machine—and then a BAM, where you're looking at the reads of the sequencing itself.

Gosh, I'm not going to explain this very well. But then you have the VCF, which is really processed data, and it's basically every single variant. It's huge— a VCF is a relatively huge file. It is orders of magnitude smaller than a FASTQ or a BAM, but it's still quite big.

Then you're putting the BAM or FASTQ into these pipelines, which process the data along with the phenotype. Then they're ordering the variants based on the HPO code related to the gene, the variant within that gene, and how likely it is to be positive for disease. The more sophisticated ones can take in relational kinds of things, where it's known that this gene binds to another gene, and gene A is related to the phenotype but gene B isn't yet, while there's a huge variant in gene B and the patient has the phenotype associated with gene A.

My actual first-ever success story was one of those cases. It's called episodic ataxia, and the patient would get really stiff and couldn't move—they would get locked in position. We did sequencing and saw that it was a variant in KCNA1, which wasn't the gene we were thinking of, but it was related to the gene we thought it was going to be. So KCNA1 just rose to the absolute top of the list, which was really, really cool.

Nathan Labenz

But that challenge of basically understanding the graph of interactions—what affects what in the cell, at the tissue level, or at the system level, whatever—has been a fascinating area for me recently. I've been really interested to see some new projects. I don't know if you've come across these yet, but there are some that are now trying to predict the evolution of essentially the transcriptome or cell state from one timestamp to the next. I think that really suggests a major revolution coming soon.

How much would you say—I don't think there's any answer to this, because I don't think we know how much we don't know—but when it comes to those interaction-type things, my sense has been that we have a relatively small amount of that space illuminated today? Of all the interactions, of all the things where something in this gene interacts with another thing and could cause a third thing downstream, my sense is that we have a pretty small percentage of those pathways mapped out and well enough understood that we could do this kind of analysis. Is that a good summary, or how would you improve on my summary?

Dr. Catherine Brownstein

No, I think that's totally right. Every time I try to look at the impact of a variant on the protein, I'm surprised at how, first of all, user-unfriendly a lot of these tools still are. It's because they're really tough. They're cutting-edge, and protein folding has come a long way. Definitely super cool, and the people who work on that are totally hardcore, but there's still a lot to be learned, and we're still folding certain proteins. We don't have everything worked out yet.

I just keep thinking about when we first got genome sequencing and how difficult it was to use some of these browsers. They would crash the computer. I think protein folding and some of these tools, like STRING—STRING-DB, for protein-protein interaction—they're amazing, and they're going to continue to get more and more amazing and more useful as time goes on. Especially when they get more user-friendly for people like me.

Nathan Labenz

Yeah, it sounds like that might be a real low-hanging fruit. This has come up on a couple of different episodes, where the general observation has been: biologists are not programmers, and doctors are not programmers. There's a missing layer that would unlock a lot of value if we could just make it a lot easier for doctors and biologists to use the models and other information tools that have recently been created. A lot of times, those are still put out there in open-source project form, and they need a UI layer or an orchestration layer on top to really make that accessible and useful for a lot more people. That could be an interesting area for somebody to dig into more.

Dr. Catherine Brownstein

Mm-hmm.

Yeah, and just little things. I got some sequence back from a new company, and they're like, “Okay, here's the commands to download your data.” I'm like, “Whoa, whoa, whoa, what?” They had no intention of helping me, either. I had to learn the command line and how to get my data from their server down to mine, or I didn't get my data.

I had to have a crash course on getting onto the Harvard/Boston Children's supercomputer in order to get my data, and it was a huge waste of time. I think they're assuming a level of literacy for some of these programs that people just don't have. You can argue that I should, being in the job that I'm in, but it's hard. It's a learning curve, and I think there's a lot of opportunity there for making things a little more friendly.

It goes back again to: you don't know what you don't know. If you make your tool accessible to a wider audience, they're going to apply it in ways you never dreamt of. Gatekeeping it to only people who know Unix is kind of tough on everybody.

Nathan Labenz

Let's circle back to that in a second, because this sounds like one of the candidate areas where you might be getting some good value from your o1 Pro Grant. Are these pipelines using any sort of predictive AI technology, like classifiers and things like that, or are they working off a sort of accepted, known literature of findings?

I could imagine—and maybe it varies across providers—that one form of pipeline is, “We want to be really grounded in things that are very well-established, and we're going to run down this super-long checklist programmatically for you and try to find things that fit.” I could imagine another pipeline that would be like, if these models exist—and I'm not sure to what degree they do—you could say, “Hey, here's my genome. Predict and give me guesses.”

Are there models like that? And I guess, to what degree is this all deterministic versus whether those existing pipelines are already starting to lean into certain kinds of AI?

Dr. Catherine Brownstein

I think you need both. You need to be confident that you've looked at a genome with all the known things and that nothing funny was missed—just very validated best practices. Then you need the exploratory pipelines, and that's what we're developing as part of my grant with OpenAI. What's the limit? Where can we take this? Where can we make shortcuts where, before, we were taking a ton of compute and a ton of time? How do we solve cases faster? What's the minimum required data set in order to make a diagnosis? What's the minimum compute necessary in order to get a diagnosis? How do we diagnose new things? How do we come up with new hypotheses faster, all using AI?

Nathan Labenz

Well, that's probably a perfect tee-up for your application of the latest models.

Yeah. Maybe for calibration, before we get into workflow specifics, when did large language models start to be useful for you? Was it just with o1, or were you already starting to see some value with earlier versions?

Dr. Catherine Brownstein

We had been using it along with the phenotyping areas more than anything else. I had a Picory grant working with Ingrid Holman and Melissa Haendel, where we were trying to take a patient phenotype, map it to HPO codes, and get the layperson to HPO faster and more accurately.

One thing that we used at one point was working with 7 questions, asking what system was affected, and drilling down that way. We were seeing the ability to get an accurate phenotype through an interactive model using your own words, compared to traditional self-phenotyping, like surveys and things that are on the web now. We’re still analyzing that.

There are a lot of publicly available tools that I was using, as I mentioned before, like AlphaFold, STRING-DB, and a lot of these protein-impact prediction models that are required to do our jobs. We need to be able to predict the impact of a variant on a protein.

We can’t treat it as gospel. People who rely too heavily on these algorithms sometimes get tripped up because some of the known gene-disease relationships wouldn’t pass those filters now. There’s just something about that gene where you perturb it a tiny little bit and it causes a phenotype that you wouldn’t even think it would cause, but we know that’s true. So, if you looked at it at face value, you would have skipped over it.

I think a lot of people are using these models and don’t even know they’re using them. They don’t really know what’s behind them; they just know that you look at the CADD score, SIFT, PolyPhen, and protein impact, and then that’s a cutoff, along with allele frequency. They don’t really realize that aggregation of allele frequency is powered by a lot of these models, with a ton of stuff happening behind the scenes. If you took it away, we would be struggling.

Nathan Labenz

So, do I have it right, then, that with an AlphaFold-type model, this is after a standard pipeline basically comes back negative? Then you would say, “Okay, let’s go into essentially anomaly-detection mode for this person’s sequence?”

Dr. Catherine Brownstein

Mhm. Exactly.

Nathan Labenz

Yeah. And you have tools for that as well that can say, “Hey, look, here’s a giant deletion,” or, “This gene has stopped prematurely,” or, “This one has been copied over a bunch of times,” whatever. There are, of course, plenty more ways things can be weird than those, but you identify those and then say, “Hmm, I wonder if that maybe is the thing. I’ll use AlphaFold to take that genetic sequence, see what that protein actually looks like, and then do a structure comparison. Does that look like that protein is really mangled?” If so, that becomes a place to go deeper?

Dr. Catherine Brownstein

Yep. Exactly. A lot of that comes with experience, too. There are some genes that are really mutated in pretty much everybody. If you don’t know, you’re like, “Oh, look at that. That’s so cool,” and then some veteran is going to be like, “No, it’s not that. It’s never that.” Or it’s never lupus.

Then you see a gene that you’ve never seen before, and it has a variant in it that’s conserved down to zebrafish and C. elegans worms. You look at it in AlphaFold, and you see that it’s royally messing up the protein, and you get excited.

It’s a roller coaster a lot of times. Even that will fall apart somewhere, and then you’ll find out that it’s only really common in one specific ethnicity that’s hardly ever sequenced, but the patient is from that rare ethnicity. It goes to show that we need to sequence the whole world in order to understand what is actually disease-causing and what is just background variation in isolated populations.

Nathan Labenz

Yeah, there’s another fork in the road here. Which question to ask? We’ll come back to the data, because that is a can’t-miss area, but just take us a little bit further down this path. We’ve identified some anomalies. Now we run the folding model and see that the structure looks off. Where do we go from there? What’s the next investigation after you’ve identified that?

Dr. Catherine Brownstein

Back in 2011, you would get really excited about it and want to publish it.

Nathan Labenz

But in 2024, is the bar always rising, for sure?

Dr. Catherine Brownstein

Yeah, the waterline is rising, and now people are like, “Wait a second. That might just be random.” So then you want other families or other cases with the same type of thing—variants in the same gene. There are all these sharing tools to be able to do that.

One is called Matchmaker Exchange or Beacon, where you put in the variant and the patient phenotype, and you see if anyone else has put in that same gene attached to the same phenotype. Then you match, and you collaborate. Or somebody has already started a paper with 19 cases of variation in this gene causing intellectual disability. If you have one, you can add it to that case series and get a better publication out of it, one that is much more convincing than if you just publish your one case, which looks pretty cool and you’re convinced by, but other people might not be after reading it.

Nathan Labenz

The bar is continually being raised on this stuff. So that brings us back to data naturally. How would you characterize the data environment? I was struck, in reading through a couple of the papers—I don’t have the vocabulary to go as deep as I might wish to on all of your papers—but I was able to see quite clearly that the n is small in a lot of these papers, with single-digit numbers of cases.

I’ve also noticed a few times that you’ve spoken about the hospital as sort of the data unit, it seems like. I’ve heard from a bunch of people over time that we have this sort of scarcity of data, and I’ve always wondered: Is it a true data-scarcity problem, or is it a sort of man-made, for lack of a better term, data-scarcity problem that’s really more about barriers to access and sharing?

Dr. Catherine Brownstein

It’s a tough situation. I don’t want to fault the young researcher who doesn’t want to share their super-cool case because they’re hoping they’ll find another one and be able to publish it as their finding, not as somebody else’s finding in a giant research-group facility across the world, where they’re just going to be a middle author and it’s not going to make their career the way it would if they held on to it tightly, did everything themselves, and got it out there.

The problem with that is that a lot of times it doesn’t work out that way. If that’s not benefiting patients, you’re not thinking of the patient; you’re thinking of yourself. It’s much better for science, and much better for patients in general, if everyone shares their data and has it open. If you see something in someone else’s case, you should be allowed to match it with the group that’s already working on that gene and put it out together.

It’s tough. Boston Children’s is really great in that we have this CRDC, this cohorts committee, where you can see other investigators’ data—patient data and genetic data. Not the phenotype, not their name, or anything identifiable. Sorry, I need to make that extremely clear.

But if you have a gene that you’re working on, you can put it into the CRDC and come up with all the patients who were seen in the hospital and their genetic variation in that gene. The physician has a de-identified ID number, and you can email the physician to find out more information about that patient.

I’ve joined national and international studies that way by having a candidate gene. I go on to the Gene Dx browser now and query the entire hospital—everyone who’s been sequenced and has their data up there. I found 4 other patients, emailed the investigator, and they were like, “Oh yeah, this person in the Netherlands is putting together a case series. Email them.”

I got my patient’s information into that case series, and now it’s awesome. They’re linked to experts, and we’re publishing an accurate, comprehensive view of what that condition looks like. But it’s hard. I understand the dilemma, and for the young investigator who really just wants to get credit for what they’ve been working on, they don’t want to hand everything over. But it’s important that they do, and that everyone does.

Nathan Labenz

You’re identifying a barrier to progress here that I had not even considered, which is the investigator holding information more closely than it sounds like they should in some cases. I guess if we were to imagine an ideal data-sharing scenario, exactly how do we square the circle on sharing versus privacy? That’s obviously a tough question.

Maybe there’s a cryptography-based solution that we could imagine, or maybe we just need to change our norms a little bit around how willing we are to share genetic data. I’ve always felt like it doesn’t seem to me like a huge risk that I’d be taking to share my genetic information with some international pool of information.

There are multiple different angles here, but I guess I’m wondering: If we were to move from today’s data-sharing reality to an ideal data-sharing reality, how much of a difference would that make for people who have these rare diseases?

Dr. Catherine Brownstein

I’m just spitballing here, but I think it would be huge. I think there are a lot of cohorts in the back of the freezer that just haven’t been sequenced and haven’t been shared, more because—not apathy, but because—it’s harder to do so. Also, sometimes at a very superficial level, it’s hard for the investigator to get there mentally and do that.

But I think if they did, there would be a lot more discoveries and a lot more diagnoses for patients, that’s for sure. That’s why I always tell patients—or people, if they email me and they’re like, “Okay, my child has this,”—“Well, here, enroll in this program and this program and this registry.” And they’re like, “Why not just one?” I’m like, “You want to do as much as possible.”

Registries are really important because when there’s a new discovery, they go straight to the registry to find patients. That way, you’re ensuring that your sample isn’t being left in the back of the freezer until they get to it, because you’re just hitting it from multiple sides, multiple angles.

Nathan Labenz

Yeah, is this sort of akin to—I mean, there are a few of these pivot points, maybe, in the medical system where a lot of data is, of course, locked up in electronic health records, and we sort of have this nominal interoperability requirement that somehow gets cashed out as everything getting faxed around. It’s like, “What the hell is that?” That seems like not what we intended, and yet it hasn’t been fixed.

Then there’s price transparency, which is outside the scope of this conversation but is definitely the kind of thing people have high hopes for. If you could get a price menu on the wall, maybe that would help in certain ways. There’s also right to try, which is a big movement where people are like, “You’re not going to let me try this experimental drug even though I’m dying? I should have that right.”

This feels like it could be another candidate for similar reform. If I was going to try to whisper into somebody in the new administration’s ear, I might say, “Hey, look at the requirements around sharing this information. Could we change the defaults here in a way that would move the needle in a big way?”

Dr. Catherine Brownstein

It’s interesting that you say that. Going back to 2011, one thing I lobbied for was shifting it so that being in the biobank—your samples, your discards, tissue, urine, anything that wasn’t used that they took from you—was an opt-out, not an opt-in. I still think it’s an opt-in, how many years later. There’s a lot of inertia around this: being able to facilitate broad sharing, especially for these cases where privacy isn’t really the number-one thing on anyone’s mind. It’s about moving as rapidly as possible and making as many discoveries as possible in a short amount of time.

I really think decreasing the barriers to sharing and to right to try is important. Mew who made Milusen, is 2 floors down from where I’m sitting right now, and it’s just this incredible story of him seeing an opportunity to make an N-of-1 drug and an extremely motivated family breaking down barriers to make it happen. They were so brilliant, motivated, and smart, and they were able to do it. You just think, “Okay, if you made the hurdles less extreme, how much more would be possible?”

Nathan Labenz

That’s an incredible story. If you don’t know it yet, I don’t know it, but here’s hoping that we might have fewer of those stories and more healthy defaults going forward.

Dr. Catherine Brownstein

Those stories are inspirational, but they sort of represent the dark matter of probably 100 other families that just couldn’t, for some reason, overcome those barriers. Some things are just so simple and maddening. We have a bunch of cases at the Manton Center where we find the diagnosis, and then we need to get it confirmed. We do stuff in the research realm, and then you have to get a new sample and verify it in a specialty lab, a Clea accredited lab, and then have the finding returned to the family through a genetic counselor or physician.

Sometimes we’ll call the physician and they won’t play ball with us. They don’t care, they don’t want to deal with it, and they don’t see the value or what it’s going to change. In my own family, I haven’t been able to Clea confirm a finding in one of my relatives because the doctor is like, “Well, I don’t have email. What’s the value of this?” It’s just like, “Oh, my God.” This is what we’re up against.

Then you multiply that by people not counseling correctly and not getting the families into research programs. As hard as we’re trying, there are still so many barriers. To bring it back, I’m really hoping that AI can break some of this down and put some of the autonomy and our ability to act into the hands of families and patients so that they’re less reliant on some of this infrastructure that doesn’t work as well as it should.

Nathan Labenz

Yeah, I mean, this is an eye-opener for me. I think often about whether we’ll end up in a similar spot with respect to AI as we seemingly have with respect to nuclear power, where somehow we have thousands of nuclear weapons deployed, but we’re still burning a lot of fossil fuels because we haven’t been able to get nearly as many nuclear reactors as we have nuclear weapons. Something seems very off about that outcome.

I can imagine an analogous version for AI where we sort of have what we need, but through a combination of errors, barriers, and abstinence, we never quite get to the actual benefits that we could get. It sounds like there is definitely some work to do here to make that change in this area.

Dr. Catherine Brownstein

There are a lot of rabbit holes.

Nathan Labenz

Yeah. So how do you think this changes going forward? We could talk about this from the patient level and what they can do. I always say that if it’s me, at this point I would go with both the human doctor and the AI doctor. I would always have the conversation with Claude or ChatGPT in advance. If they don’t want to talk to me, I say, “I’m preparing for a conversation with my doctor,” and that gets them to open up and not worry about providing unlicensed medical advice.

The patient experience could be quite different. You could talk about that. I’m also really interested in how you’re applying these latest models in your own work—where they’re saving you time and what they’re allowing you to do that you couldn’t do before. Pick your favorite approach for that, but I’m definitely interested in the AI-enabled future of all this.

Dr. Catherine Brownstein

This isn’t really that crazy or anything, but I’d say the biggest impact AI has made on my research is summarizing articles and genes. Being able to eliminate the time I spend going down rabbit holes—looking up a paper, realizing it’s paywalled, logging into the Harvard library, getting the paper, skimming the abstract, and finding that it’s not at all what I want—has changed my life. Being able to ask for a summary and get it, and either be like, “Oh, yeah, this sounds good,” or move on with my life, has given me hours back in a day.

I think there’s going to be a whole host of new tools, or new reasoning. I find it funny that sometimes it will clamp up and doesn’t want to do something because you’re getting too close to medical advice. Maybe just because there are specialty things that help, it would be really cool if you didn’t have to ask the same question 4 times to get it to answer.

Boston Children’s also launched ChatGPT behind the BCH firewall, which is great because then you’re not worried about things going out, and they’re able to maintain much more control. It stays much more accurate. I still can’t get citations to work properly, which is kind of hilarious, but it’s getting way better. The hallucinations are getting way better. I just think it’s going to be moving at light-year speed.

Going back to what we were talking about before, I think there’s a lot of fear around it that’s going to have to be addressed. Hopefully, the 1 bad situation isn’t going to be the only thing people read about it, and some of the really great things that come out of this will also be properly publicized to give a more balanced viewpoint. Again, keeping in mind that a lot of times these are really severe cases and really severe patients, they’re making huge strides and having a huge impact. Keeping that in perspective is really important, too.

Nathan Labenz

Tell me more about some of the things that you actually throw into ChatGPT. You mentioned 1: here’s my situation and here’s this paper, almost like relevance filtering—“Is this relevant?” What other sorts of tasks do you find yourself bringing to especially the latest models?

Dr. Catherine Brownstein

I also run the core facility here, so I’m tasked with learning a lot of new genetic techniques really quickly. If something comes up and I don’t know what they mean, I could Google it and find the 1 obscure paper. I could put it into ChatGPT and learn about this new type of sequencing that’s only launched at Children’s and has 1 paper attached to it, and get a nice summary that I can understand as opposed to weeding through everything.

I meet with investigators all the time, and being able to summarize their work really quickly allows me to do a much better job in my one-on-one consultations than I would have otherwise. Also, considering there are close to 20,000 genes, anytime I get a case where they think it’s this, sometimes I know what that is, and other times I don’t. I’m able to print out a summary of the condition really quickly and nicely, get the latest information on it, see who’s working on it, and go into a meeting much more prepared in much less time.

Also, when you get a paper back, a lot of times—­for some reason, it always seems to be reviewer number 2—is like, “There’s a whole body of literature on this,” and you don’t really know what they’re talking about. Being able to address some of the critiques and put them into context is really helpful.

I mean, it's all cutting down on this mundane, time-consuming, really tedious part of the job, and I'm getting back to the fun part, which is gene discovery and going through a list of 20 possible candidates and narrowing it down to 3 that you're going to present in an hour and a half. True story.

Nathan Labenz

So, what's that true story, maybe in more depth? Is that another thing where you're using ChatGPT to help?

Dr. Catherine Brownstein

Yeah. Why not? If you have 20 genes and you have the phenotype, and they all seem pretty interesting, you can go through and look at the protein impacts, so order the CADD scores or conservation and be able to order it that way. But then doing a really quick relevancy assessment using ChatGPT saves a lot of time.

Nathan Labenz

So, how do you set that up? Do you have a prompt template that you go back to over and over again? How much have you had to develop that? How much do you have to give in terms of detailed instructions or examples?

We're getting into the nitty-gritty here, but this is the part where I think both people can hopefully learn from your experience. If nothing else, demonstrating that this is possible is quite useful, because there are just so many people, including software developers. You'd be amazed—maybe you have seen this—but you'd be amazed by how many software developers tried GitHub Copilot 18 months ago, when it first came out with the GPT-3.5 model behind it, and were like, “Eh, it wasn't that good. It can't help me.”

So, I think there's just a lot of value in object lessons of, like, here's hard work that highly skilled, highly educated professionals are doing that ChatGPT—or obviously other models, perhaps similarly, but we're focused on ChatGPT in this case—can really help with. So, yeah, I love just as much detail as you can get into in terms of how you actually go about setting these things up, how you've iterated on them, et cetera.

Dr. Catherine Brownstein

Okay, so, for an example, I work on bladder pain, undiagnosed bladder pain in individuals. It's really severe. Sometimes they can't leave their house. It's called interstitial cystitis, bladder pain syndrome. There's no real gene attached to it. We've found a couple of genes where it seems like there's way more variation in those genes than you would expect, given the general population. So, it's a candidate gene. It's in no way a slam dunk, but I have around 500 patients in a cohort with that.

I've done, in conjunction with Josh Motalo at Columbia and Ali Gharavi, assessments of my cohort and other cohorts to see what genes have way more variation in them than you would expect. You can come up with lists, and then you can also come up with gene pathways, like multiple genes. These pathways are interesting because a lot of times they have a label, like the small-molecule transport pathway. There's like 12 genes in it.

Then you want to know: Are any of these genes tied to bladder pain? Are they tied to the bladder? Are they tied to bladder cancer? Are they tied to anything? Being able to ask those questions really quickly—and sometimes it's a simple yes or no, just putting them in a string and then coming out with yes or no—saves a huge amount of time.

And then the ones that are yeses, you can drill in. I always check the notes, too, just in case. It's still early yet, but I was doing that last night and was able to get through these pathway lists and be like, all right, this one has 60% of the genes that have a tie to bladder cancer—specifically bladder cancer—which means that they're expressed in the bladder and there are known perturbations that cause bladder dysmorphology or bladder conditions. So, this is more interesting than anything else.

One, I almost screamed because the gene was linked to urothelial issues, which is a great mechanism of disease, and I'm definitely going to follow up on that. I have a meeting tomorrow morning to discuss it. So, it really just helps. I only started working in genetics after the genome was published, so I don't know how people did it beforehand, and I think there's going to be this whole generation of geneticists who aren't going to know how things were done before all this was available, because it's going to be a huge game changer and time saver.

Nathan Labenz

So, how much difference would you say you see between, for example, GPT-4o, o1, and o1 Pro when you bring those kinds of questions? Because I can see interesting different trade-offs, right? In ChatGPT today, if I recall correctly—maybe they've just updated this—but certainly with GPT-4o, you can enable web search. With o1 Pro, search is unavailable.

Dr. Catherine Brownstein

That's what I thought, and that is still the case.

Nathan Labenz

So, if you have these sorts of questions, GPT-4o could go out online and find information that's maybe more recent than the knowledge cutoff, which could be really useful, but it isn't going to reason about it in the same way. With o1 Pro, you have more reasoning, but you have knowledge-cutoff issues and an inability to go out and supplement at runtime.

Do you have a taxonomy of what models you use for what things, how you know when to trust what it's saying versus when you need to fact-check, and how much the reasoning adds over 4o for your purposes?

Dr. Catherine Brownstein

I think I'm becoming more and more convinced over time that this is going to revolutionize things. I was skeptical at first. I was like, oh, we're going to have to check every single thing. Is this actually saving any time? It's just getting more and more accurate. The reasoning is getting better, and sometimes you'll be so pleasantly surprised.

You'll ask it a question, and it'll say, “Okay, answering it in the form of a genetic counselor is this.” And then it'll completely surprise you and be like, “Another way to look at it is this.” It's doing an amazing job.

I know I'm a convert and a relatively early adopter, but I think the sky's the limit, really, and we're going to get to a place where it's going to be solving cases, shortening the diagnostic odyssey, democratizing access to genetic interpretations, and sidestepping a lot of the barriers that we have right now.

It just needs to convince everyone that it's accurate and that the reasoning is good a high percentage of the time. It's kind of hypocritical, in a way, that I think we're going to have a higher bar for it than we do ourselves. We can say, like, “Oh, sorry, I missed it. I shouldn't have,” and we're not going to forgive it if it misses something. I guess that's the way it should be.

Nathan Labenz

Yeah, I'm not sure if that's the way it should be. It does seem like it's the way it is. In self-driving, my general working assumption is that it's going to have to be 10 times safer, or have 1/10 the danger rate, to be acceptable to people. I would guess probably something similar will happen here, at least when it comes to actually putting it in a more forward-facing role where patients could access these sorts of things themselves.

If it's a tool for the professionals, then maybe we're a little bit more—put the responsibility on the professional—and can use it earlier. But, yeah, I would probably advocate for going for it before it gets to 10 times better. Nevertheless, that does seem like the sort of mentality that we have.

So, just honestly, for me—maybe for the audience, but for my benefit—how are you managing those trade-offs between needing to go out and search? Because if you wanted to use o1 Pro, you'd have to go do your own search, copy and paste it in, and let it do its thing. GPT-4o can do its own web search. So, in the very nitty-gritty, what model do you go to, and how do you set it up for success?

Dr. Catherine Brownstein

I'm not using the web search right now. I'm more using o1, I think. I mean, I'm playing with 4o. It's moving so quickly that there's no sophisticated reason for that. It's just what I'm comfortable with, and then moving from there.

I'm really impressed with 4.0 reasoning. I think web search still kind of scares me a little bit, just because there's a lot of garbage on the web, and I have to really be confident in any answer I'm getting out. I think checking everything is still paramount here, but hopefully it won't be that way in the near future.

Nathan Labenz

How long does it tend to think on the questions that you're giving it?

Dr. Catherine Brownstein

At first, I think it was shortening, too, by the day. At first, I remember it would just be hanging there for a while. I'd be like, “Are you okay?” Now it's just really fast.

Nathan Labenz

Or maybe under a minute in most cases, it sounds like.

Dr. Catherine Brownstein

Also, I think I'm getting better at the prompts. As you said, you have to learn how to ask it things, too, for it to come out with the right answer right away.

Nathan Labenz

Yeah, I would love to—we'll trade. One thing that I imagine you've probably also found, but I've definitely found, even in low-stakes situations, is that I try to be really neutral in the way that I ask questions. One of the most common failure modes, at least from what I experience, is the model running with a preconception, which might have been a misconception on my part, and mirroring that back to me.

I'm not doing genetic analysis, even in terms of how to solve a programming problem or how I should think about architecting my application or whatever. A lot of times, if I give it a sense of where I'm leaning, it will lean in that direction, too, perhaps without good reason. So, that's one. What else have you found to be important in prompting?

Dr. Catherine Brownstein

I've found that it actually thinks a little too much, where I'm like, “Is this gene related to this phenotype?” It'll bring up a study, and I'll look at the study, and it's the gene related to it.

Is this step too far down the chain? It's really impressive because it made that intellectual leap, but I need something simpler. I need a paper that's just linking that gene to that phenotype. I was like, “No, too far, too far.” How do I ask it so it's not thinking as much?

Nathan Labenz

And when you describe that, it's bringing up a study out of its pretraining knowledge. You give it a question, and it says, “So-and-so et al. found this.”

Dr. Catherine Brownstein

Yeah.

Nathan Labenz

It sounds like it is also marshaling knowledge of the graph of interactions and saying, “Well, this paper showed this, and then I know from other...” And I'm like, I'm not making that case in the paper. I just want you to say it's upregulated in cancer. That's all I want.

Dr. Catherine Brownstein

That was actually kind of wild because then you're like, “Okay, it's thinking. It's really thinking and making conclusions.”

Nathan Labenz

That is quite interesting. Do you think those things are real and we're just not there yet, or is it going off in a direction that is fundamentally not super productive when it does that?

Dr. Catherine Brownstein

That's the million-dollar question. I don't know. Maybe I'm not smart enough to understand it, and it's right. I don't know. We'll find out. We just need to keep playing with it, keep working with it, and keep using it.

Nathan Labenz

It sort of is like the Move 37 equivalent.

Dr. Catherine Brownstein

Yeah, of course.

Nathan Labenz

I'm sure you're familiar with the AlphaGo championship from years ago, where Move 37 is AI shorthand for an output from an AI system that is very surprising to human experts and nevertheless proves to be a genius move. It was one of these moves where it was like, “Oh, wow, this thing is playing Go in a way that we never thought Go could or should be played, and we actually have something to learn from this system.” They initially thought it made a mistake, and then it turned out it was a genius move. Do you think there's at least some possibility that some of these weird analyses you're getting back might be Move 37-like brilliance, but we just don't yet have easy ways to resolve whether it's going in the right direction?

Dr. Catherine Brownstein

I have faith in it. I think we just have to keep an open mind and keep playing with it and see what it can do.

Nathan Labenz

How much data do you have? Do you actually throw whole cases in?

Dr. Catherine Brownstein

It's kind of too hard to do that right now. I'm not throwing in a medical record, even if it's behind the firewall. I'm just summarizing. We're building up to see the limits. That's actually part of the grant that I have with OpenAI: to see how far we can take this, how much it can handle, and how much it can replace me.

Nathan Labenz

And the barrier to doing that right now is the context? I could imagine multiple different reasons that it might not work to just take the simplest thing I would try. But I'm sure there's going to be a barrier. If I just said, “Okay, here's my whole medical record and my whole genetic summary of all the strange variations,” took the top chunk of that file, copied and pasted it, and said, “Analyze this,” what makes that not viable today?

Dr. Catherine Brownstein

I mean, the amount of compute needed for that, and then also the opportunity for tangents and all the utility. Not everyone who smokes gets cancer, so we have to figure out what we're asking, what's relevant, what's meaningful, and what's a good use of resources. I think we're forgetting that every question takes energy.

We can't just throw everyone's medical record in there and everyone's genome and see what comes out. We have to be thoughtful about it and see, in the cases where it does have utility, what information is necessary. And also, to your point, I think it'll be really interesting to look at trajectory and predictions. All the medical-record mining we're doing now—can it do it on steroids and come up with predictive models and interject, “Okay, I know normally you would want to see a colonoscopy at 40; maybe you need one at 24,” just based on genetics and everything that it's able to see that we're not smart enough to see yet?

I think the possibilities are really exciting. It's a really exciting time.

Nathan Labenz

Yeah. Sam Altman recently said they're losing money on the o1 Pro subscriptions even at $200 a month. So it sounds like people may not, in general, be conscious of the compute they're consuming and are just throwing a lot at it. The budget's got to be pretty high, right?

Dr. Catherine Brownstein

Yeah. I mean, from the medical system—from what people would be willing to pay, or what insurance is prepared to pay, compared with a $200-a-month o1 Pro subscription—I assume that looks very cheap by comparison with hiring professionals and teams of people like yourself.

Nathan Labenz

Yeah, and people like using it. Everyone I know loves using it. I don't know if that's a random sample. So what else is going on with this grant and your relationship with OpenAI? Are you working with them closely and iterating on use cases and giving them feedback, or what is the dynamic there?

Dr. Catherine Brownstein

Yeah, exactly. They've been wonderful and super cool, and it's fun meeting and working with smart, motivated people. I get emails at 11:00 at night on a weekend. They're working hard; there's no doubt about that.

We're just getting back into it after the holidays, so hopefully, in a few months, I'll have something really exciting to talk about. I'm just blown away that they're so forward-thinking and able to support this type of work. It's happening a lot faster than I would have thought, even just a couple of years ago, that's for sure.

Nathan Labenz

Maybe in terms of wrapping up, if I try to summarize everything here, it sounds like we have a data-sharing problem. We have a limited capacity for analysis as humans. One of those is going to require a non-AI solution; the other one, AI is increasingly ready and able to do a lot of analysis.

But then you still have a number of practical issues around the knowledge cutoff and search. You do your own kind of curating of the search, and you can't throw everything into it because maybe it's a little too big for the context window. You don't have all the workflows you might like because it's all in a browser, so you have to paste stuff in and get stuff back, and it seems like it's still fairly manual.

This is almost like your OpenAI customer interview, but what do you imagine the experience being like, say, a year from now, when we refine things a little bit more and integrate these systems a lot more? What do you think that could be for you and for patients?

Dr. Catherine Brownstein

That's a great question. I think they might wrap things so it's less free-form, so you'll be able to guide patients and make it user-friendly with a point-and-click interface. You're not just working with a prompt and then meant to come up with the correct question to get it to answer what you're thinking about. I think that's pretty low-hanging fruit and will be very useful for patients.

I think for researchers, too, with prompt engineering and use cases, we're still figuring out, at least at our institution, where it will be most impactful and what people want to use it for. So that's going to be clearer in a year. It's just getting people in the door and communicating. We have surveys, like, “Okay, what would you use it for? What have you been using it for?” and clearing that up, and then making that better.

I'm trying to be realistic here. I would love to say that we're using it to solve cases. I don't know if that'll be true, but I hope it is. I think it's just going to be more intertwined in our day-to-day existence. What do you think?

Nathan Labenz

All bets are off. I don't know. I mean, o3—that's the other question I had. Are you on the review team for o3 at this point? If not, I assume it'll be coming your way before too long.

Dr. Catherine Brownstein

Hope so.

Nathan Labenz

Yeah. It looks like that is another significant step up in raw reasoning ability. The FrontierMath results in particular were notable. Everybody was citing the 25% success rate, but that's the very high-compute level that costs maybe thousands of dollars a problem or whatever.

It was also really notable that the low-compute setting was still 10%, which was 5 times better than anything that had come before, with the previous best maxed out at 2%. So it seems like the o3 series is going to be another pretty serious step change in terms of just how hard of a problem these things can solve.

We might start to get into context-window limits being binding if there's just too much information in a medical history or in a genetic file. But I suspect that those can both be filtered and summarized and boiled down to what matters most, such that even in a couple hundred thousand tokens, which is what they currently have, I would honestly take the “we probably will be solving cases with an o3 in a year's time” side of that—or, for that matter, an o4—because the gap in time between an o1 and o3 was so small.

The signals we're getting are that we don't really see this slowing down. There's going to be more progress on this front. So I'm always kind of wondering what's missing, and increasingly it's harder and harder to find and pinpoint the things that are really missing. That doesn't mean there's nothing missing. I'm sure there are still some things, but it is increasingly hard to say what they are.

My best guess would be that you probably see at least some cases that you could just throw into an o3 and get meaningful, insightful conclusions back in a year.

Maybe we should get together again in a year and review the progress.

Catherine Brownstein

I’d love that.

Nathan Labenz

Cool. Well, anything else on your mind today? I really appreciate the introduction to all your work and how you’re using AI in it, but anything else on your mind before we break?

Catherine Brownstein

Just thank you so much. I think this time next year might be kind of different, so hopefully.

Nathan Labenz

That seems to be the new normal. Change is the only thing we can really count on.

I’ll look forward to putting it on my calendar now to get back together in a year and see where we’re at. But for now, Dr. Catherine Brownstein, MPH, PhD, assistant professor at Boston Children’s Hospital and Harvard Medical School, and a recent recipient of the ChatGPT Pro Grant, thank you for being part of The Cognitive Revolution.

Catherine Brownstein

Thank you.

用 o1-pro 猎基因:ChatGPT Pro 资助项目获得者 Brownstein 博士谈如何推理罕见病 — 文字稿与摘要 | BidClub