Universal Medical Intelligence:OpenAI 提升人类健康的计划,与 Karan Singhal 对谈
OpenAI 正把健康作为让使命落地、将利益规模化普惠大众的重要路径,目前每周已有超过2.3亿人使用 ChatGPT 查询健康与保健问题。 ChatGPT Health 计划接入病历、Apple Health 和可穿戴设备,同时对所有用户免费且不限速;Karan Singhal 表示“目前”没有投放广告的计划,正在形成 Nathan Labenz 所称的“普惠基础智能”的早期版本。
Labenz 表示,前沿模型在他儿子的病例中已达到主治医生水平,而人类剩余优势主要来自具身化的临床语境。 HealthBench Hard 在创建时 GPT-4o 得分为0%,当前 OpenAI 模型已升至约40%,竞争对手目前大致处于“20分上下”。儿子住院期间,Labenz 发现前沿模型“几乎每一步都跟得上肿瘤科主治医生”,但医生仍能直接观察呼吸、肤色和整体状态。
OpenAI 的健康项目依靠专家反馈、评测基础设施和真实工作流证据推进,而不是用联网患者数据训练模型。 超过250名医生——按 Singhal 的估计约260名——通过顾问、持续进行的 Slack 红队测试和紧密合作的研究转译团队参与其中;ChatGPT for Healthcare 在6个月内经历了9轮测试。HealthBench 覆盖5,000段对话和约49,000条评判标准,肯尼亚 Penda Health 的随机研究则显示,诊断和治疗结果出现统计显著改善。
瓶颈正从医学知识转向语境获取、多模态能力和产品整合。 Singhal 表示,除部分亚专科外,模型的文本表现已经很强;但模型在获得完整病历时效果最佳,仍需更好地接入影像、语音、纵向测量数据和细微的生理信号。可能的架构是原生多模态表征结合工具调用、Python 和专用模型:“随着模型智能提升,人们自行收集的数据价值也会不断增加。”
OpenAI 预计临床应用将在2026年加速,真正的约束是固化的工作流,而不是医生保护主义。 ChatGPT for Healthcare 增加了 HIPAA 合规、医学证据检索和临床工作流,首发即接入8家头部机构,收到的需求超过团队承载能力。Singhal 的目标是让 AI 辅助医疗在年末成为“医疗常规的一部分”,但他也提醒,医疗体系按月变化,而不是按周变化。
隐私正在被当作推广基础设施:健康数据获得额外加密,与普通 ChatGPT 活动隔离,也不会用于训练基础模型。 Singhal 并未声称更多数据没有价值;他的判断是,通过清晰的隐私交换降低用户的“启动能量”,将带来更大的长期健康影响。下一步政策前沿可能是经同意的数据共享,用于试验匹配、N-of-1 治疗,以及 Labenz 提出的“AI 与尝试权”。
医疗也是 OpenAI 的应用对齐实验室,但 Singhal 并不认为最难的监督问题已经解决。 医疗模型已经在一些狭窄领域超过单个医生,迫使团队继续推进可扩展监督、校准不确定性、专家聚合和模型性格研究;与此同时,随着强化学习扩展,思维链尚未出现普遍滑向“神经语”的迹象。乐观情景是出现医疗领域的“Move 37”——意料之外的诊断或治疗方案抬高健康水平上限——但能力快速扩张与更罕见安全失误之间的平衡仍难以判断。
1. 健康是 OpenAI 让 AGI 普惠大众的路径
Singhal 大约4年前开始研究医疗 AI,并在2022年底前后转为全职投入。他当时的判断是,AGI 很可能在自己有生之年到来,而自己可以从两方面改善结果:一是“做安全”,二是“做收益”,其中医疗是最直观的收益落点。
在 ChatGPT 出现前,他看到了一个“能力过剩”:经过规模化和指令微调的语言模型已经能完成惊人的任务,但临床 AI 仍把它们视为推测性技术。因此,他最初有两个目标——先说服医疗行业相信 LLM 能够工作,再建立让它们值得信赖所需的安全与可靠性体系。
OpenAI 的健康项目起初有3个目标:让医疗专业知识普惠化,用医疗为安全和对齐研究提供现实锚点,并通过合作伙伴、产品和政策制定者推动社会共同适应。Singhal 的态度转变颇有意味:两年前听起来雄心勃勃的目标,如今“还不够有野心”。
2. OpenAI 依靠医生群体判断,而不是静态规则手册
Singhal 将项目分为3个阶段:打好安全基础、承接快速增长的采用、扩大影响。采用阶段来得很快——健康已成为 ChatGPT 增长最快的用例之一,在专门的健康产品完全铺开前,每周用户已超过2.3亿。
第一性原理式的规范透明,但 Singhal 认为,无论是自己还是 Labenz,都无法可靠枚举医学各种边界案例中的正确行为。“不伤害”划出了重要边界,却“无法告诉你在大多数场景下该做什么”,所以 OpenAI 转而汇总约260名医生的判断。
医生网络分为3层:制定战略的高层顾问;在 Slack 上紧密协作、比较输出、进行红队测试、识别盲点和测试产品的社区;以及一个规模更小的团队,负责将更广泛社区的判断提炼为评测、训练数据和可供研究员使用的行动信号。
ChatGPT for Healthcare 展示了这套流程的深度。医生红队成员在6个月内进行了9轮测试,而更广泛的社区则找出了人口统计、临床和沟通层面的失误,而不是只完成孤立的偏好排序任务。
3. “不伤害”是底线,不是完整的医疗决策政策
Labenz 的反驳值得保留:机构医疗可能变得过度保守。一位医生朋友私下对他说:“我不需要随机对照试验来告诉我这件事有道理。”但医院医生经常因为生物学复杂、无法保证成功,而拒绝合理的假设或行动。
Singhal 认为问题存在于两端:医生将警示信号常态化后,患者可能需要更积极地争取关注;医生则面临证据爆炸、文档负担过重和普通人的能力边界。Labenz 的案例中,异常血检结果和最初的安慰性判断,说明单纯把决定交给现有系统可能失败。
AI 可以在一个语境中整合患者病史、当前状态和最新医学证据。医生可能为便于沟通只给出一条治疗路径,模型则可能列出3至5条——前提是明确说明证据有限,并区分合理选项与已经确立的结论。
因此,OpenAI 的校准工作分别追问两个问题:模型能否识别自身不确定性,以及能否以有用的方式表达这种不确定性。Singhal 认为,这正是模型分享早期证据、又不至于鲁莽行事的机制,也能避免模型过度保守到不提供自己的“最佳猜测”。
4. 可靠性如今衡量的是多次推理后的表现,而非下一个 token 的置信度
旧式校准图会比较模型对选择题 token(例如“A”)的概率与“A”实际正确的频率。Singhal 表示,这种方法已经无法覆盖现实任务:健康问题的答案远比选择题丰富,而且推理模型会在最终回答前先输出思考 token。
HealthBench 加入了细致的临床评分标准;“worst-of-N”测试则通过反复采样模型并记录最弱结果来检验一致性。采样20次,可以暴露不同推理路径是否偶尔会产生明显更差的医疗回答。
对于高难病例,用户可以采样10个答案,通过“LM council”合并,再让另一个模型进行综合。Singhal 预计,这相较于一次强推理运行只能带来边际提升,并表示 GPT-5 Pro 在底层大致执行着类似机制;提高推理投入也能带来相当的收益。
他的实际建议是,健康问题优先使用 GPT-5.2 Thinking,而不是 GPT-5.2 Instant;如今默认推理已经能服务大多数人,只有最难的病例例外。最差的 o3 输出也胜过最好的 GPT-4o 输出;GPT-5 nano 和 OpenAI 的开源模型如今表现接近 o3,而 GPT-5.3-Codex 和 GPT-5.2 Thinking 能以更少思考得到更好的答案。
5. 随着强化学习扩展,思维链仍意外地易读
推理模型带来了一个有用的安全副作用:它们的思考 token 往往用普通英语解释自己在做什么。研究人员可以检查这段轨迹,寻找谋划、不可取的策略,或模型表面推理与最终答案之间的差异。
Labenz 担心,更密集的推理会产生一种内部方言或“神经语”。Singhal 表示,没有大规模证据显示强化学习扩展导致可解释性持续下降;确实存在“奇怪的小波动”,但没有稳健趋势,他也不认为这一有利结果会无限期保持。
人类可读的思考并非训练直接保证的结果。Singhal 的解释是,模型继承了英语语言先验,并使用英语,因为这仍是通向有用答案的最便捷路径;OpenAI 避免对思维链施加优化压力,也有助于保留这一经过实证、对研究有用的特性。
6. HealthBench 把医疗质量推向尚未饱和的前沿
医疗评估已从考试题转向真实互动:Med-PaLM 测试广泛的健康回答,AMIE 探索追问式对话,HealthBench 则评估与普通人和专业人士的完整交流。其5,000段对话包含约49,000个不同评估维度,覆盖事实、不确定性、升级处理、推理和床边沟通。
3个 HealthBench 版本分别对应3个设计目标。完整基准必须有意义——“如果这个数字上升”,现实世界的健康结果就应改善;HealthBench Consensus 必须值得信赖,由多名医生确认评分标准适用;HealthBench Hard 则必须保持挑战性,不能滑向95%至100%。
HealthBench Hard 采用对抗式构建,挑选各家模型表现最差、但本身质量较高的样本。创建时 GPT-4o 的得分“确实是0”;当前 OpenAI 模型约40%,Singhal 认为当前竞争模型大致在20分上下,距离饱和仍很远。
评分层本身也接受了医生测试。医生将 HealthBench 的模型评分器与其他医生的判断进行比较,自动评分器的表现优于平均医生;Singhal 认为,这说明该基准的细化评分具有异常高的可靠性。
7. 在床边,模型达到主治医生水平,但没有具身判断
Labenz 的个人基准来自一次30天、高风险的住院经历,当时他的儿子正在接受癌症治疗。前沿模型“几乎每件事都与肿瘤科主治医生同步”,知识水平明显高于住院医生,既改变了他对治疗的信心,也改善了他的心理状态。
他最初使用 GPT-5 Pro,随后用 Gemini 3 重复提问,再用最新的 Claude 提问3次。AI 彼此之间的分歧甚至少于它们与主治医生的分歧,整个月只出现大约6次争议,通常集中在电解质补充等小决定上。
事后回看,Labenz 对这些分歧的判断大约是“医生六、AI四”:约三分之二的时间里,听从医生似乎是正确选择。这一有限的人类优势并非来自他能上传的信息,而是来自 PDF 中没有的证据。
主治医生实际上可以补充一句:“看着他现在的样子。”医生能观察呼吸、肤色和细微的整体状态。Labenz 也对结论做了限定:他儿子的癌症并不特别罕见,且遵循成熟方案,因此这是一场情感上极其艰难的治疗,但未必代表临床判断的最前沿。
8. 更好的语境如今与更好的文本推理同等重要
HealthBench 测试模型能否在不通过制造恐慌挤兑医疗系统的情况下,对紧急病例进行升级处理。其最大单项重点是全球健康:根据性别、地区流行病学、医疗可及性,以及用户所在地区是否常见结核病等事实调整推理。
模型在不确定时浏览资料、综合最新资源和提出有优先级的追问方面也有所进步。Labenz 学会要求模型采访自己,并判断他能否完成部分体格检查;新系统越来越能主动识别并提出这一工作流。
Singhal 表示,除部分亚专科外,模型的文本表现已经很强。随着 o1、o3、GPT-4.1、GPT-5 及后续版本持续改进,健康工作如今已进入 OpenAI 模型的每个主要训练阶段;下一步限制往往是缺少语境和模态,而非医学记忆不足。
ChatGPT Health 通过接入病历、可穿戴设备和 Apple Health,直接解决语境瓶颈。Singhal 的一般性建议是,模型从更多语境中受益,但他不会声称某一张床边照片一定会改变 Labenz 儿子的病例结果。
9. 医疗多模态将结合工具调用与原生表征
模型长期以来就擅长解决“医生的填字游戏”——所有相关语境、偶尔包括多模态数据,都在一开始完整提供的复杂临床病例。更难的产品问题,是在给出诊断或治疗建议前,通过对话主动获取缺失的观察信息。
生物医学数据存在大量尴尬的长尾格式:千兆像素病理图像、PET 扫描以及3D或4D MRI 数据,不能简单粘贴进聊天框。一条路径是模仿临床医生,借助查看工具检查选定切片,而不是一次性感知整个高维扫描。
Singhal 预计,模型会通过 Python、专用库或专门的感知模型处理这些数据,而其他模态则会像图像、音频和视频现在一样,直接编码进 token 空间。最终的医疗技术栈很可能是混合式的,按模态分别选择方案。
Labenz 提到 SleepFM,称其结合了一个人单晚睡眠期间测量的约6至8种模态,用于预测多种疾病。他的判断是“所有潜在空间最终都会连接起来”;Singhal 同样乐观,但强调时间表取决于各模态的数据可得性、相关性和影响力。
10. 真实部署缩小评测与结果之间的差距
个性化带来了巨大的测试面:即便使用同一个模型,两个用户也可能拥有不同的记忆、接入工具、指令、人口特征和沟通偏好。Labenz 担心,这条长尾会制造迎合等失误,也使得单一的“床边沟通方式”无法被明确规定。
OpenAI 将上线前评测与保护隐私的生产监控结合起来,通过分类器扫描流量,识别健康及更广泛的安全模式。Singhal 表示,在敏感心理健康对话上的工作显示,离线评测结果与真实使用中观察到的模式高度对应。
Penda Health 提供了更强的测试:在节目所称的首个 LLM 临床副驾驶随机真实世界研究中,肯尼亚医生在电子病历中记录诊疗时,获得了一个 AI 安全网,用来标记可能重要或错误的决定。
使用副驾驶的医生所治疗患者,在诊断和治疗结果上出现统计显著改善。Singhal 当前的评估阶梯,从医学考试、HealthBench 的现实离线对话,延伸到前瞻性临床研究和回溯性生产监控。
11. ChatGPT Health 优先建立用户信任,而非用联网数据训练
Labenz 个人更倾向于分享数据:公开讨论儿子的病情让他接触到有用的专业知识,而在他看来,除科幻情境外,很少有其他场景会造成同等程度的伤害。Singhal 的回答有意保持多元——隐私偏好因人而异,因此基础设施必须支持不同选择。
Singhal 认为,10年前,许多患者既无法访问自己的病历,也无法控制病历流向何处。如今访问权限有所改善,但数据仍分散在各家机构之间,患者无法形成完整图景,也无法利用数据为自己争取更好的治疗。
接入 ChatGPT Health 的数据不会用于训练 OpenAI 的基础模型。Singhal 并未说数据已经不重要;他的判断是,现有模型已经很强,而隐私优先的交换降低了用户的“启动能量”,避免那些担心数据被用于训练的人因此无法获得产品价值。
产品增加了健康专用加密,并将健康信息与普通 ChatGPT 活动隔离。健康对话可以受益于用户许可的记忆或已连接语境,但普通对话不会反向获得用户的健康语境。
12. 模型智能越高,量化自我的回报越大
ChatGPT Health 改变了 Labenz 对可穿戴设备的看法:因为 AI 终于可以处理数据,而不必把量化自我分析变成另一份工作,他买了一块 WHOOP。他衡量 AI 的更大指标是:AI 能否让他少坐在办公桌前,多花时间锻炼。
Singhal 的判断异常直接:“随着模型智能提升,人们自行收集的数据价值也会不断增加。”更好的研究和产品能力应能提取今天尚不存在的洞见;因此他说,“如果有过适合买手表的时点,就是现在。”
Labenz 每天用同一批实验室 PDF 比较 Claude、Gemini 和 ChatGPT;在他的测试中,“Grok 还没有进入第一梯队”。Gemini 3 在倡导和安慰方面最为主动,ChatGPT 回答最长、临床中立性最强,Claude 更简短,也比 Gemini 更克制。
Singhal 没有在全面性和易读性之间指定一个普遍赢家。OpenAI 明确训练模型判断用户是专业人士还是普通人,并调整术语和细节程度;他表示 GPT-5.2 Thinking 在这方面有所改善,而竞争对手也在帮助推动医学更广泛的“奥弗顿窗口”走向接受。
13. AI 可能迫使医疗围绕实验性治疗与数据建立新交换
Labenz 设想,当 AI 识别出一种实验性治疗,认为它是患者剩余的最佳选择,而医疗体系仍拒绝提供时,压力会不断上升。他提出的“AI 与尝试权”交换,是扩大治疗准入,同时系统性记录每位患者的结果。
这可以支持 N-of-1 决策——基于现有证据对个体患者的最佳判断——而不必取消临床试验。关键是将结果回流到集体医学认知中,让人类和模型能够从那些原本隐秘发生、甚至根本不会发生的干预中学习。
碎片化伤害所有参与者:患者无法整合自己的病史,医生无法跨系统调取完整记录,研究人员也难以找到符合条件的参与者。Singhal 表示,招募困难导致大多数临床试验在完成前失败,因此在获得同意的前提下进行匹配,是加快研究的直接路径。
Singhal 保留了制衡条件:传统试验标准“经过实战检验”,实验性治疗准入必须建立在主动且知情的同意上。他偏好的顺序是先以隐私优先方式部署,再增加同意机制,让患者自愿共享数据、寻找试验、探索治疗并推动科学进步。
14. 健康服务正与广告和订阅收入脱钩
Labenz 的40页住院同意书,以及明尼苏达大学后来发来的一份看似重复的请求,展现了今天公益型数据收集的摩擦。即便参与者高度积极,也可能无法寄回材料,说明依赖反复填表和人工跟进的系统产出有限。
Singhal 在变现问题上划出清晰界线,同时保留了重要的时间限定:“广告不会进入 ChatGPT Health,而且我们目前没有这个计划。”其逻辑是,将健康影响与用户有理由怀疑会扭曲医疗建议的激励机制分开。
ChatGPT Health 还计划对所有用户免费且不限速,这是产品其他部分无法提供的访问级别。Singhal 承认实际限制可能出现,但眼下看不到必须依赖广告或以数据换订阅的交易;Labenz 将这一模式称为“普惠基础智能”。
15. 医生采用 AI 是工作流问题,而非保护主义问题
消费者采用医疗 AI 的速度快于医生,这意味着医生往往先通过患者接触到医疗 AI。Labenz 的经历很典型:最初的怀疑,可能在他向医生展示模型的实际回答后迅速消失,对方会说:“哦,还不错。”
ChatGPT for Healthcare 在 ChatGPT Health 之后于1月8日宣布,是面向机构的对应产品:符合 HIPAA,专为临床医生打造,配备医学指南检索和企业级写作工作流。产品首发接入8家头部机构,收到的兴趣超过 Singhal 团队的承载能力。
Singhal 表示,医生保护主义远没有 Labenz 预想的严重。亲自使用过模型的医生和医疗系统管理者往往很容易沟通;真正的约束不是缺少自上而下的需求,而是固化的工作流无法以软件行业的速度改变。
Penda 的部署需要主动变更管理,包括培训课程、同侪学习和帮助医生将工具整合进实践。尽管如此,Singhal 仍预计这会成为“医疗史上更快的软件部署之一”,并希望 AI 辅助医疗在2026年底前成为常规的一部分,但不会在几周内完成。
16. 医疗同时也是具体的对齐实验室
Singhal 加入 OpenAI 的部分原因,是中长期对齐研究过去常停留在玩具环境或数学问题中。医疗提供了即时利害关系、专家反馈和短期激励,把关于可信超级人类系统的抽象问题,转化为研究人员必须为已部署产品解决的现实问题。
可扩展监督已经在医疗中以缩小版形式出现:模型可以在狭窄任务上超过医生,但仍需要医生监督。当单个专家无法可靠评估每个输出时,训练和评测就必须汇总多名专家,提高他们批评模型的能力,并识别任何单个审阅者都看不见的失误。
因此,HealthBench、医生网络和高算力强化学习既是医疗工作,也是对齐工作。目标是从专家判断的集体结构中学习,而不是假装一名医生或一份书面规范就能完整定义安全行为。
Singhal 最终将问题概括为模型的人设、性格或“灵魂”。随着模型获得更多患者语境并拥有更多行动能力,决定性问题变成它们能否稳定地做正确的事;训练必须精准保留有用的主动性,同时避免鲁莽自信和反射式过度保守。
17. 可扩展监督正成为普通的后训练工作,但远未解决
Singhal 如今将可扩展监督分为“评分者扩展”和“价值监督”:前者是从人类和专家那里获得更好的判断,必要时借助 AI;后者是投入足够算力,将这些判断灌注进模型。术语发生了变化,但原始研究问题仍在。
规范、宪法式规则和系统化遵循训练,都是价值监督的例子。进展越来越像“平淡无奇的后训练”,而不是一次独立的超级对齐突破,但 Singhal 对这种落地化趋势持积极态度:系统卡显示,模型对选定人设和安全要求的遵循能力及更广泛泛化都在改善。
Singhal 指出,另一个相关的乐观理由是,批评或辨别似乎比产出正确答案更容易。模型往往可以监控自己,尤其是在获得思维链特权访问时。讨论也涉及这样一个问题:能力较弱但拥有更多推理算力的模型,能否帮助监督更强模型;但 Singhal 并未声称这一扩展机制已经被理解。
Singhal 没有把闭环画上句号:让可信、能力较弱的模型对齐能力显著更强的系统,仍未被充分理解。他对更大推理规模下安全泛化增强感到鼓舞,但能力面不断扩张、单项任务失败率下降所产生的净结果,仍然“难以预测”。
18. 医疗领域的“Move 37”将目标从抬高下限推向抬高上限
Labenz 追问,医学领域的 AlphaGo Move 37 会是什么:一个人类绝不会做、后来却被证明极其高明的选择。他还提出,可以训练系统预测患者在变量改变后的未来健康状况,从而实现干预预测或计算机内实验,而不只是回答问题。
Singhal 认为,原始版本可能已经出现:患者看过多名医生后,ChatGPT 标出一个被忽略、最终导向诊断的可能性。这是否算 Move 37,“见仁见智”,但随着模型更好地理解个体生物学和纵向健康轨迹,更清晰的案例应会出现。
完整患者模拟成本高昂,也不是当前模型的主要设计目标。更可行的路径,是基于丰富语境预测结果或干预效果;Singhal 认为,预训练、推理式训练、增加模态以及现有范式的延伸,仍然“有相当多的潜力”。
OpenAI 的3条路径分别是消费者健康导航、AI 辅助医疗系统和抬高医学上限的研究。Singhal 的总结是,许多生物学突破并不缺少物理前提,只是缺少足够的聪明才智;连接正确数据的长期运行代理,可能创造出一名 AI 生物医学科学家,与 AI 医生并行存在。
Today, my guest is Karan Singhal, who leads health AI at OpenAI and was just named to the TIME100 Health list for his pioneering work. This episode began to come together last year on Thanksgiving, when I emailed Karan, whom I'd met a couple of times at AI events, to thank him for all of his work on AI for health and let him know what a difference ChatGPT had made for me and my family in the context of my son's cancer diagnosis. As it turned out, that was just as OpenAI was preparing to make a major product push with ChatGPT Health, which allows users to connect ChatGPT to data sources including electronic medical record systems and consumer wearables, plus a physician-facing ChatGPT for Healthcare, both launching in early 2026.
In this episode, we dig into how Karan and his team have achieved attending-physician-level performance with their latest models, their plan to ensure that this capability benefits all of humanity, and their vision to raise not just the floor but also the ceiling of human health with continued research and even better models to come. Highlights of this conversation include how OpenAI works with more than 250 human doctors to ensure accurate, robust, and culturally appropriate responses; how they built HealthBench, which contains some 49,000 evaluation criteria to measure models' performance; and how models have already gone from a 0% score on HealthBench Hard by GPT-4o when the benchmark was first created to 40% today.
We also discuss my experience using large language models to navigate a health emergency, including the critical importance of giving models as much context as possible about your situation, and how that's about to get dramatically easier as ChatGPT Health rolls out globally. We discuss how 230 million people are already using ChatGPT for health questions on a weekly basis. We cover the first randomized trial of AI copilots for physicians, which OpenAI conducted with Kenya's Penda Health system and which did show a statistically significant improvement in outcomes for patients whose doctors used AI. And we discuss why Karan believes, based on the reception OpenAI is getting from health systems, that 2026 will be the year using AI becomes a standard part of medical practice.
From there, we cover the steps OpenAI is taking to ensure the privacy and security of users' health information, and how they're using defense-in-depth measures to make sure models first do no harm while striving to maximize value by training AIs to acknowledge their uncertainty as they offer their best guesses. We discuss how Karan understands the relationship between AI for health, AI safety plans such as scalable oversight, and AI alignment more broadly. We also discuss Karan's report that OpenAI's models' chain-of-thought reasoning has not drifted toward neuralese as much as some reports had previously caused me to believe.
We discuss the future of medical multimodality, which will do a much better job of converting data to value and which inspired me to buy a Whoop wristband to start collecting data on myself. We also cover the compounding effect of parallel advances in AI for science, the growing potential for N-of-1 treatment plans and medical Move 37s, and the possible need for an update to the rules governing access to experimental medicines and information sharing.
Finally, Karan describes OpenAI's utopian plan to make ChatGPT Health available to all users globally, for free, with no ads—an early form of universal basic intelligence that I really think everyone ought to celebrate as a triumph of human ingenuity and goodwill. Zooming out, in the grand scheme of AI development, I think it is fair to say that we have far more questions than answers. In my mind, all outcomes, from a post-scarcity utopia to literal human extinction, absolutely remain on the table. I signed up for a recent call for a ban on superintelligence because I do worry that an AI arms race, driven by recursive self-improvement loops, could easily get out of control.
And yet, at the same time, capabilities like this, which have been so valuable for me and my family and which will undoubtedly save millions of lives in the coming years, are for me both an incredibly inspiring accomplishment and a practically irrefutable argument for the upside of AI. The question at this point is not whether we will create powerful AI systems, but exactly what form they will take and under what circumstances and incentives they'll be developed and deployed. Karan's work demonstrates that, for the moment at least, we can have it all: AI systems meticulously crafted to minimize downside risk that are both capable and efficient enough to meaningfully improve the human condition globally. There is a ton of work left to be done, both inside and outside of the frontier companies, to make sure that these lofty standards don't slip in the face of intensifying competition.
But today, if you or a loved one are facing a complex health challenge, you owe it to yourself to take full advantage of the incredible medical expertise that Karan and others have managed to build into systems like ChatGPT Health. With that, I hope you enjoy this inspiring look at the frontier of medical AI with OpenAI's head of health, Karan Singhal.
Karan Singhal, head of health AI at OpenAI, welcome to the Cognitive Revolution.
Thanks for having me. I'm super excited about this.
It's rare that I've had so much impact from a guest's work on my life as I've had from the impact that your work in health at OpenAI, and at Google before that, has made on the last 3 months for me. Regular listeners know the story: my son got cancer, and I've been an intensive user of all the frontier language models over the last 3 months to advise us as we've gone through this process. Boy, have they been a game changer.
Thank you for all your hard work and for making the last 3 months, on the mental-health side of the equation for me, dramatically better than they otherwise would have been. You've also moved the needle a bit on my son's treatment and our confidence that we were actually doing the right thing for him. It has been invaluable, and the consumer surplus has been off the charts.
Amazing. Thanks for the kind intro, and thank you for sharing the story. I think it resonated with a lot of people, so thank you for that. A big takeaway, I think, of this conversation will be that if you find yourself in a medical emergency, or even just want to do a better job of managing your health in general, the frontier models today are getting really good at that.
I always go to the example of the AI doctor. There's obviously a relative scarcity of medical expertise, even in a country like the United States, for a privileged person like myself. You broaden your worldview and look around the world, and the shortage is extreme. I've always felt like, boy, this would be just an absolutely killer use case that everyone could agree on.
When I started talking about it a couple of years ago, it felt like it was getting close but still a ways off. I've done a bunch of episodes over time with Vivek and some of your former teammates at Google DeepMind who work on similar topics. They're always very appropriately cautious, or have been over time, saying, “Yeah, it's not quite there yet. We're not quite ready to roll it out to production,” but encouraged by the progress, that sort of thing.
But, again, it is really getting there. The first question I wanted to ask is: how did you get into this, and what were you expecting? How big a dream were you daring to dream when you first got into AI for healthcare some years ago?
Yeah, I started working on AI for healthcare about 4 years ago and transitioned to doing it full-time around that time. I was thinking a lot about a few fundamental research problems when I was at Google. A few problems were around foundational work in representation learning and privacy-preserving learning. There were a lot of interesting applications in healthcare.
But I think the background for this was a conviction that I had since undergrad that AGI would be a pretty big deal and that it probably would happen within our lifetimes. I thought there were probably 2 things that I could do to make that better. One was to work on safety, and the other was to work on benefits.
I saw healthcare as the most obvious area for benefit, like you. And like you said, we've been on this amazing exponential curve over time, both with model capabilities and, I think more recently, with people's adoption and the shift in the Overton window around trust in these models. We're seeing a bunch of people start to use these models across individual patients and individual conditions, and researchers are seeing a lot of the benefits become much more tangible.
And all this comes down to, for us at OpenAI, thinking about what it means to make our mission real. Our mission is to ensure AGI is beneficial for all of humanity. There are three parts to that: one is to build and deploy AGI; the second is to prevent downside risks, whether those are short-term risks or long-term frontier risks; and finally, to think about how we can make benefits happen. And like you say, I think health is one of the most obvious and tangible ways those benefits can happen.
I think I started working on health around the end of 2022; that was when it became full-time for me, right before ChatGPT. For me, it was a lot of thinking around this capability overhang between where LLMs were at—you were scaling these models up, seeing that you could instruction-tune them and they could do amazing things—and how they were being adopted and thought about in the clinical AI and healthcare worlds.
My ambition at that time was really two things. One was to get the healthcare and clinical AI worlds to think about LLMs as something that could work. Again, this was prior to ChatGPT. The second was to think about the work we would need to do in safety and reliability to make the models trustworthy for the setting.
Then over time, I think those ambitions became less and less ambitious, and my ambitions at OpenAI became larger. For our work at OpenAI, we started out with three goals. The first was to make access to medical expertise more universal. The second was thinking about the ways in which we can think about this setting as a way to advance our work in safety and alignment. I can talk more about the safety motivation there.
The third thing was thinking about how we can bring society along with this high-stakes technology: working with partners, rolling out products, and working with policymakers to think about the right ways to iteratively deploy a setting like this. I think, over time, these things, which sounded really ambitious about 2 years ago when we started setting out to work on health at OpenAI, now actually feel like they aren't ambitious enough. So we're really excited about what's to come.
I think you did a great job there of laying out a taxonomy, or a sort of scaffold, for this conversation. Let's maybe talk about the capabilities first, but they are sort of inseparable. I'm wondering, from a capabilities perspective, is there also a Hippocratic Oath kind of mindset that you bring to the table, focused on making sure that the AI performs well in medicine, that's maybe still distinct from the bigger-picture safety agenda that motivates you?
How do you think about the way the model should perform in terms of doing no harm? I think GPT-4, while it did add value, also could definitely do some harm by giving you wrong ideas. I do think we've come a long way since then, but I wonder how you think about that.
100%. The way we think about the health work that we've been doing at OpenAI, we've been operating in three phases. The first is laying the foundation for the work. A lot of that has been around safety research and making the models not just better reasoners, but also perform better, have better bedside manner, convey uncertainty well, and escalate to a doctor when needed. I can talk a little bit about that work in a second.
We're now in this phase of adoption. As the foundations have solidified, a bunch of people have been using it. It's been one of our fastest-growing use cases. We shared recently that over 230 million people a week are using it for various health- and wellness-related queries. This year, we're really focused on scaling the impact of the work, and all of that, I think, comes from the foundation.
Your question of how we think about imbuing the models with the right kind of prior for how to behave and how to ensure that minimal harm is done—I think we were very thoughtful about this as we were laying the foundation for this work and starting to work on health at OpenAI. A lot of this comes down to our really close partnership with this cohort of 260 physicians or so that we've been working with for about 2 years.
One way of imbuing models with a certain kind of behavior is to write a spec from first principles. Maybe you or I just sit down and write down a spec and say, “This is how a model should behave. It should say X, Y, and Z in this situation, and this in another situation.” There are pros and cons to this. One pro is that it's very easy to explain and understand what's going on, and it's easy to be transparent about model behavior.
A con, though, is that it's very hard for you or me to say what should or shouldn't happen, and it's difficult to say anything beyond “do no harm.” That's an excellent thing to avoid, but it doesn't tell you what to do in most scenarios. What you want to do is figure out a way to move from really large-scale principles that matter to what you do in very specific scenarios, and how you make sure that's guided by the expertise of not just 1 or 2 people, but actually hundreds of experts.
That's the approach we've taken, and you can see it in our approach to evaluation. One way to think about how we encode model behavior in models is: What are the evals that we care most about? Evals are really the lifeblood for any researcher, and so we recently put out, back in May of 2025, this work on HealthBench, which is an evaluation of how large language models perform in realistic health conversations between users and models, where users could be either lay users or health professionals.
We went about it in a way that leaned on the expertise of these 250-plus physicians, rather than writing down specs from first principles. I can explain that more, but I'm sure you have many more questions as well.
Yeah, go on as long as you'd like. To double-click on the physicians for a second, what's the nature of the relationship with them? I could imagine anything from them sitting there doing side-by-side, RLHF-style comparisons—“prefer this to that”—to some of them being much more deeply integrated with the research team.
We have three different layers through which we think about physician expertise. We were very thoughtful about how to bring physician expertise into our team in a way that balanced and combined well with the research expertise we have on our team.
The first layer is high-level advisers, who are more informal and help us with strategy in various ways. We share our roadmap with them and things like this. The second layer is folks who we work with in what you can think of as a human data operation, in the way that you're describing, but a little bit more closely.
We're on Slack with them all the time, and we work with them closely. They're not just going off and doing tasks; we ask them for advice and ask them questions all the time. We basically have a Slack community where we're interacting with these folks. Some of what they're doing is comparing model outputs and red-teaming model outputs. Some of it is looking for ways in which we might have blind spots today and listing those so that we can prioritize them in the future. Some of it is testing new products and things like this.
As an example, in the work on ChatGPT for Healthcare, which we announced on January 8, we had red-teamers test this product over 9 waves for 6 months. This was in close collaboration with this physician community.
The final, top-of-the-pyramid way that we rely on physician expertise is through really close advisers who work most closely with our team. They're the ones who work on channeling the voice and combining the voice of these hundreds of physicians, interfacing most closely with our research team, and translating that into evals and model training data so that we can then improve our models.
We talked about channeling the voice of the physician and calling back to the Hippocratic Oath. One thing that I do find kind of frustrating, honestly, about my experience in the medical system recently is that I think there's a little bit too much emphasis on “do no harm.” This also connects in a pretty deep way to questions about how we should even conceptualize how to talk to our AIs about what they should do.
We have the very detailed, Talmudic-detail rule set as one extreme possibility, where, for every corner case, we try to map out what you should and shouldn't do, and hopefully it learns that and lives by it. And then, at the other extreme, the Anthropic Constitution—the Claude Constitution—has recently at least demonstrated that you can get pretty far with something that is less rule-based and more about trying to teach the model to have good character and use good judgment as it goes through all the situations that it finds itself in.
I would say my critique of the human doctors that I've engaged with, who have generally served us really well, is that they definitely want to do no harm, to the point where they're sometimes too reluctant to engage in a hypothetical or to act on something. I do have one good friend who's a doctor who also happens to have had the same, or very similar, kind of cancer as my son. He's notably behaved differently with me in private, one-on-one. He's like, “Look, I don't need a randomized controlled trial to tell you that that makes sense.”
But you don't get that too much when you're actually at the hospital. There's a reluctance among physicians—certainly one I've experienced, and I think it's a pretty commonly shared perception—to act on things that make sense because that isn't guaranteed to work out. Obviously, biology is super messy, and there's a ton of diversity, and there's just a lot that we don't know.
So how do you think about that kind of challenge? I kind of wonder where you guys want to land in terms of only adhering to the most rigorously defensible advice versus doing a little bit more of the Amanda Askell thing and being like, well, it does take—you've got to be willing to take some risk to help people sometimes. Maybe we should have the models do that.
Yeah, it's a great question. You're pointing out problems, I think, on two sides of the healthcare ecosystem. One is that, as a patient, you have this challenge of needing to advocate potentially pretty hard for yourself. I think in your son's story, you pointed out a couple of false starts where you saw a doctor, they said it was probably normal, you had an abnormal blood test reading a couple of times, and then they were just kind of like, “Yeah, it's probably okay.”
This is the kind of experience that patients often have when they're having something that they feel or know to be an issue, needing to advocate for themselves, and doctors often feeling like their doctor isn't hearing them. On the clinician-facing side, you have this challenge that medical evidence is increasing rapidly. It's very hard to keep up with the latest of what's going on, and you're overloaded with documentation and a bunch of other burdens. Doctors are human, too, right? So there are a bunch of challenges on both sides.
You have this amazing thing, which is AI, that's able to do a few things. You have AI that's able to talk to patients, understand the concerns that they have, and integrate knowledge and information across both their previous history and the latest medical evidence. One of the things that these models are obviously very incredible at these days is taking in a huge amount of health context—not just on you, but also on the latest medical evidence and things like this—integrating that all together into one context, and doing something that I think is very difficult for a human to do.
Then, on the physician-facing side, you again have that same capability to integrate information and interface with the physician. That's a lot of why we're doing this work: because we see this gap between where the models are at and how people are using them, and we think that that's really important with our upcoming products.
To get closer to answering your question, you mentioned where we see how the models should navigate places where there's potentially a lack of medical consensus or where physicians would potentially disagree about what to do. This is pretty fundamental to our approach, and this is why we don't have one, two, or three experts determining what the models' outputs are. We have a pretty multipronged approach.
A lot of this comes down to presenting information to the user, but being sure to present uncertainty when uncertainty exists. I mentioned the safety-research motivation for a lot of the work that we're doing. One of the directions that we've been exploring is whether these models are well-calibrated in their uncertainty and whether we can get better at having these models verbalize their uncertainty.
For example, if there are 3 to 5 potential paths for your son's next treatment, a doctor might, just for the sake of simplicity and clear communication, communicate one or focus on one. A model can potentially communicate 3 to 5 of them, but mention that the state of the evidence may be somewhat limited and mention caveats.
One of the things that we've been investing in is, first, whether models can become better at understanding their own uncertainty and whether that can be something that we measure and improve. Second, can they verbalize it better? This is, I think, a big part of the right way to thread the balance between being more aggressive in sharing potentially early results or early evidence and being overly conservative. That's how we think about that.
What are you seeing in terms of trying to get the models to understand their own uncertainty? I remember that famous graph from the GPT-4 model card where the pretrained model seemed to be much more calibrated with respect to its own uncertainty than the trained model. We've only scaled post-training since then, right?
I haven't seen an update to that kind of research in a minute, but it seems like there was at least a fundamental challenge opening up there with respect to that kind of introspective self-awareness—how confident am I in this answer? Is that something you guys have solved? Is that why I haven't seen that sort of graph in a while?
Well, I think there's a measurement challenge. The plot that you're referring to, and the kind of plot that existed back then, was: given the next token—for example, in a multiple-choice question—was the model's probability of that next token, say A for a multiple-choice question, corresponding to how likely it was to actually be correct in choosing A?
What you're seeing now is, I think, that it's become more difficult to measure that for 2 reasons. One is that we have higher expectations of our models than answering multiple-choice questions, and so it's hard to say when they're correct or not correct.
And the second is a little more technical. The models now emit reasoning tokens, or thinking tokens, between initially outputting something and their final answer. The result is that you can’t ask the model for the log probability of that next token being A in exactly the same way that you could before.
There are a couple of things that you can do to handle this. One is to go in the direction of richer ways of measuring whether a model is correct or is doing the thing that you want, rather than measuring the log probability of a certain letter. The second thing you can do is repeatedly sample from a model and then see whether performance stays the same or degrades as you repeatedly sample from it.
Our work on HealthBench is a good example of doing both of these things at the same time. In HealthBench, we did this work around measuring not just 1, 2, or 3 different aspects of model performance in health, but across more than 250 physicians and 5,000 conversations, measuring about 49,000 different axes on which model performance could differ. Part of this is whether or not the model expresses uncertainty in the right way, whether it says the right fact, whether it escalates to a physician when needed, or things like this.
There are 49,000 different things that are measured in HealthBench. One of the things that you can do there is measure correctness in a way that’s less about multiple-choice accuracy and more about whether the right facts are included—facts that are really important to emphasize to the user, and things like this.
The second thing you get out of that is this metric, which we call worst-of-N. When you repeatedly sample from the model and try to measure performance on HealthBench, what is the worst performance you get on N samples? You can sample from the model 20 times—what is the worst performance you get?
Now you have a way of measuring, instead of using the log-probability-based approach, whether a reasoning model produces a consistent result conditional on the different kinds of thinking that it’s doing. What I would say is that it’s harder to produce plots like we could before, because now the thing that we’re measuring is so much more complicated. But when we do produce the plots, as we did in HealthBench, they also look pretty promising, and models have improved quite a bit at that.
So how should you, as a user, think about the worst-of-N thing? Is there a way to sort of translate—maybe you could just describe the result? How much worse is the worst-of-N than, say, the next-worst-of-N or the average of N? If I’m a user, which I am, can I get certainty just by running it twice? Is there a practical upshot of that work that could give me confidence that I’m not getting something way worse than what the model typically outputs? If I do X, can I be sure I’m not getting something that’s way worse than usual?
The way I think about it is that the more compute you spend on things, the better results you’ll get, and the results may be marginal over time. One thing you could do as a user of the model is sample from a model 10 times, combine those outputs, have an LLM synthesize the outputs of an LLM council, and then produce that as an answer. I think that will be marginally better than the answer that you get from just running one model, and this is not so dissimilar from what GPT-5 Pro and things like this do under the hood anyway.
I think you can do a few things if you’re a user and you want to make the best of this. One is that you can use GPT-5 Pro or something similar to do a thing like this. A second thing you can do is actually increase the amount of reasoning, because I think it has a very similar effect to just running the model multiple times.
In both cases, my current sense is that we’re getting to the point where current models are performing incredibly well for most people most of the time. I know, for example, that you using GPT-5 Pro was an important part of working through your son’s situation.
I think we’re reaching a point where, except for the most complicated cases, most people are best served by just using the model on a default reasoning setting. I would recommend using the reasoning models instead of the more instant models—using GPT-5.2 Thinking rather than GPT-5.2 Instant—for a lot of health-related things. But I think most people can get the best of both worlds between latency and performance by doing that.
I think the way to think about the worst-of-N result, broadly, is that if you sample from something 20 or 50 times, you’ll have varying performance across model outputs. What we saw in the worst-of-N results is that recent models have improved pretty significantly. The worst performance of o3 at that time was way better than the best performance of GPT-4o, and we’ve continued to see that over time.
We’ve been shipping model improvements in health pretty rapidly over time, and the model improvements in the last year have been greater than in previous years since ChatGPT launched. As an example, today the nano models—the GPT-5 nano models that you can get through the API—as well as our open-source models are performing similarly to o3, which was our best and greatest model not so long ago.
The latest reasoning models continue to push the frontier of how much you can do with less and less reasoning. This is true for GPT-5.3-Codex as well as GPT-5.2 Thinking. If you try using them by default on health queries, they’ll actually think a little less but produce better results.
We’re continuing to try to push that frontier—not just needing to pour more compute into a model to get a good result, but also getting better performance at a given level of compute. The result is that the models are way better than they were even a year ago.
Yeah, it’s been crazy in my just 3 months of intensive use. Nothing has brought home the pace of shipping quite as much as how many updates there have been in this one chapter of my life. It’s been wild to see.
I did an episode with the folks at Apollo. I’m sure you know their work, if not know them personally. One of the really interesting things that they observed when they got access to the chain of thought is that, at least for o3 at the time—although I’m not 100% sure which model it was off the top of my head—there was this seeming development of a new dialect internal to the model. The famous sort of “watchers, watchers,” whatever.
Can you share anything about how you’re balancing the obvious good of efficiency and denser thinking per token—more value per token created—with the seeming tendency to have the internal chain of thought go off in weird and potentially hard-to-parse directions? I guess there’s also the commitment from OpenAI not to train on the chain of thought, or at least not to apply certain kinds of pressure to it. I don’t know if that’s an absolute ban on any feedback on the chain of thought, but I’m interested to hear your thoughts on that, because that’s something that I’ve been like—
Well, the world moves past these big stories so quickly these days.
That one seems to have come and gone, and I'm not really sure what the state of it is now.
Yeah, it's a super interesting question, and near and dear to my heart as a safety researcher. Chain-of-thought interpretability has been one of the nice advances in safety over the last year or two, which has been really cool to see.
As models have become reasoning or thinking models, they've also emitted tokens that effectively explain their work and what they're thinking. This provides a form of interpretability for researchers who want to understand what the models are doing, which is exactly what safety researchers want. This has been a really cool way of measuring whether models are doing things like scheming or producing outputs that are undesirable in various other ways.
This has been relevant beyond health to a bunch of other domains as well. I think a lot of the results that people have shared and studied have actually been in coding. I've been pleasantly surprised.
The danger that you're pointing out is: as you put more pressure into reinforcement learning to get the models to produce a good output, will they slip away from the prior of having their thinking tokens be simple English that's easy for researchers to understand? I think what we've seen is actually pleasantly surprising. At least until now, we haven't seen a lot of large-scale evidence of a slip into what's called “neuralese,” of using chain-of-thought tokens in a way that is not explainable and understandable in English.
In general, as we've been trying to understand the monitorability of our models over time, we haven't really seen that effect as we've scaled up our RL. I'm not sure if that'll continue to be the case in the future, but so far I've been fairly pleasantly surprised that this kind of side effect of the reasoning paradigm has continued to be useful and hasn't robustly been seen to become unreliable. I think the result that you're pointing out has been continuous over time, with these weird blips in the interpretability of the models at times, but we haven't seen a continuous increase in that.
We haven't seen clear evidence, even though we've actually tried to study it, that scaling RL causes that to happen more. We haven't seen that yet. I would expect that in the limit, it does.
Okay, that's quite interesting. The way I interpreted that, first of all, I thought it was happening more than it sounds like you're saying it is happening. I naively assumed that there was some sort of brevity reward signal being applied in addition to ultimate correctness. Certainly, it's intuitively comprehensible why something like that would start to happen if you did that.
Should I infer from what you're saying now that there isn't really a brevity signal, and that this is more of an emergent, weird phenomenon that doesn't happen that often? Is it one of those weird language-model things that we keep an eye on but don't obsess about too much because it's rare? Is that a fair summary of the state of play?
I do think it's important to pay attention to. I think the right way to think about it is that there's nothing reinforcing that this should happen during training. There's no reason that models should produce chain-of-thought that is human-interpretable during training. The reason it happens is because they have a prior of using the English language, and when you give them the space, with thinking tokens, to produce a more correct and helpful answer to the user, they're actually using it in English because that's just the easiest thing for them to do.
I think this is basically an empirical phenomenon that's extremely useful for safety research. I would love to keep it that way as much as possible, and I'd love to see more research into how we can maintain that. I think OpenAI's commitment to avoid optimization pressure as much as possible, as well as the commitments of the labs, is really exciting progress. I do think it's important to watch out for it, so I think it's a great question.
Yeah, to be continued. You mentioned the incredible complexity of the evaluations that you're running. By simple math, if there are 5,000 conversations and almost 50,000 criteria of evaluation, there's a lot under that. I imagine some of those criteria are reused across conversations. At a minimum, we've got something like 10 evaluation criteria per conversation, and probably a lot more criteria per conversation than that.
It's hard to summarize how good they are. Nevertheless, I'm going to ask you: How good are they, and how should we think about that? There's also this HealthBench Hard thing that might be a way of saying, “They're good at most things, but here are some things that they still struggle with.” Maybe defining the frontier is one way to say how good they are. How do you communicate to the world how good the latest models are when it comes to health?
I think it's good to keep in context the arc of work around LLM evaluation in health. A couple of years ago, people were mainly focused on evaluating LLMs by thinking about performance on multiple-choice questions, like medical exams and things like this. Then, with the work that collaborators and I did at Google, we started increasingly investing in what it looks like for specialized or unspecialized LLMs to answer general health questions, with some of the Med-PaLM work, and potentially moving in the direction of asking follow-up questions to a user with the AMIE work.
Increasingly over time, we've moved toward higher- and higher-fidelity evaluations, and I think HealthBench is the latest big step in that direction: wide coverage of LLM performance and safety, but covering it in a way that includes many different axes of performance that actually matter for the real world.
You see these 5,000 conversations and these 49,000 different axes of performance. There's actually 3 different versions of HealthBench. One version is the full data set. The second version is HealthBench Consensus, and the third version is HealthBench Hard. These mirror what we view as the 3 high-level principles when designing evals for health.
The first is that you want it to be meaningful, which means that if the number goes up, hopefully human health will improve. The second is that you want it to be trustworthy, which means that it's backed by the consensus of doctors, for example, or other experts. The third is that you want it to be challenging, so you don't want it to be at 100%.
One thing that's happened over the years is that all the previous benchmarks that meant anything at all have gone to 95% or 100% over time. HealthBench actually remains unsaturated to this day. With HealthBench, HealthBench Consensus, and HealthBench Hard, you have a little bit of focus on these individual axes.
HealthBench overall is a number that we think, if it goes up for a model and people are using that model, human health will improve. We feel like that's a statement that we can defend with the rigor behind the work. The second is HealthBench Consensus, where we specifically focus on a bunch of criteria for evaluation where, for each example, a majority of multiple physicians agreed that the criterion was applicable.
Not only did physicians write these criteria, but a bunch of physicians also checked whether the criteria held for a given conversation and whether they were the right things to be evaluating about a given model and a given conversation. The final thing is HealthBench Hard, where we took a somewhat adversarial approach against a bunch of different models across all the different model providers. We chose the examples that existing models fared the worst on but that still seemed high quality, and then we turned that into a benchmark.
HealthBench Hard has been my favorite external benchmark for whether an open model is doing really well. When it came out, GPT-4o was literally at 0 on this benchmark. It's an incredibly hard benchmark because that's just how we chose the examples. Over time, we've improved the performance to around 40%, which is still nowhere near saturation on this benchmark. I think this benchmark has a lot of room to go for OpenAI's models. Competitor models are more in the 20% range for current models.
That's the way I think about the HealthBench family of evals. We have this commitment to work on evals that are meaningful, trustworthy, and challenging. If you want to focus more on the evals that are super trustworthy, then we have this subset, HealthBench Consensus, which is really focused on that.
In addition, as part of the HealthBench work, we did a couple of additional analyses. One of them was that HealthBench involves grading these individual 49,000 different rubric items using a model-based grader. We had physicians compare the model-based grader to the grading of other physicians, and we actually found that the model-based grader was doing a better job than the average physician. What that tells you is that the grading for HealthBench is pretty high quality compared to what you'd expect from a physician.
Recursive self-improvement—signs of recursive self-improvement. Alert, alert. I have just a little bit of an intuitive sense for what's in HealthBench. I can give you just my experienced sense of the frontier from the bedside over the last couple of months.
And I would summarize it pretty simply. This is what I say to my neighbors and stuff when I’m saying, “Hey, by the way, you should really use a language model next time you’re facing a health challenge.” I basically say, “Look, I was in the hospital for an initial 30 days of really intense treatment, where everything felt super high-stakes. We didn’t always know what was going on, and we also didn’t know how much we could even trust our doctors at that point. Everything was so new and stressful at the same time.”
What I found was basically that the frontier models were step for step with the attending oncologist on almost everything. And that means they’re a lot better than the residents. They’re just much more knowledgeable. They’re at the attending level, for sure.
There were maybe a half-dozen times over the course of that month where there was some disagreement. Initially, I was just using GPT-5 Pro, whatever exact version of GPT-5 it was. Then, as these other frontier models came out, I started to do everything in duplicate with Gemini 3, and then triplicate with the latest Claude. Interestingly, I would say that the AIs disagree with each other even less than the models disagree with the attending.
It’s quite limited disagreement between the models and the attending. Typically, when there is disagreement, I found it to be a very minor thing: “Okay, his electrolytes have gone a little low. Should we give him electrolytes today or not?” Of those half-dozen things, there’s not really a major trend.
I would probably score it 6–4 for the doctors, with the benefit of hindsight. We usually followed what the doctors said, and in the end, did we feel like they were right, or did we kind of wish we had gone with the AIs? I’d say maybe 2 out of 3 times we have felt like they were probably right.
And I think the one—if I tried to chalk it up to something that they have that gives them an advantage—it almost always was one of those situations where, whatever I was putting into ChatGPT and whatever data I was providing, I was always exporting the latest results from the EMR and dropping in the PDFs that they gave me. Now we’ve got integrations for this, but at the time, that was what I was doing.
The difference usually came down to the physician saying, “In view of all that—even taking all that into account—but also just looking at him right now, watching how he’s breathing and looking at his color, I’m pretty sure he’s fine.” It was that kind of very intuitive, very multimodal, very subtle sense that these folks have developed over quite a few years of clinical practice that, on these very fine margins, seemed to give them a slight edge over the models.
That’s kind of my account. But I guess what that means maybe is that my situation isn’t that hard, and I think that is actually true, in the sense that the cancer that my son has is not a super-rare one, and the treatment protocol for it is quite well-established. So it’s not like a super-hard call in terms of the main line of what to do. Even though it’s a hard thing for him to go through, maybe it’s not that hard in terms of clinical judgment.
I haven’t experienced anything that I would say the models were only 40% on. So with that, maybe you can tell us a little bit about what’s out there still on the frontier where we do have ground truth, or at least some sort of consensus that we feel is solid enough to grade models on, where they’re still only at 40%.
Yeah, it’s a great question. Let me describe a little bit more about what HealthBench is evaluating and the ways in which models have improved over time, then talk a little bit about the next frontiers for model improvement.
HealthBench measures many, many different things. It has these themes, which are the focuses of different evaluation examples. A few examples of focuses that we had: one was whether models are appropriately escalating to care when needed versus not escalating to care unnecessarily. You want to balance this because you don’t want to be, for example, overwhelming the health system with a bunch of patients who are worried because of alarmist medical advice, versus not escalating to care when needed.
Another aspect that we measured was the ways in which models can adjust to different demographics, different epidemiological conditions, or different levels of access to care globally. This is both making sure that you adjust to a user who’s male or female, but also, if somebody asks a question in a region where tuberculosis is more common versus less common, making sure you adjust for that, and things like this.
We call that global health, and that was the biggest single focus of HealthBench, because that’s, I think, one of the biggest ways they can be most impactful. And then a bunch of other things.
We talked a little bit about calibration. One of the ways that the models have gotten significantly better over time is that, when they know they’re uncertain, they not only flag that uncertainty but actually browse to get more information—for example, getting the latest resources and being able to synthesize that information together.
Another thing is asking follow-up questions and asking the right kinds of follow-up questions over time. Initially, ChatGPT would almost never ask follow-up questions in health settings. Now it does much more often, and it’s much more likely to prioritize the right follow-up questions for you.
You had this story around when you were using the model for your son, where you actually learned over time to ask the model to interview you and figure out whether you could do the physical exam yourself. The models have gotten better at knowing when that would be useful and flagging that as well.
So it’s everything from pure reasoning, solving benchmarks, medical calculations, and figuring out the diagnosis, all the way to how the model behaves. What is the bedside manner? Is it comforting and things like this? Again, it’s a balance of both difficult and high-trustworthiness signals that we’re getting in working with our experts.
Over time, we’ve continued to see improvement as the models have improved from o1 to o3, GPT-4.1, GPT-5, and so on. All these models have improved significantly in health, and today, every major stage of model training for every model that we ship at OpenAI actually now benefits from our work in health. That’s going to continue for future models as well.
I think the frontiers are in a few different areas now. One is that a lot of text-based performance, outside of subspecialty areas, is actually pretty good. So if your goal is to keep up with the latest evidence, even as a physician, I think the models are doing an incredibly good job today, and people are finding a lot of value in them. I think that’s one of the greatest data points toward that point.
I think models have continued to improve in their ability to integrate information. You mentioned this thing of taking a bunch of different information as a physician—looking at the color of your son’s skin over time—and integrating that all into one context. I think a big challenge for models today, and this is less of a model issue and more of a question of how you surface information to them, is actually getting the right context.
One of the challenges that you faced in your situation is pulling the right information from the health system so that you could upload it into ChatGPT and ask the right questions yourself and advocate for yourself. We’re looking to make that better with the release of our ChatGPT Health offering, which I can talk a little bit more about later.
In short, this is basically an experience within ChatGPT that enables you to connect to your health information from your medical records, any wearables, or Apple Health, and provides additional, purpose-built privacy protections for that. We can talk about that a little bit later.
But I think on the model-related side, there are a few challenges. One is getting that context again. I think the models are incredible when they have the right context.
The second is thinking about various modalities that are not well captured by things like HealthBench, which is really focused on text. So thinking about really good performance in multimodal settings, including image and voice, I think people will start to rely on the models in more and more modalities, and I think there’s still significant room to improve there.
I hope that in the future, the best models in the world for doing various imaging modalities in health are actually the models that are most easily and readily available to people today.
Now, what occurs to me as you say the multimodal thing, I’m like, “Hey, maybe I was even undershooting it.” I had the intuition that taking a picture of my kid as he sits in his bed wouldn’t necessarily help and might even confuse, but maybe I’m wrong. In the future, would you advise me to start just including cellphone camera pictures with my daily synopsis?
Well, I think you’re right to point out that there is still a little bit of a challenge there. For the average user, it’s difficult to know exactly what data you can connect and how to connect it, right? And you were highly motivated in this case.
There’s a little bit of a gap between the most motivated, most expert user who’s willing to wade through signing into their patient portal and things like this, and take screenshots or copy and paste things manually, as you did, versus the person who wants their health data integrated and wants to be able to understand their health and advocate for their health or that of a loved one, but has challenges in doing so. I think that’s the thing that we do hope to work on, on the product side of ChatGPT Health: really lowering the activation energy to do that kind of thing.
I think, in general, the models will benefit from having more and more context. I’m not sure if specifically that additional information would have helped in that case, but I would generally advocate for people to try putting in various kinds of context that they can. I think people have been surprised—I think you were surprised in your own story—by how useful it can be if the models have more context.
An interesting tidbit here is that studies have shown, for the past 2 or 3 years, that in these clinical case challenges—which are effectively crossword puzzles for doctors, where you get all the patient context up front, potentially multimodal context as well—the models did an incredible job of figuring out the next steps for diagnosis or treatment. These are extremely challenging puzzles, really difficult for doctors, and people even found that the models improved on the performance of these experts. That’s been true for some amount of time, and I think a lot of the challenge in the ensuing time has been figuring out what it looks like to get the models to have a back-and-forth conversation to solicit the right kinds of context from you as a user.
With ChatGPT Health, we’re taking that a step further and thinking about the right ways that the product interface can make it a lot easier for you to do that, in a way that’s secure, where we’re not training on your data, and people can trust that as well.
How about for these larger modalities? When you get a PET scan or whatever, my understanding—I haven’t really been able to get my hands on too much of that data—is that it’s like gigabytes, and I’d have to get it on a disk. I didn’t even have a computer that takes a disk, so I’d also have to get a disk drive to read that. Obviously, that data can be compressed. How do you think about the pipeline of raw scan-type data and feeding that into tokens? Are we talking about certain specialized Perceiver model modules that ingest those and do the reduction, or some other third thing—a mysterious third thing?
Yeah, it is a super-great question. One of the interesting things about biomedicine from a modeling perspective is that there’s a long tail of modalities that are interesting and relevant, and often they’re fairly difficult to put into models today. I think there are going to be 2 broad approaches, and this is talking more about external research. There are 2 broad approaches that I’ve seen in external research.
One is thinking about the right ways for models to call specialized tools or run code over various kinds of data. If you think about how a human views a gigapixel scan—let’s say a pathology image—how does a human view it? They basically view slices of that image. Or even with a 3D or 4D scan, like your son’s MRI, they view slices of that image, right? No human can see a 3D or 4D modality.
So what they’re effectively doing is using a tool to understand which slices of the image might be most important, and then taking a look at slices of that image and manipulating it. That’s actually a thing that models can do with tools and Python and things like this. That’s number 1. There’s a sub-bullet of that: they can also use very specialized, fit-for-purpose tools, like specific Python libraries or even additional models that specifically encode these modalities.
The second approach is to have these models encode the modalities themselves—to have a way of basically putting them into token space in some way. This is kind of what’s been done for images, video, and audio for various models. I think what you’ll see in biomedicine is a bit of a mix of both. I think there are benefits to both approaches; there are pros and cons for both. Then I think you’ll see a little bit of a hybrid approach where researchers will end up doing a mix of both, depending on the modality.
It’s striking how good AI is—not necessarily even the general-purpose AIs, but we have these specialized AI models that can fold a protein in a superhuman way. What seems to be missing is the latent-space joining of that modality to text. But my expectation—and you can tell me if you think I’m right or wrong, or how long I have to wait—is that we’re going to see more of that, and that it’s going to be a major driver of truly superhuman performance in a lot of domains, because we just don’t have people who can intuitively fold a protein.
If you had a person who could do that and they had their general reasoning capabilities, it seems like you would have a really different kind of—a qualitatively different kind of—intelligence. So I’m really interested in what you see as the roadmap, the timeline, and the expectations for that kind of deep integration.
I think it’s a really optimistic picture. I totally agree: if you think about how we lean into the natural capabilities of these models, these models can take in vast amounts of information into their context, and they have the ability, at least in theory, to take in a lot of modalities of data and effectively merge that together. I think the research for that is becoming increasingly solid for these other modalities.
I think there’s significant research to do depending on the biological modality, and as I said, there’s a long tail of biological modalities that matter. So depending on the modality, the relevance and impact, the availability of data, and all these kinds of factors, I think it’ll take more or less time. It’s hard for me to give you a universal answer, right? I share your optimism.
I just saw one in the last couple of days: SleepFM, from Professor James Zou and others. He’s at Stanford, and the collaborators are probably at multiple institutions. They had a really interesting finding. The idea there is that they use, I think, 6—maybe up to 8—different modalities that are all measured during 1 night of sleep: how you’re breathing and various things that are easily measured in a kind of sleep-study setting.
It just takes 1 night of sleep to gather all this data, and then they’ve gotten really good at predicting all kinds of different diseases based on that data. They’re integrating all these modalities with each other into a holistic understanding and using that for this sort of narrow set of predictions, which are very high-value predictions. It’s not yet going all the way to text, but it’s just another data point that’s top of mind for me right now.
That’s like, man, it’s all happening. All the latent spaces will be joined, and I can’t shake the idea that that’s going to be a big part of what superintelligence ends up looking like.
Going back for a second to bedside manner, you mentioned global health and context being a big part of HealthBench, and what you’re trying to do to make sure it benefits all humanity. It was funny when you said, “With 200 million users, when we scale, we want to make sure we’re doing this right.” It’s like, yeah, one day you’ll scale—and only 200 million users today. Clearly, there’s already some scale.
The thing that I want to address, though, is the bedside manner and also how that relates to the sort of n-of-1 context of the individual user: their preferences, their memories. Now, of course, all these products have an integrated memory module, which I’m sure is also going to become a deeper kind of integration over time. Right now, I don’t think it’s disclosed exactly what it looks like. People generally understand that it’s kind of a scratchpad of key highlights that the model can reference, I believe still in mostly text modality as it goes.
But regardless of exactly how that works today, it’s safe to say it will continue to become a deeper and deeper integration as everybody pursues various strategies for continual learning. But even just that—there’s also integration with tools. Gemini now has access to my Gmail, and I’ve got even just system prompts. I can tell the model that I want it to behave a certain way or a different kind of way. It strikes me that that creates an impossible surface area for you to manage.
We’ve seen weird emergent things like sycophancy here and whatever else there. How do you think about making sure—because this is one thing I do think human doctors generally do a pretty good job of—sizing the person up and being like, “Okay, how do I cut through what I need to cut through to get this person to understand what they need to understand?” That’s something I think models broadly have not been as good at yet, but it really seems to matter a lot in the medical and health care domain.
So is there some additional suite of testing that you do, or how do you think about that—just this insane long tail of idiosyncrasy that 200 million health users bring to the product?
Well, you pointed out 2 problems, and I’ll talk about them in reverse order.
One of them is that people are bringing in an incredible array of experiences, different settings and ways that they use ChatGPT, and different memories. All these different things mean that the experience of one person using ChatGPT can be different from the experience of another person using ChatGPT for reasons that are not just because they’re using a different model or something like this.
We do a few things as a company to understand and improve model behavior. One of them is evals that we run on models before launching, and HealthBench is a great example of that. We have additional evals that we run as well. Another thing that we do is monitor usage of production traffic in privacy-preserving ways and run classifiers over it to understand if there are any safety risks, whether it’s anything from health to frontier risks.
We’re actually able to measure these things, and there’s a great example of this kind of work in the blog post we had about sensitive mental health conversations. We were able to show that the evaluations we were doing in mental health, as well as patterns we were able to see in production traffic—again, logged in a privacy-preserving way, with models running over them—showed a correspondence between how people were using the models, what that looked like, and what our evals were measuring. They were actually very well correlated.
That’s how we think about closing the gap between evals that you run on models and what people are doing in the real world. I would add an intermediate step as well, which is doing a real-world study. Here, I’ll plug our work with Penda Health.
This was, I think, the first real-world study of an LLM-based co-pilot for clinicians. We had some clinicians in this group of clinics in Kenya use AI as a co-pilot or a safety net. As they were typing in their electronic medical record, it would flag things if they were interesting, alarming, or potentially incorrect, while other clinicians did not have that.
For the patients in the group who were treated by clinicians with the AI versus those treated without the AI, there was a statistically significant improvement in diagnosis and treatment outcomes. This is another example of how you move from offline evaluation that’s potentially not super realistic, like medical multiple-choice exams, to increasingly realistic evaluation that you can run offline, like HealthBench, and then move increasingly into the real world.
You can view the HealthBench study as a forward-looking study, and you can view our analyses of production traffic as more retrospective. Again, they’re capturing a lot of the differences and variance that you’re pointing out.
Yeah, that’s really interesting. Can you talk a little bit about the privacy-preserving nature of it? I would preface by saying that I think people, at least as individuals, typically over-index on this, if they think about it at all.
I’ve had occasion to think on my son’s behalf, “Should I be putting this information out there?” Then I see what somebody like Sid from GitLab has done with his cancer—literally open-sourcing all of his own biology, down to the DNA level—and an incredible amount of very individualized data that he’s put out there. I think that’s the way, because I can’t really come up with too many ways that anybody would use that against you. You can get really sci-fi about it, but unless you’re getting really sci-fi about it, it’s hard to come up with a way that anybody would really use that against you, and there are a lot of ways that it might stand to benefit you.
In telling my son’s story on the podcast, I also had a similar question: Am I doing him a disservice in some way? Maybe something could still happen, but what has happened is that people have reached out to me with interesting opportunities for connection. I’ve been able to tap into expertise, including the team that Sid has built to help him, and they’re starting all kinds of individualized therapy companies. There are fascinating developments there.
I think my advice to people would be: seek the benefit, and don’t worry too much about whether or not your data is sitting in some log somewhere. That doesn’t seem to be a huge concern. But anyway, that’s my role to say that. It’s your role to build a product that deals with people as they come, and if people do seem to really worry about that, what would you want people to know about the privacy-preserving nature of the infrastructure that you guys have built?
Privacy is incredibly important to a lot of people, and I think it’s a very personal thing. People have different preferences for what data they share and how they share it.
If you zoom out for a second, think about what the level of ownership of your own data is as a patient today, and what it was in the past, even before things like the 21st Century Cures Act. Most patients, let’s say 10 years ago, had no real way to access their own health data. Not only did they have no way to control who they shared it with or what they shared, it was just out there and shared anyway in various ways. A patient didn’t even have a right to look at their own data or an implementable path to actually access it.
Now, we’re in a slightly better place, I would say, but still not an ideal place, where most patients feel like it’s incredibly difficult to get access to their data, or the data of their loved ones, and take control of it and how it’s used. People like you and many others with incredible stories have found incredible ownership and advocacy enabled only by having access to this data and being able to use it in the ways that they see fit. We want to respect that.
We want people to be able to access their data, and this is part of what we hope to enable with ChatGPT Health: a lower-activation-energy way for you to connect to your health data and make that useful. At the same time, it’s super important to a lot of people that this be done in a way where it’s clear that there’s not some kind of competing incentive here.
One of the things that we’re very clear about with ChatGPT Health is that none of the data you connect there is used to train our foundation models. The first reason we do that is because we think our foundation models are already great, and this isn’t actually the most important way we can improve them. The second is because we think this will further lower the activation energy for people to do incredible things with our models.
If people think that there’s some tension between privacy and utility, fewer people will go for the utility. We don’t want that to be the case. We want people to see the value of it and use it.
Specifically, ChatGPT Health has additional, purpose-built layers of encryption just for health data. ChatGPT already has a bunch of really good privacy protections built in, including encryption and things like this. The result is that you have encryption of the data, and you also have isolation of the data from other data that you have in ChatGPT.
For example, if you have other apps or memories from ChatGPT, those are kept separate from Health. The things you do in Health can benefit from that, but your other conversations in ChatGPT don’t have your health context or health information. You can keep that completely separate and have it live in a separate experience.
As we continue to get feedback and people continue to use it while we roll it out on the waitlist, we’ll continue to improve these protections. We think it’s really important for users, and we think it’s actually not as important as it seems for improving the models.
Yeah, interesting. Is that segregation of data a new feature with Health? All this stuff I’ve done to date has been in one product experience, but that fork has just now been introduced?
Exactly.
Yeah, I’ve been motivated to get a WHOOP to wear on my wrist to start collecting data. I’ve been fortunate throughout my life to have generally good health and haven’t created much of a paper trail in terms of medical history.
I’ve always looked at these quantified-self things and thought, “Yeah, it’s sort of interesting, but who’s got time for that? Am I really going to do all the processing on this?” The answer has always been no, so I haven’t worn anything to date. But now I’m thinking, “All right, it’s time. I can actually get value from wearing this thing because now I’ll have a product that runs in the background that’s smarter than me, frankly, at this point, and will really bring to light what I need to know.”
I’m motivated to change my behavior, at least on that level, and I honestly think it’s probably also going to end up encouraging me to exercise more.
In fact, that's my own North Star metric right now for all things AI. And not just health products either, but I will know that I am succeeding with AI when I'm spending less time in this chair and moving my body more. When I feel like I can untether myself from the desk, I'll feel like I'm really winning. I think I'm getting close to getting there.
I know I have one friend who uses voice mode nonstop and says he does a ton of it from the gym while he's working out. I'm like, "All right, I want to be like you." I'm not quite there yet, but I do think it's really interesting to think about both the ability to crunch all this data and the freedom—the liberation from the desk—to go out and make more exercise records in the first place.
I'm really looking forward to those. Those are major things I think on the horizon in 2026 that I'm super excited about.
Yeah, I'll just add on that. I think people should expect that the value of the data they collect on themselves should increase over time as model intelligence increases. The ability for models, both from a research perspective and a product perspective, to analyze that data and come up with useful insights that could be useful in the future—I expect that to increase over time. If there was ever a time to get a watch or anything like that, I would say it's now.
So, I do at this point basically everything in triplicate. For my son, my morning routine is: when the lab results come back, export them from the EMR and drop the PDF into all 3 of the frontier models. Grok has not cracked that top tier. Maybe it should, but it didn't in my initial evaluation, and I haven't gone back. So it's Claude, Gemini, and ChatGPT.
The tasting notes on how the models interact with me are, in short, that Gemini is by far—and this is quite surprising for a Google product—the most inclined to push me to advocate for something. It's also the most confident when it says I don't have something to worry about. Generally, I would say it's very accurate. Sometimes I'm a little uncomfortable with how opinionated it is because I'm thinking, "What if you're wrong?"
That's a surprising behavioral profile, I would say, for a Google product, especially because everybody historically has said, "Oh, they're going to be so conservative. They can't build AI products because they can't live with this sort of risk," or whatever. Gemini 3 does not reflect that analysis.
I would say ChatGPT is kind of on the other end of the spectrum. It generally gives me—and again, there's no system prompt. It does have my memories and stuff, but there's no intentional attempt by me to shape how it's responding to me—the longest answers and the most information. It's also the most clinical and neutral in tone.
It's sort of a report style. There are 9 issues that we go through and address one by one, and then summarize them all again at the bottom. Claude is somewhere in the middle. It's generally much briefer than ChatGPT, more like a Gemini-length response, but more measured in a more ChatGPT-like way, certainly compared to Gemini's more opinionated persona.
What do you think of that? How do you think about the balance between thoroughness and digestibility? That's probably the main tension I would zero in on there, specifically for ChatGPT.
Yeah, I think it's hard to know what's right and what's wrong. One of the things that we've aimed to prove and also measure via HealthBench, but also improve in our recent models—especially our most recent models—is having the ability to distinguish when the user is likely to be a health professional versus a lay user and tailoring the response accordingly.
Part of this is making sure that you're applying a level of detail or technical jargon that makes sense for the user's level of expertise. This is actually a specific thing that we train for and evaluate. It's also in our open-source HealthBench eval. This is a part of it focused on this, and it's something that we've seen improve, especially recently with GPT-5.2 Thinking. It's definitely an area that we've been investing in.
Overall, I think it's very hard to say what is ideal and what is not ideal. Personally, I think it's amazing if other competitors can be in this space and push the Overton window along with us, because I think it's hard for any one company to do.
A large percentage of the work that we're trying to do is think about where the models are at today and where the Overton window for user trust in the models is. Can we move the Overton window along in the right ways? I think additional models and additional products from other competitors, other companies, and other players in the health ecosystem actually go a long way toward helping shift that as well.
I think it would be a hard battle for us to win if it were just us. So I actually think it's amazing that there are other folks here.
Okay, so it is a striking thing to say that we don't need your data. Basically, we've got enough. I am interested in whatever you can tell us about where that data is coming from. This seems like an area where synthetic data—maybe I'm naive in saying this—but it seems like you really want some real ground truth from actual human medical trajectories.
I would feel a lot more comfortable synthesizing chats and trying to dial that in via a synthetic process than I would feel comfortable fully synthesizing data all the way to superhuman doctor performance. So I'm interested in where the data is coming from. I have this notion that you may think it is beside the point, given the tricks that you have up your sleeves.
I've been imagining a possible new social contract, not just between patients and AI providers, but also between patients and the medical regulatory system. I don't think my son is going to end up in this position, and I certainly hope not. But there are a lot of people who are in a spot where they're like, "I would try anything if it had a chance to save my life." A lot of times, they can't even get access to it.
Now we're in a spot where the AIs increasingly are going to know about it. You might not know about it as a person, but the AIs are going to know about it. They're going to tell you about it, and they're going to make a pretty damn compelling argument in a lot of cases that this is probably the best thing out there.
I've been down this path in a contingency-planning frame of mind. Everything has gone well for my son so far, but what if there were a relapse? What would we do? I've been really impressed by what the models have been able to give me there as well.
I imagine a situation where there's mounting pressure from patients who are like, "Look, AI knows about it, and it's telling me this is the right thing to do, and you're telling me I can't have access." It seems like that is going to be a very hard gate for the establishment to keep for much longer.
But the other thing I would love to see on the other side of that trade is this: if you're going to be given access to unproven treatments, then what we as a society, more broadly than just the AI companies, want back are your results, because we want to fold this into our general understanding.
It seems to me that we're getting to a point where—and this probably is going to be mostly driven by AI—we could envision not necessarily doing away with clinical trials, but moving beyond clinical trials and doing a lot of N-of-1 things that are just, "This is the best guess that we have for you based on all available information." If we can capture the result of that and train on it—as humans, as AIs, as society collectively—it seems like there's so much room to learn so much more and to move so much faster, all while delivering outcomes to at least some people who are trying these things that they couldn't otherwise get their hands on.
I'm feeling really motivated to advocate for that. I guess there are 2 data questions: how do you not need more run-of-the-mill data, and what sort of data do you still need? And what would you think about an evolution of the social contract? I call that AI and a right to try.
Yeah, that's a super interesting idea. We talked a little bit about the patient-facing problem of data fragmentation, right? The experience of this as a patient is that your health care data is just in the ether somewhere. You don't know where it's going or how it's being used. It's governed by HIPAA for the most part, and the result is that there are pathways for that data to go places that don't involve your consent.
The experience today for a patient is that you have both limited access to your data and limited control over where that data goes. If you think about it from the provider perspective, this is also a problem, right? If I'm a health care provider, I want to understand the whole picture of a patient's health and integrate across modalities, as doctors are so expert at doing.
How do we improve that experience? Right now, doctors have a lot of trouble—if you go to a new health system—pulling the relevant records from other health systems, from your entire context.
You can imagine a future where ChatGPT actually has a lot of context on you and can help the doctor pull that data in as well. I think that could be a really cool feature.
The third point that you're talking about is research, right? The majority of clinical trials that fail do so prior to conclusion due to recruitment challenges. We basically have a failure to recruit the right patients who meet some eligibility criteria. A lot of this comes down to the fact that the data is not in one place: what these patients are, who they are, and whether they are eligible.
That is also a problem facing researchers, and it is hampering our ability to raise the ceiling of human health as well. So you're pointing out a potential future where not only do patients have more access to data and providers are able to access their patients' data in a way that's unhobbled, as long as the patient consents, but researchers are also potentially able to access a patient's data, again, if the patient consents.
I think it'd be really cool to imagine a future where patients are able to opt in to additional consent and experience the benefits of AI-enabled AI and the right to try. And the right to try.
I do think clinical trials and the standard of evidence there are battle-tested, and there is something important about that as well. But I'm optimistic about a future where, because data is a little bit more centralized and a little bit less fragmented, we have a better ability to advance the science. Because patients have access to their own data, they're more able to advocate for themselves and, because of AI, figure out what clinical trials they may be eligible for or what experimental treatments may be interesting to them.
If they do have that active and informed consent, they can do that. Our understanding of science can improve, our models can improve, things like this. I think that would be a really optimistic future.
My previous comment, by the way, is not that the models have run out of data. I think data is always a helpful thing for models, but it's more a question of what is the right way to roll this out to society and where can we get the most impact. What would be the most impactful path? Is the most impactful path to lean into privacy and lowering the activation energy for users, or is it to lean into getting data?
I think the decision we made here is the right one: for users, to lean into not using the data to train our foundation models. I think that's actually what's going to have the most long-term impact if we think about the arc of improving AI and improving human health with AI—moving the Overton window along in a way that's trustworthy, that shows that we respect the value of people's privacy and data. In the long run, I think things like additional consent and different changes to the contract of how research is done are really cool things to explore.
Yeah. Just as you were thinking about this—or as you mentioned additional consent—I did get a packet of 40 pages of paper at the hospital one day, and they clearly really do want to collect this data. There's a person whose job it is to go around and visit the parents of these kids at the children's hospital, ask questions, and explain and answer questions.
I was probably the easiest customer she'd had in terms of—I was predisposed to sign. So I signed and initialed in a ton of places and whatever, right, to share whatever data we could share for the general public good. Still, I just got this in the mail from the University of Minnesota that said, “Can we have your data?” And I'm just like, “Man, I thought I agreed to this already.” What is the yield on those things that they're sending out? It's got to be quite poor. I haven't sent mine back just because I've been busy, and I fully intend to.
Those barriers are really tough. When you think about doing something like—obviously, much has been made recently about ads coming to ChatGPT—and I think the argument that, hey, we need some way to support this for billions of people, and ads are a proven way to do that, is a pretty compelling argument. I don't know if it extends to health. We have a lot of pharmaceutical advertising in the world today. Would there be such a thing as pharmaceutical advertising in the world of ChatGPT Health at any point?
Would you consider something like, “You can get ChatGPT Pro for free if you'll give us your data”? That seems like something that a lot of people probably would opt into in a way that I would think they would still feel good about. I don't know—any other sort of trade-offs you could see OpenAI making to support that reach? Because, again, we do want this to go to billions, right?
Yeah, ads aren't coming to ChatGPT Health, and we don't plan for that right now. Again, we think it's really important to create a clear separation between our health impact work and things that could be seen as contributing to other incentives for the company. That's the line that we've taken there.
I think you're right that access is a really important point for this, and this is why we've made ChatGPT Health free. This was not the default path: to provide it for free, without rate limits, to all users. But that is the path that we're charting with ChatGPT Health. That's not otherwise a thing that's available in the product.
We are deliberately optimizing as much as we can for access and doing as much as we can. I think there are going to be some limits to how we can do it, but I don't see any trade-offs in the immediate future.
Cool. Yeah, that's admirable. I think my sense of OpenAI on this dimension is very appreciative. The pains that have been taken to support a free user base of hundreds of millions of people, at obviously not insignificant cost, when probably a lot more revenue could have been extracted from those people, are pretty admirable.
To see that—I hadn't even heard that there was this plan to go free without rate limits for health for all. I've had this idea for a while of universal basic intelligence, which is basically the sort of—again, I'm obsessed with ideas of new social contracts—but this is one version of it. It's a pretty huge needle mover, right? It's going to be an unbelievable needle mover for a huge, huge number of people. That's awesome.
Do you want to talk a little bit more about the medical establishment's response? At the hospital, my experience is still mostly a lack of awareness among the providers. As I've built up rapport with certain people, I tell them, “I'm consulting with AI on this,” and they're usually okay with that. They're not hostile to it. They're sometimes skeptical.
I've had a couple of interactions where they're like, “Tell me what it said.” I'll tell them, and they'll be like, “Oh, okay, pretty good.” The doubt can turn around pretty quickly when you get the right answer. There is certainly mention of OpenEvidence. Broadly, though, I think there's just a lack of awareness of how good the systems have become.
My guess would be that this sort of awareness is probably the biggest concern that you have. I think in a lot of professions, we should expect to see the closing of guild ranks and the raising of barriers to entry. The cynic in me thinks maybe that'll happen in medicine, but the idealist is like, maybe not. Maybe this is a chance to live up to the actual mission of the profession and do the right thing.
Maybe it's also the case that people are just so overworked in medicine that they'd be happy to take whatever help they can get. But how would you characterize the broad reaction from doctors writ large?
I think broadly, you're right to point out that the Overton window shift for consumers has actually been faster than the Overton window shift for doctors. Already in the last year, the rate of adoption of ChatGPT for health—even before we launched the recent products—was incredible, and it's been one of the most amazing things to see. Again, over 200 million people a week are using ChatGPT for health and wellness questions.
On the doctor side, you do see a rapid increase in adoption, but it's not quite as fast. I would expect that to come with a couple of things. One is interesting interactions between patients and doctors who are maybe at different timescales or at different points in their journey of adopting AI.
It sounds like you've had a couple of them, and I think it's interesting because a lot of physicians first hear about AI through patients who are using AI, rather than AI that is actually built specifically for them. That's one of the things that we're doing with ChatGPT for Healthcare, which I can talk about.
This is not super interesting, and I think one of the things that we've learned over time is that the best way to shift the Overton window, like we've done over my time here at OpenAI and at Google, is through real-world studies. I've done a bunch of work studying things and doing real-world studies, things like this.
I think the thing that I found most effective for shifting people's opinions about AI and health, especially as the models have become more capable and safe, has been just putting the technology in people's hands. And so that's actually what we've been doing on the ChatGPT for Healthcare side.
ChatGPT for Healthcare was announced the day after ChatGPT Health and is more of an industry-facing announcement. This is basically a version of ChatGPT that is purpose-built for the workflows of health professionals, and specifically clinicians. This includes HIPAA compliance, additional features like specific evidence retrieval for medical guidelines, and additional workflows specific to enterprises and the writing workflows that doctors do.
We launched that with 8 of the leading institutions across the country. One of the most important things for building trust and credibility in the medical establishment is actually working with these leading partners. By working with these partners, we've been receiving amazing feedback. We've gotten a ton of inbound interest since that announcement, more than our team can actually handle.
My hope is that this announcement, along with some of the preceding work on the research side and our study with Panda Health, will hopefully lead to a wave of adoption—not just among individual clinicians, which is what we're starting to see, but also at the level of health systems and potentially even governments. I think we're just at the beginning of that wave. That's part of what I meant by we're seeing adoption, but not scaled impact yet.
I think scaled impact will start to happen, especially on the medical-establishment-facing side, more this year.
Yeah. I wonder if there's a question—I'm sure you've thought about this—that's sort of analogous to the privacy question on the patient side. What is the form factor or mode of rollout or use that's going to be best received by doctors?
Because if I'm a patient, I might say, “What I want is for the hospital, or even some bigger organization than that, to architect some workflows that just grind through everything. I want nothing missed. I want every record of mine examined, and I don't really care if that offends a doctor at some point. I just want the best results.”
But I could also imagine that if you were to do something like that, you might ruffle some feathers and create some immune response, so to speak, from the profession that you would rather avoid. So do you have a way, similarly or analogously to the privacy question, where we might take a little bit off the fastball in terms of how much immediate value you could create in order to make the rollout more acceptable to decision-makers, so that they hopefully work with you as opposed to, in some cases, being more inclined to fight you over time?
I think the trade-offs here are actually less than one would expect. You mentioned the possibility of protectionism in the industry, and we've actually seen much, much less of that than one would expect. I think the reason is that a lot of these doctors actually use it for themselves or for people they take care of.
When you do that and see the value of it, and see it getting incredible things correct and maybe pointing out things that you hadn't thought of, then that becomes the easiest kind of conversation. When we talk to health system executives, you can tell instantly who's used it and who hasn't used it for this use case. The conversation just becomes incredibly easy when people have used it.
What we see today, I think, is not a lack of top-down interest in adopting. We actually see a huge wave of top-down interest in adopting. The reality, though, is that in health care, workflows are fairly entrenched in a lot of ways, and it takes some time to make changes.
This is again the kind of thing that we found in our work with Panda Health, where, in addition to rolling out a cool new tool, we did this work on active change management. We brought people along, showed them how to use the technology, and had sessions where we hosted a bunch of them together and had them learn together about how to use the technology in the future. That was a really important part of rolling this out, and that kind of change management will be really important here as well.
I do think that takes a bit of time, but my expectation is that this will be one of the faster rollouts of software in health care history. I think that's already been happening with people's use of AI, and I think that's just at the beginning. I wouldn't expect it to take years and years for people to adopt AI more and more in clinical workflows.
One of our goals for this year is for AI-assisted care to become more part of the norm of care. I think by the end of the year, that'll be the case, we hope. But I don't think it'll happen over weeks, which is something that we AI innovators are used to, I think.
Let's talk about the connection between health and AI safety more broadly. That, I think, is super interesting. I remember the classic question from Ilya once upon a time: How do we teach AI to love humanity?
I understand that notion is not entirely gone from OpenAI and is still part of what you're thinking about. I'd love to hear more about that.
Yeah, absolutely. I think it starts with the foundations of the work that we've been doing on health at OpenAI and how we started it. I have a bit of a background as a researcher who cares about safety and has worked on safety research in the past. Part of the motivation for coming to OpenAI and working on health for me was thinking about the setting as a place that can provide concrete grounding for technical work on safety and alignment.
I had a feeling—and this was about 2 years ago—that a lot of the most ambitious work, the most medium- and long-term work on safety and alignment, was going on in toy settings or with math problems or things like this. It felt like if there was a setting where the problems that people were working on were well-motivated, that provided more concrete feedback loops to researchers and more short-term incentives for the research, the research could happen better.
That was part of the thesis for our approach to work at OpenAI. We've done a few kinds of research, and I'll talk about one of them, which is our work on calibration. Another problem that we've thought a lot about is scalable oversight: How do we supervise AI systems that are potentially more capable than us in certain ways?
This is a problem that we've actually had for some time in our work with physicians. In many ways, AI doesn't match your expectations of a doctor in its ability to integrate across a bunch of different modalities and things like this. But in many specific, narrow ways, models can even outperform physicians.
In those specific, narrow ways, when we evaluate models with physicians or have physician signal be part of the training signal for training models, we have to invest in research around the problem of how you supervise systems that are potentially more capable than you. That's one version of the problem of scalable oversight.
This is a problem that goes beyond health and is actually, I think, a really important problem for thinking about AI alignment—one of the more important problems there. Our work in health has given it a bunch of concrete grounding because we do have models that in some ways are more capable.
A lot of our focus has been thinking about the right ways to approach that problem and make it better. Again, that has included some of the work we've done on evals and some of the ways that we've pushed the high-compute RL paradigm in this setting, so that we can learn from the aggregate of the opinions of experts and things like this. That's been an important part of our motivation.
Broadly, I think a lot of where this has been heading for us has been thinking about how we extract the right personas or characters from the model. If the models become more and more capable over time, they have more and more inputs. Let's say they have more and more context about a patient's medical record because users are more proactively uploading them or the activation energy is lowered via the product. They also have more ability to take output actions, whether that's telling you things, outputting a note that you can give to your doctor, or other things that they can do.
Probably the most important thing in the very long term is whether the models are the kind of models that would do the right thing for the patient, the user, the clinician, or the researcher. That's the kind of thing that we've been investing in a bunch: How do we think about extracting the right kinds of personas from models? How do we do so in a way that's surgical, that gets the best aspects of the personas and gets less of the parts that we don't like?
For example, avoiding a bias toward being overconservative in challenging situations involving unclear medical consensus, while also navigating the uncertainty extremely well.
And so that work, which is closer to persona, character, or soul in its form, we're grounding a lot in work in health, which has been a really exciting advance. I hope we'll have more to share on that in the future.
Can you talk a little bit more about scalable oversight? This was, I think, at one point kind of the plan, right? The Superalignment team, I think, was premised on this idea. Maybe there were other ideas, too, but certainly my understanding was that scalable oversight was going to be a big part of it.
I think one of the things that stood out was that the strong student sometimes was just ignoring or overriding the instruction of the weak teacher. We're the weak teacher in this situation, right? How are we supposed to feel about the already emerging trend, as of probably at least 2 years ago now, that at times the strong student just decides that it should ignore or exercise its judgment despite the fact that it may contradict what the weak teacher—in this analogy, humans—is saying?
Have there been paradigmatic advances in that that give you confidence this is really going to work? Or is it the usual understanding I have today from OpenAI broadly around safety: it's going to be a defense-in-depth strategy, where everything will work a little bit, we'll gradually chip away at the problem, and hopefully, with enough layers of defense, it'll be okay?
Do you have more ambitious ideas about what scalable oversight can still achieve than that?
I think my thinking about how scalable oversight will proceed has changed a little bit over time. I think of the problem as having 2 parts. One is what I call rater scaling, which is how you think about the right ways to elicit opinions and values from people or experts.
That's an example of something we've been investing in in health. You can imagine schemes where AI is part of that loop and helps improve the ability for humans to critique, for example, AI outputs. That's one area of work that the company has been investing in.
The second area of work is: Given some idea of what the values are, whether they're elicited using rater scaling or not, how do you spend a lot of compute on training the models to have those values? The framing for this has become more expansive over time, and sometimes we don't use the words “scalable oversight” to refer to it anymore. I think the same is true at Anthropic.
I call that internally value oversight. So there's rater scaling, and there's value oversight.
The second problem, even though we don't use the words “scalable oversight,” is also one where I think things have advanced quite a bit over time. One example of this is the work people are doing on specs and constitutions, and on measuring and improving models to adhere to them.
This is one way of saying we have values that we care about, and we want models to adhere to them very persistently. If you check various system cards or things like this, you'll see that the models have gotten much better at doing this. People have also been finding increasing generalization in training models to have certain personas or characters, whether they're complying with certain safety requirements, specs, constitutions, or things like this.
What I'd say is that the research has actually been advancing, but in a way that looks different from what it did before. It looks a little bit more like humdrum post-training than people had imagined it would look like, and I think this is a good thing. It's a little bit more grounded in how systems will look and how we will train them.
I don't think we've solved all the research problems, but I think sometimes we're working on the research problems without referring to them in the same way as we used to.
Would you say the core reason to think that this will work is that a somewhat less capable model with a super-big inference budget can be expected to catch—or, if we're thinking of a deceptive failure mode, detect when—the smarter model is doing something wrong, even if it's more capable per token? A somewhat less capable model with a much bigger token budget is kind of the idea.
This idea has been out there for a while, but it really sparked from what you said: spend lots of compute to make it work. Just giving that thing a lot of time to think can be enough, as long as the delta isn't so bad that it can't detect anything that's going wrong in the more capable model.
As long as you've got a small enough delta and a sufficiently aligned current model, then in theory you can continue to bootstrap your way into ever more capable models without losing the alignment. Would you say that's kind of the core of the idea?
I think that's part of it. I think another part of this is that the task of discrimination or critique seems to be easier than the task of generating good outputs.
When we study the performance of monitors, a given model can monitor itself pretty well, especially given privileged information like the chain of thought. That can be a little surprising on its face, but if you think about the fact that discrimination has been better than generation—and that this underlies a lot of work in things like RLHF and Constitutional AI—then it's not so surprising.
I think that's another part of it. But I don't think we have a full understanding of how this will scale, especially when we're thinking about the regime of having trusted but less capable models align a more capable model.
So far, I've mainly been talking about this: If we have the values in some format, whether they're from humans or models, how do we instill those into models in a way that's trustworthy? That's a part of the problem. I don't think the whole problem is solved, for sure.
How do you think we're doing on safety as a whole? I have been an Eliezer reader since 2007, and actually read a lot more of his stuff way back in the day than I have more recently.
Broadly speaking, it's like, “Hey, this has in some sense gone amazingly well relative to baseline expectations.” We do have models that undeniably have a pretty good sense of human values, and that's manifestly obvious on a day-to-day basis. That was considered to be a very unlikely outcome years ago.
If you were to teleport back to 2007 and drop a post on Overcoming Bias indicating as much, it would be shocking—or considered laughable—that this would be accomplished in this way. And yet we still see these problems.
My mental model of the seesaw that it seems like we're on is that every generation of new models has new capabilities generally, and then it also usually seems to have some new emergent problem. Whether it's deception, or now eval awareness has become front and center, we sort of say, “Okay, that's a big problem.”
The next generation gets more powerful. We tamp down that last problem. It doesn't go to zero, but it's at least reduced. It seems like we're headed toward this strange world where, if you extrapolate both the measured trend and the roughly two-thirds to 1-order-of-magnitude reduction in bad behaviors that seems to happen once a given bad behavior is recognized and addressed, then if you extrapolate that out to 2028 or whatever, you're like, “Okay, I can imagine a model now that can do a month's worth of human work or a couple of months' worth of human work.”
But it also maybe has, for any given run, a 1-in-1,000, 1-in-10,000, or 1-in-100,000 chance of actively screwing me over in some super-bizarre way. That seems like a really weird thing to contemplate. It's a really weird world to live in.
If we believe in straight lines on log graphs, that seems to be where we're going, right? Do you think that's where we are going? If not, do you have a different mental model of what the balance between alignment and emergent problems would be in a couple of years' time?
I think the future world will definitely be pretty weird. It's very hard for me to predict what will happen. I do think models will become more capable, and I do think they'll be able to do things over longer time horizons. I think that's very important.
I also think we've been pleasantly surprised by the extent of safety generalization, or alignment generalization, which underlies the trend you're pointing out of various safety benchmarks or failure modes decreasing over time. That's been a relatively pleasant development.
If you think about it, that was the case during the 2 major scaling laws we've had for deep learning. One is the pretraining scaling law. What we found is that when we scaled up pretraining, we had models that were relatively good world models and could be relatively easily tuned via SFT, or small amounts of supervised learning, to exhibit a certain persona and be helpful, useful assistants.
And that's a lot of what the early work on instruction tuning and things like this did, and also what the early work on Med-Prompt did for the health setting. Which is really great. And I think we saw a lot of generalization there. I think there was a question for the reasoning setting about whether that generalization would hold or whether we would see similar generalization. So far, my sense is that at scale, we do see it, and at small scales, we saw less of it.
That's a promising development, which is that if you're able to figure out the right ways and patterns to get these models to be broadly beneficial and not harmful, then even if they're put in settings where they're doing things that you didn't foresee—doing a month's worth of work in days or things like this—they can continue to generalize and be safe in those settings, and those safety curves will continue to go down.
That said, I think you're right in pointing out there's a rapid increase in capabilities and surface area in which these models will be used, and also a rapid decrease in safety failures that we're seeing over time. It's a little bit hard to predict how both those curves will net out. I do think it's really important—and this is why I care so much about the parts of our mission that are about proactively making the tangible benefits happen, and also about working on mitigating safety risks—that we really stay ahead of these curves and think about the right ways to shape them in the right direction.
What does a Move 37 look like in health, and are we going to see that in the current paradigm, or would it take a deep integration of modalities, like we were touching on earlier? Another idea that I'm enamored with is this idea of maybe a different kind of training objective.
Instead of getting right answers to questions, what if it were more about, at a fundamental level, predicting what the state of health is going to be for a patient in a way where you could even begin to do in silico experiments? I've got this whole profile of this patient—what if I change this variable? What would their health look like in that case? That seems like it might be a pretty different paradigm from reading the whole internet and making sure you give me the right answer. What do you think? Touch on any of those that you want, but maybe also just take the opportunity to zoom out and give us a sense of what your big visions and ambitions are, and what you think people can expect as you guys continue to do your thing over the next year-plus.
Yeah, I love the Move 37 framing. Just to explain the reference, this is a reference to the famous Move 37 during a game between AlphaGo, the AI system playing Go, and Lee Sedol during one of their now-famous games. The interesting thing about this move was that everybody agreed humans would not have made it, but in hindsight, it was brilliant and was key to winning the game.
I think your question is: Can we imagine a world where models are able to do something and maybe make some kind of interesting prediction that a human probably would not have been able to make, but that was, again, brilliant and impactful in hindsight? My view is that this is not too far away. In many ways, many people report to me that they saw many doctors for their case, and only after talking to ChatGPT was ChatGPT able to flag the thing that they then shared with their doctor, after which they were able to come to a diagnosis together and things like this.
I think that seems to happen somewhat routinely, and whether it rises to the level of a Move 37 or not, I think, is a matter of taste depending on the case. What I'd say is that you should expect world models in health to improve. What I mean by world models is that we should expect our models to have a much better understanding over time of us, our health, our trajectories of health, and how that intersects with our biology.
I think that's going to be really key to thinking about what this Move 37 could look like. You mentioned this idea of simulations of people and simulating things in silico. I think full simulations are obviously very expensive and difficult, and I don't think they're what models are most designed to do in their current form. So I would say that they probably would look a little bit different from the current kinds of training that people are doing.
But if it's more along the lines of, given a lot of context about a user, predict something interesting about them or predict the results of some intervention, I think this is something that models, first, are already getting better at and, second, could get a lot better at over time. I think that would be potentially quite impactful.
I think we'll start to reach a point where there'll be more and more clear demonstrations of this. You're starting to see interesting and increasingly clear demonstrations of models doing interesting science and math in public. I think you'll see a little bit more in this space as well. But I do think it'll take a little bit of time.
My view on paradigms is that you can get a lot out of extending and tweaking the current paradigms, and you can, in fact, view everything that has happened so far as just one paradigm. Between pre-training and scaling reasoning-style training, I think you can get quite a lot of juice, and I think even integrating additional modalities and things like this doesn't require too many additional changes or tweaks.
I think broadly, our team's mission today is to do whatever it takes to ensure AGI is beneficial for human health and all of humanity. We see that happening through 3 channels. The first is helping consumers understand and navigate their health, which is already happening through ChatGPT Health.
The second is empowering the health system and thinking about the ways in which AI-assisted care can become part of the standard of care. It can reduce the extent to which clinicians are bogged down by paperwork, so they're spending more time seeing patients and thinking about the problems that matter most for improving care.
The final thing is really pushing up the ceiling of research. I found your vision pretty inspiring. I think you can expect a lot of the core problems in health care that we've talked about—the fragmentation of data and fragmentation of the patient experience—to improve. As those happen, you should also expect an acceleration of the application of intelligence to that data.
I think biology and health are areas in which marginal gains in intelligence have the most obvious value in solving more problems for humanity. There are many examples of previous breakthroughs in biology where there was nothing stopping that breakthrough from happening 5 or 10 years earlier except more human ingenuity applied to it. There's no physical blocker.
When that's the case, I think you can assume that long-running models and long-running agents connected to the right data can do really incredible things. The hope is that, in addition to our existing work, which has really been focused on how we raise the floor of human health, we start to raise the ceiling of human health as well.
Yeah, it's going to be not just the AI doctor, but also the AI biomedical research scientist. Wow, okay. It's an exciting time to be alive. I'm increasingly saying that at the end of these conversations these days. Anything else you want to make sure people are aware of that I didn't ask you about?
Like I mentioned earlier, I think the right way to think about our work for health at OpenAI is really as operating in 3 phases. One is really laying the foundation, and a lot of that is focused on our work on safety. Examples of this are our work on HealthBench, where we had this evaluation that you can run offline of large language model performance and safety in health.
We had this study of the first AI clinical copilot with Penda Health. We had a bunch of model improvements over the last year, and all of these laid the foundation for the work that we're doing today.
The second thing is this rise in adoption, which we've been seeing. We've been seeing an incredible rate of individual users, especially patients, adopting the technology, and we've seen hundreds of millions of people asking health questions a week, which has been rapid growth since last year.
The final thing is really the future of scaling the impact of this work. That's where our work on ChatGPT Health, which is this consumer-facing product that enables you to connect your health data with additional privacy protections, and also our work on ChatGPT for Healthcare, which is for health systems and really enables health professionals to use AI as a copilot in their workflows, come in.
Between all of these things, I'm really excited about the work that we're going to be doing to scale the impact of this work in the next year, and I'm looking forward to what's to come.
Karan Singhal, head of health at OpenAI. Thank you, legitimately, for all your hard work. It really is incredibly valuable and incredibly impactful, and that's obviously going to continue to grow exponentially along with so many things in the AI space. Thank you for being part of The Cognitive Revolution.
Thanks for having me.