[BidClub_]
The Cognitive Revolution · · 103 分钟

Bioinfohazards:Jassi Pannu 谈如何管控 AI 模型学习的危险数据

Jassi PannuNathan Labenz

YouTube
TL;DR
  • Pannu 的方案只限制约占数据 1% 的一小类数据——这一比例只是粗略估计,尚未被测量;这些数据直接把病原体序列与传播性、毒力、宿主范围或免疫逃逸联系起来,从而保住开放生物学。 绝大多数原始序列数据仍会开放;敏感的“功能数据”则置于分层的生物安全数据等级中,研究人员把代码带入可信环境,而不是下载底层数据。Nathan Labenz 尽管自称是终身信奉技术乐观主义的自由意志主义者,相信数据“想要自由,也理应自由”,仍支持这些管控,因为未来的智能体会利用网上留下的任何高信号信息。

  • 选择性排除训练数据,可能提供一种罕见的低成本安全干预,而且不会广泛削弱生物模型的有用能力。 Evo 2 排除了感染人类和其他真核生物的病毒序列,ESM-3 的开发者则对比了过滤版和未过滤版模型;危险病毒任务上的表现大幅下降,但其他领域能力仍然保留。在 Evo 2 的一项评估中,相关表现“基本等同于随机”,而非只是逐步变差——这表明范围窄的数据管控不必把模型整体变成“笨模型”。

  • 相关威胁正从只能由专家驾驭的生物学,转向极端组织、单独行动者,并最终转向自主 AI 系统。 大流行病原体不适合作为民族国家武器,因为它们无法定向攻击;要保护本国人口,就得事先接种疫苗,而且很难不引人注意,因此主要大国对管控存在共同利益。迫切性在于,已有系统能根据手机照片为实验室工作排障,Opus 4.6 还会自发在 Hugging Face 上找到加密的基准数据集并解密其中答案;Nathan 认为,对危险生物信息设置的壁垒应假定可被发现,而不是长久有效。

  • 2012 年的禽流感实验仍是最清晰的警示:物理实验及其公开发表的结果都可能制造不可逆的风险。 两组研究者让一种当时被认为病死率约为 60% 的禽流感病毒株在雪貂之间具备哺乳动物间传播能力;其中一组“只需要 5 个突变”。此类研究后来基本被撤销资助,但仍属合法;私人实验室的透明度很低,而天花序列已经可以在概念上与一份公开发表的马痘逐步合成方案结合。

  • AI 可以压缩计算疫苗设计的时间,但 COVID-19 的经历表明,发现阶段“不是瓶颈”。 SARS-CoV-1 的研究早已确认刺突蛋白的重要性,因此 SARS-CoV-2 测序后可以迅速完成设计;临床试验、监管、制造和全球分发耗时远长得多。因此,更被忽视的机会不只是提升分子生成能力,而是开发能提高临床试验成功率、扩大实体部署规模的工具。

  • 真正具有决定性的平台跃迁,是把研究智能体、生物基础模型、机器人、实验和新生成的因果数据连成一条闭环。 Pannu 对更深层多模态融合给出的保守时间窗口是“5 至 15 年”,因为自主实验室的吞吐量仍是约束;仅仅汇总药企的专有数据,帮助也只会有限。如今的数据集仍是“手工艺化”的、观察性的、有偏且不一致;机器人扰动有望生成系统性的因果证据,让 in-silico 模型真正具备预测能力。

  • 没有任何单一管控能让生物学安全,因此纵深防御的基础设施路线图覆盖“延缓、威慑、检测和防御”四个环节。 优先事项包括强制 DNA 合成筛查、安全的数据与算力环境、跨供应商订单监测、被动式全球病原体监测、PPE 储备,以及远紫外线空气灭菌等建筑环境防御。核心论点是纵深防御:数字进攻越来越便宜,而反制措施仍受物理世界瓶颈限制;因此,韧性要求多道相互独立的防线同时起效。

摘要 · 为研究而整理的核心内容

1. 疫情检测仍要等病人发病

  • Pannu 的出发点是一个制度性空白:社会没有针对陌生病原体的国家或全球被动式预警系统,无法像雷达监测 ICBMs 那样提前预警。检测只有在患者出现症状、前往医院并引起临床医生警觉后才开始。

  • 医生首先会用鼻拭子检测流感、RSV、鼻病毒等熟悉病因。只有当这些检测均为阴性且病情严重时,医院才会升级到宏基因组测序;因此这套系统应对已知病毒尚可,却难以及时识别真正的新型病毒。

  • 研究认为 SARS-CoV-2 大约在 2019 年 11 月下旬或 12 月出现,但直到 1 月才完成完整测序。随后,一名中国研究人员未征得政府许可便公布了序列,推动诊断设计和产能扩张;相比之下,流感的全球序列网络依赖各国实验室主动贡献序列。

  • Nathan 一再被要求分享儿子癌症数据的经历,暴露出其中的错位。Pannu 说,医学对个人隐私有很强的保护,但面对“整个地球都是你的患者”这类社会性风险,相关机制要弱得多;美国医院、州和医疗网络彼此割裂,使问题进一步恶化。

2. 受保护的临床数据无需复制也能保持可用

  • 英国的 OpenSAFELY 是 Pannu 认为可对照美国碎片化体系的案例。它于 2020 年创建,让研究人员可以访问覆盖英国 95% 人口的临床记录,同时将底层私人信息留在受控平台内。

  • 它最关键的设计选择,是“迫使研究人员来找数据”。研究人员提交代码,从不需要查看或持有底层私人记录——Pannu 后来把这套模式从隐私保护延伸到了危险生物数据。

  • 这一架构的重要性在于,访问控制不必等于科学瘫痪。共享的安全环境既能减少泄露,又能让研究人员不必逐一与每个掌握相关数据的机构谈判。

3. 快速疫苗设计依赖多年积累的科学基础

  • Nathan 指出,第一份 COVID-19 mRNA 设计似乎在拿到序列后几天内就完成,而且在成为最终接种的疫苗前几乎没有改变。Pannu 认同从刺突蛋白反推 mRNA 序列的速度极快,但强调计算“显然不是瓶颈”。

  • 研究人员并非从一张白纸开始:SARS-CoV-1 的研究早已确认刺突蛋白的重要性。正是这些积累下来的知识,让团队在 SARS-CoV-2 测序后立即选定靶点,尽管 SARS-CoV-1 本身在演变成全球大流行前就已被控制。

  • Nathan 关于 gain-of-function 的追问紧随其后:当临床试验、监管、制造和全球分发主导时间线时,缩短靶点选择的时间价值有限。Pannu 转而强调 AI 在这些被忽视的实体环节中的机会,包括提高临床试验成功率,以及改善反制措施的生产和部署。

4. 大流行武器更偏向非理性行动者,而非民族国家

  • Pannu 将毒素、有可靠抗生素反制手段的非传染性生物体(如炭疽),以及最极端的情形区分开来:一种缺乏诊断、治疗手段或疫苗的新型大流行病毒。最后一类的后果远为严重,因为它会自我复制并扩散到攻击者无法控制的范围。

  • 这种不可控性使大流行病原体不适合作为民族国家武器。要保护本国人口,就必须设计疫苗并在不被察觉的情况下大规模接种;因此 Pannu 把重点放在恐怖组织、小型极端主义团体和单独行动者身上——他们更不受理性威慑约束。

  • 这一威胁模型也为美中协同创造了可能:两国都能从防止下一场大流行中受益。若从匿名下载转向身份核验、用途审批和使用记录,就能形成“差异化授权”——为病毒学家和反制措施开发者保留访问权限,同时提高恶意用户的门槛。

5. 功能数据,而非原始序列规模,才是瓶颈

  • 在 Pannu 所说的 Carlson 曲线下,测序成本下降速度快于 Moore 定律,生物学已经被海量 PB 级数据淹没。仅 GenBank 就包含超过 40 PB 的数据,其中大多是未注释、原始且质量参差的测序数据。

  • Protein Data Bank 则完全相反:多年来靠艰苦结构实验积累出的数据集还不到 1 TB,理论上放得进一只 U 盘。过去,解析单个结构就可能耗费一名研究生整个博士阶段,说明数据量可能掩盖信息价值的差异。

  • 序列数据大多是观察性的。Evo 2 等模型表明,序列可以支持蛋白质相关任务、基因组生成,并在不同生物尺度上运行;但 Pannu 说,更关键的输入可能是因果性的“功能数据”:基因敲除、系统性扰动,以及对病毒蛋白如何与人类蛋白结合的测量。

  • 因此,拟议的管控应针对把病原体与传播性、毒力、宿主范围或免疫逃逸联系起来的数据集——这些正是政府在湿实验研究中已经审查的属性。清洗现有互联网数据将是“赫拉克勒斯式”的苦差事,且大概率徒劳;真正可操作的瓶颈在于未来由 Genesis Mission、OpenAI Foundation 数十亿美元级投入以及 Chan Zuckerberg Biohub 等项目资助产生的因果数据。

6. 5 个突变同时暴露实验室和信息风险

  • 2012 年,两个独立团队对禽流感展开实验;当时这种病毒被认为病死率约为 60%,但尚不具备高效的人际传播能力。研究以雪貂为哺乳动物模型,提高了病毒的传播性;其中一组发现“只需要 5 个突变”。

  • Pannu 认为,两篇论文当时同时投给 Science 和 Nature;两家期刊在论文披露使这项工作成为可能的突变后,向美国政府发出警报。这一事件同时包含两重风险:增强后的病毒可能因普通人为失误逃逸,而论文公开则可能向他人提供复现实验所需的突变、操作流程和反向遗传学知识。

  • 物理控制体系并不完美:CDC 的冷冻柜里曾发现被遗忘的可存活天花样本,尽管当时据称全球只有 2 个地点存有活体样本。信息控制更难——天花序列已经公开,研究人员逐步合成马痘这一近亲的操作方案也已公开。

  • Nathan 认为,这类研究的防御性收益很弱,因为更快设计疫苗并不能消除主要响应瓶颈。Pannu 也同意,常规疫苗研发不需要让病原体变得更易传播、更具毒力或更善于免疫逃逸;新冠之后,此类研究已被大幅撤资,但并未被明确列为违法,私人实验室仍是一个可见性盲区。

7. 生物学模型目前彼此分立,但正走向同一闭环

  • Pannu 主要把通用 LLM 视为知识“赋能”工具:它们可以教非专业人士生物学,也可能回答如何非法获取武器或天花的问题。因此,前沿开发者使用分类器和拒答机制,拦截一小类操作性信息。

  • 专门化的生物设计工具提供的是能力,而不只是知识。它们通常要求使用者具备计算生物学专业能力,却能让专家完成此前不可能完成的工作——例如从序列快速推断蛋白质结构、探索替代设计,或预测分子结合。

  • Evo 2 和 ESM-3 属于第三类:更大型的生物基础模型,试图在蛋白质、基因组和调控系统之间推断可迁移的规律。它们对数据和算力的需求,使 Google DeepMind 和 Evolutionary Scale 等大型组织比单个学术团队更有优势。

  • 这些类别的界线可能在“整合式工作流”中消失:智能体设计实验,自主机器人执行实验,结果数据更新基础模型,智能体再选择下一项实验。本期的框架还包括利用手机图像排查实验室故障、持续数日的自主研究,以及 Opus 4.6 绕过加密恢复基准答案。

8. 机器人,而非企业数据共享,是下一次跃迁的瓶颈

  • Nathan 追问,生物学何时会获得类似图像与文本联合理解模型所展现的深层多模态融合——不只是 LLM 调用蛋白质模型,而是共享权重,能够在语言、序列、结构和实验性证据之间流畅推理。

  • Pannu 给出的诚实预估是“也许在未来 5 至 15 年”。目标很明确,但因果数据仍要靠物理实验获得,而实验室自动化仍受机器人技术、样本搬运、流程一致性以及扩大现实世界工作规模的难度约束。

  • 药企拥有有价值的专有数据集,也在训练内部模型,但强制共享只会让行业“再向前一点”。现有生物学数据“有点手工艺化”,且多为观察性数据,因而带有偏差、杂乱不一;系统性的机器人扰动和谨慎复现代表着一种根本不同的数据生成范式。

  • 最终的愿景是让 in-silico 模型替代部分湿实验工作,预测药物在细胞和动物中的作用,并把临床结果预测到足够准确,从而减少试验次数、提高成功率。Pannu 反复强调,这仍是愿景,而不是当前能力。

9. 当进攻仍停留在数字世界,双重用途能力会带来失稳

  • Pannu 愿意为虚拟细胞模型、临床试验优化、反制措施的生产与分发等相对低风险应用“踩下油门”。更难的问题不只是哪些能力应当存在,而是谁能获得这些能力、在什么条件下获得,以及能力成熟到什么程度。

  • 她刻意设定的未来主义红线,是一种能够快速合成病原体、且成本低廉的自主机器人。这类设备不应作为现货产品出售;即便其底层自动化也服务于合法生物医学研究,访问和使用仍需要被追踪。

  • 病毒设计体现了双重用途陷阱。更好的病毒工程可以推动基因疗法,或推动感染细菌的噬菌体研发,但一个能通用设计这类病毒的系统,也可能被改用于人类大流行病原体;而疫苗及其他反制措施仍被卡在物理世界的时间线上。

  • Nathan 反驳称,即使是有益的全细胞模型,也可能被用来暴力枚举病毒变体。Pannu 承认,几乎每一项生物医学进步都存在有害路径,但她以直接性、后果和所需专业能力划线:一个只需有限生物学知识就能生成针对人类病毒设计的模型,比需要多模型协同、且所需专业能力高到让民族国家大概率会选择其他武器的工作流,更值得担忧。

10. BDL 管控数据,同时保留开源模型

  • 拟议的生物安全数据等级框架包括 5 个等级,从 BDL0 到 BDL4,参照实体实验室的生物安全等级设计。它把遏制对象设为数据,而不是病原体或模型,因为学术生物 AI 高度依赖开放模型,而有价值的功能数据仍然昂贵且集中。

  • BDL0 覆盖绝大多数生物信息,完全不设访问限制。BDL1 增加基础身份和账户要求;更高等级逐步要求研究人员说明拟开展的项目,并在使用与危险病原体属性更直接相关的数据前获得批准。

  • 更高等级的触发条件包括传播性提高、宿主范围扩大、动物到人传播、更高毒力,以及免疫逃逸或疫苗逃逸。任何使用 BDL3 或 BDL4 数据训练的模型也必须安全处理——否则发布模型只会把受限能力重新包装并传播出去。

  • Pannu 称 99% 的开放数据是合理估算,而非测量事实,因为不存在全面清单。最高等级可能只影响几十家或更少的专业实验室,其中许多本来就在非正式限制访问;最大的未知数,是政府资助视野之外私人产生的数据。

11. Evo 2 和 ESM-3 表明选择性排除数据可行

  • 核心实证问题是:一个强大模型在学会生物学更深层规律后,能否通过插值绕过训练数据中的微小缺口?如果可以,排除危险数据只会增加管理成本,并不能真正移除相关能力。

  • Evo 2 团队在保留感染细菌的病毒信息的同时,剔除了感染人类及其他真核生物的病毒序列。预训练后,评估显示其病毒蛋白相关任务的表现显著变弱,生成对应功能性病毒的序列能力也下降,但其他生物学领域仍保持有用。

  • Evolutionary Scale 训练了过滤版和完整数据版 ESM-3,使研究人员能够直接比较其在病毒蛋白功能预测等任务上的表现。Pannu 形容 Evo 2 受影响任务的表现“基本等同于随机”,而不是小幅下降——这是本期最有力的证据,说明有针对性的排除可以消除危险能力,同时不让模型整体失灵。

12. 可信环境可以让安全机制服务研究人员

  • 可信研究环境通过把敏感数据集中存放、让获批研究人员把代码带入其中,来把这套方案落地。理想情况下,它们还应提供足够的算力用于模型训练,不过 Pannu 说这是否可行尚不确定。

  • 她不确定政府是否应该直接建设并运营所有环境;研究人员发现,一些公共系统使用起来很笨重。大学和大型合作项目可以按政府制定的标准自行建设平台,把责任放在机构而不是单个研究人员身上。

  • 集中化也可能是科学收益,而不只是合规负担。AI 研究人员本来就希望获得整合的数据和算力,而不是把数据集分散在各个实验室;设计良好的可信研究环境可以在增加身份认证、监测和限制的同时,为防御性病原体研究“踩下油门”。

  • 数据管控在隐私和人类基因组治理中也有先例。Pannu 认为,Biden 和 Trump 两届政府在 gain-of-function 监督上正在出现跨党派推进,并希望为美国 AISI 配足资源和人员,把开发者目前各自临时开展的能力评估标准化。她还肯定了美国和英国 AI 安全机构在生物安全方面的有益工作。

13. DNA 筛查只有在剩余的 20% 不能掉队时才有效

  • 大约 80% 的基因合成供应商会自愿筛查订单。自动序列匹配先检查订单材料是否类似埃博拉、天花或其他受关注病原体;被标记的订单再交由人工专家处理,同时通过客户身份核查确认下单者是谁、订单是否符合合法研究。

  • 筛查成本下降已经消除了大部分原有负担,但 Nathan 指出了攻击者的明显优势:恶意客户只需转向不筛查的供应商。Pannu 说,当前推动的方向是让所有供应商都必须筛查,而不是把 80% 的参与率视为足够。

  • 分散下单又造成另一处弱点。有人可以从多家公司分别购买小片段,但供应商没有实时合并这些信号的系统;解决它需要为安全信息共享提供法律授权,也需要在制度上明确由谁负责协调。Pannu 提到 FBI 只是一个可能的例子,并非已经敲定的方案。

  • 全自主云实验室会形成未来的网络攻击面,但 Pannu 强调,目前实验室离这一情形还很远,仍依赖人类搬运样本。如果具备病原体能力的自动化系统出现,系统必须能够抵御恶意行为者远程部署“1000 个智能体的大军”来夺取控制权。

14. 生物学需要分层防御,而不是单一制胜理论

  • Pannu 不赞成寻找生物学版的单一核威慑理论。生物学是分布式的、双重用途的,其价值恰恰在于很多人都能从事;可行策略是在四层之间建立纵深防御:“延缓、威慑、检测、防御”。

  • 延缓包括数据管控和强制合成筛查。威慑包括国际社会对生物武器的禁令,尽管执行乏力仍有价值——但对不理性或自杀式行动者无效,因为他们不会对传统惩罚作出反应。

  • 检测意味着建立被动式“生物雷达”,能够在不等待患者出现症状或依赖自愿上报的情况下发现陌生病原体。对具有较长无症状期的感染而言,这一点尤其重要:Pannu 提到早期 HIV 传播,认为这类事件正是社会希望提前得多识别的对象。

  • 防御不应止于疫苗,还包括人们已经习以为常的环境防护:过滤水可以阻断霍乱,纱窗可以阻挡蚊媒疾病,但建筑物没有针对空气传播病原体的同等默认防线。远紫外线和乙二醇蒸气或许可以在无需识别病原体的情况下被动灭菌空气;Nathan 的家人已经在儿子住院期间使用 Aero Lamp,既是个人防护,也是在支持这一新兴市场。

Nathan Labenz

Today my guest is Jassi Pannu, an assistant professor at Johns Hopkins who recently co-authored an important paper calling for the creation of access-control systems meant to prevent the dissemination and misuse of functional biological data from which AI models could learn extremely dangerous capabilities, such as the modification or even de novo design of highly contagious and deadly viruses.

We begin with an overview of the biosecurity landscape today, including how new viruses are detected, how patient data is aggregated and analyzed in the context of a new threat, and what the pipeline from DNA sequence to vaccine candidate looks like today. The good news is that we are able to design new vaccines amazingly quickly, at least for viruses that are similar to others we've seen. But there is unfortunately a lot of bad news as well.

In 2012, for example, 2 research groups independently published results showing that wild-type bird flu, which already had an estimated 60% fatality rate but couldn't spread between humans, could become mammal-to-mammal transmissible with just 5 mutations. Such gain-of-function research has been broadly defunded since the COVID-19 pandemic, but it does remain legal, and visibility into the experiments private labs are conducting is low.

Governments, Jassi says, aren't likely to develop bioweapons capable of causing pandemics for the simple reason that, short of vaccinating their populations in advance of an attack, they can't realistically expect to control them. But with AI capabilities crossing critical thresholds month by month, the threat from extremist groups and even lone actors is quickly moving from theoretical to deadly practical concern.

Consider that Jeffrey Irving, chief scientist at the UK AI Safety Institute, recently highlighted for me that today's frontier models can troubleshoot laboratory experiments from a cellphone picture better, on average, than PhDs. And in just the 10 days or so since we recorded this conversation, we've seen Andrej Karpathy's autoresearch framework demonstrate that AI agents can run and make research progress for days on end.

Even more to the point, Anthropic just reported that Claude Opus 4.6, when faced with a benchmark challenge it couldn't solve, spontaneously located the full benchmark data set on Hugging Face and then figured out how to decrypt the solutions, which were encrypted in the first place to prevent the answers from leaking into training data. It did this all in order to get a single question right.

With reasoning AIs already capable of spontaneously overcoming such barriers to information, I think we should expect that future research agents will find and exploit any signal-rich data that exists anywhere on the internet. And with the smallpox sequence and the horsepox synthesis protocol already published online, and biological data poised to grow superexponentially in the coming years, we have real reason to worry and ample cause to get serious about implementing data controls before the situation gets truly out of hand.

Again, though, there is good news. Recent work by the teams behind the Evo and ESM families of biofoundation models showed that strategically excluding key data sets, such as the DNA sequences of viruses that infect humans, dramatically reduced models' performance on dangerous tasks while leaving their desirable capabilities intact.

This means that the vast majority of biological data can remain open source and open access. Indeed, Jassi and co-authors' proposal for a biosecurity data-level framework, which echoes the existing Level 0 to Level 4 biosafety framework for physical labs, would subject only an estimated 1% of data that connects pathogen sequences to dangerous properties to any additional restrictions.

Even then, structures such as trusted research environments, which allow researchers to run code on data without transmitting that data from its secure location, would still support valuable research. Once again, despite my personal history as a lifelong techno-optimist libertarian who broadly believes that data wants to and ought to be free, I find myself eager to support these control measures.

Of course, that's not the only opportunity we have to improve biosecurity. Toward the end, we also discuss the broader defense-in-depth strategy that biosecurity experts recommend: delay, deter, detect, and defend. This includes mandatory pre-synthesis screening of sequences by DNA manufacturers, investment in wastewater monitoring and other passive global pathogen surveillance, and practical frontline defenses like PPE stockpiling and far UV sterilization.

All of this is in everyone's shared interest, but it does require leaders to see beyond the current news cycle for long enough to make it happen. I certainly hope they do, but I also recommend taking individual action where you can, both to improve your own personal safety and to support the consumer market for biosecurity products.

My wife and I, for example, at our friend Jeff Kaufman's recommendation, recently purchased an Aero Lamp far UV light for use in my son's hospital rooms throughout his cancer treatment. I'd welcome additional suggestions for other products that could help us minimize disease burden today while also serving as a sort of private insurance against pandemics, if anyone has any recommendations.

For now, I would simply emphasize that, by default, we are fast approaching a world in which a rapidly growing number of people—and perhaps autonomous AIs as well—will have the ability to create deadly, transmissible, self-replicating viruses that could dramatically alter the trajectory of human history. It really does seem like we should do something about it.

With that, I hope you are properly alarmed by this scary, but solutions-oriented conversation about the sorry state of biosecurity and the rapidly rising threat from bio-savvy AI systems with Johns Hopkins Professor Jassi Pannu.

Today my guest is Jassi Pannu, an assistant professor at Johns Hopkins who recently co-authored a call for controls on biological data. We've heard quite a bit about the possibility that AI systems of various kinds could create new sorts of biorisks, and this is one attempt to put some controls in place to hopefully cut that off at the pass before it becomes a major downstream issue.

I want to get into that from a bunch of different angles, but I think it would be helpful to take a step back and lay some foundations. People who follow this feed know a lot about the AI side. They probably don't, on average, know nearly as much about the current state of play when it comes to biological data writ large.

I thought it would be helpful, and I'm actually very curious about some of these things I realized I didn't know as I was preparing for this. For starters, if you could take us into the moment that we had not so long ago—and hopefully won't have again, but very well might—when all of a sudden there is a new outbreak of something. Something we've never seen before, we don't know what it is, people are concerned, and patients in a particular hospital or city are showing up with concerning symptoms.

What happens? How do we turn that initial small patient population into actual knowledge about what we're dealing with?

Jassi Pannu

Yeah, I think that's a great place to start. To take a step back, it's important to realize how society comes to the conclusion that there is a new virus circulating. What is our mechanism for making that decision? We don't have a global or national alert system for this kind of thing in the same way that you have a radar system for ICBMs.

Right now, we largely rely on symptoms. Patients get sick, they visit a hospital, doctors get concerned, and they run a series of tests. Doctors always start with tests that are common, like influenza, RSV, and rhinovirus. We're just collecting samples by swabbing their noses and sending them to the hospital lab.

When these tests are negative and the patient is very sick, then we'll go ahead and run further tests. We learned during COVID-19 that this whole process is pretty lengthy. It works well for familiar viruses, but it does not quickly detect new viruses.

In the case of influenza, the situation is a little bit unique. We've had several influenza epidemics in the past, so we do actually have a global system to collect new influenza sequences. But it operates on a contribution model. The national lab has to send that sequence up to the global repository, and it's an active submission process.

We don't have a passive global alert system, and we can talk about the benefits of having something like that later. But once the community has decided, “Okay, it looks like there's a cluster of patients with a new virus, and we need to figure out what this virus is,” that's when we start to do more exotic tests, like metagenomic sequencing, to try to figure out the sequence of that pathogen.

That's what happened during the COVID-19 pandemic. We think, based on research, that COVID likely emerged around late November to December of 2019, but it was fully sequenced in January. The decision to share that sequence—the way it happened—was that a researcher sequenced the virus. This was a researcher in China, and they publicly shared the sequence.

They didn't request permission from the Chinese government, which is pretty different from how it might have happened for an influenza virus.

Jassi Pannu

And that's when we kicked off a process of designing diagnostics and trying to scale those.

Nathan Labenz

I recently had an experience where my son had cancer. People know about this if they've listened to this feed. I went through this process of sharing data with the broader medical establishment. Somebody, I think it was a big part of their job based on how much time they spent with me, came by and said, “Hey, I'm from the sort of data-sharing world, and I'm here to answer all your questions and get 1,000 signatures on 1,000 pages so we can hopefully use this data for the betterment of all humanity.” And I said yes to all that stuff.

I've still been getting stuff in the mail asking me to share data again with particular studies or groups or whatever. I'm wondering if it's different when it's a pathogen, because what I'm sharing in my son's case is information very specific to him, whereas you could think about it differently if it's a pathogen. It's not your pathogen, right? It's the world's pathogen in some sense. How much consent or opting in do patients have to do, if any, to get this data from them and into the higher levels of analysis?

Jassi Pannu

Yeah. What you're pointing to is the fact that we have very different systems for dealing with individual-level risk versus societal-level risk. We have a really strong set of protections around making sure that your son's data isn't inadvertently shared, that you consented to everything, and we're thinking about the privacy risks of all of that data. That's really to protect your son, the individual patient.

But when it comes to societal-level risks about a pathogen, where basically the whole globe is your patient, we don't have as good mechanisms for figuring out how to protect society from those risks. We don't have as good mechanisms for protecting the data that could lead to those consequential outbreaks.

Again, just to focus on the clinical data, that very much depends on where you are, what country you're in, and what system you're operating in. In the US, as it sounds like you've experienced, our system is very fragmented. It's different within different states and different networks, whether it's a public community hospital or a private hospital. That fragmentation results in different people having to ask you for permission over and over again, because there isn't a single unified repository for all of that data.

It is a bit different in other countries. If you look at the UK, they have the National Health Service. What the UK has been able to do is create a research platform called OpenSAFELY. This was something that was stood up during 2020, and it provides researchers access to clinical data for 95% of the UK population. It's really interesting because it doesn't require giving the data to the researchers. It forces the researchers to come to the data.

The code comes to the data set, and you can submit your code to the platform. You never have to see the private data. It's secure the entire time, and there's been a lot of positive reception to the OpenSAFELY platform. But we don't have an equivalent in the US where everything's unified in that way.

Nathan Labenz

Yeah, that's interesting. We'll come back to the different levels of protection or restriction that certain data sets should have in your mind. But let me stay on the response narrative for a minute.

I know that it sounds like it was actually maybe longer between when the virus first emerged and when it was first sequenced than it was from when it was first sequenced to when the vaccine was initially designed. I've read a couple of these magazine pieces that tell the story of a 48- to 72-hour period where certain researchers got the sequence, were able to do some analysis on it, identify a particular protein, which I think was the spike protein, and then plug that into an mRNA platform.

My understanding was that the final vaccine that I got wasn't that different from the initial design that they put together in a couple of probably pretty long days. I think it was even as early as February of 2020. Has that pipeline changed at all? Obviously, there have been so many different machine-learning models that have come out over the last 5 years that predict shapes better and predict what's going to bind to what better, but that's pretty hard to beat in terms of timeline. Would you say that if this were to happen again today, that process would look much different from how it looked then?

Jassi Pannu

Maybe first I'll talk about what happened during the COVID-19 pandemic and those different timelines, and then what we can reflect on regarding how this will look in the future with AI. You're completely correct that during the process of designing COVID-19 vaccines—and I'll focus on the mRNA platform vaccine as a primary example—the design process, the computational steps where you were looking at the spike protein and working backward to figure out what the mRNA sequence was that you were going to use for your vaccine design, was very fast.

But that was clearly not the bottleneck. There were many other steps that took much longer. So, yes, the process of knowing there was an outbreak, sequencing the pathogen, and figuring out what pathogen we were dealing with took quite a bit of time.

There was also a whole body of research that researchers relied on, which I think doesn't get enough airtime. Once researchers realized that we were dealing with SARS-CoV-2, they were able to look at this body of research that had been done on SARS-CoV-1, which was the original virus that caused severe acute respiratory syndrome many years prior. It didn't spread into a global pandemic, but there were patient cases that were isolated across the globe, and we actually managed to prevent that from becoming a global pandemic.

It was that research that resulted in researchers knowing that the spike protein was extremely important. They were then able to say, “Okay, we know the spike protein is important. Let's computationally design the vaccine.” That part was very fast.

Then it was about turning that into an actual vaccine: all the clinical components, the regulatory process, the different clinical trials that you have to go through, and then distributing that vaccine globally. Those steps took a lot longer than the actual computational design.

I think that now, with advancements in AI models, people are very optimistic about being able to design new proteins, antibodies, and vaccines. I would say that there's still going to be a bottleneck in scaling in the physical world. Clinical trials still remain a huge barrier, and scaling and deploying vaccines across the globe is a huge barrier. Figuring out if AI can speed up those steps will be really useful, and it's perhaps the neglected component of the overall pathway.

Nathan Labenz

One thing I've learned in looking at the question of security at frontier-model developers in the AI industry is that, in looking at responsible scaling policies, there are a lot of levels to the game of security. The different levels seem to correspond to different actors that you might be concerned about and how hard it would be to prevent them from doing bad stuff.

For the likes of DeepMind, OpenAI, and Anthropic, the general consensus seems to be that if a determined nation-state actor, such as China, were to want to steal the model weights, it would probably be able to do it. There's not too much that could be done to prevent that. How would you map the threat landscape when it comes to the bio-risk side?

Is it a similar thing where we have random crazy people versus somewhat more sophisticated groups of people, all the way up to nation-states? What do you think is reasonable to expect we can actually stop with all of the measures that we'll talk about potentially developing?

Jassi Pannu

Yeah, it's a really important question. Within the realm of biosecurity, there are a lot of different threats that people refer to when they're thinking about chemical and biological weapons. There are things like toxins—small-molecule toxins and protein toxins. Then there are organisms that are not transmissible between humans, things like anthrax, where we have reliable countermeasures, such as antibiotics that work against those things.

Then there's the more extreme end of pandemic threats: pandemic viruses that are novel, that we've never seen before, and for which we don't have diagnostics, therapeutics, or vaccines. When you think about that spectrum, the potential consequences of a pandemic virus are far higher than those of many of the other threats. This is all pretty obvious to us now after having lived through one.

It's also interesting that a pandemic virus is not a particularly desirable weapon for a nation-state. It's not targeted, and it's not easy to protect your own population. You'd have to design a vaccine and vaccinate your entire population. It's hard to do that without someone noticing.

So I think, in general, nation-states are not the primary actors that one is considering when thinking about pandemic threats. It is more likely to be people who are not motivated by rationality: smaller groups, terrorist groups, and potentially lone actors. Those are the folks that people are really concerned about.

That's why a data-control mechanism is most likely in the interests of all countries. I would say that China is equally invested in making sure there isn't a future pandemic as the United States is. So I'm hopeful that there can be some international cooperation—or, if not cooperation, at least an acknowledgement that data controls benefit both the US and China and nation-states globally.

Jassi Pannu

That's what I would say. In terms of data controls, how can you prevent those kinds of actors that I outlined—lone actors and smaller groups—from getting access to your data sets? I think that controls are meaningful there. Currently, the default is sharing that data publicly and making it available for anonymous access. It's extremely easy to access, and even putting minor barriers in place would make a difference.

The other important thing to consider is that we want to make sure defenders, or people who are advancing countermeasures research and virology research, have access to that data while limiting access to malicious actors. Controls can do that differential privileging, where you're privileging defensive use cases and limiting offensive use cases. You can track who's using it, give access to that crowd, and limit access to others.

Nathan Labenz

So, can we map out the data landscape as it exists today? I want to do this for both the data landscape and the models that are obviously spawned from the data. On the data side, I've heard it said many times that you can find the smallpox sequence on the internet. Then there's the question of: If that's true, why hasn't that turned into a crisis already?

I've heard various accounts, and I'd be interested in yours. More broadly, that's a known sequence for a known problem. You can tell me, but it strikes me that it's a small enough amount of data that's probably pretty hard to control or clean up from all the places where it might already have been replicated. It's a little bit hard for me to imagine a world where that's been scrubbed so thoroughly that somebody who wanted to find it wouldn't be able to, but maybe you have a plan for how we could get there.

But if we expand the scope of data of concern, there's tons of biological data in general, right? I know most of that you're not looking to restrict. So how would you draw the—I don't know if they're concentric circles or not—from the narrowest category, the smallpox sequence, which we probably shouldn't be passing out too freely, to somewhat larger categories, and then beyond that, everything that would be fine? How would you characterize those classes of data?

Jassi Pannu

Yep. Let's start with just the general, broad categories of data. The most abundant biological data that currently exists is sequence data. We are currently swimming in petabytes and petabytes of sequence data, a lot of which we don't know the function of. It's become extremely easy to collect and sequence that kind of data.

There's something called Carlson's curve, which is the equivalent of Moore's law for biology and DNA sequencing, and it has actually shattered Moore's law because it has become exponentially cheaper to do DNA sequencing. That's resulted in a lot of passive DNA sequencing and collection—just sequencing everything. That kind of data is available in government-supported repositories, things like GenBank, which is supported by the NCBI, part of the NIH in the US. GenBank alone has more than 40 petabytes of unannotated, raw, poor-quality, frankly, DNA sequencing data.

That is a large part of why there are efforts to build AI models using that data, because it's abundant. When we think about other types of data, there are protein sequence databases, and then there's the Protein Data Bank, which formed the basis of AlphaFold and contains protein structure data.

That data was collected by hand over the years. I'm sure a lot of us have heard the story of painstaking efforts to do experiments to figure out protein structure, where one grad student would spend their whole PhD project on it. That data set is actually very small. It would definitely fit on a thumb drive. It's less than 1 terabyte of data.

Those are the different types of data we have access to for biology. They're disparate, they're different types, and they're different sizes. When we're thinking about what in that whole landscape might be of particular concern, I would say that it's currently an open question.

There are some who think that you could train an AI model on genetic sequence data alone and get a pretty functional model. I think that models like Evo 2, which have done this by training just on DNA sequence data, can perform well on protein-related tasks. They can span scales, and they can do genome generation. So there's optimism that you could do quite a bit with just sequence data.

But I think that there's another view in the community that genetic data is observational. What you really need to advance biology is some kind of data that gives you insight into causality. That's where you have things like perturbation data or knockout data sets, where you're systematically knocking out different genes of a virus and then looking at how that impacts its function, or you're systematically looking at how viral proteins bind to human proteins.

I'll broadly call that functional data. It gives you some insight into causality, and there's a view that incorporating that kind of data into training AI models is what's really needed to get you over to making functional biological constructs that are viable in the real world.

And so what we're proposing in terms of our data controls is that, as you rightly said, the vast majority of data should not be under control. I think that there's actually been a huge effort in biology to make data open access and more openly shared, because that advances research overall. I'm fully supportive of that.

The controls that we're proposing are really on functional data that gives you insight into important features of viruses that, frankly, the US government has recognized as relevant to whether a virus is pandemic-capable or not. Those features include transmissibility, virulence, or how deadly it is, and things like immune evasion. Can you modify a virus so that it gets around an existing vaccine or gets around your own immune system?

Those are the kinds of features the US government already tracks for wet-lab research. If you're proposing a wet-lab experiment that intends to enhance a pandemic virus—to make it more transmissible, more virulent, or to evade your vaccines—that's something the US government wants to know about and is going to ask you whether or not it's a good idea. Doing it in the computational domain is just extending that a little bit further. We're really proposing focusing narrowly on that kind of data.

The other thing that I'd add is that I completely agree that going out and scrubbing the internet of data that already exists is not going to be possible. It would be a Herculean effort and probably not worth the effort. What we're proposing is controls on data that's generated in the future: new data sets.

My thesis is that now that we know AI is quite promising for biology, there will be huge investment in creating new data sets. We're already seeing this with the US government's Genesis Mission, the OpenAI Foundation's commitment to spend billions on data sets, and the Chan Zuckerberg Biohub as well.

As really large-scale efforts get underway to generate not just observational but also causality-related information and perturbation data sets, that's where we're suggesting that, if this is done on pandemic pathogens, that data should probably not be shared for completely anonymous access.

Jassi Pannu

You should track who has access to it and have some controls around it.

Nathan Labenz

So, going back to the people doing wet-lab experiments on gain-of-function-style premises, I'd be interested in your take. I'd put my cards on the table: I think that's not a good idea. When I look at the timeline from the rise of the virus to the sequencing to the vaccine design in the COVID-19 case, I do take your point that there was some prior knowledge that accelerated things.

I'd also be interested to hear how much, if it weren't for that sort of COVID 1 knowledge, that timeline would have moved. But my zoomed-out and somewhat ignorant view is that it didn't take that long. Certainly, the clinical-trial part took a lot, and the distribution and manufacturing took a lot longer.

So, if the argument is that we want to do these experiments now because we'll be able to shorten the timeline in the future if something like this does happen, to be able to respond to it, I would say you're not really taking the bulk of the time out by quickening the pace to vaccine design. I'd hate to see you let it loose. So maybe we just shouldn't do that. Do you see it the same way?

And then I guess another question would be the obvious extension: should we apply the same reasoning to certain kinds of data generation in the first place? Is there a certain kind of data set that we should just say maybe we're better off not scaling? It would be interesting to know exactly what it is in these viral sequences that causes transmissibility or whatever, but we are creating something there that, in a sense, can—it's obviously a little bit more upstream—but could, in theory, escape in a similar way.

Yeah, let's start with your take on gain-of-function research in the wet lab, and then whether that same analysis applies to the data-generation side.

Jassi Pannu

Yeah, excellent. I think you can state your view with more confidence because I think it's a very reasonable view. There was a lot packed in there, so if I miss anything that you just asked, feel free to flag it to me.

To provide some background on what gain-of-function research is—perhaps, I don't know if you talked about this before on the podcast—I'll give an example from 2012. In 2012, there were 2 experiments done by 2 different research groups looking at avian influenza, which at the time was thought to have a 60% case-fatality rate, so it was highly lethal, but it was not human-to-human transmissible. There were cases of humans getting avian influenza from animals, but it was not resulting in a global pandemic because it didn't efficiently transmit between humans, which is a happy accident for us, frankly.

What the researchers were doing was conducting animal experiments in ferrets where they intentionally increased the transmissibility of that virus between ferrets. Ferrets are the known mammalian model; they are meant to represent human immunity. The hypothesis was that they were creating a human-to-human transmissible version of this highly lethal virus.

When these experiments were submitted for publication—I believe they were simultaneously submitted to both Science and Nature—the journals received the publications and alerted the U.S. government, wondering, frankly, what they should do with these results, because the manuscripts included the specific mutations that would be required to create that level of transmissibility. One of the groups found that it was only 5 mutations in the avian influenza virus that got you to human-to-human transmissibility.

That work is the kind of work that people refer to as gain-of-function research. The technical terms are dual-use research of concern, or enhanced potential pandemic pathogen research. At the time, it was recommended that those publications, or at least the details of the mutations, not see the light of day.

There were 2 concerns there. One is the concern that you raised: Humans are working with these kinds of pathogens in the lab. We know that there's human error, and what happens if someone is dealing with a pathogen that they just created to be more transmissible, gets infected, leaves the lab without knowing it, and triggers a global pandemic? That's a legitimate concern.

There have been instances in the past where samples of really concerning viruses have been found in settings like the CDC. In one past example, the CDC found vials of smallpox in a freezer. They didn't know they were there, and they were still viable. There are supposed to be only 2 places in the world that have active samples of smallpox, and these were not known to be there. So I think overall there's concern about lab accidents and human error, and that was one of the major concerns generally falling under the category of biosafety.

The second concern was that, aside from dealing physically with the pathogen, there was concern about the information related to the experiments—not only how the researchers did this and how someone else could replicate that same effort, but also the exact mutations that would be needed and whether that was information that should be in the public domain.

This relates very much to what you described: The horsepox synthesis protocol is in the public domain, and the smallpox sequence is in the public domain. Theoretically, someone could put those 2 things together and try to create smallpox, even if they weren't able to get access to the physical specimen themselves.

Overall, there's a lot of information about protocols for doing reverse genetics or other ways of rescuing live, infectious virus for pandemic pathogens. In the wet-lab field, this has actually been a huge debate for years and years with regard to what we do about this information and whether it should be controlled in some way.

So far, where policy and regulation have come down is that there's a focus on controlling the physical specimen, and there's a focus on preventing experimental work from increasing these concerning characteristics of pathogens. But it's considered infeasible to try to go back and scrub the internet of data that's already out there. We really have to try to figure out a mechanism for deciding this in advance, before it's already out there and we can't do anything to pull it back.

Nathan Labenz

But it's still not illegal to do this. Is that right? Can you do it? Do you need any special permission, or is it an ask-for-forgiveness-not-permission regime that we're on with this kind of gain-of-function research, even in the physical realm?

Jassi Pannu

Yeah, this is a really good point. In the case of the 2012 experiments, there was no law that those researchers were breaking. The mechanism that the U.S. government, at least, has used in the past has been regulation. Essentially, if you receive funding from the U.S. government, you therefore have to follow certain policies.

This policy around not doing research that enhances pandemic pathogens is one way that the U.S. government has tried to do this. There are some laws on the books for dealing with controlled pathogens through the Federal Select Agent Program. For example, if you want to handle anthrax samples, you have to be a registered lab that's tracked under this program. But, yeah, that's slightly separate.

The other thing to consider is work that, for example, seeks to go into bat caves. People are collecting samples of viruses where there's a suspicion that those viruses could be pandemic-capable. They sample them in those caves, bring the samples back to the lab, manipulate them, and try to characterize them. This is also something that the U.S. government and other governments used to spend money on.

But after the COVID-19 pandemic, the fallout from that, the lab-leak hypothesis, and the political dynamics around that, a lot of that work has been defunded. It's not explicitly illegal.

Nathan Labenz

I don't want to get too bogged down in this particular point, but is there a good reason for that? I do understand, of course, that we benefit tremendously from biomedical research broadly. I could imagine you might say, “Actually, the border is a little harder to define and can be a little fuzzier, so it's hard to legislate.”

But if this is something we're sleeping on without a really good reason, it might be time to start writing our representative. Is there a good reason that this isn't more controlled than it currently is?

Jassi Pannu

I would love to see it more controlled than it currently is. I think the real reason that it isn't is because governments are good at legislating things that happen often.

Jassi Pannu

What we're dealing with are pretty rare instances that certainly would lead to extremely high-consequence harms—global pandemics and things that we don't want to see. But they just don't happen very often. And so, the push for policymakers to treat this as a live issue, as something that needs to be legislated, comes and goes very quickly.

We already saw with the COVID-19 pandemic that there was a lot of concern. There still is debate as to what the origins of COVID-19 were, and we haven't resolved that question. The WHO director actually just a couple of days ago put out a statement saying that we still need to do work on resolving this question. But policymakers have moved on to more pressing issues, because that's just the nature of policymaking: they have to put out fires today.

I think that, in reality, we need both national and international rules. Right now, the WHO has rules saying that only Russia's Vector Institute and the U.S. CDC are allowed to have access to smallpox. That's a great initiative, but it did not prevent a researcher from unilaterally publishing the step-by-step protocol for how to synthesize horsepox, the close relative to smallpox.

What that highlights is that, as synthetic biology, virology, and biomedical capabilities advance, we need a better way to make sure our regulations keep up. That's a concerning topic. But I think so far, the reason that this isn't a live issue, or why we're not thinking about it day to day, is because it still requires a lot of expertise to synthesize any of these pathogens from scratch. It really is something that you need to have a lot of background in, but this is where the concerns related to AI come up.

In terms of the different types of AI models, whether or not they provide uplift, and what kinds of biological models could be used to do this, this is such an evolving and open question that people are trying to figure out. I think the hard part is that it's moving quite quickly, and so it's hard to see how policymakers can keep up. But we're working on it.

Nathan Labenz

How much actual wet-lab gain-of-function research do you think is going on today? Has it been dramatically curtailed by these sorts of strings attached to funding and general awareness in the community that it's maybe not a good idea, or do you think there's still a lot going on?

I guess another reason that there might not be a law is, well, everybody quit doing it because they realized that it's a bad idea. Who needs to make a law against nobody doing it? But is that the case, Jassi? Do we have any way of really knowing how much is going on?

Jassi Pannu

With regard to wet-lab gain-of-function research, I would first want to say that the kinds of research we need for future vaccine design, like determining the spike protein sequence, are important. That's how we advance our ability to create vaccines for new pathogens, and that kind of work doesn't require gain-of-function research.

Gain-of-function research, the kind that we're talking about, is very narrowly scoped, and it does not require enhancing the transmission of a pathogen, making it more virulent, or making it escape the immune system. Those kinds of experiments are really not needed for the vast majority of the advancements we would want in biomedicine.

With that in mind, I would say that over the past few years, since the COVID-19 pandemic, a lot of this work has been defunded and reduced by U.S. government funding mechanisms. Our blind spot is the work that's happening in private labs. We don't actually have any legal mechanism for going into a private lab and determining whether they're doing certain kinds of pathogen research, other than whether they have registered under the Federal Select Agent Program.

There have been instances of laboratories in California, most recently, where they're handling certain types of pathogens that they really, frankly, shouldn't be, and they don't have the containment protocols for. I'll pause there. Overall, I think we're in a better spot than we were. I think people have recognized the downsides of this research and the risks, and certainly governments are paying a lot more attention to it.

Nathan Labenz

It's funny—it echoes, in a way, the reduction in bad behavior that we usually see from one generation of large language models to the next, where it's like, "We recognize that this was a problem, did some stuff to try to curtail it, and reduced it by 90%. Great news." The other 10% is out there for future work to contend with. Dizzying in the gain-of-function case.

Okay, you mentioned—let's go with your segue. There are different kinds of models, obviously, in the AI space that people might be concerned with. First of all, I was thinking ahead to this conversation and I was like, "Well, of course we've got the large language models, which output text and can reason about things." They might just know facts that could be problematic. They can use tools.

I just did a conversation with Jeffrey Irving, who's the chief scientist at the U.K. AISI, and I had not realized before that frontier models these days are getting quite good at troubleshooting lab experiments from cellphone pictures. They've now gotten to the point where you can just snap a picture of what you're working on, tell the AI that it's not working, and it will coach you through how to get it working. That's the know-how, the reasoning, and the procedural stuff.

Then we've got models that, as you alluded to with things like AlphaFold and that whole genre, are very good at making very specific predictions. What shape is this going to be? What's going to bind to what? So on and so forth.

Then there's the middle-ground hybrid: things like Evo and Evo 2, where they're trained kind of like large language models on these vast datasets. In many cases, they are literal next-token predictors, albeit in the DNA or protein-sequence domain. I probably have the least intuition for those.

I guess you could complicate that taxonomy for me if you want, but then maybe just go through and tell me how concerned I should be in a world where there's no data controls and the models have actually learned on everything we have. How concerned should I be about those different kinds of models, or possibly how they might be stitched together?

Jassi Pannu

Got it. I think, in general, I like this taxonomy. I was not as creative as you in terms of coming up with the different capabilities that these groups have, but I think about them largely in the same way.

LLMs are trained on lots of biological information, from textbooks to scientific papers. They can give that information to someone who doesn't already have it and isn't already a biology expert. That's usually called uplift, and broadly, that's a really great thing. LLMs teaching someone new biology, helping students learn, and really providing a quite useful capability.

There are a subset of instances—for example, "How do I illegally obtain an automatic weapon?" or "How do I illegally obtain smallpox samples?"—where that kind of information is clearly not something that should be provided to the general public. That's where frontier labs are working to apply classifiers and refusals to make sure that kind of knowledge is not widely shared.

Then, when you think about tools that can be used for biology tasks, I think people often call these biodesign tools. These are specialized models that are trained on biological data and are used to do specific things. I think of this as a model that gives someone a new capability. It's not about knowledge; it's about what you can do.

These types of models really require someone to already be an expert. You have to already be a computational biologist working with models to really leverage these kinds of capabilities. But the interesting thing is that they can allow those researchers to do something that was just not possible before.

Before the world of protein design, AlphaFold, and structure prediction, it was just not possible to take a protein sequence and then, immediately through computational methods, play with new designs or try to infer its structure and function. Those are really interesting capabilities. Again, the risks here are less about providing uplift to someone who didn't already know how to do that and more about giving experts the ability to do new things with biology.

And then this third category—I agree, it's kind of a middle ground—is models like Evo 2 and ESM-3. These are what I would call biology foundation models. They're trying to be general-purpose in the same way that LLMs are, and they're also often trained on different kinds of data.

AlphaFold, obviously, is trained on the PDB, but it also has MSA data. ESM has different types of data that it's trained on as well. These models are often much larger than biodesign tools, which can be small enough for an individual research group to train and host locally.

Biology foundation models require more data and more compute. They're more expensive for groups to develop, and so you often see these models developed by larger organizations. AlphaFold, obviously, is part of Google DeepMind, rather than an independent academic lab, and ESM is from EvolutionaryScale.

These types of models are trying to infer the fundamental laws of biology. They're trying to understand how biomolecular components interact across different scales and really elicit the underlying laws that govern protein function and protein structure and, in the case of Evo 2, operate across different scales.

Evo 2 is a model that's trained on just DNA sequence, and what they were able to show is that it can actually help with tasks across sequence, protein, and genetic regulatory circuits.

Jassi Pannu

So that's operating at different scales in biology, and it's inferring laws that transfer between those different scales. Overall, there's lots of interesting work being done in all of these. I think the risk considerations can be separated across the different types of models, but what I would argue is that that will collapse over time because what organizations are working toward are integrated workflows.

Ultimately, the dream is to be able to have your AI agent design your experiments. It will be connected to your autonomous robotics, which will conduct those experiments. The data from those experiments will be collected and then fed back into your biology foundation model, which will then be used by your agent to design future experiments, and so on, in a loop.

These kinds of iterative feedback loops, where you're getting data from the real world, I think, are where people are most hopeful about how this process can advance biology. Right now, we have these huge data sets that are messy and collected in an observational way, but what these feedback loops would allow you to do is systematically perturb systems, systematically try to assess causality, and then use that information to further and further improve your in silico models.

The dream would be, ultimately, that you get to a point where your in silico models perform so well that they start to replace some of the wet-lab biology that you've done. You get better and better predictions of what different drugs, for example, will do in cellular models and animal models. Ultimately, the dream would be that you get better predictions for clinical trials, so you have to do fewer of those and they have higher yields.

Nathan Labenz

So, what do we want to take from that? I'm obsessed, by the way, with that idea of both the closing of that loop and a little hobbyhorse of mine that I'd be interested in your take on: the sort of latent-space integration of these different modalities.

Obviously, we've seen this with image and text, in the sense that I can now go to a Nano Banana model or whatever and give it an image and also some text instruction, and it is understanding those in a joint way to a degree that wouldn't be possible if it were just prompting an external image model with text. I could have a language model that uses an image generator as a tool—we've seen that—but this sort of deeper integration gives you much higher fidelity to the original, and you can do text and image prompting in a very natural, integrated, cohesive way.

I've been wondering, assuming we're going to see it, on what timescale you think we see that kind of thing in the natural sciences, and specifically in biology as well. You might say, "Not, 'Okay, hey, language model, you can call this protein model as a tool,' but rather, you are both, and what I want is: working from this protein as an example, give me another protein that could do the same thing," and just have that all be understood in the same set of weights. What do you think the outlook is for that sort of system?

Jassi Pannu

I agree that that is where the field wants to go, and it would be really useful to be able to develop that. I feel like there are a couple of bottlenecks along the way.

One is the data-generation piece, which is still something that requires scaling in the physical world and doing experiments in the physical world. That itself will be bottlenecked by advances in robotics. If we were to suddenly see robotics speed up and were able to do a lot more laboratory work autonomously using robotics, then the overall picture in terms of data generation also would advance.

So I guess I'm hedging. I'm not really giving you a timeline—perhaps in the next 5 to 15 years. These are the kinds of advancements we would expect.

Nathan Labenz

Is that data—I mean, you said obviously we have huge amounts of just raw sequence data—but we are short on the causal graph, if you will, of "I did this, and this resulted." Presumably, a lot of that is maybe locked up at pharma companies that have done some of this stuff, or even just in the clinical data.

Do you think that if we had full access to all the data that exists, regardless of who owns it, how it was created, and where it's sequestered, would we have enough data already for that kind of thing to happen? Would we essentially be recreating, for IP reasons or privacy or whatever reasons, something we essentially already do have as a society? Or would you say, no, not really? Is the clinical data too messy, and maybe pharma doesn't have it? I don't know.

Jassi Pannu

Yeah, pharma definitely has lots of highly valuable data that they do not share in the public domain, for reasons that are obvious, and pharma is actually using its internal data sets to develop proprietary models in-house. They're certainly trying to do that.

If we were to suddenly wave a magic wand and say, "The government says everyone has to play nice and share their data sets," how far would we get? I think we would get a little bit further, but I think that these feedback loops and a new way of generating data are fundamentally a different approach.

You can think of the existing way of approaching biology as a bit artisanal and a bit observational, and that results in data sets being messy, having a lot of bias, and being hard to work with. What we really need to do is shift toward a much more systematic approach, where we are generating data that is comprehensive and systematically probing every single aspect.

That's where you really need the robotic aspect to scale that data and replicate it carefully, rather than having multiple different humans trying to do the protocol. There's always differences between them when they're collecting data. Just transitioning to a systematic approach that's enabled by robotics, I think, is not something that you would get just by enabling data sharing across private companies.

Nathan Labenz

Got you. Okay. Well, let's pop out of that rabbit hole and come back to the main topic. So we've got language models that can tell people things they maybe shouldn't know. They can increasingly use all kinds of tools, including design tools.

It's not clear to me at this point how well they could use something like Evo 2, but when we think about those models and their capabilities, what sort of capabilities should not be created in the first place? Is it about—I guess it's probably multiple things—but maybe I'll just leave it there: What capabilities do you think those different kinds of models should not have in order to reduce the risk to society broadly?

Jassi Pannu

I think that the fundamental challenge of biology is that a lot of these capabilities would be useful on the defensive side, but it's when they're used offensively that they pose concerns. So it's the question of what capabilities we want, but it's also the question of what capabilities we should provide access to, how broadly, and when.

When we're thinking about advancing the future of AI for biology, I think the way I like to think about it is that we should really try to step on the gas for things that are clearly good and clearly do not have a lot of risks. Things in that bucket, to me, are virtual cell models, ways of advancing clinical trials, or ways of making sure we can do better countermeasure manufacturing and distribution. There are lots of things that we could do that are clearly beneficial.

Then there is a bucket of things that, if that capability were broadly accessible right now, would be quite destabilizing. I think this is just a hypothetical example, not trying to say that this is actually the current state of capabilities, but let's say there were a breakthrough where suddenly it's very easy to use an autonomous robotic system that's quite cheap to build or get access to, and that robotic system could very quickly synthesize a pathogen.

That's obviously a futuristic scenario, but let's say that were possible. Then that's something that probably we wouldn't want anyone to be able to buy off the shelf. We'd want to know who has access to that device and what they're using it for.

Other things that are trending in the more concerning capability bucket would be things like viral design. Even there, there are considerations: We know that gene therapy based on viruses is actually an advance that we would love to see, or there are other purposes for viral design, like designing bacteriophages, which are viruses that only infect bacteria—they don't infect humans.

The challenge is that when it comes to artificial intelligence, a lot of the approaches are general-purpose. So if it becomes quite easy to have an AI model that can design a bacteriophage, then the question is, well, it seems quite easy to repurpose that for pandemic human pathogens. How many people do we actually want to have access to that kind of capability?

It's probably a subset of legitimate researchers who are using that. It's not something that you would want widely accessible on the internet, especially in a world where we don't have easily accessible countermeasures. It really becomes an offense-dominant capability where the design and acquisition of a pathogen become easy, and it's facilitated by AI.

That capability exists in the digital world—it is being uplifted in the digital world—but our countermeasures to a pandemic remain very physically world-bottlenecked. That's a world where it just becomes very offense-dominant.

Nathan Labenz

Even with things like a whole-cell model, would I be right to worry that one of the things you would want to do with a whole-cell model is throw stuff at it and see what happens to the cell? If you had that, all kinds of great things might be possible, but then presumably you could also start to throw in your virus of choice and start evolving that in whatever way you want to, potentially just brute-force throwing all these little permutations of a virus at the whole-cell model. It strikes me that these things are vulnerable to a brute-force attack.

If they're going to be good, they're going to be vulnerable to that sort of brute-force attack. Is that right, or is there any way around that conclusion?

Jassi Pannu

Yeah, you are embodying the debates that people constantly have in the biosecurity community. I think what you're saying is correct. There are ways to envision every kind of biomedical advance, especially in the AI domain, as being used for harm. Because of that, you have to think about how direct the harm pathway is and how consequential the ultimate harm would be, and you have to try to draw a line somewhere.

If you compare, for example, a generative language model that had no data filtering, had no data exclusion, was highly performant on viral genome design, and could do that for human pandemic pathogens, the pathway for harm is quite direct. Someone with very limited biology knowledge could use that model to generate thousands of potential designs, sequence them in the wet lab, see which one is the optimal candidate, and then use that candidate. There is still a lot of work going into that, but it's a direct pathway.

What you're describing would require plugging in multiple different AI models, generating candidates with one model, then running them through a different predictive model, seeing the consequences, and trying to figure out which viral candidate would cause, for example, a systemic inflammatory response, or would target certain organ cells. I think there is a pathway to harm there; it's just that when you game it out, there are more steps involved and more expertise required.

The ultimate question you ask yourself is, what kind of actor would choose that pathway over an existing weapons pathway? If it really requires high-level expertise that only a nation-state has access to, is that nation-state really choosing a biological weapon, or are they more likely to choose something else that a nation-state would have access to? Those are the kinds of questions that security professionals try to game out.

Nathan Labenz

Okay, maybe let's get to the proposed solution. I've been coming at this from a lot of different angles. You've got a whole taxonomy of five levels of biological data. Obviously, this is inspired by, or at least pattern-matched to, the levels of security around bio facilities. Maybe just take us through zero to five. What are the kinds of data that fall into these different levels? What would the access look like? What would the precautions look like? Paint a picture of the world that you envision.

Jassi Pannu

Great. I'll try to paint a somewhat visual picture. For those imagining this, it's a five-tiered system that goes from level zero to four. This is modeled on what some of you may be familiar with: biosafety levels, or BSL levels. These are the famous safety levels that biological laboratories use to determine whether I have to wear a spacesuit when I go into the lab or can just use a fume hood to deal with my samples. The system determines the different containment approaches that are required.

Actually, the BSL system was the basis for a lot of the frontier safety policies that different frontier AI labs have, in terms of the idea of having a system and mechanisms for controlling it, roughly on four tiers. What we're proposing here is applying this not to the model and not to the physical pathogen, but rather applying it to data.

The reason we're proposing this is because the entire biomedical research ecosystem, when it comes to AI, is built on open-source models. Academics build open-source models, they share those models openly, and other researchers manipulate and change those models. There are a lot of benefits to that fully open-source ecosystem, and those benefits are what have prevented security approaches from being applied to models.

It seems like that's going to be a pretty intractable approach, and we were looking for a different approach that could be applied to ensure that you could still preserve this open-source model ecosystem, but not distribute capabilities that would be particularly concerning, like viral design. That's how we ultimately settled on biological data.

Biological data, especially the kinds that I described—the more functional data—is expensive to produce, requires a wet lab, and requires expertise to produce. That's why it's a potentially useful choke point. The tiering system that we described would preserve the vast majority of biological data as fully open access, and that's what we're calling BDL0, where most data would be available to researchers.

As you go from levels 1, 2, and 3 up to level 4, you have increasing levels of control based on how potentially concerning the data is. The way it's essentially broken up is that BDL1 is data that would allow you to infer viral patterns. It's pretty basic security, just requiring an account and understanding who the person is and whether they're a legitimate researcher.

As you go up, you're getting more focused on properties of pandemic pathogens that would directly lead to harm. These are the properties that I mentioned to you, which governments globally already pay attention to for wet-lab research. This includes making a pathogen more transmissible, making the host range larger, allowing it to infect more animals, or allowing it to move from infecting just an animal species to also infecting humans. It also includes manipulating the pathogen so that it evades the immune system.

These are the kinds of properties where, if your data set has data directly linking those properties to pathogens, it requires things like use approval. You would go to the repository and say, “I intend to do this kind of model development based on this data for this purpose.” If you're a legitimate researcher who has a good purpose for doing this, then you will get approval, versus just having this data openly accessible for anonymous access. Maybe I'll pause there.

Nathan Labenz

How would you describe the magnitudes of those? Is the outer ring, BDL0, like 99% of the actual raw data? How small does it get when you get up to the uppermost levels?

Jassi Pannu

Yeah, I would say 99% is a pretty good guess. We don't actually have numbers to base this on because there isn't a comprehensive tracking system for these kinds of data sets. But my guess, based on the research that we've done, would be that the highest security tier, the BDL4 tier, is a very small subset of all data.

It would be a very small number of specialized virology labs, for example, that would be affected. Frankly, those labs are probably already limiting access to the data sets that they produce in some way. It's just not a formalized system.

BDL-0 would cover the vast majority of the data we're talking about—petabytes and petabytes of data—with the vast majority being uncontrolled. The controls we're proposing are on a very narrow slice of data, particularly getting up to BDL-3 and BDL-4, with perhaps dozens or fewer laboratories affected.

I am making that statement based on my knowledge of the field, but there probably needs to be a more comprehensive effort to try to figure out who exactly is generating this kind of data. The visibility bottleneck we currently have is what's happening in the private ecosystem. There's pretty good visibility in terms of government-funded work, and less so on the private side.

Nathan Labenz

In terms of the impact that this would have, one of the things I thought was really interesting in reading the recent paper calling for these kinds of controls was the report—which I hadn't realized—that there has been data holdout work done on a couple of the leading models, ESM-3 and Evo 2, specifically. Could you talk us through a little bit of what that has looked like?

I did put in a good word for me with Alex Rivers, please, to get him on the show. I've tried, but we did do one with Brian Hie on Evo. I have a general sense of what that looks like, but I didn't get into what was held out, how much of the overall data it represented, or how that affected performance in the areas of concern.

I assume one thing people would be really worried about here is, “I don't want to have a dumb model in general.” If I slice out this data, what does that mean in terms of what it can't do that I want it not to be able to do? But also, are there things that it can't do that I would wish it still could do? What costs are we paying for the benefits?

Jassi Pannu

Yeah, this is an important question. I think we all have an intuitive sense that AI model capabilities are based on the data an AI model is trained on. It intuitively makes sense to us, but the degree to which that is true is an empirical question.

Especially in biology, there's a reasonable reason to question whether, if I were to remove a very small subset of data, my model could just interpolate around that gap. If your model has really internalized a fundamental understanding of the laws that govern biology, does it really matter if you start segmenting out different small pieces of data?

This was an empirical question until, as you said, some of the leading biological AI model developers actually went about doing this. I'll just describe the two examples you mentioned: ESM-3, which is a generative protein design model made by EvolutionaryScale, and Evo 2, which is a generative DNA language model made by Brian Hie.

Both of those groups had decided that they wanted to share their models, but they didn't want to disseminate the capability for others to use their models for viral design.

Jassi Pannu

So the way they went about limiting that capability was by limiting what went into the training data. In the case of Evo 2, for example, given that I was involved in that work, we decided to remove the sequences related to viruses that could infect humans and viruses that could infect eukaryotic organisms. But there was still some information related to other types of viruses—for example, those that could infect bacteria—included in the training data.

After doing the big pretraining run, we then did some evaluations where we checked to see: Is the model limited in its capabilities on certain tasks? A common task that people use in this case is looking at how well the model can do certain viral-protein-related tasks and how well the model can generate sequences that correspond to functional viruses. Those were all things that the team checked, and they showed that the model's capabilities were significantly less than they were in other domains. Those evaluation results were all published as part of the Evo 2 manuscript.

The interesting thing about the work that the EvolutionaryScale team did for ESM-3 was that they actually had both versions of the model. This wasn't something that the Evo 2 team did, but the EvolutionaryScale team had both the trained model that had been trained on all the data they had chosen to include, as well as the data-filtered model. They were able to show a delta in performance on the same tasks with regard to, for example, viral protein function prediction.

So, that's an inkling of some of the empirical work that has been done and could be done in the future to try to suss out how much it matters when you remove these kinds of data and what particular kind of data matters. These are all questions that could probably be explored a lot more.

Nathan Labenz

Could you give a sense of the order of magnitude of capability reduction? Are we talking about it just not being able to do that stuff at all anymore, or is it in the uncanny valley somewhere? I don't have an intuition, honestly, for what I should expect.

Jassi Pannu

Yeah, I think so. For example, in the case of Evo 2, it was much more along the lines of the model's function being effectively random rather than just being reduced by a small amount.

Nathan Labenz

Cool. That's great. Great news. I love it when something works.

Nathan Labenz

Okay, maybe let's zoom out and take stock of all this stuff. We've got a ton of new data coming online. We want to facilitate sharing, and we want to facilitate all this discovery, but we've got to have some of these different classes of data controlled so that models aren't trained and disseminated on them.

Presumably, this also means that the models themselves that are trained by the ESM team and the Evo team also need to be controlled. I guess we're not—we shouldn't be expecting government control. If we don't have government controls on actual wet-lab gain-of-function research, it doesn't sound like we're going to get government control on this sort of thing.

So, what are we doing? We're campaigning and trying to build private agreement and consensus? Is that the play, and how's it going?

Jassi Pannu

Yeah, great questions. Just to take a step back in terms of what happens with the model and whether the models should be controlled, I think in the case of Evo 2 and ESM-3, they had effectively neutered the concerning capability from their models. That's why they felt more confident in disseminating those models. Ultimately, it made sense to be able to share those models openly because they had worked to reduce the concerning capability.

What I was more so referring to is that, if you do end up implementing data controls as we propose and you get access to BDL-3 or BDL-4 data for purposes of training a model, then what you wouldn't want to happen is that model being shared publicly, because effectively then your mitigation didn't really do anything. We do make sure to mention that any AI model trained on that kind of secure data should also be shared in a secure way if it needs to be shared.

In terms of what should be done next, I think the interesting thing about data controls is that there's a lot of policy precedent. We already do this in other domains. We do it for privacy-related data, and we do it for human genomics data. So, there's reasonable precedent for extending that same approach to data controls.

My collaborators and colleagues who worked on this are optimistic that perhaps this could get picked up by policymakers, but of course it's a new concept, and so we'll keep plugging away at it.

More broadly, in terms of what will happen with wet-lab gain-of-function work and model capabilities, I think that both the Republicans and the Democrats have decided that this is a bipartisan issue. Under both the prior Biden administration and the Trump administration, some work on advancing regulation of wet-lab gain-of-function work went ahead. I was really glad to see that progress.

I also think that both the US AISI and UK AISI are doing really great work with regard to biosecurity and biological model capabilities. Speaking given that I'm more familiar with their recent work, the US AISI has put out RFIs—requests for information—from the scientific community to better understand the issue.

We know that the developers are doing some of these data-filtering steps and trying to limit model capabilities. Maybe we should have a more systematic approach to figuring out what all the capabilities are that we should be concerned about. How can we test for those capabilities? How do we effectively mitigate them?

I would just love to see AISI be resourced and staffed to advance that line of work, because I think a lot of developers are trying to do it themselves in an ad hoc way, but they don't have the same security access and intelligence access that something like the US AISI has.

Nathan Labenz

Do you envision a sort of centralized—I’m still a little bit fuzzy on the workflow of something like this. If I, for whatever reason, am out here generating some sensitive data that maps viral sequence onto viral capability, and now I'm like, “Okay, I've got this data. I want to be a good citizen. What do I do?”

Are you envisioning a scenario where I take it to a central data bank, give it to them, scrub it from my computers, and go on with my life? Because we can't have everybody doing these sorts of security levels for themselves, right? So there would have to be some—maybe not totally centralized, but at least no more than a countable, relatively small number of organizations or entities that would have custody of the data.

Is that kind of the idea—that people would feed it in and then those organizations would be expected to actually delete it from their own servers? How do you— is that realistic to expect people would do that? How do you expect that to actually go?

Jassi Pannu

Yeah, this is all about how you operationalize the system that we're proposing. The way that this has been done in the past is through things called trusted research environments, or TREs, and OpenSAFELY, the system that I mentioned at the beginning, is one of these.

The idea would be that you have a secure environment that hosts the data and also enables researchers to bring their code to the data and answer questions. Honestly, in the age of AI, the ideal would be that you have compute resources available to that, although we'll see if that would be possible.

I think the way that you would set this up is that we considered different approaches. One approach could be that the government sets this up. After we looked into this, it seemed less ideal. A lot of researchers have been unhappy with some of the trusted research environments that the government set up, and I think there are private actors that could probably do a better job of setting up really savvy trusted research environments that work well for researchers. We have some examples of this that I've mentioned and that are included in the paper.

What we suggest is that institutions—not individual researchers, but institutions like universities or larger research collaborations—set up a trusted research environment if they want to do this kind of data generation. They would be given standards that the government would set, but they would be in charge of actually building and maintaining the environment, just because they're probably going to be better at doing that.

The interesting thing about a trusted research environment is that you can actually imagine this also being a boon for researchers. When you're trying to do AI research, what you want is a centralized, integrated platform where all the data sits and you have access to it all at once. It's much less useful to have individual data sets housed with individual researchers.

So, if we really wanted to step on the gas of advancing countermeasures research for pandemic pathogens, maybe this is something we would want anyway. We would want an integrated system that hosts all the data and really makes it easy for our researchers to interact with.

What we're proposing to layer on top of that is that you do have some security that goes along with that system. Not only are you hopefully getting the benefits of the integrated platform, but then you're also getting the security controls, given the sensitive nature of the data.

Nathan Labenz

Got you. When it comes to monitors, I know that you had alluded to some rare points of bipartisan agreement. One of those, I understand, is the insistence—or I don't know if it's fully a requirement, but I think it's verging on a requirement—that DNA synthesis companies apply certain classifiers or whatever to try to detect if somebody is trying to get a harmful sequence synthesized through their company.

There’s also things like wastewater monitoring. I have a couple of questions on this. One is, is there an equivalent of wastewater monitoring for data? Would it make sense—I don’t know if the data of concern is the kind of thing that could be identified if it’s just hanging out on somebody’s lab website.

Potentially, people might not even fully realize if they’ve generated problematic data. So, is there any system you could imagine to go around and identify data that’s out there and then classify it as possibly harmful, or is that such a difficult classification problem that it would be doomed from the start?

Jassi Pannu

It’s a really interesting concept. I have not thought about it too much, to be honest, but I feel like you would probably need—there are 2 approaches to figuring out what the data landscape is for this kind of data. One is a more active, contribution-based approach, where the government says, “If you think you’re creating this data, you have to actively contribute it to these repositories.”

What you’re describing is a passive approach, where no one has to take any individual action. Some system is flagging the data, and then it perhaps even automatically collects it and scrubs it from the original source. That would be the dream. I’m not aware of any system like that, and I think it would probably take a lot of work to do. I think that the default approach has been the first one, which is voluntary, or active contribution by the researchers who are generating the data. But what you’re describing would be cool, and it would be the analogy to wastewater surveillance, essentially.

Nathan Labenz

When it comes to the quality of monitors in general, a big thing I always think about in the rest of the world as it pertains to AI is that we’re in this sort of weird in-between phase where things are coming online, but the world hasn’t really reacted that much yet. For example, AI agents are coming online. There are a lot of concerns around things like prompt injection and whatever, but the world hasn’t really become a very dangerous place for an AI agent so far, right? Not that many people have set up honey traps or prompt-injection attacks to try to throw my OpenClaw off its path and talk it into doing something that they want it to do. I assume that’s going to happen a lot more, and there’s going to be some arms race between techniques to prevent my OpenClaw from falling for it and ever more sophisticated jailbreaks, whatever.

Is there—I assume there’s got to be something analogous in the biological domain where, for starters, these classifiers that the DNA synthesis companies are running—I’m guessing that nobody really has tried to evade them yet. I wonder if you see that kind of dynamic developing on the horizon. Nicholas Carlini, who was a past guest, has said that the attacker usually gets to act last. The defenses are set up, and then the attacker always has that advantage of knowing what they’re up against. Maybe not always, but often.

Do you see any of those kinds of dynamics now, or do you worry about them in the future? And I guess, if you extrapolate them out, do you see us getting to a place where we have a high level of confidence that we’re in a defense-dominant world and we’re going to be able to keep all this stuff under control, or is that itself still a very open question for you?

Jassi Pannu

Yeah, I have some ideas, but maybe let’s start with gene synthesis screening and the different—there were a few different attack surfaces. I’ll just describe them that way. The first one is gene synthesis providers.

So, what you’re describing is the system that is currently voluntary, where gene synthesis providers—companies that make pieces of DNA and sell them to researchers—have implemented systems where they check to make sure that someone didn’t just order pieces of smallpox or pieces of Ebola. They have a twofold mechanism.

The first is an automated mechanism that is essentially sequence matching. It’s looking at whether the order matches a sequence from Ebola or matches a sequence from smallpox. If the automated system thinks that there is some degree of concern there, it gets escalated to a human researcher, a human expert, who then looks at it and determines whether it’s something of concern or not.

What’s also happening in parallel is some degree of KYC—know your customer. Who did the order come from? Is it a researcher? Does this researcher work with these kinds of pathogens all the time? Have we spoken to them before, and has this issue come up before? Those are all the kinds of questions that are being addressed as part of gene synthesis screening.

Gene synthesis screening is something that 80% of companies already implement, and over the past few years it’s become much more cost-effective to do. Initially, it was a bit of a cost barrier. They really did a lot of work to make this a cost-effective system, because you can imagine that if this is something being run on every single order that’s coming in, it has to be cheap enough to do. Otherwise, it’s really burdensome for these companies.

Currently, it’s a voluntary system. Like I said, the vast majority of companies already do it, but the concern is that if I’m a bad actor, then a voluntary system that 80% of companies use isn’t really going to stop me from obtaining the sequences I want to get my hands on, because I’ll simply go to the companies that don’t do the screening.

What’s now being advanced is the idea of making this a mandatory rule that all companies have to follow. That would really limit the access that someone has to physical specimens that they might try to use to turn into an infectious pathogen.

But I guess, given where the research is going and what capabilities people want to achieve, the dream for the future of biology is that I, as a biologist, no longer even have to step foot in the lab. I have my autonomous cloud lab that I can fully control. I’m sitting at home using Claude Code to help me, and I can just program some designs to be done on certain pathogens.

To be completely frank, we’re nowhere close to this world. It’s still going to take a lot of work. Right now, the cloud labs that you hear about still require a lot of human input. It might not be specialized biologist input, but there is still a person picking up samples from one bench and moving them to the other, and that’s a real bottleneck.

But in a world where you do have a fully remote, highly sophisticated cloud lab, then you can imagine that if there are agents—if I’m a bad actor and I have an army of 1,000 agents that are just trying to hack their way into this cloud lab—you want to make sure that if your cloud lab has sophisticated capabilities related to pathogen creation and design, you have some cybersecurity around that. So, that’s a much more future-oriented thing, but something you could imagine becoming applicable later.

Right now, there’s a fundamental information-infrastructure layer that we’re missing. So, both for Palantir Labs and for gene synthesis screening, if I, as a bad actor, try to obtain sequences of Ebola from one company, small pieces from one company, and a few pieces from a different company, and I split up my order across different companies, there isn’t some kind of system where all those companies are easily communicating that information to each other and checking those orders against each other in real time.

This is something that’s bottlenecked by information sharing between companies. That’s also a similar concern to why the Frontier Model Forum was created: you want private companies to be sharing security-related information. You need a legal infrastructure for that. Who would house that? Would that be the FBI? Who’s facilitating this? These are all kinds of policy questions that need to be addressed.

Perhaps we can end on a more optimistic note. I’m happy to give you my vision of what a defense-dominant world looks like and see what you think of it. I think that, overall, I break this up into 4 broad buckets of interventions.

Well, perhaps to take a step back: I think people often think about what the theory of victory here is. How do we have a single, unified strategy, like our strategy for nuclear deterrence? We have a single, unified theory of victory for nuclear deterrence that has served us well for decades. How do we get there for biology?

I guess, after having thought about this for some time, it just feels like biology is very different. It’s a distributed technology. It’s dual-use. You want to give a lot of people access to it. You’re trying to limit it to a subset. There are all these aspects that really make a unified theory of victory seem much harder to get to.

So, I think the most successful approach is likely to be defense-in-depth: a layered approach, with multiple different defensive strategies applied at once. The 4 buckets that I divide this up into are deter—oh, sorry. I should start with delay, deter, detect, and defend.

Delay is essentially limiting access to concerning capabilities. Gene synthesis screening would fall into that bucket. You’re delaying the dissemination of that capability to get access to DNA fragments, for example. You could also imagine what we’re describing for our data controls as part of delay.

Then there’s deterrence. Deterrence is figuring out how you can punish someone for using a biological weapon. We do live in a world where there’s an international treaty against biological weapons. I’m glad that we live in a world that has that treaty versus not, even if, overall, that treaty is on the weaker side in terms of the actual mechanisms we have to ensure people are complying with it.

But that’s one thing. I think where deterrence breaks down is if your actor is not rational and doesn’t respond to typical punishment mechanisms. So that’s a challenge. But then, in terms of detection, this is also something we talked about: How do we have a distributed passive surveillance system, like our radar system for ICBMs, that will just detect when there’s a new pathogen without us having to go out there and look for it, perhaps even for pathogens where there are no symptoms?

When HIV was spreading early on, it would have been amazing to have known that much earlier. And that’s particularly important for pathogens that take a long time for patients to develop symptoms. So, having some kind of global surveillance system—people often refer to this as bio-radar or bio-threat radar—it’s not something we currently have.

And then the last pillar is defenses. Really, what are our defenses once there is a pathogen online already? I think that people often think of defenses as things like vaccines and countermeasures, but I would encourage folks to be much broader in what they envision defenses to be.

Because I would argue that I’m sitting in my home right now, and I actually have defenses all around me. I’m drinking water that’s been centrally filtered. I know there’s no cholera, no pathogens in that water. I have screens on my windows. Mosquitoes can’t get through them. I know that I’m not going to get malaria if there was malaria outside.

And so there’s already a lot of public health defenses built into our environment, but we don’t have that for airborne transmission. One thing that’s being explored by organizations—one that comes to mind is Blueprint for Biosecurity—is built-environment defenses to sterilize the air. This is using approaches like far-UVC and other approaches like glycol vapors.

So, could you passively sterilize the air so you don’t even have to detect the pathogen? You don’t even need a vaccine. You just always know that you have passive protection around you. So hopefully, I would say that’s a pretty comprehensive approach. If we manage to do all of that, it would take a lot of work and investment to get there, but it would make us a lot safer.

Nathan Labenz

…above his hospital bed. Every time he’s been in the hospital, we’ve hopefully taken some of the risk off the table of him getting any kind of infection while he’s going through all this. That’s a great vision. I hope we implement it.

You are doing God’s work spending your precious time and energy on this. Anything else we didn’t touch on, or any other calls to action or ways that people can help you, that you would want to leave people with before we break?

Jassi Pannu

I think we covered everything, and I really appreciate your interest. It sounds like you’re one step ahead of everyone else in terms of already getting your kid outfitted with all the defenses he needs. It’s great to chat with you.

Jassi Pannu, thank you for being part of The Cognitive Revolution.