[BidClub_]
The Cognitive Revolution · · 103 分钟

解锁细胞的秘密:与 Squidiff 和 CORAL 的 Siyu He 探讨扩散、解卷积与发现

Nathan LabenzSiyu He

YouTube
TL;DR
  • Squidiff 将一个细胞包含 30,000–60,000 个基因的表达状态转化为软件对象,让研究人员在投入稀缺的湿实验资源前先进行扰动。 该系统联合训练语义编码器和 DDIM 扩散模型;Siyu He 明确纠正了原始描述——「其实并不是 VAE」(it’s not actually the VAE)——并表示,一个约 5,000 个细胞的数据集大约 15 分钟即可完成训练。这个架构说明,为图像生成开发的技术同样可以迁移到连续生物状态建模。

  • 短期回报在于更好地选择实验,而不是实现自主生物学。 一些类器官需要数月培养,脑类器官甚至需要「1 年甚至更久」,而获取单细胞 RNA 数据至少需要 1 周;Squidiff 可能在大约 1 小时内生成另一种条件。Siyu He 警告不要完全信任模型:它的价值在于对生长因子、药物和基因扰动进行初筛,让固定的湿实验流程能够运行信息量更高的实验。

  • Squidiff 承重的技巧是语义向量运算,承重的风险则在于生物学并不线性。 通过比较同一细胞类型中的基线细胞和扰动细胞学到的扰动方向,可以迁移到另一种细胞类型,类似熟悉的「man 到 woman」嵌入方向;随后,扩散模型将被操纵的潜在状态扩展为完整的转录组。Siyu He 称这是一种近似:潜在空间可能比基因表达空间更平滑、更接近线性,但多个生长因子和方向可能通向同一个最终状态。

  • 目前披露的验证强于仅展示合理性的演示,但仍暴露出模型需要更多信息的地方。 团队在一个从 day zero 到 day three 的分化序列中隐藏了 day one 和 day two 的状态,训练时隐藏了组合基因扰动、只使用其组成部分,并将推断出的类器官中间状态与公开时间序列研究进行比较。测试案例总体有效,但中间状态预测「没有我们以为的那么好」,凸显了有用信号与真实答案之间的重要区别。

  • 要泛化到未见过的药物,模型必须把条件变成一等输入。 如果训练数据只有药物 A 和 B,基础版 Squidiff 无法预测药物 C,因为它没有 C 的语义向量;一个拟议中的变体会增加编码药物分子结构和剂量的 adapter。另一个类似视频生成的方向,则会学习完整时间序列,而不是在两个端点之间画一条直线。

  • CORAL 处理的是互补瓶颈:分子丰富度和空间分辨率通常来自彼此错位的不同测量。 单细胞 RNA 测序会摧毁组织位置;空间转录组学可以覆盖数以万计的基因,但空间分辨率较低;空间蛋白组学的空间分辨率更高,却只能覆盖少得多的蛋白质,而且测量结果可能来自相邻、发生偏移的组织切片。CORAL 的目标是把它们组合成生物学意义上的「彩色高分辨率图像」。

  • CORAL 的图架构将粗粒度组织测量解卷积为细胞级状态,同时建模邻域效应和边界。 图神经网络对每个细胞及可配置数量的最近邻进行建模,从而支持细胞间相互作用分析、功能域发现和空间变异性绘制。重构损失和 KL 散度目标函数支持保真度,图平滑项则约束局部平滑;二者的平衡很重要,因为肿瘤可能渐进式浸润,而骨骼和血管边界可能非常锐利。

  • 数据质量和验证基础设施,仍是所有虚拟细胞平台背后的战略瓶颈。 已有模拟器生成的合成数据可以提供可知的真实答案,用于基准测试,随后再与真实组织数据和生物学知识比较;Siyu He 强调「没有闭环」(there’s no loop)。Siyu He 提到 virtual-cell、Billion Cells、biological-model 和 digital-twin 等方向,并展望它们未来 1–2 年的发展,但表示,进展和临床可信度需要生物学家、机器学习研究者、统计学家、高质量数据、实验和临床测试协同推进。

摘要 · 为研究而整理的核心内容

1. 转录组是强大的状态读出,但不是完整的细胞

  • Siyu He 从生物学的中心法则讲起:细胞大体共享同一套 DNA,但转录会把 DNA 的不同片段转化为 RNA,RNA 随后再被翻译成具有不同结构和功能的蛋白质。这使基因表达成为细胞可观测差异的主要来源。

  • Siyu He 的一位博士后导师 Stephen Quake 提供了一个令人印象深刻的简洁说法:「细胞就是一袋 RNA」。单细胞 RNA 测序将这袋 RNA 转化为一个向量,包含约 30,000–60,000 个基因的表达测量,为定量生物学提供了异常丰富、覆盖整个基因组的起点。

  • 研究人员通过降维、流形方法和聚类来解读这些向量。具有相似共表达程序的细胞会聚成不同状态或类型;高度特征化的标志基因随后把这些聚类连接到生物学解释上,包括判断某一细胞群是否呈现已知的癌症相关特征。

  • Siyu He 明确了边界:转录组学遗漏了表观基因组和染色体变化、蛋白质丰度、线粒体以及其他细胞过程的信息。它之所以有吸引力,是因为当前技术能够提供全基因组覆盖,而不是因为它讲述了一个细胞的全部故事。

2. 破坏性且缓慢的实验,为虚拟转录组打开空间

  • 单细胞测序首先要把组织解离成单个细胞,因此获取分子读出会摧毁样本,也会移除其空间组织。研究人员无法在每个中间阶段反复测量同一个样本;一旦采样进行测序,样本就被摧毁了。

  • Siyu He 的现实动机来自湿实验经历。培养一个类器官可能需要数月,一些脑类器官甚至需要「1 年甚至更久」,而一个错误可能迫使实验重新开始。从解离组织到拿到可用的 RNA 文件,测序成本高昂,所需时间也可能至少为 1 周。

  • Squidiff 探索的是,生成式 AI 能否创建对应化学刺激、基因上调或下调、分化或其他条件的「虚拟或数字转录组」。它也可以研究那些实验根本难以捕捉的状态。

  • Labenz 的框架把资源约束说得很清楚:科学进展部分取决于有限的实验室空间、受过训练的人员和培养时间,能否被分配到正确的实验上。Siyu He 的答案不是取消测量,而是提供快速预测和有用直觉,引导昂贵的实验转向信息量更高的条件。

3. 生成的表达状态仍然有用,因为现有生物学可以解读它们

  • Squidiff 的最终输出是一份合成转录组,其整体形式与实验获得的单细胞 RNA 数据相同。因此,研究人员可以直接使用熟悉的分析方法——细胞注释、聚类、标志基因检查、通路分析和条件间比较——而不必重新设计一套完全不同的解读层。

  • 相似状态应当表达相似的基因组合,这使预测细胞能够被归入已知细胞群。生物学界已经认识的标志基因,则把一个 30,000 维输出连接到一系列问题:出现了哪种细胞类型、哪些信号通路被激活、癌变状态是否持续存在。

  • 这种可解释性是有条件的,并不会自动出现。只有当预测的标志基因程序和通路关系经得起与隐藏实验或外部生物学发现的比较时,生成向量才有价值;一份看起来在统计上真实的转录组,不足以证明某种生物学机制。

4. Squidiff 从适度、任务特定的数据集起步,而不是直接构建通用模型

  • 底层数据是细胞×基因矩阵,辅以元数据,包括已标注的细胞类型、组织来源、疾病阶段以及其他标签或实验条件。Siyu He 使用社区公开发布的数据集,也与能够生成新样本或执行定向验证的湿实验室展开合作。

  • 不同项目的数据规模差异极大,从数千个细胞到整个领域的数百万乃至数十亿个细胞不等;但当前的 Squidiff 实例会在特定子集上训练。Siyu He 表示,约 5,000 个细胞大约 15 分钟即可完成训练,同时提醒,细胞太少可能导致模型欠拟合或过拟合。

  • 尽管数据已经公开共享,数据质量和缺失信息仍是主要挑战。公共数据集可能缺少对齐实验所需的条件或元数据,使一致化建模变得困难。

  • Siyu He 的扩展路径是聚合相关系统:构建覆盖许多类器官的基础模型,或者单独构建一个面向真实人体组织的模型。当前的贡献是一种灵活的方法,研究人员可以用自己相关的数据训练,而不是一个已经完成的通用 checkpoint。

5. 架构由联合训练的语义编码器和 DDIM 构成

  • Labenz 自然会问:「为什么不用 Transformer?」Siyu He 解释说,许多生物医学 Transformer 把基因视为有序序列,并学习用于分类或预测的嵌入;Squidiff 则直接建模连续表达值,目标是从复杂分布中生成转录组。

  • Siyu He 还纠正了原始论文和主持人设定中的一个关键错误:「其实并不是 VAE;它是一个语义编码器。」Squidiff 的两个主要组件是该编码器和一个 DDIM 扩散模型,编码后的语义变量为去噪过程提供条件,而去噪过程从高斯噪声开始。

  • 与先单独训练自编码器进行降维、再将其接入生成器的潜在扩散系统不同,Squidiff 会联合训练编码器和扩散过程。该过程在生成真实感的单细胞表达数据时,同时学习噪声信息和语义信息。

  • 语义表示的目标不只是容纳一个细胞类型标签。疾病阶段、实验条件以及其他相关信息都可以进入同一潜在空间,让研究人员无需直接修改数以万计的基因测量值,就能操纵不同条件。

6. 转录组需要不同于图像的扩散机制

  • 图像扩散通常在两个空间维度上运行,并常用 U-Net。Squidiff 的样本则是一维基因表达值向量,因此其去噪器使用带残差连接的多层感知机,并将扩散时间和语义条件纳入其中。

  • 加噪过程本身相对简单,但单细胞表达矩阵很稀疏,包含大量 0。团队会过滤无用基因,把建模能力集中在高变基因上,而不是把每个测量到的基因都视为同等重要。

  • Siyu He 选择扩散模型的原因,在于它能够表达复杂分布。基因表达数据可能包含复杂模式,而扩散模型具备表达这些模式的能力。

  • Labenz 看到了一条更广泛的架构启示:通过图像和视频生成普及的技术流程,正在直接应用于生物学。他认为,这是 AI 系统帮助解决人类难以处理、无法在全量信息上进行推理的问题的又一个例子。

7. 语义运算可以跨细胞类型迁移扰动

  • 核心的计算机内实验从一个对比开始。如果研究人员拥有细胞类型 Z 的基线样本和扰动样本,编码器就可以估计该处理对应的语义方向;把这个方向应用到细胞类型 X,相当于在不先培养和处理 X 的情况下,询问 X 可能如何响应。

  • Labenz 将这一操作类比为嵌入运算:把从 man 到 woman 的向量加到 king 上,可以接近 queen。在这里,减法估计刺激的影响,加法则把它迁移到另一个起始状态;带条件的扩散模型再把改变后的潜在点转换成完整的表达谱。

  • Siyu He 接受这个类比,但强调其中的假设:生物过程复杂,通常也不是线性的。这里的判断只是,语义空间可能比原始基因表达空间更平滑、更有结构,同时扩散过程可以表达非线性的生物过程。

  • 分化提供了一个直观例子。给定 day zero 的 iPSC 测量值和 day three 更成熟的中胚层细胞测量值,系统利用语义插值推断 day one 和 day two 的状态。另一项相关实验则把基因扰动 A 和 B 各自学到的影响结合起来,预测 A+B。

8. 留出测试支持该方法,同时暴露轨迹误差

  • 在分化测试中,模型使用 day zero 和 day three 训练,而真实的 day one 和 day two 测量值被留出。模型通过语义插值生成这些中间转录组,再将预测结果与隐藏的实验测量值比较。

  • 扰动评估遵循同样的逻辑:基因 A、基因 B 及其组合的实验数据都存在,但组合数据在训练时被隐藏。模型必须根据单独的语义方向构建 A+B 状态,而不能直接使用观测到的组合结果。

  • 在更现实的类器官案例中,每天采样会摧毁样本,因此研究人员可能只在 day zero 和最终阶段采集数据,再通过插值生成中间阶段。Siyu He 还介绍了与公开时间序列数据的比较:团队识别出新的细胞状态,研究分化触发的基因,并发现其结果与其他论文中的实验数据保持一致。

  • Labenz 要求讨论失败案例,而不是只展示成功。Siyu He 表示,测试案例总体有效,但中间预测「没有我们希望的那么好」:多个生长因子和方向可能通向同一个最终状态。如果每个生长因子都有对应的语义变量,轨迹可能得到改善;当前的插值仍然是一种近似。

9. 未见过的药物和时间动态定义 Squidiff 的下一批变体

  • 基础版 Squidiff 可以在不同细胞类型之间迁移已知药物的方向,但如果训练数据只包括 A 和 B,它无法从中推断药物 C;因为模型不存在 C 的语义向量。

  • 拟议的解决方案是增加一个药物 adapter,对药物的分子结构和剂量进行编码,再将这一嵌入与细胞的语义变量结合。只要提供相关成分和剂量,模型就可能据此对未见过的药物进行条件生成。

  • 第二个变体受到视频生成的启发。它不再只学习端点状态、假设端点之间存在按比例缩放的方向,而是使用时间序列数据训练并表达随时间发生的发育过程,从而可能捕捉数据中更多的非线性。

  • Siyu He 认为,近期应用包括选择类器官的生长因子配方,以及筛选工程化组织对药物化合物的响应。对癌症患者而言,更长期的愿景是在患者的细胞状态上比较候选药物,并在漫长等待结果之前,预测药物可能的有效性和潜在副作用。

10. 临床价值需要实验、责任机制和复合型团队

  • 运营层面的对比很有吸引力,但边界也很清楚:获取一份新的单细胞数据集可能需要 1 周或更久,而生成另一个有条件的状态可能只需约 1 小时。这种速度可以为项目方向提供「一些直觉」,但 Siyu He 明确表示,研究人员不应「完全信任模型」。

  • Labenz 区分了发现和部署。在实验室里,错误预测可以通过额外实验验证;到了患者护理场景,准确性及其后果都很重要,因此在真实患者身上使用之前,必须完成实验测试、初步验证和临床测试。

  • 当被问及路线图是否只是「一个简单的编程问题」时,Siyu He 强调,进展取决于社区协作:生物学家理解疾病机制,机器学习研究者构建模型,统计学家进行评估,其他专家则帮助提高研究的严谨性。

  • 因此,Siyu He 的信心是有条件的。高质量数据、领域专业知识和外部验证仍然与模型改进不可分割,尤其是因为这些系统正被应用于医疗和健康领域。

11. CORAL 重建单细胞测序摧毁的空间语境

  • CORAL 从 Squidiff 主要模态缺失的信息出发。组织解离保留了单个细胞的分子测量,却移除了细胞的位置;但细胞会彼此以及与邻近细胞发生相互作用,细胞在组织中的位置也可能影响发育和疾病。

  • 空间转录组学及相关空间技术保留了位置,却引入了分辨率取舍。空间转录组学可能以较低空间分辨率测量 30,000–60,000 个基因;空间蛋白组学则可能更清晰地解析位置,但只能测量数量有限的蛋白质。

  • 不同模态还可能来自相邻而非完全相同的组织切片,因此不同测量中的结构可能发生偏移。Siyu He 的比喻是同一个人的两张照片:一张分辨率高但只有灰度,另一张色彩丰富却像马赛克,而且拍摄位置略有不同。

  • CORAL 的目标,就是得到一张彩色高分辨率图像。它整合不完整的不同模态,把低分辨率测量解卷积为预测的细胞级表达谱,从而以原始测量并未直接提供的分辨率分析组织。

12. 图模型、合成真实答案与真实组织共同完成 CORAL 的证据链

  • CORAL 使用图神经网络对单个细胞及其相互作用进行建模。每个细胞不仅连接紧邻的细胞,还连接可配置数量的最近邻;增加邻居数量会扩大建模范围,不过 Siyu He 主要关注附近细胞及其通信。

  • 细胞级潜在特征支持识别「功能域」、组织架构、相互作用模式和空间变异性。肿瘤可能在没有清晰形态边缘的情况下浸润周围正常组织,而骨骼或血管可能呈现锐利边界;分子模式还可能揭示形态学上不可见的结构。

  • CORAL 通过复合目标函数在这两类场景之间取得平衡。重构损失支持准确性,也有助于保留真实的锐利边缘;图平滑正则项则鼓励合理的局部平滑。目标是在平滑与锐利之间取得平衡,而不是要求所有组织边界都变得模糊。

  • Labenz 对合成数据提出的质疑值得保留:如果模拟器能够生成答案,「那我们是否在某种意义上已经知道了需要知道的东西?」Siyu He 回应说,Splatter 以及合作者的 scDesign 等模拟器能够提供已知的域、相互作用和分布,用于受控验证,但它们并不是最终要从真实组织中宣称的发现。

  • 合成数据与实验数据会一起用于验证模型。由于合成数据拥有真实答案,团队可以测试 CORAL 是否识别出功能域、相互作用和更高分辨率的结构。合成数据的生成过程与被评估的模型并不相同,因此 Siyu He 强调「没有闭环」。

  • Labenz 最后提出,相比一个没有明确含义的超级智能或价值数万亿美元的算力建设,把生物数据作为更具体的国家级投资或许更有意义。他提到受隐私限制的临床记录和昂贵的测量;Siyu He 则列举了 virtual-cell 工作、Chan Zuckerberg Initiative 的 Billion Cells Project、能够回答细胞问题的 biological models、virtual doctors 和 digital twins,同时对 AI 在未来 1–2 年加速医疗发展「非常乐观」。

Siyu He

Thank you, Nathan. It’s great to be here, and I’m very excited to share my research with a broader audience.

Nathan Labenz

Cool. We’ve got a lot to unpack, so let’s start with some big-picture context. I’ve been obsessed with AI for the last few years and studying it intensively, so I feel like I have a pretty good general understanding of the landscape. On the biology side, though, the landscape is even bigger and more complicated, and I haven’t spent nearly as much time in it. I always like to start by getting a sense of where you think we are in the big picture, where you think this work fits, and how you’ve decided to orient yourself within it.

The Squidiff paper is about the transcriptome state of the cell, operating at the level of an individual cell. The CORAL paper zooms out to a larger scope, looking at tissue samples. I’ve had the general sense that the grand challenge in biology involves a couple of different things. One is figuring out the intricate web of what causes what: what upregulates what, what downregulates what, and how an intervention ultimately leads to a change in outcome.

Unfortunately for all of us, that’s not only incredibly complicated, but it also works at multiple levels of scale. We can do a pretty good job of saying, “Here’s a sequence. What does that translate to in terms of the shape of a protein?” We’re even starting to get decent at asking how those proteins potentially fit together. But it seems like we still have a long way to go in understanding how all of that adds up to the function and evolution of a cell over time, and probably even further to go in understanding how that aggregates up to the tissue, system, and whole-organism levels.

I’ve been watching this space because it seems like a critical frontier, and you’ve taken a bite out of each of those problems with these 2 papers. With that preface, tell me more about how you see the big-picture landscape, how you understand the grand challenge, and how you’ve oriented your work toward solving it.

Siyu He

The major motivation behind these 2 projects is understanding whether we can utilize AI models and apply them to current biological technologies. The goal is to understand cellular systems and disease mechanisms, and then provide treatment strategies for disease.

We want to use AI models to work with different types of biological data. For example, I focus on single-cell RNA transcriptomics in the first paper and spatial transcriptomics in the second paper. These are 2 different studies, but they are aimed at a very similar goal: using AI models to accelerate our understanding of cellular activity, figure out disease-related processes, and understand how cells respond to drug treatments.

Nathan Labenz

Let’s get into the first one. The first paper that caught my eye was Squidiff. Am I saying that correctly?

Siyu He

Yes, Squidiff.

Nathan Labenz

I’m interested in where the name came from, just as a bit of trivia before we get into the science.

Siyu He

I had a previous paper called CellOracle, and I was thinking that I wanted more models named after sea animals. We have “squid,” but it’s a diffusion model, so we called it Squidiff. It can also be interpreted as “simulation-related” or “stimuli-related quantitative inferences of transcriptomics using diffusion models.”

The name captures the idea of the model and the goal of the project. CORAL, the second one, is also named after a sea organism. Hopefully, I’ll have more.

Nathan Labenz

You’ve got a theme. You’re building an entire cinematic universe of models.

Tell me about transcriptomes in terms of what we need to know that isn’t obvious. My understanding is that a transcriptome is basically a measure of which genes in a cell are being actively transcribed at a given time. It’s amazing that we can do this at the level of a single cell. The technical achievements and resolution of these techniques are remarkable.

We can take a single cell—I assume this is a totally destructive process, but correct me if I’m wrong—and pull out the RNA that has been transcribed from the DNA at a particular point in time. We can look at how much of these different RNA segments are there, and that tells us the state of the cell at that moment: what active processes are happening in that cell. That’s a rookie explanation, though, so tell me what I’m missing.

Siyu He

A good way to start is with the central dogma of biology. The cell is the smallest functional unit of a living system. In our cells, the chromosomes carry DNA. Almost all the cells in the human body share the same DNA, so what makes them different is the transcription process.

DNA is transcribed into RNA, and RNA is further translated into protein. The protein folds into a specific structure, which is what AlphaFold is trying to predict, and then it performs various functions. The process related to transcription is also called gene expression, and it indicates a lot of the variation between cells in the human body.

One of my postdoctoral advisors, Stephen Quake, has said that a cell is a bag of RNA. That indicates that the complexity of cells can be distinguished by gene expression. Protein expression can also distinguish cellular complexity, but current technology is still more prevalent at the transcriptional level. Proteins are probably noisier, even though we have increasingly advanced technologies for measuring them.

People are mostly studying transcriptomes first because they help us understand cellular activity and disease mechanisms.

The second point I want to mention is the development of single-cell RNA sequencing. I won’t go into all the details of the process, but the major step is that we dissociate tissues into individual cells and then measure the molecular expression, or gene expression, in those cells. We can profile around 30,000 to 60,000 genes, which is a very high-dimensional dataset.

It’s also a high-throughput technology, which makes AI models a good way to analyze the data. We have a matrix with cells and genes as the rows and columns, and it indicates a lot of information about individual cells. This is a phenomenal technology that has opened up new areas of quantitative biology.

Nathan Labenz

Thirty thousand to 60,000 different genes can be measured in a single cell. That’s an amazing level of resolution. It’s always helpful to consider the dimensionality of the inputs, both in biology and in the AI architectures we’re going to discuss.

To be very clear, the starting point for this whole process is a transcriptome: 30,000 to 60,000 numbers representing the expression intensity of individual genes, measured through the RNA pulled from the cell when it’s processed in this way.

Do we have any reason to think that the transcriptome is a complete picture? Are there things that are known to be missing? Is there a way in which the transcriptome doesn’t tell us the whole story of what’s happening in a cell?

Siyu He

Transcriptomes capture some of the heterogeneity between cells, but there is more information. For example, we can look at epigenomics in cells and at chromosomal events that differ across individual cells. There is also protein expression, as I mentioned, which can distinguish cellular states. Sometimes the transcriptome and protein expression share similarities.

There are also mitochondrial-related and ribosomal-related processes. Cells are very complicated, and that’s why people haven’t fully figured out exactly what they’re doing. The transcriptome is one type of data that we can obtain at the whole-genome level, so it’s a very good way to start understanding what is happening.

But there is other information that could also be considered in the model design and in the overall analysis.

Nathan Labenz

Let’s get into the Squidiff architecture. First, what are we trying to make the model do? Both of these papers involve carefully designed architectures that are more complicated than what we see even in frontier language models today. I want to get into why the bitter lesson may not have fully arrived for this kind of work, but let’s start with the goal. What do you want Squidiff to be able to do?

Siyu He

The motivation is that I’m a wet-lab researcher. Obtaining single-cell sequencing data through the whole experimental process is painful because you need to wait for cells to grow, and then you need to prepare the entire experiment. It usually takes months, and it’s also very expensive.

I started thinking about whether generative AI models could create virtual or digital transcriptomes. Then I realized that there might be a way to manipulate the major information in those models. For example, we have many conditional generative AI models. What if we could manipulate the conditions, even when we can’t directly observe the resulting state?

That’s why we designed Squidiff. It isn’t only about generating transcriptomes. It also addresses interesting biological questions, such as what happens when a chemical perturbation or gene perturbation makes cells different, and what those cells look like. In many experiments, we can’t measure those states immediately, and experiments at large scale are difficult. That’s how the model can help the field.

Nathan Labenz

You said it can take months to culture a particular cell type and condition of interest, and that sequencing is expensive. Can you put some numbers around that? Is there an established dollar amount, or is it still a relatively bespoke process where the main expense is the time required to perform the sequencing?

Siyu He

It depends on the tissues or samples you’re working with. I was usually working with organoids, which are engineered tissues derived from induced pluripotent stem cells, or iPSCs. We create engineered brain tissues or blood-vessel tissues and use them for disease modeling.

The whole process usually takes months. Some brain organoids need a year or more to create. There can also be mistakes during the experiments, in which case you need to start over. That makes the process very painful, and it’s one reason I wanted a model that could take over some of the task.

There’s another concern: in some cases, the experiments simply cannot be performed. That’s another reason such a model could be useful.

Nathan Labenz

Let’s do a little more on the actual runtime. You might have a particular cell type and a question of interest: what happens if I apply a drug, apply radiation, or upregulate or downregulate a particular gene?

You want to know what happens, and the answer is ultimately a transcriptome. The system outputs a transcriptome—not one that has been measured, but one generated or predicted by the AI model. Is that the right way to understand it? What does that do for you when you have a transcriptome generated by the model?

Siyu He

When we have transcriptomes from the data, we use different types of analysis for single-cell RNA sequencing to understand what is happening. The first step is usually annotating the cells. People often use dimensionality reduction and manifold-learning methods to see how the cells look, as well as clustering methods to classify cells into different cell types.

Because we have gene-expression information for each individual cell, we know which specific genes are expressed. We can identify the related signaling pathways and understand what is happening. That is the usual way people analyze the data.

There are multiple approaches in the field for quantitatively understanding the data.

Nathan Labenz

How well do we understand that process? If you have a transcriptome, regardless of its source—whether it was measured in a wet lab or generated by a model—what level of granularity, specificity, or confidence can we have about what is actually happening in that cell?

We have numbers for 30,000 to 60,000 genes, but can we aggregate that into the things that really matter? Could you say that a cell is definitely cancerous, that it’s no longer cancerous, or that it’s healthy or unhealthy?

Obviously, you can do more specific things because you can analyze the data gene by gene, but what is our ability to aggregate those results into biologically meaningful conclusions?

Siyu He

If cells are similar and in a similar state, they will have similar groups of expressed genes. We call that coexpression. In that case, we can classify or group cells into different clusters, and each cluster usually corresponds to a cell type.

We can identify highly expressed genes in each group. One method people use is marker genes, which are genes identified by the biological community as important markers for a particular cell type. For example, there are cancer-related markers. If researchers identify those markers, they can infer that the cells are probably in a cancerous state.

Pathologists use similar markers, often involving proteins and antibodies, to determine what is happening in tissues.

Nathan Labenz

There’s clustering, and there are indicator genes that are reasonably well understood. Let’s talk about how you train a model to produce these results.

Where does the dataset come from? How much of it is available in the community from other researchers, and how much do you have to gather yourself? In simple terms, how large is the dataset you used for this project?

Siyu He

It depends on the project. We have a matrix of cells by genes. There are 30,000 to 60,000 genes, covering the whole genome, and we can have thousands, millions, or even billions of cells, depending on the project.

In my models, we use smaller, more specific datasets rather than putting everything together. That is partly because of the complexity of the models and the current stage of development. In the future, I would like to use larger-scale datasets and make the models more like foundation models.

The community shares data by publishing it in public repositories, and we use those publicly available datasets. At the same time, we collaborate with wet labs that help us create novel datasets or provide validation for the model.

The data source is one of the major challenges in the field. We need high-quality data, and although most datasets are good, they sometimes lack important information. That makes it difficult to develop consistent models.

Nathan Labenz

Is the data annotated? When you describe a grid of cells and transcriptomes, I understand that there’s one vector for every cell indicating the strength of activity for each gene in the list of 30,000 to 60,000 genes.

Is there additional metadata about what kind of tissue the cell came from, the patient’s condition, or anything along those lines? Or is it just the raw information?

Siyu He

There is metadata, including cell types annotated by me or other authors. Some annotations are already publicly available. The metadata can include what kind of tissue the cell belongs to and what disease stage it represents. There are multiple ways to label this information, and that is one of the requirements of the model.

Nathan Labenz

How many individual cell transcriptomes constitute the training data for this project?

Siyu He

The model can take a large amount of data for training. I’m not providing a pre-trained model, so researchers need to provide their own training datasets. It depends on the dataset they have.

With around 5,000 cells, training takes about 15 minutes.

Nathan Labenz

You can get this working with as few as 5,000 cells?

Siyu He

Yes, although too few cells can cause the model to be underfit or overfit.

Nathan Labenz

That’s a small number. Is that where it starts to work, while in practice you used more? In the fullness of time, could you use much more?

Siyu He

The future direction is to build a larger model that can cover different types of organoids. In the paper, we use organoids as a case study, but we could encompass different organoids in a foundation model for organoid data. We could also create a foundation model for tissues or real human tissues.

It depends on the exact assets and datasets we want to use. The model is flexible and can take different types of data to address researchers’ needs.

Nathan Labenz

Let’s talk about the architecture. In broad terms, there are 2 core parts: a variational autoencoder and a diffusion-model component.

Before getting into how those work, why not use a Transformer? People in our audience are increasingly familiar with Transformers. If someone came to me and said, “You want to predict the next transcriptome state for a cell given the current state and a perturbation,” my first instinct would be that we could probably use a Transformer. Why did you choose a different architecture?

Siyu He

First, I should correct something in the original manuscript. It isn’t actually a variational autoencoder. It’s a semantic encoder connected to a conditional diffusion model.

There are 2 major components. One is a DDIM model, and the other is the semantic encoder. The generative process takes the semantic variable from the encoder, combines it with Gaussian noise, and uses the denoising process of the diffusion model to produce a final transcriptome.

Nathan Labenz

Why use a diffusion architecture rather than a Transformer?

Siyu He

Transformers and diffusion models are both popular generative AI models, and there are many applications of both in biomedicine.

Diffusion models may capture more stochasticity in the data. Their generation process is flexible, and they can model more complex distributions. Transformers are also powerful, and many related studies use them to generate biological data.

The difference is that the model inputs are somewhat different. Transformers usually take sequence data. In our area, researchers often rank genes first, so the input is not necessarily the continuous-valued data itself. The output is also a ranking of genes, which can be different from the original data.

Many Transformer-based foundation models take an embedding of the data and use that representation for downstream classification and prediction tasks. Squidiff has a different purpose: it directly models continuous gene-expression data and uses diffusion to generate transcriptomes.

Nathan Labenz

On the architecture itself, you described the components briefly, but I want to make sure I understand them. I’ve looked at many different architectures over time, and the one this most reminds me of is MindEye, a project that partially came out of Stability AI. MindEye reconstructed images that somebody was looking at based on an fMRI scan of their brain while they were looking at the image.

That system involved a diffusion model and a diffusion prior. There’s something similar happening here. The first step is trying to understand what really matters in the transcriptome of a cell.

As I understand it, a variational autoencoder typically passes raw data through a narrow bottleneck so that the data can be recreated at the other end, capturing as much of the original data as possible. It is trained with a reconstruction loss, with the goal of compressing the data into its most semantic form while preserving the information that matters.

That gives you a highly semantic representation that can be used for other tasks. Is that roughly correct?

Siyu He

There are different architectures, and a combination of a variational autoencoder with a diffusion model is often called a latent-diffusion model. Related models have also been used for transcriptomics.

The difference is that those models are usually trained separately. The autoencoder reduces the dimensionality and produces a latent representation first, making the generation process more efficient.

Squidiff instead trains a unified model. The encoder and diffusion model are trained together. The whole process includes learning the noise and encoding the semantic information, so the model can extract semantic information from the data while also taking advantage of the stochasticity of diffusion models.

The result is more realistic expression data for individual cells.

Nathan Labenz

So this system is trained end to end under one loss function, whereas in other contexts the variational autoencoder is trained separately and then combined with other components downstream.

That brings us to the diffusion model. People are familiar with this from image generation. I like to explain it procedurally: if the goal is to train a model to generate an image, you can give it an image and ask what the image would look like if it were slightly less noisy. With conditioning, you can also ask what it would look like if it were slightly less noisy and more like the thing you’re trying to create.

The genius of this approach is that it allows you to produce training data at scale. You can take images from the internet, gradually add noise, and train the model to denoise them. Eventually, the model can start with pure noise and go step by step until it produces a high-resolution image. With conditioning, it can produce a high-resolution image of something specific.

The same mechanics are at work here, with the major difference being that the conditioning comes from a semantic encoder. The encoder says, “This is the kind of cell we want,” in a lower-dimensional, semantic representation. The diffusion model then maps that back into the fully detailed picture of a particular transcriptome.

What am I missing?

Siyu He

The data structures are different. Most image data is 2-dimensional, with x and y directions. Single-cell transcriptomic data is 1-dimensional: the object is the cell, and the features are the gene expressions.

Because of that, we can’t use the same image-based diffusion architectures, such as U-Nets, to learn the noise. We use a multilayer perceptron, or MLP, to learn the noise. We also use residual connections to incorporate the time embedding and the semantic features.

The semantic features provide implicit information. They don’t only include cell type; they can also include disease stage, experimental conditions, and other related information. They form a unified condition that can be manipulated in the latent space rather than directly in gene-expression space.

Nathan Labenz

Is there anything special about the noising process? In images, it’s relatively simple. In protein folding, though, you need a specialized noising process because adding noise naively to a protein structure can produce something incoherent. Does that apply here, or can you use a relatively simple noising process?

Siyu He

The way we deal with the noise is relatively simple. We didn’t need to address major problems in the noising process.

The main issue is that most expression data is sparse, with many zero values. We need to account for that. We adjust the data by filtering out genes that aren’t useful and focusing on highly variable genes.

Another advantage of diffusion models is that they aren’t limited to simple distributions. They can generate very complex distributions, which is important because gene-expression data can have complex distributions. The model has the capacity to represent those patterns.

Nathan Labenz

Let’s talk about how you run an in silico experiment. You have a semantic encoder that takes a raw transcriptome and converts it into a more compact, semantic representation that abstracts away details and tries to capture what matters.

The key thing I found in the paper is that, to ask what would happen to a cell type if you gave it a particular stimulus, you need to encode that stimulus as a direction in the semantic latent space.

People may have seen examples of this from language or image models. If you have embeddings for “man” and “woman,” you can subtract them to create a direction from one to the other. Then you can apply that direction to another embedding. The classic example is that if you take the embedding for “king” and add the direction from “man” to “woman,” you may get something close to the embedding for “queen.”

It seems that this is the core mechanism here. You might have cell type X and want to know what happens when you apply stimulus Y. It could take months to culture cell type X and run the experiment. But perhaps you have cell type Z and can apply the same stimulus to it. You can calculate the difference between the original and perturbed state of cell type Z, then apply that vector to cell type X and run the diffusion process.

That would let you predict the transcriptome of cell type X after the stimulus without ever directly performing the experiment on cell type X. Is that what is happening?

Siyu He

Yes, that is essentially what we’re doing. We use vector operations to manipulate the semantic variables.

However, this is an assumption and not exactly what happens in every biological process. Biological processes are complex and usually nonlinear. The latent variables are probably more linear than the gene-expression space, so we use linear approximations to resemble the real cases.

For example, in one of our applications, we study cell differentiation. On day 1, we may have iPSCs, which are stem cells. On day 3, we may have more mature mesoderm cells. If we have semantic variables for day 1 and day 3, we can use linear interpolation to estimate the transcriptomes on days 2 and 3.

Another application involves gene perturbations. We can upregulate or downregulate gene A, upregulate or downregulate gene B, and then ask what happens when A and B are perturbed together. The result may be nonlinear, and the diffusion model can represent nonlinear processes. The latent variables may still be manipulated through vector operations because the latent space is smoother and more structured.

That is an approximation, so it’s risky to rely on the model without validation. We applied it to organoid cases and saw exciting results. We identified transient cell states during development and differentiation from iPSCs into epithelial cells, fibroblasts, and blood-vessel structures.

We’re also considering a variant of Squidiff that uses a time-series analysis. Instead of a simple semantic encoder and semantic variables, we would use a more complicated dataset with temporal information. This is similar to video generation, where the model learns the development process over time and captures more of the data’s nonlinearity.

Nathan Labenz

It’s interesting how techniques initially developed for image and video generation are finding direct applications in biology. I was following diffusion models in 2021, when they were generating strange but sometimes intriguing art. It was unclear how useful they would be.

Now we’re seeing direct applications to biology. Image generation has advanced to the point where it can be difficult to distinguish model outputs from photographs, and video generation is moving in the same direction. It’s not perfect yet, but it has come a long way.

Projecting that kind of progress into the biological domain over the next 2 or 3 years would be a paradigm shift. If these models could produce accurate biological simulations, they could completely change the field.

What do you think is the main obstacle? There’s the nonlinearity issue, which seems fundamental. How extensively has the model been validated? I understood that there was a 3-day cell-differentiation process. Because you can scale the vector, are you literally calculating the intermediate days as one-third and two-thirds of the vector?

Did you measure cells on days 1, 2, and 3 and compare the model’s predictions with those real measurements? How closely did they align?

Siyu He

We have several validation strategies. In the differentiation experiments, we have real experimental data from day 0 through day 3. During training, we use day 0 and day 3, but hold out days 1 and 2 as test data. Once the model is trained, we use semantic-variable manipulation to generate predictions for the intermediate stages and compare the predicted gene-expression values with the real measurements.

For gene and drug perturbations, we use a similar approach. We may have experimental data for perturbing gene A, gene B, and genes A and B together. We hold out the combined perturbation and train on the individual perturbations. Then we test whether the model can predict the combined result.

The third strategy involves real organoid data. We need to culture the organoid, so we can’t stop the process and collect a sample every day. Collecting the sample for single-cell sequencing destroys the entire sample. We might collect data at day 0 and at the final day, then use interpolation to generate intermediate stages.

We also have publicly available time-series data. We compare our predictions with that data to determine whether we can identify similar biological patterns. We discovered novel cell states that had not been previously reported. We then studied which genes were triggered by the differentiation process and compared our findings with experimental data from other publications. We found consistent results, which supports the model’s ability to identify new biological states.

Nathan Labenz

The semantic-space manipulations are linear vector operations, which seems like a problem because biology is so nonlinear. But the diffusion process can capture nonlinearities, so perhaps the combination can work. It remains to be seen where those linear operations break down in the semantic space.

Have you seen validation failures where you concluded that a process was fundamentally nonlinear or that the semantic space wasn’t capturing it?

Siyu He

The cases we’ve tested so far have generally worked, but the intermediate-stage predictions haven’t been as good as we hoped. That may be because we are modeling semantic variables that approximate a factor—for example, differentiation from iPSCs to mesoderm cells.

There may be multiple growth factors and multiple directions that can lead to the same final state. We’re making a linear approximation, but if we had more information—such as the semantic variables for each individual growth factor—the model would be much better.

At this point, we don’t have that information, so this is the way we address the problem.

There is another limitation. Suppose we have data for drug A and drug B and want to understand the result of applying them to another cell type. If we have the vectors for A and B, we can apply them to other cell types and predict the responses.

But what if we have a new drug, drug C, for which we don’t know the semantic vector? We can’t model that unseen drug directly.

To address this issue, we’re considering another version of the model with an adapter that encodes drug information. We use the molecular structure of the drug and its dosage information to create an embedding. That embedding is combined with the semantic variable to create a new semantic representation.

In that case, even if the model hasn’t been trained on an application of drug C, it may still be able to make predictions if we provide the components and dosage of drug C.

Nathan Labenz

Do you think this is already capable of accelerating important scientific work? I understand how it could. In the short term, we have a finite amount of wet-lab work that can be done. We have finite lab space, finite numbers of graduate students and postdocs, and a limited amount of time for cell culturing.

If that throughput is relatively fixed, then what determines how much scientific progress we make is whether we run the right experiments—the experiments that produce interesting results we can learn from. The hope for in silico experiments is that they run orders of magnitude faster.

How long does it take to perform one diffusion process?

Siyu He

The model can help researchers accelerate their work. Obtaining single-cell RNA data may take at least 1 week, from dissociation through sequencing and obtaining the RNA files.

If you want to understand what happens under other conditions, you could use the generative model and spend perhaps an hour generating predictions. That gives researchers some intuition about the direction of a project. I wouldn’t say they should trust the model completely, but it provides useful information.

For example, organoids are useful drug-screening platforms. Researchers study different ways of creating organoids. If we have vectors for individual growth factors, we may be able to predict what happens during organoid development and how organoids respond to different drug compounds.

Another application involves an individual patient. If a patient has cancer, they want to know which drugs will be most useful. It’s risky to randomly try drugs, so we need to test them. The model could predict how the patient’s cells respond, whether a drug is likely to be useful, and whether it may have side effects.

That could provide efficient predictions without requiring the patient to wait a long time to see the outcome.

Nathan Labenz

The individual-patient context is especially interesting. I was mainly thinking about general scientific inquiry: run many simulations, identify the most promising ones, and prioritize the wet-lab work accordingly. That alone could accelerate discovery.

But this becomes even more transformative when applied to an individual with an idiosyncratic situation in their own cells.

Do you think the next steps are mainly scaling up, separating the interventions into distinct semantic encoders, and adding a time-series or video-like component? That seems like plenty of work.

I’m always struck by the pace of research in this area. We’re not talking about just one paper today; we have another entire paper to discuss. When you look ahead to those next steps, do you feel that it’s mainly a matter of doing the work?

At my software company, engineers used to say, “It’s just a simple matter of programming.” They meant that it would take time, and there would be challenges, but that we would eventually overcome them. Are you similarly confident that these next steps will lead to better models and predictions?

Siyu He

It depends on the broader community and on collaboration between scientists. We need a large amount of data, and people have different areas of expertise. Collaboration will be very helpful for accelerating the work.

We need biologists to understand the mechanisms of disease, machine-learning researchers to create the models, statisticians to evaluate them, and other experts to make the models and their applications more rigorous.

This is a collaborative effort. Another important point is that we are applying AI models to health care and medicine, so accuracy matters. We are responsible for patients, and it’s critical to understand exactly what the model can do and not use it in the wrong way.

A lot of work is required before these models can become clinically available and testable. We need experimental tests, clinical tests, and a great deal of preliminary validation before applying them to real patients.

In laboratory settings, there is more flexibility because we can perform additional experiments to test the process.

Nathan Labenz

Any last thoughts on SquiDiff before we move on to CORAL?

Siyu He

I think that covers the main points. I really appreciate the questions. You captured the key ideas of the model.

Nathan Labenz

Let’s talk about the CORAL paper. This one was even more challenging for someone without a biology background to understand.

The fundamental challenge is connecting different levels of resolution or scale into an integrated understanding. In many ways, that is the central challenge of science. We can go down to the lowest level—perhaps string theory or quantum mechanics—and develop theories about how a small number of particles interact. Then we move up through atoms, molecules, proteins, cells, tissues, and whole organisms.

There’s often a gap between those levels. We don’t know how to aggregate the smaller units of analysis into the larger thing we care about. Economics has a similar divide between microeconomics and macroeconomics. We don’t have a complete way to aggregate individual economic actors into an economy, so we use a separate top-down approach based on aggregate statistics.

Biology seems to have this problem at a very high level of difficulty. What motivated this work, and what are you trying to do with it?

Siyu He

This project addresses a technical problem involving spatial data. I’ll start by introducing spatial transcriptomics, which is the main technology in the project.

We’ve discussed single-cell RNA sequencing, but it has an obvious limitation: we dissociate the tissue into individual cells, which means we lose the location of each cell. That information is important because cells aren’t isolated. They interact with one another, and we need to know which cells are neighbors.

Losing spatial information makes it difficult to study tissue structure and cellular communication. That led to the development of spatial transcriptomics and spatial proteomics. These technologies emerged roughly 7 years ago, and now an increasing amount of data is being generated. We need computational methods to understand it.

One significant issue is that spatial data does not always have the same resolution or cover the same tissue section. You can think of it like having 2 photographs of yourself: one is a high-resolution grayscale image, and the other is a lower-resolution, colorful mosaic image taken at a different time or from a slightly different location. The goal is to combine them into a high-resolution, colorful image.

The analogy is imperfect because photographs generally have 3 channels—red, green, and blue—whereas our data can have tens of thousands of channels. Spatial transcriptomics can measure 30,000 to 60,000 genes, but it may have lower spatial resolution. Spatial proteomics can have higher spatial resolution, but it can measure only a limited number of proteins.

Neither dataset is perfect. We want to combine them to obtain a more comprehensive understanding of tissue architecture, how cells coordinate with one another, how they respond to disease, and how we can develop new treatment strategies.

There are newer technologies with high resolution for both proteins and RNA, but there is still a significant need for methods that address these differences in resolution.

There is another difficulty: when we measure spatial proteomics and transcriptomics, we often cannot use the exact same tissue slice. The slices are adjacent and close together, but they can still be shifted relative to one another. That makes the data integration more difficult.

Nathan Labenz

That’s a good primer on the difficulty of doing this, especially when the data comes from a living person. It may be easier when the sample comes from an organoid developed in a lab, but with a living person you have the procedure required to obtain the tissue, measurements taken at different times, and samples that may be nearby without being identical.

Then you have different types of measurements, each with different strengths and weaknesses. All of that creates a confusing picture that isn’t integrated into a single understanding.

This reminds me of applications of the Mamba architecture to biomedical imaging. The goal there was to integrate different types of scans—MRIs, ultrasounds, and other imaging modalities—that each provide different information. If the scans are taken at different times and the person is in a slightly different position, it becomes difficult to determine how they should be aligned.

One application involved deforming scans in space so that they could be combined into a coherent view. A clinician could then see the different signals in an integrated way.

It’s a similar problem here, involving different perspectives on what is happening in the body. This is another example of how AI can help with problems for which humans don’t have good intuitions. At best, a small number of experts can develop those intuitions with a great deal of time and effort, but they can process only a limited amount of information.

AI systems can learn the intuition needed to make sense of these data and then apply it at a much larger scale. What are the major contributions of CORAL? One is taking a higher-level, lower-resolution measurement and deconvolving it—taking something that is lower resolution and determining what each individual cell in the larger sample may look like.

What practical value comes out of this whole setup?

Siyu He

Deconvolution is a major task that the model addresses. But CORAL also provides a comprehensive analysis by integrating different modalities, such as transcriptomics and proteomics. The same approach can be transferred to other modalities.

Once we have high-resolution data, we can explore it at the single-cell level and examine the latent features. We can organize the spatial regions and identify structures in the tissue. We call these functional domains.

Another important goal is understanding how cells communicate with one another. CORAL uses a graph-neural-network-based model to infer interactions between individual cells.

The model can also investigate spatial variability within tissues. Together, these capabilities allow us to deconvolve the data, study tissue architecture, and analyze cellular interactions.

Nathan Labenz

I don’t have a strong intuition for how differentiated the body is. It’s remarkable that we start as a single cell, then become a group of relatively undifferentiated cells, and eventually develop many different types of cells.

Some cell types seem to transition gradually, while others appear to have more abrupt transitions. From my basic understanding of biology, segmentation can occur through gradients of signaling factors.

A key assumption in this work seems to be that these transitions are fairly gradual. The graph neural network models cells interacting with their neighbors. I wasn’t sure whether it also models longer-distance interactions.

There also seems to be a component of the loss function designed to keep things smooth, so that as you move through a tissue, the properties change gradually.

I have trouble reconciling that with certain parts of the body where the boundary seems very clear. If I think about a bone and the tissue surrounding it, it seems like we know exactly where the bone stops and the other tissue begins. There doesn’t appear to be a smooth transition, although perhaps I’m simply not zooming in closely enough.

What am I missing, and what would help me develop a better intuition for this?

Siyu He

It depends on the type of tissue being studied. As you said, bones have clearer boundaries, and blood vessels also have boundaries. In tumors, however, the boundaries may not be as clear because tumor cells can infiltrate normal tissue.

In some cases, it’s difficult to identify structures directly from morphology. If we look at molecular information, though, we can identify interesting patterns that are not visible from the tissue’s shape alone.

The model’s ability to capture boundaries depends on the quality of the spatial data, including both the spatial resolution and the gene resolution. The model represents interactions between neighboring cells. It doesn’t only consider the single most adjacent cell; it can also consider the nearest neighbors.

The number of neighbors is controlled by a parameter. We can choose the value of that parameter and create a larger or smaller graph. It isn’t necessary to look too far away because we’re primarily interested in the closest neighboring cells and their communication.

There are many complicated processes in tissue, so we simplify the problem by focusing on each cell and its neighbors.

The graph loss function encourages smoothness. It is a regularization term intended to make the model more realistic. We also have reconstruction losses and a Kullback–Leibler divergence term.

Together, these losses balance smoothness and sharpness. With single-cell-level data, we can identify clear boundaries because the resolution allows us to distinguish individual cells. We can see the boundary between bone and blood vessels, for example, and potentially identify even more detailed internal structures within organs and tissues.

Nathan Labenz

The reconstruction loss incentivizes accuracy, so when there is a sharp boundary, it encourages the model to capture it. The smoothness term tries to smooth things out and acts as a regularizer. A carefully designed compound loss can give you the best of both worlds.

Did I understand correctly that the parameter you mentioned controls how many nearest neighbors each cell connects to? You can use a small number of adjacent cells or expand the radius to include more neighbors.

Siyu He

Yes, that’s correct.

Nathan Labenz

Let’s talk about synthetic data. In language modeling, the community has been on a bit of a roller coaster. There was concern about a data wall, followed by optimism that synthetic data would solve the problem. Then people worried that synthetic data would produce bad models through mode collapse.

That idea eventually lost traction, and it became clear that having models solve problems and applying reinforcement learning to those solutions could work very well. But I have much less intuition about synthetic data in biology.

What kinds of synthetic data exist? How much can we trust it? How much of it did you use in this project, and what are its current limitations?

Siyu He

We use synthetic data along with experimental data to validate the model.

For simulation, there are well-known models and technologies in the field. For both single-cell and spatial data, these simulations are designed to resemble experimental observations. They model distributions and sample structures that are similar to what we observe in real data.

In Squidiff, we use Splatter, a well-known method for simulating single-cell RNA data. For CORAL, which uses spatial data, we work with a collaborator who is now a professor at the University of Connecticut. He designed a model called scDesign, which was published in Nature Biotechnology around 2 years ago.

That model can sample from designed distributions with spatial information and generate observations that resemble spatial gene-expression data.

We also use real data. The advantage of synthetic data is that we have ground truth. If we are interested in cell interactions or functional domains, we know what those domains are in the simulation and can test whether the model identifies them.

In real data, we don’t have that ground truth. We can still use biological knowledge to validate the model, but synthetic data lets us perform a more direct evaluation.

Nathan Labenz

The reason we need synthetic data is that real data is difficult to gather. But this can feel like a hall of mirrors, and you hope it isn’t a house of cards. If we can generate synthetic data, it can seem circular to ask whether we can learn from it. If we already know how to generate the synthetic data, haven’t we already learned what it contains?

What are you learning from the synthetic data that wasn’t already encoded in the process that generated it?

Siyu He

We use the synthetic data to validate whether the model can capture the results we expect. We don’t use exactly the same generation process in the model and in the simulation.

We generate synthetic data with a well-known or common method, and then test whether our model can recover the underlying structures. For example, we may know what a spatial domain looks like or what the true interactions are. We can evaluate whether the model identifies those domains and deconvolves the data into higher resolution.

In real data, we don’t have that ground truth. Once we validate that the model performs well on synthetic data, we can transfer it to real data. That gives us more confidence in the results, even though we don’t have direct ground truth for the real samples.

The synthetic data is generated independently from the model we are evaluating, so there isn’t a loop in the process.

Nathan Labenz

When ground truth is difficult to obtain, things can become tricky.

There’s a lot of discussion about a Manhattan Project for superintelligence. Sometimes it’s unclear what that superintelligence is supposed to look like or do, and whether we’ll be able to control it.

When I talk to people doing detailed work in biology, I sometimes think a worthy Manhattan Project would be obtaining the data needed to scale this research. A lot of relevant data is probably locked up in medical-record systems and difficult to share because of privacy rules. Those rules have an important purpose, but responsibly making more of the data available could unlock enormous progress.

A major investment, potentially involving government support, could also help. If we’re going to subsidize AI development, perhaps we should make the biological data available so that researchers like you can make discoveries more easily.

The most compelling answer I hear when people get excited about superintelligence is that it will cure diseases and help us live longer, healthier lives. I want that. But perhaps instead of building a trillion-dollar data center to create an undefined superintelligence, we should make sure researchers have the data they need to make discoveries directly. Maybe we end up with both.

How different could your work be in 1 or 2 years if there were a strategic, national-level project to make biological data available and facilitate the next stages of scaling and improvement?

Siyu He

Many projects are trying to make models larger and more capable so that they can represent the human body. For example, Google DeepMind has projects involving virtual cells and unified models that represent different types of cells and organs across different scales.

The Chan Zuckerberg Initiative has also launched projects involving very large numbers of cells, such as the Billion Cells Project. These efforts aim to build large-scale models of biological systems.

This is an exciting period because researchers are using foundation models and large datasets to address these questions. We have GPT models, and people are creating biological versions that can answer questions about cells.

Other researchers are interested in clinical and health questions, so we may eventually have virtual doctors. There are many possible applications.

People are also working on digital-twin projects. We may be able to create a digital copy of an individual and understand what might happen if that person made different choices at some point. That is another very interesting direction.

There are many exciting questions, and I’m looking forward to seeing how these models develop over the next 1 or 2 years.

Nathan Labenz

It’s an exciting time, although also a little scary. This kind of work feels like it has mostly pure upside. Obviously, any technology could theoretically be misused, but I don’t think we need to worry about SquiDiff or CORAL getting out of control.

These are focused systems answering specific problems. They’re a reminder that this work still happens and still has value. It isn’t all about embracing the bitter lesson and maximizing the size of the cluster you can apply to a problem.

I appreciate that about both projects. Is there anything else you want to tease about the next steps, or any final thoughts you want to leave us with?

Siyu He

I’m very optimistic about how AI can help health care and medicine. I know people may be worried about how quickly AI is developing, but at least for now, I don’t think there is a reason to be afraid that AI will destroy the world.

I hope researchers working at the intersection of AI and health care will make the development of health care faster and help people live better lives. That’s what I wanted to share.

Thank you for this great opportunity. I’ve really enjoyed the discussion.

解锁细胞的秘密:与 Squidiff 和 CORAL 的 Siyu He 探讨扩散、解卷积与发现 — 文字稿与摘要 | BidClub