AI驱动的Biohub:Mark Zuckerberg 与 Priscilla Chan为何投资数据——来自 Latent.Space
Erik Torenberg × Nathan Labenz × Mark Zuckerberg × Priscilla Chan
CZI正将未来10年的慈善事业集中于Biohub,因为科学——尤其是AI与生物学的结合——已证明是其前10年影响力最大的方向。 Chan说,他们先后尝试了教育、社区支持和科学,最终得出结论:“天啊,就是这个。” Zuckerberg如今称Biohub是这项慈善事业的核心,同时谨慎界定其使命:帮助科学家治愈或预防疾病,而不是宣称由自身完成治愈。
制度性瓶颈在于,变革性科学工具可能需要10至15年、数亿美元投入,而传统资助体系主要支持规模更小、彼此独立的团队。 因此,CZI运营研究机构,让生物学家、工程师和AI研究人员共同办公,并连接 Stanford、UCSF 与 Berkeley。其押注是,新型观测工具能够像望远镜和显微镜当年那样,打开整个领域。
Biohub的核心资产是不断复利的数据飞轮,而不只是模型集合。 其单细胞项目已达到1.25亿个细胞,CZI约贡献25%,更广泛的生态贡献75%;如今,10亿细胞项目只需数月即可推进,成本也仅为此前的一小部分。AlphaFold依赖30年公共数据的经历说明了战略约束:“这些数据集不会自己生成。”
一个真正有用的虚拟细胞,必须连接分子、蛋白质、细胞、空间结构、时间,以及最终包括免疫系统在内的各类系统。 Biohub的潜在优势,是让“前沿生物学与前沿AI同步推进”,围绕模型所需的数据来设计仪器和实验。对投资者而言,这条产业链从仪器和算力,延伸到数据集、模型、湿实验验证及临床合作方。
即使AI大幅提升科研效率,湿实验室仍将是验证闭环。 Chan不知道验证能力是否会成为瓶颈;生物学目前还无法像语言模型那样,低成本完成数万次测试。近期更现实的收益,是由模型生成假设并降低试错风险,让受制于资助规模的科学家不再只能尝试“单打或双打”。
EvolutionaryScale团队的加入,使AI在Biohub下一阶段的组织架构中居于核心位置。 ESM3蛋白质模型团队将与Biohub现有研究人员合并,由 Alex 领导整个项目——一名AI研究人员与顶尖生物学家并肩负责总体研究。Zuckerberg也承诺投入前沿模型和大规模生物学算力,但 Chan 强调,单有模型并不足以构成一个令人满意的10年结果。
临床终点是N-of-one医疗:预测个体的遗传因素和暴露经历如何改变细胞、疾病通路及治疗反应。 Chan给出的典型案例是抑郁症:患者可能要先用数月熟悉的抗抑郁药,之后才知道是否有效:“与此同时,如果药物无效,就意味着这个人在持续受苦。”这并不是要打造一个“CZI应用”,而是由Biohub建设基础工具,再由合作方将其转化为临床影响。
这项使命需要多长时间,或许更多取决于AI进展而非生物学,但AI无法绕过缺失的实证数据。 Zuckerberg说,目标需要10年、20年还是40年,可能取决于强AI的发展速度,前提是前沿生物学投资持续推进。现实中的桥梁包括虚拟免疫系统和工程化细胞,但“治愈和预防所有疾病”意味着让疾病能够被早期发现并得到控制,而不是消灭每一种感染。至于长生不老,双方没有给出结论。
1. CZI正将慈善布局收窄至AI驱动的生物学
Zuckerberg 和 Chan 很早就开始做慈善,是因为他们得到的建议是:慈善和任何学科一样,也需要实践。两人先后尝试教育、社区支持和科学;近10年后,Zuckerberg说,科学带来的影响最大,推进速度也最快。
Chan 的医生经历与 Zuckerberg 的工程师背景共同塑造了这一方向:把湿实验室、算力、科学家、AI研究人员、医生和患者连接起来。他们的结论带有少见的强烈直觉——“天啊,就是这个”——而Biohub将成为这项慈善事业的核心。
Nathan Labenz 从投资者角度提出,联邦科研预算收缩时,私人资本可以为公共产品提供资金,尤其是那些传统融资机制可能无法支持的登月式项目。他同时警告政府滥权和掠夺性征税,认为社会需要“多元且彼此独立的权力中心”。他还暗示,如果政府被证明不配掌握如此强大的技术,AI公司可以释放出拒绝向其提供技术的意愿。
2. 长周期工具建设填补资助体系的结构性缺口
Zuckerberg将Biohub的基础科学角色,与 Gates Foundation 更强调公共卫生和成果转化的定位区分开来。部署一种已知疗法本身就是巨大工程;Biohub要做的,是创造让新疗法成为可能的上游工具和知识。
NIH的资金规模“基本碾压其他所有机构”,但其典型模式是通过相对较小的拨款支持个体研究者。工具开发可能需要10至15年的周期、数亿美元资金和协同执行,这些条件与现有结构并不匹配。
历史类比体现了这套战略:望远镜打开了天文学,显微镜打开了生物学。因此,Biohub选择建设并运营研究机构,而不是单纯发放拨款,将新的计算、成像和人体仪器能力视为整个领域的基础设施。
Chan始终在约束使命边界:治愈或预防所有疾病,不可能“只在我们自己的四堵墙内”完成。Zuckerberg同样更愿意表述为“帮助科学家做到这一点”——建设数据、模型和工具,让每一位研究者都更有效率,再依靠其他人完成成果转化和部署。
3. 物理距离本身就是科学技术的一部分
Biohub的第一次干预简单得甚至有些尴尬:让不同学科的人坐在一起。Zuckerberg说,物理上的接近可以消解误解和分歧,并培养出“半个生物学家、半个AI工程师”的研究人员。
最初的Biohub还促成了 Stanford、UCSF 和 Berkeley 之间的合作,而这些合作此前受到机构边界限制。回头看似乎理所当然,但在当时,这是一次围绕共同问题而非独立拨款来组织科学的全新实验。
Chan认为,学科之间的不适感具有生产力。AI研究人员迫使生物学家明确真正的障碍;生物学家则让模型构建者明白,“数据不只是数据”——数据的采集方式、采集地点及实验背景,决定了它能否支撑可信的生物学推理。“相信是第一步。”
4. 生物学数据先缓慢复利,再突然变得廉价
Human Cell Atlas 最初从一份资助单细胞转录组测量方法的RFA起步。大约10年后,其数据集已达到1.25亿个细胞;CZI或许贡献了其中25%的数据,而通过 CELLxGENE 降低参与门槛,生态系统贡献了其余75%。
Chan用“慢与快”来概括这一过程。AlphaFold受益于过去30年积累的公共数据,而Biohub如今的10亿细胞项目只需数月推进,成本也只是此前的一小部分。每一种新的测量技术都可能先昂贵,随后实现规模化。
规模并不等于完整。现有数据集主要来自健康细胞,而有意义的变量还包括物种、祖源、年龄、性别、环境暴露和疾病状态。一位主持人的判断毫不含糊:“我们才刚刚开始。”
5. 仪器必须加入空间、时间和活体语境
Biohub定制的显微镜仍然稀缺,Chan估计全球可能只有数十台。Biohub正在推进断层图速度、对比度和分辨率的提升,激光相位板也是其中一环。
时间是缺失的维度:人类不是“被冻结的切片”。Chan希望实现动态、无染色、无染料成像,在无需人工干预的情况下观察系统;Zuckerberg的理想,则是越来越多地直接成像活细胞和活体内的过程。
无法直接观察时,策略就是建立关联。高强度X射线可以在分子层面检查死亡肺部,研究人员再将这些结构与活体患者的MRI和CT图像关联起来,以“环绕视图”逼近目前尚无法直接看到的东西。
透明斑马鱼提供了一个直白得近乎可爱的解决方案:“找一个透明的东西。”接下来,AI必须判断哪些机制在人类生物学中具有保守性和相关性,避免进行无法从模式生物迁移到人类的实验。
6. 虚拟细胞需要嵌套模型和湿实验事实
Zuckerberg描述了一套层级结构:先理解分子和蛋白质相互作用,再理解细胞行为,然后是相互作用的细胞,最终延伸到虚拟免疫系统等更高层级。每一层都可以进行抽象,但准确推理仍依赖下层尺度上的“基本功”。
现有模型已经覆盖不同观察视角。Zuckerberg提到专注于基因表达的模型,以及提供空间视角的冷冻电镜模型。他还介绍了 R-Bio model,试图从相关性进一步走向对生物学组件如何组合、如何运作的理解。
一种扩散模型已经能够根据描述的条件生成合成细胞,但Biohub仪器的意义正在于提供事实锚定。其独特主张是,让“前沿生物学与前沿AI同步推进”,专门为研究人员希望构建的模型采集数据。
与语言模型评估不同,生物学无法低成本完成数万次测试。预测结果必须回到湿实验室,由实验判断提出的效应是否真正发生,再将答案反馈回来。对于未来的验证能力,Chan坦率地说:“我现在还不知道答案。”
Zuckerberg否定了模型在近期取代湿实验室的幻想。更早期阶段,模型可以生成假设,而科学家运用判断力选择实验。Chan说,如今高昂的成本迫使经费有限的研究人员倾向于尝试“单打或双打”;基于模型的风险降低,可以让他们不再那么犹豫,更愿意探索具有重大影响的想法。
7. EvolutionaryScale让Biohub成为前沿AI实验室
以 ESM3 蛋白质模型闻名的 EvolutionaryScale 团队将加入Biohub现有的模型研究人员。Zuckerberg称,这或许是AI与生物学交叉领域最有才华的团队之一,兼具深厚的生物学能力和多年的蛋白质模型经验。
Alex 将负责合并后的项目,并与顶尖生物学家合作。Zuckerberg有意让“AI的人”负责整体工作,以此释放明确信号:AI不是辅助职能,而是整个研究项目的基础组织层。
这项承诺覆盖人才、算力和产出。Zuckerberg说,Biohub可能是最早为生物学研究建设大规模算力集群的机构,并计划发布前沿模型;但 Chan 要求的成绩单是,科学家能否利用这些模型加速临床影响。
8. 精准医疗是终点,但医生仍不可替代
Chan从意义未明变异讲起:基因检测可能发现“3处不寻常的东西”,却无法说明是否需要恐慌。生物学模型可以在不同细胞中模拟每种变异,将行为变化与疾病通路连接起来,使原本无法开展的个体化实验变得可行。
抑郁症揭示了经验式医疗的代价。医生选择一种熟悉的抗抑郁药,患者等待数月,而药物无效就意味着痛苦持续。Chan希望模型能够根据一个人的遗传因素、暴露经历和生物学特征,预测哪种疗法更适合:“这是真正的精准医疗,是N-of-one。”
当主持人设想一个 CZI 应用或身体“编译器”时,Chan纠正道:“这不是我们现在要做的。”Biohub正在建设分子和细胞层面的基础,合作方必须将其推进到临床影响。Zuckerberg设想的是一个融合多种模态的生物学全能模型,可能需要5或10年。
Chan不认为模型会取代医生:“模型不会把你一路带到终点。”AI在皮肤病灶和视网膜异常识别上已经表现出色,但医生仍贡献判断、临床输入、解释、照护和同理心——也就是陪伴患者、建立信任的工作。
9. 可控的疾病、工程化免疫与数据决定时间表
Zuckerberg将医疗体系从被动治疗重新定义为主动检测。在癌症可能转移之前发现潜在致癌突变,会改变最终结果;“治愈和预防”并不意味着消灭每一种细菌,而是让疾病能够被早期发现并“某种程度上得到控制”。
虚拟免疫系统是一个重要候选方向,因为B细胞、T细胞和NK细胞会彼此作用、在全身移动,同时还必须维持微妙平衡。现有 CAR-T 疗法已经表明,免疫细胞可以被重新编程,用来攻击癌症。
在 New York Biohub,Chan描述了一种工程化细胞:进入心脏,检测斑块,将信号记录进DNA,自溶,并通过无细胞DNA输出二进制结果。后续版本可能派出患者自身的工程化免疫细胞清除这些斑块;相关工作还可能帮助理解MS、狼疮,以及痴呆症中可能存在的自身免疫成分。
当被问及这项使命需要10年、20年还是40年时,Zuckerberg说,答案可能取决于AI的发展速度——但前提是前沿生物学提供实证基础。Chan指出,分析第1亿个细胞并不会自动带来一篇显眼的终身教职论文;Zuckerberg则称,为模型训练生成数据代表着“思维方式的倒置”。最终约束很直接:“这些数据集不会自己生成。”
Today I'm excited to share a special crossover episode from the Latent Space podcast. Latent Space, I would say, is the number-one podcast for AI engineers, and I find hosts Swyx and Alessio to be consistently outstanding sources of insight into the latest trends in AI-powered coding and AI application development.
Today's episode is a bit different: a conversation with Mark Zuckerberg and Priscilla Chan, who are celebrating the 10-year anniversary of the Chan Zuckerberg Initiative, about why they are doubling down on the interdisciplinary Biohub, with the goal of leading a new era of AI-powered biology and ultimately equipping scientists to cure or prevent all disease in the coming decades.
In this conversation, Mark and Priscilla describe their perspective on the current state of biology, the role they see AI playing going forward, and the strategy underlying their investments. Highlights include why the traditional funding model fails to bring scientists, engineers, and AI experts together to tackle the most important problems in the way we might hope; their vision for a frontier biology lab that works in sync with a frontier AI lab; the acquisition of EvolutionaryScale, creators of the leading protein model ESM3; and the appointment of CEO Alex Rivas to lead the combined program.
They also discuss their plan to develop new data-collection techniques, which will naturally give rise to massive datasets on which new AI models can be trained; the roadmap to a virtual cell capable of simulating biological responses in silico, potentially revolutionizing not just drug discovery but our understanding of biology in general; and their ultimate vision for precision medicine, moving from clinical trial and error to true N-of-1 treatments designed based on each individual's unique biology.
While the conversation itself focuses on the intersection of AI and biology, for me it also serves as an important reminder of the unique role that private capital often plays in scientific progress and, considering the current moment, the importance of classical liberal values more broadly. Ironically, for all we hear that the US must win the AI race to ensure that the best AI models project American rather than Chinese values around the world, I see actors across the political spectrum pushing America toward a more Chinese model of state dominance.
The civil rights violations and abuses of power we're seeing right now from the federal government are plainly un-American. I've been glad to see prominent voices in the AI space, including Jeff Dean at Google, Dario and Chris Olah at Anthropic, and various researchers at OpenAI, speaking up against them. If anything, I think the AI industry ought to consider doing more, starting by signaling that they would be willing to withhold their technology from a government that proves itself unworthy to wield such power.
At the same time—and certainly to a much lesser degree right now—I also worry that recent proposals for confiscatory taxes, if enacted, would make moonshots like the Biohub much rarer. I do agree, of course, with Warren Buffett that it's absurd that he pays a lower tax rate than his secretary. But at a time when the federal government is cutting research budgets and generally acting against medical advice, society stands to benefit tremendously when self-made tech billionaires turn their formidable talents and immense resources to solving global problems and providing public goods.
More generally, as fears of the concentration of AI power grow, it only seems prudent that society should maintain a diverse set of independent power centers that can hopefully balance and exist in equilibrium with one another. Such checks and balances were, of course, central to the Framers' vision for the US government. In my view, they remain essential for societal dynamism and resilience today.
Hey, everyone. Welcome to the Latent Space podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Latent Space.
Hello. We're so delighted to be in the Imaging Institute of CZI with, literally, CZI. Welcome, Mark and Priscilla. [laughter]
Thanks for having us. Thanks for getting nerdy.
Yeah, we're excited to do this. We so don't often get to see this side of you, so thank you for taking some time out to talk about this. It's sort of the 10-year anniversary of CZI, so I just wanted to introduce people. If people have not been caught up, one of the interesting things that we found out just from talking to your teams is that there's an interesting difference between how you guys started CZI and the Gates Foundation. I heard that Bill Gates is a mentor of yours, so maybe you could tell that story of deciding to start CZI and deciding to pursue basic science instead of translational work.
One of the core things for us with CZI was just getting started earlier. We got some advice that philanthropy and doing science, just like any other discipline, requires practice, and you're not going to be good at it overnight. We should just dig in and start doing a few different iterations on it, see what we enjoy and where we think we can have an impact, and go from there.
We're coming up in November on the 10-year anniversary of when we started CZI. There's a lot of work that we've been really proud to be a part of, including work in education and supporting communities. But when we reflect on it, we feel like the work that we've done in science really has had the biggest impact and, in a lot of ways, is accelerating.
Especially with all the advances in AI that are coming, I think the ability to have an even bigger impact over the coming decade is becoming really clear. This is coming into focus. For the next period, we really want to make science the main focus of what we're doing, and specifically the Biohub organization.
We're really proud of this model that we've helped pioneer, which we can go into detail on. It's going to be the main focus of our philanthropy, and it's something that we're very excited about.
When we started 10 years ago, we had this idea: I bring experience as a physician, Mark's an engineer, and he builds things, and we have an opportunity to give back resources to make an impact on this world. We tried a bunch of things.
The thing about running a philanthropy that makes me incredibly envious of people who run companies is that you guys can have a dashboard, there are financial results, and people tell you if you're on the right track or the wrong track. There's clarity. But in philanthropy, there's so much you can do, and it takes a long time for you to get a sense of what has momentum and what we're doing that is actually bringing all of our skills and resources to maximum impact.
Over the past 10 years, I would say we've been getting a sense of what that thing is that really allows us to have the most impact and makes the most of what we bring to the table. It's really been around AI and biology, where we're like, "Oh my gosh, this is it."
The ecosystem is big, and we really think our ability to bring great scientists and great AI researchers together between the wet lab and the compute, as well as our ability to bring physicians and patients into the picture, is a unique niche for us at the Biohub.
We need others to take the work to translation. The Gates Foundation has a strong focus on translation, and we have had a number of really awesome collaborations and continue to do so. We really look at the basic, fundamental research and are able to partner with someone who's thinking about the translation layer. That's incredible.
We kind of see the first decade, and I would love to get your take, as a decade of creating data.
Creating a science ecosystem and then starting to work on some of the models, and the next decade maybe as more of the applied modeling side. At what point did you decide that just doing the tooling was what mattered, versus—you could have worked to cure malaria in Africa too, or some other disease?
Take a step back, and this is kind of related to your first question too. The space is huge. There are lots of other philanthropies, including Gates, who I think would say that they're primarily focused on public health. Once you know what a cure is, just getting it out to the world is a huge thing too, and someone needs to do that. That's a lot of work and a lot of resources, and it's good that they're doing that.
Basic science is another completely different part of the innovation funnel to enable that. Our view is that the federal government basically dwarfs everyone else in terms of how much they invest through the NIH. But there's a certain pattern to how they invest, which is really enabling a lot of individual investigators to do work.
Our observation was that if you look at the history of science, a lot of major advances are basically preceded by new tools or new ways of observing things. The initial telescope allowed a lot of advances in astronomy.
The invention of the microscope allowed a lot of understanding of biology. Similarly, I think we're at a point in history where a lot of new tools are being built: computational tools, tools to instrument the body in different ways, and tools to understand things.
Often, that tool development just takes a longer-term time frame and sometimes a larger commitment of capital, including because the way to do it isn't necessarily just to make grants to a lot of different people. You need to really operate it yourself, which I think is one thing that's different about the way that we've operated from others. Most of the time, when you think about philanthropy, you think about giving money away in terms of grants.
A lot of what we're doing is actually building up these institutes and building labs to do that kind of research ourselves by bringing in leading scientists and engineers and all that. But that's the strategy: we feel like there are a lot of new tools to develop. There's sort of been a hole in the ecosystem where tool development, and the 10- to 15-year runway that you need to do that, often requires hundreds of millions of dollars to build things like the microscopes and imaging that you're seeing in this institute here.
I think that's been underfunded, and that's where we think that if we do that kind of work, it can just give all these other scientists way more tools to accelerate the pace of research, hopefully discover cures, and then you have folks who are focused on public health who bring that out to the world and deploy it to everyone.
Yeah. I mean, our mission is to cure or prevent all diseases, and that's not going to happen just within our four walls. So the strategy has to be: how do we make every single scientist and everyone better and more effective? The strategy Mark talked about is where we landed on how to actually maximally move the field forward.
Yeah. The mission is to cure or prevent all diseases. By the way, a lot of people outside of the CZI world still find this concept very alien, but talking to the CZI people, they truly believe it. It's impressive how you pick the right mission to motivate everyone to work toward this enormous task.
Well, it's kind of a funny thing. We like to talk about the mission as helping scientists do it, right? Because we're not actually curing the diseases. We're just trying to build the tools, advancing the—
Data models.
Yeah. Basically accelerating scientists' work toward that. But a funny thing about it is that we initially had a time frame of the end of the century. When you ask biologists, there's a lot of questioning around, "Okay, that's really ambitious. Are we going to be able to do that?" And then when you ask AI people, it's like, [laughter] "That should be really easy. Why are you so unambitious that you're shooting for just the end of the century?"
I do think that, at the pace that AI is improving things, it might be possible significantly sooner than that. I don't think it's necessarily worth putting a number or a date on it, but to your point about the first decade, it was sort of about doing work like the Human Cell Atlas to help understand basically all of the specifics and data about all the different configurations of every cell in the body.
When we did that, we had this vague notion that it would be useful to advance science. But I think that, like a lot of people in the tech industry, we've even been impressed by how quickly AI has accelerated. That ended up being a really valuable thing to have done over the last 10 years, especially for where AI is now and for the models that can be built with that.
But the thing that's interesting, don't you agree, is that from a tech perspective, I totally agree that in our intersection of AI and biology, the AI folks are like, "Yep," and the biologists are like, "Hmm."
I think it's actually that confluence of conversations that leads both the biologists to be like, "Okay, I'm really uncomfortable about this idea and timeline, but if I'm really pinned down to think about it, what are actually the barriers? What would you need to do?" You're really forcing people to think through those questions.
You're forcing that conversation from the biology side, and from the AI side, you're really getting a sense of what data is. It's not just data—you guys know this. You need to know how the data was collected and from where, and being able to connect the AI researchers to the folks who are actually gathering the data on a daily basis makes their work better.
It's that conversation that's happening here that I think makes people outside so excited about this, because it's credible. They have really dug in and thought through how that would work, and they're excited and they believe. Believing is the first step.
Believing is the first step. There's a general pattern of software eating the world, and I think AI eating the world is kind of like the next version of this. I was talking with Garrett outside, who says he's a biologist, but I think he's using models like SAM from Meta.
You're like, "You don't look like only a biologist." [laughter]
What does a biologist look like?
I know, he's working on models out there, and that's like—biologists are working using models, right? They're not just in the wet lab.
Yeah, that's totally. I think one of the things that, referencing the wet lab, is one of the key approaches that you're pursuing: turning things from mostly wet lab into something in silico. How far along are we?
I think the first step, which I think is easy to overlook, is basically what Priscilla was talking about: just getting these folks together. It's worth taking a beat to talk about this, because I think most people assume that this is obviously what you would do.
But it's somewhat novel in science because of the way a lot of funding has been done. You grant individual teams relatively small grants, and people do a lot of science independently. It is pretty amazing how much progress you can make if you just have people from different disciplines sit together.
Over my career, both at Meta and here, teams have not been working together for some reason, or they disagree on something. It's like, okay, physically just have them next to each other, and it actually is super helpful.
So here, what are we doing? It's not just bringing together the biologists and the engineers, which was a core part of the initial Biohub model. It was also unlocking the ability for people to work together across institutions.
The first Biohub that we started here, between Stanford, UCSF, and Berkeley, allowed a lot more collaboration between scientists and engineers at those universities than was happening in practice before. You can look at this and be like, "All right, that seems really obvious," but it actually was sort of an interesting and novel experiment, and one that I'm really happy to see others also implementing because I think it's such a clear win.
That's the human side of bringing people together and having them sit together. So I would say that's step 1, or step 0, and it's probably quite overlooked, but it's a fundamental part of the model. I guess it also goes back to this idea that we're not just granting funds to other people. We're building an institution. We're having people sit together.
Then you get these people who are half biologist, half AI engineer because they have some experience doing it. We can talk through the specific models, and there's a lot of exciting stuff there, but I'd say it's an early glimpse of where this is all going.
You want to build up these models hierarchically. You give them a lot of data about specific proteins, and they can model specific proteins in the cells. Then you can model different cell behavior, and eventually you kind of zoom out and you're modeling a virtual immune system or something like that.
It's hard to simulate the immune system without having a good understanding of how a cell might work. It's hard to understand or simulate how a cell might work if you don't really understand how the proteins interact. So you need systems that understand data at all different levels of this, and then you pull them together.
If you look at the different models, there are versions that are focused on which parts of the genome are being expressed in different ways. The cryo-EM model, which I think is very interesting and is built off of the data here, is the only model that I'm aware of that's a spatial model of basically how these cells work.
You want to be able to look at stuff from different perspectives and then put them together, and you build a richer and richer model of how these cells work. But we are definitely at the beginning of this journey.
But it's slow and fast. Slow and fast, right? When we built the Human Cell Atlas, we started 10 years ago. It was one of our first RFAs, and the first RFA was actually to fund the methodologies for how you would get a single-cell transcriptome.
It took us about 10 years to get to a place where we now have one of the largest corpora of RNA transcriptomes: 125 million cells. It cost a lot of money, and the really cool thing we discovered through that process was that if we could seed the effort and make it easy for people to contribute.
It happened. That’s CELLxGENE. We were responsible for maybe 25% of the data, and the rest of the ecosystem contributed 75% of it. That’s an incredible asset and has been very important in modeling work. Similarly, if you look at AlphaFold, they built off publicly available data that was collected for 30 years prior, right? So that takes a long time.
But now we’re doing the Billion Cell Project, and that is taking months and at a fraction of the price. Really slow to fast, but it’s a single dimension, and cells are so complicated. Here, we’re looking, like Mark said, at the 3-dimensional imaging structures, and that’s slow and expensive. But with cryo-EM, it will get fast again, and you just have to repeat it. I think we’ll get growth spurts, but it’s all happening just faster and faster.
How do you think about the layers? You have compute, and we’ll talk about that later. On the data side, you build these amazing microscopes. I learned that they’re all built for you to spec. They’re not off-the-shelf things that anybody knows.
How much of a bottleneck is that still? Can we convert the world of atoms into bits now at the right rate, or do we need more work on the microscopes themselves, too?
And you’re never done, right?
Right. Yeah. Speed here has been a big question of how to get the process through. We’ve worked on the speed at which we can look at tomograms, as well as the contrast and resolution, and that’s where the laser phase plate comes in. It’s about making the data better and getting the data faster.
It’s a bottleneck insofar as there are only—I don’t know the exact number—maybe tens of these microscopes in the world. So that’s one bottleneck. And I think, really, it’s like when I was saying it’s slow and then fast. There are so many other dimensions that we don’t have yet. The cool thing here is that with the transcriptomic work, we’re looking at cellular expression, and with the imaging work, you’re able to localize it in space. Now you want to connect those 2, but that’s still 2 dimensions connected. Time is another dimension. We need to get dynamic imaging in place.
Oh, God. So much resolution. Yeah.
Right. But there’s really cool biological innovation. We need innovation in the way we can look at things, like stain-free and dye-free, so we can look at things without human intervention. Time as a dimension is another one, because we are not frozen slices.
I think it’s just continuously looking at what the next dimension is that we want to be able to either understand deeply or connect to our existing corpus of data and knowledge.
And obviously, the ideal would be to increasingly be able to image things inside living cells, right? You can simulate it a bit: you can take a cell out or some culture—
Destructive. It’s like, okay, it’s living for a little bit or something, but you really want to be able to, as much as possible, actually understand what’s going on in living organisms.
Can that be done? What approaches are there?
Well, the better it gets.
There’s this cool methodology. There is a really high-intensity X-ray methodology you can use. The organ has to be dead, so you can just shoot high-intensity X-rays at a lung and understand, at a molecular level, how the lung is assembled. Then you can correlate that with living imagery—MRIs of the lungs, CTs of the lungs—and look at the associations between the living images in real patients and the sample that you put into the high-intensity X-ray.
That’s another example of correlating data types so that we can get that high-level specificity with clinical data that impacts humans. In some ways, that’s the point about building these AI biological models: you can have a lot of data and interpolate that in that space and understand it. Then there’s one of the models—
Again, I mean, this is really early work, but the R-Bio model—the idea of doing reasoning—is that you don’t just get correlation, but you get some understanding of the logic of how these things get together, too.
So, yeah, I think it’s probably going to be a while, and people don’t have great hypotheses on how you’d actually do molecular imaging of a cell deep inside a living organism. But the goal is to approximate that as much as possible with this kind of surround view of different things that you can image.
You guys like to see cool stuff. It’s not here, but at our San Francisco Biohub, we do image see-through fish called zebrafish.
Zebrafish.
That’s another good example. It’s another good model.
It’s like, all right, what’s a good way to image a living thing? Take a see-through thing.
Take a see-through thing. [laughter]
And then use a model to say, how does this see-through thing actually relate to us, right? I’m not that interested in curing disease—cure, prevent, manage all disease for zebrafish, I am very interested.
Speak for yourself—for zebrafish. Yeah.
Mark’s pro-zebrafish. I’m okay on zebrafish. [laughter] But another application of large language models is looking at what is conserved and what is actually relevant and important to the way human biology works in a fish model. Being able to have that translation be more effective, so we don’t waste our time on things that won’t apply in a model organism, is another really interesting way to elevate biology.
On the data side, can you just give an overview of how far along we are? What percentage of all cells do we image, and what’s the distribution of them? When you say 150 million to 1 billion cells, is that a lot? Is that 10%?
The funny thing is that until recently, we didn’t know how many cell types. Wild thing. This was a big part of the Human Cell Atlas project. There wasn’t even a catalog. It’s kind of like imagining the periodic table in chemistry, but it doesn’t end.
You don’t have the squares.
We know it’s billions. We know there are billions of cells in a human, and we’ve only truly looked at a fraction of them. We’ve looked at it largely in healthy cells, so the number of permutations—age, species, because not all research is in humans, ancestries, what is your genetic background, age—babies are different from old people, gender, environmental exposures—all of those things are actually permutations on the cell that you want to be able to understand in healthy and disease states.
I feel confident that we are at the beginning of this. I'll ask a bit of an obvious question about the intersection of AI and biology: don't we want precision in biology? Don't we want some grounding in a world model that we don't normally get in a language model?
Yeah, I think that's sort of the point of doing all the measurement and being able to have all this real data. You have the diffusion model for generating cells that we put out; it's one of the recent models. It's cool because you basically have a model now that you can describe the conditions to, and it'll give you a synthetic cell.
But yeah, you want it to be increasingly grounded, and that's a lot of the point of the biology and the engineering that we're doing: to be able to have these different facets of that. The imaging institute is one part that gets you the spatial data, which is very helpful, as is the work that we're doing in the other Biohubs on cellular engineering, instrumenting inflammation, and things like that. Basically, it's scientific work to build new types of tools that allow us to measure new types of things, which generate data that allow us to ground the models in different ways.
One framing that we have on this that I think is pretty interesting is that there's this concept of a frontier AI lab that's building AI models at the frontier of what's possible. I think you can think about biology in that way, too. There's sort of a concept of a frontier biology lab: labs that are at the cutting edge of building the most advanced imaging, measuring inflammation, or doing cellular engineering in the most advanced ways, whatever the problem space is that you're working in.
Then I think there's this interesting problem space of what happens if you're at the intersection of those 2 areas. You mentioned the work that DeepMind did on AlphaFold, which was great. That's an example of a frontier AI lab using a data set that was generated by other scientists over decades.
But I think part of what we're trying to unlock here with Biohub is the idea of what happens if you do frontier biology and frontier AI in sync together. You're designing the tools on the frontier-biology side in order to specifically collect and learn the types of data that you then want to feed into specific types of models that you want to build, so that they can understand the cells and the body at different types of resolution.
I think it's a much more integrated approach that allows you to design the things you need, which should eventually get us toward more grounding, rather than just allowing folks who are good at AI to do the best they can with whatever biological data happens to be available.
What's the hill-climbing objective in this scenario? With language models, you have benchmarks. You look at the benchmark and just make it go better. With these things, you have to bring it back to the real world. As you build these models, how do you bring the 2 teams together to give feedback?
I think it's very similar to what Mark just said. You want to be able to validate the accuracy question. We expect that these models will get increasingly accurate, but you want to be able to have feedback. It's not as easy as saying, "This output doesn't make sense." You have to actually take it to the wet lab, run the experiment, find out if it actually happened as predicted, and feed it back into the model.
That's the virtuous cycle we want to build to help the AI best serve the biologists and the biologists be part of continuously improving the models.
From a numbers perspective, in a language model, you can run tens of thousands of tests.
Very fast.
Yeah, and we have to build a lot of them out.
Yeah.
And then, going to the wet lab, what do you think the feedback cycle is going to be like? As you start to have more of these things to be tested in the wet lab, do you feel like that's going to be a bottleneck—that we can't take that many?
I don't know the answer to that yet. I think the throughput on established metrics in the wet lab is actually getting quite fast. You can run a lot of experimentation in parallel. But it's not easily at the level of tens of thousands of verifications, so we'll have to—well, we actually have to see, and we'll probably need to be smart about how we do it.
But there's a lot of people who often take these things to the extreme and are like, "Pretty soon, if you have these models, you're just going to be able to run experiments with the models without even having to go to a wet lab." But no, I think that's kind of the biological version of, "Eventually, AI is going to automate every single thing in society." Maybe you get there, right?
I think there's some chance over time, but well before you do, you're going to be able to have models that can help generate hypotheses, and scientists can apply their taste to which ideas or suggestions from this are worth testing. Then you test them and feed them back into the model, which I think is basically the way that every AI model is deployed, even in other places.
Totally. Right now, because the wet lab is so expensive and relatively slow compared to computational experimentation, people are choosing, "I need something to hit." People are going for hypotheses or ideas that are, to use a sports analogy, singles or doubles, but it's just too risky. They only have so much grant funding, and they need something to help move their work along.
But if we have a model that can help de-risk some of the bigger, riskier ideas, that's going to move science faster. I think it makes the science—and those ideas can be sourced with AI as a tool—but really, it's about making the scientist less hesitant to explore big ideas.
Yeah. Obviously, that's a lot of the success of the CZI model, which is serving this part of research that's underserved because there was basically no benefactor or no funding mechanism by which to do this. One thing that we're announcing when we release this podcast is this unification of the Biohub model.
I think it's very analogous to the foundation-model and frontier-lab approach, where you bring together people from different disciplines. You have much longer time horizons than anyone else. Are there any other key elements to the strategy of the Biohub that you're taking?
Well, one thing that we haven't talked about is the EvolutionaryScale team and Alex Reeves and his team joining us.
Let's talk about the announcement. Yeah.
Yes. This is probably the most talented team working on AI and biology, right? At the intersection of basically having a good biology background, they've just been working on—
ESM3.
Yeah, they've been working on some of the top protein models for a long period of time. I think if you want to build an organization that's doing frontier biology and frontier AI, you need to have world-leading AI researchers. We're doing that by basically combining the team that we have that's already put out all the models we're talking about today, plus having the EvolutionaryScale team, which is very renowned, join us.
Alex is basically going to be running the program. I think it's an interesting decision to have the AI person be running the overall program. Partnering with these leading biologists gives a sense of how optimistic we are about the AI work being very fundamental to this, but we're very serious about building out a leading lab on the AI side as well.
That goes for both the talent and the compute. I think we were probably the first to build out a large-scale compute cluster for biological research. I think now there are some others who are doing it too, but we're also building on that, and we really plan to release frontier models.
Do you see that as the 10-year output, like in the next 10 years?
Yesterday. They say it’s faster than that. They’re always in a hurry. We have AGI in 2 years.
So, would that be a satisfactory result for you guys? You fast-forward 10 years, you have the 3 best models in biology, or is there a further goal that you want to have as an output of the foundation?
I have to bring it back to the patient. I think we will be very excited both if we have great models and scientists are using them, but you really want to make sure that it’s accelerating clinical impact. That’s the goal, right? Getting to the AI models is a very challenging milestone that we are working very hard on, and we will get there.
But how do you actually take those models and apply them to change the way people live? There are 2 variants that I think about in the application of these models. Why are they important?
One is that each one of our genomes is incredibly diverse and different. First of all, all 4 of us are unique people, but we also have things that are sort of known indicators of disease and unknown indicators of disease. I actually find the variants of unknown significance to be the most interesting and the most frustrating. Say someone that you love is sort of a diagnostic mystery. They need to go in and look at the genetics. Most likely, they’ll come back and be like, “There are these 3 things that are not usual, but we also don’t know why.” And you’re like, “Okay, should I panic? Should I not panic? What do I do now?”
What you really want to do, and I think these models will be able to do, is look at those variants and actually model out what the impact is in the different cells, how it influences cellular behavior, and whether or not that is tied to a pathway to disease. That’s a big deal, and I think we should be doing that. That is actually the future of medicine, where we think about each individual’s biology based on their genetics, their exposure, and how that predisposes them or not to disease. That’s huge.
We want to be able to see that clinical application, but we can’t. It’s too expensive, too hard to model each person—impossible to model each person in the lab. But if we can build models around this, it is possible. And then we can start thinking with extreme precision. I’m not just talking about rare disease. There are common diseases. I’ll just say depression.
Right now, it’s empirical, right? We just say, “You’re depressed. Here, let’s try this antidepressant.” And it’s usually the one that the doctor’s more familiar with, or maybe one that you’ve heard of. Then you have to try it for months before it’s, “Did it work? Did it not work?”
Months?
Yes.
That’s the cycle. I don’t have familiarity with this. That’s horrible.
Meanwhile, if it doesn’t work, it means the person’s suffering. This applies to almost every disease, right? There has to be some biological explanation as to why some medications work and don’t. Can we actually then look at each patient and say, based on who you are, we think this medication is going to work best for you?
That’s the future I want to live in, where we can actually understand individuals as individuals and use the biology and science very directly to keep them well.
Yeah. If there’s a name for this tool that has the clinical impact that is on the scale of the electron, how do you envision it? I guess I feel like it’s almost going to be the CZI app.
Oh, well, it won’t be. First of all, that’s not what we’re building right now. [laughter] We’re building the basics. We’re understanding cells and molecules. We’re painting a picture. We need partnerships.
You asked about the ecosystem before. There are experts along the way of this pathway, so we are sort of at the fundamental research side.
Yeah.
You need to be able to partner with folks to bring this all the way through impact. But the way I think about it, people call it different things, but essentially you want to get to “you” medicine, where it’s truly precision medicine. It’s N of 1. We’re understanding you and designing therapeutics for you.
Yeah, I like the mission of Rare As One as well. That’s a great framing. Do you feel like that’s possible—almost treating the body as a compiler? Because I know exactly what it looks like, I know exactly what’s going to happen. Or is the body just like there are too many outside inputs, and over time it kind of deviates from what you have?
Well, I think we’ll see how far we can get, but I’m pretty optimistic that we’ll be able to make a bunch of progress. What form does this take technologically? I would imagine you’re taking these different types of virtual cell models and eventually merging them into the equivalent of a biological omni-model, kind of like how, on the language-model side, you had people who did language and then people who did different kinds of media models and perception and all that.
Then eventually you just kind of merge that, and then you aim to get positive transfer by merging it, so that way it’s not just combining capabilities but getting everything else to be stronger. So, technologically, I think that’s basically what it looks like: over whatever it is, a 5- or 10-year period, we’re building up a series of Biohub models that increasingly get all these different dimensions of data and capabilities that can be used to help run individual science experiments and potentially eventually help with finding individual therapies for patients.
Although we’re going to be less on the clinical side, we’re going to be more on the scientific tool-development side, and the main tool, if you will, is these Biohub virtual cell models.
I would say 5 years ago, without the large language models supporting this, I don’t think it would have been possible, really, because biology is incredibly complex. What we’re essentially trying to do is break it down from a discovery-based science, where you get lucky, you get clever, and you figure out a hack to learn something new, to really making it closer to an engineering problem: This is how the system works, and when this breaks, what happens to the rest of the system.
But like you said, there are far too many dimensions for us to hold in our brains. That’s why we’re so excited about this intersection at this moment, because it is possible to consider so many more dimensions, matching the complexity of biology.
What is the role of the doctor in that future? If you can predict everything out and if you take personal superintelligence seriously, do you distribute some of the diagnosis and all of that work, or how do you envision that?
I’ve been thinking about this a lot. I think, first, the model’s not going to take you all the way. You’re still going to need to really look at individual clinical situations, and the doctor is going to be a form of data input into the model, right? There’s some judgment that comes into place, but there are already a lot of models that make doctors really good at what they do.
For instance, looking at your skin, AI is really, really good at detecting lesions in your skin that are concerning. It’s excellent at retinal issues. It is excellent. So, the AI modeling and mapping is really, really good. It’s already happening.
I think about what future doctors should be trained to do, and I really think care and compassion, and sort of walking patients through understanding. I think understanding why leads to trust in both the science and in the clinical pathway. Really walking alongside patients on that journey—it was the original calling of physicians to be healers and to be using great tools to heal patients.
Well, so, ultimately, I also think you can zoom out, though, from the role of a doctor. I think everyone wants the health system to be more proactive and less reactive.
Right. So today, it's like you show up when you're sick, and then you have someone treat you or understand what's going on. I think the goal with a lot of these systems is to be much more proactive about this.
So when we say that the vision is to try to help scientists cure and prevent all diseases, it doesn't mean that there's going to be no bacteria in the world and no one ever starts to get an infection. It's just that, ideally, you can understand all of that really early, right?
Similarly, if someone gets a mutation and it looks like it might become cancerous, you can treat it a lot better if you know that early rather than showing up to a doctor when it's already metastasized and you have a bunch of issues with that. I think there are going to be a lot of opportunities to fundamentally improve the health care system overall, but I agree with everything that you said on this.
I also think that when we say that we think it's going to be possible to prevent and cure all diseases, it's not literally that no one ever gets the beginning of a sickness. It's just that it can be managed in a way where everything is manageable.
I think we discover more diseases the longer we live. Is it possible not to die? Obviously, that's a meme that's coming to fruition. If we theoretically cure all diseases, maybe death is a disease.
Mark just said we had extreme alignment, which I love. Thank you, honey.
This is—
This is one that we don't.
This is one that I'm not sure we have extreme alignment on. In fact, I just haven't thought about this one very much because I think there is so much—
There's other things to do—
There.
There's so much to do in terms of—I’m a pediatrician. I think about babies, and very sad things happen to very small people. I think a lot about that and how we maximize quality of life and address the things that harm small people.
I'm biased, and I haven't thought as much about the other end of the spectrum. But I don't know—I'm 40. Maybe I should, but I feel like I can still focus on the little ones.
I think the strategy is the same, right? I mean, we're basically choosing not to focus on any specific disease and verticalize. Our strategy is one of trying to accelerate scientific progress overall. I think there are a lot of people who are going to focus on each of these individual things, so I don't know—
But we don't have to, because that's not our strategy. Our strategy is to make sure that we have tools that make people do the best science possible out there.
Now, I'll put to you that because aging, environments, and mutations are so diverse, you actually have a high concentration of grouping in the early years, and it should have more diversity in terms of the cell types and the problems that you face in the later years. There might be some imbalance in terms of where all these things happen, but I'm not pitching in any particular direction.
No, I mean, I think it's clear if you look at the trend over the last 100 years. There was this flip, if you pay attention to the history of science, where it changed to a hypothesis-driven scientific method: we're going to run tests and have controlled experiments. Since that happened, average life expectancy has basically increased by about a quarter of a year every year over the last 100 years.
A lot of that, like Priscilla said, is basically making it so that a lot of people don't die young.
Yeah. So far, it has had somewhat less of an impact on extending maximum human life expectancy. Although the oldest people today, I do think in general, are older than the oldest people 20, 30, or 40 years ago.
But I think there's been a little bit less of an increase there and more just making it so that people don't suffer and die prematurely from things. There are other things that you want to focus on here, too. It's not just how long you live; it's the quality of life while you're—
I think it's like—
You can live a full life and have that be high quality, or you can get sick in different ways that kind of add up over time. I think there are lots of different ways to improve. There are all these different analogies that you can throw at this, but I think there's just a lot of room to improve here.
Yeah.
And then the other element I wanted to come back to on the engineering side is that when you're presented with a high-dimensionality problem, you want to reduce things into little boxes that you can manipulate at a higher abstraction. That's something I tried to do with the folks outside, and we really struggled because over here you're imaging at the atomic level, and then you're also worrying about proteins, and then you're also trying to build a cell model.
Is every abstraction leaky? Where are the boxes I can move around and not worry about it? My physics analogy is that in the regular world, you don't have to worry about quantum physics, but here we kind of do.
I think you want to build it up a little bit hierarchically. When you're trying to understand proteins, understanding molecules makes a big difference, but at some level you can kind of just look at correlations in cells.
If you want to have the most accurate model and be able to reason about things, then you probably also want to understand proteins well. I think that kind of extends. That's part of the interesting challenge: it's not just one resolution that you're looking at.
In order to do it well, you kind of have some amount of abstraction, but I think you want the models—just like language models, or how our brains work—to basically build up different levels of abstraction and pattern matching. I think that's here, too, and you basically just need to have some basic excellence and understanding at each of these different levels.
It's weird, the number of levels at which you have the telescope up and down. It's mind-boggling, and I think when people say dimensions, they typically mean orthogonal dimensions, but here it's sort of nested, and yeah—
There are different scales that are oddly different disciplines to understand each specific scale. In a way, the people who are good at understanding one scale are like—
I've never spoken to people at the next scale.
Yeah. Yeah.
Yeah. Physics is there, chemistry is here, biology is there. It's nice to hear about it, but when you see it and you meet the people, you're like, “Oh, this is real.” They are actually working together.
Yeah.
And then there's this goal of the virtual immune system that you're working toward. I would love for you to chat about that. Also, if that happens, what should other people build? There are obviously CRISPR and some of those technologies that people should maybe ramp through. How do you think about the future?
The virtual immune system is obviously a subset of the generalized model we'll eventually get to, but it's super interesting for a couple of reasons. One, it's individual cells interacting with each other. There are a number of cells that we don't even fully understand what they do: B cells, T cells, and NK cells. We can use our current technologies to understand these cells at a more granular level.
That's cool from a biology standpoint, but the clinical impact of understanding the immune system is huge because biology turns out to have already given us a way to keep the body healthy. It also sometimes goes awry and causes disease with autoimmune disease, right? It's a very complex system that has to stay in balance, and if it goes out of balance in either direction, you get sick.
It can also go into your body, and it's a privileged system that is mobile and can go into places like your brain, your pancreas, or your heart to either do maintenance or collect signals. That's built in. So if we can understand this system, we can use it to keep people healthy. We already kind of do.
There are CAR T cells, where we reprogram T cells to go in and fight cancer. In our New York Biohub, we're doing cellular engineering to say, “Can you go into this person's heart, check if they have plaques that are causing problems, write it into your DNA, self-lyse, and then we can read out the signal as cell-free DNA and give us a binary answer: yes or no?”
Then we can put in other engineered cells and imagine that they go in and clear out the plaques using engineered immune cells that are your own. That is incredible. That is a tool that is realistic, too. I know it sounds sci-fi. It is realistic. It is happening.
On the other end of understanding the balance, there are so many autoimmune diseases—MS and lupus, for example. Those are examples of ones we know. I think there are other things where autoimmunity can play a large role that we don't understand. Dementia, for example, can have autoimmunity play a large role in it.
If we can understand the fine balance in which the system needs to be kept, then we can actually impact a lot of the ways the human body is maintained.
So I think it's both interesting from a biology perspective and feasible to model, and probably one of the highest-impact systems now if we can learn how to manipulate.
Amazing.
It's only one system, right?
If you're focused on curing and preventing diseases, the immune system is a pretty important one. I think it's also interesting and unique in a lot of the ways that you said, but there are lots of other parts of the body to understand, too.
I think we're running out of time, so we have 2 questions to close. One, again, 100 years maybe is too long, right? What would it take to do it in 50, in 25? And to make those happen, what should other people build to support your work?
A lot of this is going to end up coming down to how far a lot of these AI methods get, right? There's just this constant ongoing debate around what the time frames are for getting to very strong AI. If you get that, then I think it's pretty optimistic that, with the right investments in frontier biology, you should be able to get these systems that can allow you to have virtual cells, which allow you to do the kind of precision treatments and preventative care that can achieve this kind of mission significantly sooner.
But at the end of the day, I think a lot of that time frame will probably come down to the AI time frame. There's obviously a ton of stuff to do in biology. I think other people doing more frontier biology and helping to collect this type of data and solve these problems is super helpful to that, too. It doesn't automatically happen. But I guess if we're predicting whether it's going to take 10, 20, or 40 years, that is probably more a function of the pace of AI development than it is the pace of the pure biology side.
Yeah, I was going to agree with you. I think we're on a path to gather a lot of important biological data through advances in laboratory technique, but it's not a given. There are different groups that are experts at this all across the nation and across the world, and so we need to be continuing to push the research and the methodologies.
I want to say that the cell atlas was not glamorous work. People were not going to get their tenure-track paper by analyzing the 100 millionth and 120 millionth cell. That's just not it, right? Rethinking the way that this work gets done, and doing big things together in science, is what's going to need to happen to get the knowledge we need to build models that give us this type of insight.
I guess one thought on the type of biology that I think should get done is that there is a certain orientation around choosing problems that will help generate data that can help make the models a lot smarter. You do that when you are very optimistic about the pace of progress and what AI is going to enable, because the classic reason that scientists generated datasets is so that they could basically look through the datasets to make advances.
It is a little bit of an inversion in the thinking, which is, “I'm now going to do this so I can help train this other thing to be better and create more advances.” In a world where you really believe that there's going to be very significant AI progress, I think more frontier biology should be done in that way.
But these datasets aren't going to get created by themselves. There's a lot of work that needs to get done and a lot of investment there. At some level, you could probably have the smartest AI model in the world, but if it doesn't actually have the data to understand this stuff, it's like, okay, you can't just reason from first principles about all these things. A lot of human knowledge comes empirically, not from first-principles reasoning.
I think this is kind of the whole Biohub Network idea that we're building. I've been really happy to see other folks, especially a lot of people in technology, have this orientation, too. They believe a lot in AI. They believe in technological progress. They've generated some significant wealth building their companies, and now they're investing in science research. I think that's great, and I think doing it in this way, where you're building up these networks to solve specific problems—to basically build specific tools that generate data that make the models better—is one approach. It's not that all science should go in that direction, but it's one of the things that I'm quite optimistic about that I think is going to make a very big difference.
Cool. Well, that's probably all the time we have, but I'll just leave it to you guys for any calls to action, anything that you want biologists or engineers to check out.
I mean, check out the models. Check out the models, the tooling.
Yeah, I mean, they're early, but I think they're kind of an interesting sense of where things are going. We'd love feedback on them, and it'll help this feedback loop of what we should build next.
Yeah, I would say let's do this together. We need lots of people coming together to do this work.
Well, thank you for organizing it and solving and curing all diseases.
Trying to help others do it.
All right. Thank you.