Mark Zuckerberg 与 Priscilla Chan:AI 将如何治愈所有疾病
Mark Zuckerberg × Priscilla Chan × Ben Horowitz × Erik Torenberg × Vineeta Agarwala
Biohub 的核心判断是,共享科学工具——而不是再发一轮小额资助——才是加速治愈疾病的最高杠杆。 Mark Zuckerberg 表示,重大突破通常伴随着观察现象的新方法,但成像系统、虚拟细胞模型等工具,可能需要在10–15年内投入1亿美元至10亿美元。「我们不会治愈所有疾病」;Biohub 的策略是帮助科学界做到这一点。
只有当前沿模型与前沿生物学形成闭环数据飞轮,AI 才会真正改变时间表。 Biohub 计划围绕模型的盲区设计实验和仪器,生成专门数据,重新训练模型,再重复这一过程。这也弥合了两种截然不同的文化判断:生物学家认为治愈疾病「雄心大得近乎疯狂」,而 AI 研究者认为这「有点无聊——这件事本来就会自动发生」。
CELLxGENE 说明,开放基础设施可以形成超越最初出资方的网络效应。 它最初是为了解决细胞注释瓶颈,后来统一了单细胞实验室之间的格式和元数据;CZI 只为这一共享资源提供了25%的资金,整个社区贡献了75%。Vineeta Agarwala 的概括是:「为注释而来,因虚拟细胞模型留下」(Come for the annotation, stay for the virtual cell model)。
在昂贵的湿实验开始之前,虚拟细胞可能先扩大生物科技可承受的风险预算。 Vineeta Agarwala 认为,即便只能给出方向性预测,研究人员也可以先在计算机中检验更大胆的假设,不必拿数年的工作、论文或终身教职冒险。Zuckerberg 将这一模型比作「新的果蝇」,并表示不需要完美模拟;方向性信号同样可能有用。目标是建立与人类相关的模型,同时接受这一准则:「所有模型都是错的,但有些模型有用」。
这里所说的精准医疗,是把常见疾病视为一组分别罕见的生物学问题。 Chan 的表述非常明确:「大多数疾病都应该被视为罕见病」,因为今天对高血压和抑郁症的治疗仍高度依赖试错。若能把突变与下游细胞、蛋白表达、药物靶点以及预测的脱靶效应连接起来,就有望实现更精准的诊断和治疗。
CZI 正将慈善事业集中到一个统一运营的 Biohub,以整合 AI、数据生成和生物学研究。 EvolutionaryScale 的研究人员此前在 Meta 从事蛋白质折叠模型研究,如今将加入 Biohub;其负责人也将负责更广泛的科学项目。Chan 表示:「Biohub 确实将成为我们慈善事业的主线」,教育和本地社区工作仍会继续。他们并未将集中化描述为所有科学研究的模式;去中心化工作和外部实验室依然重要。
下一轮实验室扩张将发生在计算领域:Biohub 计划从大约1,000张 GPU 增加到10,000张左右。 外部科学家可以申请使用这部分算力,研究那些单个实验室仅有几十张 GPU、无法处理的问题。正如 Zuckerberg 所说:「GPU 在某种程度上是零和的,数据不是。」
1. 从治愈一切到建设基础设施
Chan 将这项使命追溯到自己在 UCSF 接受儿科训练的经历:有些家庭得到的可能只是一个基因名称或一份泛泛的 PDF,却没有多少可执行的指导。基础科学缺失的,正是那条「希望的管线」:如果不了解疾病机制,临床治疗就无法超越对少数不确定事实的转译。
Zuckerberg 解释了双方的分工:「我们不会治愈所有疾病」;Biohub 的目标是帮助科学家治愈它们。他的历史判断是,显微镜、望远镜等工具通常先于重大突破出现;否则,研究人员就像「无法逐步执行代码、调试问题,却在写代码」。
传统资助不可能靠给每个实验室发下一笔最合适的基金,来兑现百年目标。共享工具、标准化数据集、成像系统和虚拟细胞模型,需要在10–15年内进行约1亿美元至10亿美元的协同投入;这正是短期 NIH 项目与百年尺度目标之间长期资金不足的地带。
2. 大挑战把雄心转化为可投资的时间周期
Chan 为 Biohub 挑战设定的筛选标准既严格又有边界:「我看得到一条路径」,要有可信的负责人,同时保留足够的不确定性来证明承担风险的合理性。一个已经完全解决的问题不够有雄心;一项成功的10–15年计划,最终应当带来超过最初预期的回报。
3个基地对应3种不同能力:纽约负责让细胞检测、记录或响应信号;芝加哥研究组织构建和细胞间通信,包括炎症;旧金山开发深度成像和转录组学。合作大学提供互补实验室,Biohub 团队则跨越传统边界协同工作。
成功的标准,是让精准医疗开发者出现一次「爆发」,而不是在内部制造每一种疗法。Chan 描述了从一个遗传变异——包括令人头疼的「意义未明变异」——到下游细胞效应和蛋白表达,再到药物靶点的完整链条,同时预测药物可能在身体其他部位与什么发生作用。
3. CELLxGENE 从修复注释问题变成共享基础设施
Zuckerberg 认为,到了2025年,生物学仍然没有一个「相当于元素周期表的东西」,这「很疯狂」。单细胞图谱始于10年前,当时重点是方法学资助和种子数据集;那时,虚拟细胞模型还没有成为明确终点。
CELLxGENE 本身源于一个再普通不过的工作流故障:研究人员生成单细胞数据的速度,超过了给数据做注释的速度。Biohub 搭建了注释界面,发布收集到的数据;由于参与实验室都采用同一工具,格式和元数据也在无意中实现了标准化。
随之形成的反馈循环超出了 CZI 自身的贡献。Mark Zuckerberg 表示,CZI 为这一共享资源提供了25%的资金,整个社区贡献了75%;由于共享变得有用且操作简单,社区贡献了数百万个细胞。
Vineeta Agarwala 提供了一个下游案例:一家研究特发性肺纤维化的被投公司,通过 CELLxGENE 图谱检索数百万个健康和患病细胞,识别出成纤维细胞,再分析其基因表达,寻找潜在药物靶点。这个案例说明,开放工具已经触达初创公司、药企和研发社区,而这些用户未必会专门发表论文介绍这些工具,或直接给予工具署名。
4. 虚拟细胞降低提出高风险生物学问题的成本
设想中的架构将从蛋白质延伸到细胞内结构,再到完整细胞,最终走向虚拟免疫系统等复杂系统。Biohub 各基地生成新数据,单个模型分别学习不同能力,再逐步将这些模型组合成更通用的模拟器,用于提出假设和发现药物。
Zuckerberg 介绍了几项早期能力:Variantformer 从经过 CRISPR 编辑的细胞中学习,并预测编辑后的状态;扩散模型根据要求生成某种细胞类型的合成版本;冷冻电镜模型则补充空间理解。目标不是打造一个单体模型,而是逐步整合对罕见细胞状态的不同视角。
Vineeta Agarwala 的论点既关乎科学,也关乎经济:湿实验消耗资金和时间,因此研究人员必须优先选择更可能保住经费、论文和职业生涯的项目。高质量模拟可以先在计算机中检验风险更高的想法。至于准确度,Zuckerberg 表示模型不需要达到100%准确;即便只是方向性信号,也可能帮助降低下一项实验的风险。
生物学推理模型不仅要返回相关性,还需要解释状态如何以及为何发生变化,不过 Zuckerberg 强调,这项工作「还处于非常早期」。他的层级判断很重要:没有蛋白质层面的理解,细胞预测仍然浅显;蛋白质模型与细胞模型结合后,才能支撑更复杂的模拟。
5. 团队统一,模型到实验的闭环得以形成
Biohub 的重组把此前分散的生物学研究、软件、数据集和 AI 纳入同一个运营团队。Chan 所说的机制是一套「完整飞轮」:当模型暴露出数据缺口或盲区,科学家就能设计下一批数据集,并将其中异常丰富的元数据直接反馈给模型训练。
主持人指出,即使数据开放,领域专用模型也已经展现出意外强的能力;明确问题和用户是谁,仍然重要。Biohub 将这一经验延伸到界面设计:CELLxGENE 有意降低计算门槛,未来的虚拟模型也应让免疫学家能够把自己的专业知识带入神经退行性疾病研究。
Zuckerberg 提醒说,这并不是认为所有科学都应当集中化。Biohub 模式要填补的是一个可以通过协作和共享资源释放价值的空间;去中心化工作和外部实验室仍不可或缺。
6. 算力与可衡量的使用场景,成为下一个10年的锚点
Biohub 可能继续增加基地并组建中央 AI 团队,但扩张的主要单位已经不再是建筑面积,而是「我们正在扩充算力」。单个学术实验室可能只有几十张 GPU;拥有1,000张、最终达到10,000张的共享集群,可以支持性质不同的问题,并与外部科学家建立合作。
Chan 承认,自己在慈善事业早期曾羡慕市场拥有清晰的反馈。10年过去,她看到了可用的信号:Biohub 实现了自己承诺的目标,而其项目往往交付了超出预期的结果。运营原则是,在模糊环境中快速迭代,做到「时间跨度要长,但行动要急」。
Zuckerberg 表示,AI 的进步应能让治愈疾病这一世纪目标「显著提前」实现,这也推动了他们继续集中投入科学。收尾讨论将 Biohub 放在基础研究、 biotech 疗法创造、规模化药企交付和公共卫生可及性之间的基础设施层。Ben Horowitz 的检验标准是:「如果我们不存在,这会成为一个问题吗?会。」
This is a space where there's going to be a huge amount of leverage with AI. It still seems like there could be a lot more effort in this space around building tools, and it's a crazy thing that we're here in 2025 and there's not the equivalent of a periodic table of elements for biology. We think that this is probably one of the most important sets of tools that you need to build.
When we first set out with the goal to cure and prevent disease by the end of the century, most scientists honestly couldn't look at us with a straight face.
And that's crazy.
Yes. And it was true because if you just decided to spend the money funding the next best grant for every single lab in the country, there was no pathway to that being true. The biology folks, I think, looked at it as if it were crazy ambitious. And then the AI folks were like, “Well, that's kind of boring. That's just automatically going to happen.” I know. It's like, okay, there's something in between there that needs to be bridged.
Mark and Priscilla, welcome to the a16z podcast.
Thanks for having us.
Yeah, great to be here. Excited.
All right. Excited to have you. You're doing exciting stuff. To that end, almost a decade ago, you guys started the Chan Zuckerberg Initiative with the mission and intent to cure, prevent, and manage all disease by the end of this century. There are a lot of missions that you guys could have poured your time and resources into. Why don't you take us behind the conversations about why you picked this one? Maybe, Priscilla, why don't we start with you and hear your side of the story?
It always surprises people when I talk about how we work in basic science research. I trained as a pediatrician, and people always think, “Oh, it must be about medicine.” For me, I went into medicine because I wanted to improve people's lives. I wanted to make a difference. I wanted to be able to help others.
I think training as a pediatrician at UCSF, I met a lot of patients—frankly, little kids and families—for whom we just had no idea what the problem was. They might have had a specific gene that they could name, if they were lucky, or they could be grouped into a bunch of other diseases, and there'd be a general sort of PDF they'd print out: “This is what we know.” Then it was my job as an intern or resident to try to translate a few lines of information into how we were supposed to take care of the patient. For me, that's when I realized the power of basic science and how we need to work on basic science to advance the forefront of what's possible. And without that, there's sort of—I think of it as the pipeline of hope.
And why did you think you could cure all disease? Because that's a very aggressive goal.
Do you want to answer that one?
Well, we're not going to cure all diseases, to be clear. The strategy is to help scientists and the scientific community cure all diseases. The strategy is really one of accelerating the pace of basic science, and the theory that we had was: if you look at the history of science, most major breakthroughs are basically preceded by the invention of a new tool to observe phenomena in a new way. Think about things like the microscope, being able to observe bacteria, or, in other fields, the telescope.
Just to use an engineering example, it's like you're coding without being able to step through the code and debug things. That's like the old days. Our whole approach is basically: let's help build tools that will accelerate the pace of the whole field. I think that there's a niche that fits that because, if you look at how funding works in science, the vast majority of funding comes from the government and NIH grants. It's parceled out into relatively small grants that allow individual investigators to investigate usually pretty near-term things.
The development of these kinds of new tools, whether it's imaging or building a lot of AI things like virtual cell models, is longer-term and often more expensive to develop. Think on the order of $100 million to $1 billion over a 10- to 15-year period, and then you try to unlock those tools and give them to the scientific community to accelerate the pace. So that's kind of the theory.
Right, and it seems like there's also something that you don't really get credit for the tools in a lot of ways. We have companies that use your tools, and they're very happy about it. But I didn't even know that was the case.
That's why it's philanthropy.
Yeah. Well, it is, but most people do philanthropy to get credit, too. That's a part of it. Did you think about that, or were you just like, “No, this is going to work, and if it works, that's all we need”?
We're super focused on actually making every scientist better, and beyond science, startup founders, because the point is we can't do this alone. When we first set out with the goal to cure and prevent disease by the end of the century, most scientists honestly couldn't look at us with a straight face.
And that's crazy.
Yes. And it was true because if you just decided to spend the money funding the next best grant for every single lab in the country, there was no pathway to that being true. The biology folks, I think, looked at it as if it were crazy ambitious. And then the AI folks are like, “Well, that's kind of boring. That's just automatically going to happen.” I know. It's like, okay, there's something in between there that needs to be bridged.
So basically, you're like, “We're going to cure all disease,” and they're like, “Yeah, can't be done.” Why can't it be done? “Well, because we don't have the tools.” Okay, that's a pretty cool sequence.
Yeah. Yeah. There's also this funny thing where the biology folks, I think, looked at it as if it were crazy ambitious. And then the AI folks are like, “Well, that's kind of boring. That's just automatically going to happen.” I know. It's like, okay, there's something in between there that needs to be bridged. If you can use modern AI tools to build the types of tools that biologists need, that's a big part of how we think about our work.
AI has got to be the most overestimated and underestimated technology ever, simultaneously. So weird. I mean, we'll probably be like the internet early on, but we kind of think about ourselves and the work that we're doing at the Biohub as frontier biology paired with frontier AI, right? There are labs that do frontier AI that are building the most advanced models, and then there are lots of biological research organizations that effectively do very leading-edge research to either discover new data sets or look into certain challenges.
But so far there hasn't been anyone who's tried to do both of those at once. And when you look at something like AlphaFold, which is amazing, it was built off of a public data set that had been produced decades ago. What I think you have the opportunity to do if you do both of those together is produce specific data sets for the purpose of training AI models to build virtual cells that can do specific things.
Right. So I think that's a pretty interesting zone to be in.
And of all the things that we've worked on, when we started CZI, we focused on a number of areas, and what we found is that the science research has had by far the biggest return. So we've just doubled down on it over and over and over until now we're at the point that we're 10 years in, and Biohub is really the main focus of our philanthropy at this point.
But, yeah, I mean, that's basically the focus. Maybe you're not giving yourselves enough credit because you're sort of saying, “Well, there's bite-size science. We didn't want to do that. There's century-scale science, and that seemed like a long time horizon, but achievable and ambitious.” But you've actually identified what I think are really fantastic grand scientific challenges that are right in between. They're 10- to 15-year horizons, at least according to the way you communicate about them and the way you energize the scientific community around them.
Ten to 15 years is an interesting time horizon, similar to the time horizon of a venture-backed company, and similar to the time horizon over which a team can work together. How did you get to that number, and then how are you thinking about the challenges that you take on in each 10- to 15-year wave? That's concrete and achievable. You build a lot of credibility around it, the way that you've announced those challenges.
I'm curious how you guys think about it, but for us, when we looked at the grand challenges on the 10- to 15-year time horizon, it needs to be that when you look at it, you're like, “I see a path,” right? Not everything needs to be solved for us to take it on. In fact, if everything's solved, then that feels like it should just go—
Ambitious enough.
Yeah. We have some risk appetite. We want things where we're like, there's a credible pathway, someone who is at the helm who can do this. And there's enough ambiguity where we feel like we could take on that risk, and if we do it, the returns could be higher than even expected. The way we modeled that in the Biohubs is we have 3 Biohubs.
We have one in San Francisco, one in Chicago, and one in New York. The one in New York works on cell engineering. Can we engineer cells to go in and detect signals, read them out, or take certain actions? In Chicago, we're building tissues and looking at tissue-cell communications within tissues. And then in San Francisco, we're looking at deep imaging and transcriptomics.
The locations are not by accident. We also look at the partner universities, because we have folks who come to the Biohubs to do this work collaboratively, interdisciplinarily, and sort of unconstrained by the traditional lab. But we also build off the labs at these academic institutes that support the work. And so that's how we choose the grand challenge and the locations.
The layering of large language models and AI coming into the picture has been so interesting because we were already building tools to measure interesting data and building the data sets, but we didn't really know what to do with them yet. When large language models came onto the scene, we were like, wow, we can make sense of all of this now.
I'm curious what you view success as in the therapeutic realm. We think a lot about understanding biology, and sometimes we bet on startups that want to unlock completely new biological areas—diseases where we don't know what's going wrong. And then there's another group of folks who say, "Okay, now that we understand what's going wrong, let's fix it. Let's come in with a drug. Let's come in with a new type of chemistry, a new type of antibody." What do you think success for the CZ Biohub looks like 10, 20, or 50 years from now, in terms of the new medicines that you've enabled?
We want there to be an explosion of a community that is building this new wave of what it means to be deploying precision medicine. For rare diseases and common diseases alike, you're really talking about individual biology that we sort of lump together. We often don't know how it happens, right? We know that you have this mutation, or the worst nightmare is that you have a variant of unknown significance. What does that even mean?
The horrible VUS.
Yes. Horrible. You're like, you tell someone you know something, but we don't know what it means. But if you look at the way we've been able to look at variants and single-cell transcriptomics, we're starting to be able to say, "Okay, this variant actually impacts this set of downstream cells." Then we start looking at the proteins that get expressed and how it looks similar or different to what a healthy cell would look like.
Then you can start targeting it. "Okay, let's look at that as a target." You know both the specificity of the target you want to build, based on the ability to connect mutation to protein expression, as well as the ability to predict off-target effects. What are the side effects? You also know where else that drug will be able to interact with the body.
Those are rare, but I really think most diseases should be thought of as rare diseases because each of our biologies is different. Right now, we just get lumped based on age, demographics, and ancestry, if we're lucky to have that level of understanding. But truly, each of our biologies is different. If you look at hypertension or depression, we kind of just go by trial and error, saying, "Let's just try that drug and see what happens."
What should really happen is being able to precisely, accurately, and quickly treat people by looking at individual biology. We want to enable the basic science, and we would be thrilled if people picked up the models that we build to build the diagnostics and therapeutics that need to come.
You've built amazing data sets. You may not hear the feedback from the startup community, the pharma community, and the R&D community, but it's there because you've committed to open source. People may not all be writing papers, but they are using those tools.
There's a startup in our portfolio working on idiopathic pulmonary fibrosis. The name tells you how vexing the disease is. It's idiopathic. We don't know why it happens. IPF is named that way. He was telling me that he used your CELLxGENE atlases to look at millions of single cells in patients with disease and without disease, try to pinpoint the fibroblasts, and double-click on the fibroblasts and their gene expression. He's trying to use that to inform where he could go after a new drug target in this disease that's fundamentally a strange clump of idiopathic origin.
I think there's a huge group of innovators who love the tools, the visualizations, the query systems, and really the software approach that you built to make that data incredibly accessible.
CELLxGENE is almost an accident, though.
Tell us more.
Do you want to share a little bit about CELLxGENE, or do you want me to start?
I don't know which part you want to get into, but the cell atlas work overall is kind of this crazy thing. Here in 2025, there's not the kind of periodic-table-of-elements equivalent for biology, right? That was a lot of the inspiration for it: How do we, both through work that we're going to do in the Biohub and through other grants, pull together and standardize a format where you can have all this data?
When we were starting off, we didn't even necessarily have in mind that we were going to use that to build virtual cell models. I think that's just come into focus as the AI work has advanced, but that's a very exciting thing. We should definitely spend a bunch of time on the virtual cell models, but I'm not sure what you wanted to get into on the cell atlas.
The single-cell work was one of our first RFAs. We started 10 years ago, and we were like, "Okay, we think this is possible." We actually funded the methodology for it to standardize how it was going to be done. That was 10 years ago, and we then seeded a few labs to start building out that data set.
But we were like, there are millions or billions of different cell types and different permutations. How are we going to do this, especially with a burgeoning technique? We ended up seeding a few groups, and they started doing work. Then they told us they had a problem: There was a bottleneck in their workflow because they couldn't annotate the data fast enough.
And so we built CELLxGENE. It was an annotation tool. That's the original source of this. We built the annotation tool to make it easy for people who are doing single-cell science to annotate the data. Then we put the data that we collected publicly so people could share. Because everyone started using the same annotation tool, everyone was standardized on the same data formats.
Then there started being a community around the tool, and they wanted to share back and build the atlas. Now, after 10 years, there are millions of cells that have been built into this shared resource for the entire scientific community. We've only funded 25% of it. 75% came from the broader community saying, "This is useful, and there's an easy way for us to standardize and build the same metadata."
That's right.
It's an interesting example of what you'd call a network effect, right?
Yeah. I was going to say it sounds like the internet.
Come for the annotation, stay for the virtual cell model.
It was very important when we were getting started with the work to have everyone who was doing it use a consistent format, so that way it could be used and portable. Once that took off as the way it would get done, other people just found it valuable.
Yeah. Even relative to prior databases like GIO and whatnot, they're simply not as standardized or quality-controlled.
Yeah, controlled.
Yeah.
Let's get into virtual cells, one of the great challenges that the Grand Challenge would focus on. Maybe talk about what the promise or the hope is, and maybe some of the challenges or where we're at with it.
Yeah. We think that one of the most important tools at this point is basically building up the kind of hierarchy from proteins to different structures within the cell to a whole virtual immune system, or different levels of hierarchy. We think that this is going to end up being a very important set of tools for people to effectively generate hypotheses for different scientific work.
Even before you get to the point where you're really running full experiments in it, you can come up with some estimate of how that might run. It will be useful for some of the precision-medicine-type examples that Priscilla was talking about a few minutes ago, but we think that this is probably one of the most important sets of tools that you need to build.
It's not a single thing, so there are different angles to come at this from. The cell atlas data is helpful for understanding things on a cellular level. There's this great company, EvolutionaryScale, which has a bunch of researchers who formerly worked at Meta on protein-folding models, joining a Biohub. Alex Reeves, the leader of it, is actually going to be the head of the whole science program, which is actually kind of interesting.
Yeah. When you think about it, you have AI and biology coming together, and really, it's an AI person who understands biology running it rather than a biologist who has some understanding of AI. I think it just kind of speaks a little bit to where we think the relative weight of these things is. But we basically view, as Priscilla was saying, the different Biohubs this way. With New York doing cellular engineering, you can have cells that record different things going on around the body and share that data, and then you can build that into models. The Chicago Biohub being able to record inflammation and basically study that in order to help understand it—that's a different data set.
We have the Imaging Institute, where we just trained our first set of models around that. They are the first spatial models for understanding the way that cells look in different states. Eventually, just like you have this analogy on the industry side around language models, where you have different capabilities and then over time you train them into models and they get more and more general.
That's kind of the idea here. We'll build the Biohubs around grand biological challenges. The Biohubs will build tools that will generate novel data sets. We will build models based on those and then eventually combine the models into an increasingly general view of a virtual cell that will be useful both for scientists and, hopefully, startups and companies that are working on finding drugs, which is not our part of the whole thing, but I think is obviously a really important part of what needs to happen.
Yeah. You guys think about risk all the time in terms of when you make investments. I think the promise of being able to do virtual biology using a virtual cell model is that you can actually take on riskier ideas. Right now, grant funding can be hard to come by, and wet-lab work is expensive and slow. It's not just money; it's also time.
You have to choose something that you think is going to have some likelihood of success to keep your lab career going. It naturally leads people to take on some risk, but not a lot of risk, because they need to make sure that they are hitting a certain percentage of the time to make tenure or publish or whatever they need to do. But if you had a virtual cell model where you could simulate really high-quality biology, you could then start testing and tinkering on the computational side and ask riskier questions—things that would have been expensive and costly in terms of time and resources to do in the lab—and actually see if there is promise in doing the experiments in silico before you make the time and money investment in the wet lab.
Do you think of it kind of like a model organism?
Yeah, like it's the new fruit fly.
Yeah. [laughter] I was going to ask, given the complexity of a cell, how close—how accurate do you think you'll get the model to? I mean, just assuming maybe you get it to a perfectly accurate representation of a cell, but how accurate does the virtual cell have to be to be useful?
I think it will obviously iterate and get better and better, because right now we're still just talking about transcriptomics. We're expanding into different ways of looking at the cell, but you get more and more accuracy. I don't think it needs to be 100% accurate to be useful, because you just want to be able to de-risk the idea on the front end a little bit.
The more and more you de-risk it, the more efficient it gets, obviously, but it will be useful if you even get a directional signal. And yes, we do think about it as a model organism, but in a way that has fidelity to the human body. I don't want to—
All models are wrong. Some are useful.
Yeah.
Hopefully, this has utility on certain axes.
Exactly. And just like with language models, you build in specific capabilities. So, for example, one of the models that we're publishing is Variantformer. Basically, it makes it so that it's trained on a bunch of effectively paired examples: You have a cell, you apply CRISPR to it in a place, and you see what comes out the other side. So it is basically able to make that kind of prediction: If you have this edit that you're doing to a cell, what is likely going to happen?
Another one of the models is a diffusion model. Basically, you can describe a type of cell that you would like it to simulate, and it will just produce a kind of synthetic model of the cell. Again, it's kind of interesting because, to Priscilla's point before about how everyone is different and different cells have—you want to be able to simulate these rare configurations. Having at least a synthetic version of what that could look like is interesting, and then you can test against that.
The cryo-EM model, I think, is interesting because it's spatial. It kind of gives you a sense that there are all these different models that you can have that allow you to basically look at different kinds of things, and then you just train them to be increasingly general over time.
Is the modeling technology basically LLMs, or is there a reasoning model?
Oh, that's actually—yeah, I know, that's a fascinating one too, because one of the new models—I think this one is very early—is basically the first reasoning model over biology. So the idea is that you effectively have these models that simulate world models in different ways, and then you want them to be able to not just spit out correlations, in terms of what they've found, but actually be able to reason through how things would evolve and why things would happen.
I think that one's quite early, but it is interesting conceptually, as I think it's clearly going to be an important direction in terms of how these models evolve.
Yeah. No, because that's what I was thinking: If it doesn't work, the next question you have is why?
Yeah.
But I think what you find in reasoning—the analogy—
You're married to your hypothesis. [laughter]
Well, yeah. Sure. I thought you were saying, if the reasoning model doesn't work, why? I think the language-model analogy for that would be that you need better world models or better pre-trained models in order to get the reasoning to be good. But, yeah, you just build more capabilities into it.
I think there's probably an order, too. The work that Alex and the EvolutionaryScale folks worked on is a lot of it is protein, which is interesting because that's at a kind of smaller resolution, obviously, than the cellular data, the Cell Atlas. Part of the hypothesis is that you can look at all these different cells and kind of simulate how they might behave, but you're going to have a somewhat shallow understanding unless you actually have this hierarchical understanding of how the subcomponents of the cells are going to interact.
Our view is that you basically want to build up a state-of-the-art protein model and then have that be a part of the state-of-the-art cellular model. Once you have that, you build things like the virtual immune system, which allows you to simulate much more complicated systems. It's sort of this hierarchical approach to building up these virtual models.
That makes a lot of sense, because also, as you get into personalization, you've got common proteins combining into a unique cell. From a systems standpoint, that makes it much more manageable. That makes a lot of sense.
Yeah.
Yeah. No, it's very fascinating stuff.
Yeah.
So you guys are announcing some big news this week. Do you want to give us a sneak preview?
Well, the big news is thinking about how we are going to be coming together as 1 team. In the past, we've run Biohubs, we've built software, and we've done some AI research, but all of it has been a little bit decentralized. Now, under Alex's leadership, we are going to come together as the Biohub, an operating philanthropy where we are doing the science in service of a singular goal together: How do we actually advance the state of biology and research at the intersection of AI and biology?
Amazing. Alex is amazing.
Yeah, no, he's great. And then the other thing is the piece that I mentioned earlier. CCI has focused on a number of different things. We've really just found over time that we feel like we've been able to make the biggest difference in science, so we've just kept on doubling down on it.
We're going to continue doing work in education. We're going to continue supporting local communities and those different pieces. But going forward, the Biohub is really going to be the main thrust of our philanthropy, and we're very excited about that because I think that, when we started the mission to see if we could help the scientific community cure and prevent diseases by the end of the century.
I do think that, with the advances in AI, it should be possible to do that significantly sooner, and that is a very worthy, important, and exciting goal. We think we have a unique place in the ecosystem where we can help empower others to make fast progress on that.
There are obviously plenty of advantages to decentralization, from management and communication overhead and so forth. What are you trying to add by adding this kind of new layer or unification on top? What are the outputs, and then, I guess, what are the complexities to that? I'm sorry to ask a CEO question.
No, no. I mean, do you want to go for it, then I can jump in.
Yeah. So there are obviously amazing groups doing frontier AI and a lot of groups doing great frontier biology. Where we think we can uniquely contribute is by tying these 2 together. We've funded datasets, we've built datasets, and we're building the instrumentation now to be able to look at the cell—whether it's for tissue-cell communication or cryo-EM, where we can look at the cell at a nearly atomic level.
We have the ability not only to build the datasets but actually to shape and form them the way we want, based on what we see as necessary to complement the existing body of knowledge. We have amazing teams doing that work, and we're building these AI models. The reason to do it together is that we can actually complete the flywheel: the model is looking like it has some gaps and blind spots in this area. Okay, who do we talk to? How do we build the next dataset? We're seeing this in the lab—the metadata is going to be so rich that we can feed it back into the way that we do this modeling.
Yeah. I think it's going to be incredibly powerful. And it's more than just writing down a spec and saying, “Please deliver this.” These people need to be working shoulder to shoulder and shaping each other's work for this to actually be the more and more accurate model of how the human cell works.
Well, yeah. It's so interesting, because that is exactly—it has been the biggest surprise in the industry for us in the AI world. Forget biology for 1 second: the domain-specific models have been super interesting. The original thesis was that some AIs are going to get so smart they're going to be smarter than everybody at everything, but—
Like, on video models, every video model is best at something but not everything. And so knowing what problem you're solving actually turns out to be, ironically, very important in AI, because you can actually get to a way better result. Yes.
If you put the 2 together, we're seeing that over and over again in a way that—
I would say it's very counterintuitive to the whole narrative going into it.
And in biology, it used to be—or at least one assumption was—well, the datasets aren't on the internet. So part of the reason you need a domain-specific model is that the datasets are not public. You guys are kind of bucking that trend, too, by creating a lot of open-source access to the data, and even then it sounds like you're betting on the trend that we're seeing in other industries. But still, there will be nuance in how you annotate that data and curate that data.
Well, and how you talk to a scientist, right? Because you have to not only know the data and the model and so forth, but the conversation is what we keep finding ends up being very, very important, right?
So rich and so important—how you actually—
A scientist isn't going to talk to it like I talk to ChatGPT or whatever. So this is the fly you can talk to.
Yeah. That's really super exciting.
And the user interface is actually really important. You talked about how you guys have a founder who's using CELLxGENE. That user interface was intentionally designed not to require a computational or really deep biological background to be able to use, because you want people coming from different fields to look at the problem. It's like, “Look here. Help us solve problems here.”
Building that user interface in a way where there's not a very high barrier to entry to be able to poke around and learn something and bring knowledge back to your work—that's intentional. We're really hoping that, when we build these virtual models, we get to a place where we can allow a lower and lower barrier to entry for people to say, “I have some knowledge about this. Maybe I can contribute.”
It seems like immunology is behind all this, so it might be part of your century vision.
You need to be able to allow the immunologists to come in and understand neurodegeneration and understand how their world fits in. The more you lower the barrier to entry, the more you allow people to actually think in a truly collaborative and interdisciplinary way.
Will the Biohub grow as a team? Will you employ more people at the Biohub proper, or are you moving toward more of a network model with more sites, more labs, and more community-driven datasets? Which is the thrust? Or maybe it's both.
Probably a little of both. We've added new Biohubs over time, and then we're also building up more of this central AI team.
Cool. But I think these organizational questions of how you set this up are fascinating, and a lot of your approach is informed by what the rest of the field is doing. You kind of think about science as this portfolio, right? Society has a portfolio of things that it's trying to do, and in terms of philanthropy, you want to—
Be the most additive that you can be by trying to figure out what else is underrepresented. Science by default is very decentralized, right? It's kind of the way that grantmaking has worked, the way that I think scientists by default want to work.
So I think a lot of what we've found is that figuring out ways to encourage collaboration in ways that otherwise seem very simple, but weren't happening before, can unlock a lot of value. So the very first Biohub, what we did—there were 2 interesting things. One was this collaboration between UCSF, Stanford, and Berkeley. There are all these really smart people at all these different places who previously, I guess in theory, could have figured out a way to work together, but there wasn't really a formal construct for them to do that, and this just allowed a lot more collaboration.
The other one is cross-disciplinary: basically having biologists sit next to engineers, and this view that these 2 disciplines are things that need to—
There are so many interesting—
In companies, they always set them apart.
Well, it's interesting—no, it's interesting how many organizational questions or problems you can fix just by having 2 teams sit together, right? It doesn't matter what the org chart is or whatever. It's like, you guys need to sit next to each other until you get this thing to work, and—
That's something I really believe in. So—
And you have 10 to 15 years.
Well, no, it's all—communication is such an underrated problem in general, in building anything or solving anything.
That's pretty neat.
Yeah. It's really kind of simple stuff, but I think it's—
It's sort of novel as a model.
And one of the things that's neat is that we've now copied this from the first Biohub to the Biohub Network and expanded it to other models. It's also been neat to see other folks who are working in the field adopt similar models, because it's a pretty intuitive thing.
But you know, at some point you'll reach the point where it's actually really good to have decentralized work, too, right? It shouldn't be that we're saying this is the way that all science should work. We're just saying that there's a space for this. It can unlock a lot of value because, for whatever reason, it hasn't been the default.
Yeah. And we still rely on—
Yeah. There are famous stories in the MIT lab about that. That's how they invented lasers and so forth: they put a bunch of people from different departments in the same—
The lab. Yeah. Well, actually, physics is where we got a lot of the inspiration. Physics has historically been a field where labs have rallied around big projects and big shared resources.
We're relatively centralized, but we still depend on a lot of labs that are doing exact frontier work or complimentary work to come together to support this. There's that. But one more thought on your expansion question is maybe this is like the modern AI lab. We are not expanding a lot of square footage per se, but we're expanding our compute.
The research—they don't want employees working for them. They don't want space. They just want GPUs—
Agents. So it's, in a sense, new lab space. It's much more expensive than wet-lab space.
And you guys have always been creative on that. Even in the last few years, you've created ways to share access to compute. You've enabled academic labs to—I forgot the name of your program—kind of like scientists in residence or something like that, rental, kind of hoteling.
The core of it is clusters. If you look at individual labs, they'll have—
Like, a large lab would have tens of GPUs.
And we were the first to really build a large-scale compute cluster—1,000 GPUs. Now we have plans to move to the 10,000 range, and that requires a different type of project. Obviously, you're able to ask different types of questions.
It's a resource that we use, but we've also invited scientists to apply and say, “What question do you have that could use this amount of resource?” and be able to seed collaborations that way.
And so, if a scientist is out there listening who's not employed by or working at the Biohub but wants to collaborate with the Biohub, you're going to create interesting—
Interesting doors to utilize the resources. That's awesome.
Yeah, the GPUs are somewhat zero-sum, right? The data isn't. [laughter] Yeah.
Yeah. Fair enough.
Yeah. So you're about to celebrate 10 years doing this. As you look out at the years to come, what else can you tell us about either things that you're thinking about for the future, or maybe even principles or a north star that's going to guide how you grow and evolve going forward?
You know, it's been really interesting in the past 10 years because I actually spent the first few years completely envious of people working for for-profit companies because there's so much clarity. The market will tell you whether or not you're doing a good job, whether it's private or public.
If they think you're doing a good job—
If they think you're—[laughter]—they're not always right.
They're not always different.
But I was still envious, because I craved that feedback: Am I doing a good job?
And, you know, 10 years in, the reason why we're doubling down on biology is that not only did we achieve what we said we were going to do, but when we set out on these projects, they actually delivered more than we thought they were going to. I was like, “Okay, that's a signal I can latch on to. That's a signal we can really continue doubling down on and doing more of.”
I think it's about continuing to tolerate the early ambiguity, when you're like, “Okay, I'm going to do more of this,” and being patient, but being willing to have a long time horizon and be impatient at the same time.
Because it's all those iterations along the way that have allowed us to get to this place where, to get lucky, you're ready. We've built data sets to take advantage of AI and large language models, and that's because of all the work that we have been doing. Being able to continue moving forward in this ambiguity, and sometimes lack of signal, on a big goal—I think we've sort of set the DNA for that.
Amazing.
Oh, no pun intended. [laughter]
Yeah. But we get to see how many people use the tools and the feedback.
Yeah. You have customers, which is pretty cool.
Yeah.
For philanthropy. That's awesome.
Yeah. No, it's one of the fun things about building tools: You kind of get to see—
Yeah.
How valuable do people find the tools? Do people use the tools in order to publish important work?
Right, right, right, right. Yeah. And, well, I mean, our feedback is that they're awesome.
Feedback and completely unique, by the way. The other thing is, what would you use if you didn't have this? It's like there's nothing.
No. Yeah. It's a real void. I mean, there's this whole pipeline that needs to exist, from accelerating basic science to funding a lot of people to use it. Then you can get into the biotechs that can start working on coming up with novel therapies, and then you get the pharma companies that do them at scale.
Then there's a space for philanthropy on the other side of public health, basically taking the therapies and bringing them to everyone in the world. But this is a space where there's going to be a huge amount of leverage with AI, and it still seems like there could be a lot more effort in the space around building tools and accelerating the whole thing a lot better.
Yeah. And I do think it is the place where you are completely unique, right? The other things—there are other people who can do that, but there's nobody doing what—
That's got good founder-market fit.
Yes, founder-market fit. [laughter] I mean, if we didn't exist, would it be a problem? Yes. Those questions really land, you know, as a VC.
Like, one of us is an engineer, and the other one is a scientist and doctor.
Yeah, very happy with this direction.
Yeah.
We thank you very much, not only for our companies but for us as humans, for working on this work. It's amazing work. Thank you.
Thank you guys.
Thank you so much.