No Priors 第121期|对话 Chai Discovery 联合创始人 Jack Dent 和 Joshua Meier
- Chai-2 的核心成绩是:最多尝试20次,抗体设计命中率约为20%,而早期计算方法约为0.1%或更低,传统筛选则要在数百万或数十亿个候选分子中搜索。 在超过50个留出靶点上,Chai-2 约有一半实现命中,将开放式搜索重新定义为“工程问题”,并把实验验证周期压缩至约2周。
- 这项基准测试之所以格外有说服力,是因为 Chai 优化评估时追求的是广度,而不是精心挑选的治疗性展示案例。 团队抓取供应商目录中的现货产品,剔除了与 SAbDab 中任何条目序列相似度超过70%的靶点;在相似度降至25%的子集中,成功率依然接近,Joshua Meier 的总结是:“模型并不在意”。他仍认为公布的靶点成功率可能只是下限,因为仓促的实验设置可能引入了错误。
- Chai 的技术跃迁,是从“原子级显微镜”走向分子生成器:结构预测负责看清锁,Chai-2 则逐个原子设计钥匙。 整体误差有时低于一个原子的宽度,而在分布外靶点上的成功说明,模型学到的可能是分子相互作用的某些基本规律,而非死记蛋白质家族——不过联合创始人也承认,他们仍未完全理解模型为何有效。
- 这套逻辑的重点是扩张市场,而不只是降低研发成本:AI 不同的失败模式,可能打开传统筛选无法触及的靶点。 但 Chai-2 目前只在约一半基准靶点上奏效,Meier 强调,药物发现“只是冰山一角”;制造、稳定性、资本市场和临床风险仍待解决。
- 湿实验室会成为规模化的配套环节,而不是显而易见的受害者。 由于模型采样具有概率性,从20个设计扩大到100,000个,可能找到更好的药物,进而改变患者需要注射还是可以皮下注射。与人类逐一审阅10,000条聊天机器人答案不同,实验室可以测试全部10,000个候选分子,CRO 和高通量筛选仍有存在空间。
- 一个合作案例让经济账变得具体:一个由5至10人组成、几年间预计花费$5M-$10M的项目,一直没能找到同时结合人和猴版本靶点的分子。 Chai 下单14条序列后,4条命中人源版本,1条命中猴源版本,其中1条重叠设计同时命中两者;联合创始人称,这已经足以推动项目继续前进。
- 可防御性正在转向一体化设计产品:多目标提示、治疗属性优化、实验室反馈和专家工作流,而不只是一个模型。 Chai-1 已开源;Chai-2 是更大的流水线,用户必须指定表位、多种物种以及规避脱靶。Sarah Guo 的判断很尖锐:顶尖抗体工程师会成为提示词专家,而 Chai 认为,领域专长的重要性只会增加,不会下降。
- 联合创始人“看多生物科技”的依据是已经观察到的研发速度,而不是当前市场环境:XBI 已经低迷5年,他们称这是生物科技数十年来最差的市场之一。 他们将一年内从低于0.1%升至接近20%的跃升,与迷你蛋白在全部5个测试靶点上接近70%的结果相对照;随后只是推测、并非承诺,某些分子类别或许能在另一个年份达到50%以上,甚至接近100%的成功率。
1. Chai 成立于从预测走向设计的拐点
Meier 对时机的判断是:AI 药物发现一直停留在研究设想,直到团队看到一个“1、2年”的窗口——足够早,可以在成果被验证前搭建公司;又足够晚,不至于让公司等上10年。他们押注扩散模型和语言模型的进步,会把蛋白质折叠能力推进到分子相互作用和分子设计。
AlphaFold 2 约在2020年实现突破,能够根据序列、以实验级精度预测单个蛋白质的结构,但药物的作用机制是调节其他分子。早期预测器只能返回“蛋白质的一种视图”,就像扩散模型出现前的图像系统;Chai 想生成多样化的抗体—抗原以及小分子—蛋白质相互作用。
Meier 对早期 AI 生物科技公司的批评是,实验室整合往往“几乎过于紧密”。Chai 仍然做实验,但目标是打造一个可跨数百乃至数千个项目部署的“可移植 AI 平台”;开源的 Chai-1 是其通用性的第一份证明。
Jack Dent 的职业选择遵循同一逻辑:当以“原子级精度”工程化分子变得可信,未来就“再也无法视而不见”。对他而言,此后“几乎很难再用人生做其他事情”。
2. 20次尝试取代开放式抗体搜索
抗体占近期获批药物的近50%,也是10款最畅销药物中的7款,因此 Chai-2 的设计结果具有商业意义,而不只是一次狭窄的科学演示。
工作流从一个明确靶点开始,为一个24孔板生成最多20个抗体,然后送入约2周的验证周期。下单设计中接近20%实现预期结合,而在约一半测试靶点上至少出现了一个命中。
Chai 曾将1%设为全公司的年度目标,因为此前计算方法的成功率约为0.1%或更低。因此,这一结果较内部目标高出约一个数量级,较此前方法高出多个数量级。
传统筛选可能在酵母或噬菌体文库中搜索数百万乃至数十亿个候选分子,像是在“淘金”。其他路径包括免疫小鼠或美洲驼,再提取抗体;COVID 期间,研究人员则从感染者体内提取的抗体中寻找能够中和病毒的那一个。
3. 这套基准测试就是为了击破单靶点展示
大多数 AI 药物设计论文只测试1、2或3个靶点,Chai 测试了超过50个。Meier 将只解决一个问题的结果比作声称某个大语言模型解决了整场竞赛:“你需要一个真正的基准测试”,尤其是当随机性可能让1、2次试验产生误导时。
靶点选择刻意不追求亮眼:Chai 从供应商目录中抓取当时有现货的蛋白质,确保所有材料可以一次性订购。团队提取这些蛋白质的序列,与 SAbDab——蛋白质数据库(PDB)中的抗体结构集合——进行比对,并剔除序列相似度超过70%的靶点。
最终靶点组合并非按治疗重要性挑选,其中一些蛋白质已经有药物项目。Meier 因此将这项工作视为模型评估,而不是展示性项目组合;他还表示,50%的靶点级成功率可能只是下限,因为仓促的实验设置可能引入错误。
更难的一组测试将其与训练数据的相似度降至25%,但成功率基本不变。Meier 的结论是:“模型并不在意”,这挑战了人类认为彼此不同的生物学分类。
4. Chai-2 以原子级精度生成分子钥匙
结构预测提供了一台“原子级显微镜”:输入序列,输出三维空间中的原子位置预测。Chai-2 随后从观察走向生成,同时产出一条新序列和一个旨在结合指定位置的结构。
Dent 用“ImageNet 时刻”和“分子领域的 Midjourney”作类比。如果靶点是一把锁,Chai-2 就通过逐个放置原子来构造钥匙,整个结构的误差有时低于一个原子的宽度:“如果看不见锁,又怎么可能设计出钥匙?”
对陌生靶点的泛化能力说明,数据中存在一种可被模型学习的蛋白质相互作用特征。Meier 认为这很深刻,但问题仍未解决:团队“仍未完全理解”模型为何有效,也没有声称已经提取出模型可能编码的生物学原理。
5. 采用路径将把生成模型与更多实验采样结合起来
Meier 将价值来源分成两类:在计算机上更快地筛选已有项目,以及攻击传统方法无法触及的问题。由于 Chai-2 只在约一半靶点上奏效,近期最合适的场景可能是那些其失败模式与实验室不同的案例。
Dent 向学术团队和产业合作伙伴开放访问权限,因为药物发现资源消耗太大、机会又太广,不可能由一家公司覆盖所有靶点。数百个请求在数小时内涌入,迫使这支约12人的团队进行优先级排序。
产业界的质疑很直接:企业本来就能发现药物,更快的发现是否会改变哪些分子是可能的?联合创始人的回答是重新审视停滞项目;社区中更广泛流传的回应则是,这只是“工具箱里的又一件工具”,企业可能需要使用它,以免被甩在后面。
Meier 预计,增加采样量会改善概率性输出:相较20个设计,扩大到10倍、100倍甚至几个数量级以上,应该能探索更好的区域。CRO 已经在询问100,000个设计的运行规模;届时,一个更优分子可能决定患者需要接受注射,还是可以皮下注射——实现“AI 的最佳能力”与“生物学的最佳能力”结合。
6. 生物科技的看多逻辑不止于更快的单抗
Dent 将 XBI 连续5年的低迷和漫长投资周期,与 Chai 在1年内从低于0.1%跃升至接近20%的结果进行对比。迷你蛋白实验实现接近70%的设计成功率、皮摩尔级亲和力,并在全部5个测试靶点上命中,支撑了他提出的“看多生物科技”棒球帽。
长期类比是一套面向生物学的计算机辅助设计软件:分子工程领域的 SolidWorks,或创意工作者使用的 Photoshop。除了结合能力,模型还必须优化可制造性和稳定性;更容易的设计也可能让结合两个旁位点的双表位抗体,比传统单一位点单抗更容易推进。
一个合作伙伴花了几年时间,投入5至10人和预计$5M-$10M,寻找一种能够同时结合人源和食蟹猴版本蛋白质的交叉反应分子。在仅14个 Chai 设计中,4个命中人源版本,1个命中猴源版本,其中一个重叠命中同时命中两者。
Guo 的反驳值得保留:许多业内人士认为,制药行业的高成本和瓶颈在临床,而非药物发现。Guo 认为,如果发现风险下降,行业就可能变得更高效、更有效。Meier 的回应是,资本市场、临床失败以及制成药物所需的一切,仍是 Chai 当前结果之外“冰山的一角”。
7. Chai 的产品必须把命中转化为多目标药物候选物
Dent 明确划定边界:“这些还不是药物,只是命中。”Chai 必须刻画治疗属性,最终实现端到端、零样本设计完整药物候选物;在20次尝试中看到抗体出现后,这种可能性显得没那么未来,但它仍是未来的投资方向,而非已经证明的结果。
Chai-1 可以被视为一个模型:输入序列,得到结构。Chai-2 则是更大的流水线和产品,其界面必须表达复杂的设计意图,并纳入实验室结果,让系统成为后续轮次的副驾驶。
提示可以指定两个靶点、同时结合人和动物版本,或要求对一种蛋白质保持亲和力、同时避开另一种蛋白质。Meier 那句令人印象深刻的概括——“现在是生病小鼠的好时代”——指出,多物种设计可能消除单独开发替代抗体的需要,并降低动物证据对应的是一个细微不同分子的风险。
给抗体工程师的建议是获取访问权限、学习提示,并“开始梦想新的可能性”。选择跨物种共享的表位,或有意选择不同区域以实现选择性,都会成为高杠杆判断;Dent 预计,一旦这种假设能力真正可用,专业人士会“惊掉下巴”。
8. 一支小型跨学科团队把研究代码当作基础设施
Chai 团队约12人,覆盖化学、物理、生物学、AI 和软件工程;令人意外的是,拥有传统计算机科学学位的人很少。团队的运行约束是共同聚焦:所有人都攻克同一个公司的问题,而不是各自推进偏好的研究项目。
Dent 将抵达新领域前沿描述为“一场彻底的战斗”,伴随“一波波兴奋和痛苦”。他更快的路径是让各领域专家围绕在自己身边,并跨学科学习;同时确保研究人员也是足够强的工程师,能够把发现交付给合作伙伴。
他的架构原则是,必须有人始终掌握整个系统,在软件熵阻断进展前保持简单和模块化。深度学习中的 bug 可能在昂贵运行中隐藏数周;Chai 曾通过反复提交训练任务对 Git 历史做二分,定位代价达数万美元的回归问题,这也是它甚至会为研究代码编写单元测试的原因。
算力投入遵循一种精打细算但进取的门槛:Chai 从免费云额度起步,看到“生命迹象”或 scaling law 后才扩大规模,避免为未经验证的方向扩容。Chai-2 上线后,团队正在招聘产品工程、抗体工程、业务拓展和客户经理岗位。
Hi, listeners. Welcome back to No Priors. Today, I'm excited to speak with Joshua Meier and Jack Dent, two of the co-founders of Chai Discovery and former BioAI and engineering leaders at Meta, OpenAI, and Stripe.
This week, Chai released their industry-leading Chai-2 zero-shot antibody discovery platform. At its core, it's a generative model that can design antibodies that bind to specified targets with 100-fold the hit rate of prior computational approaches. We'll talk about their product, the next frontier for Chai, why they're bullish on biotech, and why the most effective antibody engineers will soon be working as expert prompt engineers.
Jack, Josh, congrats on the Chai-2 launch. Thanks for doing this.
Thanks for having us, Sarah. We're excited to be here.
Great to be here.
Josh, I'll start by asking: You and several of the scientists on the team have been working on AI drug discovery for about a decade now in different settings. I've also been looking at this area for over a decade. We haven't yet seen drugs designed with these AI computational techniques reach the market. What made you believe in this? Why start the company when you did?
That's a great question. Many of us had been working in this space for a while and didn't start a company because it was really a research idea, I think, until very recently. There were signs of life that someday this was going to work, but it wasn't really on the timeline of a company. You can't really start a company thinking that 10 years from now things are going to work, but you also don't want to start a company after it's already working and miss the boat.
The sweet spot is, "Okay, we have maybe 1 or 2 years where we have to really get this off the ground." We made a bet when we started the company that it was going to work. There were really a couple of things that fueled that decision.
The first one was that we made a bet that structure prediction and protein folding were going to get a lot better. Obviously, protein folding was considered solved a couple of years ago, around 2020, with the breakthroughs of AlphaFold 2 and the ability to predict protein structures with experimental accuracy. But it was just a single protein structure at a time. We could take a single protein sequence and see what that protein looks like. That's very useful for basic biology, so we can understand what the proteins we're looking at look like.
But if you think about drug discovery, which is where we're really focused at Chai Discovery, you need to understand how multiple molecules interact with one another. You need to understand how a small-molecule drug is going to modulate a protein, or how an antibody protein is going to modulate an antigen protein.
We started to see early signs of life that that was going to be possible, and again, we made a bet that we would be able to take this to the next level with the kinds of breakthroughs that we were seeing around diffusion models and language models. The previous generation of structure-prediction models would really just predict one conformation of a protein at a time. It was kind of like one view on a protein. It's like the early image models: They didn't have diffusion models, so you weren't really able to look at the diversity of generations that could come out. We thought the same thing would impact drug discovery and protein folding as well.
That's a bit of color on how we decided to start the company, and we did. Lastly, I should say that almost every AI bio company before us had some kind of very tight lab integration with what they were doing—and almost too tight, I think. The lab integration is great. We do a lot of lab experiments at Chai. But the thing that was missing was whether you could actually have some kind of portable AI platform, something that would be generalizable and could be applied to lots of different areas.
If you could do that, it means that your impact could really be taken to the next level. We can take Chai-2, the model that we've just released, and deploy it to hundreds of different projects, thousands of different projects. Chai-1, which we open-sourced, is already being applied throughout the industry to tons of different applications. We don't even know everything it's being applied to because it's open-source.
That was also really important to us. If we were going to see this transformation of biology from a science into more of an engineering discipline—which is ultimately the goal of the company—we needed that kind of platform.
I want to come back to what you said about lab integration as we talk more about the technical approach here. But Jack, you and I met in the context of you being a beloved engineering and product leader at Stripe, coming from the engineering side and looking for the most interesting problems to work on in AI. Why did you decide to work on this versus some of the other things we were talking about, like code generation and such?
As you know, Sarah, I spent quite some time thinking about my next steps and what I wanted to do with my life after the period I was at Stripe. I give a lot of credit to Josh for this. We were good friends going back even to college—we were study buddies at Harvard and in many of the same classes together.
While I was maxing out the computer science curriculum, Josh was doing the same with chemistry, physics, and all the other scientific curricula. We had landed in a lot of the same classes, and as we went our separate ways after college, we really just made a point of keeping in touch every 3 to 6 months. Josh would always talk to me about his research.
Once it became clear that the research Josh and others were doing in this space was really no longer just a toy, but was really going to impact and change the entire industry, that idea became infectious. It sort of became impossible to unsee the future once you had that glimpse.
Although we didn't know until very recently that any of this was going to work—and, of course, there's still a lot left to prove—once you start to grasp the implications of the fact that, over the next few years, we're going to have the ability as a human race to engineer molecules with atomic precision, it's almost hard to work on anything else with your life.
The impacts for society broadly, human health, and not just health—there are a ton of other areas that this will touch, which we can get into—but that is a platform shift in an entire industry. Put that together with the belief or conviction that you might just be able to get it working, and I think it was impossible to say no to working on this in many ways.
There's a breakthrough result in Chai-2. Can you give us a layperson's explanation of what the result was, the model itself, and what you think is the most valuable part?
Sure. Chai-2 is our latest series of models, which are state-of-the-art across a number of different tasks. Specifically, the one we're most excited about is design. What we've shown is that we can design a class of molecules known as antibodies, which are some of the most therapeutically interesting molecules as well.
These account for close to 50% of all recent drug approvals, and 7 of the top 10 bestselling drugs out there are actually antibodies. What we've shown with Chai-2 is the ability to design antibodies against targets that one wants to go after in a small, 24-well plate, in just 20 attempts.
What this means is that we take a target, run our models, and ask the model to design an antibody. We then ship that antibody to the lab. We have about a 2-week validation cycle in the lab, and 2 weeks later we see that roughly close to 20% of these antibodies actually bind to their targets in the intended way.
Chai-2 is a major breakthrough for the field. When we set out on this project, we were actually only targeting a success rate of 1%. That was the company-wide goal for the entire year. The reason we set that goal of 1% is that previous attempts at this problem were maybe successful around 0.1% or even lower of the time—and those are the computational techniques.
If you look at the traditional lab-based high-throughput screening techniques, people are really screening between millions or billions of compounds just to find 1 molecule that sticks. There's a reason we call it drug discovery: It's a discovery problem. It's a search problem. People are really panning for gold in these massive yeast or phage libraries.
Alternatively, you might inject a mouse or a llama. You might wait a couple of weeks for them to get really sick. You might then bleed them, take their plasma, take the antibodies out, and isolate them. This is actually what we did for COVID. We took some humans who had already gotten COVID, took their antibodies out of them, and tried to find one that actually then neutralized the virus.
You can imagine: Not an ideal, or the most efficient, or the most principled process. What we've shown with Chai-2 is that we've been able to increase the success rates in discovering antibodies computationally by multiple orders of magnitude compared to the prior computational state of the art, and by many, many, many orders of magnitude compared to the traditional lab-based alternatives.
What this means is pretty profound for the industry, in our view.
You know, there are 2 ways to look at this. There's, of course, the faster, better, cheaper. This is going to allow us to make drugs against targets and get them turned around faster. But I think the thing that we're really excited about, and what's more important, is the entire class of targets that this will unlock in the future, which have just been inaccessible to previous methods. And I think, in general, the biotech industry is—everybody's a little glum right now.
XBI hasn't done that well over the last 5 years. I think we're in one of the worst markets in biotech over the last few decades. But I think with Chai-2, we're starting to see those first early signs of a real platform shift in biotech, the sort that comes around only so often. We had one in the 1970s with all sorts of new techniques. The idea that in the next 5–10 years there are going to be entire new classes of molecules that we're going to be able to discover, entire new targets that we're going to unlock, markets that we can open up, and therapeutics we can get to patients to really cure diseases that have had no cure before—that's just an incredibly exciting prospect for us.
I want to come back to impact because I think the ramifications here are really huge. But if we just think about the first problem—design—you looked at 52 problems. Why that many? And how do you specify a target? I'm picturing, “Bind to epitope X,” but I'm sure there are other requirements you'd want to have as drug designers.
It's a great question, Sarah. So in the Chai-2 paper, we look at over 50 targets. Most of the existing papers in this area of doing AI for drug discovery are usually looking at 1, 2, or 3 targets. But again, it was important for us, if we were seeing this as an engineering problem, to make sure that this is going to be generalizable.
Imagine you had a new LLM paper and you said, “Oh, I solved 1 problem in the USACO contest. Really, really cool.” No, you need a real benchmark, and you need to actually have that benchmark at scale. You need to have enough problems to convince yourself that the system is working. So that's why, whenever we do these experiments, sometimes we'll try 1 or 2 targets just to make sure there's not a huge bug and make sure not everything fails. But even if everything fails in 1 or 2, the hit rate's 50%; you could have just gotten unlucky. So that's one of the reasons why we decided to do a big benchmark here, to really convince ourselves things are working.
The way we selected the 50 problems—the biology people would laugh at this and engineering people would love it—we actually just went to the vendor catalogs to see what was in stock because we wanted to turn around this experiment quickly. We ordered all of these designs at the same time. So we actually wrote a scraper that would go and see what was in stock. We would go and pick out the protein, and we would go look up what that protein sequence was.
Now, we need to make sure this is held out from training as well, right? So we would take that protein sequence. We would compare it to everything in a database called SAbDab. It's a collection of antibody structures in the Protein Data Bank, and we'd make sure that none of these sequences were in there or even close to anything in there. We removed things that had more than 70% sequence identity—really, things that are a bit different from what we could have trained on. Then we selected those, made our designs, and shipped everything off to the lab.
So we actually think that the 50% is possibly a lower bound, because we might have just messed things up because of how we set up this experiment. We did not think about the biology. These are not necessarily things that are even that useful for therapeutics. Some of these already even have drug programs against them. We were just doing this really from a model-assessment perspective: let's understand how well the model is working. Let's convince ourselves. Let's convince the community that Chai-2 is working.
In terms of applying this to problems, I think now we've got hundreds of people who want to go and try the model tomorrow and apply it to the various drug programs that they're working on. So that was really how we came up with those 50 tasks: let's benchmark this and treat it as an engineering problem.
We have a broad audience for No Priors that ranges from business people to engineers, machine-learning researchers, and some scientists in other fields. What intuition can you give listeners for how the model works under the hood, especially for anybody who might start with some familiarity with structure-prediction models?
Yeah, well, structure prediction is really a key part in making these models work, and it's actually the first thing we did when we started the company: we sprinted to build a state-of-the-art structure-prediction engine. We actually open-sourced the first version of that. It's called Chai-1, and scientists around the world are using that now.
But structure prediction basically gives you an atomic-level microscope, and it allows you to see where atoms are placed in 3D space. So once you can do that and you have this microscope, the next question is: can we start moving those atoms around? We can now start to make changes in a sequence, and then we can see the ramifications of those changes in 3D space.
So the actual design model—you can think of it as—you prompt it with some information: here's a target that we want to design an antibody against. Then the model will try to place these atoms in 3D space in order to satisfy that constraint. We tell the model, here's a target, and I want you to make a molecule that binds to that location. Then the model will generate both a sequence and a structure that fits into that. So that's the high-level intuition for this.
One piece of intuition around that is that you can almost think about structure prediction as the ImageNet moment for the field, where with structure prediction, we are asking a model to go from sequence to a predicted structure, and it's sort of like a classification task. Then design, where you're trying to design binders, is much more like a generative task; that's sort of like Midjourney for molecules, whereas with structure prediction, you're looking to predict the placement of atoms in 3D space. With design, you're taking an existing placement of atoms and you're trying to craft a new set of atoms that is complementary to that original set.
So one analogy that people like to use is that of a lock and a key: when designing a protein or a drug, you have some target, which is your lock, and you're trying to design a key using a generative model that fits that lock. And the way that the models work is actually pretty interesting. They reason quite literally by placing individual atoms in 3D space. Often, they're getting the resolution of these structures—the error—down to less than the width of 1 atom when we look at the error across the entire structure. So when we talk about an atomic-level microscope, you can see now why that might be important for design, because how can you hope to design the key if you can't see the lock?
Yeah, that's completely wild from a precision-of-prediction perspective. If we analogize to LLMs, you have learned grammar, syntax, and semantic capabilities that emerge in the model that you can measure. Is there anything that would be analogous in terms of emergent vocabulary or concepts that you think Chai-2 has?
I think this whole point about the atomic-level microscope is actually that point right there. There's something really deep—we still don't fully understand it—about why these models work. Again, we didn't even know this was possible. Obviously, we tried, so we thought there was a chance, but I think it just tells you something: maybe the signature of how proteins interact with one another is really embedded in the data, right?
We're generalizing to a new setting. So it's not like the model has seen specific binders against the target and then we're just trying to do some in-domain generalization and walk through that space. That's actually quite an impactful application as well, and that's already being done throughout the biotech industry. Our team published work on that years ago already. But I think this really new frontier about generalizing to a new space tells us—and, again, the model is learning something really fundamental about how the molecules interact with one another.
Again, it's able to generalize to problems that look very different in terms of how we would actually organize them in biology. I think the whole question of what we think about as a protein family being different—these targets that we tested on are, again, to a biologist, very “dissimilar” from what we saw during training. But it doesn't seem like the model thinks that way.
We even have a plot in our paper's supplement where we look at an even harder subset—not looking at things with up to 70% sequence similarity to the training data, but actually pushing all the way down to 25%. So we're really looking at tasks that are very different. We saw in training that the success rate was basically the same. The model didn't care. And again, I think that indicates something very profound about what the models are learning here.
My assumption is the same here: Obviously, the fastest path to immediate impact is going to be antibodies in the clinic, or whatever other therapeutics Chai and its partners work on. But it does raise a question: If the model has learned something that, fundamentally, the biology research community doesn't yet know from a principles perspective, will we also learn those rules from these models, or whatever the principles are of structure and interaction? So I think that's super exciting.
Yeah, totally agree. Well, how would you characterize the overall hoped-for impact of Chai-2 in terms of bringing it to industry or your own programs?
It's a great question, Sarah. There are maybe 2 main areas that we can break it down into. The first one is, again, we turn this into an engineering problem and spend months—or sometimes even years—trying to discover some molecule. Now we can actually do it way faster because the screening, if you will, is happening on the computer instead of in the lab.
But the second area that we're actually even more excited about is: How do we actually solve problems that just weren't even reachable with traditional methods? The model is not perfect. It worked in 50% of the targets that we tried. Maybe it would have been more, right, for the caveats we talked about before, but it works in 50% of cases. The failure mode of the models is going to be different than the failure mode in the lab today.
And I think that's really going to be the sweet spot to focus in on. What are the areas that were not possible a few months ago, where now we'll be able to actually generate potential molecules really quickly against those targets? So those are the 2 areas. You have things that you can do today; let's do them a lot faster, a lot cheaper. But I think really the breakthrough opportunities are things that just weren't possible before.
One other thing that we've announced is that we will be opening up access to both academic groups and industry partners. I think when you think about how this space is just going to evolve in the next few years and the amount of opportunity that's out there given this platform shift, there is way too much opportunity for any 1 company to capture alone. Drug discovery itself is just an incredibly resource-intensive process, and I think it would probably be a conceit to assume that we could go after and pursue every target, every program ourselves, even if we wanted to.
And so when we think about impact and think about what is going to move the needle for the company, of course, but also for the world, we think that the way to do that is to go out and bring this to life with a really exciting set of partners. And so we've opened up access. There's an access page on our website which people can go to and fill out. We're currently working through those. We've been inundated with requests, but my hope is that we can really enable quite a few use cases with this and do that quite quickly.
What has the reception been like so far? What is the biggest objection? Because this is a significant challenge to the ideas of high-throughput screening, or even the workflow that innovative pharma and biotechs have today.
Yeah, it's a great question. Usually when these kinds of papers come out—again, people have tried to do this many times—the critique is often, “Does this really work?” You show this on maybe COVID, for example; is this going to work for a case where we have less training data? Are the molecules going to be high quality? Do we really believe the data?
I think the approach we took, benchmarking this at scale, has really helped a lot with that reception. People really appreciated that approach, which has been great. Some of the questions people have are, “Okay, I can already discover drugs. So now I have AI that can do it a lot faster, but does that actually change the kinds of molecules I can work on?”
And it goes back to what we just discussed before. I think there are other folks that are responding to that saying, “No, the transformation here is: How about those projects that didn't work for you, or where you're really struggling today? Now you've got another tool in the toolkit, and you kind of have to use this tool now or you might be left behind.”
So I think that it's been really interesting to see the community digesting this. Of course, a lot of the AI folks are really excited, right? We're getting artificial antibodies before we're getting maybe other breakthroughs we would have expected earlier, but it's overall been really exciting to see that reception.
Our inboxes are just flooding up. The early access has gotten hundreds of people within hours of launching reaching out to us. We just announced, so I think we're still kind of digesting all that. We're a small team, so we're prioritizing early access to the right people, but we're really excited to get the models out there and for them to start solving some really hard problems in the drug discovery space.
Is there an important future for large-scale wet-lab screening? Does it just become a data collection exercise to fill out the distribution for Chai models? Are there areas where you think we'll need that in 10 years or 20 years?
Yeah, I think if you just take the models and then you sample more, you probably will get a better result. So we tested only 20 molecules per target in the paper—up to 20 molecules. If you were to do 10 times that, 100 times that, orders of magnitude more, you probably just get into spaces with better molecules.
So the machine-learning model is probabilistic. It's like using ChatGPT: If you're trying to solve a math problem and then you look at the top 1 response, or if you look at the top 10,000 responses, you're going to get a better result if you look at the top 10,000. You can't really do that with a product experience on ChatGPT. I'm not going to look through 10,000 math responses. I won't even know which one is correct.
The cool thing with a lab, actually, is we could just test all 10,000 of those in the lab. So I don't know if you have to. But that's definitely something that is, I think, going to be tested out with these models. And I think the future of high-throughput screening and how they interact with the models—the question is still open.
But I expect that people will be creative and will find ways to actually take the best of AI and marry that with the best of biology to push the bounds forward.
And just to add one thing to that, there's a whole host of really amazing CROs and other players with this incredible expertise running those traditional methods. And to Josh's point, we have many, many companies asking us, “Can you run this not just 20 times, but can you run this 100,000 times, even if it's going to work in 20, because I just might find something better?” And that something better can result in a better drug.
That could be the difference between getting a patient an antibody that requires an injection or something that requires subQ dosing, for example. And so I think with these tools you can sample research space sort of ad infinitum, and that marrying of traditional technique and models will actually hopefully get us into areas of the space where we can just find better products for patients.
I want to ask one more question generally about predictions for biotech, and then I want to talk about the future of Chai as well. What do you think biotech looks like 25 years from now? I realize that's a ludicrous question to anybody working in AI, where you're like, “Hey, I didn't know this was going to work at all last year.”
As I mentioned before, there is a lot of doom and gloom in the biotech industry right now due to macro factors, with rates where they are and the long-term investment cycles that are required to make biotech viable. There is just a real pessimism in the industry right now. It's sort of hard to imagine what the market will look like in a couple of decades.
And I think that it's moments like this—breakthroughs like this—which give us these flashes of light and these reasons for just immense optimism about the future of this industry, not just in terms of improving timelines and reducing costs, but also in terms of fundamentally enabling those new products.
And so if we think ahead over the next 25 years, we've gone from a less-than-0.1% success rate to a close to 20% success rate in a year. Well, who's to say that in another year that can't be a 50%-plus or even a close-to-100% success rate?
I think if you see our mini-protein results, we are, I think, close to 70% on those with picomolar affinities—really, really tight binders—for every single target that we tested. So all 5 targets we tested worked, and 70% of the designs that we ordered worked. I think that there's no reason that other classes of molecules, those success rates can't be that high as well.
And I think once you have that, you really enter this era where you sort of have a computer-aided design suite for molecules, in a way that we have maybe SolidWorks for mechanical engineering, or we have Photoshop for creatives. That entire software suite will exist for biology.
I think the implications of that—the ability to design, program, and understand the interactions between atoms and molecules at the most fundamental level—are pretty vast and should just give us a lot of hope and excitement about what's about to happen.
We were just talking last night about whether we should get baseball caps saying “Bullish on Biotech” on them. I think this is one of those special moments that can really shift opinions. We've heard from many others writing into the company that this has really shifted their opinion.
If you think about going from antibodies to obviously better success rates, and then also to other therapeutics, is there a difficulty hierarchy we should have in our minds? Or is it just unexplored space in terms of enzymes and peptides, small molecules, and other domains?
Yeah, it's actually a lot more than just success rates. There are lots of properties that need to be optimized for a molecule. Finding a drug is like looking for a needle in a haystack, and I think we've really passed through massive swaths of that sequence space with Chai-2, right, by really focusing on the things that bind.
That's where a lot of the search space has to be searched in the lab today. We're going deeper into other properties as well. Let's make sure these antibodies can be manufactured well, and let's make sure that they can be really stable. So there are lots of other properties that we're excited about. Stay tuned for that.
The other thing is that there are next-generation antibody formats as well. What we predict will happen is that people probably won't be as interested in the clinic in things like monoclonal antibodies. These are antibodies that are hitting, for example, a specific epitope on a protein.
But if we can make antibodies much faster and more easily, you can imagine a future where, if I want to hit a target, I can choose 2 different parts of that target, make 2 different proteins that are hitting them—basically 2 different primitive antibodies—and bring them together. This is called a biparatopic antibody: 2 paratopes, basically 2 different antibody interactions. That kind of thing is going to become a lot easier to do.
I think these days there are a lot of trade-offs that get made in biotech around risk on your target, risk on your discovery process, and how hard it's going to be to make your molecule. I think AI is going to raise the bar across the board. I think the “Bullish on Biotech” movement that Jack is announcing here as well could represent something significant.
There's a lot of risk in biotech right now. There's a lot of crowding on the same kinds of targets. The risk actually starts to go down in terms of discovering some of this stuff. Maybe there's still clinical risk if you try something that's totally new, that people haven't done before, but we've just opened up the aperture of opportunities that can be pursued here. That's something that I think is really exciting.
There's still a lot more work to do for us to validate that all this is going to be possible, but I think the pace at which the field is moving gives us a lot of optimism for what's going to be possible next. Maybe I can just share one anecdote about why we are so optimistic.
We had a partner come to us as we were in the process of building these models. We didn't even really know—we hadn't had back our first few batches of data—so we didn't know if it was going to really work yet. This partner had been working on this problem for a few years, and they had a team of, I think, 5 to 10 people working on it.
They estimated that, fully loaded, all of those people might have set the company back, with the experiments they had done, maybe $5 million to $10 million. It was a problem where they wanted to build a molecule that cross-reacts against 2 different species: both a human form and a cyno, or monkey, form of this protein.
When they put this molecule into animal testing, they didn't want it to fail because the monkey has a slightly different version of that protein than the human does. They were really struggling to get this to work for whatever reason, so we put this into the model and prompted the model to design for these 2 targets at the same time, not just 1 target. You can imagine that this is a slightly more sophisticated challenge than just designing against 1.
We ordered only 14 sequences for the lab, and I think 4 of those were hits to humans. 1 of those was a hit to the cyno, and 1 of them actually overlapped and hit both. That one now allows us to move forward with that program and gives us a whole host of diversity around that molecule to explore as well.
First of all, that's very cool. Second, I think it's interesting that a lot of industry observers would say that the bottleneck in pharma, and the expense in pharma, is clinical, not discovery.
I think you're pointing to the fact that we can design for the clinic, right? It's intuitive, but it's an argument from people bearish on biotech or concerned about the ability to make progress in programs and reduce costs for any given successful drug: if discovery had less risk, as Joshua Meier was pointing out—which is a huge claim—then the entire industry is more efficient and more effective. That's the hope.
Yeah, and I think we've got a lot of reason to be optimistic. I also don't want to oversimplify things. There are lots of other things that go into making a drug. There are capital markets that go into this, and there's tons of clinical risk. This is really just the tip of the iceberg, but we're really excited about the progress that this could represent.
I want to ask strategically where Chai invests from here. You talked about other attributes that you want to be able to design in Chai models, but if we just look at this generically as an AI model company, where do you think the defensibility is?
There are 2 key areas of investment for the company. Firstly, what comes out of these models—these just aren't drugs yet. They're hits. They're antibody hits. But there's a lot more work to be done to actually turn these into viable molecules that we can put into humans.
We have early data, which we put in our preprint, to suggest that these molecules actually have a lot of the properties that one might want from a drug. But we need to do a lot more characterization and assays to convince ourselves that we can do that.
I think the next stage beyond that is actually designing entire drug candidates in zero-shot, right out of the models. A few months ago, we might have said this was a pretty futuristic idea, and nobody in the company was really talking much about this. But once you see these results and grapple with the implications—the fact that we can get antibody hits in just 20 attempts—there's no reason that we couldn't generate entire drug candidates in that same number of attempts.
I think there are going to be some key investments there. The model right now is a model; it's not really a product. It is a product, and it's certainly useful today, but there's a lot more that product can become with more investment into making sure that we can optimize all the therapeutic properties that people care about.
Then, of course, there's the entire interface and software layer around that to make this really easy to use, and the platform that goes around supporting it. If you want to hit 2 targets, how do you design a molecule that hits both? How do you specify that in the software?
This is going to be a sufficiently advanced piece of software. It's going to become as advanced as Photoshop over time. As we build that out, I think we're going to need to make some really core investments into the engineering and the product to ensure that we are building software that we ourselves and others will really love to use.
Yeah. One thing to add on to that as well: we released Chai-1 open source. We thought of it as a model, and I think Chai-2 is a lot more than a model, right? It's become a product. It's actually a bigger pipeline that comes together to make this happen, and it also becomes trickier to use these models.
Protein folding is one thing: you put in your sequences, and you get out a structure. Design is a different story, right? Specifying the prompt on its own—we did that programmatically in the paper to assess this thing at scale. But a scientist who wants to use this to initiate a drug discovery program probably isn't using a script to come up with that prompt. They're probably going to be really thoughtful about it.
I think that's why investing in the product layer here is really important. Not to mention, it's only going to get more complicated from here, right? As we start to support more advanced drug modalities, as various properties come online, and as we actually show some early evidence of this in the white paper, you might want to optimize for multiple proteins at the same time.
Sometimes it's actually a good time to be a sick mouse. In order to have a human drug, it usually needs to work in animals as well. Sometimes drug programs actually get stuck there. It's like, “Okay, guys, we either have a mouse drug or a human drug. It's really hard to get both.”
There are actually some cases where people have to discover 2 different drugs. They have what they call a surrogate antibody: “I'm going to make the mouse version. I'm going to study that and convince the FDA that this mechanism works.”
But you’re even taking a risk. You’re saying, “Maybe this molecule works slightly differently.” We literally show that example in the paper on optimizing. We don’t do mouse; we actually do monkey—monkey and human together. But you can throw other species into that as well.
Sometimes you’ve got the opposite problem: I want to hit this protein, and I don’t want to hit this other protein. We’ve got some early evidence as of late that that’s possible as well. These sorts of things—the prompts are just a lot more complicated, and it means that you need to have the right product. What happens when you start doing those experiments in the lab? We want the models to learn from that and then help us really be a copilot, driving the next stage of those designs as well.
All of this, again, is more than just the models. It’s really thinking about those workflows as well. It’s even about just getting the word out to people and having them think about this as a new tool in their stack. What happens if you’re an antibody engineer and you’ve been doing things in a certain way for the past 30 years, and now there’s a new paradigm in discovering drugs? That itself is actually a problem that a company needs to solve. So these are all different areas that we’re investing in right now.
That actually begs a question I was going to ask you. If you’re an antibody engineer or a biologist today, what advice would you give them, let’s say, given that they believe you about how much is going to change and that this CAD for biology—a software suite—is coming into existence? What should they learn or be good at? What should they go study?
Well, number 1, get access to Chai-2. Number 2, figure out how to get your prompts right and actually take full advantage of it. And then I think, number 3, start dreaming about the new possibilities.
You know, it’s interesting: We’ve talked to a lot of antibody engineers since starting the company, and we’ve been alluding sometimes to what we’re doing here. Sometimes you do the market-research question; you ask, “Suppose you had a 1% success rate for antibodies—what would you use that for?” The conversations are changing now. First of all, it’s not 1%; it’s 10%, and people see that it’s working.
I think that creativity is really being unlocked, even in ourselves, right? When people are thinking about the answer to that question, there’s always some big doubt in your mind. It’s like, “Ah, it’s a hypothetical question.” Your neurons are not activating in the same way as doing something with it. It was the same thing with LLMs. Imagine asking someone 5 or 10 years ago, “If we could predict the next word in a sentence perfectly, what would you do with that?” It’s actually very hard to imagine until you start playing with the models.
Even our team internally, now, without sending anything to the lab, we can again choose some targets, choose a prompt, and generate stuff against them. You start to look at the generations that are coming out of the model and you’re like, “Oh, wait. I can actually solve this problem by choosing the right epitope on a target, choosing the right part of the target.” These 2 targets are different. Sure, we have an engine that the model can optimize for 1 or optimize for both, or selectivity-optimize for 1 and not the other. But you can actually get a lot of that by choosing your prompt in a smart way.
So let me hit part of that protein that is actually quite different between the 2 things, or quite similar between the 2 things. These are the sorts of realizations that, in retrospect, are quite obvious, but they don’t really hit you until you actually start to use a product like this yourself. So I think people are—once they get their hands on this, I think they’ll start to dream of the new possibilities.
I think it just really raises the bar. The people who are most excited about that are often these antibody engineers and these biologists. A lot of the work that they’re doing today is painstaking, and they’re not the biggest fan of these slow feedback loops and these intractable problems, because many of them that we speak to are just really motivated to solve a particular task.
And so you give them—I’m an engineer. You give me a tool which says I have to write less code. I love that. I can now think more about system design and architecture, more complex products, and all these other things. But it’s really going to raise the bar for a lot of these people, and I think people are only really now starting, as Josh said, to think through all the possibilities.
I was on calls with people a matter of a few weeks ago where people were saying, “When do you think this is going to happen?” And they’d say, “Oh, not for 3 to 5 years. This is a really futuristic idea.” Then, a couple of weeks later, you show them what they have, and they sort of fall off their chair.
And so there’s going to be a sort of joint effort with us alongside these real domain experts to actually figure out these key application areas, because biology is so vast and so complicated that there is so much knowledge that so many of the practitioners, the specialists, have that no one company will ever possess. Which is why we’re so excited to go out and partner with people to really bring this to life.
I want to ask a couple more questions, specifically about company building, before we run out of time. Maybe, Jack, I’ll start with you. You’re an amazing engineer, and you guys also have a very software-oriented team working on biological problems. Some of those people come from long-term research in that space in particular. But for yourself, Jack, as you said, you’re a software person. How do you get up to speed on the bio area to go do leading work?
Well, I think it’s 2 things. First of all, ramping up on any new field is always just a total fight. You have to get to the frontier, have read the right papers, and be knowledgeable about the areas that you need to learn. You just have to put your head down and push through, and there are waves of excitement and misery in that experience, but you can get there fast if you really set your mind to it.
And I’d say the second part is that surrounding yourself with just the most incredible team is the best thing that you can do, far beyond anything that you can learn by yourself. We have certainly the most special group of people that I’ve ever worked with within the company. Our co-founders, Matt McPartland and Jack Bo, are just rare talents. And then the entire team beyond that: some of the former heads of AI at other drug-discovery companies, some of the top open-source contributors. The team is so multitalented. It’s small—it’s around a dozen people—but mighty. And I think, as we’ve seen in other areas of AI, small but mighty teams can go a really long way these days.
And so, there are actually surprisingly few people on our team even with a computer science degree. Josh himself got a chemistry degree. Alex got his PhD in physics, and a whole host of others. But this work is so interdisciplinary that really having that breadth of knowledge across biology, chemistry, physics, artificial intelligence, computer science, and engineering—it really takes a village, and everybody is learning from each other every day because of just how vast that subject matter is that one has to have a command of.
I think we’ve also just benefited from such immense focus as well. Everyone has been so passionate about trying to solve this problem, and I really credit that to being a huge reason why we were able to achieve it. And we’ve also got a team that, because of that focus, is very engineering-centric as well. So if you look at the whole team, we have a very research-oriented team right now, but everyone is a stellar engineer as well and takes that very seriously.
It’s not that everyone’s solving their favorite pet problem; we are all going after the same problem and solving that together. And even just 10 people solving a problem together, there’s a lot of code being written every day. You have to be very thoughtful about how that all comes together and interacts, and I think especially in our next phase of growth for the company, as we start to invest more and more in product and the velocity around that and getting this into folks’ hands, that’s just going to become even more important.
How do we make sure that the latest research breakthroughs that we’re shipping internally are actually making their way into partners’ hands? That’s something that, again, we are very thoughtful about at Chai and take very seriously.
Yeah. I also remember, Jack, in our office at Conviction, debating the merits of dev containers with some of your scientist teammates at the very beginning of the company. Both of you, from the beginning, talked a lot about platform investment, and so I actually think that’s a little bit unconventional in terms of such a research-oriented team—to say, “We need to make this platform investment.” Can you talk a little bit about that?
Yeah. I’ve gone through the experience of going from 0 to 100 on large engineering products before. I worked on Stripe Link, which was kind of a multiyear project, and again Stripe Capital, where engineering teams scaled from 0 to 25 or 50 people by the time we were done there.
Same for Link, maybe more. And I think you just learn that unless somebody is really taking care to keep the entire system in their head and is an effective technical steward of the architecture, things just devolve, and the entropy of software takes over and slows your rate of progress to zero because nobody can get work done anymore. Somebody needs to keep the entire system in their head and understand the interaction between all those components, and make sure that people who are working on individual subcomponents of your codebase have to minimize the amount of context they need to load into their heads to understand how to accomplish that task.
These are just the principles. It’s pretty basic: simplicity and modularity, but making sure that’s a practice and a kind of cultural practice, and that everybody’s on the same page about investing in that and that people aren’t cutting corners. They see it as their responsibility to lay the groundwork for the next person.
This is doubly hard to do in deep-learning codebases because often, if you introduce a bug or write a regression, you won’t know for weeks that it has shown up. It’s sort of terrifying because you could spend $1 million on a training run with a bug that crept in 4 weeks ago. We’ve literally had to do this in Chai’s history: we’ve had to go and bisect Git history, launch training runs with a sort of binary search to identify a small enough range of pull requests to identify a bug, then go to that pull request and identify the bug.
And I think it’s those sorts of experiences—and the cost of finding that bug. I’m not sure if it was millions of dollars, but it was certainly tens of thousands of dollars of compute time to go back and find it—that make rigor, and engineering rigor, such an important practice in the company. Some people are surprised to learn that even though we do deep learning, we’re pretty rigorous about writing unit tests for everything.
These basic software engineering practices are actually sorely lacking from most research codebases. Bringing in some of those basic principles has allowed us to move very fast—not just in the short term, but in a way that should give us a mechanism to compound on that investment over time.
Well, and it’s overall very aligned with your mission of turning biology from science to engineering, right? It makes sense that it would go through the core of the company’s practice, too. I have 2 more questions before we run out of time here.
The first is: You talk about the expense of training experiments. What’s your decision framework for how quickly to scale compute or parallelize experiments here?
Yeah, we tried to set up the company in a pretty scrappy way. Actually, when we were getting started, we should have talked about this as well. We were based in San Francisco, but it wasn’t even clear that the company would be in San Francisco when we started, and back then we hadn’t really raised capital for the company yet. We were kind of using free compute credits from the cloud providers.
For us, it’s just about being laser-focused on solving the problem and really making the case: Why are we doing something? I think that if it’s reasonable, we’ll go and invest in it. In an engineering problem, if it’s clear you’re seeing signs of life, if you’re seeing some scaling law or whatever it is, let’s go as fast as possible and make that work. But let’s also not get distracted by scaling something out if we’re not convinced that it’s going to work yet.
I think that scrappy culture on the “where are we spending?” side goes hand in hand with making really fast progress because it means we have a high bar for where we’re spending our time. Everyone on the team works extremely hard. There are people in the office at all times of day and all times of night. It’s pretty beautiful to see that.
We work hard, but I think we also work really smart, and I think you have to do that to make progress with how fast the field is moving right now.
You now see signs of life. You’re very bullish on biotech. That also means, given that you’re going to try to scale to support demand from the industry and your own efforts, who are you looking to hire now?
We’re really hiring across all functions right now. We’ve made some really big breakthroughs here on the AI research side. As we take that to the next level and try to get Chai-2 in front of the right partners, we’re hiring for product engineering, antibody engineering, business development, and account executive. There’s a whole host of roles that are open on our site right now.
Again, this work is extremely interdisciplinary, and we really wanted to build this in a thoughtful way so that we can make Chai-2 as useful as possible for the industry.
Well, thanks for doing this, guys, and congratulations on progressing the frontier of AI discovery.
Thanks so much for having us on, Sarah. It’s been really fun.
Thank you, Sarah.
Find us on Twitter at No Priors Pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briers.com.