Patrick Hsu
I want to make science faster. Our moonshot is really to make virtual cells at Arc and simulate human biology with foundation models. We can figure out how to model the fundamental unit of biology—the cell—and then from that we should be able to build.
Jorge Conde
My goal is to really try to figure out ways that we can improve the human experience in our lifetime. There are a few things that, if we get them right in our lifetime, will fundamentally change the world.
Erik Torenberg
Patrick, welcome to the podcast. Thanks for joining.
Patrick Hsu
Thanks for having me on.
Erik Torenberg
I've been trying to have you on for years, but finally I could get your time.
Patrick Hsu
Here I am. I'm excited to do it. It's going to be great.
Erik Torenberg
For some of the audience who aren't familiar with you and your work at Arc and beyond, how do you describe your moonshot? What are you trying to do?
Patrick Hsu
I want to make science faster, right? We can frame this in high-level philosophical goals like accelerating scientific progress. Maybe that's not so tangible for people.
I think the most important thing is that science happens in the real world. If it's not AI research, which moves as quickly as you can iterate on GPUs, you have to actually move things around and transfer liquids from tube to tube to make life-changing medicines. These are things that take place in real time. You have to actually grow cells, tissues, and animals.
I think the promise of what we're doing today with machine learning in biology is that we could actually accelerate and massively parallelize this. Our moonshot is really to make virtual cells at Arc and simulate human biology with foundation models.
We'd like to figure out something that feels useful for experimentalists—people who are skeptical about technology. They just want to see the data and see the results, and for it to be the default tool that they go to when they want to do something with cell biology.
Erik Torenberg
Okay, well, hold on. Let's back up. Why is science so slow in the first place? Whose fault is that?
Patrick Hsu
Whose fault is that? Now, that is a long one. We should get into it. It's really multifactorial.
Erik Torenberg
Okay.
Patrick Hsu
Right. It's this weird Gordian knot that ultimately comes down to incentives. It comes down to how science funding can be better, but it's also about how the training system works, how we incentivize long-term career growth, how we try to separate basic science work from commercially viable work, and generally the space of problems that people are able to work on today.
I think things are increasingly multidisciplinary. It's very hard for individual research groups or individual companies to be good at more than 2 things. You might be able to do computational biology and genomics, or chemical biology and molecular glues, but how do you do 5 things at once? That's increasingly hard.
We really built Arc as an organizational experiment to try to see what happens when you bring together neuroscience, immunology, machine learning, chemical biology, and genomics all under one physical roof. If you increase the collision frequency across these 5 distinct domains, there would hopefully be a huge space of problems that you could work on that you wouldn't otherwise be able to.
Now, obviously, in any university or any kind of geographical region, you have all of these individual fields represented across different campuses. But people are distributed, and you want everyone together.
Erik Torenberg
Okay. But if I may, a university—I would have thought a university was an attempt to bring multiple disciplines under one roof. You're saying it's not. It's too diffuse.
Patrick Hsu
It's across an entire campus.
Erik Torenberg
Okay. So the physical distance literally creates inefficiency.
Patrick Hsu
That's part of it. And I think the other part is that folks have their own incentive structures. They need to publish their own papers. They need to do their own thing and make their own discovery, and you're not really incentivized to work together.
I think in many ways, in the current academic system, a lot of what we've done is to try to have people work on bigger flagship projects that require much more than any individual person, group, or idea.
Erik Torenberg
Yeah, that's cool. So the original hypothesis for the Arc Institute is that if you can bring multiple disciplines together to increase the collision frequency, as you said, and if one could remove some of the cross incentives that may exist in traditional structures, the combination of those 2 things will make science faster.
Patrick Hsu
Yeah, these are absolutely part of it. We have 2 flagship projects: one trying to find Alzheimer's disease drug targets, and the other to make these virtual cells.
I think it's not just the people and the infrastructure, but also the models will hopefully literally make science faster—that you could do experiments at the speed of forward passes of a neural network if these models could become accurate and useful.
Erik Torenberg
Mm-hmm. Yeah. So that will be one thing that shortens the length of discovery: You compress the time discovery takes naturally by just throwing technology at the problem, at the risk of oversimplifying.
Patrick Hsu
Well, we're techno-optimists here, though.
Erik Torenberg
We are.
Patrick Hsu
Yeah.
Erik Torenberg
Why has AI progressed so much faster in image generation and language models than in biology? If we could wave a wand, where are we excited to speed certain things up?
Patrick Hsu
To be honest, it's a lot easier, right? Maybe that's a hot take.
Erik Torenberg
I mean, technology is easier than biology.
Patrick Hsu
Natural language and video modeling are easier than modeling biology. Correct?
To some degree, if you understand and learn machine learning, you have already learned how to speak and you already know how to look at pictures. So your ability to evaluate the generations or predictions of these models is very native.
We don't speak the language of biology. At very best, we speak it with an incredibly thick accent. When you're training these DNA foundation models, I don't speak DNA natively. I only have a sense of the types of tokens that I'm feeding into the model and what's actually coming out.
Similarly, with these virtual cell models, I think a lot of the goal is to figure out ways that you can actually interpret the weird, fuzzy outputs that the model is giving you. I think that's what slows down the iteration cycle: You have to do these lab-in-the-loop things where you have to run actual experiments to test with experimental ground truth. Increasing the speed and dimensionality of that is going to be really important.
Erik Torenberg
Yeah. How much of this is the fact that you talk about how we speak biology poorly, or with a very thick accent? How much of this is that, if you're training on an image, we can see the image, and so we can see how good the output is?
What about all the things in biology that we can't see or don't even know exist yet? How can we create a virtual cell? Maybe we should come back to what a virtual cell model is, by the way, for the lay audience. But how can we create a virtual cell model when we're not even sure that we understand all of the components that are in a cell and how they function?
Patrick Hsu
People talked a lot about this in NLP as well. There's this long academic tradition in natural language processing, and then it was just weird, unintuitive, and intensely controversial that you could feed all this unstructured data into a transformer and it would just work.
We're not saying this will just work in all the other domains, including biology, but I think there is this controversy around what it means to be an accurate biological simulator. What does it mean to be a virtual cell?
It's true. We can't measure everything. We can't measure things like metabolites at really high throughput with spatial resolution. There are going to be different phases of capability where initially they model individual cells, then they model pairs of cells, then they model cells in a tissue, and then they model cells in a broader, physiologically intact animal environment.
Those are length scales and layers of complexity that will aggregate and improve over time. I think the other kind of nonintuitive thing, in many ways, is the scaling laws that you get in data and in modeling.
I'll give you an example. There's a lot of discussion in molecular biology about how RNAs don't reflect protein and protein function. We don't have proteomic measurement technologies that are nearly as scalable as transcriptomic measurement technologies today, particularly at single-cell resolution, but we're getting there.
You can layer on certain nodes of protein information and add them on top of the RNA information. But in many ways, the RNA representation is a mirror. It might be a lower-resolution mirror for what's happening at the protein layer, but eventually what is happening in protein signaling will get reflected in a transcriptional state.
For an individual cell, this may not be very accurate. But when you imagine the massive data scale that we're generating in genomics and functional genomics, you start to gather tremendous amounts of RNA data that will reflect what's happening at the protein level in some sort of mirror echo. That can be the case for metabolic information as well, and so on.
So it's a low-pixel image, but if we can get zoomed out far enough, we'll get a sense of what's going on. You have to bet on what you can scale today, right? We're able to scale single-cell and transcriptional information today. We're able to add protein-level information over time. We'll need spatial information, spatial tokens, and temporal dynamics as well.
I kind of bucket things into 3 tiers: invention, engineering, and scaling. There are certain things today, biotechnologically, that are scale-ready, and then there are things that we still need to invent. That's part of why we felt we needed a research institute to tackle these types of problems. We weren't just going to be an engineering shop trying to scale single-cell perturbation screens, right? That would be interesting, but in 3 years would feel very dated, I think. There's a lot of novel technology investment that we're making that we think will bear fruit over time.
Erik Torenberg
Yeah. Can we flesh out the virtual cell concept? Why that's the ambition we've landed on, and what's it going to take to get there? What are the bottlenecks?
Patrick Hsu
I would say the most famous success of ML in biology is AlphaFold, right? This solved the protein-folding problem: when you take a sequence of amino acids, what does the protein look like? It's pretty good. It's not perfect—it certainly doesn't simulate the biophysics and molecular dynamics—but it gives you a sense of what the end state is with 90%+ accuracy. That's the AlphaFold moment that people talk about, right? Anytime you want to work with a protein, if you don't have an experimentally solved structure, you're just going to fold it with this algorithm.
We kind of want to get to that point with virtual cells as well. The way that at Arc Institute we're operationalizing this is through perturbation prediction, right? The idea is that you have some manifold of cell types and cell states. That can be a heart cell, a blood cell, a lung cell, and so on. You know that you can move cells across this manifold. Sometimes they become inflamed, sometimes they become apoptotic, sometimes they become cell-cycle arrested, they become stressed, or they're metabolically starved—they're hungry in some way.
If you have this representation of universal cell space, can you figure out what perturbations you need to move cells around this manifold? This is fundamentally what we do in making drugs, right? Whether we have small molecules, which started out as natural products from boiling leaves, or antibodies, when we injected proteins into cows, rabbits, and sheep and took their blood to get those antibodies, we were basically trying to get to more and more specific probes. We had experimental ways to cook these up. Now we have computational ways to do zero-shot design of these binders.
Ultimately, what you're trying to do with these binders is inhibit something and, by doing so, click and drag it from a toxic, gain-of-function, disease-causing state to a more quiescent, homeostatic, healthy one, right? The thing that's very clear in complex diseases, where you don't have a single cause of the disease, is that there's some complex set of changes. There's a combination of perturbations, if you will, that you would want to make to be able to move things around.
People talk about this classically as polypharmacology, right? But I think we're moving from, “This thing happens to have a whole bunch of different targets by accident,” to, “We have the ability to manipulate these things combinatorially in a purposeful way.” To go from cell state A to cell state B, there are these 3 changes I need to make first, then these 2 changes, and then these 6 changes over time, right? We want models to be able to suggest this.
The reason why we scoped virtual cell this way is because we felt it was just experimentally very practical. You want something that's going to be a copilot for a wet-lab biologist to decide, “What am I going to do in the lab?” We're not trying to do something that's like a theory paper that's really interesting to read, where the numbers go up on an ML benchmark, but you practically can decide what are the 12 things that you're going to do in the lab in 12 different conditions. You want to actually test them, right?
That's how we enter the lab-in-the-loop aspect: model predictions to experimental measurements to improved or RL or whatever model predictions again. The goal is to be able to do in silico target ID, where you can basically figure out new drug targets and then figure out the drug compositions you would need to actually make those changes.
If we could do that, we could make a new, vertically integrated, AI-enabled pharma company, right? I think that's obviously a very exciting idea today, but in many ways the pitch and the framing of these companies precede the fundamental research capability breakthroughs. That's what we're really invested in at Arc Institute: just making that happen, along with many other amazing colleagues in the field, to make this possible for the community.
Erik Torenberg
If the goal is—I'm going to oversimplify it for you—if we wanted to get to the AlphaFold moment, where it gives you a useful folded structure 90% of the time, to use your data point, we wanted to take that comparison in the virtual-cell model and say, “Okay, 90% of the time, if I ask the model, ‘I want to shift the cell from cell state A to cell state B,’ it's going to give me a list of perturbations, and 90% of the time those perturbations in fact result experimentally in shifting from cell state A to cell state B.”
How far away are we from that AlphaFold moment for virtual cells?
Patrick Hsu
I find it helpful to frame these in terms of GPT-1, GPT-2, GPT-3, GPT-4, and GPT-5 capabilities, right? I think most people would agree we're somewhere between GPT-1 and GPT-2. A lot of the excitement was that we could achieve GPT-1 in the first place—that you could see a path, with scaling laws of some kind, to make successive generations where capabilities would improve.
But with our Evo DNA foundation models that we developed at Arc Institute with Brian He, one of the things that we've seen is that these genome-generation models are, quote-unquote, blurry pictures of life. We don't think that if you synthesize these novel genomes, they would be alive, but we don't think that's actually impossibly far away. We'll just have to follow these capabilities.
We're taking a very integrated approach to attack this problem, right? You need to curate public data, generate massive amounts of internal private data, build the benchmarks, train new models, and build architectures—doing these things full-stack. We'll just attack this hill climb over time.
Erik Torenberg
What's the GPT-3—I'll say GPT-3—moment going to look like? By that I mean a public release that alters the public's conception of what's possible here from a capabilities perspective and also inspires a whole new generation of talent to rush into biology.
Patrick Hsu
The good thing with biology is that we have a lot of ground truth. There are entire textbooks that describe cell signaling and cell biology and how these things work. Even without a virtual-cell model at all, if you went into ChatGPT or Claude and asked it some question about receptor tyrosine kinase signaling, it would have an opinion on how that works, right?
I think you would want the model to be able to predict perturbations that are famous, canonical examples of biological discovery. For example, if you loaded into the model an iPSC—an induced pluripotent stem cell state—or a human embryonic stem cell state and a fibroblast cell state, could it predict that the 4 Yamanaka factors would reprogram the fibroblast into a stem-like state? Essentially, could it rediscover something that won the Nobel Prize in 2009? That would be 1 really classic example.
Then you could do the inverse. If you have a stem cell, can it discover Neurogenin-2, ASCL1, and MYOD? Can it find differentiation factors that will turn that into a neuron or a muscle cell, and so on? These are classic examples in developmental biology.
You could also use this to try to discover or recapitulate the mechanism of action of FDA-approved drugs. For example, we know that if you inhibit HER2 in breast cancer cell states, you would get this type of response. Or it could predict certain clones that will be more metastatic or more resistant and lead to minimal residual disease.
I think there are lots of biological evals that you can add onto these models over time that are really tangible textbook examples, as opposed to what the early generations of models do today, which is very quantitative things like mean absolute error over the differentially expressed genes and stuff like that. Those are ML benchmarks, and we want to increase the sophistication into something that you could explain to an old professor who has never touched a terminal in their life.
Erik Torenberg
By the way, you talk about textbooks as ground truth. Do you think we're going to find that a lot of the textbooks are wrong?
Patrick Hsu
I would say textbooks are compressed, right? So, for example, when you look at these classic cell-signaling diagrams of A signaling to B, which inhibits C, right? That's a very two-dimensional representation of our understanding of a complex system.
Erik Torenberg
Right. Right. Right.
Patrick Hsu
Yes, textbooks are what they are. They represent the corpus of reliable knowledge, but everyone knows that there are an incredible number of exceptions. Part of what discovery is is finding new exceptions, right?
Erik Torenberg
Why don't you talk about the difference between simulation of biology and actual understanding, and what it would take to actually be able to model the extremely complex human body?
Patrick Hsu
Some people don't like the phrase “virtual cells” because it sounds too media-friendly. It's not rigorous enough, right? But I've always found it funny that many people are okay with “digital twins” and “digital avatars,” which talk about modeling biology at a way higher level of abstraction. I think virtual cells, if anything, are way more scoped and rigorous than modeling a digital twin or avatar.
I think these are useful words because they describe the goal and the ambition. In the long run, we don't care about predicting the perturbation responses of an individual cell at all, actually. We want to be able to predict drug toxicity, predict aging, and predict why a liver cell becomes cirrhotic when you repeatedly challenge it with ethanol molecules or whatever, right? These chemical or environmental perturbations should be predictable.
I think you just have to layer on the complexity. Why are we so worried about modeling entire bodies over time when we can't do it for an individual cell, where we broadly believe that this is a fundamental unit of biological computation, if you will? Let's just start there, just like you have to start with things like math, code, and language modeling—things that are easier to check. You can build to superintelligence over time.
Erik Torenberg
Yeah, I think that makes sense. That's a very laudable, ambitious goal: that we can figure out how to model the fundamental unit of biology, the cell. From that, we should be able to build—
Patrick Hsu
Like in early AI, we just started with language translation and basic NLP tasks. This was long before the tremendous, ambitious scope that we have today, and hopefully we can mirror that type of trajectory if we're lucky.
Erik Torenberg
It seems that biotech and pharma have been shrinking in terms of their rate of growth. What's it going to take for these innovations in science to be reflected in business models and in growth for the industry?
Patrick Hsu
A lot of these biotech startups would initially try to sell software to pharma companies, and then they would realize, “Oh, wow, we're competing for SaaS budgets,” which aren't very large. Now they're realizing, “Oh, we have to compete for R&D budgets.”
I think there's this narrative from the current generation of these companies that our biological agents will compete for R&D budgets and replace headcount or something like that, just like we're seeing with AI agents across different verticals. Whether or not that will pan out depends on whether these things meaningfully allow us to build drugs more effectively in the pharma context. I think that's the most important thing in this industry.
We believe in virtual cells not just because we think they will be a fountain of fundamental mechanistic insights for discovery, but also because, in the case of success, they could be industrially really useful. We'll have to see over time.
If 90% of drugs are failing in clinical trials, that means two things, and you're not sure what percentage of each. One is that we're targeting the wrong target in the first place. The second is that the composition—the drug matter that we're using—doesn't do the job. It's not clear for each individual failure which one it is, whether it's both, or what proportion of each. We'll have to sort that out over time.
You can imagine that even in the case of success, when we have 90%-accurate virtual cells, you'll probably end up with suggestions like, “Okay, now you need to target this GPCR only in the heart, but not in literally any other tissue.” We don't have the drug matter that can do that today. That's also why you probably need research to figure out novel chemical biology matter that allows you to drug pleiotropic targets in a tissue- or cell-type-specific way.
Part of why biology is slow is because there's this Russian nesting doll of complexity in terms of understanding, perturbation, and safety. The crazy thing is that the progress in just the short time that I've been doing this is insane. I did my PhD at the Broad Institute during the heyday of developing single-cell genomics, human genetics, CRISPR gene editing, and so many other things.
The early-2010s papers on single-cell sequencing would have 20 cells or 40 cells. At Arc, in the next relatively short amount of time, we're going to generate 1 billion perturbed single cells. How's that for Moore's law?
Erik Torenberg
Yeah, that's remarkable.
Patrick Hsu
Yeah.
Erik Torenberg
Jorge, I want to hear your answers to a couple of these questions too, as the lead of our bio practice—both on the GPT-3 moment, what that could look like, and also whether you think it's G1's or sort of building off that, or if it's going to be something different. I'm also curious what it's going to take for the science to reflect itself in the business and for the industry to grow.
Jorge Conde
Yeah. I'll take the second one first, if I could. In terms of where the industry is right now, one of the big challenges we have is, as Patrick describes very nicely, discovery is hard and it takes time. The failure modes are exactly as you described. Oftentimes, when drugs fail—which they do 90% of the time in clinical trials—it's because we're going after the wrong thing, or we made the wrong thing to go after the right thing. Those are the 2 failure modes, and that happens all too often.
A lot of what Patrick is describing is going to improve our hit rate, or our batting average, in figuring out what to go after and then making the right thing to go after.
Erik Torenberg
And we have to be good at making them in the first place.
Jorge Conde
And we've got to make them too. Exactly. That bottleneck is necessarily important. We want that bottleneck to exist. I'm not suggesting we've got to remove it, but are there ways to reduce the cost and time associated with getting through the bottleneck of human clinical trials?
We talk about all of the various stakeholders when you're making a drug. There are the companies, there's of course the science that supports the company that's trying to commercialize a product, and there are the regulatory agencies. Everyone is trying to ensure that, first and foremost, we can discover and commercialize drugs that are safe and effective for humans.
That middle part—actually getting through that bottleneck—is hard to speed up in a very obvious way. You can increase the rate at which you enroll clinical trials. You can use better technology to change the way we design these clinical trials, so maybe they can be faster or shorter. But some of them just have a natural timeline you have to go through.
If you want to demonstrate that a cancer drug promotes survival, guess what? It's going to take some time to demonstrate a survival benefit. If you want to do a longevity drug, that is by definition a lifetime trial in terms of length. A lot of these bottlenecks are really hard to get through.
So, what helps the industry? I think there are a couple of things that help the industry. One is capital intensity. Hopefully, that will go down over time as technology gets better. Capital intensity is something that our industry faces. In some ways, it looks a little bit like AI now, in terms of the cost of training these models, but the capital intensity is very, very high. That has not come down. We have to get the success rates up to bring capital intensity down.
The second thing is: where can we compress time? Good models can help us compress early-discovery time.
We still haven't seen—and I think it's coming, but it hasn't happened yet—artificial intelligence or other technologies massively compress the amount of time it takes us to do the clinical development, the clinical trials, the enrollment of patients, all of those things. We're seeing some interesting things coming. We haven't seen the payoff there yet.
The third thing is, if we can make better drugs going after better things, the effect size should be higher, so the answer should be obvious sooner. If we can get those 3 things right—reduce capital intensity, compress timelines, and effectively increase effect size in some very tough, intractable diseases—that is what I think fixes the industry.
And from where we sit at the early stage, in terms of being early-stage investors, the reason why that helps us is, if the capital intensity goes down and the value creation goes up, it becomes easier to invest in these companies in the early days because you get rewarded for coming in early. The problem we have right now is that, for most companies, you're not seeing rewards happening when there's value inflection. So you come in early, you bear the brunt of the capital intensity, and even if a company's successful, that success isn't reflected in the valuation. So we're not seeing the step-ups that you see in other parts of the industry. And that's just really, really hard from an investing standpoint. So I think we need to see those various factors addressed for this space to really get fixed, to use your word.
Patrick Hsu
Yeah, that was great. I have a lot to add to this.
Erik Torenberg
Please add away.
Patrick Hsu
A few simple observations. The first is, the amount of market cap added to Eli Lilly and Novo Nordisk based on the development of GLP-1s is over $1 trillion. I mean, Novo's stock has decreased a lot, so let's say $1 trillion. That's more than the market cap of all biotech companies combined over the last 40 years.
One of the interesting corollaries of this is that, when we have a 10% clinical trial success rate for preclinical drug matter, you tend to circle the wagons a bit and try to manage your risk. The way that you do this is you try to go after really well-established disease mechanisms, where if I develop new drugs that go after well-understood biology, they should work the way that I hope they will in the trial. That's really, really expensive and costs a lot more in many ways than the preclinical research.
The problem with this is you go after very well-validated disease mechanisms, but with really small patient populations. So then the expected value of this is actually relatively low. One of the things that we've seen with GLP-1s is the value that you can create when you go after really large patient populations. I think that has really increased the ambition of the industry, both from the investor and from the drug developer side, and that's something that we should keep our foot on the gas for.
Jorge Conde
Look, I would argue the trend on that is positive. You're absolutely right: the demonstration of the value that has been created with the increasing use of GLP-1s and the value transfer that's gone to companies like Lilly—which I would argue is very merited, right?—because they've cracked an endemic social problem in terms of managing diabetes and eventually helping manage obesity, is remarkable. There's a lot of value that goes to that because they tackled and cracked a very, very challenging problem for society beyond just science.
That's great, and I agree with you: the prize—the juice—needs to be worth the squeeze, right? A lot of biotech has been around going after the low-hanging fruit because it's low risk and we have to eat today, right? So you go get it, and you push off the big, ambitious indication, the large population, or the really tough-to-crack disease.
Erik Torenberg
But I do think we're seeing more and more of that. By the way, we can get into some of these genetic medicines, but some of these genetic medicines are going after some of the hardest problems—the things that you quite literally couldn't address but for editing DNA. I think that's incredibly remarkable, laudable, and frankly inspiring.
Patrick Hsu
Yeah.
Erik Torenberg
But the fundamental elements of the industry have to work so the capital formation is there to support those kinds of things. Right now it's hard because of the issues we talked about before. Fifteen years from now, we're back in this room. We've barely escaped being part of the permanent underclass, and we're reflecting on the GPT-3 moment, or maybe the legacy of GLP-1s beyond where they are now.
What do you think it could be? I'm curious to get your take on what you think is going to be the technological breakthrough that we're going to point back to and say, “Oh, this is really what set it all off.” Or do you think it's going to be a multifactor combination?
Jorge Conde
Yeah, look, I think it's going to go back to where we started this conversation. GLP-1s as a drug are 4 decades in the making, or something like that. These are not overnight successes.
But I do think what we are going to see more of—and our hope is that when you combine the fact that we're getting better at understanding what to target and getting better at designing medicines to hit those targets, by the way, in a whole array of new, creative ways—we're going to see a lot of progress. We have small molecules, the natural products that we got from boiling leaves, as you said earlier. We're getting really good at designing smarter and better small molecules that do new things, that function in ways that they didn't before.
We've gotten quite good at designing biologics or proteins, with a lot of help from things like AlphaFold that help us understand how proteins fold. We're going to get a lot better at designing some of the more complex modalities, like gene therapies or gene editors.
When you can do that and combine that with our ability to hopefully use things like virtual cell models to really understand what to go after, we're going to have drugs—I would hope and expect that the industry will continue to bring forward drugs—that have very large effect sizes for very difficult diseases that hopefully affect a lot of patients. If that's true, then we'll start to see some of these really, really difficult diseases that affect all of society get tackled, hopefully one by one by one by one.
We have obesity, we have metabolic disorders, we're dealing with cardiometabolic disease, and we're starting to see interesting, promising things happening in neurodegenerative diseases. If we can tackle cancer, or at least several cancers that have now begun to be treated more like a chronic condition than the death sentence that they were in the past, the more we see of that, the more I think that value to society will accrue over time. This should be an industry that is extraordinarily valued by society and, candidly, by the markets. We have to deliver.
Patrick Hsu
If we play this out right, and let's say these AI models work, and you can make a trillion binders in silico that will yield exquisite drug matter, we still need to make these things physically and test them in animals and, hopefully, predictive models, and then actually in people. I think that will increasingly be the bottleneck in many ways.
My friend Dan Wang recently released a book called Breakneck, which talks about the U.S. and China, the difference between the 2 countries, and their philosophy—the way they approach markets.
Erik Torenberg
We're a country of lawyers or a country of engineers.
Jorge Conde
That's right. China is an engineering state, right? It's kind of a polity of folks who have engineering degrees. You need to build bridges and roads and buildings, and these are the ways that we solve our problems.
Whereas, I think, from the first 13 American presidents, 10 of them practiced law. From 1980 to 2020, all Democratic presidential candidates—both vice president and president—went to law school. You see the echoes of that in the FDA and the regulatory regime, and in all the bottlenecks that people talk about when developing drugs stateside.
Increasingly, you see folks thinking about how we can run Phase 1 trials overseas, build data packages that we can bring back domestically for Phase 2 efficacy trials. I think that's an interesting direction, but it's not enough. We need to figure out these 2 bottlenecks: the designing and the testing.
Erik Torenberg
Yeah.
Patrick Hsu
Even if we can solve the designing part.
Jorge Conde
Oh, I agree. Yeah, yeah. That's the bottleneck. We joke about it: you have to get a molecule that can go first in mice, then in rats, then in monkeys, and then in man. It takes a long time, and it's so hard to compress that. So when you do, you should make the journey worth it, right? When you fail on the other end of that, that's obviously horrible.
Patrick Hsu
And so finding ways to make sure that when you walk that path, it'll be a successful journey as often as possible is what this industry desperately needs.
Erik Torenberg
Mhm. AlphaFold solved a protein-folding problem, but what didn't it solve—drug discovery? Or, more broadly, what would it take to get a drug discovered? What's sort of the bottleneck on the tech side, at least?
Patrick Hsu
On the tech side.
Erik Torenberg
Yeah, maybe another way to ask the question is that I always ask founders a version of this question—the AI ones that are like, “Oh, we're going to do AI for life, for drug discovery.” My question that I always like to ask founders is: give me examples where you think AI is hyped, potentially overly hyped; where there's real hope—what do we expect, what's next; and where we already see real heft. So if I asked you, in AI, where is there hype, where is there hope, and where are we seeing heft today?
Patrick Hsu
I would say there's hype in toxicity prediction models.
Erik Torenberg
Okay, so that's the idea that we will say, “I'm going to show you a molecule, and you're going to tell me”—the model is going to tell me if it's going to be toxic or not?
Patrick Hsu
That's right. There's heft in anything to do with proteins, right? Obviously, protein binding, but increasingly in protein design, right? I think there's real heft there. And then where there's hype is in multimodal biological models, whatever that means, right? Pick your favorite layers: it could be molecular layers, it could be spatial layers. Actually, I would say there's also heft in the pathology AI prediction models, like automating the work of pathologists and radiologists.
Erik Torenberg
That's a powerful use case. Sure.
Patrick Hsu
Yeah. And there's a lot of stuff where you don't have to train weird biology foundation models, and you can write regulatory filings and reports and things like that. That's impactful and important.
So now, go back to Erik's question: Why hasn't AI turned out drugs yet? I think that was your question, right? AI for drugs is one of these weird things where everyone who works in the industry is trying to claim that their drug is the first AI-designed molecule. I feel like, increasingly, in just a few years, this will just be a native part of the stack. Just like we use the internet and we use phones, we're going to have AI in all parts of the stack. It's just going to become a native part of everything that we do.
So why hasn't it worked yet? It's this long, multifactorial process that we've been talking about today. There's designing, there's the making, there's the testing, there's the approvals side of it. And I do think safety and efficacy, as the 2 pillars in the industry, are the 2 things that we need to get right. We need to be able to figure out faster ways that we can predict whether or not a molecule will work and if it's going to be safe or not.
And there are ways that AI can operationalize this. If you designed a small molecule, you could now computationally dock it to every protein in the proteome and see if it's likely to bind to off-target molecules. You can use this to tune binding selectivity and affinity. Those might be ways to predict safety and efficacy, right? And how well will that work? Well, that's a feedback loop that we'll have to actually test in the lab.
And that's part of what's slow: the testing takes real hours, days, months, years. And that's really why we've picked, at Arc, the virtual cell models as our initial wedge, because we think it can integrate a lot of these different pieces.
Erik Torenberg
In Dario Amodei's essay, “Machines of Loving Grace,” he predicts, among other things, the prevention of many infectious diseases and the doubling of lifespans, perhaps in as soon as the next decade. What's your reaction to the bullishness of his essay and some of his predictions?
Patrick Hsu
I think the core intuition that Dario had was the idea that important scientific discoveries are independent, or they're largely independent. And if they are statistically independent, then it would stand to reason that we could multiparallelize. If we had models that were sufficiently predictive and useful, you could have not just 100 of them but millions, billions of these discovery agents or processes running at a time, which should compress the timeline to new discoveries and turn it into a computation problem.
I think that is a very futuristic framing for something that is actually very tangible today. And if we can have virtual cell models at work, for example, that can start to do these kinds of things that we've been talking about and help us, we can have molecular design models. We can have docking models. We can then ask, when you bind to this thing in this cell versus all the other off-target proteins, whether a cell will be corrected in the right way.
These kinds of layers of abstraction and complexity start to get to things that feel very tangible through drug discovery. If you could actually traverse these steps reliably and in sequence, you could start to see how you can get the compression right. And so I think in the long run, this should be possible.
Erik Torenberg
One of the core positions in building a good virtual cell model is that we are feeding it all the relevant data.
Patrick Hsu
The right data. Yeah.
Erik Torenberg
The right data. What if we're missing a core element? What if we just haven't discovered the quirk or whatever? We just don't know what we don't know, and therefore what we're feeding the model is fundamentally or importantly incomplete.
Patrick Hsu
I think that's almost certainly true. It seems almost obvious that we're not measuring many of the most important things in biology, right? And you can, of course, find many important exceptions for any of these measurement technologies. In biology, we ultimately have 2 ways to study it in high throughput: imaging and sequencing, right? But there are so many other types of things that you would care about that those things aren't necessarily going to do at scale, right?
And that's really why I think the stuff that we're talking about—the RNA layer as a mirror for other layers of biology—is something that we've spent a lot of time thinking about. There's a difference between a mechanistic model and a meteorological simulation type of model. For example, if you want to predict the weather, you can build AI models that will predict whether or not it will rain next Tuesday. It won't explain physically or geologically or whatever why and how that happens. But as long as it knows if it's going to rain next Tuesday, you're probably happy, right?
Similarly, with a virtual cell model, it may not tell me literally why—just like AlphaFold doesn't tell me literally why the protein folded this way and how—but it just told me the end state, and it was reasonably accurate. I think that would already be very important.
Erik Torenberg
Shifting gears a little bit, we've been talking about science and biotech, but in addition, you're an elite AI investor more broadly. So I want to talk about where your investment focus is right now, just as it relates to AI more broadly. Where are you excited? Where are you spending time? Where are you looking forward to?
Jorge Conde
Oh, yeah. My goal is to really try to figure out ways that we can improve the human experience in our lifetime. I think of it like this: if I think about the future that we're going to leave to our children, there are a few things that, if we get them right in our lifetime, will fundamentally change the world and how we live in it.
I think synthetic biology, obviously, is one. Think GLP-1s, things that improve sleep, things that can improve longevity. These are all things that are easy to get excited about. I think brain-computer interfaces are another area where we're going to see really important breakthroughs over the decades to come.
And then I think the third is robotics, both industrial and consumer robotics, that allow us to basically scale physical labor in interesting ways. You can kind of see how each of these 3 things, even in the medium cases of success, really kind of changes the world. And so I'm very interested in helping make these kinds of things possible.
In the techno-optimist vision of the world, there are a few different types of scarcity. It's very easy, when you do research, to come up with important ideas. The hard thing is to tackle them in the right timeframe. Writing futuristic sci-fi things is not that hard; being able to actually execute on them in the next 5 years or 8 years is much, much harder.
Academic discovery is littered with plenty of ideas that are interesting and important, but long before their time. In many ways, the story of technology development is trying to use new technologies to solve old problems. Most of our tools are for productivity, whether that's the Industrial Revolution, the computing revolution, or the current AI revolution. We're trying to do the same stuff.
And so there's a relatively small set of very powerful ideas. New technologies give us new opportunities to attack them, and there's a set of people and teams that are going to be positioned to be able to do that.
They need to have technical innovation and then an intuition about product and business. In the RPG dice roll of the skills you get in these 3 domains, people start at different base levels. You might have an incredibly technical founder who doesn't know how to think commercially, or someone who's just natively a very commercial thinker who doesn't have very strong product sense, even though they could sell the crap out of it. I think these 3 broad categories of capabilities need to come together in a way that you can allocate capital at the right times to make these ideas possible in a really differentiated way.
This thing literally wouldn't happen if we didn't get these people together and fund them at the right time in the right way. That's really what motivates me, and these are the kinds of things I've been excited about backing: longevity companies like Newman, BCI companies like Nudge, and robotics companies like The Bot Company. These are examples of things that I think must happen in the world and therefore should happen. How do we find the right people and the right time to go on the Fellowship of the Ring hunt?
Erik Torenberg
Yeah. If it's not too difficult, I want to ask Jorge a question adapted to these additional spaces—robotics, sort of BCI, and longevity. The 3 questions, I believe, were: what's overhyped, where do you see opportunity or a path, and what's got heft already?
Jorge Conde
I think the cool thing about agents generally is that they do real work. Compared to SaaS companies that came before, agents replace real productivity. They have a lot of errors today, and I would say computer-use agents will probably trail coding agents by maybe a year. But it's coming, and we'll follow the trajectory as these go from doing minutes of work without error to hours to days.
I think you're going to get a completely different product shape as we march through that across legal, back office, medicine, healthcare, whatever. We'll follow that as an industry, and that's going to be really exciting. I think that's where we're going to see real heft, because most of the economy is services spend; it's not software spend. The reason we're all excited about this stuff is that it can attack the services economy.
I would say there's a tremendous amount of hype; that's no doubt. The hype is in the model capabilities, and we're working with an architecture that dates back to 2017. If you look at the history of deep learning, every 8 years there's something really new, and in 2025 we're really overdue for some net-new architecture. I think there are lots of really interesting research ideas bubbling up that could do that.
In many ways, there's a set of really interesting academic ideas, especially from the golden age of machine-learning research, from 2009 to 2015. There are so many interesting little arXiv papers that have 30 citations or less. As the marginal cost of compute goes down year on year, I think you're going to be able to take all of these ideas and scale them up. You don't see the scaling laws when you're training them at 100M or 650M parameters, like back then, but if you can scale them up to 1B, 7B, 35B, or 70B, you start to see whether or not these ideas will pop.
I think that's very exciting because there's going to be a lot of opportunity for new superintelligence labs to do things beyond what the established foundation-model companies are doing today. In addition to these research teams, these are in many ways becoming applied AI companies. They need to build product for all kinds of different enterprises, do RL for businesses, and make money—or build coding agents and make API revenue. That's important, and I think it's a timely race to survive today.
But I'm just very bullish on the research of, say, Sakana AI, which was founded by one of the authors of attention is all you need, Ian Jones. They're doing incredibly interesting stuff on model merging and how you can have evolutionary selection of different models. I think the opportunities here in the long run to move beyond just RL gyms, for example, and also to figure out new ways to learn and find reward signals, are going to be really exciting.
Erik Torenberg
I think it's a great place to wrap. As we're gearing toward closing, is there anything upcoming for Arc that you'd like us to know about? Anything you want to tease? Anything for people who want to learn more—what should they know about?
Patrick Hsu
So, AlphaFold, in many ways, came out of a protein-folding competition called CASP, or Critical Assessment of Structure Prediction. We created our own Virtual Cell Challenge at virtualcellchallenge.org, where we have $100,000 prizes sponsored by NVIDIA, 10x Genomics, Ultima Genomics, and others. It's an open competition that anyone can enter, where you can train perturbation-prediction models, and we can openly and transparently assess these model capabilities both today and in subsequent years, following them to get to that ChatGPT moment.
I'm extremely excited about this. We'd like more people to train models and apply to the challenge—both bio-ML experts and engineers in any other domain. I want this thing to exist in the world. Hopefully, we're important parts of making that happen, but I'd be happy if someone does it.
Patrick Hsu
Yeah.
Erik Torenberg
That's an inspiring note to wrap on. Patrick, Jorge, thanks so much for the conversation.
Patrick Hsu
Thanks so much, guys. Appreciate it. Thanks.
Jorge Conde
For having me.