Nathan Labenz
Today my guest is Jassi Pannu, an assistant professor at Johns Hopkins who recently co-authored an important paper calling for the creation of access-control systems meant to prevent the dissemination and misuse of functional biological data from which AI models could learn extremely dangerous capabilities, such as the modification or even de novo design of highly contagious and deadly viruses.
We begin with an overview of the biosecurity landscape today, including how new viruses are detected, how patient data is aggregated and analyzed in the context of a new threat, and what the pipeline from DNA sequence to vaccine candidate looks like today. The good news is that we are able to design new vaccines amazingly quickly, at least for viruses that are similar to others we've seen. But there is unfortunately a lot of bad news as well.
In 2012, for example, 2 research groups independently published results showing that wild-type bird flu, which already had an estimated 60% fatality rate but couldn't spread between humans, could become mammal-to-mammal transmissible with just 5 mutations. Such gain-of-function research has been broadly defunded since the COVID-19 pandemic, but it does remain legal, and visibility into the experiments private labs are conducting is low.
Governments, Jassi says, aren't likely to develop bioweapons capable of causing pandemics for the simple reason that, short of vaccinating their populations in advance of an attack, they can't realistically expect to control them. But with AI capabilities crossing critical thresholds month by month, the threat from extremist groups and even lone actors is quickly moving from theoretical to deadly practical concern.
Consider that Jeffrey Irving, chief scientist at the UK AI Safety Institute, recently highlighted for me that today's frontier models can troubleshoot laboratory experiments from a cellphone picture better, on average, than PhDs. And in just the 10 days or so since we recorded this conversation, we've seen Andrej Karpathy's autoresearch framework demonstrate that AI agents can run and make research progress for days on end.
Even more to the point, Anthropic just reported that Claude Opus 4.6, when faced with a benchmark challenge it couldn't solve, spontaneously located the full benchmark data set on Hugging Face and then figured out how to decrypt the solutions, which were encrypted in the first place to prevent the answers from leaking into training data. It did this all in order to get a single question right.
With reasoning AIs already capable of spontaneously overcoming such barriers to information, I think we should expect that future research agents will find and exploit any signal-rich data that exists anywhere on the internet. And with the smallpox sequence and the horsepox synthesis protocol already published online, and biological data poised to grow superexponentially in the coming years, we have real reason to worry and ample cause to get serious about implementing data controls before the situation gets truly out of hand.
Again, though, there is good news. Recent work by the teams behind the Evo and ESM families of biofoundation models showed that strategically excluding key data sets, such as the DNA sequences of viruses that infect humans, dramatically reduced models' performance on dangerous tasks while leaving their desirable capabilities intact.
This means that the vast majority of biological data can remain open source and open access. Indeed, Jassi and co-authors' proposal for a biosecurity data-level framework, which echoes the existing Level 0 to Level 4 biosafety framework for physical labs, would subject only an estimated 1% of data that connects pathogen sequences to dangerous properties to any additional restrictions.
Even then, structures such as trusted research environments, which allow researchers to run code on data without transmitting that data from its secure location, would still support valuable research. Once again, despite my personal history as a lifelong techno-optimist libertarian who broadly believes that data wants to and ought to be free, I find myself eager to support these control measures.
Of course, that's not the only opportunity we have to improve biosecurity. Toward the end, we also discuss the broader defense-in-depth strategy that biosecurity experts recommend: delay, deter, detect, and defend. This includes mandatory pre-synthesis screening of sequences by DNA manufacturers, investment in wastewater monitoring and other passive global pathogen surveillance, and practical frontline defenses like PPE stockpiling and far UV sterilization.
All of this is in everyone's shared interest, but it does require leaders to see beyond the current news cycle for long enough to make it happen. I certainly hope they do, but I also recommend taking individual action where you can, both to improve your own personal safety and to support the consumer market for biosecurity products.
My wife and I, for example, at our friend Jeff Kaufman's recommendation, recently purchased an Aero Lamp far UV light for use in my son's hospital rooms throughout his cancer treatment. I'd welcome additional suggestions for other products that could help us minimize disease burden today while also serving as a sort of private insurance against pandemics, if anyone has any recommendations.
For now, I would simply emphasize that, by default, we are fast approaching a world in which a rapidly growing number of people—and perhaps autonomous AIs as well—will have the ability to create deadly, transmissible, self-replicating viruses that could dramatically alter the trajectory of human history. It really does seem like we should do something about it.
With that, I hope you are properly alarmed by this scary, but solutions-oriented conversation about the sorry state of biosecurity and the rapidly rising threat from bio-savvy AI systems with Johns Hopkins Professor Jassi Pannu.
Today my guest is Jassi Pannu, an assistant professor at Johns Hopkins who recently co-authored a call for controls on biological data. We've heard quite a bit about the possibility that AI systems of various kinds could create new sorts of biorisks, and this is one attempt to put some controls in place to hopefully cut that off at the pass before it becomes a major downstream issue.
I want to get into that from a bunch of different angles, but I think it would be helpful to take a step back and lay some foundations. People who follow this feed know a lot about the AI side. They probably don't, on average, know nearly as much about the current state of play when it comes to biological data writ large.
I thought it would be helpful, and I'm actually very curious about some of these things I realized I didn't know as I was preparing for this. For starters, if you could take us into the moment that we had not so long ago—and hopefully won't have again, but very well might—when all of a sudden there is a new outbreak of something. Something we've never seen before, we don't know what it is, people are concerned, and patients in a particular hospital or city are showing up with concerning symptoms.
What happens? How do we turn that initial small patient population into actual knowledge about what we're dealing with?
Jassi Pannu
Yeah, I think that's a great place to start. To take a step back, it's important to realize how society comes to the conclusion that there is a new virus circulating. What is our mechanism for making that decision? We don't have a global or national alert system for this kind of thing in the same way that you have a radar system for ICBMs.
Right now, we largely rely on symptoms. Patients get sick, they visit a hospital, doctors get concerned, and they run a series of tests. Doctors always start with tests that are common, like influenza, RSV, and rhinovirus. We're just collecting samples by swabbing their noses and sending them to the hospital lab.
When these tests are negative and the patient is very sick, then we'll go ahead and run further tests. We learned during COVID-19 that this whole process is pretty lengthy. It works well for familiar viruses, but it does not quickly detect new viruses.
In the case of influenza, the situation is a little bit unique. We've had several influenza epidemics in the past, so we do actually have a global system to collect new influenza sequences. But it operates on a contribution model. The national lab has to send that sequence up to the global repository, and it's an active submission process.
We don't have a passive global alert system, and we can talk about the benefits of having something like that later. But once the community has decided, “Okay, it looks like there's a cluster of patients with a new virus, and we need to figure out what this virus is,” that's when we start to do more exotic tests, like metagenomic sequencing, to try to figure out the sequence of that pathogen.
That's what happened during the COVID-19 pandemic. We think, based on research, that COVID likely emerged around late November to December of 2019, but it was fully sequenced in January. The decision to share that sequence—the way it happened—was that a researcher sequenced the virus. This was a researcher in China, and they publicly shared the sequence.
They didn't request permission from the Chinese government, which is pretty different from how it might have happened for an influenza virus.
Jassi Pannu
And that's when we kicked off a process of designing diagnostics and trying to scale those.
Nathan Labenz
I recently had an experience where my son had cancer. People know about this if they've listened to this feed. I went through this process of sharing data with the broader medical establishment. Somebody, I think it was a big part of their job based on how much time they spent with me, came by and said, “Hey, I'm from the sort of data-sharing world, and I'm here to answer all your questions and get 1,000 signatures on 1,000 pages so we can hopefully use this data for the betterment of all humanity.” And I said yes to all that stuff.
I've still been getting stuff in the mail asking me to share data again with particular studies or groups or whatever. I'm wondering if it's different when it's a pathogen, because what I'm sharing in my son's case is information very specific to him, whereas you could think about it differently if it's a pathogen. It's not your pathogen, right? It's the world's pathogen in some sense. How much consent or opting in do patients have to do, if any, to get this data from them and into the higher levels of analysis?
Jassi Pannu
Yeah. What you're pointing to is the fact that we have very different systems for dealing with individual-level risk versus societal-level risk. We have a really strong set of protections around making sure that your son's data isn't inadvertently shared, that you consented to everything, and we're thinking about the privacy risks of all of that data. That's really to protect your son, the individual patient.
But when it comes to societal-level risks about a pathogen, where basically the whole globe is your patient, we don't have as good mechanisms for figuring out how to protect society from those risks. We don't have as good mechanisms for protecting the data that could lead to those consequential outbreaks.
Again, just to focus on the clinical data, that very much depends on where you are, what country you're in, and what system you're operating in. In the US, as it sounds like you've experienced, our system is very fragmented. It's different within different states and different networks, whether it's a public community hospital or a private hospital. That fragmentation results in different people having to ask you for permission over and over again, because there isn't a single unified repository for all of that data.
It is a bit different in other countries. If you look at the UK, they have the National Health Service. What the UK has been able to do is create a research platform called OpenSAFELY. This was something that was stood up during 2020, and it provides researchers access to clinical data for 95% of the UK population. It's really interesting because it doesn't require giving the data to the researchers. It forces the researchers to come to the data.
The code comes to the data set, and you can submit your code to the platform. You never have to see the private data. It's secure the entire time, and there's been a lot of positive reception to the OpenSAFELY platform. But we don't have an equivalent in the US where everything's unified in that way.
Nathan Labenz
Yeah, that's interesting. We'll come back to the different levels of protection or restriction that certain data sets should have in your mind. But let me stay on the response narrative for a minute.
I know that it sounds like it was actually maybe longer between when the virus first emerged and when it was first sequenced than it was from when it was first sequenced to when the vaccine was initially designed. I've read a couple of these magazine pieces that tell the story of a 48- to 72-hour period where certain researchers got the sequence, were able to do some analysis on it, identify a particular protein, which I think was the spike protein, and then plug that into an mRNA platform.
My understanding was that the final vaccine that I got wasn't that different from the initial design that they put together in a couple of probably pretty long days. I think it was even as early as February of 2020. Has that pipeline changed at all? Obviously, there have been so many different machine-learning models that have come out over the last 5 years that predict shapes better and predict what's going to bind to what better, but that's pretty hard to beat in terms of timeline. Would you say that if this were to happen again today, that process would look much different from how it looked then?
Jassi Pannu
Maybe first I'll talk about what happened during the COVID-19 pandemic and those different timelines, and then what we can reflect on regarding how this will look in the future with AI. You're completely correct that during the process of designing COVID-19 vaccines—and I'll focus on the mRNA platform vaccine as a primary example—the design process, the computational steps where you were looking at the spike protein and working backward to figure out what the mRNA sequence was that you were going to use for your vaccine design, was very fast.
But that was clearly not the bottleneck. There were many other steps that took much longer. So, yes, the process of knowing there was an outbreak, sequencing the pathogen, and figuring out what pathogen we were dealing with took quite a bit of time.
There was also a whole body of research that researchers relied on, which I think doesn't get enough airtime. Once researchers realized that we were dealing with SARS-CoV-2, they were able to look at this body of research that had been done on SARS-CoV-1, which was the original virus that caused severe acute respiratory syndrome many years prior. It didn't spread into a global pandemic, but there were patient cases that were isolated across the globe, and we actually managed to prevent that from becoming a global pandemic.
It was that research that resulted in researchers knowing that the spike protein was extremely important. They were then able to say, “Okay, we know the spike protein is important. Let's computationally design the vaccine.” That part was very fast.
Then it was about turning that into an actual vaccine: all the clinical components, the regulatory process, the different clinical trials that you have to go through, and then distributing that vaccine globally. Those steps took a lot longer than the actual computational design.
I think that now, with advancements in AI models, people are very optimistic about being able to design new proteins, antibodies, and vaccines. I would say that there's still going to be a bottleneck in scaling in the physical world. Clinical trials still remain a huge barrier, and scaling and deploying vaccines across the globe is a huge barrier. Figuring out if AI can speed up those steps will be really useful, and it's perhaps the neglected component of the overall pathway.
Nathan Labenz
One thing I've learned in looking at the question of security at frontier-model developers in the AI industry is that, in looking at responsible scaling policies, there are a lot of levels to the game of security. The different levels seem to correspond to different actors that you might be concerned about and how hard it would be to prevent them from doing bad stuff.
For the likes of DeepMind, OpenAI, and Anthropic, the general consensus seems to be that if a determined nation-state actor, such as China, were to want to steal the model weights, it would probably be able to do it. There's not too much that could be done to prevent that. How would you map the threat landscape when it comes to the bio-risk side?
Is it a similar thing where we have random crazy people versus somewhat more sophisticated groups of people, all the way up to nation-states? What do you think is reasonable to expect we can actually stop with all of the measures that we'll talk about potentially developing?
Jassi Pannu
Yeah, it's a really important question. Within the realm of biosecurity, there are a lot of different threats that people refer to when they're thinking about chemical and biological weapons. There are things like toxins—small-molecule toxins and protein toxins. Then there are organisms that are not transmissible between humans, things like anthrax, where we have reliable countermeasures, such as antibiotics that work against those things.
Then there's the more extreme end of pandemic threats: pandemic viruses that are novel, that we've never seen before, and for which we don't have diagnostics, therapeutics, or vaccines. When you think about that spectrum, the potential consequences of a pandemic virus are far higher than those of many of the other threats. This is all pretty obvious to us now after having lived through one.
It's also interesting that a pandemic virus is not a particularly desirable weapon for a nation-state. It's not targeted, and it's not easy to protect your own population. You'd have to design a vaccine and vaccinate your entire population. It's hard to do that without someone noticing.
So I think, in general, nation-states are not the primary actors that one is considering when thinking about pandemic threats. It is more likely to be people who are not motivated by rationality: smaller groups, terrorist groups, and potentially lone actors. Those are the folks that people are really concerned about.
That's why a data-control mechanism is most likely in the interests of all countries. I would say that China is equally invested in making sure there isn't a future pandemic as the United States is. So I'm hopeful that there can be some international cooperation—or, if not cooperation, at least an acknowledgement that data controls benefit both the US and China and nation-states globally.
Jassi Pannu
That's what I would say. In terms of data controls, how can you prevent those kinds of actors that I outlined—lone actors and smaller groups—from getting access to your data sets? I think that controls are meaningful there. Currently, the default is sharing that data publicly and making it available for anonymous access. It's extremely easy to access, and even putting minor barriers in place would make a difference.
The other important thing to consider is that we want to make sure defenders, or people who are advancing countermeasures research and virology research, have access to that data while limiting access to malicious actors. Controls can do that differential privileging, where you're privileging defensive use cases and limiting offensive use cases. You can track who's using it, give access to that crowd, and limit access to others.
Nathan Labenz
So, can we map out the data landscape as it exists today? I want to do this for both the data landscape and the models that are obviously spawned from the data. On the data side, I've heard it said many times that you can find the smallpox sequence on the internet. Then there's the question of: If that's true, why hasn't that turned into a crisis already?
I've heard various accounts, and I'd be interested in yours. More broadly, that's a known sequence for a known problem. You can tell me, but it strikes me that it's a small enough amount of data that's probably pretty hard to control or clean up from all the places where it might already have been replicated. It's a little bit hard for me to imagine a world where that's been scrubbed so thoroughly that somebody who wanted to find it wouldn't be able to, but maybe you have a plan for how we could get there.
But if we expand the scope of data of concern, there's tons of biological data in general, right? I know most of that you're not looking to restrict. So how would you draw the—I don't know if they're concentric circles or not—from the narrowest category, the smallpox sequence, which we probably shouldn't be passing out too freely, to somewhat larger categories, and then beyond that, everything that would be fine? How would you characterize those classes of data?
Jassi Pannu
Yep. Let's start with just the general, broad categories of data. The most abundant biological data that currently exists is sequence data. We are currently swimming in petabytes and petabytes of sequence data, a lot of which we don't know the function of. It's become extremely easy to collect and sequence that kind of data.
There's something called Carlson's curve, which is the equivalent of Moore's law for biology and DNA sequencing, and it has actually shattered Moore's law because it has become exponentially cheaper to do DNA sequencing. That's resulted in a lot of passive DNA sequencing and collection—just sequencing everything. That kind of data is available in government-supported repositories, things like GenBank, which is supported by the NCBI, part of the NIH in the US. GenBank alone has more than 40 petabytes of unannotated, raw, poor-quality, frankly, DNA sequencing data.
That is a large part of why there are efforts to build AI models using that data, because it's abundant. When we think about other types of data, there are protein sequence databases, and then there's the Protein Data Bank, which formed the basis of AlphaFold and contains protein structure data.
That data was collected by hand over the years. I'm sure a lot of us have heard the story of painstaking efforts to do experiments to figure out protein structure, where one grad student would spend their whole PhD project on it. That data set is actually very small. It would definitely fit on a thumb drive. It's less than 1 terabyte of data.
Those are the different types of data we have access to for biology. They're disparate, they're different types, and they're different sizes. When we're thinking about what in that whole landscape might be of particular concern, I would say that it's currently an open question.
There are some who think that you could train an AI model on genetic sequence data alone and get a pretty functional model. I think that models like Evo 2, which have done this by training just on DNA sequence data, can perform well on protein-related tasks. They can span scales, and they can do genome generation. So there's optimism that you could do quite a bit with just sequence data.
But I think that there's another view in the community that genetic data is observational. What you really need to advance biology is some kind of data that gives you insight into causality. That's where you have things like perturbation data or knockout data sets, where you're systematically knocking out different genes of a virus and then looking at how that impacts its function, or you're systematically looking at how viral proteins bind to human proteins.
I'll broadly call that functional data. It gives you some insight into causality, and there's a view that incorporating that kind of data into training AI models is what's really needed to get you over to making functional biological constructs that are viable in the real world.
And so what we're proposing in terms of our data controls is that, as you rightly said, the vast majority of data should not be under control. I think that there's actually been a huge effort in biology to make data open access and more openly shared, because that advances research overall. I'm fully supportive of that.
The controls that we're proposing are really on functional data that gives you insight into important features of viruses that, frankly, the US government has recognized as relevant to whether a virus is pandemic-capable or not. Those features include transmissibility, virulence, or how deadly it is, and things like immune evasion. Can you modify a virus so that it gets around an existing vaccine or gets around your own immune system?
Those are the kinds of features the US government already tracks for wet-lab research. If you're proposing a wet-lab experiment that intends to enhance a pandemic virus—to make it more transmissible, more virulent, or to evade your vaccines—that's something the US government wants to know about and is going to ask you whether or not it's a good idea. Doing it in the computational domain is just extending that a little bit further. We're really proposing focusing narrowly on that kind of data.
The other thing that I'd add is that I completely agree that going out and scrubbing the internet of data that already exists is not going to be possible. It would be a Herculean effort and probably not worth the effort. What we're proposing is controls on data that's generated in the future: new data sets.
My thesis is that now that we know AI is quite promising for biology, there will be huge investment in creating new data sets. We're already seeing this with the US government's Genesis Mission, the OpenAI Foundation's commitment to spend billions on data sets, and the Chan Zuckerberg Biohub as well.
As really large-scale efforts get underway to generate not just observational but also causality-related information and perturbation data sets, that's where we're suggesting that, if this is done on pandemic pathogens, that data should probably not be shared for completely anonymous access.
Jassi Pannu
You should track who has access to it and have some controls around it.
Nathan Labenz
So, going back to the people doing wet-lab experiments on gain-of-function-style premises, I'd be interested in your take. I'd put my cards on the table: I think that's not a good idea. When I look at the timeline from the rise of the virus to the sequencing to the vaccine design in the COVID-19 case, I do take your point that there was some prior knowledge that accelerated things.
I'd also be interested to hear how much, if it weren't for that sort of COVID 1 knowledge, that timeline would have moved. But my zoomed-out and somewhat ignorant view is that it didn't take that long. Certainly, the clinical-trial part took a lot, and the distribution and manufacturing took a lot longer.
So, if the argument is that we want to do these experiments now because we'll be able to shorten the timeline in the future if something like this does happen, to be able to respond to it, I would say you're not really taking the bulk of the time out by quickening the pace to vaccine design. I'd hate to see you let it loose. So maybe we just shouldn't do that. Do you see it the same way?
And then I guess another question would be the obvious extension: should we apply the same reasoning to certain kinds of data generation in the first place? Is there a certain kind of data set that we should just say maybe we're better off not scaling? It would be interesting to know exactly what it is in these viral sequences that causes transmissibility or whatever, but we are creating something there that, in a sense, can—it's obviously a little bit more upstream—but could, in theory, escape in a similar way.
Yeah, let's start with your take on gain-of-function research in the wet lab, and then whether that same analysis applies to the data-generation side.
Jassi Pannu
Yeah, excellent. I think you can state your view with more confidence because I think it's a very reasonable view. There was a lot packed in there, so if I miss anything that you just asked, feel free to flag it to me.
To provide some background on what gain-of-function research is—perhaps, I don't know if you talked about this before on the podcast—I'll give an example from 2012. In 2012, there were 2 experiments done by 2 different research groups looking at avian influenza, which at the time was thought to have a 60% case-fatality rate, so it was highly lethal, but it was not human-to-human transmissible. There were cases of humans getting avian influenza from animals, but it was not resulting in a global pandemic because it didn't efficiently transmit between humans, which is a happy accident for us, frankly.
What the researchers were doing was conducting animal experiments in ferrets where they intentionally increased the transmissibility of that virus between ferrets. Ferrets are the known mammalian model; they are meant to represent human immunity. The hypothesis was that they were creating a human-to-human transmissible version of this highly lethal virus.
When these experiments were submitted for publication—I believe they were simultaneously submitted to both Science and Nature—the journals received the publications and alerted the U.S. government, wondering, frankly, what they should do with these results, because the manuscripts included the specific mutations that would be required to create that level of transmissibility. One of the groups found that it was only 5 mutations in the avian influenza virus that got you to human-to-human transmissibility.
That work is the kind of work that people refer to as gain-of-function research. The technical terms are dual-use research of concern, or enhanced potential pandemic pathogen research. At the time, it was recommended that those publications, or at least the details of the mutations, not see the light of day.
There were 2 concerns there. One is the concern that you raised: Humans are working with these kinds of pathogens in the lab. We know that there's human error, and what happens if someone is dealing with a pathogen that they just created to be more transmissible, gets infected, leaves the lab without knowing it, and triggers a global pandemic? That's a legitimate concern.
There have been instances in the past where samples of really concerning viruses have been found in settings like the CDC. In one past example, the CDC found vials of smallpox in a freezer. They didn't know they were there, and they were still viable. There are supposed to be only 2 places in the world that have active samples of smallpox, and these were not known to be there. So I think overall there's concern about lab accidents and human error, and that was one of the major concerns generally falling under the category of biosafety.
The second concern was that, aside from dealing physically with the pathogen, there was concern about the information related to the experiments—not only how the researchers did this and how someone else could replicate that same effort, but also the exact mutations that would be needed and whether that was information that should be in the public domain.
This relates very much to what you described: The horsepox synthesis protocol is in the public domain, and the smallpox sequence is in the public domain. Theoretically, someone could put those 2 things together and try to create smallpox, even if they weren't able to get access to the physical specimen themselves.
Overall, there's a lot of information about protocols for doing reverse genetics or other ways of rescuing live, infectious virus for pandemic pathogens. In the wet-lab field, this has actually been a huge debate for years and years with regard to what we do about this information and whether it should be controlled in some way.
So far, where policy and regulation have come down is that there's a focus on controlling the physical specimen, and there's a focus on preventing experimental work from increasing these concerning characteristics of pathogens. But it's considered infeasible to try to go back and scrub the internet of data that's already out there. We really have to try to figure out a mechanism for deciding this in advance, before it's already out there and we can't do anything to pull it back.
Nathan Labenz
But it's still not illegal to do this. Is that right? Can you do it? Do you need any special permission, or is it an ask-for-forgiveness-not-permission regime that we're on with this kind of gain-of-function research, even in the physical realm?
Jassi Pannu
Yeah, this is a really good point. In the case of the 2012 experiments, there was no law that those researchers were breaking. The mechanism that the U.S. government, at least, has used in the past has been regulation. Essentially, if you receive funding from the U.S. government, you therefore have to follow certain policies.
This policy around not doing research that enhances pandemic pathogens is one way that the U.S. government has tried to do this. There are some laws on the books for dealing with controlled pathogens through the Federal Select Agent Program. For example, if you want to handle anthrax samples, you have to be a registered lab that's tracked under this program. But, yeah, that's slightly separate.
The other thing to consider is work that, for example, seeks to go into bat caves. People are collecting samples of viruses where there's a suspicion that those viruses could be pandemic-capable. They sample them in those caves, bring the samples back to the lab, manipulate them, and try to characterize them. This is also something that the U.S. government and other governments used to spend money on.
But after the COVID-19 pandemic, the fallout from that, the lab-leak hypothesis, and the political dynamics around that, a lot of that work has been defunded. It's not explicitly illegal.
Nathan Labenz
I don't want to get too bogged down in this particular point, but is there a good reason for that? I do understand, of course, that we benefit tremendously from biomedical research broadly. I could imagine you might say, “Actually, the border is a little harder to define and can be a little fuzzier, so it's hard to legislate.”
But if this is something we're sleeping on without a really good reason, it might be time to start writing our representative. Is there a good reason that this isn't more controlled than it currently is?
Jassi Pannu
I would love to see it more controlled than it currently is. I think the real reason that it isn't is because governments are good at legislating things that happen often.
Jassi Pannu
What we're dealing with are pretty rare instances that certainly would lead to extremely high-consequence harms—global pandemics and things that we don't want to see. But they just don't happen very often. And so, the push for policymakers to treat this as a live issue, as something that needs to be legislated, comes and goes very quickly.
We already saw with the COVID-19 pandemic that there was a lot of concern. There still is debate as to what the origins of COVID-19 were, and we haven't resolved that question. The WHO director actually just a couple of days ago put out a statement saying that we still need to do work on resolving this question. But policymakers have moved on to more pressing issues, because that's just the nature of policymaking: they have to put out fires today.
I think that, in reality, we need both national and international rules. Right now, the WHO has rules saying that only Russia's Vector Institute and the U.S. CDC are allowed to have access to smallpox. That's a great initiative, but it did not prevent a researcher from unilaterally publishing the step-by-step protocol for how to synthesize horsepox, the close relative to smallpox.
What that highlights is that, as synthetic biology, virology, and biomedical capabilities advance, we need a better way to make sure our regulations keep up. That's a concerning topic. But I think so far, the reason that this isn't a live issue, or why we're not thinking about it day to day, is because it still requires a lot of expertise to synthesize any of these pathogens from scratch. It really is something that you need to have a lot of background in, but this is where the concerns related to AI come up.
In terms of the different types of AI models, whether or not they provide uplift, and what kinds of biological models could be used to do this, this is such an evolving and open question that people are trying to figure out. I think the hard part is that it's moving quite quickly, and so it's hard to see how policymakers can keep up. But we're working on it.
Nathan Labenz
How much actual wet-lab gain-of-function research do you think is going on today? Has it been dramatically curtailed by these sorts of strings attached to funding and general awareness in the community that it's maybe not a good idea, or do you think there's still a lot going on?
I guess another reason that there might not be a law is, well, everybody quit doing it because they realized that it's a bad idea. Who needs to make a law against nobody doing it? But is that the case, Jassi? Do we have any way of really knowing how much is going on?
Jassi Pannu
With regard to wet-lab gain-of-function research, I would first want to say that the kinds of research we need for future vaccine design, like determining the spike protein sequence, are important. That's how we advance our ability to create vaccines for new pathogens, and that kind of work doesn't require gain-of-function research.
Gain-of-function research, the kind that we're talking about, is very narrowly scoped, and it does not require enhancing the transmission of a pathogen, making it more virulent, or making it escape the immune system. Those kinds of experiments are really not needed for the vast majority of the advancements we would want in biomedicine.
With that in mind, I would say that over the past few years, since the COVID-19 pandemic, a lot of this work has been defunded and reduced by U.S. government funding mechanisms. Our blind spot is the work that's happening in private labs. We don't actually have any legal mechanism for going into a private lab and determining whether they're doing certain kinds of pathogen research, other than whether they have registered under the Federal Select Agent Program.
There have been instances of laboratories in California, most recently, where they're handling certain types of pathogens that they really, frankly, shouldn't be, and they don't have the containment protocols for. I'll pause there. Overall, I think we're in a better spot than we were. I think people have recognized the downsides of this research and the risks, and certainly governments are paying a lot more attention to it.
Nathan Labenz
It's funny—it echoes, in a way, the reduction in bad behavior that we usually see from one generation of large language models to the next, where it's like, "We recognize that this was a problem, did some stuff to try to curtail it, and reduced it by 90%. Great news." The other 10% is out there for future work to contend with. Dizzying in the gain-of-function case.
Okay, you mentioned—let's go with your segue. There are different kinds of models, obviously, in the AI space that people might be concerned with. First of all, I was thinking ahead to this conversation and I was like, "Well, of course we've got the large language models, which output text and can reason about things." They might just know facts that could be problematic. They can use tools.
I just did a conversation with Jeffrey Irving, who's the chief scientist at the U.K. AISI, and I had not realized before that frontier models these days are getting quite good at troubleshooting lab experiments from cellphone pictures. They've now gotten to the point where you can just snap a picture of what you're working on, tell the AI that it's not working, and it will coach you through how to get it working. That's the know-how, the reasoning, and the procedural stuff.
Then we've got models that, as you alluded to with things like AlphaFold and that whole genre, are very good at making very specific predictions. What shape is this going to be? What's going to bind to what? So on and so forth.
Then there's the middle-ground hybrid: things like Evo and Evo 2, where they're trained kind of like large language models on these vast datasets. In many cases, they are literal next-token predictors, albeit in the DNA or protein-sequence domain. I probably have the least intuition for those.
I guess you could complicate that taxonomy for me if you want, but then maybe just go through and tell me how concerned I should be in a world where there's no data controls and the models have actually learned on everything we have. How concerned should I be about those different kinds of models, or possibly how they might be stitched together?
Jassi Pannu
Got it. I think, in general, I like this taxonomy. I was not as creative as you in terms of coming up with the different capabilities that these groups have, but I think about them largely in the same way.
LLMs are trained on lots of biological information, from textbooks to scientific papers. They can give that information to someone who doesn't already have it and isn't already a biology expert. That's usually called uplift, and broadly, that's a really great thing. LLMs teaching someone new biology, helping students learn, and really providing a quite useful capability.
There are a subset of instances—for example, "How do I illegally obtain an automatic weapon?" or "How do I illegally obtain smallpox samples?"—where that kind of information is clearly not something that should be provided to the general public. That's where frontier labs are working to apply classifiers and refusals to make sure that kind of knowledge is not widely shared.
Then, when you think about tools that can be used for biology tasks, I think people often call these biodesign tools. These are specialized models that are trained on biological data and are used to do specific things. I think of this as a model that gives someone a new capability. It's not about knowledge; it's about what you can do.
These types of models really require someone to already be an expert. You have to already be a computational biologist working with models to really leverage these kinds of capabilities. But the interesting thing is that they can allow those researchers to do something that was just not possible before.
Before the world of protein design, AlphaFold, and structure prediction, it was just not possible to take a protein sequence and then, immediately through computational methods, play with new designs or try to infer its structure and function. Those are really interesting capabilities. Again, the risks here are less about providing uplift to someone who didn't already know how to do that and more about giving experts the ability to do new things with biology.
And then this third category—I agree, it's kind of a middle ground—is models like Evo 2 and ESM-3. These are what I would call biology foundation models. They're trying to be general-purpose in the same way that LLMs are, and they're also often trained on different kinds of data.
AlphaFold, obviously, is trained on the PDB, but it also has MSA data. ESM has different types of data that it's trained on as well. These models are often much larger than biodesign tools, which can be small enough for an individual research group to train and host locally.
Biology foundation models require more data and more compute. They're more expensive for groups to develop, and so you often see these models developed by larger organizations. AlphaFold, obviously, is part of Google DeepMind, rather than an independent academic lab, and ESM is from EvolutionaryScale.
These types of models are trying to infer the fundamental laws of biology. They're trying to understand how biomolecular components interact across different scales and really elicit the underlying laws that govern protein function and protein structure and, in the case of Evo 2, operate across different scales.
Evo 2 is a model that's trained on just DNA sequence, and what they were able to show is that it can actually help with tasks across sequence, protein, and genetic regulatory circuits.
Jassi Pannu
So that's operating at different scales in biology, and it's inferring laws that transfer between those different scales. Overall, there's lots of interesting work being done in all of these. I think the risk considerations can be separated across the different types of models, but what I would argue is that that will collapse over time because what organizations are working toward are integrated workflows.
Ultimately, the dream is to be able to have your AI agent design your experiments. It will be connected to your autonomous robotics, which will conduct those experiments. The data from those experiments will be collected and then fed back into your biology foundation model, which will then be used by your agent to design future experiments, and so on, in a loop.
These kinds of iterative feedback loops, where you're getting data from the real world, I think, are where people are most hopeful about how this process can advance biology. Right now, we have these huge data sets that are messy and collected in an observational way, but what these feedback loops would allow you to do is systematically perturb systems, systematically try to assess causality, and then use that information to further and further improve your in silico models.
The dream would be, ultimately, that you get to a point where your in silico models perform so well that they start to replace some of the wet-lab biology that you've done. You get better and better predictions of what different drugs, for example, will do in cellular models and animal models. Ultimately, the dream would be that you get better predictions for clinical trials, so you have to do fewer of those and they have higher yields.
Nathan Labenz
So, what do we want to take from that? I'm obsessed, by the way, with that idea of both the closing of that loop and a little hobbyhorse of mine that I'd be interested in your take on: the sort of latent-space integration of these different modalities.
Obviously, we've seen this with image and text, in the sense that I can now go to a Nano Banana model or whatever and give it an image and also some text instruction, and it is understanding those in a joint way to a degree that wouldn't be possible if it were just prompting an external image model with text. I could have a language model that uses an image generator as a tool—we've seen that—but this sort of deeper integration gives you much higher fidelity to the original, and you can do text and image prompting in a very natural, integrated, cohesive way.
I've been wondering, assuming we're going to see it, on what timescale you think we see that kind of thing in the natural sciences, and specifically in biology as well. You might say, "Not, 'Okay, hey, language model, you can call this protein model as a tool,' but rather, you are both, and what I want is: working from this protein as an example, give me another protein that could do the same thing," and just have that all be understood in the same set of weights. What do you think the outlook is for that sort of system?
Jassi Pannu
I agree that that is where the field wants to go, and it would be really useful to be able to develop that. I feel like there are a couple of bottlenecks along the way.
One is the data-generation piece, which is still something that requires scaling in the physical world and doing experiments in the physical world. That itself will be bottlenecked by advances in robotics. If we were to suddenly see robotics speed up and were able to do a lot more laboratory work autonomously using robotics, then the overall picture in terms of data generation also would advance.
So I guess I'm hedging. I'm not really giving you a timeline—perhaps in the next 5 to 15 years. These are the kinds of advancements we would expect.
Nathan Labenz
Is that data—I mean, you said obviously we have huge amounts of just raw sequence data—but we are short on the causal graph, if you will, of "I did this, and this resulted." Presumably, a lot of that is maybe locked up at pharma companies that have done some of this stuff, or even just in the clinical data.
Do you think that if we had full access to all the data that exists, regardless of who owns it, how it was created, and where it's sequestered, would we have enough data already for that kind of thing to happen? Would we essentially be recreating, for IP reasons or privacy or whatever reasons, something we essentially already do have as a society? Or would you say, no, not really? Is the clinical data too messy, and maybe pharma doesn't have it? I don't know.
Jassi Pannu
Yeah, pharma definitely has lots of highly valuable data that they do not share in the public domain, for reasons that are obvious, and pharma is actually using its internal data sets to develop proprietary models in-house. They're certainly trying to do that.
If we were to suddenly wave a magic wand and say, "The government says everyone has to play nice and share their data sets," how far would we get? I think we would get a little bit further, but I think that these feedback loops and a new way of generating data are fundamentally a different approach.
You can think of the existing way of approaching biology as a bit artisanal and a bit observational, and that results in data sets being messy, having a lot of bias, and being hard to work with. What we really need to do is shift toward a much more systematic approach, where we are generating data that is comprehensive and systematically probing every single aspect.
That's where you really need the robotic aspect to scale that data and replicate it carefully, rather than having multiple different humans trying to do the protocol. There's always differences between them when they're collecting data. Just transitioning to a systematic approach that's enabled by robotics, I think, is not something that you would get just by enabling data sharing across private companies.
Nathan Labenz
Got you. Okay. Well, let's pop out of that rabbit hole and come back to the main topic. So we've got language models that can tell people things they maybe shouldn't know. They can increasingly use all kinds of tools, including design tools.
It's not clear to me at this point how well they could use something like Evo 2, but when we think about those models and their capabilities, what sort of capabilities should not be created in the first place? Is it about—I guess it's probably multiple things—but maybe I'll just leave it there: What capabilities do you think those different kinds of models should not have in order to reduce the risk to society broadly?
Jassi Pannu
I think that the fundamental challenge of biology is that a lot of these capabilities would be useful on the defensive side, but it's when they're used offensively that they pose concerns. So it's the question of what capabilities we want, but it's also the question of what capabilities we should provide access to, how broadly, and when.
When we're thinking about advancing the future of AI for biology, I think the way I like to think about it is that we should really try to step on the gas for things that are clearly good and clearly do not have a lot of risks. Things in that bucket, to me, are virtual cell models, ways of advancing clinical trials, or ways of making sure we can do better countermeasure manufacturing and distribution. There are lots of things that we could do that are clearly beneficial.
Then there is a bucket of things that, if that capability were broadly accessible right now, would be quite destabilizing. I think this is just a hypothetical example, not trying to say that this is actually the current state of capabilities, but let's say there were a breakthrough where suddenly it's very easy to use an autonomous robotic system that's quite cheap to build or get access to, and that robotic system could very quickly synthesize a pathogen.
That's obviously a futuristic scenario, but let's say that were possible. Then that's something that probably we wouldn't want anyone to be able to buy off the shelf. We'd want to know who has access to that device and what they're using it for.
Other things that are trending in the more concerning capability bucket would be things like viral design. Even there, there are considerations: We know that gene therapy based on viruses is actually an advance that we would love to see, or there are other purposes for viral design, like designing bacteriophages, which are viruses that only infect bacteria—they don't infect humans.
The challenge is that when it comes to artificial intelligence, a lot of the approaches are general-purpose. So if it becomes quite easy to have an AI model that can design a bacteriophage, then the question is, well, it seems quite easy to repurpose that for pandemic human pathogens. How many people do we actually want to have access to that kind of capability?
It's probably a subset of legitimate researchers who are using that. It's not something that you would want widely accessible on the internet, especially in a world where we don't have easily accessible countermeasures. It really becomes an offense-dominant capability where the design and acquisition of a pathogen become easy, and it's facilitated by AI.
That capability exists in the digital world—it is being uplifted in the digital world—but our countermeasures to a pandemic remain very physically world-bottlenecked. That's a world where it just becomes very offense-dominant.
Nathan Labenz
Even with things like a whole-cell model, would I be right to worry that one of the things you would want to do with a whole-cell model is throw stuff at it and see what happens to the cell? If you had that, all kinds of great things might be possible, but then presumably you could also start to throw in your virus of choice and start evolving that in whatever way you want to, potentially just brute-force throwing all these little permutations of a virus at the whole-cell model. It strikes me that these things are vulnerable to a brute-force attack.
If they're going to be good, they're going to be vulnerable to that sort of brute-force attack. Is that right, or is there any way around that conclusion?
Jassi Pannu
Yeah, you are embodying the debates that people constantly have in the biosecurity community. I think what you're saying is correct. There are ways to envision every kind of biomedical advance, especially in the AI domain, as being used for harm. Because of that, you have to think about how direct the harm pathway is and how consequential the ultimate harm would be, and you have to try to draw a line somewhere.
If you compare, for example, a generative language model that had no data filtering, had no data exclusion, was highly performant on viral genome design, and could do that for human pandemic pathogens, the pathway for harm is quite direct. Someone with very limited biology knowledge could use that model to generate thousands of potential designs, sequence them in the wet lab, see which one is the optimal candidate, and then use that candidate. There is still a lot of work going into that, but it's a direct pathway.
What you're describing would require plugging in multiple different AI models, generating candidates with one model, then running them through a different predictive model, seeing the consequences, and trying to figure out which viral candidate would cause, for example, a systemic inflammatory response, or would target certain organ cells. I think there is a pathway to harm there; it's just that when you game it out, there are more steps involved and more expertise required.
The ultimate question you ask yourself is, what kind of actor would choose that pathway over an existing weapons pathway? If it really requires high-level expertise that only a nation-state has access to, is that nation-state really choosing a biological weapon, or are they more likely to choose something else that a nation-state would have access to? Those are the kinds of questions that security professionals try to game out.
Nathan Labenz
Okay, maybe let's get to the proposed solution. I've been coming at this from a lot of different angles. You've got a whole taxonomy of five levels of biological data. Obviously, this is inspired by, or at least pattern-matched to, the levels of security around bio facilities. Maybe just take us through zero to five. What are the kinds of data that fall into these different levels? What would the access look like? What would the precautions look like? Paint a picture of the world that you envision.
Jassi Pannu
Great. I'll try to paint a somewhat visual picture. For those imagining this, it's a five-tiered system that goes from level zero to four. This is modeled on what some of you may be familiar with: biosafety levels, or BSL levels. These are the famous safety levels that biological laboratories use to determine whether I have to wear a spacesuit when I go into the lab or can just use a fume hood to deal with my samples. The system determines the different containment approaches that are required.
Actually, the BSL system was the basis for a lot of the frontier safety policies that different frontier AI labs have, in terms of the idea of having a system and mechanisms for controlling it, roughly on four tiers. What we're proposing here is applying this not to the model and not to the physical pathogen, but rather applying it to data.
The reason we're proposing this is because the entire biomedical research ecosystem, when it comes to AI, is built on open-source models. Academics build open-source models, they share those models openly, and other researchers manipulate and change those models. There are a lot of benefits to that fully open-source ecosystem, and those benefits are what have prevented security approaches from being applied to models.
It seems like that's going to be a pretty intractable approach, and we were looking for a different approach that could be applied to ensure that you could still preserve this open-source model ecosystem, but not distribute capabilities that would be particularly concerning, like viral design. That's how we ultimately settled on biological data.
Biological data, especially the kinds that I described—the more functional data—is expensive to produce, requires a wet lab, and requires expertise to produce. That's why it's a potentially useful choke point. The tiering system that we described would preserve the vast majority of biological data as fully open access, and that's what we're calling BDL0, where most data would be available to researchers.
As you go from levels 1, 2, and 3 up to level 4, you have increasing levels of control based on how potentially concerning the data is. The way it's essentially broken up is that BDL1 is data that would allow you to infer viral patterns. It's pretty basic security, just requiring an account and understanding who the person is and whether they're a legitimate researcher.
As you go up, you're getting more focused on properties of pandemic pathogens that would directly lead to harm. These are the properties that I mentioned to you, which governments globally already pay attention to for wet-lab research. This includes making a pathogen more transmissible, making the host range larger, allowing it to infect more animals, or allowing it to move from infecting just an animal species to also infecting humans. It also includes manipulating the pathogen so that it evades the immune system.
These are the kinds of properties where, if your data set has data directly linking those properties to pathogens, it requires things like use approval. You would go to the repository and say, “I intend to do this kind of model development based on this data for this purpose.” If you're a legitimate researcher who has a good purpose for doing this, then you will get approval, versus just having this data openly accessible for anonymous access. Maybe I'll pause there.
Nathan Labenz
How would you describe the magnitudes of those? Is the outer ring, BDL0, like 99% of the actual raw data? How small does it get when you get up to the uppermost levels?
Jassi Pannu
Yeah, I would say 99% is a pretty good guess. We don't actually have numbers to base this on because there isn't a comprehensive tracking system for these kinds of data sets. But my guess, based on the research that we've done, would be that the highest security tier, the BDL4 tier, is a very small subset of all data.
It would be a very small number of specialized virology labs, for example, that would be affected. Frankly, those labs are probably already limiting access to the data sets that they produce in some way. It's just not a formalized system.
BDL-0 would cover the vast majority of the data we're talking about—petabytes and petabytes of data—with the vast majority being uncontrolled. The controls we're proposing are on a very narrow slice of data, particularly getting up to BDL-3 and BDL-4, with perhaps dozens or fewer laboratories affected.
I am making that statement based on my knowledge of the field, but there probably needs to be a more comprehensive effort to try to figure out who exactly is generating this kind of data. The visibility bottleneck we currently have is what's happening in the private ecosystem. There's pretty good visibility in terms of government-funded work, and less so on the private side.
Nathan Labenz
In terms of the impact that this would have, one of the things I thought was really interesting in reading the recent paper calling for these kinds of controls was the report—which I hadn't realized—that there has been data holdout work done on a couple of the leading models, ESM-3 and Evo 2, specifically. Could you talk us through a little bit of what that has looked like?
I did put in a good word for me with Alex Rivers, please, to get him on the show. I've tried, but we did do one with Brian Hie on Evo. I have a general sense of what that looks like, but I didn't get into what was held out, how much of the overall data it represented, or how that affected performance in the areas of concern.
I assume one thing people would be really worried about here is, “I don't want to have a dumb model in general.” If I slice out this data, what does that mean in terms of what it can't do that I want it not to be able to do? But also, are there things that it can't do that I would wish it still could do? What costs are we paying for the benefits?
Jassi Pannu
Yeah, this is an important question. I think we all have an intuitive sense that AI model capabilities are based on the data an AI model is trained on. It intuitively makes sense to us, but the degree to which that is true is an empirical question.
Especially in biology, there's a reasonable reason to question whether, if I were to remove a very small subset of data, my model could just interpolate around that gap. If your model has really internalized a fundamental understanding of the laws that govern biology, does it really matter if you start segmenting out different small pieces of data?
This was an empirical question until, as you said, some of the leading biological AI model developers actually went about doing this. I'll just describe the two examples you mentioned: ESM-3, which is a generative protein design model made by EvolutionaryScale, and Evo 2, which is a generative DNA language model made by Brian Hie.
Both of those groups had decided that they wanted to share their models, but they didn't want to disseminate the capability for others to use their models for viral design.
Jassi Pannu
So the way they went about limiting that capability was by limiting what went into the training data. In the case of Evo 2, for example, given that I was involved in that work, we decided to remove the sequences related to viruses that could infect humans and viruses that could infect eukaryotic organisms. But there was still some information related to other types of viruses—for example, those that could infect bacteria—included in the training data.
After doing the big pretraining run, we then did some evaluations where we checked to see: Is the model limited in its capabilities on certain tasks? A common task that people use in this case is looking at how well the model can do certain viral-protein-related tasks and how well the model can generate sequences that correspond to functional viruses. Those were all things that the team checked, and they showed that the model's capabilities were significantly less than they were in other domains. Those evaluation results were all published as part of the Evo 2 manuscript.
The interesting thing about the work that the EvolutionaryScale team did for ESM-3 was that they actually had both versions of the model. This wasn't something that the Evo 2 team did, but the EvolutionaryScale team had both the trained model that had been trained on all the data they had chosen to include, as well as the data-filtered model. They were able to show a delta in performance on the same tasks with regard to, for example, viral protein function prediction.
So, that's an inkling of some of the empirical work that has been done and could be done in the future to try to suss out how much it matters when you remove these kinds of data and what particular kind of data matters. These are all questions that could probably be explored a lot more.
Nathan Labenz
Could you give a sense of the order of magnitude of capability reduction? Are we talking about it just not being able to do that stuff at all anymore, or is it in the uncanny valley somewhere? I don't have an intuition, honestly, for what I should expect.
Jassi Pannu
Yeah, I think so. For example, in the case of Evo 2, it was much more along the lines of the model's function being effectively random rather than just being reduced by a small amount.
Nathan Labenz
Cool. That's great. Great news. I love it when something works.
Nathan Labenz
Okay, maybe let's zoom out and take stock of all this stuff. We've got a ton of new data coming online. We want to facilitate sharing, and we want to facilitate all this discovery, but we've got to have some of these different classes of data controlled so that models aren't trained and disseminated on them.
Presumably, this also means that the models themselves that are trained by the ESM team and the Evo team also need to be controlled. I guess we're not—we shouldn't be expecting government control. If we don't have government controls on actual wet-lab gain-of-function research, it doesn't sound like we're going to get government control on this sort of thing.
So, what are we doing? We're campaigning and trying to build private agreement and consensus? Is that the play, and how's it going?
Jassi Pannu
Yeah, great questions. Just to take a step back in terms of what happens with the model and whether the models should be controlled, I think in the case of Evo 2 and ESM-3, they had effectively neutered the concerning capability from their models. That's why they felt more confident in disseminating those models. Ultimately, it made sense to be able to share those models openly because they had worked to reduce the concerning capability.
What I was more so referring to is that, if you do end up implementing data controls as we propose and you get access to BDL-3 or BDL-4 data for purposes of training a model, then what you wouldn't want to happen is that model being shared publicly, because effectively then your mitigation didn't really do anything. We do make sure to mention that any AI model trained on that kind of secure data should also be shared in a secure way if it needs to be shared.
In terms of what should be done next, I think the interesting thing about data controls is that there's a lot of policy precedent. We already do this in other domains. We do it for privacy-related data, and we do it for human genomics data. So, there's reasonable precedent for extending that same approach to data controls.
My collaborators and colleagues who worked on this are optimistic that perhaps this could get picked up by policymakers, but of course it's a new concept, and so we'll keep plugging away at it.
More broadly, in terms of what will happen with wet-lab gain-of-function work and model capabilities, I think that both the Republicans and the Democrats have decided that this is a bipartisan issue. Under both the prior Biden administration and the Trump administration, some work on advancing regulation of wet-lab gain-of-function work went ahead. I was really glad to see that progress.
I also think that both the US AISI and UK AISI are doing really great work with regard to biosecurity and biological model capabilities. Speaking given that I'm more familiar with their recent work, the US AISI has put out RFIs—requests for information—from the scientific community to better understand the issue.
We know that the developers are doing some of these data-filtering steps and trying to limit model capabilities. Maybe we should have a more systematic approach to figuring out what all the capabilities are that we should be concerned about. How can we test for those capabilities? How do we effectively mitigate them?
I would just love to see AISI be resourced and staffed to advance that line of work, because I think a lot of developers are trying to do it themselves in an ad hoc way, but they don't have the same security access and intelligence access that something like the US AISI has.
Nathan Labenz
Do you envision a sort of centralized—I’m still a little bit fuzzy on the workflow of something like this. If I, for whatever reason, am out here generating some sensitive data that maps viral sequence onto viral capability, and now I'm like, “Okay, I've got this data. I want to be a good citizen. What do I do?”
Are you envisioning a scenario where I take it to a central data bank, give it to them, scrub it from my computers, and go on with my life? Because we can't have everybody doing these sorts of security levels for themselves, right? So there would have to be some—maybe not totally centralized, but at least no more than a countable, relatively small number of organizations or entities that would have custody of the data.
Is that kind of the idea—that people would feed it in and then those organizations would be expected to actually delete it from their own servers? How do you— is that realistic to expect people would do that? How do you expect that to actually go?
Jassi Pannu
Yeah, this is all about how you operationalize the system that we're proposing. The way that this has been done in the past is through things called trusted research environments, or TREs, and OpenSAFELY, the system that I mentioned at the beginning, is one of these.
The idea would be that you have a secure environment that hosts the data and also enables researchers to bring their code to the data and answer questions. Honestly, in the age of AI, the ideal would be that you have compute resources available to that, although we'll see if that would be possible.
I think the way that you would set this up is that we considered different approaches. One approach could be that the government sets this up. After we looked into this, it seemed less ideal. A lot of researchers have been unhappy with some of the trusted research environments that the government set up, and I think there are private actors that could probably do a better job of setting up really savvy trusted research environments that work well for researchers. We have some examples of this that I've mentioned and that are included in the paper.
What we suggest is that institutions—not individual researchers, but institutions like universities or larger research collaborations—set up a trusted research environment if they want to do this kind of data generation. They would be given standards that the government would set, but they would be in charge of actually building and maintaining the environment, just because they're probably going to be better at doing that.
The interesting thing about a trusted research environment is that you can actually imagine this also being a boon for researchers. When you're trying to do AI research, what you want is a centralized, integrated platform where all the data sits and you have access to it all at once. It's much less useful to have individual data sets housed with individual researchers.
So, if we really wanted to step on the gas of advancing countermeasures research for pandemic pathogens, maybe this is something we would want anyway. We would want an integrated system that hosts all the data and really makes it easy for our researchers to interact with.
What we're proposing to layer on top of that is that you do have some security that goes along with that system. Not only are you hopefully getting the benefits of the integrated platform, but then you're also getting the security controls, given the sensitive nature of the data.
Nathan Labenz
Got you. When it comes to monitors, I know that you had alluded to some rare points of bipartisan agreement. One of those, I understand, is the insistence—or I don't know if it's fully a requirement, but I think it's verging on a requirement—that DNA synthesis companies apply certain classifiers or whatever to try to detect if somebody is trying to get a harmful sequence synthesized through their company.
There’s also things like wastewater monitoring. I have a couple of questions on this. One is, is there an equivalent of wastewater monitoring for data? Would it make sense—I don’t know if the data of concern is the kind of thing that could be identified if it’s just hanging out on somebody’s lab website.
Potentially, people might not even fully realize if they’ve generated problematic data. So, is there any system you could imagine to go around and identify data that’s out there and then classify it as possibly harmful, or is that such a difficult classification problem that it would be doomed from the start?
Jassi Pannu
It’s a really interesting concept. I have not thought about it too much, to be honest, but I feel like you would probably need—there are 2 approaches to figuring out what the data landscape is for this kind of data. One is a more active, contribution-based approach, where the government says, “If you think you’re creating this data, you have to actively contribute it to these repositories.”
What you’re describing is a passive approach, where no one has to take any individual action. Some system is flagging the data, and then it perhaps even automatically collects it and scrubs it from the original source. That would be the dream. I’m not aware of any system like that, and I think it would probably take a lot of work to do. I think that the default approach has been the first one, which is voluntary, or active contribution by the researchers who are generating the data. But what you’re describing would be cool, and it would be the analogy to wastewater surveillance, essentially.
Nathan Labenz
When it comes to the quality of monitors in general, a big thing I always think about in the rest of the world as it pertains to AI is that we’re in this sort of weird in-between phase where things are coming online, but the world hasn’t really reacted that much yet. For example, AI agents are coming online. There are a lot of concerns around things like prompt injection and whatever, but the world hasn’t really become a very dangerous place for an AI agent so far, right? Not that many people have set up honey traps or prompt-injection attacks to try to throw my OpenClaw off its path and talk it into doing something that they want it to do. I assume that’s going to happen a lot more, and there’s going to be some arms race between techniques to prevent my OpenClaw from falling for it and ever more sophisticated jailbreaks, whatever.
Is there—I assume there’s got to be something analogous in the biological domain where, for starters, these classifiers that the DNA synthesis companies are running—I’m guessing that nobody really has tried to evade them yet. I wonder if you see that kind of dynamic developing on the horizon. Nicholas Carlini, who was a past guest, has said that the attacker usually gets to act last. The defenses are set up, and then the attacker always has that advantage of knowing what they’re up against. Maybe not always, but often.
Do you see any of those kinds of dynamics now, or do you worry about them in the future? And I guess, if you extrapolate them out, do you see us getting to a place where we have a high level of confidence that we’re in a defense-dominant world and we’re going to be able to keep all this stuff under control, or is that itself still a very open question for you?
Jassi Pannu
Yeah, I have some ideas, but maybe let’s start with gene synthesis screening and the different—there were a few different attack surfaces. I’ll just describe them that way. The first one is gene synthesis providers.
So, what you’re describing is the system that is currently voluntary, where gene synthesis providers—companies that make pieces of DNA and sell them to researchers—have implemented systems where they check to make sure that someone didn’t just order pieces of smallpox or pieces of Ebola. They have a twofold mechanism.
The first is an automated mechanism that is essentially sequence matching. It’s looking at whether the order matches a sequence from Ebola or matches a sequence from smallpox. If the automated system thinks that there is some degree of concern there, it gets escalated to a human researcher, a human expert, who then looks at it and determines whether it’s something of concern or not.
What’s also happening in parallel is some degree of KYC—know your customer. Who did the order come from? Is it a researcher? Does this researcher work with these kinds of pathogens all the time? Have we spoken to them before, and has this issue come up before? Those are all the kinds of questions that are being addressed as part of gene synthesis screening.
Gene synthesis screening is something that 80% of companies already implement, and over the past few years it’s become much more cost-effective to do. Initially, it was a bit of a cost barrier. They really did a lot of work to make this a cost-effective system, because you can imagine that if this is something being run on every single order that’s coming in, it has to be cheap enough to do. Otherwise, it’s really burdensome for these companies.
Currently, it’s a voluntary system. Like I said, the vast majority of companies already do it, but the concern is that if I’m a bad actor, then a voluntary system that 80% of companies use isn’t really going to stop me from obtaining the sequences I want to get my hands on, because I’ll simply go to the companies that don’t do the screening.
What’s now being advanced is the idea of making this a mandatory rule that all companies have to follow. That would really limit the access that someone has to physical specimens that they might try to use to turn into an infectious pathogen.
But I guess, given where the research is going and what capabilities people want to achieve, the dream for the future of biology is that I, as a biologist, no longer even have to step foot in the lab. I have my autonomous cloud lab that I can fully control. I’m sitting at home using Claude Code to help me, and I can just program some designs to be done on certain pathogens.
To be completely frank, we’re nowhere close to this world. It’s still going to take a lot of work. Right now, the cloud labs that you hear about still require a lot of human input. It might not be specialized biologist input, but there is still a person picking up samples from one bench and moving them to the other, and that’s a real bottleneck.
But in a world where you do have a fully remote, highly sophisticated cloud lab, then you can imagine that if there are agents—if I’m a bad actor and I have an army of 1,000 agents that are just trying to hack their way into this cloud lab—you want to make sure that if your cloud lab has sophisticated capabilities related to pathogen creation and design, you have some cybersecurity around that. So, that’s a much more future-oriented thing, but something you could imagine becoming applicable later.
Right now, there’s a fundamental information-infrastructure layer that we’re missing. So, both for Palantir Labs and for gene synthesis screening, if I, as a bad actor, try to obtain sequences of Ebola from one company, small pieces from one company, and a few pieces from a different company, and I split up my order across different companies, there isn’t some kind of system where all those companies are easily communicating that information to each other and checking those orders against each other in real time.
This is something that’s bottlenecked by information sharing between companies. That’s also a similar concern to why the Frontier Model Forum was created: you want private companies to be sharing security-related information. You need a legal infrastructure for that. Who would house that? Would that be the FBI? Who’s facilitating this? These are all kinds of policy questions that need to be addressed.
Perhaps we can end on a more optimistic note. I’m happy to give you my vision of what a defense-dominant world looks like and see what you think of it. I think that, overall, I break this up into 4 broad buckets of interventions.
Well, perhaps to take a step back: I think people often think about what the theory of victory here is. How do we have a single, unified strategy, like our strategy for nuclear deterrence? We have a single, unified theory of victory for nuclear deterrence that has served us well for decades. How do we get there for biology?
I guess, after having thought about this for some time, it just feels like biology is very different. It’s a distributed technology. It’s dual-use. You want to give a lot of people access to it. You’re trying to limit it to a subset. There are all these aspects that really make a unified theory of victory seem much harder to get to.
So, I think the most successful approach is likely to be defense-in-depth: a layered approach, with multiple different defensive strategies applied at once. The 4 buckets that I divide this up into are deter—oh, sorry. I should start with delay, deter, detect, and defend.
Delay is essentially limiting access to concerning capabilities. Gene synthesis screening would fall into that bucket. You’re delaying the dissemination of that capability to get access to DNA fragments, for example. You could also imagine what we’re describing for our data controls as part of delay.
Then there’s deterrence. Deterrence is figuring out how you can punish someone for using a biological weapon. We do live in a world where there’s an international treaty against biological weapons. I’m glad that we live in a world that has that treaty versus not, even if, overall, that treaty is on the weaker side in terms of the actual mechanisms we have to ensure people are complying with it.
But that’s one thing. I think where deterrence breaks down is if your actor is not rational and doesn’t respond to typical punishment mechanisms. So that’s a challenge. But then, in terms of detection, this is also something we talked about: How do we have a distributed passive surveillance system, like our radar system for ICBMs, that will just detect when there’s a new pathogen without us having to go out there and look for it, perhaps even for pathogens where there are no symptoms?
When HIV was spreading early on, it would have been amazing to have known that much earlier. And that’s particularly important for pathogens that take a long time for patients to develop symptoms. So, having some kind of global surveillance system—people often refer to this as bio-radar or bio-threat radar—it’s not something we currently have.
And then the last pillar is defenses. Really, what are our defenses once there is a pathogen online already? I think that people often think of defenses as things like vaccines and countermeasures, but I would encourage folks to be much broader in what they envision defenses to be.
Because I would argue that I’m sitting in my home right now, and I actually have defenses all around me. I’m drinking water that’s been centrally filtered. I know there’s no cholera, no pathogens in that water. I have screens on my windows. Mosquitoes can’t get through them. I know that I’m not going to get malaria if there was malaria outside.
And so there’s already a lot of public health defenses built into our environment, but we don’t have that for airborne transmission. One thing that’s being explored by organizations—one that comes to mind is Blueprint for Biosecurity—is built-environment defenses to sterilize the air. This is using approaches like far-UVC and other approaches like glycol vapors.
So, could you passively sterilize the air so you don’t even have to detect the pathogen? You don’t even need a vaccine. You just always know that you have passive protection around you. So hopefully, I would say that’s a pretty comprehensive approach. If we manage to do all of that, it would take a lot of work and investment to get there, but it would make us a lot safer.
Nathan Labenz
…above his hospital bed. Every time he’s been in the hospital, we’ve hopefully taken some of the risk off the table of him getting any kind of infection while he’s going through all this. That’s a great vision. I hope we implement it.
You are doing God’s work spending your precious time and energy on this. Anything else we didn’t touch on, or any other calls to action or ways that people can help you, that you would want to leave people with before we break?
Jassi Pannu
I think we covered everything, and I really appreciate your interest. It sounds like you’re one step ahead of everyone else in terms of already getting your kid outfitted with all the defenses he needs. It’s great to chat with you.
Jassi Pannu, thank you for being part of The Cognitive Revolution.