Nathan Labenz
Today, I’m speaking with Dr. Catherine Brownstein, MPH, PhD, an assistant professor at Boston Children’s Hospital and Harvard Medical School whose research focuses on identifying the genetic causes of previously unexplained rare and orphan diseases, and who was recently awarded a ChatGPT Pro Grant from OpenAI.
You might be surprised to learn, as I was, that so-called rare diseases are not necessarily all that rare. Any disease affecting fewer than 1 in 2,000 people, or fewer than 200,000 people in the United States, is classified as a rare disease. Often, families spend painfully frustrating years bouncing around the medical system in search of an accurate diagnosis before ultimately reaching Dr. Brownstein’s elite team at Boston Children’s.
Of course, considering the radical cost reduction we’ve seen in genetic sequencing in recent years—with nearly a 10,000× improvement in affordability, one of the very few cost curves ever to rival that of large language models—there’s been an ongoing revolution in this space even before the current AI moment. In 2007, a genome sequence cost upwards of $1 million. In that era, it was used only in the most challenging cases and was often a difference-maker. Today, it’s just a couple hundred dollars and has become commonplace for individual patients.
But that creates new challenges for specialists like Catherine, who now have to comb through a vast and still exponentially growing literature to find candidate diagnoses for their most challenging cases. This new wealth of information, which, as you’ll hear, could be growing even faster with improved regulations and incentives, makes information-processing capacity relatively scarce and valuable. And you can probably guess where this is going: a great target for the latest generation of reasoning models.
This conversation is, above all, a window into how frontier large language models are starting to become useful in highly specialized fields. Dr. Brownstein is pioneering the application of AI to rare-disease research in real time. She’s using AI to triage potentially relevant research and, in some cases, to connect the dots between subtle clues. She’s working directly with OpenAI to develop use cases and provide feedback.
Considering that every case represents a real person with a life-altering or even life-threatening condition, she’s constantly working to find the right balance between enthusiasm for AI’s capabilities and a healthy skepticism about any specific AI output. As you’ll hear, she’s still figuring out where AIs can be the most valuable, how best to use them, and how much to trust them. That such an established expert is bringing what amounts to a beginner’s mindset to such high-stakes cases may be surprising to some, but I really don’t think it should be.
Even the most AI-obsessed folks like me have only managed to log a few thousand hours with large language models, and nearly all of that was with earlier and less powerful models. So, for the current frontier, we’re all still figuring this out together, and there’s currently an unprecedented opportunity for people with deep experience in specific niche domains to become the leaders in applying AI to their particular fields.
Catherine Brownstein, MPH, PhD, an assistant professor at Boston Children’s Hospital and Harvard Medical School, you specialize in the discovery of new genes for rare and orphan diseases, and you’ve recently been awarded a ChatGPT Pro Grant. Welcome.
Dr. Catherine Brownstein
Thank you so much for having me.
Nathan Labenz
Yeah, I think this is going to be really exciting. As regular listeners know, I have a growing obsession with the intersection of AI and biology. When I saw your name on the o1 Pro blog post announcement, I was excited to reach out and learn more about how you’re applying the latest AI tools to some of these very challenging and pressing problems.
Maybe we could start with a zoomed-out overview of your work, because our listeners are definitely following AI developments. They know about o1, they know about o1 Pro, and probably quite a few have subscribed, even at the $200-a-month level. But they probably don’t know a lot about rare diseases, what the state of knowledge is, or what sorts of techniques people use to try to figure these things out.
I’d love to get a layman’s introduction to your advanced work. This may be a very tough question, maybe the toughest question, but what’s the layman’s introduction to your work?
Dr. Catherine Brownstein
When I’m asked a question, I usually answer it by saying that rare diseases are quite common, actually. There are more people with rare diseases in the United States than there are natural blondes, so it’s actually quite common to have a rare disease.
A lot of it is how you define disease. Autism is really common, but autism due to a de novo variant in KCNJ8 is quite rare. It’s a tricky definition, and it’s always evolving as we learn more, but basically, I consider myself a gene hunter trying to diagnose the undiagnosed.
Nathan Labenz
I read in preparing for this—I think it was Perplexity that gave me this answer—that the definition of a rare disease is one that affects fewer than 200,000 people in the United States. That was a surprisingly large number, wasn’t it?
Dr. Catherine Brownstein
It’s always wild to me because when you think of a city that has 200,000 people, that doesn’t seem like a small town, or at least it doesn’t to me. But that’s the definition of rare in comparison to common disease, which can affect millions. Epilepsy, for example, affects 1% of the population.
Nathan Labenz
Yeah, that’s really interesting. Maybe a little bit more background on the patient’s journey through the medical system to get to you, and then your experience of encountering new patients. I know it’s impossible to give just one story because I’m sure they’re extremely varied, but how do you end up coming into contact with patients? What have they gone through to get to you? And what do you do once you get a new case?
Dr. Catherine Brownstein
I’m really lucky to be at Boston Children’s, which is an internationally known tertiary hospital, so we get really interesting cases from all over the globe. Usually, a patient or family starts out by going to their local medical provider. They can’t figure out what’s wrong with the child or person, so they get referred from specialist to specialist, and they still can’t figure out what’s wrong. Eventually, they get to us.
A lot of times, patients and families have been bounced around for years, trying to figure out what’s going on, what’s next, what they can expect, and just looking for answers. Sometimes that’s not the case. We have a lot of really medically savvy families who know their child, know something’s wrong, and need the best right away. They search on the web, find the person who works on that phenotype, and call every day until they get an appointment.
A lot of times, though, it’s a more circuitous route, going from doctor to doctor to doctor and then finally somehow ending up at Boston Children’s. If they see a clinician who doesn’t know what’s going on, they often refer the case to the organization I work with, the Manton Center for Orphan Disease Research.
We get a lot of the negative cases throughout the hospital where they think it’s genetic in origin. Then we’re able to get the medical records. We’re a philanthropically supported center, and patients can self-refer. We get all the medical records and all the genetics that have been done before. Then we have a huge multidisciplinary team, and we review the case, go through it, and do a reanalysis.
Sometimes we resequence or use a new technology if one is available, like RNA-seq or long-read sequencing. Then we work together to try to figure out what’s going on. When I first started in 2011, genome sequencing and exome sequencing were quite rare. If patients were able to get it, a lot of times it was like shooting fish in a barrel. We would have something like an 80% diagnosis rate.
But now genome sequencing and next-generation sequencing are so common that we only see the families if they've already had a negative sequencing test. So we go from diagnosing roughly 80% of cases to roughly 10%, just because we're getting the most difficult of the difficult cases. They've already been reviewed by really good geneticists and are getting to us because they just can't figure it out.
But that's one thing that I think AI can really address: shortening this diagnostic odyssey for patients who have just been jerked around—not through anyone's fault, but just by the nature of how these things go. Maybe AI can help in analyzing symptoms: maybe you should see this doctor right away, maybe you need this test, or maybe you need to go to this specialist, and just make things happen a lot faster.
Nathan Labenz
That callback to 10 years ago, I think, is quite interesting. Maybe you could give us a little bit of a sense of the relative pass-through rates at these different levels of the filter. People initially go to their local doctor, the local doctor doesn't know what's going on, they get referred, and eventually they get to your hospital, where you've got the best of the best.
There's another related but distinct line of research that has recently been comparing AI's ability to diagnose through a natural-language conversation with patients against doctors. It seems like, against at least the average doctor, the latest models are now very much holding their own. I don't know if that would be true if we were looking at Boston Children's elite clinicians and their ability to diagnose, but they still don't know what's wrong.
In the past, if I understand things correctly, because sequencing was rare, you could often just do a full genome sequence and then be like, “Oh, okay, well, there's your problem.” The literature has characterized this: now that we have this additional information, there's a pretty clear match. Today, that low-hanging fruit is getting absorbed somewhere else in the system before it gets to you, and you're now seeing things that are basically not characterized in the literature at all, or maybe just a little bit.
I'm not sure why the connection wouldn't be one that others could make, but use that prompt and fill in a little more detail, if you would.
Dr. Catherine Brownstein
When I started, I was actually hired as a project manager at Boston Children's to help clinicians get their patients sequenced. Clinicians, even though they weren't geneticists, were really good at identifying cases that were probably genetic in origin. They had freezers full of this DNA, just waiting for the technology to come online so they could analyze it and figure out if there was something genetic that could be discovered.
When I started, an exome, which is just 1% of the genome—it's just the coding region, just the genes—was $3,800, or close to $4,000. It's a good place to start if you're being economical, because a lot of the variants are within the coding region. Now, I just priced out an exome, and it's $160 for that exact same test.
It was so expensive that they went through a rigorous selection process if you were going to get an exome done. If you were going to bet money, you were going to bet money that it was genetic and that you were going to be able to figure it out by doing a trio—that is, the patient and the parents—and that you were going to see something that was de novo, which is basically not in the parents but is in the child. It's like lightning striking, an error happening during development, which causes disease.
The first cases of that were the 80% that I was talking about, and it was because these patients had been collected, in some cases, 20 years ago. They had the DNA there, and sure enough, there was a premature stop codon or a huge deletion of 1 gene that was already hypothesized to be related to this condition or a similar condition. You could point at it and be like, “Yep, that's it.” You would also have multiple cases of the same type, where you'd see 4 families with the same gene missing and the same phenotype, and then you're really confident that that gene is causative of the condition.
As the price dropped, it became less of a thing that happened. It's not because you couldn't get an exome done anywhere; there are a lot of geneticists, a lot of really savvy clinicians, and a lot of for-profit companies that you could send it off to, get a report back, get diagnosed, have more precision-medicine treatment, and go on your way and do very well.
So it was the negative cases that were getting referred up the chain to Boston Children's, because they had already had a genome and it came back negative. That is, there was no obvious variant in a known gene that could explain what was going on. Then it becomes a little trickier.
We start forming cohorts. At the Manton Center, we work with clinicians. We have clinicians in every department of the hospital who are able to refer patients to us. We consent them to our protocol, and then we collect samples and medical records. Sometimes we wait, and we reanalyze.
When we have 4 to 10 patients with the same thing, we're able to look at them together as a whole group and be like, “All right, are there things in the same gene, the same family of genes? What can we come up with as a hypothesis here?”
In 2014, I think Zack Kohane, who had previously been a guest on your podcast, had the idea of having an international competition to solve undiagnosed families. We got 3 families with seemingly Mendelian disorders. That is, we thought they were genetic, and we thought there was something going on with a clear relationship between gene and condition.
We released their data all over the globe to 23 different teams, and we had them compete and each submit a report on what they thought the cause of each family's condition was. It was really interesting. A lot were actually diagnosed from this—I think 2 out of 3 walked away with diagnoses from the process.
We were also able to show that diverse teams did much better. You can't just have a bunch of bioinformaticians in a room together looking at cases and expect them to come up with the right answer. It was teams that had a mix of research assistants, genetic counselors, researchers, clinicians, research clinicians, and clinical geneticists working together. All those diverse perspectives, on the whole, were able to solve more cases.
That was really interesting, and I think that's a recurring theme here when we're talking about LLMs, large language models, and AI. None of this exists in a vacuum. It's helping us along, and maybe it will be enough, but right now, having multidisciplinary teams with multiple strengths all working together means we do much better as a whole.
The other thing I wanted to add is that we still see those slam dunks. We just had a case a little while ago where it was 1 family and 3 generations, all with a rare bone disorder. The matriarch or patriarch was in their 90s, and we were able to give a diagnosis to this person in their 90s, which I thought was really, really cool and shows the power of just having an answer.
They had already gone through surgeries they didn't need to go through and had their whole life with this condition. But something as simple as being able to explain what's going on in 10 seconds, as opposed to 3 minutes of describing symptoms, means a lot to the family.
Nathan Labenz
Yeah, I imagine, especially if you've been dealing with something like that for 90-plus years. That's crazy to think about.
So, a lot of different questions I have about all this, but in these cases where you're getting all the way through the entire medical system, basically, and finally getting to one of these cross-functional teams—
Dr. Catherine Brownstein
Mm-hm.
Nathan Labenz
Can you tell us a little bit more about what the process looks like when that team gets to work? In AI prompting, we talk about thinking step by step and breaking problems down. Maybe one way to frame it would be: What is the sort of collective chain of thought that the group goes through to start with inputs?
Inputs would at least be symptom descriptions and results of genetic testing sequences. I don't know if there's any other inputs that you get at that level. I guess you have the whole scientific literature also as sort of an input. Then you do some thinking and reasoning, maybe some additional testing, and finally you get to a result. What are you doing when you're doing that?
Dr. Catherine Brownstein
Okay. So, when a case comes across my desk, usually there's a medical record that comes along with it because, again, they've been bounced around for a long while. Usually, at this point, they've had some genetic testing that gets transferred to us.
More and more patients are coming with it on a thumb drive, like, “Here's my genome,” which I think is really cool and didn't even happen a few years ago. We run it through our genomic pipelines, and usually we run it through more than 1 because they all have their strengths and weaknesses. Some are more comprehensive but harder to use, and you'll get more false positives because they rule fewer things out. Then you have others that are really easy to use—my kids can use them and understand intuitively what it means—but sometimes they're black boxes, and you don't know the reasoning behind why a variant was eliminated or not.
I'm a PhD, non-MD, so I usually like things to stay anonymous. I don't want to be a walking HIPAA violation, so I kind of don't want to know the names or meet the families, but sometimes I do. I know who they are.
We go through everything case by case and line by line. There are certain phenotypes where I think more information is better. You'll get the occasional phenotype that's only linked to 1 condition, like lack of tears in 1 condition, and that's a really important clue. Then we'll look at that gene.
For what the patient is experiencing overall, generally there are gene lists of what's already been discovered, and you can look at the genomic information for any variation that could be causing disease. We'll call it, for simplicity's sake, pathogenic variation, though suspected pathogenic variation is probably more accurate to say in those genes.
You get the new analysis done, and then you're looking at what's known. If you don't see anything, then you start looking at your special sauce. How am I going to approach this? Where else in the genome is notable? Is there a huge structural change that hasn't been linked to disease, or a translocation where chromosomes break and reattach in the wrong spots? Is there some other deletion or duplication? What's rare? What's unique to this patient?
Now there are also all these new technologies, like looking at epigenetics, where you can kind of predict which genes are turned on and off. Even if you can't see a mutation, is the gene of interest's expression perturbed somehow, or is it constitutively on even though it's not supposed to be? Can you take a look at that?
Sometimes, in the back of your mind, you're thinking, “Is it multifactorial?” It's not just 1 gene impacting it. It's not some big error in 1 gene; it's a bunch of tiny little things scattered throughout the genome. Then there are different types of tests, like looking at a GWAS, or genome-wide association study, or SCAT, where you can look at rare variation weighted by how rare the variation is and how damaging it's predicted to be to a protein. You look at that and see, “Okay, is there some reason why you think that this is going on?”
A lot of the time still—let's say 25% of genetic testing comes back positive—what does that mean? Sixty-six to 75% are negative. Then you go through this whole process, and still most are negative. You put it on the shelf, wait a little bit, and analyze it again a year later.
Reanalysis is actually really, really important because things get discovered all the time. You can't be an expert in every gene, every condition, every structural variation, and other people are actively working on it. A lot of times, you'll take something off the shelf and look at it again, and it rises right to the top. The number-one thing in the genome browser is the answer, and you stared at it a year ago and didn't make that connection. Now, all of a sudden, there is an answer.
Actually, I was asking Alan Beggs and Monica Wojcik, who are the director and medical director of the Manton Center, for success stories—if they had any that stuck out. One was from 20 years ago. There were 3 siblings who all passed away from a type of myopathy, and they couldn't figure it out.
They kept testing and testing and testing, and eventually ran out of DNA. Then we had a pilot grant at the hospital to do RNA-Seq, and Alan and Monica submitted this family because we had some RNA left. We found a variant in CFL2, I think that's the gene name.
Even though it was 20 years ago, the surviving siblings were now planning families, and they had an answer. They could do genetic testing to make sure that there weren't 2 variants and that they were each carriers of 1 variant. They hadn't passed away, so they only had 1 variant, not 2. They could also make sure that their partners didn't have a variant in the same gene and ensure that the next generation wasn't going to have this horrible, fatal myopathy.
In some ways, we had an interesting discussion: “Okay, is that really a success story?” Whenever there are multiple deceased people, is that really a success? Yes, you diagnosed it, but it's not changing anything. But it is changing the future. They're going forward with their eyes wide open and are able to plan as a result.
Nathan Labenz
Yeah, that sounds like certainly some form of success to me. I have 3 young kids, and fortunately, no crazy medical conditions in my family, but we still did a little bit of genetic testing. I would say I was probably never more nervous than when opening that report, just to make sure that I wouldn't have to see something really weird or strange, or that it would change the course of my life.
To be on a potentially negative course and get the assurance that you could confidently get on a path where you'd be able to have healthy children, I think sounds like, to put it mildly, a life-changing development for those folks.
Dr. Catherine Brownstein
That definitely resonates with me.
Nathan Labenz
Okay, let me dig back in at a few points along the way. I'll try to summarize a little bit and interject a couple of questions.
The pipelines that you're describing—I guess those are maybe a mix of commercial options or things that other academic groups have put out. The inputs to those, are they highly structured data?
I mean, I'm thinking here: My sequences are, of course, structured; my symptoms are not, right? I describe myself in words, and the doctor I'm talking to notes that in words. Is there a way that gets translated into specific coded sets of symptoms, or what is the intake of these pipelines?
And then are they basically doing deterministic work, where they're essentially running down a long checklist and saying, “If you have this, we check this. You don't have that, so that's out,” and working down a long set of known possible conditions? Or how would you characterize what those pipelines are doing internally?
Dr. Catherine Brownstein
So, I think you're exactly right. A lot of them let you input the phenotype, and it's coded to ontologies—sometimes HPO codes, sometimes ICD-9 or ICD-10, sometimes SNOMED. There are a bunch of different ontologies. I like HPO the best.
Nathan Labenz
Mm-hm.
Dr. Catherine Brownstein
Let me be real clear: In those sorts of ontologies, something like “no tears” would be a single alacrima HPO item. My condition might be summarized by a set of those. If I had no tears, hair falling out, and loose teeth, that would be 3 things. It would be, “Okay, the patient presents with this bundle of things.”
It's also a huge field of research. My friend Melissa Haendel works with HPO and her site, Monarch Initiative, mapping that onto animal phenotypes and making sure it's one-to-one. Humans don't have paws, but the phenotype that's closest to that can be translated.
Then there's layperson HPO, where we're not saying “alacrima,” but we say “no tears,” or “lazy eye” and “strabismus.” There's a whole mess of work that goes into that, making sure that it's accurate and also culturally sensitive, like “fit” for epilepsy. It's all this stuff that you never think of, and if you don't make those translations, then all of a sudden your phenotype is way less accurate than it could be.
That gets incorporated into the model. Then, when you input that with the genetics, you can have raw data, which is FASTQs—the zeros and ones that come off the machine—and then a BAM, where you're looking at the reads of the sequencing itself.
Gosh, I'm not going to explain this very well. But then you have the VCF, which is really processed data, and it's basically every single variant. It's huge— a VCF is a relatively huge file. It is orders of magnitude smaller than a FASTQ or a BAM, but it's still quite big.
Then you're putting the BAM or FASTQ into these pipelines, which process the data along with the phenotype. Then they're ordering the variants based on the HPO code related to the gene, the variant within that gene, and how likely it is to be positive for disease. The more sophisticated ones can take in relational kinds of things, where it's known that this gene binds to another gene, and gene A is related to the phenotype but gene B isn't yet, while there's a huge variant in gene B and the patient has the phenotype associated with gene A.
My actual first-ever success story was one of those cases. It's called episodic ataxia, and the patient would get really stiff and couldn't move—they would get locked in position. We did sequencing and saw that it was a variant in KCNA1, which wasn't the gene we were thinking of, but it was related to the gene we thought it was going to be. So KCNA1 just rose to the absolute top of the list, which was really, really cool.
Nathan Labenz
But that challenge of basically understanding the graph of interactions—what affects what in the cell, at the tissue level, or at the system level, whatever—has been a fascinating area for me recently. I've been really interested to see some new projects. I don't know if you've come across these yet, but there are some that are now trying to predict the evolution of essentially the transcriptome or cell state from one timestamp to the next. I think that really suggests a major revolution coming soon.
How much would you say—I don't think there's any answer to this, because I don't think we know how much we don't know—but when it comes to those interaction-type things, my sense has been that we have a relatively small amount of that space illuminated today? Of all the interactions, of all the things where something in this gene interacts with another thing and could cause a third thing downstream, my sense is that we have a pretty small percentage of those pathways mapped out and well enough understood that we could do this kind of analysis. Is that a good summary, or how would you improve on my summary?
Dr. Catherine Brownstein
No, I think that's totally right. Every time I try to look at the impact of a variant on the protein, I'm surprised at how, first of all, user-unfriendly a lot of these tools still are. It's because they're really tough. They're cutting-edge, and protein folding has come a long way. Definitely super cool, and the people who work on that are totally hardcore, but there's still a lot to be learned, and we're still folding certain proteins. We don't have everything worked out yet.
I just keep thinking about when we first got genome sequencing and how difficult it was to use some of these browsers. They would crash the computer. I think protein folding and some of these tools, like STRING—STRING-DB, for protein-protein interaction—they're amazing, and they're going to continue to get more and more amazing and more useful as time goes on. Especially when they get more user-friendly for people like me.
Nathan Labenz
Yeah, it sounds like that might be a real low-hanging fruit. This has come up on a couple of different episodes, where the general observation has been: biologists are not programmers, and doctors are not programmers. There's a missing layer that would unlock a lot of value if we could just make it a lot easier for doctors and biologists to use the models and other information tools that have recently been created. A lot of times, those are still put out there in open-source project form, and they need a UI layer or an orchestration layer on top to really make that accessible and useful for a lot more people. That could be an interesting area for somebody to dig into more.
Dr. Catherine Brownstein
Mm-hmm.
Yeah, and just little things. I got some sequence back from a new company, and they're like, “Okay, here's the commands to download your data.” I'm like, “Whoa, whoa, whoa, what?” They had no intention of helping me, either. I had to learn the command line and how to get my data from their server down to mine, or I didn't get my data.
I had to have a crash course on getting onto the Harvard/Boston Children's supercomputer in order to get my data, and it was a huge waste of time. I think they're assuming a level of literacy for some of these programs that people just don't have. You can argue that I should, being in the job that I'm in, but it's hard. It's a learning curve, and I think there's a lot of opportunity there for making things a little more friendly.
It goes back again to: you don't know what you don't know. If you make your tool accessible to a wider audience, they're going to apply it in ways you never dreamt of. Gatekeeping it to only people who know Unix is kind of tough on everybody.
Nathan Labenz
Let's circle back to that in a second, because this sounds like one of the candidate areas where you might be getting some good value from your o1 Pro Grant. Are these pipelines using any sort of predictive AI technology, like classifiers and things like that, or are they working off a sort of accepted, known literature of findings?
I could imagine—and maybe it varies across providers—that one form of pipeline is, “We want to be really grounded in things that are very well-established, and we're going to run down this super-long checklist programmatically for you and try to find things that fit.” I could imagine another pipeline that would be like, if these models exist—and I'm not sure to what degree they do—you could say, “Hey, here's my genome. Predict and give me guesses.”
Are there models like that? And I guess, to what degree is this all deterministic versus whether those existing pipelines are already starting to lean into certain kinds of AI?
Dr. Catherine Brownstein
I think you need both. You need to be confident that you've looked at a genome with all the known things and that nothing funny was missed—just very validated best practices. Then you need the exploratory pipelines, and that's what we're developing as part of my grant with OpenAI. What's the limit? Where can we take this? Where can we make shortcuts where, before, we were taking a ton of compute and a ton of time? How do we solve cases faster? What's the minimum required data set in order to make a diagnosis? What's the minimum compute necessary in order to get a diagnosis? How do we diagnose new things? How do we come up with new hypotheses faster, all using AI?
Nathan Labenz
Well, that's probably a perfect tee-up for your application of the latest models.
Yeah. Maybe for calibration, before we get into workflow specifics, when did large language models start to be useful for you? Was it just with o1, or were you already starting to see some value with earlier versions?
Dr. Catherine Brownstein
We had been using it along with the phenotyping areas more than anything else. I had a Picory grant working with Ingrid Holman and Melissa Haendel, where we were trying to take a patient phenotype, map it to HPO codes, and get the layperson to HPO faster and more accurately.
One thing that we used at one point was working with 7 questions, asking what system was affected, and drilling down that way. We were seeing the ability to get an accurate phenotype through an interactive model using your own words, compared to traditional self-phenotyping, like surveys and things that are on the web now. We’re still analyzing that.
There are a lot of publicly available tools that I was using, as I mentioned before, like AlphaFold, STRING-DB, and a lot of these protein-impact prediction models that are required to do our jobs. We need to be able to predict the impact of a variant on a protein.
We can’t treat it as gospel. People who rely too heavily on these algorithms sometimes get tripped up because some of the known gene-disease relationships wouldn’t pass those filters now. There’s just something about that gene where you perturb it a tiny little bit and it causes a phenotype that you wouldn’t even think it would cause, but we know that’s true. So, if you looked at it at face value, you would have skipped over it.
I think a lot of people are using these models and don’t even know they’re using them. They don’t really know what’s behind them; they just know that you look at the CADD score, SIFT, PolyPhen, and protein impact, and then that’s a cutoff, along with allele frequency. They don’t really realize that aggregation of allele frequency is powered by a lot of these models, with a ton of stuff happening behind the scenes. If you took it away, we would be struggling.
Nathan Labenz
So, do I have it right, then, that with an AlphaFold-type model, this is after a standard pipeline basically comes back negative? Then you would say, “Okay, let’s go into essentially anomaly-detection mode for this person’s sequence?”
Dr. Catherine Brownstein
Mhm. Exactly.
Nathan Labenz
Yeah. And you have tools for that as well that can say, “Hey, look, here’s a giant deletion,” or, “This gene has stopped prematurely,” or, “This one has been copied over a bunch of times,” whatever. There are, of course, plenty more ways things can be weird than those, but you identify those and then say, “Hmm, I wonder if that maybe is the thing. I’ll use AlphaFold to take that genetic sequence, see what that protein actually looks like, and then do a structure comparison. Does that look like that protein is really mangled?” If so, that becomes a place to go deeper?
Dr. Catherine Brownstein
Yep. Exactly. A lot of that comes with experience, too. There are some genes that are really mutated in pretty much everybody. If you don’t know, you’re like, “Oh, look at that. That’s so cool,” and then some veteran is going to be like, “No, it’s not that. It’s never that.” Or it’s never lupus.
Then you see a gene that you’ve never seen before, and it has a variant in it that’s conserved down to zebrafish and C. elegans worms. You look at it in AlphaFold, and you see that it’s royally messing up the protein, and you get excited.
It’s a roller coaster a lot of times. Even that will fall apart somewhere, and then you’ll find out that it’s only really common in one specific ethnicity that’s hardly ever sequenced, but the patient is from that rare ethnicity. It goes to show that we need to sequence the whole world in order to understand what is actually disease-causing and what is just background variation in isolated populations.
Nathan Labenz
Yeah, there’s another fork in the road here. Which question to ask? We’ll come back to the data, because that is a can’t-miss area, but just take us a little bit further down this path. We’ve identified some anomalies. Now we run the folding model and see that the structure looks off. Where do we go from there? What’s the next investigation after you’ve identified that?
Dr. Catherine Brownstein
Back in 2011, you would get really excited about it and want to publish it.
Nathan Labenz
But in 2024, is the bar always rising, for sure?
Dr. Catherine Brownstein
Yeah, the waterline is rising, and now people are like, “Wait a second. That might just be random.” So then you want other families or other cases with the same type of thing—variants in the same gene. There are all these sharing tools to be able to do that.
One is called Matchmaker Exchange or Beacon, where you put in the variant and the patient phenotype, and you see if anyone else has put in that same gene attached to the same phenotype. Then you match, and you collaborate. Or somebody has already started a paper with 19 cases of variation in this gene causing intellectual disability. If you have one, you can add it to that case series and get a better publication out of it, one that is much more convincing than if you just publish your one case, which looks pretty cool and you’re convinced by, but other people might not be after reading it.
Nathan Labenz
The bar is continually being raised on this stuff. So that brings us back to data naturally. How would you characterize the data environment? I was struck, in reading through a couple of the papers—I don’t have the vocabulary to go as deep as I might wish to on all of your papers—but I was able to see quite clearly that the n is small in a lot of these papers, with single-digit numbers of cases.
I’ve also noticed a few times that you’ve spoken about the hospital as sort of the data unit, it seems like. I’ve heard from a bunch of people over time that we have this sort of scarcity of data, and I’ve always wondered: Is it a true data-scarcity problem, or is it a sort of man-made, for lack of a better term, data-scarcity problem that’s really more about barriers to access and sharing?
Dr. Catherine Brownstein
It’s a tough situation. I don’t want to fault the young researcher who doesn’t want to share their super-cool case because they’re hoping they’ll find another one and be able to publish it as their finding, not as somebody else’s finding in a giant research-group facility across the world, where they’re just going to be a middle author and it’s not going to make their career the way it would if they held on to it tightly, did everything themselves, and got it out there.
The problem with that is that a lot of times it doesn’t work out that way. If that’s not benefiting patients, you’re not thinking of the patient; you’re thinking of yourself. It’s much better for science, and much better for patients in general, if everyone shares their data and has it open. If you see something in someone else’s case, you should be allowed to match it with the group that’s already working on that gene and put it out together.
It’s tough. Boston Children’s is really great in that we have this CRDC, this cohorts committee, where you can see other investigators’ data—patient data and genetic data. Not the phenotype, not their name, or anything identifiable. Sorry, I need to make that extremely clear.
But if you have a gene that you’re working on, you can put it into the CRDC and come up with all the patients who were seen in the hospital and their genetic variation in that gene. The physician has a de-identified ID number, and you can email the physician to find out more information about that patient.
I’ve joined national and international studies that way by having a candidate gene. I go on to the Gene Dx browser now and query the entire hospital—everyone who’s been sequenced and has their data up there. I found 4 other patients, emailed the investigator, and they were like, “Oh yeah, this person in the Netherlands is putting together a case series. Email them.”
I got my patient’s information into that case series, and now it’s awesome. They’re linked to experts, and we’re publishing an accurate, comprehensive view of what that condition looks like. But it’s hard. I understand the dilemma, and for the young investigator who really just wants to get credit for what they’ve been working on, they don’t want to hand everything over. But it’s important that they do, and that everyone does.
Nathan Labenz
You’re identifying a barrier to progress here that I had not even considered, which is the investigator holding information more closely than it sounds like they should in some cases. I guess if we were to imagine an ideal data-sharing scenario, exactly how do we square the circle on sharing versus privacy? That’s obviously a tough question.
Maybe there’s a cryptography-based solution that we could imagine, or maybe we just need to change our norms a little bit around how willing we are to share genetic data. I’ve always felt like it doesn’t seem to me like a huge risk that I’d be taking to share my genetic information with some international pool of information.
There are multiple different angles here, but I guess I’m wondering: If we were to move from today’s data-sharing reality to an ideal data-sharing reality, how much of a difference would that make for people who have these rare diseases?
Dr. Catherine Brownstein
I’m just spitballing here, but I think it would be huge. I think there are a lot of cohorts in the back of the freezer that just haven’t been sequenced and haven’t been shared, more because—not apathy, but because—it’s harder to do so. Also, sometimes at a very superficial level, it’s hard for the investigator to get there mentally and do that.
But I think if they did, there would be a lot more discoveries and a lot more diagnoses for patients, that’s for sure. That’s why I always tell patients—or people, if they email me and they’re like, “Okay, my child has this,”—“Well, here, enroll in this program and this program and this registry.” And they’re like, “Why not just one?” I’m like, “You want to do as much as possible.”
Registries are really important because when there’s a new discovery, they go straight to the registry to find patients. That way, you’re ensuring that your sample isn’t being left in the back of the freezer until they get to it, because you’re just hitting it from multiple sides, multiple angles.
Nathan Labenz
Yeah, is this sort of akin to—I mean, there are a few of these pivot points, maybe, in the medical system where a lot of data is, of course, locked up in electronic health records, and we sort of have this nominal interoperability requirement that somehow gets cashed out as everything getting faxed around. It’s like, “What the hell is that?” That seems like not what we intended, and yet it hasn’t been fixed.
Then there’s price transparency, which is outside the scope of this conversation but is definitely the kind of thing people have high hopes for. If you could get a price menu on the wall, maybe that would help in certain ways. There’s also right to try, which is a big movement where people are like, “You’re not going to let me try this experimental drug even though I’m dying? I should have that right.”
This feels like it could be another candidate for similar reform. If I was going to try to whisper into somebody in the new administration’s ear, I might say, “Hey, look at the requirements around sharing this information. Could we change the defaults here in a way that would move the needle in a big way?”
Dr. Catherine Brownstein
It’s interesting that you say that. Going back to 2011, one thing I lobbied for was shifting it so that being in the biobank—your samples, your discards, tissue, urine, anything that wasn’t used that they took from you—was an opt-out, not an opt-in. I still think it’s an opt-in, how many years later. There’s a lot of inertia around this: being able to facilitate broad sharing, especially for these cases where privacy isn’t really the number-one thing on anyone’s mind. It’s about moving as rapidly as possible and making as many discoveries as possible in a short amount of time.
I really think decreasing the barriers to sharing and to right to try is important. Mew who made Milusen, is 2 floors down from where I’m sitting right now, and it’s just this incredible story of him seeing an opportunity to make an N-of-1 drug and an extremely motivated family breaking down barriers to make it happen. They were so brilliant, motivated, and smart, and they were able to do it. You just think, “Okay, if you made the hurdles less extreme, how much more would be possible?”
Nathan Labenz
That’s an incredible story. If you don’t know it yet, I don’t know it, but here’s hoping that we might have fewer of those stories and more healthy defaults going forward.
Dr. Catherine Brownstein
Those stories are inspirational, but they sort of represent the dark matter of probably 100 other families that just couldn’t, for some reason, overcome those barriers. Some things are just so simple and maddening. We have a bunch of cases at the Manton Center where we find the diagnosis, and then we need to get it confirmed. We do stuff in the research realm, and then you have to get a new sample and verify it in a specialty lab, a Clea accredited lab, and then have the finding returned to the family through a genetic counselor or physician.
Sometimes we’ll call the physician and they won’t play ball with us. They don’t care, they don’t want to deal with it, and they don’t see the value or what it’s going to change. In my own family, I haven’t been able to Clea confirm a finding in one of my relatives because the doctor is like, “Well, I don’t have email. What’s the value of this?” It’s just like, “Oh, my God.” This is what we’re up against.
Then you multiply that by people not counseling correctly and not getting the families into research programs. As hard as we’re trying, there are still so many barriers. To bring it back, I’m really hoping that AI can break some of this down and put some of the autonomy and our ability to act into the hands of families and patients so that they’re less reliant on some of this infrastructure that doesn’t work as well as it should.
Nathan Labenz
Yeah, I mean, this is an eye-opener for me. I think often about whether we’ll end up in a similar spot with respect to AI as we seemingly have with respect to nuclear power, where somehow we have thousands of nuclear weapons deployed, but we’re still burning a lot of fossil fuels because we haven’t been able to get nearly as many nuclear reactors as we have nuclear weapons. Something seems very off about that outcome.
I can imagine an analogous version for AI where we sort of have what we need, but through a combination of errors, barriers, and abstinence, we never quite get to the actual benefits that we could get. It sounds like there is definitely some work to do here to make that change in this area.
Dr. Catherine Brownstein
There are a lot of rabbit holes.
Nathan Labenz
Yeah. So how do you think this changes going forward? We could talk about this from the patient level and what they can do. I always say that if it’s me, at this point I would go with both the human doctor and the AI doctor. I would always have the conversation with Claude or ChatGPT in advance. If they don’t want to talk to me, I say, “I’m preparing for a conversation with my doctor,” and that gets them to open up and not worry about providing unlicensed medical advice.
The patient experience could be quite different. You could talk about that. I’m also really interested in how you’re applying these latest models in your own work—where they’re saving you time and what they’re allowing you to do that you couldn’t do before. Pick your favorite approach for that, but I’m definitely interested in the AI-enabled future of all this.
Dr. Catherine Brownstein
This isn’t really that crazy or anything, but I’d say the biggest impact AI has made on my research is summarizing articles and genes. Being able to eliminate the time I spend going down rabbit holes—looking up a paper, realizing it’s paywalled, logging into the Harvard library, getting the paper, skimming the abstract, and finding that it’s not at all what I want—has changed my life. Being able to ask for a summary and get it, and either be like, “Oh, yeah, this sounds good,” or move on with my life, has given me hours back in a day.
I think there’s going to be a whole host of new tools, or new reasoning. I find it funny that sometimes it will clamp up and doesn’t want to do something because you’re getting too close to medical advice. Maybe just because there are specialty things that help, it would be really cool if you didn’t have to ask the same question 4 times to get it to answer.
Boston Children’s also launched ChatGPT behind the BCH firewall, which is great because then you’re not worried about things going out, and they’re able to maintain much more control. It stays much more accurate. I still can’t get citations to work properly, which is kind of hilarious, but it’s getting way better. The hallucinations are getting way better. I just think it’s going to be moving at light-year speed.
Going back to what we were talking about before, I think there’s a lot of fear around it that’s going to have to be addressed. Hopefully, the 1 bad situation isn’t going to be the only thing people read about it, and some of the really great things that come out of this will also be properly publicized to give a more balanced viewpoint. Again, keeping in mind that a lot of times these are really severe cases and really severe patients, they’re making huge strides and having a huge impact. Keeping that in perspective is really important, too.
Nathan Labenz
Tell me more about some of the things that you actually throw into ChatGPT. You mentioned 1: here’s my situation and here’s this paper, almost like relevance filtering—“Is this relevant?” What other sorts of tasks do you find yourself bringing to especially the latest models?
Dr. Catherine Brownstein
I also run the core facility here, so I’m tasked with learning a lot of new genetic techniques really quickly. If something comes up and I don’t know what they mean, I could Google it and find the 1 obscure paper. I could put it into ChatGPT and learn about this new type of sequencing that’s only launched at Children’s and has 1 paper attached to it, and get a nice summary that I can understand as opposed to weeding through everything.
I meet with investigators all the time, and being able to summarize their work really quickly allows me to do a much better job in my one-on-one consultations than I would have otherwise. Also, considering there are close to 20,000 genes, anytime I get a case where they think it’s this, sometimes I know what that is, and other times I don’t. I’m able to print out a summary of the condition really quickly and nicely, get the latest information on it, see who’s working on it, and go into a meeting much more prepared in much less time.
Also, when you get a paper back, a lot of times—for some reason, it always seems to be reviewer number 2—is like, “There’s a whole body of literature on this,” and you don’t really know what they’re talking about. Being able to address some of the critiques and put them into context is really helpful.
I mean, it's all cutting down on this mundane, time-consuming, really tedious part of the job, and I'm getting back to the fun part, which is gene discovery and going through a list of 20 possible candidates and narrowing it down to 3 that you're going to present in an hour and a half. True story.
Nathan Labenz
So, what's that true story, maybe in more depth? Is that another thing where you're using ChatGPT to help?
Dr. Catherine Brownstein
Yeah. Why not? If you have 20 genes and you have the phenotype, and they all seem pretty interesting, you can go through and look at the protein impacts, so order the CADD scores or conservation and be able to order it that way. But then doing a really quick relevancy assessment using ChatGPT saves a lot of time.
Nathan Labenz
So, how do you set that up? Do you have a prompt template that you go back to over and over again? How much have you had to develop that? How much do you have to give in terms of detailed instructions or examples?
We're getting into the nitty-gritty here, but this is the part where I think both people can hopefully learn from your experience. If nothing else, demonstrating that this is possible is quite useful, because there are just so many people, including software developers. You'd be amazed—maybe you have seen this—but you'd be amazed by how many software developers tried GitHub Copilot 18 months ago, when it first came out with the GPT-3.5 model behind it, and were like, “Eh, it wasn't that good. It can't help me.”
So, I think there's just a lot of value in object lessons of, like, here's hard work that highly skilled, highly educated professionals are doing that ChatGPT—or obviously other models, perhaps similarly, but we're focused on ChatGPT in this case—can really help with. So, yeah, I love just as much detail as you can get into in terms of how you actually go about setting these things up, how you've iterated on them, et cetera.
Dr. Catherine Brownstein
Okay, so, for an example, I work on bladder pain, undiagnosed bladder pain in individuals. It's really severe. Sometimes they can't leave their house. It's called interstitial cystitis, bladder pain syndrome. There's no real gene attached to it. We've found a couple of genes where it seems like there's way more variation in those genes than you would expect, given the general population. So, it's a candidate gene. It's in no way a slam dunk, but I have around 500 patients in a cohort with that.
I've done, in conjunction with Josh Motalo at Columbia and Ali Gharavi, assessments of my cohort and other cohorts to see what genes have way more variation in them than you would expect. You can come up with lists, and then you can also come up with gene pathways, like multiple genes. These pathways are interesting because a lot of times they have a label, like the small-molecule transport pathway. There's like 12 genes in it.
Then you want to know: Are any of these genes tied to bladder pain? Are they tied to the bladder? Are they tied to bladder cancer? Are they tied to anything? Being able to ask those questions really quickly—and sometimes it's a simple yes or no, just putting them in a string and then coming out with yes or no—saves a huge amount of time.
And then the ones that are yeses, you can drill in. I always check the notes, too, just in case. It's still early yet, but I was doing that last night and was able to get through these pathway lists and be like, all right, this one has 60% of the genes that have a tie to bladder cancer—specifically bladder cancer—which means that they're expressed in the bladder and there are known perturbations that cause bladder dysmorphology or bladder conditions. So, this is more interesting than anything else.
One, I almost screamed because the gene was linked to urothelial issues, which is a great mechanism of disease, and I'm definitely going to follow up on that. I have a meeting tomorrow morning to discuss it. So, it really just helps. I only started working in genetics after the genome was published, so I don't know how people did it beforehand, and I think there's going to be this whole generation of geneticists who aren't going to know how things were done before all this was available, because it's going to be a huge game changer and time saver.
Nathan Labenz
So, how much difference would you say you see between, for example, GPT-4o, o1, and o1 Pro when you bring those kinds of questions? Because I can see interesting different trade-offs, right? In ChatGPT today, if I recall correctly—maybe they've just updated this—but certainly with GPT-4o, you can enable web search. With o1 Pro, search is unavailable.
Dr. Catherine Brownstein
That's what I thought, and that is still the case.
Nathan Labenz
So, if you have these sorts of questions, GPT-4o could go out online and find information that's maybe more recent than the knowledge cutoff, which could be really useful, but it isn't going to reason about it in the same way. With o1 Pro, you have more reasoning, but you have knowledge-cutoff issues and an inability to go out and supplement at runtime.
Do you have a taxonomy of what models you use for what things, how you know when to trust what it's saying versus when you need to fact-check, and how much the reasoning adds over 4o for your purposes?
Dr. Catherine Brownstein
I think I'm becoming more and more convinced over time that this is going to revolutionize things. I was skeptical at first. I was like, oh, we're going to have to check every single thing. Is this actually saving any time? It's just getting more and more accurate. The reasoning is getting better, and sometimes you'll be so pleasantly surprised.
You'll ask it a question, and it'll say, “Okay, answering it in the form of a genetic counselor is this.” And then it'll completely surprise you and be like, “Another way to look at it is this.” It's doing an amazing job.
I know I'm a convert and a relatively early adopter, but I think the sky's the limit, really, and we're going to get to a place where it's going to be solving cases, shortening the diagnostic odyssey, democratizing access to genetic interpretations, and sidestepping a lot of the barriers that we have right now.
It just needs to convince everyone that it's accurate and that the reasoning is good a high percentage of the time. It's kind of hypocritical, in a way, that I think we're going to have a higher bar for it than we do ourselves. We can say, like, “Oh, sorry, I missed it. I shouldn't have,” and we're not going to forgive it if it misses something. I guess that's the way it should be.
Nathan Labenz
Yeah, I'm not sure if that's the way it should be. It does seem like it's the way it is. In self-driving, my general working assumption is that it's going to have to be 10 times safer, or have 1/10 the danger rate, to be acceptable to people. I would guess probably something similar will happen here, at least when it comes to actually putting it in a more forward-facing role where patients could access these sorts of things themselves.
If it's a tool for the professionals, then maybe we're a little bit more—put the responsibility on the professional—and can use it earlier. But, yeah, I would probably advocate for going for it before it gets to 10 times better. Nevertheless, that does seem like the sort of mentality that we have.
So, just honestly, for me—maybe for the audience, but for my benefit—how are you managing those trade-offs between needing to go out and search? Because if you wanted to use o1 Pro, you'd have to go do your own search, copy and paste it in, and let it do its thing. GPT-4o can do its own web search. So, in the very nitty-gritty, what model do you go to, and how do you set it up for success?
Dr. Catherine Brownstein
I'm not using the web search right now. I'm more using o1, I think. I mean, I'm playing with 4o. It's moving so quickly that there's no sophisticated reason for that. It's just what I'm comfortable with, and then moving from there.
I'm really impressed with 4.0 reasoning. I think web search still kind of scares me a little bit, just because there's a lot of garbage on the web, and I have to really be confident in any answer I'm getting out. I think checking everything is still paramount here, but hopefully it won't be that way in the near future.
Nathan Labenz
How long does it tend to think on the questions that you're giving it?
Dr. Catherine Brownstein
At first, I think it was shortening, too, by the day. At first, I remember it would just be hanging there for a while. I'd be like, “Are you okay?” Now it's just really fast.
Nathan Labenz
Or maybe under a minute in most cases, it sounds like.
Dr. Catherine Brownstein
Also, I think I'm getting better at the prompts. As you said, you have to learn how to ask it things, too, for it to come out with the right answer right away.
Nathan Labenz
Yeah, I would love to—we'll trade. One thing that I imagine you've probably also found, but I've definitely found, even in low-stakes situations, is that I try to be really neutral in the way that I ask questions. One of the most common failure modes, at least from what I experience, is the model running with a preconception, which might have been a misconception on my part, and mirroring that back to me.
I'm not doing genetic analysis, even in terms of how to solve a programming problem or how I should think about architecting my application or whatever. A lot of times, if I give it a sense of where I'm leaning, it will lean in that direction, too, perhaps without good reason. So, that's one. What else have you found to be important in prompting?
Dr. Catherine Brownstein
I've found that it actually thinks a little too much, where I'm like, “Is this gene related to this phenotype?” It'll bring up a study, and I'll look at the study, and it's the gene related to it.
Is this step too far down the chain? It's really impressive because it made that intellectual leap, but I need something simpler. I need a paper that's just linking that gene to that phenotype. I was like, “No, too far, too far.” How do I ask it so it's not thinking as much?
Nathan Labenz
And when you describe that, it's bringing up a study out of its pretraining knowledge. You give it a question, and it says, “So-and-so et al. found this.”
Dr. Catherine Brownstein
Yeah.
Nathan Labenz
It sounds like it is also marshaling knowledge of the graph of interactions and saying, “Well, this paper showed this, and then I know from other...” And I'm like, I'm not making that case in the paper. I just want you to say it's upregulated in cancer. That's all I want.
Dr. Catherine Brownstein
That was actually kind of wild because then you're like, “Okay, it's thinking. It's really thinking and making conclusions.”
Nathan Labenz
That is quite interesting. Do you think those things are real and we're just not there yet, or is it going off in a direction that is fundamentally not super productive when it does that?
Dr. Catherine Brownstein
That's the million-dollar question. I don't know. Maybe I'm not smart enough to understand it, and it's right. I don't know. We'll find out. We just need to keep playing with it, keep working with it, and keep using it.
Nathan Labenz
It sort of is like the Move 37 equivalent.
Dr. Catherine Brownstein
Yeah, of course.
Nathan Labenz
I'm sure you're familiar with the AlphaGo championship from years ago, where Move 37 is AI shorthand for an output from an AI system that is very surprising to human experts and nevertheless proves to be a genius move. It was one of these moves where it was like, “Oh, wow, this thing is playing Go in a way that we never thought Go could or should be played, and we actually have something to learn from this system.” They initially thought it made a mistake, and then it turned out it was a genius move. Do you think there's at least some possibility that some of these weird analyses you're getting back might be Move 37-like brilliance, but we just don't yet have easy ways to resolve whether it's going in the right direction?
Dr. Catherine Brownstein
I have faith in it. I think we just have to keep an open mind and keep playing with it and see what it can do.
Nathan Labenz
How much data do you have? Do you actually throw whole cases in?
Dr. Catherine Brownstein
It's kind of too hard to do that right now. I'm not throwing in a medical record, even if it's behind the firewall. I'm just summarizing. We're building up to see the limits. That's actually part of the grant that I have with OpenAI: to see how far we can take this, how much it can handle, and how much it can replace me.
Nathan Labenz
And the barrier to doing that right now is the context? I could imagine multiple different reasons that it might not work to just take the simplest thing I would try. But I'm sure there's going to be a barrier. If I just said, “Okay, here's my whole medical record and my whole genetic summary of all the strange variations,” took the top chunk of that file, copied and pasted it, and said, “Analyze this,” what makes that not viable today?
Dr. Catherine Brownstein
I mean, the amount of compute needed for that, and then also the opportunity for tangents and all the utility. Not everyone who smokes gets cancer, so we have to figure out what we're asking, what's relevant, what's meaningful, and what's a good use of resources. I think we're forgetting that every question takes energy.
We can't just throw everyone's medical record in there and everyone's genome and see what comes out. We have to be thoughtful about it and see, in the cases where it does have utility, what information is necessary. And also, to your point, I think it'll be really interesting to look at trajectory and predictions. All the medical-record mining we're doing now—can it do it on steroids and come up with predictive models and interject, “Okay, I know normally you would want to see a colonoscopy at 40; maybe you need one at 24,” just based on genetics and everything that it's able to see that we're not smart enough to see yet?
I think the possibilities are really exciting. It's a really exciting time.
Nathan Labenz
Yeah. Sam Altman recently said they're losing money on the o1 Pro subscriptions even at $200 a month. So it sounds like people may not, in general, be conscious of the compute they're consuming and are just throwing a lot at it. The budget's got to be pretty high, right?
Dr. Catherine Brownstein
Yeah. I mean, from the medical system—from what people would be willing to pay, or what insurance is prepared to pay, compared with a $200-a-month o1 Pro subscription—I assume that looks very cheap by comparison with hiring professionals and teams of people like yourself.
Nathan Labenz
Yeah, and people like using it. Everyone I know loves using it. I don't know if that's a random sample. So what else is going on with this grant and your relationship with OpenAI? Are you working with them closely and iterating on use cases and giving them feedback, or what is the dynamic there?
Dr. Catherine Brownstein
Yeah, exactly. They've been wonderful and super cool, and it's fun meeting and working with smart, motivated people. I get emails at 11:00 at night on a weekend. They're working hard; there's no doubt about that.
We're just getting back into it after the holidays, so hopefully, in a few months, I'll have something really exciting to talk about. I'm just blown away that they're so forward-thinking and able to support this type of work. It's happening a lot faster than I would have thought, even just a couple of years ago, that's for sure.
Nathan Labenz
Maybe in terms of wrapping up, if I try to summarize everything here, it sounds like we have a data-sharing problem. We have a limited capacity for analysis as humans. One of those is going to require a non-AI solution; the other one, AI is increasingly ready and able to do a lot of analysis.
But then you still have a number of practical issues around the knowledge cutoff and search. You do your own kind of curating of the search, and you can't throw everything into it because maybe it's a little too big for the context window. You don't have all the workflows you might like because it's all in a browser, so you have to paste stuff in and get stuff back, and it seems like it's still fairly manual.
This is almost like your OpenAI customer interview, but what do you imagine the experience being like, say, a year from now, when we refine things a little bit more and integrate these systems a lot more? What do you think that could be for you and for patients?
Dr. Catherine Brownstein
That's a great question. I think they might wrap things so it's less free-form, so you'll be able to guide patients and make it user-friendly with a point-and-click interface. You're not just working with a prompt and then meant to come up with the correct question to get it to answer what you're thinking about. I think that's pretty low-hanging fruit and will be very useful for patients.
I think for researchers, too, with prompt engineering and use cases, we're still figuring out, at least at our institution, where it will be most impactful and what people want to use it for. So that's going to be clearer in a year. It's just getting people in the door and communicating. We have surveys, like, “Okay, what would you use it for? What have you been using it for?” and clearing that up, and then making that better.
I'm trying to be realistic here. I would love to say that we're using it to solve cases. I don't know if that'll be true, but I hope it is. I think it's just going to be more intertwined in our day-to-day existence. What do you think?
Nathan Labenz
All bets are off. I don't know. I mean, o3—that's the other question I had. Are you on the review team for o3 at this point? If not, I assume it'll be coming your way before too long.
Dr. Catherine Brownstein
Hope so.
Nathan Labenz
Yeah. It looks like that is another significant step up in raw reasoning ability. The FrontierMath results in particular were notable. Everybody was citing the 25% success rate, but that's the very high-compute level that costs maybe thousands of dollars a problem or whatever.
It was also really notable that the low-compute setting was still 10%, which was 5 times better than anything that had come before, with the previous best maxed out at 2%. So it seems like the o3 series is going to be another pretty serious step change in terms of just how hard of a problem these things can solve.
We might start to get into context-window limits being binding if there's just too much information in a medical history or in a genetic file. But I suspect that those can both be filtered and summarized and boiled down to what matters most, such that even in a couple hundred thousand tokens, which is what they currently have, I would honestly take the “we probably will be solving cases with an o3 in a year's time” side of that—or, for that matter, an o4—because the gap in time between an o1 and o3 was so small.
The signals we're getting are that we don't really see this slowing down. There's going to be more progress on this front. So I'm always kind of wondering what's missing, and increasingly it's harder and harder to find and pinpoint the things that are really missing. That doesn't mean there's nothing missing. I'm sure there are still some things, but it is increasingly hard to say what they are.
My best guess would be that you probably see at least some cases that you could just throw into an o3 and get meaningful, insightful conclusions back in a year.
Maybe we should get together again in a year and review the progress.
Catherine Brownstein
I’d love that.
Nathan Labenz
Cool. Well, anything else on your mind today? I really appreciate the introduction to all your work and how you’re using AI in it, but anything else on your mind before we break?
Catherine Brownstein
Just thank you so much. I think this time next year might be kind of different, so hopefully.
Nathan Labenz
That seems to be the new normal. Change is the only thing we can really count on.
I’ll look forward to putting it on my calendar now to get back together in a year and see where we’re at. But for now, Dr. Catherine Brownstein, MPH, PhD, assistant professor at Boston Children’s Hospital and Harvard Medical School, and a recent recipient of the ChatGPT Pro Grant, thank you for being part of The Cognitive Revolution.
Catherine Brownstein
Thank you.