[BidClub_]
Gradient Dissent · · 75 min

Curing Every Disease With Al by 2050 | Sam Rodriques, Edison Scientific

Lukas BiewaldSamuel Rodriques

YouTube
TL;DR
  • Rodriques’ core thesis is that biology is constrained by scarce scientific talent, while “capital scales, and logistics scales, and talent just does not scale.” AI can relieve that bottleneck through high-throughput reasoning: reading more evidence and testing more hypotheses than any researcher could. It cannot eliminate physical validation, especially when the decisive experiment is a multiyear human clinical trial.

  • FutureHouse’s first full-loop agent generated a treatment hypothesis for dry age-related macular degeneration that progressed from wet-lab validation into animals and, two days before the interview, publication in Nature. Rodriques calls that May 2025 result the moment “the future is here.” Its successor, Cosmos, has probably been used to generate 20,000-30,000 novel findings, though that figure describes findings proposed rather than approved medicines delivered.

  • Biewald frames the commercial tension as “progress comes from discovery, but the commercial value is in development.” Rodriques says discoveries pan out infrequently and may not reveal their value for a decade, while faster experiments can move medicines toward patients. AI can help prepare protocols and regulatory documents, coordinate trial sites, and move compounds through the pipeline; molecule-design specialists can simultaneously improve candidates. Rodriques says future pharma companies will be much leaner and pursue many more programs with the same headcount, but human trials remain the binding physical and operational constraint.

  • Edison Scientific’s proposed moat is specialized scientific reasoning, customer-specific training, and last-mile deployment—not an attempt to outperform frontier labs at everything. Rodriques says small amounts of task data can produce “enormous gains” over general models in niches such as synthetic chemistry, while proprietary pharma data preserves customer differentiation. His hedge matters: specialization may lose its advantage if intelligence reaches task saturation, and coding has not yet shown that a specialized model can beat the strongest generalists.

  • The platform deliberately exploits the frontier models’ non-overlapping strengths instead of betting on one provider. Rodriques says Anthropic led Edison’s evaluation for reproducing analyses from scientific papers, while only Gemini 2.5 could perform one early Cosmos world-model task. This multi-model architecture also answers pharma’s reluctance to become locked into OpenAI, Anthropic, or Google while leadership keeps changing.

  • Rodriques pairs enthusiasm for faster experimentation with deep skepticism toward uncontrolled biohacking and a provocative case for clinical-trial reform. Unstudied peptides may not work, may merely produce placebo effects, or may cause untracked harm; “no one is looking” for adverse events among self-dosing communities. Yet onerous trials help create that behavior, so he favors decentralized early-stage approvals like those he describes in Australia and China and, conditionally, loosening efficacy requirements after safety is proven.

  • Despite declaring that this AI cycle is different, Rodriques sets a hard industry scoreboard at pivotal Phase 3 trials and FDA approvals. AI drug discovery has overpromised since roughly 2012, with AlphaFold the notable exception, and pharma insiders will discount announcements that a model “came up with” a drug. “The proof is in the pudding, and the pudding is approved drugs.”

Digest · the substance, structured for research

1. Biology’s fixed supply of talent drove the move from physics to AI

  • Rodriques jokes that physics has “actually no unsolved problems left,” then supplies the serious distinction: most visible phenomena can be explained down to subatomic particles, but not why a mouse behaves as it does, why people become sick, or how the brain works. “Most things are pretty well understood except for biology.”

  • His MIT experience reduced scientific production to three inputs: capital, logistics, and talent. Money and experimental infrastructure can be expanded, but “talent just does not scale”; curing disease, understanding the brain, and addressing aging therefore require a way to multiply researchers’ reasoning capacity.

  • Leaving a promising neuroscience and bioengineering career felt like throwing it away and “jumping off a cliff.” AI was attractive not as an end in itself, but as the most plausible mechanism for removing talent as science’s bottleneck.

2. A nonprofit was the right incubator until demand arrived years early

  • In 2022, with GPT-3 and InstructGPT but no reasoning-model paradigm, Rodriques expected an AI scientist to require five or 10 years and saw no obvious business model. He and agent pioneer Andrew White therefore created FutureHouse as a nonprofit for basic research and broad scientific benefit, including open-source releases.

  • The structure followed Rodriques’ earlier proposal for focused research organizations: goal-driven nonprofits for projects “too big for academia” but uneconomic for venture capital. Mapping a brain’s microscopic neural fibers, for example, requires company-scale engineering yet might take 20 years to yield drug value; an FRO can pursue one defined goal and then wind down.

  • FutureHouse was announced in November 2023, but powerful agents appeared within roughly two years rather than a decade. By spring 2025, pharma executives had moved from needing agents explained to calling for deployment, making the Edison Scientific for-profit spinout a means of scaling successful nonprofit research rather than evidence that the nonprofit model had failed.

3. Robin supplied the proof point; Cosmos multiplied the search

  • Robin was a manually orchestrated multi-agent system built for drug-repurposing hypotheses. It could generate a hypothesis, plan experiments, analyze data returned from the wet lab, and propose the next experiments—the “full loop of scientific discovery,” with humans still executing the physical work.

  • In May 2025, Robin proposed a new treatment approach for dry age-related macular degeneration, which Rodriques says affects roughly 5%-10% of people over 50. The team validated it first in wet-lab experiments and later in animals; the paper appeared in Nature two days before the recording.

  • Cosmos replaced Robin’s on-rails workflow with an orchestrator able to coordinate its own subagents. Its “world models”—described less mystically as sophisticated context management—integrate field knowledge accumulated across dozens or hundreds of runs and use that evolving representation to make discoveries.

  • Rodriques estimates users have generated 20,000-30,000 novel scientific findings since Cosmos launched. That scale is central to the pitch, but the episode maintains the distinction between generating a potentially novel result and experimentally proving that it works.

4. AI excels where answers are checkable or throughput dominates

  • Rodriques readily concedes the models’ “extremely” spiky intelligence: after asking one to search Notion and Slack and draft a press release for an upcoming deal, he found the result terrible. The practical strengths instead cluster around verifiable tasks and problems where evaluating far more evidence or hypotheses creates value.

  • Coding and mathematics provide rapid feedback, while science is only verifiable in principle. Closed-loop reinforcement learning for drugs could require proving safety, dosing humans, and waiting six months—a feedback loop perhaps three years long—so the familiar recipe of abundant data and fast rewards does not transfer cleanly.

  • A giant robot laboratory helps only when its automated experiment reflects the real objective. Pipetting robots can run useful screens, but a system trained on laboratory proxies cannot determine whether a medicine works in humans: “Your model is only as good as what you train it on.”

  • What Rodriques calls “high-throughput reasoning” fills a different gap. Agents can interpret biological context across evidence that no individual could survey, provided researchers control multiple-hypothesis testing and P-hacking and ultimately run the required experiments.

5. Parasites and bacterial operons show what high-throughput reasoning buys

  • To search for autoimmune treatments, a FutureHouse fellow began with nature’s own immune suppressors: parasites that evolved to dampen human immune reactions. Agents examine every candidate protein across parasite genomes using sequence, structure, genomic context, and organism biology, then reduce thousands of possibilities to a laboratory shortlist.

  • The selected parasite proteins were synthesized on a DNA chip and screened for effects on T cells. The example captures the division of labor: machines make an otherwise impossible search tractable, while physical assays determine whether the hypotheses survive contact with biology.

  • Bacterial proteins provide a second reasoning task. When a protein of unknown function sits in an operon beside one with a known function, an agent can infer the operon’s biochemical role and propose what the mystery protein does—potentially identifying enzymes or bioengineering tools rather than merely matching sequences.

  • Biewald’s insistence that testing still requires “something physical” is accepted without qualification. Rodriques’ claim is not that science can be reasoned into completion, but that hypothesis generation itself is limiting and can now be expanded dramatically.

6. Development captures the value even when discovery creates the progress

  • Customers use the agents both to generate early hypotheses and to operate development: ordering experimental materials, assembling protocols, coordinating clinical-trial resources and sites, preparing regulatory documents, and moving a drug toward patients.

  • Biewald’s commercial framing is blunt: “Progress comes from discovery, but the commercial value is in development.” Rodriques says discoveries fail frequently and may not reveal their worth for 10 years; faster experiments, especially human trials, are what convert ideas into both scientific and economic value.

  • A separate AI capability can design the actual molecule after a biological mechanism is chosen. For an oral candidate, it needs suitable bioavailability, sufficient lifetime in the bloodstream, correct receptor activity, suitable distribution, metabolism and pharmacokinetics, tolerable toxicity, and minimal off-target effects—before confronting the possibility that the original mechanism was wrong.

  • Rodriques cites Chai Discovery, Isomorphic Labs, Boltz, Lambda Labs, and Profluent Bio among companies attacking molecule generation. Pharma companies may have strong internal AI teams, but training the world’s best antibody-generation model “is just not the pharma company’s game” in the way it is for a focused model builder.

7. Leaner pharma still collides with a human bottleneck

  • Rodriques says future pharma and biotech companies will pursue many more drug programs in parallel with the same people. That is the operational meaning of removing the talent bottleneck, and it could support a medicine-scale project analogous to the Human Genome Project across diseases currently too uneconomic to pursue.

  • Biewald presses an apparent contradiction: if clinical trials are the bottleneck, why emphasize candidate generation? Rodriques answers that drug development contains multiple bottlenecks, but the ultimate operational constraint is still people—operators to conduct trials and patients in whom efficacy must be tested.

  • In most areas, subject to diseases where eligible patients are scarce, twice the money and twice the operators could support roughly twice as many ideas.

  • Biewald notes that capital can scale if more good ideas are investable. Rodriques’ response is that the limiting factor is often the people and operators needed to advance those ideas through trials.

8. Peptide skepticism and trial reform are two sides of one argument

  • Asked about his peptide stack, Rodriques answers that he has none because he “know[s] too much about biology.” Studying drug development teaches how many subtle things can go wrong and how strong placebo effects can be, making him skeptical of compounds never tested in robust, controlled trials.

  • Feeling better after a peptide proves little: it does not establish whether the peptide is really working or whether placebo is responsible. More seriously, even a systematic 10% incidence of cardiac arrest could remain invisible because “no one is tracking those adverse events” among people who self-dose.

  • Biewald says he admires biohackers’ willingness to try things. Rodriques nevertheless calls uncontrolled self-dosing generally ill-advised; the problem is experimentation without a clear view of uncertainty, possible harm, or reliable measurement.

  • The flip side is that slow, onerous clinical trials push people toward uncontrolled testing. Rodriques says Australia and China decentralize some early-study approvals to individual centers, enabling regulated competition on speed, while even small US trials require central FDA approval; US biotechs consequently take studies abroad.

9. Looser efficacy rules invite both personalization and statistical abuse

  • Rodriques proposes considering approval based on demonstrated safety without always requiring proof of efficacy for one predefined condition. A drug that helps a genuine subgroup can fail its aggregate trial; after safety is established, real-world use might reveal who benefits, especially when standard care has already failed that patient.

  • Within-patient comparisons could provide evidence: if one antidepressant leaves symptoms unchanged and a second produces a marked improvement, the changed medicine is informative and partly avoids the placebo effect of moving from no treatment to any treatment. Comparisons against an established drug could also expose superior outcomes.

  • Biewald’s pushback is that retrospective real-world analysis invites P-hacking and post hoc subgroup selection. Rodriques agrees the risk exists but frames it as statistical power: a pattern found after searching many subgroups can be spurious, while a consistent vaccine effect confined to men aged 18-35 across 10 million people signals something worth investigating, even if confirmatory experiments remain necessary.

  • Biewald also asks how often real-world data reaches that scale without selection bias. Rodriques points to population-scale vaccines and common cancers, while acknowledging formidable data-quality problems: patients skip doses and follow-ups, clinicians write uneven notes, and records fracture across health systems. He says Edison already works with clients on such analyses but that existing real-world evidence is “never as good as you want it to be.”

10. Medical decision support trades comforting answers for explicit uncertainty

  • Biewald describes medicine as high-stakes yet persistently ambiguous: studies conflict, datasets carry different biases, and patients increasingly ask language models to synthesize research for individual decisions. Rodriques says friends and family already use Cosmos this way.

  • One friend facing a breast-cancer scare asked what delaying treatment by one or two months might do to recovery or survival odds. Rodriques says the platform surveys literature “dispassionately” and grounds its answer in papers rather than “vibes”; the friend ultimately turned out to be fine.

  • Biewald recounts entering the same description of his daughter hitting her head into two accounts: his model said she was fine, while his wife’s advised seeing a doctor, perhaps reflecting conversational history or reinforcement toward desired answers. Cosmos would more likely return explicit probabilities of serious outcomes and of being fine—more grounded in evidence, but sometimes more frustrating.

  • The larger tension is that abundant intelligence does not automatically improve judgment. “More information is not always better for making decisions”; users need the right information, not necessarily the maximum amount.

11. Specialization, proprietary data, and model diversity form the moat

  • Against OpenAI, Anthropic, and DeepMind, Rodriques argues that a specialist should beat a generalist because the latter must encode everything in its weights. Edison trains reasoning models for specific scientific tasks and claims that small datasets can generate enormous gains in niches such as synthetic pathways, while spending more test-time compute on suppressing hallucinations.

  • Biewald challenges the thesis with coding. Rodriques says many people are trying to make specialized coding models beat the frontier generalists but have not yet done so; coding may be exceptional because every major lab prioritizes it. He offers a second hedge: once a task reaches “intelligence saturation,” as humans have with optimal tic-tac-toe, added specialization or intelligence may stop helping.

  • Pharma’s proprietary datasets strengthen Edison’s position because customers want models trained on their own evidence, not identical intelligence available to every competitor. Edison combines that customization with last-mile internal deployment, work frontier providers may avoid because they do not want a different model for every customer.

  • Customers also resist dependence on one frontier vendor: OpenAI appeared far ahead one year, while Anthropic felt ahead at recording. Rodriques says the models have substantially non-overlapping spikes—Anthropic was particularly strong on a benchmark reproducing paper analyses, while only Gemini 2.5 handled one early Cosmos world-model task—so combining providers can materially improve results.

12. Long-running agents evolved toward steerability, not smaller ambition

  • The original Cosmos could run for six to 12 hours, write 45,000 lines of code, and read 1,500 papers in one run. Rodriques calls it among the most intensive agents available, but recognizes that scientists rarely want to submit a question and disappear for half a day.

  • A recent update added an interactive front end without discarding long-horizon work. Scientists can now give feedback and steer the investigation while the underlying system continues orchestrating extended research.

  • This evolution reflects a broader maturation from demonstrating an agent framework toward delivering usable scientific infrastructure. Rodriques still claims Edison’s agents lead frontier-lab offerings in science, but increasingly describes the value through deployment, proprietary training, workflow fit, and pipeline acceleration.

13. Scientific productivity is slowing, but human taste remains valuable

  • Rodriques invokes Eroom’s law—Moore spelled backward—for the historical rise in real drug-development cost, tentatively recalling a doubling roughly every nine years. He says the deterioration broke around 2011 or 2012, perhaps because the Human Genome Project made targets easier to find, but productivity has since been roughly flat rather than rising.

  • One explanation is depleted low-hanging fruit; another is the “better than the Beatles” problem. A new therapy must beat the existing best-in-class treatment, yet medicines cannot always compound prior mechanisms as semiconductors compound engineering advances; first-in-class mechanisms often deliver the dramatic gains, and each subsequent discovery becomes harder.

  • Rodriques still recommends a PhD because its purpose is learning how to research, and he expects human scientists to matter for a substantial period. Science is not cheaply verifiable, so humans contribute taste—especially deciding which problems are interesting—even if models become better at proposing ideas likely to work.

  • His uncertainty stays explicit: within five or 10 years, he does not know whether models will surpass elite researchers’ judgment. At recording, though, he sees “no question” that Nobel laureate Frances Arnold would propose better directed-evolution research than Claude 3.7.

14. Seventy million dollars buys deployment; approvals decide credibility

  • Edison’s for-profit company raised $70 million, beyond the nonprofit’s separate funding. Rodriques values investors who can open pharma doors: Spark’s Yasmin Razavi brings frontier-AI perspective, former Grail CEO Jeff Huber brings extensive industry relationships, and an unnamed large institutional biotech investor helps secure executive access and deployments.

  • The contrast in diligence is revealing: tech investors repeatedly ask why OpenAI, Anthropic, or Google will not absorb Edison’s market, while pharma insiders rarely ask. Rodriques’ interpretation is that operators who know drug development “viscerally know where the value is” and see the specialized workflow and deployment problems directly.

  • He acknowledges a decade-plus of AI overpromising in drug discovery, with AlphaFold the notable success in protein-structure-enabled molecule work. His categorical claim—“I promise you this time it is different”—is immediately disciplined by a hard endpoint: proposed discoveries will not impress pharma until a pivotal Phase 3 trial succeeds or an FDA approval letter arrives.

1. Using AI to synthesize medical research

Samuel Rodriques

If we want to go and cure all diseases, understand how the brain works, solve aging, and so on, AI seemed like the right way to do that. When it comes to the world as we experience it, most things are pretty well understood except for biology. To this day, we don't have the wiring diagram for the human brain. We kind of have the wiring diagram for the fly, and that's the best that we have.

When we got started in 2022, remember, we didn't have any notion of reasoning models or whatever, right? We showed the first multi-agent system that was capable of doing the full loop of scientific discovery, and it came up with a new hypothesis about a way to treat a form of blindness called age-related macular degeneration. It actually just got published in Nature 2 days ago. Since we launched Cosmos, people have probably used Cosmos to make 20 or 30,000 novel scientific findings, which is wild.

The pharma companies of the future will be much leaner. You're going to be able to pursue many more drug programs in parallel than you can today with the same number of people.

Lukas Biewald

There's a long history of AI overpromising in drug discovery. Is this time different?

Samuel Rodriques

Oh, yeah. I promise you, this time it is different.

2. Introduction

Lukas Biewald

You're listening to Gradient Descent, a show about making machine learning work in the real world, and I'm your host, Lukas Biewald. All right, I'm here with Sam Rodriques, the founder and CEO of Edison Scientific and Future House. He's right at the forefront of using agents for scientific innovation, and he's got a lot to say on how to use agents well and science itself. I really hope you enjoy this podcast.

All right, Sam, thanks for doing this. I've really been looking forward to it.

Samuel Rodriques

Likewise.

3. Origins and the role of nonprofits in technology

Lukas Biewald

All right. So, let's get started. You had this incredibly promising, successful career in neuroscience and bioengineering. I think you studied theoretical physics originally, and then you moved into AI.

Samuel Rodriques

Yeah, right. What went wrong?

Lukas Biewald

What were you thinking? Why did you do it, Sam?

Samuel Rodriques

Why would I give up such a promising career? Yeah, great question. I started doing theoretical physics and did quantum information theory. The problem with physics is there are actually no unsolved problems left in physics. That's a slight exaggeration. We need to build quantum computers, and there are some interesting, important materials science problems, and we still don't know how the universe works, right?

But actually, I think one of the key insights is that if you look around at any phenomenon that you can see in the room today—in this room that we're sitting in—I could probably explain to you how that phenomenon works down to the level of subatomic particles, right? Except for you, or the mouse. Why does the mouse do what it wants to do? Why do you wake up at the time you wake up? How do you get sick, and so on, right? When it comes to the world as we experience it, most things are pretty well understood except for biology.

And so that was originally what got me into biology. Then I did my PhD at MIT. I'm an inventor at heart. I invented a bunch of different technologies. The key thing that I learned at MIT about doing biology is that there are 3 things that you need to do science, right?

You need capital—you need money. You need logistics, which is: do you have the things that you need to run the experiments you need to run, in the place and when you need them? And you need talent. Capital scales, and logistics scales, but talent just does not scale. Fundamentally, in biology, we're limited by talent.

The thing I got to thinking about is: how do we remove talent as a bottleneck in science? If we want to go and cure all diseases, understand how the brain works, solve aging, and so on, we need to figure out how to scale talent, and AI seemed like the right way to do that. I had basically built up this career, and I kind of just threw it away and jumped off a cliff. That was pretty wild, but it's been working out well so far.

Lukas Biewald

Originally, the company that you started—which you should explain what it does—I think it started as a new kind of nonprofit, right?

Samuel Rodriques

Yeah. So, this was 2022. I had figured out that the most important thing that seemed like it was going to happen in science in the next 10 years was figuring out how to build an AI scientist, precisely because it was going to unblock talent.

At that time, in biology, you're used to things going really, really slowly. We were at this point where GPT-3 was out, and it could kind of say things. It was very impressive. You could kind of see where things were going. But it seemed like it was going to take a really long time to get there, and I didn't know how to commercialize it. I didn't know what it was.

We were like, “Well, we should do this as a nonprofit. Just go and do the basic research,” which was great. I teamed up with Andrew White, my co-founder, who is really a pioneer in the space of AI agents for science. He was working with OpenAI on GPT-4 at the time and was a professor at the University of Rochester.

We had the same vision for building an AI scientist. I knew the biology side, and Andrew actually knew how to do it technically. It felt like the right thing to do was to be a nonprofit because we thought it was going to take 5 or 10 years, which was complete nonsense. That was totally not right. As we all know now, within 2 years of launching FutureHouse—we announced it in November 2023—we already had extremely powerful AI scientists.

The timing wasn't something that ended up being very important. I will say that the fact that we started FutureHouse as a nonprofit was not the only reason we started FutureHouse as a nonprofit. We also really wanted, and continue to want, those technologies to benefit the entire scientific community. At FutureHouse, we open-sourced a lot of stuff, which I think is also super important.

By spring 2025, one day Andrej Karpathy tweeted or something, and all of a sudden the entire world knew what AI agents were. We started getting phone calls. We literally went from having to explain to people in our decks what an agent was and how it was different from a language model to having senior executives at pharma companies call us to say, “Oh my God, how do we use your agents? How do we use your thing?”

It became evident pretty quickly that we were going to have to have a for-profit spinout in order to satisfy that demand.

Lukas Biewald

Do you feel like there's a role for nonprofits in technology, then? We now have 2 examples of nonprofits developing interesting technology and mainly turning it into for-profits to get that technology out into the world. Is there something broken about the nonprofit model?

Samuel Rodriques

No. First of all, is there a role for nonprofits in technology development? Yes, absolutely. Is something broken about nonprofits? No, definitely not. This just feels like for-profits doing what for-profits do well, which is scaling.

Let's back up. When I was in my PhD, I was working on a connectomics project. Connectomics is figuring out how to map all the connections between neurons in the brain. A core piece of understanding how the brain works is figuring out how to get the wiring diagram, and to this day, we don't have the wiring diagram for the human brain. We don't have the wiring diagram for the mouse brain. We kind of have the wiring diagram for the fly, and that's the best that we have. It's really, really difficult to understand how the brain works without the wiring diagram.

We wanted to map the wiring diagram in the brain. It's really, really challenging to do that, basically because it involves tracing these tiny, tiny fibers—tiny meaning 1/100th the width of a human hair—through this gigantic tangle of the brain. You have to make no errors, because if you make errors, then you're connecting neurons that aren't actually connected to each other. It's like trying to imagine mapping all the roots in a field of grass, but way, way harder.

That was the problem I was interested in, and I tried to do it in an academic lab. I was like, “Wow, there's no way I'm going to be able to gather the engineering resources, the capital, and the talent that I'm going to need in order to do this.”

And you can't do that for profit, right? You actually can't go into a for-profit setting and get money to map the brain, because what are you going to do with a map of the brain? You're going to develop drugs, and that will take 20 years or something. That's not an attractive proposition for investors.

4. Is this time different for AI in drug discovery

So I was like, "Okay, I can't do this in academia. I can't do it for profit." Basically, there's no way for me to do this right now, right? So I proposed these things called focused research organizations, which is the idea that for some problems, like mapping the brain, that are too big for academia but can't be done for profit, we need a third way. We need some alternative kind of structure.

The idea of a focused research organization is that it's a nonprofit that operates a lot like a company, not like the slow-moving foundation that you think of when you think of normal nonprofits. It's a fast-moving, hard-charging research organization that has a specific goal. It pursues that goal, and then when it's done with that goal, it's done. You can actually spin down the nonprofit.

Since then, we've gotten philanthropists to fund a bunch of these focused research organizations, and they're doing projects that you just couldn't imagine happening in a for-profit setting. Sometimes, if they really work out—if they hit and go really well—then what is the right step afterward? It might just be spinning out a for-profit. That has always been the idea.

With FutureHouse, we started with the goal of building an AI scientist. We didn't know how to commercialize it, so a nonprofit made sense. Plus, we wanted to be able to share the basic research with the world. Then it went really well, and we made a bunch of progress. What is the next most sensible step? Spin out a for-profit and allow it to scale.

This was actually not the plan. At the beginning, we thought maybe we'd spin out a for-profit in 5 or 10 years. We certainly didn't think it would take us 2 or 3 years. But I think this is just a consequence of success, as it was also in OpenAI's case.

5. Where AI is strong and weak in science

Lukas Biewald

What did you see that made you feel like the technology was even more promising than you thought?

Samuel Rodriques

When we got started in 2022, remember, we had GPT-3 and InstructGPT. The language models then knew how to respond to you, right? But we were very much in a world where these models just couldn't do very much. We didn't have any notion of reasoning models or anything like that. They just seemed very rudimentary.

It was 18 or 24 months before they started to make discoveries. The first discovery they made was described in a paper about a system called Robin. It was the first multi-agent system that we showed was capable of doing the full loop of scientific discovery: hypothesis generation, experiment planning, running the experiments, analyzing the data, and coming up with new experiments.

It came up with a new hypothesis about a way to treat a form of blindness called age-related macular degeneration, specifically dry age-related macular degeneration, which affects 5 or 10% of people over the age of 50. This was in May 2025. Our agent came up with a new way of treating it and proposed a new treatment that we were able to validate in some wet-lab experiments. Subsequently, we've been able to validate it in animals. It was just published in Nature 2 days ago, and that was the thing where we looked at it and thought, "Oh, the future is here."

Lukas Biewald

That makes sense. Yeah. So where's it gone since then?

Samuel Rodriques

I mean, that was discovery number 1, back in May. We have since released an updated version of our agent called Cosmos. This is a much more powerful version of Robin.

Robin was kind of on rails. It could do this, but we manually orchestrated a bunch of different agents to get it to do 1 specific thing, which was identify new ways of treating diseases, particularly through drug repurposing. With Cosmos, we built in an orchestrator so that it could orchestrate itself.

We also built in this notion of world models, which is basically a very sophisticated context-management tool. It allows Kosmos to build up an integrated notion of the knowledge in a field over the course of dozens or hundreds of subagent runs. That is core to the way it makes discoveries.

Since we launched Cosmos, people have probably used it to come up with 20 or 30,000 novel scientific findings, which is wild.

Lukas Biewald

Wow. Yeah. I think anyone who's used AI for any technical application, and maybe even for nontechnical applications, notices that the intelligence is really spiky in surprising ways. There are some things it does so much better than a human, and some things it does shockingly worse.

Samuel Rodriques

Yeah.

Lukas Biewald

It wrote me a press release the other day.

Samuel Rodriques

Oh, my God. I asked it to go into our Notion and our Slack, look up all the details for a deal that we're going to announce shortly, and draft a press release. It was terrible. I was like, "I cannot—I don't know." Yes, extremely.

Lukas Biewald

I appreciate you saying that because I've had a few guests on this show lately who kind of refuse to acknowledge any weakness in their algorithms, and it gets really boring and weird because these models are going to struggle with some things. In practical fields, where do you feel like it's really strong, and where do your customers get surprised that it can't do something?

Samuel Rodriques

Okay, so it's strong in 2 areas. It's strong on things that are verifiable, and it's strong on things where throughput matters a lot.

Lukas Biewald

Okay, so it's strong in 2 areas: things that are verifiable and things where throughput matters a lot.

Samuel Rodriques

Verifiable means that you can tell whether or not an answer is correct. This has always historically been where AI has been strongest, because it allows you to get feedback very quickly, and then you can use reinforcement learning. AI is strongest where there's a lot of data and where the problems are verifiable. Coding is verifiable. Math is verifiable. Science is very much nonverifiable. It's verifiable in principle because you can go and run experiments.

Lukas Biewald

Right, but the loop is expensive.

Samuel Rodriques

The loop is expensive. It takes forever, right? People say, "Why don't you just do closed-loop reinforcement learning to teach it how to find new drugs?" I'm like, "Because that loop means I need to prove that this drug is safe, dose some humans, and then wait for 6 months." The loop is going to be 3 years long.

The other thing I hear a lot is, "Why don't you build a gigantic science warehouse where you just have pipetting robots doing experiments over and over again?" That's a better idea, and there are a lot of things for which that is probably great and very useful. But in general, that requires the experiments you're doing in a lab to be reflective of what you want.

In our case, we want medicines for humans. I can't test whether a drug works in a human using a pipetting robot in a lab. I have to test it in a human. Your model is only as good as what you train it on.

Basically, AI is good at 2 things. It's good at tasks that are verifiable, and it's good at tasks that require high throughput. On the high-throughput side, I think this is where it has really shined so far in science, because it's able to consider so much more evidence than any human is able to consider and test so many more hypotheses than any human is able to test.

You have to control for p-hacking and multiple-hypothesis testing and so on, but in biology, we have a problem with throughput. This is sort of about statistical synthesis of data. I'll give you an example: one of our FutureHouse postdoctoral fellows is interested in figuring out how to cure autoimmune diseases.

Lukas Biewald

Okay, in order to cure all of them?

Samuel Rodriques

Ideally, yeah. We'll start with 1, but actually we don't need to be picky. In order to treat autoimmune diseases, what you need are ways to manipulate the immune system.

If we go to nature for inspiration, where has nature figured out how to manipulate the human immune system? Parasites. Before modern hygiene, humans just lived with parasites. The way that worked was that the parasites had figured out how to manipulate the immune system to turn down immune reactions.

Wow. If we could figure out how they do it, maybe we could do it on ourselves in order to cure autoimmune diseases, right? This is the idea.

But there are many, many parasites, and each parasite has a genome with thousands or tens of thousands of proteins in it. What is the mechanism? How are you going to figure out what mechanisms the parasites use to regulate the immune system?

We're actually now using our agents on every single protein in any parasite genome. We're running our agents to look at that protein, look at its structure, look at its sequence, look at the context in which it appears in the genome, look at the biology of that organism, and figure out: Could this be a candidate for how this parasite moderates the immune system?

From that, we've come up with a short list of proteins that we need to test, and we're now testing them in the lab. That is something that previously there was no way to do. It would have been completely impossible.

Lukas Biewald

Right. But presumably, the hard part of that is building tools to—when you say “look at a genome,” it's obviously not feeding the genome into the context window. It's finding tools to—

Samuel Rodriques

Correct.

Lukas Biewald

—to look at what's going on.

Samuel Rodriques

But the biology—but I call this high-throughput reasoning. Yes, you can just run some homology. You can come up with an algorithm to look at the sequence, or you can use a model of protein structure and just look at the structure of all the proteins. That doesn't tell you what role that protein might play in the biology of the organism, which is fundamentally what we're interested in.

Another good example is bacteria. There are many proteins that have no known function, particularly in bacteria. If we knew what their functions were, we might be able to figure out new ways to create bioengineering tools, like finding new enzymes that do functions we can't do today.

One of the ways that you can do this is that bacteria organize their genomes into units called operons that are all functionally linked. If you have one protein with a known function inside an operon, and you also have a protein with an unknown function in that operon, you can reason about the function of the latter protein based on the former protein, which has the known function.

That's something where you really just need intelligence in order to think about the biochemical role of this operon. What role does it play in the bacterial life cycle? Based on that, can we come up with a hypothesis for what this protein of unknown function might do?

Lukas Biewald

Right. But when you test that hypothesis, presumably you have to do something physical, right?

Samuel Rodriques

Absolutely. In the case of the autoimmune disease parasite proteins, we've now gotten a bunch of them synthesized on a DNA chip, and we're screening them to see what effect they have on T cells.

You absolutely can't just reason your way to solving science. You have to do experiments. But coming up with hypotheses is one of the steps that is limiting, and that is something where the models are able to help us.

Lukas Biewald

Got it. So presumably, your customers are using your models mostly for drug discovery. Is that right?

Samuel Rodriques

Yeah. Customers use this in 2 ways. I just talked about hypothesis generation. The other place where you can imagine these models having a major impact is in the operational work of science.

How do we actually get all the materials together that we need to run an experiment? Once you figure out what experiment you want to do, you need to order and organize the materials and figure out what the protocol is. When you get into testing on humans, you need to coordinate all the resources you need in order to run the clinical trials. You need to figure out what the clinical trial sites are and prepare your regulatory documents.

All of that is part of the process of doing science, and we're used in both areas. We're used for early-stage hypothesis generation and for the actual operational tasks of development: How do we get this drug through the pipeline and to patients as quickly as possible?

Lukas Biewald

That's really interesting, because Weights & Biases customers do both also. At first, we saw mostly drug discovery applications, and then, post-LLMs, we started to see all these operational applications. I started to think that maybe the operational applications are more important. Certainly, I think more dollars go into that, and it's more of the bottleneck.

Do you have a passion for the drug discovery side and want to focus on that, or do you think the bigger business here is on the operational side? There's this very funny situation in science where progress comes from discovery, but the commercial value is in development.

Samuel Rodriques

The reason is that discoveries pan out so infrequently, and you can't tell whether or not they're valuable until 10 years or something after you came up with your hypothesis. What matters commercially—and frankly, what matters practically—is being able to run experiments faster.

You can come up with as many hypotheses as you want, but the experiments that matter are human clinical trials, right? That's what tells us whether or not the hypotheses actually work in practice.

If you want to accelerate the process of coming up with medicines, you need to make those experiments faster. That's where the commercial value is, and that's where a lot of the scientific value is. This is not to say that early-stage discovery is not important. It's also critically important.

6. Changing drug discovery

Lukas Biewald

We've had a whole slew of guests on this podcast doing different parts of the drug discovery pipeline. We've had CEOs and researchers. We've even had notorious pharma bro Martin Shkreli giving his take. He was actually down on the whole drug discovery-with-AI thing.

I'm curious how you think about the entire drug discovery market and how it's changing. What's working, what's not, what's changing, and what's static?

Samuel Rodriques

Yeah, great question. I think, obviously, there are 2 big areas where AI is having a major impact. The first one is what I talked about, which is all of the reasoning. That includes both hypothesis generation, which we do, and the operational aspect of development: How do you get these molecules through to patients faster? We also do that.

The second major important part is coming up with the molecule. If I have this hypothesis that agonizing the GLP-1 receptor is going to cause people to lose weight, then in order to test that, I need to have a molecule that, inside a human, is going to agonize that receptor, bind to the receptor, activate it, and so on.

That is way harder than it sounds. If we imagine that you want to do this orally, with a pill that you just swallow, odds are it's not going to go into your bloodstream. If it does go into your bloodstream, odds are it's just going to get filtered out immediately by the liver. Even if it doesn't get filtered out by the liver, it's probably not going to stay in your blood long enough to have an effect.

Even if it activates the receptor, maybe it activates the receptor in the wrong way. Even if it activates the receptor in the right way, maybe it activates 100 other receptors that you don't want to activate, leading to bad side effects. Then there's the possibility that your original hypothesis might not work.

Coming up with a molecule that satisfies all of those criteria—that it has the right bioavailability, meaning that it gets into the body; that it has the right distribution, metabolism, and pharmacokinetic characteristics; and that it has the right toxicity characteristics—is another extremely critical portion of the process of discovering and developing drugs.

There have been really exciting and revolutionary companies like Chai Discovery, Isomorphic Labs, Boltz, Lambda Labs, and Profluent Bio working on the problem of coming up with the actual molecule.

The upshot is that I think the biotech and pharma companies of the future will be much leaner.

Lukas Biewald

Right.

Samuel Rodriques

You'll be able to pursue many more drug programs in parallel than you can today with the same number of people. I think it's time for us to start thinking about what that means. Given that we'll be able to remove the talent bottleneck, we're going to be able to pursue cures for so many more diseases.

What would the Human Genome Project look like for medicine?

Lukas Biewald

But I feel like earlier you said the bottleneck was human clinical trials.

Samuel Rodriques

Absolutely.

Lukas Biewald

And what you just said makes it sound more like the bottleneck is coming up with the candidates for the clinical trials. Is that really the bottleneck? Which is it?

Samuel Rodriques

No, no. First of all, let me just say that there are many bottlenecks that all need to be overcome. But I think that fundamentally, at the end of the day, there’s an operational bottleneck, which is that you need humans in order to run the clinical trials. If we had twice as much money and twice as many operators in drug discovery, then in most areas—this is not always true, because sometimes you’re limited by patients and so on—you would just be able to pursue twice as many ideas.

Lukas Biewald

Right? So we’re actually limited. Like I said before, capital scales, right? If we had more good ideas that were investable, then there would be more money available for them. Do you think that the new upstarts that you mentioned, like Isomorphic Labs, Chai Discovery, and others, have a structural advantage over the incumbents in this process?

Samuel Rodriques

I mean, yes, in that mostly what those companies are doing is coming up with molecules that they then partner with pharma companies to develop. So they are mostly—not today, at least to my knowledge—developing their own drugs, and are instead trying to get this technology everywhere.

I think that the pharma companies don’t have the expertise in-house to do it themselves, and so that’s why they’re going out and partnering.

Lukas Biewald

I mean, they do have teams that supposedly work on drug discovery, right?

Samuel Rodriques

Oh, they do. I mean, the pharma companies have strong AI teams, right? Actually, we have been very impressed with the quality of the AI people we find inside many pharma companies.

But when it comes to training the best model in the world for generating a new antibody, that’s just not the pharma company’s game, in the way that it is Chai’s game, Latent Labs’ game, and so on.

7. Peptides and clinical trial processes

Lukas Biewald

Got it. Okay, so Sam, tell me about your peptide stack. [laughter]

Samuel Rodriques

I had to think about what you meant for a minute, because we’re actually working with one of our partners on training a model that is better at reasoning about peptides. So I was like, “Wait, how does he know about that?”

My peptide stack? Okay. I’m going to admit I’m not a biohacker.

Lukas Biewald

Wow. No peptides?

Samuel Rodriques

No, no peptides.

Lukas Biewald

Wow.

Samuel Rodriques

Yeah, I know. The reason is that I know too much about biology. And I think that you’ll find this—I won’t put this person on the record.

I was hanging out the other day with the head of research and development at a very large, top-20 pharma company. He was talking about the peptide fad, and he was just like, “Yeah, these people don’t understand what can go wrong.”

Lukas Biewald

Oh, I see. I think that’s true. When you have studied biology and drug development, you get an appreciation for everything that can possibly go wrong, including things that are very, very hard to identify if you’re not looking carefully.

Lukas Biewald

Interesting. You also gain a deep appreciation for things like placebo—the strength of placebos.

Samuel Rodriques

Those 2 things together just make you very intrinsically skeptical.

Lukas Biewald

So, are you nervous even about something like Ozempic that tons of people take?

Lukas Biewald

I mean, Ozempic has been used to treat diabetes for a long time, right? Or, at least, GLP-1 agonists have been. So I’m not that concerned about it. The GLP-1 agonists have now been through many, many robust, well-controlled trials, and so I don’t worry about that.

Lukas Biewald

Well, these are separate concerns, right? Does it not do anything, or does it hurt me?

Samuel Rodriques

Yeah, those are different concerns, right? But I’m also just a scientist. I think the issue is that when you think about trying peptides on yourself that have never been studied in a really robust trial setting, you have no way to know whether they’re really working.

People can feel better. That’s great. Placebos also make you feel better, so that doesn’t tell you anything about what’s making you feel better, right?

Samuel Rodriques

Even if you say, “Oh, look, everyone who takes it feels better,” right? Yeah, that’s the point. Yeah, I know. That’s right.

If this group of people were systematically—even if a very large group of people taking some of these peptides systematically had, say, a 10% incidence of cardiac arrest or something—we wouldn’t know. No one is looking. No one is tracking those adverse events when people just go and dose themselves with peptides, right?

I don’t want to sound too much like a—my recommendation would always be: probably don’t. I feel like I have a professional, ethical, and moral obligation to say to people, “Generally, doing this is ill-advised.”

Lukas Biewald

Right.

Samuel Rodriques

But this is totally off the record, so you can say anything you want on this podcast.

Lukas Biewald

Yeah, exactly—completely off the record. [laughter] I do admire people who go out and want to just try things, because I’m a fan of just trying things, right?

We just need to make sure that you’re clear-eyed about what the risks are, because there are no guarantees with a lot of these things, and they can do damage. Would you change anything about the clinical trial process that we have in the United States?

Samuel Rodriques

So, the flip side of what I just said about peptides is that part of the reason why people feel compelled to go and test things on themselves is that the process for testing them in humans in a robust, rigorous, and well-controlled manner is so onerous and takes so long.

If it were possible to run really high-quality, well-controlled studies in a way that is well regulated and so on, quickly, then maybe people would be doing that instead, which would probably be better.

There are a huge number of things that we should be doing. The first one, which is just the most obvious, boneheaded thing—we absolutely should be doing it; it’s crazy that we’re not—is that in Australia and China, for early-stage studies, the process of getting a clinical trial approved is decentralized relative to the way it is in the US.

In the US, you need to get central approval from the FDA even just to do a small-scale initial trial. In China and Australia, you have individual centers that run trials that are capable of approving trials. That means those centers can compete over how easy they can make it to do trials while staying within the regulatory guidelines, right? That is a drive for efficiency, and that is great.

As a result, US biotechs are going to Australia and China to do their clinical trials. This is obvious, right? We should be fixing it.

Actually, if there’s any administration that is going to fix it, you would think that the bull-in-a-china-shop kind of approach that this administration takes would be a great candidate to do it. They need to be doing that.

Then I think there are other things that we should definitely be looking at. One of the more obvious ones is loosening the requirements for efficacy.

The FDA requires you to prove 2 things in order to get your drug approved: They require you to prove that your drug is safe, and they require you to prove that your drug is effective for treating a specific condition.

Usually, drugs fail on efficacy. The thing about this that’s a little bit perverse is that often a drug will work in a specific subpopulation—it will work in 1 population of patients, but it does not work in the entire trial population. Therefore, the trial fails, and therefore it can’t get approved, even though it works on some subset of patients.

That feels like a failure of the system. The alternative is to only require that drugs be proven to be safe, and then determine that they are effective in the course of using them on patients in the field.

Right now, you can imagine that at some point this would not be ethical, and this would not be something that you would want to sign up for in all cases. If there is a standard of care that is known to be effective and then there's a drug where we know it's safe but it's not necessarily effective—we don't know if it's effective—you might choose the effective one. On the flip side, if the effective one doesn't work for you and you have no other options, you might choose to go with this other one. I think that would probably reduce the barrier to doing clinical trials in many diseases—not always, but—

Lukas Biewald

It's so hard to know what's effective, right? I feel like I have lots of smart friends who take lots of medicine or things like that that they think are effective. In my mind, I'm thinking, “Probably not,” because the science doesn't show that it's effective. I'm actually not even sure who's more likely to be right.

Samuel Rodriques

But what you can do, absolutely, is use randomized controlled trials. In the standard process, you have patients, and half of them get the actual medicine; half of them either don't get the medicine or, more often, get the standard of care, because you would usually consider it unethical to deny a patient a medication. Again, it depends on the condition and the disease.

The other way that you could do it, if we were to relax the efficacy requirement so that you only have to show a drug is safe, is to look at the extent to which a single patient improves—you could look at within-patient improvement. If you take drug A for depression—and, obviously, antidepressants are famously very variable in whether they work in any given individual—maybe you take antidepressant number 1 and it doesn't work for you and your symptoms persist, and then you take antidepressant number 2 and suddenly you get much better. That's pretty good as far as evidence goes, because the only variable there that has been changed is which medicine you're taking.

That avoids the placebo problem to some extent, right? The issue with the placebo is you go from not taking medicine to taking a medicine, and simply taking a medicine is often effective, regardless of whether the medicine works. The other way that you can think about doing it is patients who are on drug A, which is known to be effective, versus patients who are on drug B, which is not known to be effective. If drug B is more effective—or if the patients on drug B have a better outcome than the patients on drug A—that's very strong evidence in favor, right?

Lukas Biewald

Although I imagine this all sounds good in theory, I would think it would be very vulnerable to p-hacking if you're just doing these natural experiments in the wild, tracking everybody, and looking for effects like this.

Samuel Rodriques

P-hacking is not especially something you can control after the fact. Picking your subpopulations post hoc can be a problem. Although, again, this really just becomes a question of power. P-hacking is always a question of power.

Lukas Biewald

Except that you don't know all the possibilities that were considered, right?

Samuel Rodriques

Well, you need to control for that, or preregister what you're going to look at.

Let's imagine—let's just take the example of a vaccine, because vaccines' real-world evidence is very unambiguous: You either get the disease or you don't get the disease. If I go out and have a vaccine—I don't know if it's effective—and I just give it to a bunch of people, and then afterward, let's say at a population level it's not effective, which is to say people without the vaccine get the disease at a rate that's indistinguishable from people who get the vaccine, then I'm going to go in and look at, “What about just men? Women? Men above the age of 18 but below the age of 35?” I'll go and look at all the different subpopulations, and inevitably I'll find one where no one in that population got the disease. Therefore, I'm going to say it's 100% effective, and you're going to say, “No, it's just p-hacking,” and you'd be correct.

If I have 10 million patients and across 10 million patients I still find that somehow men between 18 and 35 who get my vaccine never get the disease, whereas all other populations get the disease at the same rate, no, that's not p-hacking anymore. That's obviously not p-hacking. Something is going on; we don't know what it is, but that's obviously not p-hacking.

That's what I mean: It's a question of power. If you get enough real-world data, you can robustly go back and identify these subpopulations in which a drug is effective without worrying about p-hacking. Now, you may still want to do more experiments because you may have no idea why, and you may want to do a confirmatory experiment. But you can imagine that it's better to be able to go back in and at least find those hypotheses than it is to just have the trial fail and the drug—

Lukas Biewald

I mean, how often do we have real-world data at that scale where there's not sort of weird selection bias?

Samuel Rodriques

Yeah. Vaccines are often given at population scale. But you could imagine that antidepressants are a great example. You could imagine getting enough real-world data on antidepressants in order to be able to do this.

The more common cancers—I mean, cancer is a specific thing because the treatment is obviously life or death. You need to be more careful. You never want to deny a cancer patient the standard of care, right? But the common cancers, breast cancer, colon cancer, and so on, will have hundreds of thousands or millions of patients.

So I think that, obviously, for a rare disease—

Lukas Biewald

Mm-hmm.

Samuel Rodriques

—you won't get a huge amount of real-world data, but also for a rare disease, you don't have that many subpopulations usually.

Lukas Biewald

Mm-hmm.

Samuel Rodriques

Right.

Lukas Biewald

Do you think that a model like yours could go through the existing real-world data and find new patterns you haven't seen before?

Samuel Rodriques

Absolutely. We have several clients that we're working with aimed at that kind of work. The challenge there is always the quality of the data, which is to say that real-world data is gathered in the real world, and real-world people—as probably many people watching this understand—do not always go to their follow-up appointments, do not always take the medication on the schedule they're supposed to take it on, and doctors do not always take high-quality notes.

Patients move between healthcare systems and then you lose the records; they end up disconnected. For all these reasons, existing real-world evidence is never as good as you want it to be.

There are a number of great companies trying to fix this. One of them is a company called Empower Medicine, which is focused on gathering extremely high-quality electronic health record data from patients so that you can design synthetic clinical trials. There are other companies like Tempus, Komodo, and so on that do this as well, but—

Lukas Biewald

I guess in the single-treatment case, it makes sense that there would be a lot of public research and someone would want to decide once and for all: Is this effective or not?

Samuel Rodriques

Yeah.

Lukas Biewald

I've found in my life that when I have medical issues, they feel high stakes, and when I go look at the research and the data, it never seems clear. It often seems like there's conflicting research, and every single data set seems biased in different ways. I've found myself more and more using LLMs to try to synthesize the research into making a sensible decision.

For example, fertility was a big issue for me and my family, and it's expensive, and there are upsides and downsides. I really wanted a model like yours, and I think I asked you to use your model for a recent thing that I was looking at. Do you expect that people might use your model in that way, to look at their individual, specific situation and try to dispassionately synthesize the research into a decision?

Samuel Rodriques

My mother does.

Lukas Biewald

No way. Tell me about that. That's cool.

Samuel Rodriques

Well, I don't know that my mother wants me to air her medical history on a podcast. But this is totally off the record, man.

Lukas Biewald

I know. Good. I forgot.

Samuel Rodriques

Exactly. No, look, I have a bunch of friends who use it. I have one friend who had a breast cancer scare, for example, who wanted to figure out, if she delayed treatment, what that would do to her recovery or survival odds. She had a reason why she wanted to delay treatment for a month or two. Luckily, she turned out to be fine, which is good.

But one of the things that our platform is very good at is going out and dispassionately surveying the evidence and giving you an answer that is directly grounded in hard evidence from papers and from the literature, as opposed to vibes.

Lukas Biewald

All right. Actually, you're a recent father. Have you used these models, or your own model, for your child yet?

Samuel Rodriques

Yeah. We have been lucky enough so far that we've not needed to use them for medical reasons. But, 100%, I ask it things like, “When will my baby start to walk and crawl?” and so on. Yeah, yeah, yeah.

Lukas Biewald

I'm sure you will. I actually had a recent, interesting experience with my wife where we both put in the same symptoms and incident involving our daughter. She hit her head, and we both described it identically. The model told me it was fine and told my wife to take the child to the doctor, which I think might have been a different kind of incentive in the model. I do feel like they kind of want to tell you what you want to hear. Maybe some RLHF gives it that. So I wonder if it sort of implicitly knew—

Samuel Rodriques

Or, I mean, it has the history. It probably—you know, it definitely does.

Yeah. I mean, like I said before, what we focus on is making sure that our answers are grounded in scientific fact, right? But it can be frustrating. You can ask ChatGPT or Claude whether you should take your child to the doctor, and they will say, “No, the child's fine.” If you ask Cosmos, our agent, whether you should take your child to the doctor, Kosmos will probably come back and say, “Children who hit their heads have an X-percentage chance of developing this serious problem and a Y-percentage chance of doing this.” Then it will give you something like, “On average, there's a 97% chance that if you don't take your child to the doctor, they're going to be fine.” And that can be more frustrating sometimes.

Lukas Biewald

Honestly, that sounds fantastic. Yeah.

Samuel Rodriques

But it's a question of how much information you want. It's very interesting to think about. We're entering this era where information and intelligence are just so much more abundant than they were before, but it's not always good that we're going to want that.

Lukas Biewald

Right. More information is not always better for making decisions, right?

Samuel Rodriques

Totally true.

Lukas Biewald

Although I don't like to admit that.

Samuel Rodriques

I know. But often, you need the right information, but not the most information.

Lukas Biewald

So OpenAI, Anthropic, and DeepMind are obviously working on similar things to you, at least in the sense of deep research, making the models more advanced, and making agents that think more. Do you feel like you need to keep some kind of structural advantage over what they're doing?

Samuel Rodriques

I mean, yes. If you're building a company in general, you want to try to have a structural advantage over what other people are doing.

Lukas Biewald

Okay. So what is your structural advantage?

Samuel Rodriques

Right, great question. There are a couple of ways to think about this. The most—

Lukas Biewald

By the way, were you yawning through my question?

Samuel Rodriques

Yeah.

Lukas Biewald

Is that a bad question?

Samuel Rodriques

Yeah. Sorry, man. My whole life, I get asked, “OpenAI and Anthropic are doing this,” and I'm just like, “Oh—”

Lukas Biewald

I'm giving you a softball, man. Let's go.

Samuel Rodriques

No, so—

Lukas Biewald

Um, so—

Samuel Rodriques

Man, now I'm worried. Maybe we should move on.

Lukas Biewald

No, no, no, no, no, no. It's a great question. I'm trying to tease you. It's a great question.

Samuel Rodriques

Fundamentally, at the end of the day, OpenAI, Anthropic, and DeepMind are focused on building AGI, or artificial superintelligence, or whatever, right? I think the key thing to understand is that having a specialized model for a particular task is always going to be better than having a generalist model for that task, even when you have superintelligence or whatever. The reason is that the generalist model has to do everything in the weights, and the specialist model only has to do some things.

Lukas Biewald

Wait, so do you have a specialist model?

Samuel Rodriques

Absolutely. We train reasoning models on specific scientific tasks.

Lukas Biewald

I see.

Samuel Rodriques

What we see is that with a very small amount of data, you can get enormous gains for those specific tasks over the frontier models. That's the first thing: having specialized agents and specialized models, which also, by the way, comes with specialized user experiences, like I was mentioning. We put way more test-time compute into minimizing hallucinations than the others would.

Lukas Biewald

But presumably, you also use the bigger models, right?

Samuel Rodriques

Yeah, absolutely. We use the bigger models in some areas. We don't want to be better than Anthropic at coding, for example, right? But when it comes to very niche biological reasoning, we have our own models internally that are superior.

Lukas Biewald

Doesn't the coding thing actually show an example of a specialized model? These coding models are not really specialized in coding, right? And yet—

Samuel Rodriques

They definitely—if your point is, could a specialized coding model beat GPT-5.5 or beat—

Lukas Biewald

You think it seems unlikely?

Samuel Rodriques

Well, I think a lot of people are trying to do it, and they haven't yet. Coding is very special because I think all the labs have realized at this point that coding is the thing that they need to be really, really good at, right? Maybe it's not true in coding for that reason. But I can kind of guarantee you that if you want to train a specialized model for reasoning about synthetic chemistry and synthetic pathways, the specialized model is going to do better than the generalist model.

Right now, the place where this may fall down is when you get to intelligence saturation. The thing I tell people about saturation is that if you think about tic-tac-toe, any superintelligence will be exactly as good as humans at playing tic-tac-toe. No amount of intelligence above human intelligence will improve performance at tic-tac-toe, because humans are optimal at tic-tac-toe. You can always—

Lukas Biewald

Some humans—

Samuel Rodriques

Some humans, not all humans, but there exist humans who are optimal at tic-tac-toe, right? Similarly, you should imagine that at some point, intelligence will get strong enough that we will have saturated, and more intelligence will not improve your ability to reason about chemistry. I don't know where that point is, but it's feasible that we would get there, and then my argument might fall down, right?

Lukas Biewald

Yeah.

Samuel Rodriques

No, at that point the argument might fall down. But that's the first thing. The second thing—but I think we're very far away from that point, right? All of the pharma companies have their own internal datasets, and they need to have proprietary advantages over each other, right? They need models that are trained on their data, because if everyone has the same intelligence, then no one has any advantage in R&D.

Those are the areas where we really shine. The reason right now why all pharma companies—or it feels like all pharma companies—are fighting to do a deal with us, and we're just overwhelmed with demand, is because we have the best models for science, we do the last-mile integration to get those models deployed internally and actually accelerate the pipeline, and we train on their data. That provides them with their sustainable advantage, which is something that, at least today, OpenAI and Anthropic don't want to do because they don't want to have a different model for every customer.

Lukas Biewald

It's kind of interesting. I feel like when I talked to you 6 months or a year ago, you talked more about the agent framework and the evals that you were doing. Have things changed, or is that less exciting to talk about?

Samuel Rodriques

No. I think that we have matured, or we are in the process of maturing. The way that we think about what we do and about the value proposition is maturing, but it remains the case that our agents are, I think, the best at science and are substantially further ahead of the offerings that Anthropic, OpenAI, DeepMind, and so on have. I think we'll be able to maintain that edge for a while.

But when you think about the macro dynamics of what the market will look like in 3 years, 4 years, or 5 years, there are a couple of key things happening. The first one is that these companies need models that are trained on their data to maintain their advantage.

The second one is that they don’t want to be locked into a single model provider. They don’t want to be locked into just Anthropic or just OpenAI or whatever, because last year OpenAI was way ahead, and today Anthropic feels like it’s way ahead, and so on. Those dynamics are pushing customers to work with us.

Lukas Biewald

Do you think, when you look at the major model providers, that they’re different enough that using them for specialized use cases is a valuable thing to do, or are they essentially interchangeable?

Samuel Rodriques

This is a great question. I think that, at the moment, they’re mostly—they’re spiky, right?

Lukas Biewald

But are the spikes overlapping, or are they—

Samuel Rodriques

I think there’s a lot of non-overlapping spikiness.

Lukas Biewald

Oh, interesting.

Samuel Rodriques

So—

Lukas Biewald

Give me one example.

Samuel Rodriques

Yeah, great. For example, we have a benchmark that we’re building that we’ll release shortly, which involves reproducing analyses that were previously done in scientific papers. On our evaluations as of when we’re recording this—it may change—but at least as of the last data I saw, Anthropic is particularly good at that, significantly more so than OpenAI and significantly more so than DeepMind.

Another good example, which maybe makes the point to an even greater extent, is that there was a time early in the development of Kosmos when there was a specific task that we needed the models to do in the process of updating Cosmos’s world model. At the time, the only model we could find that was able to do that task was Gemini 2.5. We could not figure out how to get any of the other models to do that task, which was wild.

There’s a lot of that spikiness, right? They have slightly different personalities, obviously, so that’s also interesting. We’ve been surprised by the extent to which you can get much better results on tasks by combining models across different providers.

Lukas Biewald

I see. When I fire off Cosmos today, how long does it run for? How many calls is it making?

Samuel Rodriques

Yeah, great question. We released the original version of Kosmos back in November. It would run for 6 to 12 hours, write 45,000 lines of code, read 1,500 papers in a single run, and was way more powerful than any agent that anyone had seen. I think it’s still one of the most powerful, most intensive agents out there.

The issue with it was a UX issue. If you’re a scientist doing research, you don’t really want to ask a question, walk away, and come back 6 to 12 hours later. We recently announced a significant update to Cosmos that puts an interactive front end on it, so it can still go and do those extremely long-running tasks, but you can give it feedback throughout. You can steer it. It’s more interactive.

Lukas Biewald

Let’s go to rapid-fire random questions here.

Samuel Rodriques

Yeah, let’s do it.

Lukas Biewald

All right. You said—I saw that you said—scientific progress has slowed down a lot. That was different from how I think about scientific progress. What did you mean by that?

Samuel Rodriques

That depends on which time you’re talking about when I said this.

Lukas Biewald

Do you still stand by that point that I pulled out of context?

Samuel Rodriques

The—okay, it is empirically the case that progress in medicine has slowed down.

Lukas Biewald

What does “empirically” mean?

Samuel Rodriques

It means that we have, in medicine, what is called Eroom’s law.

Lukas Biewald

Which is Moore spelled backward.

Samuel Rodriques

Exactly. It’s the observation that the amount of money, in real terms, that it costs to develop a new drug has doubled over the past 40 years. I forget what the doubling time is. It might be that it has doubled every 9 years or something like that.

Semiconductor prices have had this exponential drop. Drug prices have had this exponential increase. There are various reasons people have argued for this. That trend actually broke around 2011 or 2012. People think it’s largely because of the Human Genome Project.

Lukas Biewald

Oh, actually, the Human Genome Project made it easier to find drugs.

Samuel Rodriques

But now it’s about flat, and certainly productivity is not yet increasing.

Lukas Biewald

Is that because we found all the best drugs, or—

Samuel Rodriques

That’s one of the reasons. One of the reasons may be a lack of low-hanging fruit. One of the reasons may be what’s called the “better than the Beatles” problem. If you have a drug that treats condition X—if you have a drug that treats depression, maybe a bad example; whatever, you have a drug that treats colon cancer—you need to come out with a drug that is better than that drug—

Lukas Biewald

Okay?

Samuel Rodriques

—in order to gain market share, right? In order for it to get approved.

Lukas Biewald

Even just in order for it to get approved, and then in order for it to be used—to be useful.

Samuel Rodriques

Exactly.

Lukas Biewald

Why do you call it “better than the Beatles”?

Samuel Rodriques

Just because there’s this observation that, I guess, the Beatles are still extremely popular.

Lukas Biewald

I see.

Samuel Rodriques

Right. The Beatles have not been displaced as a band. They were, like, first. Any band that comes afterward—if the Beatles had not existed, maybe there would be some other band. There are still good bands, but they don’t get as much share as the Beatles, because the Beatles took that space in the consciousness.

Lukas Biewald

Basically, all diseases have been cured. So—

Samuel Rodriques

No, but this is the thing. They haven’t been cured. It’s just that we can’t find better treatments for them, right? No one is going to go out and argue that depression has been cured or that colon cancer is cured. But if you want your drug to get approved, you have to be better than the best in class.

Lukas Biewald

But, yeah, I think that’s a problem. Of course you need to be better than the best in class, right?

Samuel Rodriques

But that just means it’s getting harder, right? It’s getting harder to come up with new things, because you don’t build on previous innovations in medicine in the way you were able to build on previous innovations in semiconductors, right? Why not? Well, because if you have a drug that uses a particular mechanism to cure cancer or fight cancer, you need to come up with a different mechanism. There’s only so much that you can get from juicing that specific mechanism, right?

You can optimize the drug, make it marginally better, and so on, and there are a lot of gains that you get out of that. But it’s often the first-in-class drugs that lead to the dramatic improvements.

Lukas Biewald

I see.

Samuel Rodriques

Right.

Lukas Biewald

Interesting. So there’s some new mechanism that we’re using.

Samuel Rodriques

Yeah, and so I think science is moving more slowly for that reason. I also just think, as in the case of physics, that as you discover more things, it becomes harder to discover more things.

Lukas Biewald

Right?

Samuel Rodriques

Interesting.

Lukas Biewald

All right. Is a PhD or formal science degree still worth pursuing?

Samuel Rodriques

Right. Wow. Yeah, good question. I don’t know where to start my answer. Sorry, you can yawn. Your turn to yawn.

Lukas Biewald

Oh my God. The—

Samuel Rodriques

Equivocating. I hate it.

Lukas Biewald

The—

Samuel Rodriques

I think the answer is yes, because I don’t think human researchers are going anywhere anytime soon. Fundamentally, the point of a PhD is to learn how to do research, right? If you never learn how to do research, you’re definitely not going to be effective at doing research using the—

Lukas Biewald

But aren’t you automating research?

Samuel Rodriques

We are, or certainly accelerating it. The more interesting question is whether we’re going to need human scientists.

Lukas Biewald

Yeah, a related question for sure. What do you think?

Samuel Rodriques

I think the answer is yes for a substantial amount of time. The reason is that science is nonverifiable. Fundamentally, at the end of the day, we need human scientists for their taste. I’m not sure that it’s going to be—maybe in 20 years. On the 5- to 10-year time scale, I’m not sure that we’re going to get to a point—it’s not clear to me yet whether we’ll get to a point where the models will just have obviously better taste than humans.

8. Will AI ever replace Nobel laureates

Lukas Biewald

Well, it’s interesting, because here it says you said language models will eventually be better than humans at coming up with ideas.

Samuel Rodriques

Yeah, but—

Lukas Biewald

You take it that far out, though. I mean, eventually—and “better” exists along many different axes, right? Better can mean more likely to work. I think that’s definitely the case, right? Language models are going to do way better than humans at coming up with ideas that are better, that are more likely to work, right? But when it comes to which of these problems is the most interesting to pursue, it’s just harder because it’s not verifiable.

Samuel Rodriques

And so this is just uncertainty, right? Like I said, I’m a scientist. I admit when I don’t know things. I don’t know whether we’re going to get to a point within 5 years where we’re just like, “Oh, you know, there’s no point in asking Frances Arnold, Nobel laureate, what she thinks about evolving proteins because we could just ask the model.” I think it’s a pretty tall bar. When I look today at asking Claude 3.7 what it thinks we should be doing with directed evolution to improve chemistry or to open up new avenues in chemistry, versus asking Frances, there’s no question that Frances is going to have better ideas.

Lukas Biewald

Interesting today, right? I mean, in 2 years, we’re going to be sitting here again and you’re going to be like, “Well, you said that, and now you know…”

Samuel Rodriques

In this off-the-record podcast.

Lukas Biewald

All right. I wonder if it’s going to work going forward to tell guests it’s off the record. Chat, the house rules, people. Chat, the house rules.

Samuel Rodriques

I was literally just at an event where I was on a panel, and they were like, “It’s off the record.” I was looking around thinking, “What about all the cameras that are pointed at me?” They were just like, “Oh my God, nothing’s off the record. Are you kidding?”

9. Raising $70 million from pharma investors

Lukas Biewald

You’ve raised quite a lot of money. I think I have $70 million in my notes. Is that even accurate?

Samuel Rodriques

Yeah, that’s right. $70 million.

Lukas Biewald

And that was just the for-profit; the nonprofit raised more. I think, unlike most of the new labs and the kinds of companies that I come across, most of your investors are actually folks I don’t know. They’re coming more from the pharma world, I think. Why do you think that is?

Samuel Rodriques

Well, for us, success is getting inside the pharma companies. The pharma investors are the ones who have that capability. So we get very high value out of all of our investors.

Yasmin Razavi is an investor at Spark Capital who co-led our round, and she is absolutely incredible. She’s on the board of Anthropic. She led their first VC round and has extraordinary perspective and insight. She has really helped us with talent and so on.

When I think about the concrete business traction, the investors who have contributed the most value are Jeff Huber, who was the CEO of Grail and somehow seems to know literally every single person in pharma, and an unnamed institutional biotech investor—a very large, unnamed institutional biotech investor—who would be displeased if I said in this off-the-record forum who they are. But I think most people in biotech will know what that means because they’re very well known, and they have been extremely valuable in terms of setting us up with high-level connections into companies and getting us deployed.

Lukas Biewald

It’s interesting. So the people who actually know your field better are more bullish on you than the maniacs doing AI investment in Silicon Valley.

Samuel Rodriques

Yeah. I think the maniacs are very excited about many things, right? They’re very excited about the concept of, “We’re going to automate science.” I mean, everyone is very excited about, “Let’s go automate science. Let’s go automate drug discovery,” right? But the insiders know the problems, so they really viscerally know where the value is. Let me put it this way: I don’t get any of the insiders asking me how we’re differentiated versus Anthropic, OpenAI, and Google.

Lukas Biewald

I see.

Samuel Rodriques

Right. I get asked that by every single tech investor—

Lukas Biewald

Right.

Samuel Rodriques

—who will ask me why Anthropic, OpenAI, and Google won’t do what we’re doing. But anyone who has operated inside a pharma company, who has large positions in pharma companies and biotechs—we never get that question from them.

Lukas Biewald

Okay, there’s a long history of AI overpromising in drug discovery.

Samuel Rodriques

Oh, yeah.

Lukas Biewald

Is this time different?

Samuel Rodriques

Yes. I promise you, this time it is different.

Yeah, man, back to 2012, people have been saying that pharma AI is going to revolutionize pharma, and it’s been one flop after another, with the notable exception of AlphaFold. AlphaFold really changed the way that a lot of things are done, but it changed the way a lot of things are done in 1 area of drug discovery and development: the actual process of coming up with the molecule because you have the structure of the proteins.

But I think there’s a ton of skepticism. That said, it is crazy to think that this time will not be different. We have so much evidence already that this time it’s really going to be different.

The thing I do want to emphasize is that the proof is in the pudding, and the pudding is approved drugs. You’re going to hear a lot of stuff in the next couple of years: “My model came up with this drug,” and “My model discovered this fundamental aspect of biology,” and so on. Anyone who has been in pharma is not going to care until they see the outcome of a pivotal Phase 3 clinical trial or an approval letter from the FDA.

Lukas Biewald

So that seems like a great place to end.

Thanks so much for listening to this episode of Gradient Descent. Please stay tuned for future episodes.

Curing Every Disease With Al by 2050 | Sam Rodriques, Edison Scientific | BidClub