[BidClub_]
Latent Space · · 77 min

🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI

Joseph Krause

YouTube
TL;DR
  • Radical AI’s core wager is that materials AI gains its edge by closing the experimental loop, because composition generation is only the beginning and “the ground truth is the material itself.” A useful material must be synthesized, characterized, processed, manufactured, and qualified against cost, supply-chain, and application constraints. Krause’s differentiation from Lila, Citrine, Periodic, and others is therefore experimental data and lab infrastructure—not merely a better model.

  • The early throughput numbers suggest a real step-change, though Radical remains far from end-to-end manufacturing. Krause gave two different windows for producing roughly 1,200 alloys: five or six months in one account and three months later. About 300 were absent from the literature and “probably 10” had especially exciting performance. Current throughput is 8-20 alloys per day at roughly $60-$300 each, with a stated target of 100 per day by June or July—versus the MACH program’s cited benchmark of 500 alloys in 12 months.

  • Commercialization risk moves downstream, where qualification and manufacturing intuition could absorb much of the discovery advantage. Aerospace qualification typically takes about 10 years, and Radical currently operates at grams to 200-500 g—not the 300-pound or 10-ton scales raised in the conversation. Krause sees a plausible 3-5-year path into defense or space applications, but not manned-flight turbines; semiconductor integration remains “still pretty long.”

  • The addressable opportunity is not just discovering stronger alloys, but designing materials concurrently with the products that need them. High-entropy alloys containing five to seven roughly equal elements could target temperatures north of 2,000°C, even 3,000°C, as well as pressure, oxidation, corrosion, or neutron bombardment. Krause’s borrowed framing is “concurrent engineering”: instead of designing a rocket, turbine, or chip around decades-old materials, engineers iterate the material and product together.

  • Radical’s self-driving lab is an orchestration problem spanning software, robotics, perception, scientific judgment, and physical tooling. An automated lab is like hands-free highway driving; a self-driving lab is a Waymo that chooses the route and runs an entire research campaign. Custom grippers must pry 3,000-4,000°C alloy “buttons” from trays, models must judge whether melting is complete, and an operating system must decide whether a failed sample should proceed or be killed.

  • Krause argues that materials science is experiment-constrained rather than compute-constrained, making laboratory throughput and experimental data the economic moat. The relevant search space may contain roughly 10^40 possible alloys, yet useful discovery signals are already appearing from hundreds of experiments because high-quality experimental data are scarce. His blunt formulation: “We think in science models aren’t the moat, experiments are”—hence Radical can open-source models while treating its experiments and data as the edge.

  • The broader strategic thesis is that self-driving labs could multiply scarce scientific labor and help the US compete with China without copying China’s system in which one entity can control public and private activity. Radical says one metallurgy PhD can oversee 10 campaigns at once, reversing the traditional ratio of roughly 10 researchers to one campaign. Krause’s proposed counterweight combines national-lab data, HPC and instrumentation with private software and autonomy—but the hosts noted that China can deploy the same productivity tools, so execution and scale-up infrastructure remain decisive.

Digest · the substance, structured for research

1. Materials AI must close the loop from hypothesis to physical truth

  • Challenged to differentiate Radical AI from Lila, Citrine, Periodic, and an increasingly crowded field, Krause returned to one conviction: “the ground truth is the material itself.” Models matter, but a proposed composition is only the beginning.

  • Radical’s intended closed loop has an AI scientist propose candidates, the lab synthesize and characterize them, and experimental results feed the next campaign. Automation is extensive, but humans remain involved, especially in synthesis and scientific annotation. The objective is not a plausible digital structure; it is a material that can eventually enter an industrial application.

  • Krause’s causal argument is that performance often emerges after composition selection. Microstructure, heat treatment, post-processing, additive manufacturing versus casting, and scale determine whether the nominally same alloy is strong, ductile, oxidation-resistant, manufacturable, or useless.

  • The industry’s cited 15-30-year timelines reflect fragmentation: academia discovers, government-backed programs conduct light testing, and large companies optimize existing systems by 5% or 10%. Data rarely travel across those handoffs, severing discovery from manufacturing.

2. Radical has automated discovery-scale work, not the material’s full lifespan

  • Radical currently covers hypothesis generation, synthesis, characterization, and early property testing. Its characterization suite includes SEM, EDS, XRD, XRF, and TGA—different instruments for identifying structures, phases, chemistry, and thermal behavior.

  • Property testing includes oxidation performance, tensile stress-strain curves, and microindentation. Vickers hardness is measured directly, while the lab’s ductility signal is only a proxy; Krause explicitly cautioned that it is “not an exact measurement of ductility.”

  • The reported output is roughly 1,200 alloys. Krause described the window as five or six months in one account and three months later in the conversation. Around 300 were novel relative to the literature, and “probably 10” produced performance exciting enough for deeper industry discussions and patent work.

  • The scale boundary is material: Radical works with grams, commonly 200 g or 500 g, not the 300-pound or 10-ton scales raised in the conversation. Wind-tunnel, torch, and other expertise-heavy aerospace tests remain with third parties, while full manufacturing data have not yet entered the loop.

3. High-entropy alloys make the case for concurrent engineering

  • Radical is not merely permuting a mature recipe book, Krause argued. Its high-entropy alloys combine five to seven elements at approximately equal atomic shares, creating candidates for extreme temperatures—often north of 2,000°C, even 3,000°C—high pressure, oxidation, corrosion, and neutron exposure.

  • The opportunity exists because aerospace and other industries still rely heavily on alloys developed in the 1950s through 1970s, sometimes augmented by later coatings. Long development cycles make incumbents rationally favor incremental optimization over unfamiliar material families.

  • Krause borrowed SpaceX materials executive Charles’s phrase “concurrent engineering”: design the material while designing the rocket booster, turbine, missile, solar cell, or other product. Performance specifications become inputs to material discovery instead of constraints inherited from whatever qualified alloy already exists.

4. Qualification, supply chains, and unit economics decide what survives

  • The hosts compared downstream materials qualification with drug development. For aerospace and defense alloys, qualification under the FAA or military specifications can require multiple ingots, standardized tests, and roughly 10 years before use in safety-critical systems.

  • The hosts asked whether qualification could receive an “Operation Warp Speed” treatment by parallelizing tests. Krause pointed to DARPA work using additive manufacturing and layer-by-layer analysis, but did not claim the sequence had been solved; the goal is to achieve the same result through a newer mechanism.

  • His safety distinction is deliberately stark: a bending iPhone can be rejected or, disastrously, recalled; a turbine material on a 787 must meet a much higher bar. “No one would want that bar to be removed”—the outdated mechanism, not the safety requirement, is the target.

  • Supply-chain shocks can rewrite the objective midstream. Krause cited hafnium rising 10-15x because China controls a majority of the supply chain; C103 contains about 10% hafnium by weight, and Radical has worked on removing that element while preserving performance. Space tolerates higher cost for performance, whereas consumer electronics and medical devices are far more price-sensitive.

5. A self-driving lab chooses the route, not merely the speed

  • Krause’s clean distinction: an automated lab executes human-defined experiments at high throughput, like hands-free driving that still requires the driver to make a turn. A self-driving lab runs research campaigns, like a Waymo whose passenger specifies only the destination.

  • Radical divides that system into difficult sample manipulation and tooling, a lab operating system, and connected automation. The software tracks samples, controls instruments, consumes sensor data, performs quality checks, and can terminate a bad experiment before wasting time on XRD, SEM, and later tests.

  • Physical manipulation is deceptively difficult. Alloy “buttons” blasted at 3,000-4,000°C stick to their trays; a person intuitively uses a chisel, while Radical needed custom robotic actuators that could remove them without damaging the sample or altering its microstructure.

  • Humans remain important teachers. Metallurgists annotate SEM images—“I see dendritic formation on this image in these locations”—so the AI scientist can absorb the judgment a PhD applies almost unconsciously when reading microstructures.

6. Radical narrowed its platform ambition to earn vertical depth

  • Krause’s change of mind is central: the founding plan was seven labs across seven material systems. Customer questions about specialized tests, manufacturing, and scaling to 300 pounds exposed the shallowness of that plan, so Radical prioritized vertical depth in alloys before polymers or ceramics—possibly without ever expanding.

  • Semiconductors remain an adjacent program because new back-end-of-line interconnect materials might reduce integration losses and energy costs. Krause said some systems recommended today could deliver roughly 2x to 5x improvements, potentially exceeding 10x later, but withheld the material details and admitted, “I don’t know to what level.”

  • Integration into an iPhone or an NVIDIA GPU remains “still pretty long,” compounded by scarce chip capacity and the need to build testing infrastructure from scratch. Krause was more confident about a 3-5-year alloy path into defense or space systems, explicitly excluding manned flight as the likely first use.

7. Active learning proceeds campaign by campaign, with selective human control

  • Radical’s AI scientist designs a campaign, selects a batch of candidates at its chosen confidence level, and sends them through synthesis, characterization, and early property testing. Machine-learning models analyze some outputs; scientists annotate others; all results return to the database for the next campaign.

  • Updates occur by campaign, not after every specimen. Krause said the lab could probably run approximately seven to 10 campaigns across different systems, with results revising hypotheses daily or every other day: “We actually want to take a few shots and get enough data back to change our hypothesis.”

  • Characterization is fully automated, as are oxidation and microindentation. Tensile testing is nearly automated, while synthesis still uses metallurgy PhDs because casting requires judgments such as whether a corner has fully melted; Krause expected the custom synthesis tool to be automated by the summer.

  • Candidate generation is already assigned to the AI scientist. Human researchers occasionally submit competing compositions—effectively red-teaming it—and may be rejected as insufficiently strong, but they also learn from unexpected elements the model introduces.

8. Throughput lets the AI explore where human intuition says not to look

  • Radical’s literature map shows published alloy families overlaid with regions its AI scientist explored. Human scientists explained their omissions candidly: they expected certain elements to evaporate, fail to cast, form poor grains, or damage mechanical properties—yet some combinations synthesized successfully.

  • The hosts’ pushback is worth preserving: perhaps the machine is simply receiving more trials and a higher “temperature,” while an unconstrained human team might explore similarly. Krause conceded that literature often pulls the system toward known successes and that throughput is “an important number.”

  • His answer is behavioral as much as algorithmic. A PhD researcher might perform around 50 experiments annually, spend roughly two weeks at a time fabricating and testing each one, and therefore treat every choice as precious; an AI scientist making eight, 20, or eventually 100 samples daily can afford speculative “shots on goal.”

  • Current experiments cost roughly $60-$300 depending on elements such as platinum, palladium, aluminum, or titanium. Throughput ranges from eight refractory alloys to 20 easier systems per day, with a rough June-July target of 100 daily regardless of system.

  • The AI scientist also operates in parallel rather than in the serial way of a human researcher: Krause said it could compare 100,000 publications with 100,000 SEM images in real time, whereas a human cannot retain and directly compare that volume.

9. The bottleneck is experiments, not compute or search-space size

  • Krause contrasted Radical with the DARPA-GE Aerospace MACH program, which he described as producing 500 alloys in about 12 months after AI and simulation screening. Radical’s target is 500 in five business days, an order-of-magnitude change even before manufacturing-scale work.

  • The search space still overwhelms brute force: Krause estimated roughly 10^40 possible alloys and said humans would take seven million years to synthesize them all. AI remains valuable for screening, but its feedback quality depends on experiments that the industry historically has not captured or shared.

  • The hosts compared tens or hundreds of alloy samples with biological assays that can scale to millions or billions, questioning whether sparse local patterns would beat expert design. Krause’s empirical rebuttal was that Radical sees meaningful results from 100, 200, or 300 experiments, with 300 new alloys among the 1,200 run to date and roughly 50-150 data points per alloy.

  • “We’re not compute constrained in the materials industry. We’re experiment constrained.” Radical’s ambition is consequently a “protein data bank for materials,” but one containing processing, microstructure, properties, cost, and application context—not merely crystallographic structures.

10. There is no single AlphaFold moment for an industrial material

  • Krause agreed that “there is no AlphaFold for materials” at the full-system level. Narrow AlphaFold-like moments are possible: segmentation models can read SEM images, identify dendrites, cracks, or defects, and relate crack propagation to mechanical behavior.

  • Krause’s broader contrast with biology is that SELFIES and SMILES can represent molecular elements and bonds in a string, while an alloy’s supply chain, cost, microstructure, processing, and additive manufacturing versus casting cannot be captured that way. No single model can one-shot a material that ends up in an iPhone or on Starship.

  • That capability does not answer whether a material can be atomized into powder, additively manufactured, cast, scaled, or integrated. Worse, scale-up may reveal variables nobody knew to test, making the desired dataset incomplete by definition.

  • Krause’s standard for discovery is intentionally severe: hypothesis, synthesis, and characterization are milestones, not the finish. “We count a new discovery when you pick up your phone and there’s a new material sitting inside it.”

  • A 35-year 3M advisor crystallized the manufacturing problem: the vital data may reside in an operator who knows exactly when and how far to turn a knob. Krause’s honest non-answer—“we have not solved that problem yet”—leads to partnerships with established manufacturers while Radical learns to instrument and automate those tacit processes.

11. Hardware friction is becoming an infrastructure moat

  • One early war story involved instruments whose software did not expose an interface, forcing a two-week software sprint to work out programmatic control. Radical was paying for the software and had to be strategic about obtaining access; Krause said the team found ways to control what it needed.

  • Building autonomous inorganic science splintered the team into experimental and computational materials science, mechanical engineering, mechatronics, full-stack software, applied ML, robotics, path planning, perception, and computer vision. “It’s not about a robot in front of a tool”—every downstream integration problem appears once that arm is installed.

  • Krause sees three reasons the timing now works: machine-learned interatomic potentials can speed parts of the computational funnel; robots, grippers, and actuation are cheaper and better; and instrument vendors increasingly maintain software teams that support automation interfaces.

  • The fundamental limit remains long physical feedback loops. A facility containing 1,000 XRDs or SEMs could compress them through parallelism; the more transformational fix would be vendors rebuilding instruments “for agents and robots,” so researchers operate the scientific system rather than train individually on every machine.

12. National scale and open models reinforce an experiment-first moat

  • Krause described China’s manufacturing innovation hubs as an advantage the US “should not mimic” institutionally but must answer operationally. In his description, whether public or private, one entity can control the relevant pieces and support a new material through scale-up, directly attacking the 25-year gap Radical wants to shorten.

  • Radical’s productivity example is one metallurgy PhD running 10 campaigns, versus roughly 10 scientists historically focused on one research problem. The hosts noted that China can do the same; Krause’s answer was sustained investment plus a distinct public-private model.

  • National labs contribute HPC, researchers, instruments, and, Krause said, perhaps the world’s deepest store of experimental data. He cited self-driving or semi-autonomous work at Berkeley, Argonne, Ames, Livermore, and Oak Ridge, alongside the Genesis Mission and hundreds of millions of dollars in investment.

  • Radical’s AI scientist is itself multi-agent: an orchestrator proposes and tests hypotheses; a literature agent extracts relevant figures; paid industry-standard datasets and previous experiments ground campaigns; and MATRIX, a VLM fine-tuned on “Quinn,” reads laboratory images. Krause said the public dataset showed, he believed, 5%-16% gains in general scientific reasoning, with math the stated exception.

  • His workforce call is specialization, not retraining everyone into the same hybrid role: “Don’t try to become a materials scientist. Be an MLE that works in materials science.” Unfamiliar ML perspectives can challenge decades of inherited laboratory procedure while domain scientists supply the intuition models lack.

  • Open source follows the same logic. Radical released MATRIX and its benchmark, and spun TorchSim into a nonprofit to harness community feedback, while its own experimental data remain proprietary. Better external foundation models are welcome because “we don’t sell models”; Radical expects experiments, automation, and accumulated physical evidence to remain the edge.

Joseph Kraus

This is the difference between AI for bio and AI for materials. If you look at bio, or maybe small molecules as a broader category, you look at SELFIES and SMILES strings, right? That has been a big way to have those materials, those molecules, in text. And then you can use that because you know the elements and the bonds, so you know most of the things you need to know.

What about everything I just told you about the alloy, supply chain, cost, microstructure, how you're processing, additive versus casting? How do you capture that in a string? You can't. This is what's so hard: there is no one model that can one-shot a new material that ends up in your iPhone or on Starship. That's just not the way materials work. And so there is this really tough challenge: how do you capture all this data and try to bring that back and really improve your AI engine to encompass more than just discovery?

swyx

We're in the room with Joseph Kraus, CEO of Radical AI. Joseph, you're in a market that's getting crowded really fast. You have Lila, you have Citrine, you have Periodic, all developing AI for materials science. What are you trying to do that's different? And how are you going to beat the heavily capitalized competition?

Guys, great to be here. Thank you so much for having me. I must start with this: I'm a big fan of the show. I get to commute in New York City every day, and you're one of the top things in my rotation. I always love learning, and I'm a materials scientist by training, so the aspects that I can learn from your show are awesome. I'm super excited to be here, especially in person. Thanks for making it work.

What makes us different is our deep belief in experimental data, right? I think now you're starting to see the industry pay more attention to this. You see self-driving labs, a talked-about concept everywhere from academia to people like Google DeepMind, all the way through to pretty much every competitor that you've named in the space building an SDL.

It was not always that way. When we started the company 2.5 years ago, people would have thought we were crazy. That's CAPEX-intensive. Are you really going to be able to pull the data? Models aren't really built for that data today, and we can get into why models struggle in materials science, particularly inorganic materials science. And so why are you going to do that?

I had a deep belief, and my co-founders had a deep belief, that in materials, the ground truth is the material itself. You have to be able to make it, you have to be able to test it and characterize it, and then you have to, at one point, be able to see if it can go into a real application if you're going to have it used in industry.

That was where our thesis really started: you're going to build this loop, this closed-loop system, what we call a self-driving lab, that can actually run those experiments, capture that data, and feed that information back to your AI scientist so that it can learn and actually predict materials that are relevant to industry. That is what our whole company is built around, and for the last 2.5 years, that's what we focused on building.

Alessio Fanelli

Why do you believe that versus, pejoratively, the “think big thoughts, come up with stuff, and then try later” approach?

Joseph Kraus

So much of what makes a material real is in the latter part of the discovery process, specifically at the characterization and synthesis phases. What did we make? Does it have some cool properties in the lab? But also, what happens after that?

We work in a field called structural metals, or alloys. So much of what dictates the performance of those alloys is actually in processing. How do you manufacture it? What techniques are you using in post-processing and manufacturing that push performance or change performance?

You can generate a new composition, and that's very important to do, and we do that with AI. But it's everything that comes after that that actually impacts whether you have a new discovery, whether that new discovery is relevant to the application space you're going for, and then whether you can actually make it. Can you scale it? Can it actually go into that application space and be used by an end customer?

Those 2 and 3 don't get solved with AI today, right? A model can't figure out your way through the qualification pipeline for a new alloy for a jet turbine. You have to do experiments to do that. And so that ground truth is really important for us to bring back and understand: We know what we want to make, but are we actually making those things? Do they actually have the properties we care about? Can we actually push them to industry?

Those latter questions are the hardest questions to answer in materials. One thing you always hear about, which is true in materials, is long timelines. We've heard everything from 15 to 30 years—pick your favorite number, whatever you're feeling on this day of the week. But the point is, that's true.

The reason why is that materials is so fragmented today. Academia handles discovery and some of the lab-scale testing. You have small companies that will look at it, typically supported by the Department of War, the Department of Energy, or other government programs like NSF. And then you have late-stage companies, where bigger companies really optimize their current systems today.

They're not focused on room-temperature superconductivity, high-entropy alloys, or new ceramics. They're focused on how to take their current material system and make it 5% or 10% better, and capture the margin from that. There is so much fragmentation across this whole industry that the data never gets shared, and the connection from discovery to manufacturing is typically lost in that process.

That's the connection that we want to bring back to materials science. That is what we think the true opportunity is for AI and autonomy in materials: linking those two together in a fully closed-loop system.

swyx

I want to dig into that. I have a rule of thumb that I often follow when thinking about things: anytime you change orders of magnitude in a scaling system, your problems completely change. The orders of magnitude of the problems that you're talking about are drastically different, right? In discovery, you have an N of 1, or an N of some small number. On the commercialization side, you have an N of millions or whatever, right? Why do you think that you're capable of solving all those intermediate problems?

This is a good question, and I think I'll use a real, practical example to explain how hard it really is. In our field, with these alloys, one of the important things that determines properties is the microstructure. How does the microstructure form in this alloy so that you can see things like strength, ductility, emissivity, or whatever your other favorite mechanical property is?

At the generation level—candidate generation, hypothesis generation—you can predict a new composition. AI is actually quite good at that. All of our hypotheses are generated by our AI scientist today. You'll take that composition and synthesize it in the lab.

There, step 1, something changes. It might not be homogenized. There might be dendritic formation on the surface. You might see different phases, or it might be single-phase. Those dictate what the properties look like.

Once you move past that, you actually go to manufacturability: annealing or thermal processing, and looking at how you manufacture it. Is it additively manufactured with powders, or are you casting with actual raw metal? Both have wildly different outcomes in the performance that we like to see.

To answer your question, the first step in solving this is capturing that data. We do that at the discovery and testing phases today. We don't do the manufacturing part, to be clear, but we do that at discovery. We do synthesis and characterization. We have a bunch of characterization tools in our lab: SEM, EDS, XRD, XRF, and TGA.

swyx

Wait, real quick. That was a lot of acronyms. Would you like to explain them now? We can also pause this, and we wanted to talk about them later if it makes sense. Or we can just talk about it.

We can come back to it, and I'm happy to dive into everything that we're doing in this.

swyx

For now, those are just a lot of acronyms. You don't need to know what they mean.

They're just a lot of tools that tell you different things about a material in a lab, and we can talk about what they do. Then we do testing of properties at the lab scale. We'll look at oxidation performance in our lab today, which is really important to see how these alloys perform in oxidative or corrosive environments.

We'll look at mechanical properties, something called a tensile test, which gives you these stress-strain curves of a material. Then we'll look at microindentation. Here, you can pull what's called the Vickers hardness from a material, as well as a proxy for ductility. It's not an exact measurement of ductility; we kind of pull out whether the material is ductile from that.

So that's everything happening on the discovery side and moving into the testing side. We have not yet crossed into, “Okay, now when we go to manufacturability,” but we hope to. Back to your original question, if you can capture the data at the manufacturing side as well, now you have the whole suite of what we call the lifespan of the material.

swyx

I see the hypothesis. I see the synthesis. I see the characterization. What do we make? I see the early properties showing good results, and then I see the manufacturing and what came out of the back end of that. Does it actually make it to the end system?

Now I've seen this lifespan of the material. Now I can use that to go pick more materials targeted at the right applications. That is the North Star of what the company wants to go out and do.

Where do we stand with that today, then?

Shawn Wang

I'd say we're really good at the first part: the discovery and that lab-scale testing I've mentioned. There are some testing mechanisms that we use externally, like with third parties, where there is deep expertise required in the industry itself. Aerospace is a perfect example. If they do wind-tunnel tests or torch testing, we don't have those capabilities at Radical today, and so far we haven't needed to own those. We want to use third parties.

There's a heavy tail.

The Self-Driving Lab

Exactly right. Really heavy. Then you look at the cost of a wind tunnel and you think, “Yeah, I'm never going to get an update from that.”

Lastly, we haven't touched manufacturing. We have spoken with people who do manufacturing, and we do know some of the things they care about. Processability is a really good one. Can this material be formed? If you're going to cast it, can I actually move it into the shape that I need it to be in? We do look at that, but we look at it at the small scale, not at 10 tons.

We look at grams—200 g, 500 g of material—so not at that larger scale yet. But that first section is done; that's running today. We've probably made 1,200 alloys in the last 5 or 6 months. 300 of those alloys are new, novel, never before seen in the literature. I'd say probably 10 of those alloys have performance that has got us very excited about where they're going to be in the industry. That's a rough scale of where we are today from a company perspective.

Shawn Wang

That brings up a follow-on question about how much you're just optimizing within a well-understood space, picking permutations that are new and novel, versus trying experiments that are really pushing the frontier of science.

The Self-Driving Lab

Yeah, the latter. We really are making new materials that push the frontier. A good example is that we work in a field called high-entropy alloys. These alloys are really exotic because they have 5 to 7 elements in the system, all roughly equiatomic, give or take, and they have really exotic properties in extreme environments.

Think super-high temperatures, usually north of 2,000°C, even 3,000°C. They have very high pressures—think space, coming back from space. Then they have environments that can be corrosive, like a nuclear reactor where you're feeling neutron bombardment, or oxidative, like if you're flying in a defense application or in a jet turbine.

These alloys are really exciting, and for the past 50 years, the same alloys have been used in all these industries. The reason why is the long discovery timelines we talked about. This is a perfect example where we're working in an industry; we're not really creating a new industry. Of course, turbines—there's a big industry, it's a great one. It's having a second tailwind right now.

What we're trying to drive there is new performance that does not yet exist for the materials they have today. We stole this term from a gentleman named Charles, who's the VP of materials at SpaceX, which is called concurrent engineering. It's this idea that I can actually design my materials as I'm designing my product.

As I make a new rocket booster, a jet turbine, a missile, or a solar cell, I'm actually inventing the new materials that meet the property specs for it. I'm going back and forth as I engineer them to get to the application. We don't do that today. The alloys that are in the plane I flew here on are from the 1950s, 1960s, or 1970s. They might be coated with some CMZs from the late 1990s.

There's a real opportunity here. We're tackling industries that exist and have huge markets, but we're bringing novel materials that historically they never would have had the ability to look at in enough throughput and at enough scale.

Shawn Wang

Going back to the bottlenecks you were talking about before: you said that it takes 15 to 30 years to get material through, and part of this is because there's a disconnect between research and productization. How do you validate that this is going to be the thing that will work, and that you're not going to get killed by another unexpected bottleneck in the whole process?

I'm coming from the world of drug discovery, where even if you think you have good early-stage data, there are all these things that will get you later on in different phases of clinical trials. I wonder, are there things like this in materials? Where do you think these are going to get hit? What will stop this super-cool alloy you just developed from actually making it to market?

The Self-Driving Lab

Really good question. There absolutely are those in materials. If we talk about the alloy example I just gave, 1 of those areas is called qualification. Qualification is this process that, if it's going in manned flight, is run by the FAA. There's also a MIL-SPEC one for the US military as well.

Essentially, your alloy has to qualify to be used in aerospace or defense applications, and that process is very slow. It's typically a 10-year process today. You have to make a number of different ingots of material and run these standardized tests on them to prove your material is usable in those systems.

There are a bunch of things that have a gotcha later. I think actually trying to capture that data and understand where and why those things are happening can impact your discovery loop.

Shawn Wang

It takes 10 years because, like clinical trials, you have a series of incremental phases. Can you Operation Warp Speed this, where you do these all in parallel, or is it really like you need to do these things sequentially?

The Self-Driving Lab

There are people working on changing that, rather than doing it sequentially. There is really good work right now—for example, at DARPA—where they're looking at new ways to do qualification. They can use additive manufacturing to go layer by layer and actually look at whether we can do qualification of a material that way. I would say that's a new technique that's trying to completely redo the way we do that process today.

Shawn Wang

So it's not regulatory.

The Self-Driving Lab

This is the challenging part. Some of it is regulatory in nature, in that it's a government body like the FAA or MIL-SPEC that runs it, and you have to see almost the human side of this as well.

It is different to develop a new alloy for an iPhone. If it bends in your testing, you can get rid of that or move on, or even, in a worst-case scenario, recall iPhones. That's obviously a terrible scenario for the company, but you can. But if it's going in a jet turbine and you're going to fly it on a 787, there's a serious bar that has to be met there, and for good reason. No one would want that bar to be removed.

I'm flying back to New York tonight. I certainly don't want that bar to be removed. To be super clear, make sure that's on the record.

I think the way we go about that process is very dated. What people are trying to attack is whether we have to do it that way, or whether there's another way to get the same result via a different mechanism. That's where AI is interesting, but autonomy has been really interesting as well.

As you make the process of manufacturing more automated, there are more sensors, therefore more data, therefore more things you can capture and analyze, and therefore a bigger loop that you can build around that. That has not been deeply extrapolated in the materials-manufacturing sense today.

Shawn Wang

So maybe 1 difference between drug discovery and materials is that, in drug discovery, you have different phases, each of which is designed to basically not kill people. Whereas in materials, there is, in some sense, no reason other than budget that you can't just do the—

The Self-Driving Lab

Solve it in 1 go.

Shawn Wang

And successive levels of qualification do not depend on the prior ones, aside from just budgetary constraints.

The Self-Driving Lab

As well as some of the other metrics you have to pay attention to, this is what makes materials so hard, actually. A perfect example is supply chain.

Probably 5 years ago, maybe a little bit longer, maybe 10 years ago, there weren't the constraints we're feeling today from the metals industry or the minerals industry. Hafnium is up 10–15× in price because China owns a majority of the supply chain. Things like refractories—

Shawn Wang

Tantalum and niobium. What is hafnium?

The Self-Driving Lab

Hafnium is an element in the periodic table that's used in things like C103. It's about 10% by weight of hafnium in C103, which is a very common aerospace and space alloy that's used today.

Now we're starting to see requests in conversations around, “Can you remove that element from that material?” There, you're actually trying to maintain the same performance specs, or you're trying to completely remove the hafnium from that equation.

We've worked on that problem specifically, and we have successfully done that. This is where you get back to supply chain being a concern, cost and margin being a concern, and who is paying for that and feeling it. I can tell you the space industry has much more tolerance for high cost.

Performance is everything. When I'm designing a new heat shield or a new cone that goes on the rocket engine, performance is the number-one thing I care about. Cost is not the first thing I care about. It's not irrelevant—I don't want to pay $100 million for a nose cone—but it's still not the top thing I'm thinking about.

You think about something like consumer electronics or maybe even a medical-device application where alloys go into products. Well, now cost is definitely much more sensitive. There are probably alloys we could put inside smartphones today, but they would just make them unbelievably expensive and probably not tolerant of some of the other things we have to put in there.

There are just so many things about a material that make it so much harder. This is one of the reasons why we deeply believe in self-driving labs. Just to come back to this point for a second, this is the difference between AI for biology and AI for materials, in my opinion, from the materials lens.

If you look at biology, or maybe small molecules as a broader category—you probably could include some organic materials in that—you look at SELFIES and SMILES strings, which has been a big way to represent those molecules in text. You can use that because you know the elements and then you know the bonds, so you know most of the things you need to know.

But what about everything I just told you about the alloy? Supply chain, cost, microstructure, how you're processing, additive versus casting—how do you capture that in a string? You can't. This is what's so hard: there is no one model that can one-shot a new material that ends up in your iPhone or that ends up on Starship. That's just not the way materials work.

There is this really tough challenge of how you capture all this data and try to bring that back, really improving your AI engine to encompass more than just discovery, and certainly more than just composition.

swyx

You mentioned this sort of loop, right? You're really doing 2 things with automation. One is you're collecting data; one is you're running experiments and building stuff. How iterative is that, and what does it look like at the different steps where humans could be in the loop?

Yeah, and humans are in our loop today in a very important manner: training and teaching what I like to call the scientist about what they know. We call this scientific intuition at the company, and what that literally looks like is that we have a scanning electron microscopy image—that's an image that takes a picture of the material—and our scientist will go in, analyze that image, and make comments in our system.

“Hey, I see dendritic formation on this image in these locations.” The AI scientist goes and looks at those comments. That's one amazing example of a human in the loop, where we are trying to download the brain of a PhD in metallurgy. When you look at this image, what do you see as a PhD scientist? We need to be able to replicate that as an AI scientist.

That's one way they're in the loop. The second thing—and we can go deep into this if you guys want to—the lab is not easy to automate.

swyx

Yeah, it is.

We're getting there. It is super hard to automate, both in ways that are hard engineering challenges and in ways that are annoying. The tool vendors just don't have SDKs or APIs to work with. There are engineering things that we can talk about, but even this idea of the tool provider letting you have access to the data via their software layer was not understood 2 years ago.

I can tell you that a few very big tool vendors were not too excited about self-driving labs 2 years ago. They were not jumping to give us, even with payment, access to the software and pull the data. That tone has now changed.

swyx

Are they trying to own it? Is that why?

From my understanding—and I'm not a tool provider, so they might give you a different answer—a lot of what they sell is the ability to analyze the data coming out of their tool. One of the things that makes those tools different is how they actually use the software to generate your spectra, and they like to sell on that. That's a really big thing for them.

If they give you access to the raw data from that and you no longer need their software, why would you buy their tool? We tell them, “No, no, no, you're way wrong,” and we're getting them there. It's a work in progress.

I do think there's been a lot of momentum in AI for science. Number 1, it's having an incredible moment, which is going to be so good for the world. Number 2, you're seeing, I'd say, academic and national buy-in, as well as private buy-in.

You see the Genesis Mission. You see the Department of Energy and the national labs moving this way. You see people like Google DeepMind, Microsoft, and other places like Meta either building their own lab or running experiments at someone else's lab to get that data back.

Then you have the private companies that are forming self-driving labs and looking at the automation of scientific equipment. That has really started to push this wave to, “Oh, now we don't really debate that AI labs or self-driving labs are a part of the future anymore. It's which part are we going to play in that?”

That's a better, more helpful conversation to have now because we can get access faster.

swyx

Self-driving labs in biology have been notoriously difficult. You can automate certain parts of them, but you inevitably have people who are just walking around moving trays from one section to another, and it doesn't actually end up speeding things up. Oftentimes, it can even slow things down.

Full end-to-end automation for non-research activities—or, you know, for manufacturing—we're really good at. But when the process changes, how do you deal with that? Have you figured out some way of automating the type of problems you're solving?

The self-driving lab—first, I think it's important to talk about what a self-driving lab is, because this impacts your answer. There's a difference between an automated lab and a self-driving lab, right? An automated lab does experiments for you automatically, without humans, and at high throughput. That can be very effective.

A self-driving lab runs research campaigns for you, and there's a big difference. The way I like to describe it is, one of them is like hands-free driving. I don't have to touch the steering wheel; it'll keep me in the lane and keep my speed set. But when a left-hand turn is coming up, I have to pay attention, put on my turn signal, turn the car, and know to make a left.

Now compare that to a Waymo, which I love bringing up because every time I'm here, I'm going to go out of my way to just drive around the block sometimes. It's living in the future. You don't need to make a left-hand turn. Actually, you don't even need to know to make a left. You don't care what route it takes to get you there.

You get in the car, you can close your eyes if you want, scroll X, work on a research paper, and then end up at your destination without knowing how you got there. That is the difference between an automated lab, which you are controlling and just using automation to do throughput, and a self-driving lab, where it's actually doing this entire process for you.

In the self-driving lab, there are things that a human scientist does that are actually very hard. Sample manipulation is a perfect example. When we synthesize these alloys, we get these little pucks that come out. They're called buttons in the industry, and because you're blasting them at 3,000 or 4,000 degrees, they get stuck to the tray.

How do you get them out? You have to be careful because you don't want to mess with the microstructure or chip off part of it. They're strong enough that you're not really going to do that, but we had to design custom actuators that go on our robotic arms to be able to manipulate them.

That does not really have anything to do with the discovery of this new high-entropy alloy. That's just required if you want to run autonomous alloy science. One answer is that there are these challenges that humans either don't face or, if we do, they're very intuitive.

The button's stuck, so I take a little chisel, smack it, flip it over, and move on. I don't even think twice about doing that. Not so much for a robot. The second thing we talked a little bit about is the software.

It's not just about controlling the tools; it's about running the lab. How do I track my samples? How do I know what sample should go in a tool or should not go in a tool? Is there a quality check where, if I look at a sample after it comes out of synthesis, I actually want to kill that experiment?

I don't need to waste time going through XRD and SEM and the other tools in the lab. I want to just stop that sample, save its state, and throw it away. How does it know how to do that?

This is where you start to bring in all these different factors from vision, different sensors in the lab, and sensors on the tooling themselves to build what we call an operating system that runs the self-driving lab.

And then the third part is automation. Automation includes what I like to call the connection of the lab. I have one tool that's automated, and I have another tool that's automated. You have this operating system that's running them individually. How do you connect them? The same way a human scientist would come in, look at the results, take the sample out, and go to the next one, our robots do that today.

Those are the three parts that we really see making up this self-driving lab, and I'd say each of them has its own difficulties. We can walk through them. Some are, like I mentioned, the tool provider. Some have no actuators. Some are just really hard to load and unload, like XRD. You have to put it in a hard sample mount, and it's weird and awkward geometry, and you need a custom gripper. That's just what it is.

All of that kind of forms what a self-driving lab becomes. We're very good at that for alloys today because the tools in our lab are built for alloys. Some tools are shared, like XRD, XRF, and SEM. Even tensile testing, although mostly used in the structural metals space, can be used elsewhere.

I would say our oxidation chamber is very suited for the specific customer application we're going after. I would say our synthesis mechanism is directly for alloys. We custom-built a tool with a third party to do alloy synthesis at high throughput. That's built to do alloys. It's not built to do ceramics, polymers, or any other material system today.

swyx

Is the goal to expand to polymers and ceramics and everything, or is it that we're going to do alloys and basically get all the way to end-to-end manufacturing on alloys and then expand?

Yeah. Or never expand.

It's both, but on the right timeline. The first one is vertical integration, and this is very important. When we started the company, and I'm not afraid to admit it, we were like, "We're just going to do 7 different labs, 7 different material systems, across the board. We're going to capture all this data. It's going to be amazing."

Then we started talking to customers about it.

swyx

Yeah, exactly.

Exactly. And that'll come back to doing that over time, but we started talking to customers, and they're like, "How are you going to do this? Have you thought about this test? What about when you go to scale? What about when you need to go to 300 pounds?" And we were like, "Oh, no, we hadn't thought about that per se. We were more worried about going to polymers and ceramics and everything else."

Why is that important? Because we are a materials company. You see the company talk about this a lot. We really believe in inventing new materials that change the future of the world. We think that is the opportunity with AI and autonomy.

There are so many industries that we all care about that are blocked because of a lack of novel material advancement: automotive, aerospace, manufacturing, defense, climate, energy, semiconductors, and electronics.

Shawn Wang

What's your favorite example of that? What's a problem that they unlocked with an amazing material?

The Self-Driving Lab

Aerospace and semiconductors immediately.

Shawn Wang

Well, what specifically?

The Self-Driving Lab

In back-end-of-line integration for semiconductors, there are particular materials that we've been using for a long time called interconnects. The entire industry is doing R&D here. They are a cause of not having great efficiency and very high energy bills. At the back end of that, a new material could potentially completely remove that problem. That's probably an example.

Shawn Wang

Is there a theoretical efficiency that you could achieve, and where are we now compared to that?

The Self-Driving Lab

We have ideas for materials that would be able to solve that problem, but we'll release more on that in the coming months. We have a specific program we're working on that's directly around that problem. It's a really exciting one, and there are estimates from the industry on moving past the current material system and what new materials could bring, but there are other challenges.

Shawn Wang

Are we talking about 2× more efficient, 5×, 10×, or 1.1×, which is often quite huge?

The Self-Driving Lab

Yeah, of course. I think you could, in the near future, with some of the systems that have been recommended today, see 2× to 5× generally, especially when you think about integration. In these materials, you need barrier layers, and there are all these interface things that you have to be able to understand.

I think beyond that, you could start to see it push over 10×. I don't know to what level, but there are cool, exciting materials that we'll share more about in the future.

Shawn Wang

What's the validation time scale for this?

The Self-Driving Lab

In which way? The lab starting to work on these materials, or a material being in a new chip in the iPhone?

Shawn Wang

You have a material you just created that you think is going to revolutionize the world. How long do you think it's going to take for that thing to get into an iPhone or an NVIDIA GPU?

The Self-Driving Lab

That's still long. That's still pretty long. We're new and early in semiconductors, and actually, what's long about that is that when you start from ground zero, you have to build everything from scratch.

Even the way that we do material testing for that industry today—I would say they're one of the industries that's actually far ahead of everyone else because they invest so much in R&D, and materials really make or break some of their performance—even there, that integration timeline is very slow.

And number 2, no one can get enough chips, so everything is delayed in that industry, which is important. I even saw today, or this week, TSMC telling ASML they're going to hold off on some of those new tools until they get through their 2029 production run, or it was some story like that.

Shawn Wang

Go, go, go.

The Self-Driving Lab

That's what I believe, yeah. I'll find it afterward and send it to you guys, but I was like, "Man, this industry is really getting pushed to its limit." I would say something on the alloy side, like aerospace, offers a good opportunity on a 3-to-5-year timeline.

Shawn Wang

That's pretty short.

The Self-Driving Lab

Yeah, correct. I think it'll be an application, not manned flight. I don't think it'll be a jet turbine because of the constraints there with humans. I think defense and space systems, though, are definitely doable in that timeline.

There have been examples in the past of people who have done that in that timeline, and we feel very confident in our ability to try to execute on that.

Shawn Wang

I want to get back to the validation question because I feel like this is the crux of automation, right? Every AI engineer who uses cloud code or something has the experience of one-shotting something, and then when you look at it, you're like, "What is this garbage?"

If you're talking about what is essentially active learning, where you are hypothesizing, manufacturing, testing, and then forming hypotheses on that, and doing that over and over again, your mistakes will obviously compound as you do that. So, how are you thinking about this?

The Self-Driving Lab

We call these the negative results, these mistakes, and we do have a version of this loop built already today. It doesn't include manufacturing data, like I mentioned—we don't have manufacturability in the lab today—but it does include synthesis, characterization, and those early property tests that I described to you guys.

What this system does is have this AI scientist that's really good at designing the campaigns I talked to you about. It can come up with these campaigns, determine the number of materials that it's confident it wants to make and test, and launch that campaign.

It'll send it to the lab, it'll start running autonomously, and then it will go through the whole characterization suite, and we'll get all the data. That data is all pulled out autonomously. Some of it is analyzed with machine learning models, like computer vision. Some of it, again, has a human in the loop analyzing it as well, and it'll get put in our database. That AI scientist will look back to that data when it designs the next campaign as a follow-up.

So, it's this active learning loop on a campaign-by-campaign basis. It's not every experiment. Honestly, we don't need it to be that fast. We actually want to take a few shots and get enough data back to change our hypothesis, but it is very rapid.

I would say we probably could run 7 to 10 different campaigns right now inside the lab across different systems, and we're updating those daily, or at least every other day, with the results we're seeing from the lab when they come out.

Shawn Wang

And what does the human do in that part? I find that when I'm doing transcriptomic analysis or whatever, Claude does some of my work, and then I end up course-correcting. Is that the gist of what's going on in your lab as well?

The Self-Driving Lab

In some parts, yes, though some parts are more complex, to be honest. Synthesis is a really good example. We still have PhDs in metallurgy running our synthesis machine. That tool is not yet fully automated. It should be automated by the summer. That's the custom tool I was telling you guys about that we're building with that vendor.

It's not just opening and closing that's automated. The synthesis machine itself requires automation. If you've ever seen how you cast these alloys, you take a plasma torch and blast them, melt these raw precursors down into a liquid, and then cast it, and it solidifies into the shape that you cast it into.

So, there’s a lot of intuition in that. A scientist stares at it and looks at it: “That’s not melted yet. Let me hit that corner there.” We have models built that can actually start to learn how to do that at the same performance as a human scientist can.

So, we’re getting up there. That’s not fully automated yet. Characterization is fully automated. We just have scientists annotating images or results afterward to train the AI scientist on. All of our characterization tools can be loaded, unloaded, and controlled with our back-end operating system to do characterization.

Property testing: 2 of the 3 property tests are fully automated. Microindentation and oxidation are automated. Tensile is almost automated; it should be automated in the future. So, that’s where humans are now.

After the process for generation, all of our materials are generated by our AI scientist today. Occasionally, a human scientist will try to compete, and I love telling this story. They hate when I tell this story. They’ll throw in a composition, and then the AI scientist says, “Get that out of here. That’s not strong enough.”

Or we’ll see, “How do we think about something new?”

Shawn Wang

Actually, it’s like you’re red-teaming.

The Self-Driving Lab

Yes, that’s a good way to describe it. I don’t know if they would call it that. I think they would call it losing their job, but they’re not, actually. They’re super important to the process.

I think what’s cool is that sometimes the scientist will recognize a new learning: “Oh, interesting that you threw that element in there. That’s cool.”

The other part of that is going places where human scientists won’t go. We have this beautiful chart that shows, in publications—in the literature—what we can access, where all the places scientists have gone. Then we have a second overlay on that chart showing where our AI scientist has gone.

Shawn Wang

Mhm.

The Self-Driving Lab

It’s moved into elemental families or alloy families that no one has ever published on before. The question is, why? Why did it think to do that?

We ask our scientists, “Why did you never go there? Why did you never use that element or that element?” Their answer is, “I just didn’t think it would work with the other elements that are in that mixture. I didn’t think it would cast. I thought it would evaporate when we tried to make it, and it didn’t—we were able to synthesize it.”

“I didn’t think it would work in the microstructure, or would cause grains to be not what I was looking for, and would not get the mechanical properties I thought. So, I just never considered it.” But it actually works in that formation.

There’s this really interesting feedback where now the scientist is getting good at exploring places that, I’d say, humans have a natural bias against, even though it might be an unknowing bias. That’s a huge power of an AI scientist.

Shawn Wang

But is part of this just that you have higher throughput and that you’re letting the AI scientist do its own thing? I wonder what would happen if you took those same scientists and said, “All right, no constraints. Just go crazy. You can do whatever you want to.”

swyx

If you have an AI scientist that doesn’t have preconceived notions, I’m honestly kind of surprised that it’s not just reiterating what is known. But I wonder how much of it is that you’ve just turned up the temperature in your sampler.

The Self-Driving Lab

Yeah, that’s an important metric. Processing is really important. To be clear, there are times when it goes to what it knows, especially when it pulls in literature. It’s like, “Oh, this is where it is.” Literature is a great teacher. If things work, it’s actually a good place to ground on why they work and try to understand why they work. That’s a big problem for the materials field that we can talk about.

But it also has this good ability, because it is high-throughput, not to be afraid to test. When I was in my PhD, I probably did 50 experiments a year—rough estimate, something around that. Every experiment is important, right? Not to mention the mental load: 2 weeks at a time to fabricate this thing, synthesize it, and then go test it.

The scientist doesn’t think like that. This AI scientist is like, “I’m making 8 of them today, 20 of them today, and once that tool is done, I’m making 100 per day. I don’t really care about taking a shot on goal and learning from that shot.” So, it’s a mindset shift.

swyx

How much does 1 experiment cost?

It depends on what elements you use. Some elements, like platinum and palladium, are much more expensive than aluminum and titanium. It’s anywhere from $60 up to $300.

swyx

Okay. It’s all element-dependent, though, usually. What’s your throughput?

Today, it can be anywhere from 8 to 20. That depends on the elements as well. Refractories, particularly if we’re doing refractories, are much harder to cast, so we go down to that 8 number.

If you’re doing things like titanium and aluminum—your standard alloy, Ti-64—those are much easier to cast. They melt immediately or quickly, so we go to a higher throughput of 20. We should be at 100 per day regardless of the system by around the June or July timeframe, rough estimate, give or take.

swyx

This is across the entire lab, not per workflow?

Yes, that’s correct. That’s across the entire lab.

swyx

Okay. Eight to 20, and you could conceivably have humans inspect many or most of these, or is that—

I don’t know. You could.

One thing I wanted to touch on—you just reminded me with that answer—is that AI doesn’t operate in the same dimension that humans do. Let me explain what I mean. When I was a scientist, I went through a very serial-based process. I read a bunch of papers, made a new hypothesis, and might run some computational workflows—DFT, MD, or ML. Then I’d go synthesize it in a lab and move to characterization. I’d study my characterization for 2 weeks, get a new idea from an image or something I saw, circle back, and do that whole process again.

That’s what a human does, and it’s very serial. If I take 100 SEM images, I don’t memorize all 100. I’d love to think I could—my advisor would have loved it if I could—but I couldn’t. So, I pull one thing or a couple of things out of that that I want to learn.

Now switch over to the AI scientist in that same process. Now it’s parallel. I can read 100,000 publications and directly compare them to 100,000 SEM images at the same time, in real time. I can study, learn, and memorize all of the things I’m seeing in those 100,000 SEM images and draw direct conclusions back to my papers, back to my hypotheses, or back to my mechanical property testing, where I want to see what actually comes out.

I can’t do that as a human scientist. This parallel nature allows it to operate in a way that human scientists simply don’t have the ability to do.

swyx

Okay, but you’re still talking about on the order of 10s or 100s of materials that you’re producing per week or month. The overall scale is not that large compared to even a lot of biological methods. You have ways of scaling up to 1 million if you’re doing, let’s say, next-generation sequencing-based assays. You can do billions, whatever.

This is much more reminiscent of ligand-based modeling, where you’re really looking at a small number of examples and trying to pull out local patterns. For small molecules, you can have real predictive power. These are useful techniques. But the almost universal rule from experience in cheminformatics is that, oftentimes, by the time those become useful, the actual scientist can just go out and do it. They could have designed it or found the molecule they were looking for without using the AI model by the time the AI model gets there.

So, this is a very specific kind of regime, but it’s one that has been well established. I’m wondering, since the timescales in this sort of data seem very reminiscent of that, why is this actually that much more effective?

Yeah, 2 things there. Number 1, throughput is an important number. To our knowledge, based on what is publicly released—if there’s someone who has done it behind closed doors, I don’t know about them. Please, I would love to talk to them.

The largest alloys program was the MACH program. It was run by DARPA and GE Aerospace. They did a bunch of AI and simulations on the front end, and then they synthesized 500 new alloys in about 12 months. That’s kind of the benchmark, I would say, for how many alloys someone can do in a year.

Again, we’re trying to do 500 in 5 business days. We’ve done 1,200 in 3 months. So, that’s an order-of-magnitude step up that we’re moving to.

The second piece of this is that what’s really challenging in the alloy space specifically—and I think it’s probably specific to alloys, though I do think it will carry over to some other industries—is that there are so many variables that go into determining your end product and then your end properties from that product. That makes it harder to do discovery because there are endless potential combinations.

There are 10^40 different potential alloys that you could go out and synthesize. How do you do that? Even if you could do high throughput, to your point, it's still not that much throughput, right? We think it would take humans 7 million years to make all of them. What do you do even if you're only doing 30,000 a year?

The screening mechanisms here are very helpful, as we all know. That's where AI is great. But I think the other point is that the data is missing from the industry. We don't have experimental data.

We do see results from 100, 200, or 300 experiments quite aggressively. We see 300 new alloys that we've never seen before from the experimental results of 1,200 alloys that we've run to date. You probably think each alloy has anywhere from 50 to 150 different data points, depending on how many images you take, how many spectra you run, and so on. That's a fairly small data set.

It's funny when I talk to the ML side. They're like, "How are you going to get millions and millions of data points?" And I'm like, "I just don't think you need to. We have not seen that you need to make new discoveries yet today. We have a bunch of new discoveries, many that are going through patent protection and that we're talking to potential customers about. We just haven't needed millions of data points."

This brings up the arguments I was getting about compute. I got an argument at GTC about this: we're not compute-constrained in the materials industry.

swyx

Yeah. You're making stuff.

Yes. We're experiment-constrained.

swyx

I mean, this is even what I do, which is computational. It's oftentimes dominated by data movement, right? It's not like I can—I see these people, and I'm jealous that they have 14 Claude sessions going at night. They have all these different experiments going, and I couldn't do that just because I can't move the data around fast enough.

Yes. That's almost a good comparison to our world. It's not a model problem. It's not a language problem. We don't have the same problems there. It's an experiment problem. It's really: how can you run enough experiments to start to change the output of an AI scientist and capture the data you need to discover something new?

For us, that's really what it's about. That is our bottleneck. That is the throughput. That's why we're so bullish on self-driving labs. That's why, when I start the conversation, "What do you guys work on?" it's all about the self-driving lab, the autonomy, and the experimental data.

We're trying to build the Protein Data Bank for materials, and it's much more complicated than just crystallography structures or whatever else was in there. There are all these different properties you talked about today that have to be inside that data set to make it relevant. It's hard to do.

swyx

That reminds me of Heather Kulik’s episode, where she said that there is no AlphaFold for materials. First of all, do you agree?

I do. I think you can add AlphaFold moments for specific areas of materials, like microscopy, for example. Reading, using a segmentation model on SEM images—our team does that today. That's a cool AlphaFold moment, whether you call it that or not. I don't know, but there are real-world models that make a huge impact on being able to do that.

What I don't think you can do is go from, "I have this new hypothesis," to, "Oh my gosh, I have a new material. It's scaled, it's done, it's in products in your iPhone." You can't do that today. I would agree with that statement.

swyx

Okay, but even then, AlphaFold solved a scientific problem, which is: how do you take a protein sequence and figure out what the 3-dimensional structure of that protein is? I want to put in so many caveats so that none of my structural biologists flame me.

Yeah, yeah.

swyx

But anyway—

Luckily, I'm not a structural biologist, or everyone watching, so I get a free pass.

swyx

But the thing that seems useful here is that you put in a chemical formula and some sort of processing, and what you get is—essentially, what you call—the microstructure, which is something where you can get part of the information from X-ray diffraction, but not all of it. That problem sounds much harder in a lot of ways than the biological problem. Can you maybe explain a bit more about that?

Shawn Wang

Yeah, and it's even harder than what you just described. There are probably things you don't know to test for yet, or that you might see when you go to scale that you did not know you should have predicted or been paying attention to. That's really hard to build a data set around.

What I like to tell people when they ask about this is, "Can you build the database for materials?" Well, if you want to do SEM images—scanning electron microscopy images—and use segmentation, yes, you can build a really large data set of SEM images that are very good at finding dendrites, cracks, or defects in a material. The model will be very, very good at using that to predict and relate that to a mechanical property, because you're looking for what crack propagation does to strength.

Okay, tracked.

The Self-Driving Lab

But that's not the same thing as, "Can you make it? Can you atomize it? Can you make it with the powder? Can you use additive manufacturing? Can you cast it?" Way different thing. It's related, obviously—the microstructure relates to that problem—but just because you understand the microstructure, just because you see and can predict crack propagation, does not mean you're necessarily going to perfectly nail manufacturing.

There are just so many things that we see stack up, and we learn a lot of new things. Every time we think we know everything, we go somewhere and learn new things along the way. I do think it's multifaceted, for sure, and each inorganic material has different constraints.

I talked about the supply chain. That's very relevant for defense applications; those are a perfect example. Supply chain is not one of the things we worry about in consumer electronics, per se. Ti-6-4 is still there. It's available. People care about where it comes from, but that's not the same as the critical-minerals focus that you see in the US today: where are we getting these minerals, the vast majority of which we do not control? That's a different problem.

There are different inputs that go in to get an output. I feel like that's why materials are so hard. It's all of this other data that comes after discovery. What I always tell people is that the second you design a new material, that's a milestone. The second you synthesize it, that's a milestone. The second you characterize it, that's a milestone.

That is not a new discovery. We count a new discovery when you pick up your phone and there's a new material sitting inside it. That, I think, is a fair claim on a new discovery in a scaled material.

Shawn Wang

As you get past—there's fallout in every one of those steps, every milestone. Presumably, in manufacturing, there are separate steps where there's fallout as well. As you get closer and closer to the consumer or the application, then you have less and less data, right?

Fundamentally, how do you get over that? Because I think in pharma right now people are starting to think about rules of thumb that can be used to do reverse translation back from the clinic to the discovery process. How do you do that in materials?

The Self-Driving Lab

It's funny. I love telling the story of one of my advisors at the company. He was at 3M for 35 years, and we asked him about manufacturing: "When you guys go to manufacture, what are you paying attention to?"

This was early. This was 3 months into the company. He said, "Whew, you guys got a lot to learn." I was like, "What do you mean? I'm a materials scientist." He said, "Different worlds."

One of the challenges he pointed out was, "You know the hardest part about the data you're asking for or inquiring about? That is someone with a 35-year trajectory at the company who knows exactly where to turn the knob on whatever manufacturing tool you're talking about. What you're asking for is his or her ability to know when to turn the knob right to that spot at the specific moment. How do you capture that? How do I give you that? Even if you assign a formula to it, is it the same every time?"

Again, this gets back to intuition, which we've talked a lot about today, just in the manufacturing sense. This is the hard part. I don't have an answer for you on manufacturing because we haven't done it. What I do have a lot of answers on is the discovery side, where we've had to look at where intuition is important.

I talked about the casting of alloys, which I touched on earlier. That's one of them. Reading SEM images, that's one of them. Looking at XRD spectra and identifying phases and how strong the peaks are: that's arbitrary. That's really strong. That's kind of strong. That's not really strong. That's terrible. What do any of those mean?

I can guess what they mean, but if you look at an XRD, you might not get it perfectly compared to what you or I think about that. This intuition aspect is so important. This is why we still have humans in the loop, because you want to capture that. Now, when you go to manufacturing, we think we'll have to do the same thing.

And we think the opportunity is to rebuild those processes fully automated. You can put all the sensors and capture mechanisms—in the absence of a better phrase—in place so you can bring all of that back. That's a hard problem to solve, and to be clear, we have not solved it yet. We are certainly still at the discovery and testing side of that, but that's where we want to get to. How do we get there quickly? Partners.

We talk to a lot of companies in our field that make materials at scale, particularly in the alloy space, and are thinking about this. They look at it from a different lens. They're not all hyped about AI for science. Actually, I'd tell you that a lot of them are bearish. They're like, “You don't know what we know. We've been doing it for a long time.” And that's okay. I think that's healthy.

What they do know and bring to the table is that we have that intuition. We will tell you, when you show us a family of elements, what we think is going to work or not. We might be wrong, but we can tell you why we think that's going to happen. We can tell you why that relates to aspects of a business that are important, like supply chain, cost, and performance under certain environments that don't exist in others—extreme environments, for example, involving temperature and pressure. That's really important information that you want to bring back. That's how we get there in the near term, until we can do it ourselves: you partner with people who want to bring this discovery, this turbocharged engine, to their process.

Shawn Wang

So, okay. Right now, you're still refining that process in the lab.

The Self-Driving Lab

Absolutely.

Shawn Wang

What are some war stories from the lab? Give me your best.

The Self-Driving Lab

That's a good question. I love that question. About the first tools in the lab, I can get in so much trouble for saying it, but I'm pleading the fifth.

Some of the tools don't let us interface with the software. We now pay for that software, by the way, and we love that tool vendor. We were very strategic about how we got access to it, and the engineering team—the software engineering team—was smart about how we could do that. That was a whole 2-week sprint that we had to figure out how we could programmatically control all these tools.

Shawn Wang

So, what's the juice, man? Come on.

The Self-Driving Lab

You can look into the things that are running those tools, and you can find out how you can control what you want to control.

Shawn Wang

Fair enough. Fair enough.

The Self-Driving Lab

So, that's one. My comms team is going to be so mad at me for that one.

Shawn Wang

We can cut it.

Jill S. Becker

No, I'm kidding. I'm kidding.

I think one of the other war stories that we saw early was how interdisciplinary the team needs to be. We knew it was going to be interdisciplinary going in. I think each field we had assigned has splintered into even more fields.

Materials science at large is one we knew would involve computational and experimental work. With mechanical engineering, we have real mechanical engineers who build tools, design tools, and put them together. We have mechatronics engineers who design all our own custom mechatronics to make those tools run autonomously. Obviously, they're in the field of mechanical engineering, but they have completely different jobs.

Software certainly involves standard full-stack work and building the operating system, but then there is more of what we call applied ML. I come from a software background, and I'm applying the systems that we are building, like pulling out images from SEM into the AI scientist. That was interesting.

Robotics—things like path planning and perception—are areas we probably didn't think we would need as much as we do today, simply because we thought, “We'll use what's open source and off the shelf today.” Then we started to realize that what the scientists were doing was very intuition-based, and that's the perfect place where perception and computer vision can be really effective. I do this with PyTorch.

All of these different fields have splintered, and I think we had to build the plane as we flew it. The startup mantra was that we had to continually add people. Ironically, that has now built a huge moat. In inorganic materials science, it is not easy to build self-driving labs. If you sent me back 2 years ago, I'd have said, “That is a tough path to walk.”

Now, of course, it's a big moat for us. We're like, it's not about a robot in front of a tool. Go ahead, put a robotic arm in front of a tool, and watch what happens. Everything else I just talked about will come the second you do that. Now we feel so much farther ahead of the industry in really running self-driving labs for inorganic materials science. That war story is funny—we laugh about it—but now we see it as a huge win.

Shawn Wang

Is there something special about this moment that enabled the self-driving lab versus 5 years ago or 10 years ago? You've been doing this for a while, so why now?

The Self-Driving Lab

A couple of things. One, AI for science is important. If I take force fields—machine-learned interatomic potentials—they're what, 2 or 3 years old, or whatever the exact timeline is? I don't know. I mean, that's interesting. You can actually start to do some parts of your process faster. Although they're still computational, in the simulation sense, with things like DFT, they're still very important to that funnel, to sharpening that funnel and moving faster at the top of the funnel.

Two, robotics are just better. First of all, they're cheaper. Second, you can do more in actuation: custom grippers and different systems that we can either custom-build or acquire.

Three, there's the buy-in from the tool vendors, like I talked about earlier. Again, 2 years ago we saw a difference from what we see today. There is a lot more optionality. I'll give you a perfect example: some of the tool vendors now have software teams. Or, if they had them before, they weren't focused on this problem. Now they actually provide support: “Here's how you can work with the interface.” That's a big change for the industry.

We've started to see it become easier to build self-driving labs from an infrastructure and hardware perspective. Most importantly, I think the biggest change has also been the excitement. Everywhere I go, AI for science is a talked-about area. I was here this week speaking at Unlock—kudos to Michelle and the MSR team—and there was unbelievable energy in the room from so many different builders across so many different companies and fields talking about AI for science.

As I mentioned, even at the national level, one of the things that we did early on was spend a lot of time in D.C.: on the Hill, at the Office of Science and Technology Policy, at the Department of Energy, and at the Department of War. We let them know, “Hey, if you want to be competitive in science, you guys really need to pay attention here. AI for science is a serious field. Self-driving labs are a serious field. There are other people who have already built these systems.”

We really think it is national infrastructure that should be built out, and there's been a big buy-in at the federal level and at the state level as well. I think you have so many tailwinds pushing this industry forward, plus the excitement and the venture dollars coming in that start to solidify some of that.

I think soon we'll start to see some results. We're still waiting on a big result from someone. Hopefully we're one of the first ones there. I think that'll be a cherry on top for the final tipping point, when customers who are coming in with restraint or caution really see this and say, “We can't believe you did that. We've been working on this problem for X many years. We've never had an output like that in that period of time. I'm convinced. I'm a believer. Let's talk about it together.”

I think you're seeing that moment happen already in robotics. It feels like we're on the upswing of that for robotics. I don't know—2 years ago, there was interesting work in robotics if you were a nerd like me and were reading about it in your free time. But I still feel like, when I talk to founders in that area, they're just now getting over the hump. Supply chain and logistics companies, the big warehouse companies, and even the big humanoid companies are now like, “Oh, okay. There are real foundation models that we want to pay attention to. We want to push from 99% to 99.9999% in our foundation model technology.”

I think we're going to have that for science over the next 2 to 3 years. I do. Once the discoveries start coming out and this field continues to mature even more, remember, we're early in this field. We are a couple of years in from an energy, a community, and a new technology perspective.

swyx

So, going to the competitive landscape a bit, we talked about what people are working on in the US and in Europe. But China is both running ahead on materials development and thinking about many of these ideas about labs going all the way through manufacturing, and they do have the expertise in that.

swyx

So, what is your thought about how we stay competitive versus China, and what is the most important thing we need to focus on there?

The Self-Driving Lab

Yeah. This is an incredible question. This is actually what I spend a lot of my time talking about, especially when I'm in D.C. China has an unfair advantage that we do not want to replicate but must figure out how to defend against.

In China, they really will go out of their way—and there's incredible work by NIST and a couple of other groups documenting this—to stand up manufacturing innovation hubs where they make a new material and support, via capital or infrastructure, the scale-up of that material system or invention. They can do that because, honestly, in China, whether you're public or private, one entity owns everything.

Again, that's the part that we should not mimic or copy. In no way am I suggesting that. What I'm saying is that, because they have that, we need to have a similar focus. We need to figure out how to break the 25-year timeline, because when I come in here and tell both of you that materials development timelines are long, you're like, “Yeah, that's all we ever hear about.” That's all we hear about. Everyone I talk to says the same thing. We know the same thing. We need to change that.

How do you do that? I think number 1 is you start to teach the scientists of the future how to run science this way. You know what the most impressive thing from our lab is? I love everything we've talked about today, but what is so impressive is that we can have 1 Ph.D. in metallurgy or alloys run 10 campaigns at a time.

When I was in a Ph.D. program, we had 10 scientists focused on 1 campaign, 1 research problem at a time. That's an order-of-magnitude jump in productivity from 1 scientist. Now imagine every scientist in the United States, every scientist in North America, in the world, doing 10 times the research output. That's fundamental. That just changes the trajectory of discovery.

swyx

But China can do that, too, right?

Yeah, they can. They can. So, I think the second piece is investment. I think we're getting it in the private sector, which is great. I think we'll continue to see the government invest in this area to try to start bridging these gaps and building up this workforce.

You see the Genesis Mission as 1 area where there's a ton of—hundreds of millions of dollars of investment there. I know that groups internal to the national labs are building self-driving labs. My co-founder Herd Seeder has 1 at Berkeley. Argonne has 1. I know that Ames is looking at self-driving labs. Livermore's looking at self-driving labs, or may already have 1. Oak Ridge has a Manufacturing Demonstration Facility, which is almost fully autonomous or semiautonomous in nature.

These labs are now starting to invest in the infrastructure to start to speed that gap up, to start to show you can shorten that gap. That's important. The third one is, I think, maybe where we have to be different: public-private partnership.

This is what we talk a lot about. This is the perfect opportunity for private enterprise to work with public research to create the greatest scientific tool in the world, as the DOE likes to say. Why? Because we have all the HPC that we need—high-performance compute. We have all the researchers that we need. We have some of the best scientists in the world at the national labs, and the infrastructure from a materials-tooling or science-tooling perspective.

Then, lastly, you actually have the data. You actually have more experimental data in all of those national labs than anywhere in the world—or, we think, anywhere in the world—from all the science that we've run. If you can couple all 3 of those things together, and you can bring in private enterprise to help you make sense of that, help you close that system, and help you tie the loop together, then I think the entire national infrastructure runs this way.

The research infrastructure, whether that's corporate R&D or, again, SBIR/STTR R&D, runs this way, and private enterprise runs this way. Now you've changed the fabric of all of R&D in the U.S. Now you're at a place where you can compete with China, not by owning everything and employing forced labor and the unethical practices that they might employ, but rather by changing the mentality and the approach to how we do R&D.

That's how I think we can compete. That's the only way we can compete, if we want to move forward. If we do not do that, then they will continue to win, because they will outpace us on cost and they will outpace us on people. If you try to play that game, it feels like we're going to lose that game.

But if you play the game of changing the system and building a better workforce and a better system to do R&D than they have, then we can beat them with raw output.

swyx

We often ask our guests a question that you've kind of answered now, but I wanted to ask it anyway and see if you have anything more you want to add. If you could remove a bottleneck from the industry, what would that be?

I don't think this is removable, so that's why it's not a good answer. The hardest part about AI for science is that our feedback loops are long.

swyx

Right, that's fundamental.

Yeah, exactly. It's fundamental. That's why I don't know if you can remove it. Maybe there are ways to get it faster, but I'll give a perfect example.

You think about math, like AI and math, right? You can run a lot of experiments in hours that will take us weeks or years to run in science. How do you get around that problem? That's a really hard problem to solve, and I think 1 answer is large-scale automated systems.

A perfect example: if you build a facility with 1,000 XRDs or SEMs, you can certainly build that model that I mentioned that can do image analysis better than anyone else in the world. I think there are paths there to leapfrog the challenge of doing fast experimentation, but it's fundamental, so it's a bad answer to the question.

For something that's not fundamental, I would have the tool providers restart their stack. Their tools are built for humans; they should build them for agents and robots. I feel like this is already happening in software.

swyx

Yeah, I think you see this in CLI and MCP. You're seeing this already happen.

That's right. If you could do that at the infrastructural level for tooling, I think it would supercharge this industry. Because now you don't want to train someone on running an XRD or an SEM; you want to train someone how to run the system that can do that, and that scales so much more effectively.

Now you don't need to get a Ph.D. to analyze your alloys in an SEM. Now you just need to focus on how I can run the system to do that exact analysis for me. That'd be transformational, but it would be quite expensive.

swyx

Okay. Do you have any calls to action for AI for scientists or, let's say, AI engineers?

The Self-Driving Lab

Yeah, they can bring techniques that scientists are not aware of. I'm a perfect example: I am not a machine learning scientist by training at all. I'm a materials scientist by training. I did my graduate work in materials science. My first job out of grad school was materials investing in materials science. Now I run a materials science company. I'm a materials scientist at heart. That's what I do.

One of my other co-founders is a materials scientist. He's an academic professor in materials science. But we have learned so much from the MLEs and the AI research scientists on our team because they're not materials scientists.

They show up to a problem and they don't get stuck on dendritic formation and grain boundaries and this stuff. They're just like, “Why don't you just train a materials model for that?” And it's like, “I don't know what that is.” They're like, “That's not the first thought I have when I look at an SEM,” but they do.

One thing that they can do is supercharge the industry with their own skill set. One thing I don't like about the industry is that I see so many people trying to be the other thing. I meet ML engineers who are like, “I want to be a materials scientist,” and I'm like, “Why? We need ML engineers that work at Radical AI. You should be an ML engineer that helps us do science.”

Then I see this one way more: materials scientists try to be an MLE. “I did a Ph.D. in materials, did a master's in materials, and I got to get into the AI side because that's where the field's going.” No, you just have to learn how to use those tools to make you better at your job.

I don't care who can figure out the problem from SEM. I just want to be able to figure it out. So there's so much cross-disciplinary work that I would highly encourage ML engineers to, number 1, pay attention to AI for science. I think that's already happening, actually. I don't think they need to hear that again, especially on this podcast. I was a huge supporter.

But number 2, lean into your expertise. Bring a first-principles perspective to the way that we do science. We have been doing science the same way for hundreds of years, or 50 years, or however old the tool is that we're running on. We've been doing it for that long.

You can come to that perspective and just totally change the way something works. The field has already done that in the past, and it will continue to do that in the future. I can tell you, 100% guaranteed, that we've done that at Radical AI today.

Specialization. Exactly. Bring the specialization and lean into your expertise. Don't shy away from it. Don't try to become a materials scientist. Be an ML—be an MLE that works in materials science.

swyx

Awesome.

swyx

So, what does your AI stack look like?

Joseph Cross

At a high level, the AI scientist is really a multi-agent approach. There are multiple agents that sit within what the AI scientist is, and we really have this orchestrator agent at the top that comes up with new hypotheses and has a specific way to test those hypotheses internally. That allows us to test whether they’re going to be good hypotheses before we send them to the lab.

But also going into that scientist are a bunch of other models as well, right? We are taking in datasets like industry-standard datasets. We pay for Caltech, and we actually pull that data in so that we can use it the same way I would use it if I were a human scientist. That’s really important.

We have a literature review agent that we’ve custom-built. That benchmark is also public on our website, and it can go and extract figures and information from scientific literature that’s relevant to the hypothesis that we’re making. So, that’s in the stack.

We have custom models built as well. One of them is called MATRIX, which is on our website. I would encourage all the ML engineers out there to go check this out. The model and the benchmark are available on Hugging Face, and you can find a blog post on our website about this.

MATRIX is incredible because it’s really a VLM that we’ve fine-tuned on Quinn. It can go into images from the lab and experimental data and extract scientific knowledge from them. The obvious benefit that we saw in the model, which you can guess, is that it gets really good at reading experimental data, which makes sense. The one that we maybe didn’t see coming was that, by understanding that data, it gets better at being a scientist. This is really cool, right?

For the AI and ML engineers out there, for the AI scientist, that is how you start to capture that. We talk a lot about this intuition, the scientific intuition you get a PhD on. That’s how you start to capture it.

We have actually seen this, and you can go read the publication—it’s on archive—where the public dataset that we used is showing improvements, like 5% to 16%, I believe, on general scientific reasoning.

swyx

Is it adding math to your reasoning?

Yeah, actually, I think math is the one area we call out that doesn’t work.

swyx

No, but the theory is the same, right? Where you add math and then you end up in other domains.

In the paper, we do move outside into the bio space, I believe, and see that same improvement. You can find this all in the preprint that’s out on archive.

What’s cool about this? What’s so cool is that, as you start to build these systems that can do science like a human scientist does, you start to compound the knowledge in the way we talked about earlier. I can be looking at all of these different things as I’m making a hypothesis.

Although a human scientist would like to think we can pull in Cal Fad, pull in literature, extract the right information from literature—not just read a paper, but pull the right stuff out that’s relevant to this hypothesis—and look at all my past experiments, all of the database of our experiments is feeding into that AI scientist.

So, when we make a new campaign, it is looking at the past results to make that campaign. We have a really cool demo that we show customers where, when you look at a generated hypothesis, it’ll actually tell you what experiments it’s pulling in and what it’s using to learn from about why it made that new hypothesis. So, that’s really cool.

Then, again, with these models like MATRIX or MATRIX-PT, you can actually start to pull out intuition, and that intuition actually helps you become a better scientist at large. This is a really important concept. This multi-agent approach is, I think, what an AI scientist really means.

I don’t think it’s ever going to be one scientist. Maybe something will happen in the future. If you’re an MLE, maybe that’s a good problem to tackle. Build that, and we’d love to be a customer of yours. But if not, I think you’re going to have these specialized models, these specialized agents that are really good at one thing, and together, collectively, they make a scientist that’s better than Joseph, better than the scientist that we have today.

That’s kind of what our stack looks like at a high level on the hypothesis-generation and new-materials side.

Alessio Fanelli

And then, just a quick follow-up. I love that you’re open-sourcing a lot of your work. Why are you doing that? Not that I want to discourage it in the least.

Joseph Cross

Yep. This is a really important question. There are 3 reasons.

Number one, I talked about community in this episode. We need the community to move toward doing science this way. Open-source work is one of the best ways to do that, as we’ve seen over history.

Number two, learning. We actually get way more feedback from open-sourcing things than we could possibly work on ourselves. You guys are probably aware of TorchSim, which was this package that we open-sourced. We don’t need to go into that today—we can—but the feedback from the community and the ideas from the community have been incredible.

We’ve actually spun TorchSim out into its own organization, a nonprofit, that can continue to run with the community. We call it Ignite: a materials revolution, a simulation revolution, which we’re excited about. That’s a good example. This idea that you can build better technology with the group is number two.

Number three, we actually don’t think models are the moat. We actually think that in 5 years, most models will be open-source. There’s probably a proprietary model or two, the same way there’s a proprietary model or two today that I can run, whether it’s Claude, ChatGPT, Grok, or whatever your favorite AI is.

However, we think in science models aren’t the moat; experiments are. We actually think the more great models we can share, like MATRIX, which is out there, the better. The dataset isn’t, right? The model that we built and put in the preprint is built on public data. Of course, we have our own proprietary data on top of it. That’s what we train MATRIX on. Then the model can go out there.

I tell people all the time, “What happens if someone else comes out with a better foundation model, better MLIP, or better diffusion model?” I’m like, “That would be amazing. It would supercharge our scientists. We’ll drop it right into our stack.” That’s why we don’t want to sell models. We don’t sell models.

We think the entire community will continue to have ideas that we cannot have alone. That’s why we open-source, because we don’t think that’s the edge. We think the edge is on the experimental side.

That’s specific to Radical, and obviously I’m talking about what Radical believes in—the thesis there—but that is a big part of why we do it. If the whole community can push the whole field forward, we benefit from that, and they benefit from that. That’s a win-win for materials at large. That’s a win for AI for science. And it’s a win because our stack now has a better model that we didn’t have to produce, which is great.

swyx

So, whether it’s us—we do have custom models built internally—or someone else brings that model in and uses it, or we even pay for a proprietary model, which we do on the LLM side, obviously, all of that just goes into making a better scientist.

That’s why we open-source: for the community to get better ideas, and to really understand that we think experiments are the moat, not the model itself.

Alessio Fanelli

Yes, Joseph Cross. Thank you so much for

The Self-Driving Lab

Thanks for having me, guys. Awesome to be here.

swyx

I hope that you enjoyed the show so much. I hope that we can spread the good vibes.

I think we just talked about a great closing topic, which is how important this industry is. The impact that the world will feel from AI for science is enormous. People that can start at the grassroots and actually push that forward, like yourself, are imperative. So, thank you for everything you do.

I’m super happy to be here, and looking forward to coming back when the lab is fully autonomous. We’ll run an episode in the lab.

swyx

Yeah, I was just going to say that. Yes.

Alessio Fanelli

We should definitely do that.

swyx

We’ll come back to New York.

The Self-Driving Lab

Yeah, no problem. Pick a different time than the winter, though. We had a tough winter this year.

Alessio Fanelli

Okay, we’ll do something in the summer.

The Self-Driving Lab

Perfect. Thank you guys for having me. Thank you very much.

🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI | BidClub