[BidClub_]
Latent Space · · 35 分钟

🔬材料领域没有 AlphaFold——与 Heather Kulik 探讨材料发现 AI

Brandon AndersonRJ HonickyHeather Kulik

YouTube
TL;DR
  • 材料 AI 没有 AlphaFold 式的捷径,因为材料涉及更多种类的构建单元、变化高度复杂的成键方式,以及稀疏的实验真值。 Kulik 将 AlphaFold 在球状蛋白质上的成功与材料领域进行对比:前者主要使用20种天然氨基酸,而材料领域当前的势函数“在整个化学空间内肯定不正确”,且在缺乏明确实验验证时,失效后果可能更严重。
  • Kulik 介绍了一项清晰的 AI 赋能发现:一种聚合物网络通过出乎实验人员预料的设计,使韧性提高了约4倍,并在实验室中得到验证。 AI 从数千到数万种候选材料中进行搜索,而每种材料的单独实验可能耗时数月到数年,最终发现了一个“完全由量子力学驱动的现象”:电子重新排布,使分子组件在开始断裂时获得稳定。
  • 当材料必须同时满足多项约束时,主动学习尤其有价值。 Kulik 的直接空气捕集项目同时优化7个目标,包括成本、湿度稳定性、CO2选择性,以及机械和热稳定性;即使模型并不完美,也能让每个优化维度“至少提速100至1000倍”。
  • 认为神经网络势函数已经取代基于物理的模拟,仍然超出了现有性能证据。 某个一度引发巨大轰动的未具名模型,速度仅比 Kulik 最快的 GPU DFT 计算快约5倍,而且“并非总能工作”。她认为真正具有变革意义的门槛,是以大约快2个数量级的速度,可靠取代 DFT。
  • 通用 LLM 可以补充化学知识,但仍需要专家充当错误检测器。 ChatGPT “在维基百科级别的化学知识上非常出色”,但面对 Kulik 提出的简单要求——设计一个恰好含22个原子、并通过2个氮原子配位的配体——却反复出错。实际操作原则是“把化学学到足够好,知道这些模型什么时候对、什么时候错”。
  • 除了模型规模,实验数据、验证和制造工艺也是主要瓶颈。 从文献中提取的标签会因来源是图表还是作者解读而相互冲突;自动化实验室在处理人类觉得容易的实验时反而可能表现不佳;而材料在器件尺度上的性能取决于加工过程——Kulik 说,“我们还在起点,几乎没有进展。”
  • 算力充裕的公司正在改变学术界选择问题的方式。 Kulik 将学术资源与 Microsoft、Meta“基本无限的资源”进行对比,同时指出,被忽视的化学问题、更扎实的证据、创造性的问题选择、共享云实验室,以及机器可读的实验报告,都是学术界可以发力的方向。
摘要 · 为研究而整理的核心内容

1. AI 找到了化学家未曾预料的聚合物设计

  • Kulik 介绍了一个明确的验证案例:团队从数千到数万种材料中进行筛选,而每种材料的单独实验可能耗时数月到数年。最终选出的聚合物网络设计让实验合作方感到意外,但测试证实,它使材料韧性提高了约4倍。

  • 这一设计使用了会以特殊方式断裂的分子组件,使整体结构变得更坚韧。AI 找到的是一个“完全由量子力学驱动的现象”:电子在分子开始断裂的瞬间重新排布,从而稳定分子;这种效应类似催化或酶促化学,但此前从未在这些聚合物中得到展示。

  • 主持人将这一机制比作一根受控引线,在地震中保护海湾大桥。Kulik 的澄清很关键:牺牲性断裂本身几年前已发表于 Science;她的团队贡献在于找到一种具体的、由电子机制驱动的设计,并由此实现了这一效果。

2. 主动学习在7重权衡中体现价值

  • Kulik 在2000年代中期进入数据驱动化学领域,起因是她厌倦了“一次研究一个分子”。到2015–2016年左右,她不再把这项工作称为 cheminformatics,而改称机器学习;学生 Jean-Paul Janet 则把最初的一次设计讨论发展成了课程作业中的神经网络项目。

  • 这项聚合物搜索“原则上”属于主动学习,但团队在完成一代迭代后就停止了,因为候选空间已经耗尽。更大的价值在于,即使模型尚未准确,也可以提前启动优化,尤其当目标是从“稻草堆里找针”式的多维问题中筛出候选时。

  • 她目前围绕金属有机框架开展直接 CO2 捕集研究,同时权衡7项目标:成本、在潮湿或水相条件下的稳定性、CO2选择性、受力时的机械稳定性、热稳定性,以及其他约束。Kulik 说,即使模型并不十分准确,也能让每个优化维度“至少提速100至1000倍”。

  • MOF 是一种模块化材料,像“Tinker Toys 或 Lego”,应用于储气、传感、分离、聚合物复合材料、CO2捕集、催化,甚至药物递送。其精确的化学基团可以像一只分子“手套”一样包住目标客体分子,但稳定性仍是核心限制。

3. 化学专业知识仍是错误检测器

  • 过渡金属含有未配对电子,因此反应性更强,也更适合催化,包括 Haber–Bosch 合成氨等过程。传统上,人们通过薛定谔方程的各种近似来模拟其行为,但一次量子力学预测可能需要数小时、数天甚至数周。

  • 机器学习可以加速这些计算,并帮助选择合适的近似方法;这种选择过于复杂,无法靠简单启发式规则完成。Kulik 的团队将量子力学波函数中的信息输入神经网络,学习特定材料需要采用哪种计算方法。

  • Shawn Wang 曾提出一个挑衅性问题:如果 ChatGPT 已经具备博士级理解,为什么还要学习化学?Kulik 随后提出了22原子配体测试。即使明确告知配体必须通过2个氮原子配位,LLM 仍反复数错原子,而这件事专家化学家“1秒就能完成”。这些模型适合作为导师和知识增强工具,但不应成为盲目信赖的起点。

4. 材料既没有 AlphaFold 那样边界清晰的空间,也没有同等的真值

  • 数据稀缺或高度多样的领域,仍有大量机器学习机会,包括反应预测、过渡金属成键、温稠密物质,以及通过光照材料产生的激发态。相比之下,“非常无聊的化学”——有机分子和蛋白质结合——已经拥有成熟的数据集、基准和排行榜。

  • Materials Project 和 Open Catalyst Project 的排行榜主要建立在低保真度 DFT 计算之上,而不是实验真值。与 CASP 不同,它们没有提供可比的实验真值,因此“所有最聪明的 ML 工程师”可能都在使用无法反映实验室行为的数据进行训练。

  • Kulik 说,每一个新的“基础势函数”在被她的团队用于真实问题之前,都可能看起来表现出色;但一旦应用,分子就直接解体。那个夏天引发巨大轰动的某个未具名模型,速度仅比她最快的 GPU DFT 计算快约5倍,而且无法稳定工作;如果能以快2个数量级的速度可靠取代 DFT,才会真正改变这门科学。

  • Shawn Wang 将 AlphaFold 概括为解决基态结构问题;Kulik 则强调,它处理的是球状蛋白质,主要由20种天然氨基酸组成,而材料范围从相对简单的铝,到氧化铁、金属有机成键和高熵合金。在更大的长度和时间尺度上,实验数据的缺失或解释性会让人“没有真正的办法知道自己对不对”。Kulik 还指出,AlphaFold 也会失败,而材料模型的失败后果可能更严重。

5. 瓶颈从计算转向证据与工艺

  • 自动化实验室有些实验做得比人类更好,却在另一些人类觉得容易的实验上表现不佳;研究人员也在探索如何引入人类可能经历的偶然性或噪声。对于把材料推向器件尺度的人来说,关键“不只是材料,还有工艺”,但 Kulik 说,机器学习几乎还没有处理加工过程的影响。

  • 文献提取还带来一个更隐蔽的数据质量问题。Kulik 的团队可以从实验图表中提取某种 MOF 的分解温度,也可以依据作者的文字解读提取,但“这两者并不一致”;现代 LLM 的提取结果仍容易出现假阳性,因此验证本身也成了一项额外开销。

  • Kulik 提出了一个开放问题:用一个领域最初30年的数据训练的模型,能否预测接下来20年的发现?不确定性量化或许可以帮助识别最值得补充的材料;共享云实验室,以及标准化、适合机器学习读取的报告格式,则可能让实验结果从论文发表当天起就具备复用价值。

  • Kulik 将学术界的资源与 Microsoft、Meta“基本无限的资源”进行对比,如今会避开那些主要靠堆算力就能解决的问题。她的团队通过 molSimplify 和 MOFSimplify 提供过渡金属与 MOF 设计工具,覆盖网站、Conda 和 GitHub,并支持机器学习预测和新结构生成;团队也在收集用户反馈,其中包括已经在使用这些工具的公司。

Heather Kulik

There's a school of thought that says, "Why should I bother to learn chemistry or physics or whatever when ChatGPT has PhD-level understanding of that anyway?" ChatGPT is super good at Wikipedia-level chemistry knowledge.

I'm really interested in molecular design. How do you find a new ligand that can go into a transition-metal complex? What that means is that some combination of atoms is going to bind to the metal and change its properties. The thing I constantly do every time an LLM is updated is ask it, "Please design me a ligand that has 22 atoms." I can never get an answer that has 22 atoms.

swyx

All right, we're really excited to have Heather Kulik here. She's a professor of chemical engineering at MIT. Heather has done some amazing work in materials science and computational chemistry. But we're particularly excited to have her today because she has, for almost her entire career, been working at the intersection of data-driven methods and AI, applying them to improve materials and our understanding of materials. She has a lot of interesting opinions about what works and how to approach these problems to get the most out of them.

We're really excited to have you here. To get started, can you tell us about one of the coolest things you've done, in your opinion, for our general AI engineering audience?

1. AI Discovers Tougher Plastics

My group works a lot on accelerating the discovery of new materials. When I first started out, we were really just using AI to make predictions that we would normally make with computational models, just faster.

But the question I would often get when we were doing that was, "Okay, but what's surprising? What's something from AI that I wouldn't have already known if I were a really smart chemist or a really smart materials scientist?" You make all these computational predictions, but has anyone actually made something in the lab that you predicted?

Recently, I was able to do a really nice demonstration where the answer to both of those questions was very clear from the work. We were able to screen with artificial intelligence a set of thousands, tens of thousands, of materials where each individual experiment, if it were done in the lab, would have taken months to years.

Through AI, we uncovered an unexpected chemical phenomenon that led to an emergent property in what's known as a polymer network—plastics—that would make the polymer about 4 times tougher. When we showed the design that AI had come up with to the experimentalists, they were really surprised. They would never have come up with it on their own. Then we were able to convince them to make it in the lab, and in fact, it was a tougher material.

Where this has applications is that if we can make plastics tougher, we can get more use out of them, and it will ultimately address some of the problems we have with the overall durability and use of plastics. I think that's an example of some of the promise of AI in materials discovery.

swyx

Can you dig into that a little bit? What was the surprising chemical discovery there?

It's sort of hard for me to explain it without getting too deep into the chemistry. Basically, these are molecules that have to break apart, and when they break apart, they make the overall structure that they're in tougher.

Normally, the way you would think about making it easier to break apart these small molecular components might be to create a hinge, so they can peel open instead of sliding apart. But what we discovered was a fully quantum-mechanical phenomenon. There was really no way for us to predict this based on anything else, where the electrons just move around in a different way so that, at this moment when the molecule is going to break apart, it's a lot more stabilized.

These types of concepts are sort of similar to what's known about how catalysts and enzymes work, but it had never before been shown in these polymer materials.

swyx

This is sort of like the fuse in the Bay Bridge that allows the bridge to keep its structural integrity during an earthquake by having a controlled break. Is that kind of—

Yeah, yeah. We weren't the first ones to discover that phenomenon on its own. The general phenomenon of putting little places that could break to make the network stronger was published in Science a couple of years ago. But the specific way we came up with to design the material to do this was our new contribution.

swyx

You mentioned that you started off accelerating existing methods using enhanced computation. What caused you to take that leap to more machine-learning-based methods?

2. Active Learning Hunts Tradeoffs

I was drawn to data-driven discovery pretty early on, before I even knew the phrase "machine learning." I was really excited by what you could learn from patterns in data. Back then, we were trying to call it cheminformatics and just trying to think about in what ways you could unearth trends in data.

I started my career working one molecule at a time, or one material at a time, and I was impatient. I wanted to understand not just one molecule at a time and write one paper about it—which is something people would have been happy to do when I was starting my career in the mid-2000s—but actually unearth broader trends in how you understand how a material is going to behave.

Somewhere around 2015 or 2016, I realized it was a bad idea to call things cheminformatics and a good idea to start calling things machine learning. I had a brilliant student, Jean-Paul Janet, who's now, I think, an assistant director at AstraZeneca in Sweden, running their inverse-design program. He and I originally talked about all sorts of ways of thinking about materials design, and he very quickly adapted that into training neural networks.

I thought we were in the first hype cycle, the first wave, but compared to what's going on right now, it was a tiny baby wave.

swyx

I read in your paper that this was actually a class project or something.

Yeah, that's right. He just said, "I have to do something for my homework," and that's how we got into it.

swyx

I've also read in your paper that you've done a lot of work more recently on active learning. Can you talk a little bit about that?

Even that polymer example I was giving would have been active learning in principle, but we stopped after 1 generation because we had exhausted the space.

I think one of the areas where machine learning, with what's out there right now, has the most promise in the chemical sciences is solving multidimensional challenges. Right now, we're working on a project in metal-organic frameworks where we're trying to solve trade-offs relevant to the direct capture of CO2 from the air.

In order to find a material that's good for that, we would worry about its cost, its stability in aqueous, humid environments, its ability to take in CO2 over other molecules, its mechanical stability—whether it's going to hold up under force—and its thermal stability—whether you can heat it up and it will be okay. I'm just naming a few, but in total, right now, in an active-learning campaign, we're working on 7 different objectives.

Usually, even for a not-so-accurate machine-learning model, you get at least a 100- to 1,000-fold speedup for every dimension you're optimizing over. The real promise is going to be in searching for that needle in a haystack with, say, 7 objectives, and doing something where you're not waiting for the models to be accurate before you start doing that optimization. That's really the promise of active learning.

swyx

That has an interesting parallel in my mind to the pharma world, where you have a lot of computational work in the discovery process, but actually getting the drug into people's hands is often the bottleneck for a drug. Also, what happens to this drug when it sits on the shelf for 3 months? That kind of thing.

Are these metal-organic frameworks? What kinds of things are they used for?

They're used mostly in gas storage, sensing, and separations. They're used in combination with polymer composites. They have really strong promise for CO2 capture, especially, but people have looked at them for catalysis.

The limitation on catalysis has been how stable they are. So, one of the things we spend a lot of time on is trying to predict their stability. But they're used for all sorts of things, even drug delivery.

What they have the opportunity to do is place precise chemical groups in specific orientations that can ultimately allow for what's known as a host-guest interaction. Basically, they can create a glove to have a targeted interaction with a guest molecule in the metal-organic framework.

swyx

I see. For the non-chemists, metal-organic frameworks are like Legos for chemistry?

Yeah. Metal-organic frameworks are going to be a little bit more of a household name among some engineers because the discoverers of those materials just won the Nobel Prize in Chemistry this year. As much as that can make something in chemistry a household name, they're basically like Tinker Toys or Legos, and they have different building blocks that can be combined in basically infinite ways to create very precise chemistry.

swyx

I see. Maybe for context, could we say something like: What are the techniques you were using before you started—or maybe in parallel with—machine learning, and how does machine learning help you advance those? What are the roles of the two?

3. Quantum Chemistry Meets Machine Learning

So I started my career studying what's known as transition-metal catalysis. If you look at the periodic table, the middle of it contains a bunch of metals. A good example would be iron, and all of those things sitting in the middle of the periodic table have what's referred to as an open shell. The electrons in those materials are not paired, and as a result, they're more reactive.

Normally, the way that you understand how they're going to behave is through the different combinations of these metals that give rise to the catalysts used in a large number of transformations, including things that feed and sustain most of the world's population, such as the Haber–Bosch process for ammonia synthesis. Going back 20, 30, 50 years, the way that people understood these materials and could enable rational design was through quantum mechanical modeling.

Heather Kulik

Quantum mechanical modeling, by using approximations to the Schrödinger equation, can be very accurate, but it's very computationally costly. And so a single quantum mechanical prediction, depending on the level of fidelity used, could take hours to days to weeks. And that's what I would have normally been doing before I got started in AI.

Some of what we do these days is accelerate those quantum mechanical predictions, as well as look at an area that I'm particularly excited about: not all quantum mechanical approximations are equal, and you can actually use ML models to predict what the best approximation to use is, depending on the material studied.

Shawn Wang

Is that like, “closer is better,” or is it not really distance-related in terms of which method is the right method to use?

Heather Kulik

It actually turns out to be quite complex. You can't just determine it from heuristics. So we actually, in one area, use the quantum mechanical wave function as inputs to neural networks to predict what the right method to use is and learn that mapping.

Shawn Wang

I see. I have a spicy question I want to ask. There's a school of thought that says, “Why should I bother to learn chemistry or physics or whatever when ChatGPT now has PhD-level understanding of that anyway, and shouldn't I just focus on being really good at using AI for stuff?” So I want to hear your thoughts.

4. LLMs Hit Chemistry Limits

Heather Kulik

My personal experience is that—and this will date itself immediately—ChatGPT is super good at Wikipedia-level chemistry knowledge. But one of my favorite things to actually throw at ChatGPT, as an anecdote, is that I'm really interested in molecular design. How do you find a new ligand that can go into a transition-metal complex? What that means is that it's some combination of atoms, and it's going to bind to the metal and change its properties.

The thing I constantly do every time an LLM is updated is ask it, “Please design me a ligand that has 22 atoms.” The first time I did that, there were many ligands out there that had 22 atoms. And then I say, “I want it to bind to the metal with 2 nitrogen atoms.” I can never get an answer that has 22 atoms. So then you can try a range and see how many times you can get that.

That's maybe a trivial thing, but it's something that an expert chemist could do in a second. There are really good introductions to chemistry that I think you can get through conversations with an LLM. You can get a lot of insight into an area you're unfamiliar with, and for sure, things have improved a lot.

When I first tried typing in, “Which exchange-correlation functional should I use for this type of chemistry?” the answers were completely wrong. They looked right, but they were completely wrong. I think things have gotten better because that knowledge is out there on the internet and it's in the training data, but I think there's a lot of things that—backing up a moment—you should learn chemistry well enough to know when these models are right or wrong.

If you don't know any chemistry at all, it's hard to know if you're assessing it correctly. But I think there are a lot of things that you don't have time to do a deep dive into that you can now get from, say, an LLM that can augment knowledge. I think you have to start from somewhere and then use it as a tool, rather than starting from zero and relying blindly on what an LLM will say.

But one of my favorite things is, if someone can get an LLM to generate me a 22-atom ligand in one shot, I would love to see it.

Shawn Wang

[Laughter.] Sounds like a challenge. Cool, the 22-atom ligand challenge—go. What do you think the biggest gaps are in machine learning, from your experience? If you were an aspiring ML engineer looking to take on a new problem from the machine-learning side, what do you think someone could work on that would really help the chemistry side?

Heather Kulik

There are a lot of challenges out there where the data sets aren't large enough or diverse enough, and so I think they've attracted less interest. The ones closest to my heart are reactivity predictions: predicting which reactions will occur and why, especially in complex phenomena involving multiple elements and predicting those transformations.

Another thing that I think there's not enough data on is more diverse chemical bonding and more diverse chemistry. For me, that's transition metals, but there are also questions of warm dense matter and exotic phenomena. We have really good data sets out there for really boring chemistry. We have, probably—even if you're not a chemist, you're familiar with organic molecule data sets and organic molecules binding to proteins. Those are the common data sets out there.

There are lots of challenges out there where the physics is much more complex, including things like how matter behaves when you shine light on it and excite it into excited states. All sorts of things like that receive relatively little attention because there may not be a benchmark or leaderboard yet for them. Maybe it's on us chemists to generate more data sets so those leaderboards are out there, but there's definitely a lot of interesting chemistry for which there has been less attention.

Shawn Wang

So in the protein world, there's CASP, right? People have been working on that for a while, and this led to AlphaFold. Without CASP, AlphaFold probably wouldn't exist. Is there an equivalent to CASP in materials science?

Heather Kulik

There are all sorts of repositories of fairly low-fidelity DFT data on crystalline materials, such as the Materials Project and the Open Catalyst Project. These do provide good leaderboards, but some of the limitations are that the data comes from not-very-high-fidelity density functional theory.

I'd say that's a second challenge: all the smartest ML engineers right now are learning on data that is not going to be reflective of experiment. There aren't big experimental data sets. For example, one of the advantages of things like CASP is that it comes from an experimental ground truth, whereas that aspect just isn't available in materials as much.

Shawn Wang

We talked about CASP and the role of CASP and AlphaFold. Do you think there's a way of phrasing this problem so that we could start collecting data at scale, so that we could really have a community challenge that breaks open a problem you have in mind? And maybe, actually stepping back beyond that, what would you want to have if there was an AlphaFold for materials? What would you want it to do?

5. Materials Resist AlphaFold

Heather Kulik

One kind of murky area—so maybe I'm not going to directly answer this question—is that electronic structure calculations are expensive, and they should in principle give you the right answer. They should, from first principles, give you the right answer of how a material is going to behave.

A lot of people are scaling these up right now with machine-learned interatomic potentials on training data. And every time someone comes out with a new data set trained on it—and they call it a foundation potential or foundation model—it looks really good. Then you get it into your lab and say, “Okay, I want to use it for this problem I'm really excited about,” and it starts doing wacky things, like molecules falling apart.

I won't name names, but there was one that made a huge splash this summer, and people started declaring, “Oh, this method is dead. This method is dead. We're all going to just use these neural network models now.” The one I'm still not naming is, in my hands, only about 5 times faster than my fastest DFT calculation on a GPU, and it also doesn't work all the time.

So I would say we need a more transparent way of trying to figure out if these models can really replace conventional physics-based modeling. If I could just give up ever doing a DFT calculation again and rely on machine-learned potentials, and if they were 2 orders of magnitude faster than the traditional approach, that would change how we're doing science.

But there needs to be a little more rigor on what we consider just fitting data when that data maybe lacks quality. There needs to be a little bit tougher a requirement for how we say this model can really replace physics-based modeling.

Shawn Wang

Yeah, so one of our theses is that the interface between bits and atoms is really the bottleneck, right? Actually trying things in the lab is the bottleneck, and you've addressed that to some extent in active learning.

But I think there's also an extent to which pure process and automation, and good operational practice, are important things. If you can push the automation on one side, on the other side that creates brittleness. How do you think about bridging that gap to experimental chemistry and using that sort of thing as nature's computer to figure out things for your design process?

6. Automation Meets Experimental Reality

Heather Kulik

Yeah, there are a lot of really smart people working in high-throughput synthesis and experimentation and autonomous labs. I think the thing that's interesting to me in that space, at least as of the last conference I went to on this, is that there are some types of experiments that are really hard for autonomous, high-throughput experimentation but really easy for a human, and vice versa. Then there's the serendipity a human might experience in the lab. A couple of people have tried to think about that, like, “How do you introduce that noise into high-throughput experimentation?” So I think that's a challenge.

Your question also brought to mind another point that I am by no means an expert on. Most people who actually work on getting materials to the device scale—say, something that would be in your television or something like that—will tell you that it's not just the material, it's the process. I think we're at ground zero. We're nowhere when it comes to, “How do we machine-learn not just the structure and the properties, but also the role that processing plays?” I don't think we know anything about how to do that.

Shawn Wang

Maybe for non-experts, with protein structure, it's really easy to imagine, “Oh, you can see these proteins, and we can run some simulations and see them wiggling around, and the structures look really pretty.” What does the data look like for materials science? Is there computation, like DFT, that gives you something that looks like a crystal structure you can imagine? Is there also experimental data where you can observe that crystal structure? Or is this mostly a sort of probing where you're measuring individual properties, which are kind of collective rather than fine-grained?

Heather J. Kulik

Experimental structures are available, and the example I was giving is something we know is stable, we've seen a structure of it before, and it will fall apart with some of these models. The challenge here is that what AlphaFold has done really well is predict structures of globular proteins, primarily with 20 natural amino acids. I can actually point to lots of cases where AlphaFold fails, too, for more interesting chemistry. The challenge is that you have a lot more than 20 building blocks when it comes to materials.

There are lots of different ways to think about chemical bonding, and right now no potentials are really robustly encoding all of that bonding, especially with respect to metal-organic bonding.

Shawn Wang

Yeah, maybe a different way of saying it is that AlphaFold is solving ground-state structures. It's not looking at dynamics, which I think is consistent with some of your statements about needing quantum mechanics for catalytic enzymes. But even at just ground-state properties, you're saying there are too many parameters, and there's not a clear set of interactions that's limited to a small number of building blocks?

Heather J. Kulik

The bonding is highly variable across all of material space. There are simple regions of material space. You can pick aluminum. Aluminum's very boring, and people in the '60s could write down on paper how you need to model aluminum. That's something that's pretty easy to fit a neural network potential to.

But if you want to get over to iron oxide, and then if you want to get over to high-entropy alloys, there are definitely cases where people are using these methods. I'd say a big challenge is that there's no real way to know if, when you go to bigger length scales and time scales, you're right or wrong. The experimental data is not there. Even interpreting, say, an image of an experimental surface—which you would want to do—requires some degree of interpretation.

So it's just hard to know from experiment or from other computations if these types of models are correct. They're certainly not correct across all of chemical space. I'd say they could fail more catastrophically than AlphaFold. Obviously, AlphaFold fails, too, though there are definitely failures of AlphaFold.

Shawn Wang

Switching gears a little bit, I also read in your paper that you had done some work integrating textual information from papers into your models. It's kind of the AI that we all know and love right now. Can you talk about what kind of lift that gives the models, and how did you actually do that integration?

Heather J. Kulik

Yeah, we started, I guess, about 5 years ago. When we first started doing it, we were just doing standard natural-language processing and graph digitization. These days, we use LLMs, but the goal is just to try to extract data sets of properties from the literature wherever people are widely reporting properties.

What we noticed is that there's a lot you can learn from these models. Even on the scale of a few thousand data points, you can do things like predict the temperature at which a MOF will break apart based on experimental reports. One of the funniest things I think we noticed is that you can get the temperature at which a material will break down in 2 ways: 1, you can get it from the graph, and 2, you can get it from what the authors say about how they interpret the graph.

Those 2 things do not line up. One of the challenges with literature extraction from papers would be the obvious mistakes people make—no one's perfect. But the other would be that people interpret their results in different ways. If we're building models based on those interpretations, that's a challenge.

In terms of LLMs, they've come a long way in terms of literature extraction, but they're still definitely sensitive to false positives. The amount of time we spend checking LLMs to make sure that the data we're ingesting is accurate definitely is an overhead on those types of workflows.

Shawn Wang

I see. And what about the way that it might bias the discovery process? You have this known literature, and your job as a chemist is kind of to find new stuff. But if your computational method is pulling in literature, then maybe it's biasing you toward the previously reported results instead of something new.

Heather J. Kulik

One of the ways we try to address that is to train a model on the literature, but then apply it to new structures that have never been seen before and really look at how far we can extend the model. We are trying to answer this in general. There are repositories out there of experimental data where you can have a sense of when it was published, what the structure is, and what it was used for.

We're really trying to build generative models on top of that now to be able to say, “Well, if I know about the first 30 years of a field, can a model trained on that predict the next 20?” I think that's an open question: what model is best at that? Maybe it won't get all of them, but maybe some of those discoveries that we think are new in the most recent 20 years are trivial for a model to generalize to, whereas others are not.

I think in an ideal case where we have the available literature data, we could use uncertainty quantification to identify the most interesting materials to add to our data set.

Shawn Wang

I see. And those data sets, just for people who are interested in getting involved, what are the best ones to get started with?

Heather J. Kulik

I don't know about the best. We've curated a few thousand data points on metal-organic framework thermal stability, as well as metal-organic framework activation stability and water stability. Other groups have curated other measures of stability. They're all out there. They're on our website, that kind of thing.

Shawn Wang

Awesome. Do you imagine it would be useful to create an initiative or a multi-institutional funding source or something that is really trying to get data in a high-throughput, automated way? What would your dream be if you could organize something that really drove the field forward?

Heather J. Kulik

I think the National Science Foundation has 1 initiative. I've also heard about foundations being interested in putting together cloud labs—things that users can, on demand, make use of high-throughput automation. I definitely think having user facilities where a computational researcher like me could design an experiment and have it executed would be awesome. Having all that data collected in a sort of public way would be great.

The way that research right now gets published into papers makes it very hard to then extract back out. We spend a lot of energy trying to get it back out. Some of this is also a need for systematization of how results get reported, so that they can be machine-learning-ready from day 1 when they're published. Some research subfields are trying to do that, but it's not really developed across materials science.

For sure, I think there will be more shared facilities where people can make use of data from high-throughput experimentation, and that would be really awesome.

I don't know if it'll come from companies donating equipment, from the National Science Foundation, or from private foundations.

Alessio Fanelli

Yeah, there is a large philanthropic push in the biotech space. It seems like people haven't quite picked up on this as such an important field, especially with things like materials for climate change. You can imagine a very important problem that we could use a lot of push on. That kind of brings up the question: There's been a ton of very recent materials investment from private companies and startups. Where does that leave, in your mind, the role of the academic in chemistry?

7. Academia Finds Its Edge

Heather J. Kulik

I ask myself that all the time, or at least more recently, in the past year. In particular, there's a lot of compute that companies have access to that academics don't. So I ask myself: What can we do that's more creative and doesn't require just brute-force compute? I think there is a lot of stuff that we can still do, but we have to ask those questions.

For sure, Microsoft and Meta are the companies that have basically infinite resources, and as an academic, I don't have infinite resources. But we have an interest in problems that haven't crossed the radar of those companies yet. Whenever someone poses a problem to me now, versus a few years ago, I try to make sure that we're not just in the process of trying to do something that, if you threw a lot of compute at it, would solve it.

Alessio Fanelli

I think we're kind of running out of time, but I'd like to give you an opportunity for a call to action. What would you like our listeners to know about or do? What should they do to get involved, or is there something that you're really passionate about?

Heather J. Kulik

I think I will stick to something kind of niche.

Alessio Fanelli

Great.

Heather J. Kulik

So I think there is still a place for chemistry. I will say that. My group develops code for transition-metal complex structure generation and metal-organic framework screening. It's called molSimplify. When we're working on MOFs, we call it MOFSimplify.

There are website versions of it that you can look up without installing anything, but it's also on Conda and GitHub. If you do have an interest in transition-metal complexes, just try it out. It includes machine-learning predictions, but it also makes novel structures. I'm just really interested to hear if people are ever using it. I know a lot of companies are using it, but we sort of find out after the fact.

If you're more interested in this materials space, I'm definitely interested and open to feedback.

Alessio Fanelli

Great. So awesome—getting involved. Thank you very much. Take care, Doctor.