[BidClub_]
The Cognitive Revolution · · 71 分钟

Google 的 AI Co-Scientist “发现”了什么?来自 Podovirus 播客的人类科学家视角

Jessica SacherJoe CampbellJosé PenadésTiago Costa

YouTube
TL;DR
  • Google 的 AI Co-Scientist 独立提出了 Imperial College 科学家花数年验证的核心假设。 在只有研究问题和公开文献、没有研究团队未公开数据的情况下,它提出:形成衣壳的 PICI 会在细胞外把已封装自身 DNA 的衣壳接到不同的游离噬菌体尾部上,从而实现传播。开场旁白估算,这次推理成本仅为 100–1,000 美元,构成了一次成本极低、但极具冲击力的假设生成演示。
  • 这个系统的优势在于摆脱了整个领域的共同假设,而不是具备完整的生物学理解。 噬菌体研究者默认感染后释放的物质已经是完整的感染性颗粒;Co-Scientist 却做出了一个简单连接:“如果尾部决定宿主嗜性,也许是因为它在使用不同的尾部。”但它从未还原释放机制;当研究者追问这一细节时,据报道系统只回复了一句:“糟了。”
  • 这一得到实验支持的机制,可能发展成覆盖范围超过传统噬菌体的可编程递送平台。 CF-PICIs 先释放无尾衣壳,再利用 adaptor 和 connector 蛋白接纳尾部;尾尖决定目标细菌,替换这些蛋白则会改变尾部兼容性。Tiago Costa 认为,这可能形成一套面向不同菌株和物种定制治疗或诊断方案的工具箱,解决噬菌体疗法通常宿主范围过窄的问题,但完整的结构机制仍未厘清。
  • 真正可信的效率提升,不是自动化科学,而是减少走入死胡同的实验。 Costa 最初表示实验本身不会改变,随后澄清:如果早期就得到正确假设,失败实验可能从“90%”降至“50%”,从而节省半年或一年。“科学方法完全不会改变”:人类仍须评估假设、设计实验、解读结果,并拒绝那些看似诱人但错误的解释。
  • 在这个问题上,Co-Scientist 的方法框架明显胜过被测试的通用 LLM。 Gemini、ChatGPT 和其他系统给出了看似合理的总结或相邻想法,但没有一个识别出同一套不同尾部机制——即便其中一个似乎找到了研究团队的实验预印本。Co-Scientist 则围绕获胜的尾部假设返回 5 个排序后的、可以直接测试的假设、相关研究者以及约 15 篇相关论文,带来一种“正在和该领域专家交谈”的感觉。
  • 这是一场干净的验证,但不能证明该系统能持续发现全新的机制。 Penadés 认为它当前的强项是“以不带偏见的方式把点连起来”,而不是凭空发明完全全新的东西;这次的点异常完整,包括数十年前将分离的衣壳和尾部突变体混合的实验。第二次针对质粒移动性的反向测试则生成了复杂但错误的假设,因为文献中占主导的“自私元件”框架压过了团队更新、也更不强调自私性的解释。
  • 短期内能否使用以及是否具备投资价值,仍取决于 Google 的开发流程。 系统当时尚未向公众开放,外部实验室被引导加入可信测试者计划,Google 则继续测试其跨科学领域的稳健性;Imperial 团队没有获得资金,也没有拿到股票期权。The Cognitive Revolution 的旁白将其描述为 Gemini 2.0 时代的能力,并推测未来把 Gemini 2.5 接入这套架构、再接入 Cloud Lab API,可能让系统从提出实验延伸到指挥实验,但这仍只是旁白的前瞻性设想。
摘要 · 为研究而整理的核心内容

1. CF-PICIs 的宿主范围之谜持续了 10 多年

  • Penadés 研究噬菌体诱导型染色体岛已有 20–25 年,其中包括一场持续多年的争论:PICI 应被定义为噬菌体卫星,而不是缺陷噬菌体。大约在 2010 年,他的团队发现了形成衣壳的 PICI:这类元件编码自身的衣壳和 DNA 包装 machinery,但仍需借助噬菌体尾部才能具备感染性。

  • 规模差异具有功能意义。典型 PICI 基因组约为 10–15 kilobases,而辅助噬菌体约为 45 kilobases,因此 CF-PICI 会构建更小的衣壳,只容纳自身 DNA,排除更大的噬菌体基因组;Costa 的结构研究认为,蛋白质序列插入改变了衣壳对称性。

  • 真正的谜题在于分布:Penadés 提到,曾在 5 个属的 7 个物种中发现过一个完全相同的 CF-PICI。如果噬菌体尾部通常决定狭窄的宿主嗜性,那么同一个元件如何跨越这些相距甚远的细菌宿主?

  • 这种关系未必是寄生关系。原型 CF-PICI 并未明显损害辅助噬菌体的繁殖,而一些 PICI 携带抗噬菌体系统,在被诱导时会让噬菌体付出代价。Penadés 如今认为双方可能具有协同关系——“可能比我们 15 或 20 年前以为的更像朋友”。

2. 一个 70 年未解的接合问题,带来了反向验证测试

  • Imperial College 的 Fleming Initiative 在 Costa 的实验室与 Google 之间牵线,当时 Google 正在开发一套面向科学家的 LLM 系统。Costa 最初询问启动接合的分子“T=0”,他说这个问题已经悬而未决 70 年;Co-Scientist 返回了 5 个假设,其中排名靠前的几个让他印象深刻,但验证需要数月。

  • Penadés 看到了一个反转 Google 原定工作流的机会。团队没有从 AI 假设出发、等待实验验证,而是提出了一个自己已经通过实验确定答案的问题:CF-PICI 如何跨越不同细菌物种传播。

  • 他们采取了异常严格的保护措施。由于论文和专利相关想法一直“锁在我们电脑里的保险箱中”,这一机制不在公开领域,系统不可能只是检索到现成答案。Costa 强调,Co-Scientist 复现的是人类已经完成的发现;它没有生成那些未公开实验,也没有原创出完整的故事。

  • Penadés 形容这是一次彼此受益的幸运合作:Google 获得了实验性证据,研究者则得到了一次对自身推理的独立检验。合作没有获得 Google 资金,尽管双方曾开玩笑说要索要股票期权。

3. 获胜假设是细胞外尾部交换

  • Penadés 的机制始于一个概念转变:细胞裂解时——无论是诱导后还是未经诱导——相关实体并不是一个完成的感染性颗粒,而是一个装有已包装 PICI DNA 的无尾衣壳。在细胞外,这个衣壳可以捕获不同噬菌体过量产生的游离尾部。

  • 附着的尾部随后决定颗粒能够把 DNA 递送到哪里。因此,一个衣壳可以根据遇到的兼容尾部不同而获得不同宿主嗜性,从而解释同一元件如何抵达多个物种和属。

  • Co-Scientist 将这一可能性排在 5 个假设的前列,或至少接近前列:“需要检验这些衣壳能否与不同噬菌体的不同尾部发生相互作用。”它还独立将注意力引向遗传学已确认具有决定性的 adaptor 和 connector 蛋白;替换这些蛋白,就会改变衣壳能够结合的尾部。

  • 但系统只提供了“最后一幅图”。它没有解释这些组件是在细胞内还是细胞外相遇、释放如何发生,也没有解释实验论文中提出的其他几个概念。

4. 人类研究者被每个噬菌体生物学家“都知道”的事实困住

  • Penadés 坦言,问题在于偏见:“我们知道得太多,也太有偏见。”研究者知道尾部决定宿主嗜性,却默认噬菌体感染后释放的一切都已经同时包含衣壳和尾部;因此,转移失败被理解为受体细菌的问题,而不是供体只产生了衣壳的证据。

  • 在 Klebsiella 中,团队观察到 CF-PICI 强烈复制并完成包装,却没有发生转移,随后尝试了许多受体菌株。在 E. coli 中,一名学生逐个删除约 6 个前噬菌体;删除其中一个后,即便诱导 PICI 仍在继续,原本很低的转移率也完全消失,由此区分出负责诱导的噬菌体和提供尾部的噬菌体。

  • 组件层面的逻辑其实已经存在了 50–60 年。研究者可以诱导 lambda 或 ϕ80 的衣壳和尾部突变体,混合裂解液并恢复感染性颗粒;但正如 Jessica Sacher 所指出的,他们此前没有进一步想到:一个元件的衣壳可以接纳另一个元件的尾部。

  • 一旦 Penadés 和他的学生提出由不同噬菌体分别负责诱导和提供尾部,“一切开始变得合理”,进展也迅速加快。一个将 CF-PICI 供体、功能性前噬菌体供体和受体结合起来的简单实验,证明了这种移动能够发生,包括在接近群体状态的条件下。

5. 这场历时多年的发现,进一步凸显了偏见的教训

  • Penadés 说,他原本准备在 2010 年发表 CF-PICI,但最早发现葡萄球菌 PICI、并将其命名为 SaPI 的 Richard Nobig 警告称,衣壳基因会强化“PICI 是缺陷噬菌体”的旧误解。团队于是等待,并持续积累多年的复制、包装和无法转移的结果。

  • Penadés 通常把无知视为抵御成见的保护:“我不会读太多。好吧?这样我就不会有偏见。”他告诉学生,应相信控制良好的实验结果,而不是文献中的说法;但在这次事件中,他深厚的卫星研究经验反而制造了自己通常能够避开的盲点。

  • Sacher 将 AI 的贡献比作一种没有 100 年历史和成见负担的“新手心态”。Penadés 表示认同,但也收窄了这一说法:Co-Scientist 并没有真正理解这个系统,只是把两个熟悉事实连了起来——广泛分布与尾部决定宿主嗜性——而没有受学科边界约束。

6. Adaptors 和 connectors 可能把宿主范围变成工程变量

  • Costa 区分了真正的细胞宿主嗜性与组件兼容性。尾尖仍然负责结合细菌受体,但 portal-adaptor-connector 组成的颈部结构决定衣壳是只接受某一种尾部,还是能接受一组兼容性更广的尾部,并间接控制哪些宿主嗜性可以实现。

  • Penadés 说,团队在 2023 年提出,CF-PICI 基因可能来自 HK97 噬菌体 machinery。衣壳、portal、terminase 和 protease 组件已经进化出特异性,而 connector 和 adaptor 则必须同时结合形成衣壳的 PICI 以及相应尾部。完整机制仍未解决。

  • 研究者推测,也许只需“1 或 2 个残基”就能决定尾部兼容性是狭窄还是广泛;Costa 的实验室已经开始研究这些结构编码了什么。

  • 专利机会在于构建一种能够结合多个尾部、并将 DNA 递送到多个菌株或物种的合成 PICI。Costa 将其描述为一套可定制的治疗或诊断工具箱,可能缓解噬菌体疗法宿主范围狭窄的问题,而不是一个已经验证的产品。

  • Penadés 认为,普通裂解性噬菌体也可能交换尾部,尤其是在抗尾部防御系统阻止尾部形成、只留下衣壳的情况下。他明确保留了这一限定:颗粒可能进入另一个物种,但进入后噬菌体能否在那里存活,则是“另一个问题”。

7. Co-Scientist 的架构带来了更好的问题,而非最终答案

  • 与 Gemini、ChatGPT 及其他系统的基准测试产生了不同格式和一些看似合理的替代方案,例如一种能够将 DNA 注入不同物种的尾部。但没有任何系统触及关键的不同尾部机制;其中一个似乎找到了实验预印本,却仍未提取出答案,另一个则对 PICI 生命周期掌握得非常好,却没有完成决定性的推断。

  • Co-Scientist 的输出更为完整:包括排序后的假设、相关科学家和支持性文献。Penadés 回忆,尾部假设背后大约有 15 篇真实且合理的论文,为用户提供了一条可检查的证据路径,而不是一个无依据的答案。

  • Costa 只在较高层面理解这一系统:候选假设彼此竞争、相互质疑,并获得类似 Elo 的排名。他说,AI 专家似乎对这一架构印象深刻,尤其是它能够完成一些不那么直接的推断,例如 CF-PICI 可能通过接合实现动员。

  • 可测试性是它在实践中的一大优势。系统建议用冷冻电镜研究 adaptor 和 connector 的结构,用脂质体实验检验囊泡假设,还提出了一个错误假设:衣壳蛋白可能无需尾部就能结合受体。团队的无尾部对照排除了最后一种可能。即使是错误候选,也可能转化为清晰的实验。

8. 科学方法仍然有效,但文献偏见依然是失败模式

  • Costa 的工作模型是“像一个合作者”:假设不是“普遍真理”,研究者仍必须进行实验、解读结果并得出结论。针对 Joe Campbell 的质疑,他澄清说,早期指导可能改变实验选择——或许将失败率从 90% 降至 50%,节省半年或一年——但不会改变证据标准。

  • Penadés 将节省具体化:正确的框架本可以避免反复筛选受体菌株,因为那些颗粒根本没有尾部。但要还原准确的反事实很难,因为实验室往往更容易记住成功实验,而不是失败实验,也很少发表阴性结果。

  • 第二次反向测试暴露了系统的局限。面对“质粒为什么没有噬菌体 pac 或 cos 包装序列”这一问题,Co-Scientist 生成了有思路但错误的“自私元件”解释;团队更新后的理解是,质粒可能会避免过度移动,以维持多样性,这是一个在文献中代表性不足、也更不强调自私性的框架。Campbell 提出的 Tn10 类比——多拷贝抑制可防止转座压垮其 E. coli 宿主——支持了这一更广泛的直觉。

  • Co-Scientist 当时仍在开发中,尚未向公众开放,有兴趣的实验室被引导加入 Google 的可信测试者计划。Penadés 和 Costa 目前正与处于不同项目阶段的博士后前瞻性使用该系统,测试它的“正向版本”能否在无人预先知道答案时,保留反向实验的价值。

Speaker 0

Today, we're following up on our recent episode on Google's AI co-scientist with a special crossover episode from the Potovirus Podcast, in which hosts Dr. Jessica Sacher and Dr. Joe Campbell speak with José Penadés and Tiago Costa, the scientists at Imperial College London who recently made a surprising discovery that Google's AI co-scientist later put forward as a hypothesis entirely on its own.

For context, here's a quick crash course on the biology you'll hear discussed in this episode. Bacteriophages are viruses that infect bacteria. In general terms, these phages reproduce by inserting their genes into a host bacterial cell and hijacking the cell's protein-making mechanisms to produce copies of the virus itself until the cell ultimately bursts and releases many copies of the virus into its environment. Structurally, phages consist of a tail, which is specifically adapted to attach to specific types of bacterial cells, and a head, also known as a capsid, which stores and protects the genetic material until it's injected into a target cell.

Phage-inducible chromosomal islands, also known as PICIs, are a fascinating product of evolution. They are DNA sequences that have evolved to lie dormant in bacterial genomes until the cell is infected by a certain type of bacteriophage, at which point they become active and hijack the virus's reproduction process, replacing the virus's normal DNA with copies of themselves. The affected cell still ends up bursting, but instead of releasing copies of the virus that infected it, it releases viruses that spread the PICI DNA to its sister cells, in some cases thereby serving as a collective bacterial defense against the attacking virus.

The question that José and Tiago and their teams, as well as Google's AI co-scientist independently, had set out to answer was how a certain class of PICIs, known as capsid-forming PICIs, or CF-PICIs for short, which encode only the capsid or head portion of the virus with no tail to latch onto other target cells, had somehow managed to spread widely across many different types of bacteria. The surprising answer, which had eluded the human scientists for years but which Google's AI co-scientist was able to surmise just from its analysis of the relevant literature, at an estimated inference cost of maybe between $100 and $1,000, is that these capsids have evolved the ability to connect with different kinds of viral tails in the environment. That's how they were able to infect and ultimately become incorporated into many different kinds of bacterial cells.

Beyond serving as a vivid reminder that evolution is an eternal arms race, and that we should absolutely avoid evolutionary competition with AIs at just about all costs, this episode shows that, as of the Gemini 2.0 generation, large language models with proper scaffolding and a decent inference budget can now contribute to frontier scientific research—not just by expediting the grunt work, but in some cases by providing an unbiased perspective or even the key insight. Obviously, such hypothesis generation can and will accelerate scientific discovery, even if its hit rate is ultimately fairly modest. You can imagine how quickly this becomes even more powerful as Google plugs Gemini 2.5 into the co-scientist architecture, which they've surely already done by now, and then again as AIs begin to direct experiments and collect their own data via Cloud Lab APIs. This is all science fiction, but it's happening for real in our lifetimes right now.

And once again, I can only conclude that the singularity really is quite near.

Jessica Sacher

Welcome, everyone, to Potovirus Podcast. Today, we have a special episode, and I might say that every time, but we're doing lots of different things all the time these days. I heard about this exciting AI story that's also a phage PICI story, so it was the obviously perfect choice of topic. I'm so glad to have convinced José Penadés and Tiago Costa, who are professors at Imperial College London, to join us.

They have been working together—we're sure to hear a lot more about that—but they have been working in the space of mobile genetic elements, phages, and specifically phage-inducible chromosomal islands, or PICIs. Apparently, Google let them use its not-yet-released co-scientist tool, which is an AI tool, and they were able to use it to re-derive a body of work that they had not yet published but had been working on for a couple of years. Giving the AI tool this research question and a little bit of background, but none of their unpublished data, they saw that the AI co-scientist was able to come up with the same hypothesis that they had recently proven but not yet published.

This is a story that we want to dig into, and you might have seen a bunch of coverage of it lately. It's been in Forbes, The Economist, and the BBC. So they made it to the mainstream, which doesn't happen all the time with phages. I'm sure you'll be able to look into that coverage too, but we wanted to talk to our scientists—our phage scientists especially—and everyone who might be curious about what these AI tools are useful for and how they actually start using them.

I'm very excited to have you both here, and of course, my co-host, Joe Campbell, is here too. We'll get deep into this and hopefully find out more about the pros and cons of using this kind of tool, as well as its limits right now. To start, tell us just a little bit about what you were studying when you came across this tool. How did that even come into your lap?

1. CF PICIs Defy Species Boundaries

José Penadés

For more than 20 years, probably closer to 25, we have been working with these PICIs. They are phage satellites, so theoretically they use other phages for induction and mobility. At the beginning, we had quite a few issues because people thought that these PICIs were defective phages. For more than 10 years, we were fighting to establish PICIs as a new family of mobile genetic elements, of phage satellites.

Probably 10 or 12 years ago, we discovered a peculiar PICI family: capsid-forming PICIs, or what we call CF-PICIs. They were a peculiar family because they have all the genes required to produce the capsids, the small capsids, and to package the DNA into the capsids, but they still require tails to create active particles. This is quite unusual because the majority of satellites just hijack everything from the phage.

We realized that these capsid-forming PICIs were present with exactly the same DNA—the same element—in many different species. Sometimes we use as an example one capsid-forming PICI that was in 7 species from 5 different genera. It was a bit surprising. We thought that this was probably a new mechanism of gene transfer, allowing these elements to move between species and even between genera. This is what we have been doing in the lab for a few years: trying to understand the mechanism.

Tiago is a structural biologist. He can introduce himself later. It's quite funny because he works on another important mechanism of gene transfer: he's an expert on conjugation. We really wanted to know what these capsid-forming PICIs look like structurally. The genes involved in producing the small capsids of the capsid-forming PICIs are related to phage proteins, but somehow they have evolved the ability to produce a small capsid and package the PICI DNA inside.

We started working with Tiago because he is the expert at determining these kinds of structures, and in fact, he solved the mystery. But I think it's better that he introduce himself and explain.

Jessica Sacher

Go ahead.

2. The Smaller Capsid Strategy

Tiago Costa

We solved a fraction of the mystery, so we still don't know the full story, but I think we're on the right track. This is also supported by the AI co-scientist, which seemed to agree with our experimental data and our hypothesis.

I come from a different mechanism of gene transfer: conjugation. My lab and José's lab share this love and excitement for bacteria and how they either directly transfer DNA among themselves by conjugation or use phages to mobilize the DNA. José asked me to team up with him to understand, from different angles, the mechanism by which these capsid-forming PICIs are able to spread widely among different bacterial species.

So my role, or the role of my lab, was to understand at the molecular level how all those cf-PICIs are formed structurally, and how these capsids are able to take one protein—a phage protein—and change it to make these structurally very different entities. To go into a little more detail, we understood that there are some insertions of protein sequences into the capsid-forming proteins that dictate—we think they dictate—the different symmetry of the capsid-forming units and make them smaller, so that they can only accommodate the DNA of the cf-PICIs and not the DNA of the phages. This sort of makes it smaller.

Jessica Sacher

Oh.

Tiago Costa

So that they can only harbor the DNA of the PICI.

José Penadés

Usually, satellites, these kinds of PICIs, have one-third the size of the phage genome. Usually, these PICIs are around 10 to 15 kilobases, and the helper phages are around 45 kilobases. A classical mechanism of phage interference is for the satellites to produce small capsids, where just the PICI, or satellite, DNA can be packaged.

Jessica Sacher

Got it.

José Penadés

These capsid-forming PICIs all produce small capsids. They also have a specific terminase to package the cf-PICI DNA, not the phage DNA. So this is a very specific packaging system to package the capsid-forming DNA into small capsids that are formed by proteins encoded by the satellite. That was another unique thing compared to the phage.

Jessica Sacher

And I think I saw in one of your papers that these are not considered parasites because they don't have a detrimental effect on the phage. Is that right?

José Penadés

For many years—and I think, probably in terms of selling the papers, it was nice to tell the story of a war between parasites and phages—but in this case, with the prototypical member of this family, we didn't see any interference with phage reproduction. We don't know whether that's the case for other cf-PICIs or not. But we published, I think it was 2 years ago, that some phages might suffer from inducing the PICIs, because some of the PICIs have anti-phage systems. So there is even some cost to inducing these PICIs.

Jessica Sacher

Mm.

José Penadés

So we don't see them now as having a parasitic relationship. We might see them more as synergistic, depending on the scenario. They're more friends than we thought probably 15 or 20 years ago.

Jessica Sacher

Okay. Okay. So what made you want to use this co-scientist AI tool? Where were you when you figured out that was an option, and why did you set it to work on this problem?

José Penadés

Tiago started using it for his work. Tiago, go ahead.

Tiago Costa

Yeah. There was a Fleming Initiative at the college that put us in contact with Google because they knew that Google was developing this new AI, a large language model tailored for scientists, and they wanted to engage with scientists so that they could understand how good the language model was.

Initially, this was not related to PICIs. The work started with my lab trying to understand a particular question in conjugation that had been unknown for 70 years. We were interested in understanding how conjugation is actually initiated. We wanted to understand the molecular determinants behind the T = 0 of conjugation.

We challenged the algorithm, and the algorithm came up with several hypotheses—5 different hypotheses. We realized that it would take us many months to understand whether the highest-ranked hypothesis would be correct.

José was in that meeting and immediately saw an opportunity. Why don't we challenge the algorithm with our actual experimental data that we already had in our hands? We were about to submit—or I don't know if we were about to submit or had already submitted—to one of the top-tier journals. Basically, we turned Google's strategy upside down.

Jessica Sacher

Yeah.

Tiago Costa

Instead of having an AI-driven hypothesis that would be validated a few months down the line, in the case of the PICIs, we already knew the mechanism behind the spreading of the cf-PICIs in different bacterial species, and we challenged the algorithm with that exact same question. It was astonishing. One of the top hypotheses, or the top hypothesis, was basically a recapitulation of what we had observed in the lab.

Jessica Sacher

Wow. So were you surprised? Yeah, go ahead, José.

José Penadés

He didn't say this, but the hypothesis that the co-scientist suggested for the conjugation—I’m not an expert there—really impressed him. That's why I said, “Okay, this looks really good. I need to take some advantage here,” because if the hypothesis provided by the co-scientist for conjugation were crap, I probably wouldn't have been interested. But he was really impressed. I said, “Okay, this is looking very good.”

So we changed the approach, and it was quite impressive. The audience would probably like to know that the way these capsid-forming PICIs can be mobilized between different species is very simple. They can produce these capsids with the PICI DNA inside, and that was a mistake we made for many years. We always thought that whatever was released after phage infection was an infectious particle.

Jessica Sacher

Yeah.

José Penadés

But we realized in this paper that these PICIs can be induced, or even without induction, as soon as the cells lyse. What is released is just the PICI capsid with the packaged DNA. Outside the cell, they have the ability to take tails from different phages—free tails.

Jessica Sacher

Huh.

José Penadés

They have to be free tails. Depending on the tail, that will determine the tropism, because we have known for many years that the tails in phage biology determine tropism. So there were many things in the paper before we realized that, with or without induction, as the cells lyse, the capsid with the packaged DNA will be released with no tails, and then outside there will be free tails.

We also knew for many years that many tails are produced in excess. Even during purification, they already showed that there were more tails than infected particles or whatever. For the same capsid, the tail that binds to it will determine the tropism.

3. Google Finds The Missing Tails

One of the hypotheses that Google gave us was, “You need to check the possibility that these capsids can interact with different tails from different phages.” The system didn't talk about how that happened or how it was induced, so it was just the final picture. Many of the new concepts that we put in the paper are not in the Google document.

Jessica Sacher

Yeah.

José Penadés

But the final thing was right. Something else that was quite impressive is that we knew from the genetics—and now Tiago is solving the structure, hopefully it will be soon—that there were 2 key players in the ability of a capsid-forming PICI to bind to one tail or another. They are 2 proteins called the adaptor and the connector. If you swap the adaptors and the connectors, you always swap the ability of a capsid-forming PICI to bind to some tails or others.

The co-scientist also suggested that we should look at the adaptor and the connector.

Jessica Sacher

Wow.

José Penadés

There were so many things that were not in the output from the co-scientist.

Jessica Sacher

Yeah.

José Penadés

But the key one—the tails—was there.

Jessica Sacher

Wow.

José Penadés

There was another one. The system also suggested that we check conjugation. That was quite funny, as a new mechanism, because conjugation is theoretically more promiscuous than transduction due to the narrow specificity of a particular tail. We've started working on that as well to see if that's true or not, but that was quite an interesting hypothesis that the system provided.

Jessica Sacher

Tiago must have been excited that it chose your favorite area.

Joe Campbell

Yeah. It was back to the beginning, right? That was the first thing. We have to try our best to understand the question a little better.

José Penadés

I said this to phage biologists.

Joe Campbell

The question.

José Penadés

Tiago loves the adaptor and the connector.

Jessica Sacher

Ah.

Joe Campbell

Yeah, yeah.

José Penadés

Have dinner or lunch with him, and you just need to mention the adaptor and the connector. It's pretty cool that some adaptors and connectors bind to one specific tail, which looks to be very narrow, while others have the ability to bind to several tails from different species. He's always very excited to see how, mechanistically, probably 1 or 2 residues—

Jessica Sacher

Love it.

Joe Campbell

—to get a little better understanding of the process. So I assume, because you hadn't done this when you started the experiments, that the hypotheses generated by the computer were things that you had thought about and therefore designed experiments to test? Or did you have ideas...

Jessica Sacher

How were the experiments designed by you that ended up leading to the conclusions that agreed with the AI-generated ideas?

José Penadés

I agree with that, because the timing is very important here.

Jessica Sacher

Yeah.

José Penadés

So we have been thinking a lot about how this machine got the right answer.

Jessica Sacher

Right.

José Penadés

It's very frustrating, because we have the answer there and we didn't see it. As phage biologists, we know that the tail determines the tropism. Is that right? We've known that for many years.

Jessica Sacher

Yeah.

José Penadés

So if one DNA is using a tail in many different species, it's because maybe it uses different tails. But we were biased. That's the problem we had, and I'm killing myself over that because, for many years, I always thought—and all phage biologists would think—that after infection, what you have are infected particles.

Jessica Sacher

Yeah.

José Penadés

With the capsid and the tail. And we knew it was very strange. We have phages that can be induced, but we didn't get transfer. We couldn't understand this because we were wrong in the way we solved it. We were so biased. We always thought that after induction, whatever is released from a bacterium consists of infective particles, both for phages and for satellites. And I think that's the big thing.

Google didn't understand what was happening here, but made the simplest connection: if you are in many different species, and to go to these species you need a tail, you are binding to different tails. That's it.

Jessica Sacher

So, look, if I could just go back to the order, am I hearing you say that you had data—

José Penadés

Yeah.

Jessica Sacher

—that you weren't exactly sure how to interpret it, and then you—

José Penadés

Correct.

Jessica Sacher

—realized that—

José Penadés

Correct.

Jessica Sacher

—that this was telling you how to interpret it?

José Penadés

No, we had the data for years.

Jessica Sacher

Yeah.

José Penadés

For example, the PICI we use—the capsid for it is from Klebsiella. We use the PICI, so this is a prophage that is a helper prophage. Until the paper is published, for us, a helper prophage means a phage that will provide whatever is required to produce an infective particle: capsid and tails, or just tails.

So we have a Klebsiella 1 strain that we use. We show very nice capsid formation, replication, very nice packaging, and no transfer.

Speaker 1

Hmm.

José Penadés

So in our mind, it was because the recipients we used didn't have the receptor. But we were thinking all the time that they were infective particles.

The same happened to us in E. coli. Then, one day, suddenly, talking to my student, we realized that in E. coli there were several prophages. When we induced them, we knew that the island was heavily induced, with a lot of replication and a lot of packaging, because we could purify the capsids from the lysate, but the transfer was very low.

He was making mutants to delete, I think, 6 prophages one by one to see what was induced, or whatever. Then, one day, he realized that when he removed one, it eliminated the transfer—even the low transfer—but the island was still induced.

At some point, we thought that maybe we'd need phages to induce and phages to provide the tails. Then we started connecting the dots. It's been known for 50 or 60 years that you can get lambda or ϕ80 lysates, mutant capsids, and mutant tails. You induce them, take the lysate, mix them, and then you have infective particles.

Jessica Sacher

Mm.

Joe Campbell

So if you—

Jessica Sacher

But the people who did these experiments never thought of taking capsids from one phage—

Joe Campbell

Yeah.

Jessica Sacher

—and tails from another. Never.

Joe Campbell

Yeah.

Jessica Sacher

Because we all thought that everything was there. That's the big bias.

Joe Campbell

Mm.

Jessica Sacher

So then, that day, everything started making sense.

Tiago Costa

I was excited when we understood that the adaptor and connector are the key factors here. You should have seen José that day when he cracked this hard nut. He was very excited because he knew that this was something big and novel.

Jessica Sacher

And you did this before the AI came into play. You cracked this, but you were spending a couple of years trying to get there, and then—

José Penadés

More than a couple.

Jessica Sacher

Okay. More than a couple.

José Penadés

Yes.

Jessica Sacher

Many years, a scary number. And so then, the fact that the AI came up with that was because it did not have this bias: once you release a particle, of course it has its tail and whatever it needs to be infectious.

José Penadés

Right.

Jessica Sacher

It was unbiased in that thinking. And, of course, why not? Could you have this later meeting of capsid and tail outside the cell? And that's sort of what happened.

Tiago Costa

So we also—sorry, José—I think we need to clarify the chronology here a bit. There was another angle that was in play here, which was that we saw an opportunity to patent these new elements because these elements have a broad range, which addresses one of the limitations on phage therapy.

Phage therapy has always had a narrow range. These phages used in phage therapy either only hit a bacterial species, or sometimes there are strains within that species. With these chimeric particles, you can pretty much tailor-make them and open up a toolkit where you can customize the target bacteria that you want to hit for therapy or to make diagnostic tools.

So we kept this very secretive because, until the patent was filed, we could only speak about this in public once the patent was filed.

Jessica Sacher

Right.

Tiago Costa

So there was new public domain.

Jessica Sacher

Yeah.

Tiago Costa

That's why we are very reassured that the AI system would never have access to this manuscript or these ideas, because they were kept in a safe box on our computers.

Jessica Sacher

Yeah.

José Penadés

Also, we challenged other systems with the same input. None of the other systems even cross-related to the AI co-scientist.

Jessica Sacher

Yeah.

José Penadés

I think the important thing is that we didn't mention this in the preprint. When we got the first round with the 5 hypotheses, we asked—because, as I said, it's not really a real mechanism in terms of the capsids being released outside the cells.

Jessica Sacher

Mm.

José Penadés

It was just, “Check if this can bind to different cells,” without telling it how this interaction would happen, or when it would happen, or whatever.

Jessica Sacher

Yeah.

José Penadés

We said, “Make a second round.” And I said, “Yeah, but this doesn't make any sense. How are these particles produced? Where will they be?” And the reply was, “Crap.” You know what I mean?

So the system doesn't really understand what is proposed. It's just connecting the dots, which makes sense. Just taking this simple thing—that if the tails determine the tropism, maybe it's because it's using different tails. That's it. It was there. We all knew that.

This is the type of thing where, when people read the paper, every single person will understand it very quickly because it's obvious—until somebody needs to tell you. You know what I mean?

Jessica Sacher

Yeah.

José Penadés

So it was quite frustrating because it was quite a few years trying to understand. We have strains and a lot of mutants, and we couldn't understand why we didn't get transfer. It's because they were just tailless. We said this is a new entity. We call it a tailless, capsid-forming entity with packaged DNA. These are the entities that are released.

Jessica Sacher

Can regular lytic phages also have this sort of mix-and-match of tail and capsid, or have you started to look?

José Penadés

Yeah.

Jessica Sacher

That's so cool.

Do they—

José Penadés

Same, you know?

Jessica Sacher

Yeah. They have these adaptor-connectors. Do many phages have those? Is that conserved?

José Penadés

We have asked.

Jessica Sacher

I don't know. Okay.

José Penadés

Sorry. I think the biggest difference is that when a phage infects, the capsid and the compatible tail are already there. So we think that the number of capsids that will be released will probably be lower.

But structurally, the capsid-forming proteins involved are exactly the same as those you have in the typical HK97. But we also think that even all the myoviruses, or whatever, work in a similar way.

4. Phage Therapy Gains A Toolkit

The funny thing is that we're working on this now, and probably it's good that we share it with the public, because then, if somebody copies us, we can always get the credit.

Jessica Sacher

Yeah.

José Penadés

So now you know that there are many anti-tail systems—for example, antiphage systems that block tail formation. When that happens, you will have capsids.

Jessica Sacher

Oh.

José Penadés

It's true.

Jessica Sacher

It's completely coming out.

José Penadés

So we really think the same will happen with phages: they will have the ability to inject DNA into different species by swapping tails.

Whether these phages in the new species survive or not is another question. We also show that when the capsids form and go to another species, they have a mechanism to hijack other phages in the new species so they can also be mobilized and transferred. Mechanistically, we think that will be the case.

Jessica Sacher

Yeah. So tropism is—you’re narrowing in, getting a higher resolution on what it’s actually defined by—not necessarily just the tail, but specifically this connector point, maybe a little bit.

Tiago Costa

Yeah. I mean, true tropism will always be determined by the tip of the tail, right? Whether it binds the surface or the receptor in the target cell. But it’s true that the neck region, where the portal, the adapter, and the connectors are, will have a say in determining that tropism—not by direct binding to the surface or receptor, but by generating a promiscuous or a very specific structure that binds only one or several phage strains. That’s what we are looking at, what we are starting to understand now: what is imprinted in those structures of these proteins that determines that specificity or promiscuity in binding only one type of tail or a set of tails.

Jessica Sacher

Yeah.

José Penadés

But this is for the next story. This is the next story.

Tiago Costa

Yeah.

Jessica Sacher

Yeah.

José Penadés

This capsid-forming system—we proposed in the paper that we published in 2023 that the genes came from the HK97 phage, let’s say. But they have evolved, so there is no cross-talk between the 2 proteins, even though they are very similar in terms of sequence. The capsid, the portal, the terminases, and the protease are all very specific. Everything is very specific until they arrive at the connector and adapter. Each connector and adapter needs to bind to the capsid—the C-PICI or the phage PICI—and then they need to bind to the same tail.

So it will be quite a funny thing. You evolve something to be completely separate in terms of capsid-forming components and the proteins involved, so they cannot cross-talk, but then the tail is the same, and there will probably be competition between them as well. It’s quite interesting to know what’s happening there, because the idea we have is that probably we can create synthetic PICIs with the ability to bind to multiple tails. If that’s possible, we will solve one of the limitations for phage therapy. We can deliver DNA to multiple strains or multiple species. This is the pattern that Tiago was mentioning before.

Jessica Sacher

Yeah. Yeah, I think this is so cool. It feels like one of the first fresh phage host-range discoveries that I can remember in a while. It’s another level of looking at it as a whole, and pairing it with the AI aspect, it feels like we have this potential set of tools that we’re all going to use more, not less.

Now we can have someone almost in a beginner’s mind, who doesn’t have all the assumptions of our 100 years of history, give us input into what would physically be probable or possible. Then we can bounce that off a whole other body of knowledge built differently from ours.

5. AI Removes Scientific Bias

José Penadés

For me, it’s a kind of revenge, you know? A nice one for me. I normally don’t say these things because people say that this is not the right thing to say, but I think I succeeded—or some of the things we found in the lab happened—because I don’t read too much.

Jessica Sacher

Yeah.

José Penadés

Yeah, yeah. So many—

Jessica Sacher

You can admit it. It’s hard to admit.

José Penadés

I admit it, you know.

Jessica Sacher

Yeah.

José Penadés

So many times, I have a very good idea and go, “Wow, this looks so good,” and then I find that maybe somebody in 1985 felt the same. Then you have a huge admiration for these people.

But I wasn’t biased. That happened, for example, with lateral transduction and these kinds of things.

Jessica Sacher

That’s interesting.

José Penadés

Completely ignorant. I always say to my students, “I trust your results,” because the students sometimes say, “José, but you know…” I don’t mind what people have said. You have the controls, it works well, so let’s interpret your results, okay?

Now I was so biased. I knew too much about the satellites, and I had this thing there—the tails. It was obvious, you know what I mean? So I think that’s the good advantage of these kinds of systems: there’s no bias. It doesn’t matter if this is a phage or a conjugation. We always thought that the satellite should move by a phage because that’s the name. We never thought, “Why not by conjugation?” It’s just a matter of putting in a very small sequence.

So it was a good revenge in that way. The system gave me that because I knew too much and I was so biased. We were so biased when we talked about this thing, and then we realized.

Another funny thing is that I had no idea about tails—zero. It’s something that I never, ever thought would be interesting at all: the tail, its form, the funny things with the packaging, the capsid. These were the big papers on the capsid in 1997, and now suddenly it looks like the tail. It’s a funny thing as well, because as you said, you can now have a very simple mechanism to move things between species.

Jessica Sacher

Wow. Yeah. So were Google excited that you showed the world a nice test case? It’s pretty neat how well it worked. Is it really that clean-cut? They must have been very excited.

Tiago Costa

Yeah. Well, that’s a question for them, obviously, but I think they were very pleased to have experimental evidence that the hypothesis the system generated was sound, solid, and verifiable in the lab. There were other studies initiated by hypotheses from the AI co-scientist that are also reported in their system preprint, but they still have to go through peer review and publication.

In our case, we were a little bit more advanced because we already had a full story. We knew the answer, and the AI system was able to recapitulate it.

José Penadés

We tried to claim some stock options from them, but they didn’t agree with us.

Jessica Sacher

Try to claim what was that?

José Penadés

Some stock options from Google.

Jessica Sacher

Oh, yeah. I was going to say they should be funding your lab.

José Penadés

They didn’t agree with us, okay? I think they were lucky, and we were lucky as well, because this is the perfect system to test. As I said, everything was there. You produce the capsid, you can produce the tails, and people had already done experiments mixing things—tail mutants with capsid mutants—and then you have an infected particle. So everything was there. We just had to connect the dots, with a big bias.

I think it was lucky for them and lucky for us.

Tiago Costa

There was no funding involved, so they didn't invest a penny.

José Penadés

We are getting money from Google. We don't. We are trying, but so far nothing.

Tiago Costa

So it was mutually beneficial for both sides.

José Penadés

I don’t know if you checked the preprint, because we didn’t put this in the preprint. Sorry, Tiago. We didn’t put it in the preprint, but for each hypothesis, the system provides key papers that it used to arrive at that hypothesis. Do you think it would be nice, for example, to include that information as well for people? They were mentioning names, as you know.

Jessica Sacher

Yeah.

José Penadés

For each name, they highlighted one paper that they thought was relevant to the hypothesis, and that was pretty cool as well. For each hypothesis they ran, they listed a few papers.

Jessica Sacher

Yeah.

José Penadés

It was for the tail hypothesis we were mentioning, and I think there were 15 papers that they said, “Okay, these are the papers they found relevant to arrive at that hypothesis.”

Jessica Sacher

They were real papers.

José Penadés

Yeah, real papers. It made a lot of sense.

Jessica Sacher

Yeah.

Joe Campbell

Can you tell whether the other AI tools you used that weren’t as effective saw those papers and failed to realize their importance? Or did they somehow miss them? I guess, in trying to understand what’s different and why they’re giving different answers, I could broadly think of 2 categories.

One is that somehow they didn't find the right papers, or they found them and didn't realize or weren't able to understand their importance. So can you go back and figure out whether those other search engines—or whatever learning tools or AI tools—actually looked at those papers?

Jessica Sacher

Yeah.

Tiago Costa

So, part of the preprint was benchmarking the Co-Scientist system against other systems, such as Gemini and ChatGPT. The outputs from those systems are different, and the format is not the same. However, the hypotheses and answers are there. The Co-Scientist system outputs references and notable scientists who have worked in that field, so it's a very comprehensive report.

The other systems don't give such a comprehensive output because they were built differently. We are not AI people, and, for me at least, Co-Scientist is still a black box, but the system is built differently. Funny enough, we asked where we had published the preprint of the experimental paper in bioRxiv, and then we challenged the different LLMs. Some of them hadn't seen the preprint. They couldn't even find the data we had already published.

José Penadés

I think what the system found was the preprint.

Joe Campbell

Yeah.

José Penadés

But it didn't provide the right answer.

Joe Campbell

Mm.

Tiago Costa

So it's the same: it didn't provide the right answer even though it found it.

José Penadés

But they found the preprint. I think that was one case.

In the other systems, they also proposed things that made sense. For example, some systems proposed that maybe these capsid-forming satellites can take tails that have the ability to inject DNA into different species—the same tail.

Tiago Costa

Okay.

José Penadés

For example, in that case, that would work both for this capsid-forming satellite and for any satellite. So it is not something specific, because in the end the tails are very similar. Some of the hypotheses were kind of okay, but none of the other systems provided the right thing—swapping different tails, you know.

Even some of the systems—I can't remember which one was the last one we put in the paper—made a very nice overview of the PICI cycle, so it definitely had very good access to the bibliography. But I don't know to what extent, as I said before, that's a problem. This is the problem we have: you have very good knowledge, but then how do you interpret things as something new?

That's the question that some people ask us: to what extent are these just new hypotheses, or are they just connecting the dots? I think so far it's more connecting the dots in an unbiased way than thinking of something that is completely novel.

Tiago Costa

Okay.

Joe Campbell

Yeah. If I could ask just a couple more questions. First, are you in any way planning to rewrite the manuscript based on what it did? I guess my guess would be that it would most likely affect the discussion, but are you planning to rewrite it? I'll ask what you add to that before I ask the follow-up.

José Penadés

For the experimental manuscript, nothing changed. Basically, the output from Google was already addressed in the experimental manuscript. Even in the experimental manuscript, we provide new concepts that Co-Scientist never mentioned.

What we are doing now is creating a new manuscript about our experience with Co-Scientist, explaining what input we provide and then evaluating the outputs that we got.

Joe Campbell

Right.

José Penadés

We're explaining, for example, the tails: this is a good thing, but these are the things where we have this kind of frustration about why we think Co-Scientist was able to arrive at that point, even in this unbiased way. They even provided the picture, and we also made an evaluation of the rest of the hypotheses.

We posted that manuscript as a preprint as well, and we have sent it to the same journal as the experimental manuscript because—

Joe Campbell

Yeah.

José Penadés

—they might be interested in having both the experimental one and the Co-Scientist-related one.

Joe Campbell

Love it.

I guess the other thing I was thinking about—and maybe this is a hard one—is, if you had had the output of this LLM system before you started the experiments, what would you have done differently, if anything? Would you have done different experiments?

No. Okay.

Tiago Costa

No—nothing changes. It could just save time. The way that I see the system is that it's like a collaborator, something that you can interact with. You obtain hypotheses in the end, but they are not the final truth. You have to go to the lab, run the experiments, interpret the data, and draw conclusions. The hypotheses that the system generates are not universal truths.

The scientific method does not change at all. What it does is put you perhaps on the right path right from the beginning.

Joe Campbell

Right. I understand that the hypothesis doesn't change the fact that you need to do an experiment to address it, but I guess what I'm asking is: back when I used to do experiments, often, if you have a hypothesis, you design an experiment to test it. So when you said it was quicker, is that because you would have designed the right experiment?

Tiago Costa

Maybe, yes. It would be because perhaps you would not fail 90% of the experiments that you failed. You would just fail 50%, and this would save you half a year or 1 year of experimental work—and 1 year to me.

Joe Campbell

So it would have changed things in terms of which experiments you did and didn't do, right?

Tiago Costa

Exactly. It would change—

José Penadés

I think an example, coming back to the idea that we knew we had some strains in which we could induce the island. It had been induced and was highly packaged, and we tried to get transfer, but we didn't get any transfer because there were just capsids—no tails there. Can you imagine how many recipient strains we tried? We thought it was a defect in the recipient strain.

Jessica Sacher

Mm.

José Penadés

Can you imagine how many we tried?

Jessica Sacher

Mm-hmm.

José Penadés

And then we started relating these things because we never thought that these were tailless. We never thought that, if they have everything to package the DNA, they just need a tail. Maybe they can bind to different species.

So, in the end, we arrived at the right experiments, but we failed a lot of experiments because we couldn't understand what was happening.

Joe Campbell

Right.

José Penadés

If they give you some kind of path, the other hypotheses that you realize might be irrelevant for the biology of the satellites—but they were not really important for the question because they can apply to many other satellites, so they were not exclusive to this family. The type of conjugation is the same. They provide a hypothesis that is very easy to test.

Something that we highlight in this second preprint is that, at least in our example, all 5 hypotheses are very easy to test.

Jessica Sacher

Yeah.

José Penadés

You know?

Joe Campbell

Right.

José Penadés

It was also, as Tiago said, that we had the feeling that we were talking to an expert in the field. It would even say, “What do you think about this? I think you should take this thing.”

Joe Campbell

Right.

José Penadés

And then you think, “Okay, that makes sense.” Definitely, at the end we were very happy because we arrived at the same conclusion.

But for the other examples—for example, the one that Tiago was talking about—it has been a mystery for 70 years, and nobody thought about it. Can you imagine how many experiments people did in the past trying to understand the type zero for conjugation? This is a big question. I can't believe that these people working on conjugation didn't really know that.

So now the system provides you with a hypothesis that looks okay. It's saving a lot of time.

Joe Campbell

Right. Yeah, and I wonder if, in your paper about this method, you could give concrete examples of why you would move from 90% of the experiments not working to maybe only 50% of the experiments not working. You could say, “We would have done this, but if we had known this before, we wouldn't have gone down this rabbit hole, and we wouldn't have gone down that rabbit hole.”

I guess maybe I'm trying to think about if someone's out there trying to say, “Why am I going to do this?” I think a lot of people would be attracted to knowing which experiments you did that ended up being dead ends, and which ones you think you might just not have done if you had run this AI before starting the experiments.

Does that make any sense?

José Penadés

Yeah, it makes sense.

Tiago Costa

Yeah. It makes sense, but you tend to forget the failed experiments and keep track of the good ones, right? So I don't know, and that's why we arguably don't publish the negative results, or maybe we should. But anyway, that's another question. It's now very difficult to say which experiments should have been done, should not have been done, or could have been done differently.

Jessica Sacher

Yeah.

Joe Campbell

Yeah, yeah.

Tiago Costa

Sure, this won't be an issue because if it is an AI-driven hypothesis, you have never run those experiments before, okay? So in this case, because we knew the answer, we had the experiments—we had the portfolio of failed and successful experiments—we could make that judgment.

But in the future, if it's an AI-driven hypothesis for experiments, I think what will make for a good—or quick, or quicker—discovery is how well the human can critically interpret the hypothesis and design an experimental setup. Although the AI system already gives you some experiments that you could do to test that hypothesis, the human will definitely still have a very important role in this process.

José Penadés

So I think I can answer your question, okay? We knew this capsid-forming PICI for I don't know how many years. I was almost ready to publish this capsid-forming PICI in 2010, okay? And Richard Nobig told me, “José, you can't.”

Richard Nobig was—and is—the person who discovered the PICIs in Staph that he called SaPIs. This is the first member of the family. Let's say he's the father of the PICIs. He discovered the first element, okay?

He discovered them, I think, in 1998, and for more than 10 years—12 years—people thought they were defective phages. So when we discovered this capsid-forming PICI, that was probably in 2010, and he said to me, “José, you can't publish this thing because then people will think that these are defective phages. They have everything. They have the capsid, they have whatever, so you can't.” Okay? So we waited.

So it has been for many years that we knew they existed. They were kind of manual findings, you know. We knew that they were in many different species. We couldn't understand why a satellite would want to use half of the genome to carry genes that it could hijack from the phage. So it was one of those things. We did a lot of experiments that did not transfer, or whatever.

And the day, talking to the student, that we realized that maybe some satellites use a phage for induction and another phage for the tail, everything made sense, okay? I think when Tiago was mentioning it before, I said, “Tiago, I think we have something pretty big.” And since then, everything was very quick—very quick—because these experiments are very simple.

If you have a capsid mutant, you can have the tails, you can mix them. So these things even work in natural populations. A normal phage or prophage, when it's induced, produces a lot of tails; if you just mix them with the capsid-forming PICI, you have infectious particles. Everything—we have an experiment in the paper. We have a donor for the capsid-forming PICI, a donor for the prophage, a functional one, and a recipient, and things move, you know.

That was the day that we realized that maybe this is something for induction and something to provide the tail. Maybe we realized that everything makes sense. So you have something from Google that said, “Maybe this can use different tails.” You're thinking, you know, because the experiments are very easy to do—at least in our case. I don't know about other areas.

So I think it's because of the bias, and I said, you know, it was a failure because we didn't see the big picture. But having something that is new, it's like the same thing with conjugation. We never thought—I never thought—about conjugation, okay? Tiago never thought about conjugation, and he's an expert on conjugation.

There is a guy at the Pasteur Institute, Eduardo Rocha. He's an expert on satellites and conjugation. He has published a lot of papers and studies about oriTs that are required for conjugation. And we sent an email saying, “Eduardo, have you ever thought about this?” And he said, “No, never.” Because satellites use phages, and plasmids use conjugation, you know.

We are so biased that things that are there, we didn't realize. So I think that's the main power of this system, at least for us. There's no bias. Whatever has been published, it will highlight that. And then it's up to you, as Tiago said, to decide if that makes sense or not. The hypotheses they provided were very easy to test, all of them.

Jessica Sacher

I want to touch on that last point. I think the testability is really interesting, and I wonder if, behind the scenes, the AI co-scientist that Google made was given a lot of training related to understanding what science is like, so it could understand what would be a testable hypothesis versus what would not.

Because I don't think that's usually what you're getting when you talk to regular ChatGPT or other LLMs or even regular scientists. I think that's an advanced skill: of all the things you could do, what are the things that are the most bang for your buck and the most testable, with the least time in the lab? You have to have a lot of background to know that.

But even filtering on that lens and giving you the hypotheses and giving you testable ones, it's pretty interesting that you said they're also testable, and I wonder if that's built into the system. If so, that seems like a really nice feature. Is that your view of it?

Tiago Costa

Like we just said, we don't—I don't know. I think this is a question for the Google team. I mean, we do know the basic principles. We know that the system generates different hypotheses. They are ranked, and then they compete and challenge each other, and they are ranked with an Elo ranking, like chess players.

We know that it's a very complex system. I've heard experts on AI talking about the co-scientist, and they seem to be very surprised by the architecture of the system.

In particular, this novel hypothesis that José was just mentioning about the oriTs, and that these capsid-forming PICIs could also be mobilized via conjugation. These are what they call indirect assumptions or hypotheses, and they are not imprinted directly, or they cannot be readily interpreted from the data that is available.

So whether there is an element of reasoning in the system that will come up with these less evident hypotheses, then, yeah, maybe. But this is again something for the Google people.

José Penadés

The system sometimes provides specific experiments. For example, for the connector and the adapter, they said you should check by cryo-EM how these things look. For some vesicles as well, the system suggests using liposomes.

It was quite a funny thing. Even some of the hypotheses they found, even though they were incorrect, were quite funny because we included them in the paper as a negative. For example, one of the hypotheses was that, in the capsid-forming PICIs, there were proteins that might have the ability to bind to some specific receptors without the tail. So it's like a different mechanism of entry.

And we included these controls because we had experiments with no tails to show that the tail was absolutely required.

Jessica Sacher

Okay. So you ruled it out.

José Penadés

Yeah. It was kind of suggesting experiments in a very easy way, you know, okay? But I again have the feeling here that the phage world is a very easy, very good model for the system because, at least, the work we do in satellites is mutant complementation. There's some cryo-EM. We don't really know in other areas, you know.

Jessica Sacher

Yeah. Wow. Well, yeah. I think we should probably let you go on with your evening and close out, but this has been so, so cool.

I guess my last thing I wanted to end on was just: When can others use this? Or when are you still going to be using it? Is it still kind of in a testers-only phase, or what's that like going forward?

Tiago Costa

So the system is just under development, right? The system is not publicly available still. This is again in the hands of Google, but the discussions that we've had with them are that there is a process for the maturation of the system until it gets publicly released, okay?

I think the strategy will be to incorporate different areas of science into the system and understand how robust the algorithm is. Eventually, if it pans out as well as it has with phage biology, it will eventually be released to the public.

I think now it's still—this is written in Google's AI co-scientist preprint—that there is a trusted tester program, and labs interested in testing the system can contact Google. We will definitely keep working with them because we have a very close collaboration and partnership with them.

Speaker 1

Yeah.

Tiago Costa

and believe it will be mutually beneficial for Google and for us, as it has been—

Jessica Sacher

Yeah.

Tiago Costa

So far. Mm.

José Penadés

We are using it now with some postdocs who are in the process. They're very talented postdocs, so they can challenge the system with the questions they really want to address at different stages. There are projects that are quite new, to see how the system answers the questions, and projects where the system is better established. But this is a different approach. It's now asking for new hypotheses, and—

Jessica Sacher

Forward version.

José Penadés

Yes, a forward version.

Jessica Sacher

I like your reverse version. I think—

José Penadés

Yeah.

Jessica Sacher

Yeah, both make sense, but it's like forward and reverse genetics—like the reverse AI co-scientist. It makes so much sense. Everybody else is sitting on data that they haven't put out there.

José Penadés

We have done another thing that's quite funny. We also work with plasmids and the impact of transduction and plasmid mobility. We have a paper now, and again, nobody knows about the paper. Why do you never have PAC or CAS sequences on plasmids? Because if, as a plasmid, you really want to move, just put a pac or a cos sequence, and that's it—you will fly.

So we have an answer for that. In reality, plasmids don't really want to move too much. If there is a plasmid that moves too much, it will lose the variability, or whatever. But this idea of plasmids being less selfish than we thought is still represented in only a very small proportion of the literature.

For many years, we thought that mobile elements were selfish. So when you ask the system why there are no CAS or PAC phage sequences on plasmids, that body of literature is more important than the cooperativity or whatever. Most of the hypotheses that are pretty good are wrong because they're just based on the selfish model, you know. They do these things. It terminates with the cleavage of the plasmid, and the plasmid will not be functional. The thinking was more that it was a selfish element, something that blocked the biology of the plasmid that spread well. So it's a quite funny thing, you know. The thinking is really good.

Jessica Sacher

Yeah.

José Penadés

The system also needs to learn which papers are more important. So this is what we're working on. In sum, we try again with the reverse for the plasmid thing. The hypotheses were pretty cool, but they were wrong.

Jessica Sacher

Yeah.

José Penadés

And we are working now, trying to—

Joe Campbell

Did you try the other LLM systems with that one?

José Penadés

Not for this one. You can do that.

Joe Campbell

Would that be interesting? I mean, I guess it could be scary, but it could be interesting if—

José Penadés

No, no.

Joe Campbell

depending on the question you ask, different LLMs work better.

José Penadés

That's a good point. We are now having a chat, and we're telling the system, “Yeah, that's very cool, but you are ignoring that these are not really selfish. You can have another type of relationship with the cells, so let me know what you think if you include that possibility as well.” And we're waiting, you know. It's quite interesting.

Joe Campbell

Yeah. I guess you want to be selfish, but within reason. If you're dependent on bacteria for survival, if you get too selfish, you might kill your own host.

When I was a graduate student, I was working in a lab that studied this transposon Tn10, and it had this thing called multicopy inhibition. You would naively think that it just wants to keep jumping and making more and more, but eventually that's bad for the cell.

So it has a mechanism that prevents it from transposing and makes it transpose less when there are too many copies in the cell. That's what I was thinking about when you were talking about it.

Because again, that's another selfish DNA that you would think should jump whenever it can, but if it starts jumping too much, it trashes the E. coli it's in, and then—

José Penadés

Yes.

Joe Campbell

Correct.

Jessica Sacher

Yeah.

Joe Campbell

All right.

Jessica Sacher

AI's going to find Joe's thesis soon.

Joe Campbell

Well, that's not what my thesis was; that was someone else's work.

Jessica Sacher

Well, thank you all so much, and good luck with everything. I can't wait to see your papers come out, and we'll be posting them. And, yeah, I can't wait to get this out.

José Penadés

Thanks. Thanks a lot.

Tiago Costa

Yeah, thank you. Thank you for the invitation. Okay, bye-bye now.

Jessica Sacher

You're welcome.

Tiago Costa

Bye.

Jessica Sacher

Bye.