Cohere 首席 AI 官 Joelle Pineau:Scaling Laws 为何会延续,以及合成数据的未来
- Scaling Laws 依然成立——不要押注它会失效。 2017年至2025年间在 Meta 从事基础 AI 研究、如今加入 Cohere 的 Pineau 表示,Scaling Laws“表现得异常稳健……过去有很多人押注它会失效”,但她不会这么做。不过,算力和数据只能带来近似线性的进步;非线性跃升来自算法(transformer、Adam、reasoning),而这恰恰是投资者最难判断“该把筹码押在哪里”的地方。
- 今天的 AI 编程,就像2015年的图像生成。 Pineau 的类比是:2015年的图像生成分辨率和构图都很差;如今的 AI 编程也会产出大量糟糕代码,其中很多最终会被丢弃,但她认为再过10年,代码质量会非常出色。其结果是:当生成变得充裕,价值会转向筛选——“策展、验证,这些工作不会消失”,机器生成内容的海量产出之上,会出现一位“首席策展艺术家”。
- “10倍”重构动摇了 Stebbings 的判断框架: Pineau 用“能否让大多数员工完成10倍的工作”反驳 David Khan 的“AI 能否替代每个职能中最差的5%”指标——她认为这“未来几年内”就能实现,部分工作甚至会达到100倍。Stebbings 随即追问,这是否削弱了 VC 的基本假设:回报必须来自人力预算转化为 AI 支出;真正的机制可能不是替代员工,而是效率提升。
- 数据正在成为上升最快的成本项。 “这是猫、这是狗”式的简单标注已经过去;如今昂贵的是拥有专业领域经验的人才,以及构建合成环境来训练 agent 的创意人才。谈到标注公司(Mercor、Surge,可能还有 Turing)时,她说:“其中一些公司5年后可能已经不存在,但让人类引导并训练 AI 系统行为这一逻辑,会一直存在。”
- 不能只把 Galacticos 买齐。 团队需要愿景、执行力和社会黏合剂;“我不认为把一群 AI 超级明星全塞进一个房间,就能让团队变得高效。”在被追问 Andrew Tull、Daniel Gross 和 Alex Wang 时,她承认,如果负担得起,确实需要“少数几个这种超级人才”,但对于数十亿美元级别的价格,“时间会给出答案”。
- Agent 开辟了新的安全战线,而大量问题仍未知: LLM 对应的幻觉风险,在 agent 身上对应的是冒充——agent 代表“并不真正授权它们的实体”采取行动,包括渗透银行系统。“关于这些系统的脆弱性,我们还有很多不了解的地方。”
- 这是一个“方差更大的泡沫”——上涨更猛,下跌也更深;只有能承受风险,AI 才是好的投资。她的逆向判断同样适用于两端:她承认自己曾经怀疑神经网络而“错得相当离谱”,也认为关于生存风险的讨论缺乏科学严谨性,并表示最想禁掉的一个 buzzword 是“existential risk”。
1. RL 是基础,但效率极低
- Pineau 从事强化学习已有20多年,也见证了它随着 reasoning model 和 agent 的兴起而进入公众视野。她在 Meta 得到的教训是:验证一个假设究竟需要多长时间。针对 Andre 在上一期节目中“RL 很糟糕”的说法,她回应:“比20年前没那么糟了。”她依然“非常看好 RL”这一概念——通过奖励进行训练“如此基础,不会消失”——但认为“仅靠开箱即用的 RL 就得到 AGI,这件事没那么乐观”。
- RL 低效的原因很明确:序列决策意味着动作树的每一条分支上都会累积错误,“就像在大海捞针”;同时,模型必须通过行动来学习,这要求建设模拟器和合成数据,而要让这些模拟器具备足够多样性,成本很高。
- 只要奖励函数可以被写出来,成本就已经下降——AlphaGo、数学、定义清晰的 reasoning 和游戏都属于这一类。但“用 RL 塑造模型行为、让它们成为社会性生物,我们完全不知道该怎么做”;这就像育儿,“你重复同一件事多少次,它们却还是做别的事……你不知道如何用数学把它写出来”。
2. Scaling Laws 稳健,算法是非线性杠杆
- 她对进步的拆解是:算力和数据带来的是近似线性的提升——更多算力、更大模型、更好表现;算法则是非线性的,transformer、Adam 优化器以及 reasoning-in-the-loop 都属于范式变化。问题在于,算法创意“可能需要很长时间才能验证”;成千上万篇论文,仍未在合适的规模、数据和超参数下被真正测试。
- 对于 Scaling 是否会延续,她说:“Scaling Laws 表现得异常稳健……它们不能单独发挥作用,我们还需要这些算法创新,但大多数时候我不会押注它失效。”
- 算法也是最难下注的方向:算力和数据可以买到,但选择算法路线“有点像强化学习——不到终点,你不知道自己选的是不是正确方向”。她的结论是:这是“最有意思、最令人沮丧、当然也是从投资者角度最困难”的变量。
- 经济层面最大的结构性问题是不可预测性:“所有人都想知道突破什么时候到来……我到底需要多少 GPU……能期待什么回报”;但这项技术根本无法提供这些答案,于是每一个数据中心、劳动力和数据决策都被迫承受风险。
3. 让劳动力提升10倍,而不是替代劳动力——Stebbings 的判断开始动摇
- 针对 David Khan 的指标——AI 能否在任何职能中替代最差的5%——Pineau 反问:“能否让大多数员工借助 AI 完成10倍的工作量?”简单替代员工“其实相当不现实”,因为人和 AI 的能力是互补的。Stebbings 认为10倍的要求更令人不安,她则回应:“我完全不认为这不现实……未来几年内就可以做到。”
- 她给出的具体案例是:多页文档的机器翻译从数小时缩短到数秒。人类仍然负责提出问题、验证结果和定义任务,但一旦任务被明确,“按下按钮,几秒钟就能得到答案,而过去这件事需要几周甚至几个月”。
- Stebbings 在节目中的即时重估值得保留:他的 VC 论点原本假设,赚钱的关键在于“人力预算转化为 AI 支出”;而 Pineau 的回答意味着,收益可能体现为效率提升,“有些工作会更难……但另一些工作会达到100倍”。分界线在于任务定义的模糊程度:机器无法处理歧义,能够被精确定义的任务会最先自动化。
- 对于 Orman 所说的——年轻人把 AI 当作操作系统,年长者把它当作下一代 Google——她有温和的不同看法:“我看到很多人更多是把它当作工具,而不是伴侣”,把 AI 当成工作生活中的瑞士军刀。
4. Cohere 的本地部署切口与企业现实检验
- Cohere(她加入还不到1个月)开发的是由企业在本地部署运行的模型。针对 Stebbings 认为推理可能占据市场95%的说法,Pineau 表示,公司需要考虑训练,但企业不必承担推理成本;她认为这会对高效模型形成结构性需求。Stebbings 追问:如果客户支付推理费用,你有什么动力提升效率?她的回答是:“我们仍处在企业采用 AI 的早期阶段,所以对客户有利的事情,对我们也有利。”
- 她离开纯学术研究的原因是:“当你需要把 AI 卖给一家企业时,你会得到真实信号,知道什么有效、什么无效。”学术基准能提供一些信号,“但让 AI 真正完成生产性工作,是另一回事”。这些来自现实世界的反馈,会反过来引导研究者在想法空间中搜索。
- 企业面临的最大障碍,是 AI 与数十年积累的信息系统整合;这也是 Cohere 聚焦本地部署和数据保密的原因。另一个障碍来自文化:“很多人觉得自己必须第一次就做对”,但技术的成熟度要求的是“探索和好奇的精神”。
5. 数据正在成为上升的成本项——从标注到构建环境
- 如果给她100亿美元,她会在人才和算力之间取得平衡——“人才太多、算力不够,就是在浪费时间”——同时拿出“一大块”预算投入数据,因为“我们经常低估数据的重要性”。Cohere 自身的算力资源“相当充足”。
- 数据变贵的原因是,猫狗标注时代已经过去。企业 AI 需要真正理解工具和业务逻辑的人;训练 agent 则需要“非常有创造力的人,来为你构建合成模拟器”,也就是把机器人模拟器的路径应用到工作流程。
- 谈到标注公司(可能是 Mercor、Surge 和 Turing),以及这是否是一个能长期存在的市场,她说:“其中一些公司5年后可能已经不存在,但让人类引导并训练 AI 系统行为这一逻辑,会一直存在。”变化的是人类需要提供什么信息。她看到的更大趋势,是“从单纯标注数据转向构建环境”。
6. 只有多样性消失的地方,合成数据才会崩溃——代码可以保存
- 合成数据是否会让模型退化,“完全取决于你如何生成它”。如果无法注入多样性——例如让 LLM 彼此对话——就会发生分布坍缩。她的完整类比是:“找一群人,把他们放到一座岛上,让他们繁衍……某个时候,遗传多样性会不断缩小。”国际象棋和围棋这类封闭世界可以生成大量合成数据,但也不是无限生成,因为世界本身是封闭的。
- 代码处在两者之间,而且位于更有利的一侧:“我可以拿几个代码仓库,把它们混合匹配,再用 LLM 进行转换。”结构加上可以注入的多样性,使代码能够进行大规模合成训练,而不会发生坍缩。
- 她对代码质量担忧的标志性判断是:现在的代码生成,就像2015年的图像生成——分辨率差、构图差,但随后一路取得巨大进步,直到大约2022年。“是的,现在生成了很多糟糕代码……但再等10年,生成代码的质量会非常出色。”
- 沿着这个类比走下去,终局是产出规模:真正重要的将变成“从海量内容中挑出高质量内容”,也就是一种筛选机制、一种编辑设计选择。Stebbings 说:“所以会出现首席策展艺术家。”Pineau 回答:“是的……策展、验证——这些工作不会消失。”当 Stebbings 指出这会消除人类之间的协作时,她回应:“这就是你获得10倍生产力提升的地方。”不过,意图和批评仍然由人类掌握;一旦设计师可以从想法直接进入数字世界,团队构成也会发生显著变化。
7. Agent 安全:冒充是新的幻觉
- Agent 正在打开一条新的安全战线,而“关于这些系统的脆弱性,我们还有很多不了解的地方”。LLM 的风险向量——prompt injection、jailbreaking——正在通过 red-teaming 被逐步理解;agent 还没有经历同等程度的审视。她的类比是:LLM 会产生幻觉,agent 则会冒充——“代表并不真正授权它们的实体采取行动,比如渗透银行系统”。安全仍是一场“猫鼠游戏”;最直接的缓解方式,是让 agent 与互联网隔离,以信息访问能力换取风险下降。
- 谁来裁定 agent 是否经过验证?她认为:“政府擅长制定大家共同遵守的标准;企业更擅长大规模构建解决方案。”她以航空业为例:政府定义规范,最终带来了令人难以置信的50年安全记录。但监管“不应走在技术前面……那会把顺序搞反”。
- 对于主权 AI,她认为在美国和中国之外开发的模型“有利于思想多样性”;但 Cohere 的“愿景不是成为一家加拿大公司”,而是一家全球公司。加拿大总部让公司更敏感地意识到,一套标准无法适用于所有市场:“你去日本、韩国,他们确实希望拥有在本国语言中运行良好的模型。”
8. Galacticos:需要少数几个,不需要整套名单
- 她的团队配方是:1至3个拥有愿景的人;一批“执行力惊人”、且“不在乎这是不是自己的想法”的人;以及社会黏合剂。她见过的失败模式是,团队里只有一种人——“把一群 AI 超级明星全塞进一个房间,却没有执行机器,也没有社会黏合剂”。此外还需要聚焦:明确北极星目标,即使这个目标随着时间推移需要改变。
- Stebbings 直接反问:如果不需要把 Galacticos 买齐,为什么这个团队还要集合 Andrew Tull、Daniel Gross 和 Alex Wang?她让步说:“确实需要少数几个这种超级人才……真正深刻理解这项技术的人数量相对有限……如果你负担得起,就应该招一些”,但不需要把整套名单都买下。至于每人30亿美元的价格是否合理,她说:“时间会给出答案。我不认为一定需要达到那个规模。”
- 对于“见过成功”所带来的溢价(可能指 Mira Murati 以100亿美元估值融资20亿美元),她认为,履历不仅意味着掌握核心配方,也意味着“曾经组建出优秀团队、打造世界级模型的成就——这里面有很多细微之处”。
- 她从 Zach 身上学到的一点是:“他从不松懈……他学习 AI 时提出的问题之深。”你可以拥有最出色的团队,但作为领导者,仍然需要深入理解这项工作。
9. 一个方差更大的泡沫,以及她放弃的判断
- 对于好泡沫与坏泡沫,她的看法是:“我把它看作一个方差更大的泡沫——上涨会更大,下跌也会很深。只要人们能够承受风险,AI 就是一项很好的投资。”Evals 是“衡量系统性能的单元测试”,是很好的指标,但不应成为优化目标:“没有客户会问,你的模型能不能赢得数学奥林匹克竞赛?”
- 她最值得借鉴的一次改变看法是:“我过去相当怀疑神经网络是否真的是机器学习的终极方案。”每一次过去的规模跃升,都曾有更好的方案出现,比如21世纪初的 SVM;“而这一次,我似乎错得相当离谱。”她的认识论是:“我的信念不强,但我对科学方法有非常强的敬意。”
- 她认为别人错在:“作为科学家,我对那些预测极端情境的人缺乏耐心……比如 AI 最终成为我们的霸主……我认为你缺乏分析这类情境所需的科学严谨性。”她最想禁掉的 buzzword 是“生存风险”——“我们不是出于恐惧,才做出最好的工作。”
- 最后有两个判断:市场确实需要高效的小模型——2019年的小模型 RoBERTa 在大语言模型狂热期间每月获得2000万次下载;她希望看到能在1至2块 GPU 上运行的模型。行业转向封闭系统和封闭访问则是“一个深刻的错误……想法会流通,人们也正在让它们流通”。如果今天投资,她会选择医疗健康和科学发现,预计“5年内就能看到真正切实的进展”。
The scaling laws have been remarkably robust. There's a lot we don't know yet in terms of the vulnerability of these systems.
Maybe you don't need to buy the Galácticos. Why do you have an Andrew Tull, a Daniel Gross, an Alex Wang, and the Galácticos assembling?
I used to be quite skeptical that neural networks were necessarily the ultimate solution to machine learning. I seem to have been quite wrong on this one.
Knowing what you know, what do you not let your children do?
Eat too much sugar. I don't have a lot of patience as a scientist for people who are predicting the extremist scenarios—the catastrophic risks of AI, where AI becomes our overlord.
1. The Rising Cost of Data
If I gave you $10 billion, what would you spend it on first?
Ready to go? Joelle, it is so great to have you in the studio. I've heard many great things from Nick, Aidan, and Shrep[?]. Thank you so much for joining me.
2. How Meta Shaped How I Think About AI Research
Thank you. Happy to be here.
You spent over 6 years at Meta, and I want to start there because it's a very transformative time and place. What are the biggest takeaways for you from that time, and how did that shape your mindset and how you think today?
I was there from 2017 to 2025, and you have to see just how much AI changed over that period of time. What we were really focused on was fundamental AI research. One thing that I've learned is how long it sometimes takes to prove out a hypothesis. We feel like AI is moving at the speed of lightning, but in fact, there are some things that just take a few years to mature—to get the right optimizer, the right compute, and the right data for that to really make a difference.
I look at where we are today, and everyone says, “It's here, it's here, it's here.” Then you actually look at what a lot of the leaders have been saying recently. Andre was saying it's not the year of the agents; it's the decade of agents. Sam is pulling back, too. Have we got over our skis, and are we actually all pulling back and realizing that time is the factor we need to rely on?
3. Challenges in Reinforcement Learning
I'll give you an example. I've been in research for a couple of decades now. I've been working on reinforcement learning for over 20 years, and suddenly everyone's been talking about reinforcement learning since the advent of reasoning models, agents, and so on. Sometimes you have to be a little patient with these ideas, and the right algorithmic tweak, the right context, or the right problem domains just opens up the magic.
I was listening to Andre yesterday, and he said on this show that reinforcement learning is terrible.
Less terrible than 20 years ago.
Have we overinvested in RL-based methods at the expense of more scalable alternatives?
I'm still super bullish on RL. The concept itself is so fundamental: this idea of training a system through a system of rewards, indicating what's valuable and what's not valuable through numerical values. That is so fundamental, and it's not going away. Where we're maybe getting a little bit ahead is thinking that RL out of the box is going to give us AGI. That part is a lot less likely.
If you look at the curve of progress, RL is terribly inefficient, and the amount of signal you need to really shape the behavior of a model is far from where we are today. We'll need to figure out how to deal with this learning-efficiency problem.
You're probably thinking, “What did I get myself in for?” I don't blame you. I ask questions that I think everyone else thinks, but I'm not afraid to say I don't know. Why is RL so inefficient?
4. AI in Enterprise: Efficiency and Adoption
You're going to get me on a deep topic. There are a few reasons. One is the fact that RL is about sequential decision-making. Think about starting at a point: you need to figure out what you're going to do next, and you might pick the right side of the branch or the wrong side of the branch, and then the road keeps splitting.
Every time you make a mistake, it compounds through the length of the series of actions you're making. That means the amount of error you can make can be very large, and getting it right is quite difficult. Sometimes people compare it to finding a needle in a haystack—finding the right solution in RL.
The other part that's hard is the fact that, to train the system, to train the models, you essentially have to take actions to learn. You can't learn from static data. You can learn some things from static data, but to get the right policy, you need to test it out. That means you need a simulator, you need to get the synthetic data, and all of that can be really expensive. We have difficulty getting a variety of environments and simulations to test RL.
When we look at the cost curve for RL—you said you've been working on it for 20 years—have we seen that dramatically come down? Will we see it continue to dramatically come down, or is it fundamentally an expensive method of training?
It's come down, especially in domains where we have good reward functions. The place where most people started hearing about RL was around the time of AlphaGo. The game of Go was one of the goals for AI. Many people thought we were still a decade away from having machines play Go at the level of humans. Then a team from DeepMind went off, played against the world champion, and showed that RL could basically do it.
In cases where we clearly know what the goal is and can write down the reward function precisely, we're good. We can make a ton of progress. That's why you're seeing progress in mathematics, very well-defined reasoning tasks, games, and these kinds of things.
Using RL to shape the behavior of models and get them to be social creatures—we have no idea how to do that. I don't know if you have children, but shaping their behavior is difficult. The number of times you can repeat the same thing and they still do something else is remarkable. You don't know how to write that out mathematically, and that's where I think we're still in for some hard work.
Okay, so we still have some hard work. When we look at the training-versus-inference market today, we've put so much weight on training so far, and it's been incredibly costly and expensive. Then I hear everyone say, “Actually, inference is 95% of the market. That's where it's all going, and that's where Nvidia will make most of its money.” How should we think about the cost curve applied to training versus inference, and where it sits today?
I think there are a lot of different variants. If you'll allow me, I'll pivot to where I'm going with Cohere. I joined Cohere less than a month ago, and it's a super exciting company.
One of the things that Cohere is doing is developing AI models that run on-premises. That means enterprises bring them in and run them locally, so the company has to worry about training the models. Obviously, we want world-class models for the needs of the enterprise, but the enterprise doesn't have to worry about the inference cost. The client's customers have to figure out what's the right way for them to digest the AI.
That means there's a lot of motivation to have very efficient models so that they can run really efficiently on-premises. We get caught up in one paradigm, but there are other paradigms as well.
If they're the ones paying for the inference, is there less incentive to make the models efficient? You're not the one paying for it. If you're the one paying for the inference costs, you want it to be as efficient as possible because it's your dollar going to that. But if it's IBM's dollar, I would love it to be efficient, but we're not paying for it.
We're still in the early days of AI adoption in the enterprise, so what's good for the client is good for us.
5. Is It Possible To Be Capital Efficient in AI
Totally. What's the biggest challenge about capital-efficient AI today? I know that sounds strange when you look at the economics, so to speak. What's the biggest challenge?
There are a lot of challenges today. In terms of the economics of AI, I think one of the biggest challenges is the fact that it's very hard to have predictability. Everyone wants to know when we're going to hit the breakthrough. Everyone wants to know how many GPUs they actually need. Everyone wants to know what return they can expect.
There's just a lot of uncertainty built into the system. A lot of that is because there's a lot we don't know about this technology. That means we have to take on quite a bit of risk when we're building out a data center, building out a workforce, or trying to figure out how much data to curate. That makes it difficult for a lot of people. People want answers, and this is a world where we don't have that level of predictability compared to other industries.
Does progression happen in a linear fashion, or does it happen in step functions, like AlphaGo or DeepSeek, which, depending on what you believe, suggests a lot of efficiency in terms of model improvement? Is it a step function, or is it linear?
I tend to decompose the different ingredients that lead to progress. People often talk about the algorithms, the data, and the compute. I think, in general, compute and data have a more linear effect on progress. You build more compute, you run bigger models, and you can typically get better performance. You feed in more data; it's not just quantity—you need to worry about quality and diversity as well—but roughly, it's more linear-ish with respect to the data.
The algorithms are the ones that have the nonlinear effect. You can explore lots of ideas, and then something like the Transformer comes along and just changes the paradigm. It's not just the Transformer: on the optimization side, suddenly we hit upon Adam, which is a technique to optimize your model, and it changes the paradigm. With reasoning, suddenly we start thinking about how to put that in the loop, and it changes the paradigm. Those ideas tend to have a nonlinear effect.
The challenge with these algorithmic ideas, though, is that they may take a long time to prove themselves out. The paper can be sitting out there, and there are thousands of papers coming out. The idea is sitting out there, and we may not think to try it with the right data, at the right scale, and with the right combination of hyperparameters, so you don't notice that effect for a while. It's hard to predict, and it's nonlinear more on the algorithmic side than on the data, compute, talent, or other sides.
With respect to Google, Transformers were obviously birthed at Google and sat as papers for maybe a couple of years. When you mentioned compute, algorithms, and data, if we just go through them to understand, everyone suggests that there are 2 different worlds. Scaling laws exist—just throw more compute at it. When you look at data-center investment and the desirability of compute, and then you have GPT-5 seemingly focusing on efficiency and other signals, do scaling laws play out from here, and if so, for how long?
The scaling laws have been remarkably robust. They don't play out exactly as we expect, but they've still been remarkably robust. Lots of people have bet against scaling laws in the past, and overall, we've seen a pretty robust effect. They don't work alone; we also need these algorithmic innovations. Most of the time, I wouldn't bet against it.
On the algorithm side, is that the hardest thing to innovate on? You could think about buying more compute; it might be hard, but you can buy more compute. With data, there are different ways to get it, whether it's synthetic or human. Are algorithms the hardest thing to innovate on?
It's certainly the most creative work to be done. The space of ideas is so wide that I would say it's the hardest in the sense that, as a researcher at heart, you can move in so many different directions. Picking the right one is something you don't know until you get there.
It's a little bit like reinforcement learning. In that sense, I think it's the most interesting one, the most frustrating one, and certainly the most difficult one from an investor's point of view, because you don't know where to put your chips.
Speaking of knowing where to put your chips, and moving from purely a research lens with Meta to now also building product, is there ever this inherent conflict between intellectually interesting research and the need to productize and monetize? How do you think about that?
One of the reasons I'm really excited to be joining Cohere is that we're at a stage where AI is really starting to be useful. Maybe not as useful as people think it is, but we are there. By working on AI that's going into enterprises, I feel we're going to get such an interesting signal of what works and what doesn't work.
We keep talking about AGI and AI for the masses, but when you need to sell AI to a business, you get a real signal of what works and what doesn't work. That's what I'm most curious to see. We've been using academic benchmarks for many years. You get some signal, but it's not the same as getting this to do productive work.
I'm curious to learn from that. We're going to get new types of data and, I think, a lot of insights that are then going to drive the research ideas. That's the other thing to think through when you have a large space of ideas to explore: getting that feedback signal from the real world is super useful to guide you through that search of ideas.
I just had a great chat with David Khan at Sequoia, who said that he thinks a good barometer for utility value within enterprises is: does it have the ability to replace the work of your bottom 5% in any category? He says we overestimate a lot.
Can it replace the bottom 5% in any function? If it can, that's a very meaningful improvement. Do you think that's a good barometer, and how would you advise an enterprise on whether something is useful as a yardstick?
I prefer, in terms of a barometer of productivity, something a little bit different, which is to say: can most of your employees do 10x the amount of work with AI versus on their own? That, to me, is actually a better barometer. I think humans and AI have very complementary abilities, so to just flat-out replace a portion of your workforce is actually pretty unrealistic. Some may try, and some may be slowing down their hiring, but I actually think—
Respectfully, I think 10x-ing your work feels more unrealistic. Is that not a bigger ask? I'm almost more intimidated by 10x-ing my work.
I don't think that's unrealistic at all.
Wow.
Yes, within a timeline that is the next couple of years.
Yes. How does that actually shape out, then?
I think you have to identify very concretely the types of work that you are delivering. We're starting to see Hollywood-quality productions being made in a matter of hours. We're seeing, to take a super-concrete case, machine translation. If humans are doing the translation compared to machines doing it, you go from hours to seconds on long-form text and multipage documents.
For a lot of work, it's not that AI can do all of the work. Humans still need to ask the right question, verify the information, and shape the tasks. But once the task is well-defined, the parameters are clear, and all the design considerations are fed into the prompt, you press the button and you've got an answer in seconds for something that used to take weeks or months.
I completely hear you and understand that. I'm just trying to reevaluate a belief that I've had for the last few months. I'm a venture investor, and for all of us to make money, we need to see the transition from a human labor budget to AI spend. It's with that transition that we obviously see the TAM massively increase and make a lot of money.
But when I hear you say that, I suddenly question that assumption. The barometer for whether we make money is that you're suggesting we don't replace the human labor budget; it just makes us 10x more efficient. Is that correct?
Yes. There's a lot of nuance to all of that. Some work will be harder to get that same level of efficiency gain, whereas with other work you'll see 100x in terms of efficiency gain. But I do think that, for a lot of the work that's happening right now, that's absolutely feasible.
Where do you think the efficiency gains are most tangible? It goes back a little bit to this notion of what the tasks are that we can specify.
In any case where we can be very precise about what a great result looks like, we'll be able to make that task automatic much more easily than tasks that are much more nuanced and have a lot of complexity.
So it's ambiguity.
Ambiguity in the specification of the task is what's hard for our machines.
How have you seen enterprise reaction to this? There's fear from workers sometimes, excitement from leaders, and apathy sometimes. How have you seen and measured enterprise response?
A lot of the workforce can be reasonably fearful about job displacement. There are also a lot of individuals who have an instinctive reaction to change, and change can be hard for a lot of people. We're seeing a lot of change in a very short time span, so I think there's also a generational effect. For some generations, that change is more jarring.
I have teenagers and young adults at home. For them, they're just natives. They're going to grow up with that technology in a different way than some of the older generations.
It's interesting. Speaking of children at home and how they engage with it, Sam Orman said that young people engage with it as an OS to the world, and AI is that companion for them, while older people use it as a next-gen Google. Do you agree with that, and do you see that in your work?
I see a lot of people using it as a tool more than as a companion. People have this Swiss Army knife in their work life all of a sudden that can be super helpful, but that's really most of what I see.
Totally get you. What are enterprises' biggest challenges with AI adoption at scale?
For many enterprises, one of the challenges is making sure that the AI comes in and can be integrated into their workflows, their processes, and their information.
And so the challenge is to deploy in a way that allows them to exploit all of the information systems that they already have. Some of them have accumulated these over decades. That's some of the work that remains to be done.
Integration with existing systems and data flows.
That's, of course, something we see a lot at Cohere because we do on-premises deployments. One of the things we focus on the most is data confidentiality and security, so that enterprises can exploit all of that information. That's top of mind for us.
But it's also a huge opportunity. I would say there's a big interest in that, but making sure to get that compatibility, I think, is a challenge. In many cases, change is hardest for people, and so you have to get them curious about using the technology. Many people feel they have to get it right the first time, and I really think a spirit of exploration and curiosity is much better suited to the phase of maturity of the technology that we have today.
We don't have all the answers about how it should be used. That's going to come from people in the field.
6. Security Concerns with AI Agents
Security is a topic that we quite often glaze over, especially when investing in application-layer AI tools. What does no one know about AI security that people should know?
With respect to AI security, I think there's a new front that's opening up with the development of agents. Frankly, there's a lot we don't know yet in terms of the vulnerability of these systems.
With LLMs, we're starting to get a better understanding. We've had quite a few red-teaming exercises and jailbreaking, and so on. People have identified different risk vectors—prompt injections, things like that—which are vectors for malicious actors to interfere with a system.
With AI agents, we haven't seen that. One of the features of computer security in general is that it's often a bit of a cat-and-mouse game. There's a lot of ingenuity in terms of breaking into systems, and then you need a lot of ingenuity in terms of building defenses, so we just have to stay very active in that sense.
What are the potential vulnerabilities in an agent world?
In terms of agents, we worry a lot about hallucinations in LLMs. The parallel in agents is impersonation: agents that come along and are essentially impersonating entities they don't legitimately represent and, in doing so, taking actions on behalf of those entities. That could involve infiltrating banking systems and so on.
I do think we have to be quite lucid about this, develop standards toward it, and develop ways to test for that very rigorously. There are ways to reduce that risk drastically. You can run your agent completely cut off from the web, and you're reducing your risk exposure significantly. But then you lose access to some information.
Depending on your use case, depending on what you actually need, there are different solutions that may be appropriate.
Totally get you. That's a really hard one because then verification becomes the most important thing. But then who's the arbiter of verification? Is it governments? Is it companies?
Mm-hmm.
How does one think about that? Who says you're a valid agent versus an invalid agent?
Governments can be good for defining standards on which we all agree. Companies are much better at building the solutions at scale and deploying them.
Do you think governments are good at setting the standards when you look at AI and where we're at? And then when you look at the sophistication levels of government programs or decision-makers, with respect, they're just a little bit behind. Do you think they are actually equipped?
I don't think you should look at where governments are in terms of AI regulation necessarily. AI as a field is so incredibly young and fast-moving, and by nature—and there's some good in this—governments are moving a little bit more cautiously and usually need to benefit from our knowledge to make good policies.
I do think you can look at other fields in terms of regulation. You look at aviation: the security record for aviation today, compared to where we were 50 years ago, is just incredible. Governments have played a role in defining that in terms of standards and what the norms are, and so on.
I'm quite hopeful. I'm an optimist about this—maybe it's my Canadian side—that governments can play a useful role. In many cases, clear standards actually mean reducing uncertainty for a lot of companies in this space.
But we shouldn't expect that to be ahead of the technology. I think that would be the wrong order of things, in some sense. We need to develop that technology with enough creative space, and we need to learn fast. Then we need to develop the right guardrails for that technology from the real learnings we have.
We mentioned governments and their role. When I had Nick Frosst on the show, he was saying that there are actually benefits to not being an American company, given some geopolitical challenges sometimes. I'm just intrigued: do you think we will have these kinds of sovereign models for each geography? We have Mistral in France, and Cohere, obviously, was founded in Canada, but I know you've got global headquarters. Do you think we will have these sovereign models and regionalized winners?
I do think it's healthy that there are models getting built in different places around the world, not just in the US and China right now. I think this is healthy in terms of diversity of thought. I think it's healthy in terms of having a greater number of people with access to technology.
For Cohere, the vision isn't to be a Canadian company. The vision is to be a global AI company. We have headquarters in Toronto, and we have teams distributed around the world. We have a great team here in London, as well as in the US, in France, and other places.
Having the ability to deploy models that operate across the world is going to be an important part of the strategy for Cohere. I think there's a great opportunity.
What being headquartered in Canada gives us is a sensitivity to the fact that it's not always a one-size-fits-all solution. I go back to the research we've done. We've done leading work in terms of multilingual models, and it turns out it matters. You go to Japan, you go to Korea, and they do want models that work well in their language.
7. Can Zuck Win By Buying The Superstars of AI
People in the workforce are still operating in the language of the country. Having a company that's attuned to that, that values that internationalization of models, is actually important in the global market.
Totally get that. On the team-building side, obviously Canada has great talent. You mentioned some in London as well. What have been your biggest lessons and observations on team-building in this talent frenzy that we're in? Also, how do you analyze that?
One of the things that's important when you're building a team for AI is that you need people who have vision, who have a sense of what we can create, because we're in a space where there's so much innovation that is still needed. You need an ingredient of vision. That can be 1, 2, or 3 people who bring that ingredient of vision.
You need people who have amazing execution muscle. They don't care that it's their idea. They care that, if the team agrees on an idea, they're just going to push it and get it done. They're going to build the system and run the experiments. They have the technical rigor to execute.
Then you need people who keep the team together, who have a sense of who needs what to operate well, and who are that social glue. Humans are still social beings, and that social glue in a team matters a lot.
Where I've seen it fail is when you have just one type of person inside the team. I don't think it becomes that productive to put a bunch of AI superstars all together in a room without the execution machine and without the social glue. I don't think you necessarily get the same result.
I'm a big believer in building teams with diverse, complementary skill sets.
So you can't just buy the Galácticos?
I don't think you need to. You really have to be thoughtful about putting people in a group.
The other thing that helps a lot is for the team to have focus. If it goes in all sorts of different directions, you'll lose that power that you get from people working together. Having a lot of clarity—what's the North Star, what's the goal, where are we going—even if over time that needs to change, that level of clarity is required for everyone to be working in the same direction.
Can I be so blunt? If you don't need to buy the Galácticos, why do you have an Andrew Tull, a Daniel Gross, an Alex Wang, and the Galácticos assembling? Is that wrong?
You do need a few of these uber-talents on the team. There's a relatively small number of people who just understand this technology very deeply. You do need some of this talent, and if you can afford it, you should get some of that talent.
But you don't need all of your team to be that way. You need a team with complementary skills as well.
Does that create a good team? If I gave you $10 billion to go build a team and you could buy a couple of these luxury star players...
I feel like it's Top Trumps cards for sports teams: you can buy a couple. Does that create a good team when one is a $3 billion person and the rest are just average $50 million people?
Yeah, I wouldn't say no if someone offers me the opportunity to hire. There's definitely some really talented people in the field, and they deserve to be fairly compensated. This technology is going to probably make a lot of people very rich and have major effects in terms of society, and so we should be rewarding the talent.
But I'd be very thoughtful about what teams I put together and how they work together, rather than just hiring a roster of superstars without being thoughtful about how they're going to work together.
It's so funny, because of the impact that you can have in these teams, the multibillion-dollar price tags that you see can even be justified.
Time will tell. I don't think it's necessarily needed to go at that scale, but time will tell.
If I gave you $10 billion, what would you spend it on first?
One of the things you need is a balance between talent and compute. If you have too much talent and not enough compute, you're wasting your time. So usually, there's an equilibrium between those 2 pieces.
I think we often underestimate the importance of data, and data is getting more and more expensive, so I would certainly spend a good chunk of it on data as well.
So many things to unpack there. Do you feel you have sufficient compute today?
I think we are reasonably well-resourced in terms of compute and in building the models that we want to build. Yeah, access is not a massive problem.
No. Okay, great. Why is data becoming more expensive?
Data comes in different forms. On the one hand, the days of having data labelers who can say, “This is a cat and this is a dog,” are somewhat over. The easy tasks are things AI can do, so we're getting into a space where we need more specialized tasks.
Imagine you're building AI for enterprise. There's a particular business logic, and you need to make sure that you're catching the errors. You're going to need someone with a deeper understanding of the tools. So that's more expensive talent to come in and actually prepare the data.
There's also a lot of data that's synthetic. When you're building agents, you need to build environments. To build environments, you need some pretty creative folks who are going to build you synthetic simulators.
We've seen this on the robot side for many years, with people building robot simulators. Now you're building AI for enterprise, so you need to think about how you're going to simulate these work processes in a reasonably realistic way that the AI can train on. That generation of environments, benchmarks, and dynamic domains can be pretty expensive too.
When you look at the expansive data, and then you said, “Cat, dog, lamppost,” you've got these CAPTCHAs.
I get them wrong.
I legitimately get them wrong. I'm like, “Jesus, it's getting harder.”
They're getting so hard. They are. The other day, I called up my CFO and said, “I failed the reCAPTCHA. I'm so sorry. I'll try again in half an hour.”
Please let my AI agent answer that one for me. It's embarrassing. But the question that I have is, when you look at (likely Mercor), (likely Surge), and (likely Turing), how do you evaluate that market, which is providing a lot of that talent?
Is that an ongoing, enduring market, or is that just, “Hey, for the next 3 to 5 years, we'll need it in the training phase of these models, but I don't know what it looks like beyond that”?
I don't think it's a phase, in the sense that I do think this partnership—we'll call it—between humans and machines, where humans provide guidance to machines, is here for a long time. What will change is the nature of the information that the AI provides versus the information that the humans must provide as a complement.
Some of these firms may not be around in 5 years, but this notion of having humans guide and train the behavior of AI systems is here to stay.
It's super interesting. As an investor in one of them, I see all of them converge around needing to do 3 things now. They used to just be talent acquisition: “Oh, we'll get you these people.” Now they're like, “We'll get you these people and we'll get you high-quality data that you can really use.”
And now it's like, “Oh, shit, we need this third pillar, which is, we'll also help you implement that data into your models, do training, and help you with benchmarking.” Now they need all 3.
Are you seeing that third one, where it's the implementation of their data as well? They don't just hand it over the fence.
8. Synthetic Data and Model Degradation
There's definitely some of that happening. I think, for me, the even bigger trend we're seeing is the move from just labeling data to crafting environments to produce new tasks.
You said about synthetic data and that also being a very important segment to consider. Do you get model degradation when you get this kind of reinforcing loop of models learning on synthetic data, which creates more synthetic data? Does it actually degrade, or does it improve?
It really depends on how you're generating your synthetic data. In some domains—if you think of images, language, or LLMs talking to each other—at some point, you definitely get degradation. That degradation is due to essentially a loss of diversity in your data.
You can make an analogy: you take a bunch of people, put them on an island, and let them reproduce. At some point, the genetic diversity is going to keep shrinking. You get a reasonably similar phenomenon with models because you're not injecting diversity into the data. So there are domains where a lack of diversity means you get a collapse of the distribution.
There are other domains where you don't need diversity. If you think of playing chess or playing Go, these kinds of games, we know exactly how to generate board configurations. We can generate tons of synthetic data—not endless, because it's a closed world, but still tons of synthetic data—and through that, learn for a long time.
Then there are domains that are sort of in between. If I think of coding, we can generate synthetic code. You take normal code, and we know how to inject diversity into the code: I can take a couple of repositories, mix and match, apply an LLM to transform it, and so there's a way to generate synthetic data.
The language is predictable enough, and there's enough structure that I also know how to inject diversity so that you don't get that collapse. The hope is that, especially in these domains, we can use a lot more synthetic data and do it without suffering from degradation of performance.
9. Why AI Coding is Akin to Image Generation in 2015
Do you worry that we are creating a world with just much worse code? A lot of people are concerned about the quality of the code that's being output and about how we're relying on it so haphazardly. Do you worry about that?
Let me make an analogy in terms of the quality of generation. You ask about code generation, but let me take you back to 2015 and image generation. I don't know if you have it in your mind, but the quality of the images that were generated—we had image-generation models in 2015—was really bad. The resolution was bad, the composition was bad, and so on.
From 2015 to about 2022 or so, we saw huge progress in terms of the quality of image generation. So if you think of code generation, right now we're in the phase we were in for image generation 10 years ago. Yes, there's a lot of bad code that's getting generated. There's a lot of code that will get thrown away. But wait another 10 years, and I think the quality of the code that's produced is going to be excellent.
What will the developer world look like in 10 years when that is the case?
Well, if I carry my analogy further, I don't know if it's a reassuring scenario, because if we look at where we are today in terms of image generation, the volume of images getting generated is huge. What matters now is picking the quality out of the volume.
If I fast-forward 10 years on code generation, when we have the ability to generate a ton of code to do a ton of different things, we're going to need some selection mechanism to decide what code we actually want, where there's actually value. That's going to come.
There's still going to be some sort of editorial design choice. Someone needs to decide, of all the code we can generate, what's the code we want to generate? What do we need to be running in terms of our digital world?
So it's like a chief curation artist.
Yes.
With AI generation, curation doesn't go away. Curation, verification—this is work that doesn't go away.
Does the structure of teams fundamentally change, then? It's funny, kind of playing that back to you, and also playing back to you what you said earlier about the human and AI: if that is the case, there's not much of a partnership, is there, between humans and AI? It's a chief curation person sitting on top of a huge amount of artificially created code.
Well, that's your 10x productivity improvement there.
It is. You're ticking that box. It removes the human.
You still need people with intent. That's one thing that you need to decide: what do you want to build, and what purpose does it serve?
And so that intent is still there. That role of critique is still there. The team composition does change significantly once you suddenly have designers who, in their hands, have amazing tools to go directly from the ideas in their heads to the digital world, maybe eventually to the physical world. That equation definitely changes.
Do you think prompts, and the way that we interact today with prompts and with ChatGPT, are the enduring interface for human engagement with AI?
It’s awfully limited. Prompts can mean a few different things, but the idea of typing in a box, to me, is very limited, and we’re already going to break out of that box. We’re seeing a lot of cases where voice is a much more natural interface. I do expect we’ll see gesture, eye gaze, and these kinds of much more multimodal ways to interact with AI rather than just sticking in that prompt box.
But language is incredibly powerful. If you think of a prompt as language, as a way to express ideas and communicate with a machine, that’s a powerful paradigm. As humans, so much of our communication is based on language. I don’t think we’re going to move away from that because it encodes information. Language and words are symbols that encode so much information so efficiently, and I don’t think we’re close to getting away from that.
It’s funny, this conversation has changed a lot of previously held assumptions for me. When you think about what you did believe that you’ve now changed your mind on, what’s most prescient? I’m genuinely curious to know.
I’m a scientist who is happy to be proven wrong at any time, as long as there’s new evidence. Other scientists are much more likely to hold on to very strong convictions. I have weak conviction, but very strong respect for the scientific method and rigor—experimental rigor and theoretical rigor as well.
There are a ton of things. I used to be quite skeptical that neural networks were necessarily the ultimate solution to machine learning. I’d seen enough cycles of neural networks peaking and then becoming less useful. I used to think that every time you changed the scale of the data—from hundreds of examples to thousands, to hundreds of thousands, to millions of examples—neural networks were the first thing we tried because they’re a universal function approximator, and then something else would come out that was better. That was true for the previous generations. Some of you may remember support vector machines as being better than neural networks in the early 2000s.
I seem to be quite wrong on this one. Neural networks seem to be here to stay, and the ability to do backpropagation and gradient descent and all that seems to be a really powerful way to learn.
What does everyone else believe quite strongly that you think they’re quite wrong on?
I don’t have a lot of patience, as a scientist, for people who are predicting extreme scenarios—whether it’s the catastrophic risks of AI or the winner-takes-all, AI-becomes-our-overlord kind of scenario. I don’t have a lot of patience for that. I wouldn’t say it’s necessarily widespread, but I think you lack scientific rigor to analyze these kinds of scenarios.
I’m much more pragmatic and grounded. I’m pro-innovation. I’m excited to see where AI is going and the problems it can solve, but I’m not so interested in going around and making up science-fiction scenarios.
You’ve been on the most incredible journey. You said there about image generation in 2015 and partly how much it’s improved. We’re seeing this unbelievable capital supply go into the space in a way that we haven’t seen for many years. Is it a good bubble, where we’re getting incredible improvements and fundamentally advancing technology, or is it a bad bubble, where costs are becoming too exorbitant, teams are too difficult to build, and compute is too difficult? Is it a good bubble or a bad bubble?
I think about it as a bubble with bigger variance. The upswing is going to be bigger, and there are going to be big downswings as well. There’s a lot of variance in the system right now.
As long as people have a tolerance for risk, I think AI is a great investment. We should continue to support risk-taking, new enterprises, and new ideas. There are a ton of exciting new startups being created, and we should continue to support them. You just have to be tolerant of risk.
I’ve had some people on the show suggest that evals are, to put it delicately, bullshit, and that they don’t actually mean anything anymore. Humanity’s Last Exam—what does that really even mean? We have these new tests that come up, and leaderboards—what is this? Is that fair, or do you think they actually serve a very effective utility to the ecosystem?
I do think they’re really good indicators. You do need to take evaluation seriously in terms of knowledge, but you shouldn’t take it seriously in terms of the ultimate goals. There are lots of different benchmarks and so on. You have to decide: What type of model are you building? What are the characteristics of your system? Then think of evaluations as unit tests for the performance of your system.
Software engineers will know what that is, right? You run through that evaluation, and that gives you a signal of how the system is doing in a particular dimension. But as we’re building systems that are more and more general, you don’t optimize for these, right? We build AI systems that go into enterprises. None of our clients ask, “Are you able to win the Math Olympiad with this model?” That’s not what they care about. They care about bringing value to their business.
We’re curious to know how well we do on math problems because it can be predictive of behavior on other things, but you don’t obsess over specific benchmarks. You look at the return on investment in terms of what you’re trying to build.
Can I ask—we mentioned access for enterprises. Enterprises have money, and that’s a great luxury in a lot of cases. Research institutes and universities often don’t. With these bubble-like tendencies, people with money are able to afford the compute and the talent. Are we seeing this lack of access or democratization for great educational institutions that now can’t afford to compete in this new world?
Certainly, a lot of universities have a lot fewer resources than companies today. That’s not completely new. When I joined Meta in 2017, one of the reasons I did that was because I could already see the disparity in terms of access to compute, and I was really curious to see how you could do research with a lot more compute.
There’s still amazing research being done in universities. You go to the major international conferences—NeurIPS, ICML, and others—and often the best paper awards are actually won by researchers from universities. There are a lot of good ideas that you need to test out at small scale. In a university, you have a lot more freedom to pick pretty risky ideas at a small scale. No one is asking you to justify your research in ways that often happen in companies.
I think they play different roles in the ecosystem. What’s especially good is that talent flows between them. University students come in, do internships, and take jobs at companies. We’ve also seen a movement of people coming out of these large companies, going back to universities, teaching, and sharing with the next generation what they’ve learned.
How important is it to have seen success, and how valuable does that make you? When you look at people like (likely Mira Murati) raising $2 billion out of $10 billion, it’s like, “Well, no one’s seen the success that she’s seen with OpenAI, so it’s valid.” Help me out as an investor. Is it valid to place that much of a premium on access to people who have seen it at that level, or is that slightly overpricing it?
In many cases, when it comes to deciding where to invest very early on, when you don’t have tangible information, you look at people’s track record. There’s a part of that that’s about what they’ve learned in terms of the core recipe. But the other thing is also the achievement of having put together amazing teams who are building world-class models, and there’s a lot of subtlety to that.
I think both of these ingredients are important to consider.
10. If Joelle Was a VC Where Would She Invest?
If you were investing today and you were joining my team, which category would you most like to invest in—security, generative AI, compliance, you name it?
There are a lot of verticals, whether healthcare or scientific discovery, that I think have incredible promise. We’re going to see real, tangible progress within 5 years that is going to completely change the face of what we can do. That’s probably where I’d push.
11. Quick-Fire Round: Lessons from Zuck, Biggest Mindset Shift
That’s very exciting on the healthcare front in particular, when you think about that timeline as well. I’d love to do a quick-fire round with you, if that’s okay. I’ll say a short statement: What would you most like to do but, because of technical or financial limitations, you’re not able to?
I’m super keen to figure out how we build societies of AI agents. We’re doing it implicitly, but how do we look at populations of AI agents interacting together and have a sandbox for doing that? Maybe that’s something I’ll get to do.
Is it a lack of time, resources, or something else? There's just a ton of different things to do, but I'm keen to see what happens there. When you think about that ecosystem of agents, you have children.
Yes.
How does AI impact social friendship and connection? Do you think
It definitely does. There's a sense that we spend a lot of our time in the digital world. When I look at 2 of my children, they spend a lot of time in the digital world playing online games with their friends. It's still very social. There must be some AI, there's the digital platform, but it's still a very social experience.
Others have more individual experiences. There's definitely a shift in the time we spend towards that platform where we go look for that social element.
Knowing what you know, what do you not let your children do?
Eat too much sugar.
Totally. That's the physical diet. Completely agree with that. Is there a technical diet?
I spend some time discussing settings. You get an Instagram account. Great, you can have an Instagram account, but what are the settings on that account? Making sure they understand.
I mean, they'll go and change them, won't they? We're going to discuss settings. I know that was not a popular one. Do they listen?
The thing with children is you don't know until later.
Do you limit screen time?
I spent a lot of energy, especially in their younger years, limiting screen time. My kids did not have a cell phone until they were 14 or 15.
Did you see Adolescence?
I have not.
Okay. Watch it. It's fascinating. Basically, a little boy goes up to his bedroom and gets lost down rabbit holes of TikTok and Reddit.
It does not turn out well.
I've heard about it. I just haven't had time to sit down and watch it. Do you worry about the loneliness pandemic and the mental health crisis that we have?
I do worry a lot in general about making sure that people are mentally healthy. I think we have to be careful about taking shortcuts and saying, because suddenly we have certain platforms, we have AI, and so on, that is causing mental illness.
There are a number of people who are suffering, and they deserve to have good answers to the situation. They deserve—we deserve—to find real solutions to that. There's a lot of people looking for shortcuts and short answers, but I think more research into that is definitely warranted.
What's your biggest lesson from working with Zach?
He is incredibly deep into understanding the work. He does not coast. When he started getting into AI, the depth of the questions that he'd ask was incredible. He just gets really interested in the topic and goes super deep, and that then informs everything he does after.
You can have the most amazing team, but as a leader, you need to go deep and understand the work.
Did you see him change?
As anyone gets more knowledgeable about a topic, they get more decisive. There's a phase where you're really learning and trying to understand, and there's a phase where you understand a lot of things and then make your decisions faster. So certainly, that shift happened.
What 1 AI buzzword would you ban if you had a magic wand?
Existential risk.
Why?
Because it just makes people afraid. It's not out of fear that we do our best work and make good decisions.
Do you find the cost of talent acquisition prohibitive?
Talent is costly. Talented people deserve to be paid well. Someone coming in just because of money rarely is going to be the right person. But you do need to compensate people fairly.
Totally get that. Final one. What are you most excited for? You don't like existential risk. I don't like doomsday planning. When you think about the positivity that can come, what are you most excited for when you look forward to the next 3 to 5 years?
I do think some of the work in terms of AI for scientific discovery is going to be pretty fascinating to see, just in terms of the doors it's going to open up—the ability to explore the combinatorial space of solutions. So I'm curious about that. And then I'm super curious to see: how can we actually make our models more efficient?
There are larger and larger and larger models. No one wants to run these models. I spent a lot of my career building open-source models. I'll give you 1 example. We were in the frenzy of large language models, and I pulled the stats on the most downloaded models of last month.
We had a model like RoBERTa from 2019, a small language model, getting 20 million downloads a month. People want efficient models that they can use, that they can run. So I'm also super keen to see what we're going to be able to do at a scale that runs on 1 or 2 GPUs.
Final, final one. You said there about being open. We seem to be reverting to a closed world now. Is that the world we should predict and plan on?
That's a mistake. That's a deep mistake. I will continue to believe that, especially for research, the ideas need to circulate, and this thought that you can just close us down was absolutely false. People are circulating.
Do you not think we are, though, moving into that world? Everyone seems to be closing systems and closing access.
There are definitely a number of places where people are closing down access. I don't think that is going to be effective. Ideas will circulate, and I also think it's a mistake from a point of view of fostering innovation.
This has been such a joy. I've learned so much from this conversation. Thank you so much for putting up with my very basic questions, but I've loved having you on the show.
My pleasure. Thank you.