Cohere 创始人 Nick Frosst:如何与 OpenAI、Anthropic 竞争,以及 Sam Altman 给 AI 帮的倒忙
- Nick Frosst 的核心判断是,AGI 时间表和生存威胁叙事具有误导性,并不能证明 AGI 即将到来。他表示:「我不认为 Sam Altman 反复谈论 AGI 已经近在眼前,是在为世界做好事……他现在已经做出过几次错误预测,而且在当时显然就是错的。」Altman 周游世界、警告各国领导人 AI 存在生存性威胁,在他看来「在学术上不诚实,而且我认为这也给他所热爱的技术帮了倒忙」。Harry 更尖锐的说法是,末日叙事与融资需求相关;Frosst 有保留地回应:「你指出的相关性确实存在」,但他不知道这是否就是策略,并补充:「我们是一家由风险资本支持的公司,需要融资。但我们没有这么说。」
- Frosst 不相信单纯增加算力就能保证 AI 持续进步。Harry 问 GPT-5 比 GPT-4 好多少,Frosst 回答:「我其实觉得它更差」;Harry 随后说,这说明了把更多算力投入问题本身意味着什么。Frosst 将产品问题归因于更慢、更笨重的模型自动选择机制。他对 AGI 的定义是「一台你会像对待人一样对待的电脑」,而当前语言模型技术尚未达到这一点:「我不认为这项技术能把我们带到那里。」
- Cohere 的差异化定位是资本效率加企业市场聚焦。全球真正做 LLM 的公司不到 20 家;融资讨论提到约 6 亿美元融资和 68 亿美元估值。Cohere 的投入「确实比其他公司少了几个数量级」,并将 Command A 系列模型训练到可以在 2 块 GPU 上运行,因为企业客户受限于 GPU 获取能力。核心口号是:「ROI,而不是 AGI。」
- 围绕劳动力问题,两人公开交锋,这也是本期最精彩的一段。Harry 认为,25-26 岁的营销经理和 SDR「并不出色……他们会在 12 个月内被替代」。Frosst 的反驳不是乐观预期,而是结构性判断:「还没有任何 LLM 独立完成过真正的突破……突破仍然来自人」;而且这不是时间问题,因为「序列模型的工作方式从根本上就是这样」。
- 基准测试并不能可靠衡量企业实用价值。Frosst 说:「它们反映的是模型在这些基准上接受了多少训练。」数学推理基准和 ARC-AGI(「像素操作挑战」)与企业客户提出的需求毫无对应关系;排行榜循环对消费级 AI 的炒作有意义,却不能回答「我买了 LLM、部署了它,最后拿到 ROI 了吗」。
- 模型主权已经成为可投资的主题。语言模型是类似发电厂的基础设施,各国应拥有自己的基础设施和模型;而美国「已经表明,他们愿意出于政治原因切断技术访问」(他提到美国持有 Intel 10% 股权)。加拿大身份如今成了 Cohere 的商业资产;至于美国以外哪家 AI 公司可能达到万亿美元市值,他的答案是:「也许就在北边。」
- 他对 2026 年最大胆的预测,刻意选择了一个并不惊天动地的场景:你打开 North,说一句「帮我报销」,模型就能执行完整的多步骤流程——「我知道这算不上多大胆……但要让它真正可用、真正值得依赖,仍然不是每家公司都做到了。」谈到中国时,他说约 1 个月前曾有 7 家新模型供应商在 1 周内发布 7 个模型,这让他说出「哇」,但「我不认为他们做出了击败其他模型的模型」。
1. AGI 的控诉:炒作「在当时显然就是错的」
- 本期的主线,是 Frosst 对 Sam Altman 的指责:「我不认为 Sam Altman 反复谈论 AGI 已经近在眼前,是在为世界做好事。我认为他现在已经做出过几次错误预测,而且在当时显然就是错的」——包括暗示「AI 会在 2 年内杀死全人类」。
- 周游世界是他举出的最佳案例:Altman「与全世界每一位主要领导人交谈,告诉他们,这项技术构成生存性威胁。我认为这在学术上是不诚实的,而且给他所热爱的技术帮了倒忙」。
- Harry 的追问是全场最锋利的时刻:末日叙事与融资需求相关。他说 Demis 和 Zuck 不需要资金,因此保持克制;「其他确实需要融资的人,就必须说得更挑衅」。Frosst 承认「你指出的相关性确实存在」,但说不知道那是否就是策略。随后他给出反击:「我们是一家由风险资本支持的公司,需要融资。但我们没有这么说。」
- Harry 随后指出,即使他提到的那些领导人,也已经改变了对即将到来的劳动力变化的说法,「这确实让我感到担忧」。Frosst 承认这项技术「和个人电脑一样具有根本性的变革意义」,但坚持认为「人们花时间做的事情里,也有大量并不合理」。
2. 单纯增加算力不够——GPT-5 就是例证
- Harry 问 GPT-5 比 GPT-4 好多少,Frosst 回答:「我其实觉得它更差。」Harry 随后说,这说明了把更多算力直接砸向问题的本质。Frosst 解释,模型自动选择「更慢、更笨重……我只是想要一个快速答案,它突然就进入深度研究。好吧,博士,冷静点。」
- 他对 AGI 的定义是行为性的,而不是基于基准测试:「一台你会像对待人一样对待的电脑」;但人们并不会像对待人一样对待语言模型。他的判断也很明确:「我不认为这项技术能把我们带到那里。」他称这是一个被悄悄广泛接受的共识——如果问大学计算机科学专业学生,继续增加算力能否实现 AGI,「大多数人会说不能」。
- 关键在于,他区分了平台期和进步:让模型帮你报销——读取邮件、寻找收据、交叉核对政策、调用内部 API、获得审批——这项工作「并没有进入平台期。那是更多的建模工作、更多的产品工作……构建更好的连接器」。但架构本身停滞了:「相同的模型架构已经用了接近 10 年」,仍然是 2017 年的 transformer,仍然是在加入新的训练步骤后进行下一个词预测。
3. Cohere 的切入口:不到 20 家公司之一,专注企业市场
- 按 Frosst 的统计,全球做 LLM 的公司「少于 20 家」——大多数在美国,中国有几家,Cohere 在加拿大,法国有 1 家。Cohere 的独特之处在于「专注于把这项技术带给企业」:没有消费级应用,也不追求用户参与度,「我们不想让任何人每月花 200 美元,用于个人生活中的某些事情」。
- 企业聚焦改变的是训练数据,而不是架构:Cohere 现在生成「合成企业环境」——「虚构的公司、这些虚构公司员工之间的虚构邮件,以及这些虚构公司内部的虚构 API」——并在其中训练模型,使其真正有用。但数据依然是瓶颈:「你需要真实世界数据,才能启动合成数据的生成过程」,Cohere 仍然依靠人工标注员在内部制作数据。在算力、算法和数据三者中,算法的约束最少;高质量数据,以及从高质量数据衍生出的高质量合成数据,才是门槛。
- 他对个人生活和工作的区分解释了这套策略:「我其实不想更快地回复我妈妈的短信……但在工作中,有大量事情我不想做。」
4. 基准测试衡量的是模型接受了多少基准训练
- Frosst 用一段历史回顾说明问题:LM1B 是预测报纸文章的后续内容;大约在 2022 年出现 Hello Swag,但「现在已经没人谈它了」;如今又变成数学推理基准和 ARC-AGI。「我们的客户没有任何人要求模型做数学推理……ARC-AGI 就像一个像素操作挑战。我们的客户从来没有要求模型做这件事,我也不认为他们以后会要求。」
- 他的结论平铺直叙:基准测试「不能准确反映模型的实用价值。它们反映的是模型在这些基准上接受了多少训练」。能不能被做成游戏?「当然可以。」大公司是否确实这么做?「我不知道」——他把这个保留意见留在了那里。
- 企业真正关心的替代指标是:「我有没有把它推向生产?我买了 LLM、部署了它,最后拿到 ROI 了吗?」排行榜对消费级应用来说「很酷」,但他的客户并不在意。
5. 模型与应用趋于融合——但这是连续谱,不是二选一
- 对于基础模型公司究竟会停留在 AWS 式的商品化底层,还是会吞下应用层,Frosst 拒绝接受这种二分法:「如果你想为某个给定界面打造最好的模型,最好的方式就是在这个界面上训练模型。」这正是 Anthropic 挑战 Cursor、Cohere 构建 North 的原因。
- 最关键的洞见是连续谱:2015 年左右的机器学习意味着一个任务对应一个模型,比如用猫的图片训练一个数猫模型;transformer 则「更接近另一边」——「训练一个在所有语言任务上普遍表现良好的模型,再针对你想做的事情对它进行优化」。Anthropic 的模型普遍擅长代码,而不是分别训练重构模型和调试模型;Cohere 的模型则针对企业工具调用和海量文档进行了优化。
- 对于「你们的评测更差」这一批评,他的回答很直接:「我们关心的是,客户拿到一份我们的模型,尝试用它做某件事时,能不能尽可能轻松地跑通……这些都不会体现在每年轮换的各种基准测试里。」
6. 劳动力交锋:「他们会被替代」对上「突破仍然来自人」
- Harry 援引 2 天前一位 Salesforce 嘉宾提出的「同一个人加一个 agent」,直接切入:「大多数 25、26 岁的营销经理或 SDR——很抱歉这么说——并不出色。他们不热爱这门手艺。未来 12 个月内,他们不会比一个出色的 agent 做得更好。他们会被替代。」
- Frosst 的反驳是全场最值得引用的技术判断:「还没有任何 LLM 独立完成过真正的突破……从来没人要求 LLM 去解决一个此前无人解决的问题,然后得到答案。突破仍然来自人。」这只是时间问题吗?「不是。这从根本上就是序列模型的工作方式」——它们是文本的统计模型;而营销人员真正的工作——「理解文化、理解时代精神……运用直觉」——并不在互联网上的文本数据集中。
- Harry 用 Evian 广告活动反例回应:让模型提出 3 条故事线,再选出最好的一条。Frosst 没有否定,而是将其纳入自己的框架:「这是一个很好的用例……但工作并没有在那里结束。那只是工作的开始。你现在不是从一张白纸起步,而是从几个可供推进的方向起步。」
- Frosst 描绘的 5-10 年后公司是:你坐在电脑前,用语言把所有「信息已经存在、不需要创造力或洞察力……而且做起来有点无聊」的事情交给电脑;自己则把时间花在与人交流,以及判断模型做得好不好上。
7. 工业革命框架:政策决定 AI 是改善还是恶化不平等
- 对于「AI 会改善还是恶化收入不平等」,他的答案保留得非常准确:「这取决于政策。如果劳动力政策良好,我认为它可以改善;如果劳动力政策糟糕,它可能恶化。」前例是如今所有人都认可的工业革命:农业劳动力占比从约 90% 降至低于 5%;但在工会和劳工权利出现之前,孩子们先被送进了煤矿——「其中很多结果都来自公共政策……由企业和政府共同推动形成」。
- 这又回到了他对 Altman 的批评:生存性威胁叙事「让人们更难讨论真正的问题,比如收入不平等」——而收入不平等在语言模型流行之前就已经上升;他担心,如果部署方式不当,这项技术「有可能加剧」这一趋势。
- Adam Smith 的插曲是真正的分歧:Harry 认为看不见的手会把 Frosst 过去在餐馆做最低工资烤架工作的价值,定价为「明确低于」Cohere 那种影响数百万 Fujitsu 用户的工作。Frosst 的立场居中:经济体「相当擅长判断这些事情……但我不认为它是完美的」;而在空调坏掉的厨房里跑腿添薯条,「是一份有挑战、也有回报的工作」。
8. 2 块 GPU 的商业模式:「不是 AGI,而是 ROI」
- 融资讨论提到约 6 亿美元。Harry 说上一轮估值「大概是 67 亿美元之类」;Frosst 纠正为 68 亿美元。Cohere 在创建基础模型上的花费,「确实比其他一些基础模型公司少了几个数量级」——这条路线可以追溯到公司创立时期的论文:如何利用数据中心里「剩下的 GPU」进行训练。
- 战术核心是:刚发布的 Command A 推理模型和 Command A vision,都「训练到可以在 2 块 GPU 上运行」——因为企业在部署时受阻,「它们没有足够的 GPU」;而 2 块 GPU 恰好在性能、成本和实际 GPU 可获得性之间取得了平衡。消费级公司可以因为用户参与度带来收入,而「每次推理调用都亏很多钱」;企业模式承受不了这一点。
- 分发策略也随之明确:前置部署工程师是帮助客户进入生产环境的「关键组成部分」(这是个好主意,但「我不知道这是否适用于每一种业务」);参考客户包括 RBC、Fujitsu 和 LG;模型权重以非商业用途发布——这是一种建立可信度的中间路径,让客户可以验证「它们是否适用于我的问题」,同时把商业用户引导进付费关系。「我很惊讶没有更多基础模型采取这种方式。」
- 对于人才大战的新闻(Meta 开出 1 亿美元薪酬包),他的态度是审慎怀疑:「我读到的那些人拿到巨额薪酬的故事,和我读到的第二天就离职的故事一样多。」他确认可能从 Meta 招来了 Joelle Pineau;如果一名研究员「能带来正确的价值」,他愿意支付 500 万美元,但人们留下来是因为「稳定性、使命感和价值观一致」。
9. 主权:模型是发电厂,加拿大如今也是卖点
- 基础设施类比支撑了整套论点:「拥有一个使用本国语言的语言模型,就像为本国人民建设基础设施……我喜欢加拿大拥有几座核电站。」Frosst 认为各国拥有自己的模型基础设施是好事;使用在中国或美国构建的模型,「可能不如使用一个掌握本国语言、方言和文化语境的模型,更能让本国和本国经济获得良好发展」。
- 地缘政治上的关键一击是:「美国已经表明,他们愿意出于政治原因切断技术访问……随着时间推移,美国科技公司与美国政府之间的联系变得越来越不清晰」——他提到了美国持有 Intel 10% 的股权。结果是:「加拿大以及世界各地有很多公司希望与非美国科技公司合作,我会说这一直是我们的资产。」
- 他对监管的噩梦与对基准测试的批评如出一辙:最糟糕的做法,是监管者认为「我们正在构建的是数字神明」,然后挑选「他们认为代表 AGI 的某个随机基准」——无论往哪个方向都可以被做成游戏——最终据此叫停开发。
10. 快问快答:告别乐观主义、中国,以及 2 个 junior chicken 汉堡
- Frosst 已经不再称自己是技术乐观主义者,算起来已有 10 年。他的转变故事是:他曾经喜欢 Google Glass,后来「有一次我上了一辆公交车,有人戴着它……所有人立刻都注意到了」。至于 VR:「我其实不想把电脑绑在脸上。我想更多地投入现实世界。」但他同时保留 Harry 对孤独的担忧,以及古希腊哲学家哀叹文字出现的历史模式:「你必须把这两种相互冲突的观点同时放在脑子里。」
- 中国:他不担心。「他们做出了不错的模型……但我不认为他们做出了击败其他模型的模型」——不过约 1 个月前,7 家新模型供应商在 1 周内发布了 7 个模型,确实让他发出「哇」。他对 2026 年最大胆的预测是,「帮我报销」可以端到端真正跑通——「我知道这算不上多大胆……但这种方式最终普及到成为使用电脑的常态,仍然很疯狂」。美国以外的万亿美元 AI 公司?「也许就在北边。」如果不是 Cohere,那就是 Google DeepMind。
- 关于自己曾经判断错误的地方,他过去相信人类进步是单调向上的——「预期寿命会上升,收入不平等会下降……但事实并非如此。」他还承认了一个 2020 年的技术误判:「我记得当时我说,你不可能只用一小份来自人类反馈的数据,就把模型变得更好」——RLHF 的数据效率证明他错了。并购报价确实来过(「我们已经是一家成立 5 年的公司了——当然有」),但吸引他的仍是「打造一些能比我们活得更久的东西」,即使他知道结局会像《奥兹曼迪亚斯》:「有一天,剩下的只会是沙漠里的两条腿。」
- 这个仪式已经确认:每次完成一轮融资后——目前是第 5 轮,也可能是第 4 轮——3 位联合创始人都会去 McDonald's。「我通常会要 2 个 junior chicken。」
I don't think Sam Altman has done a service to the world by talking about how close AGI is. I think he has made several predictions now that are wrong and that were obviously wrong at the time he made them.
“I think AI will probably lead to the end of the world.” He's made allusions to things. He did a world tour where he spoke to every major leader the world over to tell them, “Hey, this technology is going to pose an existential threat.” I think that was academically disingenuous and did a disservice to the technology he loves.
Nick, I'm so excited for this, dude. When I had Aidan Gomez on the show, he was like, “You've got to have Nick on. He's the real star of the show.” And he introduced us way back then. So, I'm so excited that we can make this happen.
Yeah, man. I'm happy to be here.
1. Biggest lessons from Geoff Hinton at Google Brain?
Now, before we dive into Cohere, I have to ask: You were Geoffrey Hinton's first hire at Google Brain, and then you're put in a room with Geoffrey Hinton and get to work with him every day. What was the biggest lesson from working with Jeff, a legend of the industry?
I loved working with Jeff. I learned everything I know about research from those 3 years. I was very surprised at how creatively and playfully he approaches research.
When we would discuss algorithms, optimizers, or loss functions, we would often discuss them through physical analogy. We'd spend a lot of time talking about, “Imagine there's a ball here, an elastic band to this thing, and a pulley here. It's on this kind of a surface.” A lot of it was described in the natural physical world, and that was very playful.
A lot of it was approached with curiosity, and I didn't expect that when working with him. I expected it to be much more like, “Here's the equation. Let's figure out what the derivative is, and let's go from there.” Instead, a lot of it was based on intuition.
2. Did Google completely sleep at the wheel and miss ChatGPT?
When you look at Google Brain and DeepMind, a lot of people think that Google was asleep at the wheel, given that it wasn't at the forefront of the consumerization of the technology with ChatGPT. Do you think that's fair?
I don't know. It's certainly interesting. The Transformer was invented at Google, right? In 2017, Aidan Gomez, amongst many other brilliant people at Google Brain, published the Transformer as an architecture.
It wasn't commercialized very quickly within Google. It wasn't scaled up very quickly within Google. A lot of that work had to be done elsewhere and years later. So that's interesting. Why is that? What systems are in place to make that be the case? I don't know.
I will say there's still a ton of brilliant people in DeepMind. I think Google DeepMind has now just subsumed the rest of it and is doing great work. They continue to make good products. It is interesting that all the people who worked on the Transformer left to continue to work on the Transformer.
If we go down your tangent, for people who don't know, and just to set the scene before we dive in properly, what is Cohere, and how does it differentiate from more generalized models that are maybe more well-known, like OpenAI and Anthropic?
We're a foundational model company, like those other 2. We build foundational models; we build language models. There are maybe 10 companies in the world that are building large language models in the West. Maybe there are 15 in the West; I don't know. We'll have to figure out how many there are.
A few have popped up recently, but there's some number less than 20 in the whole world: most of them in America, a handful in China, us in Canada, and 1 in France. Those are really the companies out there.
We're unique in our singular focus on bringing this technology to the enterprise. So that means we train a model that is good at enterprise tool use. We train a model that you can give a bunch of tools and APIs within your business, give it access to your business's data, and then ask it to help you with something in your work, and it does a good job of it. That's what we train it for.
How does a focus on enterprise over consumer change the way in which you train and build a model?
The models themselves—the Transformer architecture, which was introduced in 2017—haven't changed very much, right? The whole industry is still using Transformers. We've changed the way we train them, but the model architecture itself is approaching 10 years of being the same model architecture.
When we train our model, we're not training it to be an amazing conversationalist with you. We're not training it to keep you interested, keep you engaged, and occupy you. We don't have engagement metrics or things like that. We're just training it to augment you in the workplace. We're training it to help you do your job.
3. Is data or compute the real bottleneck in AI’s future?
That means the type of data we train it on is very different. Recently, we started doing a bunch of work on synthetic data. We generate a whole bunch of data to create fake companies, fake emails between people at these fake companies, and fake APIs within those fake companies. Then we train the model in that synthetic environment to help out within that fake business.
Do you think data is a bottleneck, given the ability for synthetic data to produce an infinite supply?
Yeah, data is still a bottleneck. You need real-world data in order to start a process of synthetic data. Synthetic data has helped a lot, and it's made models better than they would be if they didn't have access to it.
Getting access to high-quality data is still something people think about. We still make a whole bunch of data in-house with annotators who are making real data, not synthetic data.
When you think about the 3 pillars of compute, algorithms, and data, which one do you think is most constrained or the biggest bottleneck?
It's interesting. The algorithms haven't changed very much. They've changed a little bit. When we started this industry, originally we were just training base models, which weren't called base models at the time. They were just called large language models.
They weren't trained from human feedback. All they would do is take in the first part of a sentence and write the second part of a sentence. If you tried to have a conversation with them, it wouldn't work because that wasn't the data they were trained on.
Since then, we now train models in a few different steps. There's a base modeling step, then there's a supervised fine-tuning step from human feedback with SFT data. After that, there might be a variety of other reinforcement learning techniques you can use.
4. Does GPT5 Prove That Scaling Laws are BS?
The algorithms, I think, are not the bottleneck in terms of making those models more useful. I do think a lot of it is still getting good-quality data and then making good-quality synthetic data from your good-quality real data.
When we think about the bottlenecks, that leads to a potential plateau that people are worried about. Everyone seems to now be on the train that more compute and scaling laws are more real than ever, and that we will continue this exponential progress with more compute.
Do you agree that we're seeing the benefits of scaling laws for the next 12 to 24 months, or do you think that more compute will not just lead to more progress? How much better do you think GPT-5 was than GPT-4?
I actually think it was worse.
So I think that tells you something about the nature of just throwing more compute at the problem.
Does it? So why do I think it was worse? I think it was worse because the way that they now do model selection is slower and more cumbersome, and it's actually a pain. It gets it wrong sometimes. To me, I just want a quick answer, and it suddenly goes into deep research. I'm like, “For fuck's sake, I just want a quick answer. All right, PhD, calm down.” Do you know what I mean?
I think it's a worse product in that respect. I think we waited for 1 year or 1.5 years.
Yeah.
That was for model auto-selection. I think, if I go back to your original question of whether I think just throwing more compute at the problem will help—some people are thinking there's a plateau—do I think there's more compute? I think we need to agree on where we think the technology is going to establish whether or not there's a plateau, right?
I think language models are incredible. I think they're super useful. I use them in my work life as often as possible. One of the reasons why we're focused on the enterprise is because that's really where I think large language models are useful.
If I look at my personal life, there's not a ton that I want to automate. I actually don't want to respond to text messages from my mom faster. I want to do it more often, but I want to be writing those. I want to be engaged.
Whereas in my work life, there's a ton of stuff I don't want to do. We need to get to a stage where I can open up North and say, “Hey, file my expenses,” and it can figure out, “Okay, cool. I've got to look through all your emails, look through photos of receipts you've taken, cross-reference that with the things you're allowed to expense via internal documentation, figure out what the API is for how to expense things within your company, and then do all of those and get approval before I do them.”
That’s a many-step process. But that’s where the technology is going. The work of making a model do that is not plateauing. That’s more modeling work. That’s more product work.
That’s building better connectors. That’s building safer data integrations so that you can trust giving a model access to the types of things I just said. That stuff’s still ongoing, and that’s what we’re working on. I think when people are talking about building towards AGI, I don’t think this technology gets us there.
When you say “gets us there,” what is “there”?
Well, yeah, great question. We’ve had many years of people discussing AGI, and not many definitions thereof.
Next to none. I mean, my definition is when Sam Altman and Microsoft decide.
Yeah, they’ve changed their definition a few times on that. When I say AGI, what I mean is a computer that you treat like a person. I mean, when you use a computer and you expect it to behave like a person and treat it that way, I’ll call that AGI.
Do you not think we’re already there, then?
People do not treat language models like they treat people.
Do you think OpenAI and Sam Altman now realize that more compute does not lead to this exponential progress when they look at GPT-5?
I don’t know. I think they’re a great company. They build a really cool consumer product. I don’t know.
I guess my question is, why does the world still think scaling laws are so prevalent when you don’t?
I think a lot of the world does not think scaling laws are super prevalent. If you go out into a university and talk to the students there who are studying computer science, or even the students who are not, and you ask them, “Hey, is throwing more compute at this problem going to get us to AGI?”
Most of them say no.
You mentioned a couple of different use cases. Expense management was one that you clearly articulated. A question that I think I have, and a lot of people have, is how far do models go in terms of value capture into the application layer? You’re seeing Anthropic now, with Claude, really challenge Cursor and the world. You’re seeing OpenAI with a lot of consumer products. You’re seeing Cohere with a lot of enterprise use cases. How do you think about whether they stay as AWS-style commodity layers, or whether they extend into value capture of the application layer?
Yeah, that’s a good question. I don’t see the two as that different. I see the two as related. If you want to make a good product with a large language model, you are best suited to training that large language model for that use, for that product, right?
That’s one of the really interesting things about LLMs: they’re phenomenal and they generalize really well, but they don’t generalize as well as you might think. If you want to make the best model for a given interface, it’s best to be training the model on that interface.
So, do we see this deeply specialized, unbundled-model world, where you have very, very specific use cases and models are trained for the ElevenLabs of the world? That would be another brilliant use case, specifically voice. Is that the world that we live in?
Well, there’s a spectrum. The old world of machine learning, back in 2015 or something, when the world was new, was that for any task you wanted to do with a neural net, the best neural net you were going to get was one trained on that task.
If you wanted to make a neural net that was going to identify pictures of cats and tell you how many cats were in an image, you were best suited to train a model on identifying pictures of cats. The world of machine learning in the first half of that decade was all about: here’s a problem, make a dataset, train a new model from scratch, or maybe take SIFT features, or maybe take some base model, but pretty much train a model on the dataset itself and go to production with that model.
That’s not the case with language. If you want to make a model that’s the best at summarization, you can’t just train it on summarization. You have to train it on all language, and that is the technological reality that has brought us to where we are today. That’s true for foundational models and not true for the neural nets of 2015 and before.
That’s super interesting. But that’s a spectrum, right? On one end is every single task: train a single model for it. On the other end is train 1 model to do everything.
I think the reality of what we’re seeing with transformers is that they’re not at this end of the spectrum. They’re a little over here: train a model that is generally good at all language and refine it on the type of stuff you want to do with it.
You’re seeing that Anthropic’s Claude models are very good at code, but they didn’t train a model specifically as a refactoring model, a debugging model, or a model just for writing test cases, right? It’s a model that’s good generically at code.
For us, for Cohere, with our focus on enterprise, secure deployments, and customizations for our customers, that means training a model that is good at helping people in an enterprise setting, using internal tools and reading through massive amounts of documentation.
Does that mean a model doesn’t need to be as good? Again, I’ve learned to be incredibly blunt over time. If someone wants to criticize, they say, “Oh, well, if you look at models or evals, Cohere is not as good.”
Well, there’s a whole long conversation to be had about evals, but effectively, what we care about is not the score, not the hype, not the discourse. What we care about is that if a customer uses our model, gets a copy of our model from us, and tries to do something with it, it works as easily as possible.
5. Are AI benchmarks just total BS?
That’s what we optimize for. None of those things are really reflected in the various benchmarks that cycle through every year, and so we don’t focus too much on that stuff.
Do you think the benchmarks, because we place a lot of emphasis on them in the Twitter sphere and the Reddit sphere, are an accurate reflection of model progress?
Let’s go back in time a little bit. When we first started in this industry, the benchmark that was used the most was called LM1B. Do you remember that benchmark?
No.
Okay, cool. That was a benchmark that took in the first part of a text, like a newspaper article, and then wrote the second part of the newspaper article.
After that, there was a benchmark called Hello Swag. You remember that one?
Yeah, I do remember that one.
All right, cool. So, okay, that’s like 2022.
That’s my introduction.
That’s when you started. All right, cool. No one’s talking about that anymore, right?
Right.
Yeah. Now, a lot of people talk about a math-reasoning benchmark.
None of our customers ask the model to do math reasoning. That doesn’t come up in the workplace that often. It comes up in a few workplaces where mathematicians work, but there aren’t a ton of people out there making a living doing math reasoning.
Stuff like the ARC-AGI challenge is a benchmark that people talk about, but that’s a pixel-manipulation challenge. It’s taking in a grid of pixels and, based on rules, predicting the next one. That’s not a thing any of our customers have ever asked the model to do, nor do I think they will.
I think, taken as a whole, they’re interesting. I don’t know. There’s good scientific work in some of them. I think it’s very interesting to evaluate emergent capabilities for models.
But they’re not an accurate reflection of the utility value of models.
They’re a reflection of how much the model has been trained on those benchmarks.
So, you can gamify them, essentially?
Oh, you can definitely gamify them. Yeah.
Do the big players gamify them?
I don’t know. Yeah, I don’t know. None of it’s too relevant to us, right? I don’t think those leaderboards are that helpful.
I think in a consumer space it’s cool. If you’re making a consumer app and it’s exciting and fun, and people like to look at it and want to try out the most recent thing, that’s fun. That’s cool.
If you’re not in a consumer space, people don’t care about the hype as much. They care about, “Hey, did I get to production? Hey, did I buy LLMs, deploy them, and then get ROI on that?”
Given the pace of deployment, we’re seeing model evolution so fast and so rapidly that you’re essentially seeing this kind of decay rate on models being greater than ever, because it’s the next one, next one, next one. And actually, they’re still being trained on H100s, or NVIDIA chips, from 18 months ago.
Is there a misalignment in terms of the progression of models versus the progression of chips?
You can cycle through new versions of models quicker. I mean, it’s still very slow. When I was training neural nets in 2011, it would take hours to days. I remember being, “This is crazy. I can’t believe this takes so long to train this model.”
Now we spend months training models. That’s a timescale I didn’t anticipate when I was working on neural nets a long time ago.
But that's still very different from the timescale of working on chips, right? That's still slow. I think when you talk about how we're seeing all these models iterate so quickly, on the one hand, we're seeing models iterate really quickly and people are releasing new models. On the other hand, they're still based on the transformer that was invented in 2017, they're still sequence models, and they still take in words and predict the next word.
We've changed how they're trained a bit. We've added steps: now there's a base modeling step, then an SFT, supervised fine-tuning from human feedback, where somebody writes a sentence and then writes the response they want, and we train on that. Then there's a reinforcement learning aspect where the model is generating and you're telling it, “That's good” or “That's bad,” or something.
There are new ways of training them, but fundamentally the technology is still the same. We keep making them better and keep iterating on them, but it's not as though anybody has trained a model that's fundamentally different from a transformer. It's an interesting dichotomy: on the one hand, there's constantly new stuff; on the other hand, we've been working on the same stuff for a while.
We have been working on the same stuff for a while. The thing that has seemingly changed is the value of the people working on the stuff. We're now seeing billion-dollar people, in terms of Zuck's willingness to pay for chief scientists.
You recently hired—and I just want to get the name right—Joelle Pineau.
Yeah. Joelle from Meta.
How do you think about the war for talent that we're seeing today?
I think there are a lot of crazy headlines out there. I don't know how much of it is real.
You don't think it's real that Anthropic are paying $10 million, $15 million, $20 million for great AI researchers?
I have no idea. I know that there are lots of people who are adding that much value, and there are lots of people who are bringing that much value into the industry. It's a super impactful industry, right?
I know that there are people who are bringing that much value, and I know there are lots of brilliant people. This is really demanding work. It's really hard work. It requires a lot of experience, a lot of ingenuity, and a lot of dedication, so I think it's a good place for people to be spending their time. I think it makes sense that many of them are rewarded very well.
That being said, when I see stories of Meta hiring people for $100 million, I read as many stories of those as I read of people leaving the next day. I don't know what's going on over there.
I know that people like to work in a place that gives them stability. People like to work in a place that gives them purpose and is value-aligned, and they like to work in places that make them feel good about what they're doing. Part of that is compensation.
6. Would Cohere spend $5M on a single AI researcher?
Do you sit down with Aidan and the team and say, “We need to step up what we pay people because we are in a war for talent, and budget is a big part of it”?
We definitely think about that. Our company is what we spend the most time thinking about and talking about, and the company is the people who are there. Functionally, we are only the people who we have the privilege of working with.
We do think about, “Hey, is this the right place? Are we making sure that this is the right place for people to work? Are we making sure that this is getting the best in their lives from a financial perspective?”
Would you spend $5 million on an AI researcher?
Yeah, if they were bringing in the right value. There are certainly lots of people who, through our equity, own portions of that—own what you're talking about. I feel great about that.
Do you worry that the industry is becoming commoditized or transactionalized with the hype around it?
I do think the hype around it is misleading sometimes. I'm in such a strange place of being caught between the technology being the most beautiful technology I've ever seen and the most transformative technology I've ever worked with. It is already fundamentally changing the way I work, and I'm very sure it will fundamentally change the way we all work soon.
On the other hand, there's a lot of hype around it. There's a lot of misleading rhetoric and misinformation, and I don't think the hype is necessarily helpful for getting to the truth.
Can I just dig in there? What do you think is the hype and misleading rhetoric that is most damaging or confusing?
I think the hype around AGI is the most damaging and confusing.
This assumption that we'll all have no work to do, we'll all be on UBI, and—
Yeah, and even—this isn't really in the discourse as much this year as it was last year and the year before that, and that's because it's pretty clearly not true. The idea that this technology poses an imminent existential threat to humanity at large was incorrect and not helpful for talking about the real ways in which this technology could be damaging.
It wasn't helpful for talking about the real ways in which this technology could shake up a system and cause rapid changes, and it was not helpful for getting people to understand what the technology is. I don't hear that as much anymore these days. I think that's because people have realized that that's not the case, but the remnants of that discourse are still in the world. They're still out there.
I think the remnants are, and I think they're most prevalent in the way internal employees in large organizations respond to AI being introduced. People do not welcome the introduction of AI in large companies.
Yeah.
Maybe it's a European slant, but a lot of people are very nervous and scared and do not embrace it wholeheartedly.
Interesting. I haven't found that as much. When we've worked with our customers and our enterprises, I find a lot of people are interested in and excited about using an LLM. Mostly, that's because I think they realize that the LLM is augmentative, for the most part, and it allows them not to do the things they don't want to do.
I mean this in a nice way: do you actually buy that?
Yeah.
Like—nice. Yeah. Yeah. I had Marc Benioff on the show from Salesforce 2 days ago, and he said, “The same human plus agent.”
Are you serious? Most 25- or 26-year-old marketing managers or SDRs—I'm sorry to say it—they're not brilliant. They do not love the craft. They are not better than a phenomenal agent will be in the next 12 months. They will be replaced.
Oh, no. Yeah, I actually believe it. Yeah, no, I believe what I said. Sorry. I fundamentally believe that this technology is—and I think this is clear if you use it for a while—there are things that it's way better at than you are. But there are still lots of things that people are better at.
LLMs are incredible. People have been using them for years now, and there has been no independent breakthrough that an LLM has made, right? Nobody has seen an LLM solve a problem no one has solved before and get the answer. The breakthroughs are still made by people.
Is that not only a matter of time?
No, it's not. That's fundamentally the way that sequence models work. We were training statistical models of text. They're phenomenal and incredible, and they're capable of generalizing across unseen tasks, which is why they're so useful.
When you're talking about the 25-year-old marketer, some portion of their work is writing text on a computer. All the information is out there. They have this document and that document, this tool and that API. They just need to take that, turn it into another form, combine it, and put it out there. That's some portion of the work. It's not the majority of it.
Most of it is talking to people, understanding the culture, understanding the zeitgeist, understanding what's going to hit and what's relevant, and using their intuition and human experience to understand how they can be helpful or what they can do. That is not in the dataset of text from the internet. So, yeah, I fundamentally—
I disagree, because you know what they do now? They say, “Hey, I'm running a campaign for Evian. Come up with 3 different storylines that would be cool for us to do. Make it relevant to news cycles today.” That's the prompt. Then they come up with 3, and they're like, “Oh, that one's pretty good.”
Yeah.
Yeah. I think that sounds like a good use of it as a starting point, and then I'm sure it is worked on. Some of those are thrown out because of things they understand, and some of them are like, “Oh, that's a great insight. I'm going to run with that. I'm going to tinker with that.”
What you just described is a good use case. I can imagine, as a jumping-off point, that would be helpful, but that's not where the work ends. That's the beginning of the work. You've just augmented it. You're now starting with, instead of a blank page, things to go for.
So, you don't mind that we're going to see this dramatic reduction in team sizes?
I think we're going to see changes to the nature of the workforce in the same way that we saw changes to the workforce when the computer was created, when the personal computer and the internet were created, and when the printing press happened, right? We've seen drastic changes to the workforce in the Industrial Revolution, and that's going to keep happening.
What will the changes be? What will a company look like, do you think, in 5 to 10 years?
I think you will arrive at work and sit in front of a computer. You will predominantly use language to interact with that computer, and anytime there's something to do that can be done, and the information is out there, and it doesn't require creativity or insight—you know that it's there, you just need to do it, and doing it is kind of boring—you will get the model to do it for you.
I think that mostly looks like sitting down, speaking to the computer, getting it to do the things you don't want to do, and then spending your time talking to other people, thinking about how it can be useful and whether or not what it did was good.
I think that shift is maybe chaotic. I would like us as a world to spend time thinking about how we can make sure that change is as easy as possible, and how we can make sure that making language models allows people to do the stuff that they're good at and that they like. How can we make sure the labor force is resilient? How can we make sure that income inequality doesn't go up as a result of that? Those are the types of things I would like us to be talking about, and those are the important things.
To go back to your earlier question, when talking about the risks of AGI, I think those existential-threat questions made it harder to talk about the real things, like income inequality.
Do you think AI does more to help or to hurt income inequality?
I think it depends on policy. If there's good labor policy, I think it could help. If there's bad labor policy, it could hurt.
Can you explain that to me? Look, when you saw the last Industrial Revolution, broadly speaking, everybody looks back on that Industrial Revolution and says that was a good idea, right? Nobody's saying, “Hey, we shouldn't have automated.” Before the Industrial Revolution, it was, I don't know, 90-something percent of people working on farms. Now it's, I don't know, 5%—I don't know what the percentage is, a small 5%, less than that.
Everybody thinks that was a good idea. It was a crazy time. If you read stories about what went on during that moment, there are lots of things that people did that they stopped doing pretty quickly, like having kids work in coal mines, right? That was a crazy thing.
Out of that Industrial Revolution came a whole bunch of really good labor policies. Unions came out of it, workers' rights came out of it, and things that I think we also think were a good idea. They resulted in not only better lives for people, but actually more productivity, a better economy, and a better world.
A lot of those were from public policy. A lot of those came from things created in unison between businesses and governments.
Why do we need to have such significant policy change if it only augments humans and it doesn't replace them?
Well, I think right now we've seen income inequality go up over the past several years. A lot of that was happening before AI, before language models were popular. I'm worried the technology has the potential to exacerbate that without being deployed correctly and without having good policy around employment.
7. Open vs Closed AI Models
Can I ask you, when we think about problems to solve? I think a lot of people also get worried about the open-versus-closed argument. How do you feel about where the future of efficient AI lands in the balance between open versus closed models?
At Cohere, we make our foundational models and then release the weights for non-commercial usage. So we're somewhere in between the open and closed, right? We're a for-profit company; we exist to make money. We release our weights for scientific and research use, and you can download them on your computer and run them.
I think that's a good sweet spot for us as a business. That allows us to build credibility within the community. If people want to check out our weights, they can go check them out, right? There are lots of companies that started out as open who no longer release the weights of their models, or who never did, right? We have our models out there. You can go look at them. You can use them. You can validate: Do they work on my problem, yes or no?
But if you're using them for commercial purposes, you've got to talk to us, and then we figure out a commercial relationship so that we can exist as a business. That works for us. I'm surprised there aren't more businesses taking that tack, more foundational models taking that approach.
Do you think Meta will move to a closed-model approach from an open one?
They've certainly hinted at that, right? It certainly looks like it, but I don't know what they're doing over there. I don't think a lot of people know what they're doing over there, and I don't spend a lot of time thinking about it.
Do you not think it's helpful for founders to be very aware of competitive landscapes in case they're asked about them by customers? In case customers are going, “Hey, why aren't you more open? Why aren't you more closed? Is our data secure?”
As with most things, a middle ground is the right place to be, right? You could spend your whole time as a founder only looking at competitors and being like, “Why are they doing that? Why are they doing this? What's going on with that?” I think that will not be helpful for you.
You could also spend your whole time with your head in the sand, only thinking about what's going on in your company, and I think that would not be helpful either. You have to find some middle ground.
The discourse around AI is inescapable. You would be hard-pressed to ignore it. It is every other headline. It's all over the place. I don't think there are many people who work in the industry who suffer from not enough information about what's going on in AI, right?
8. Future of Prompting
I think there are a lot of people who suffer from way too much of it, obsessing over the minute details of how so-and-so got 2% better on this thing, or constant small changes in businesses out there. I think that can mislead you from staying grounded in what you're actually doing, who you're actually helping, and how this is making things better for your customers.
Do you think we will still have prompting as the core user-input guidance mechanism in 5 years' time?
Prompting as in, like, you write something to a model and it writes back?
Yeah.
Yeah. What else would it be?
The way that it changes is that the way that you do it changes. You wouldn't say, “Hey, make it a funny tone,” or, “Hey, add in a light, personalized style that also is sincere.”
I think the idea of prompting as a skill will become less relevant. If you look at the trajectory, when I started doing this, if you wanted to get a model to summarize something, you wrote the first paragraph, and then you wrote “In summary:” on a new line, and then you generated it. That was the skill of prompting: figuring out how to trick a model into getting it to do what you wanted it to do.
That's because they weren't trained on feedback from people. They were only trained on text from the web. All they were were sequence models trained on text from the web. Nobody on the web wrote, “Please summarize this for me,” and then a summary. They wrote a paragraph and then wrote, “In summary:” So if you wanted to get the model to do that, that's what you would do.
We're training language models more to fit how people expect them to work, and that means that getting good at prompting is less important. I think the idea of saying, “Oh, yeah, you've got to learn how to prompt,” is going to go away.
I think the idea of saying you need to learn how language models work and you need to know what they can and can't do, in the same way you had to learn how a computer works and what it can and can't do, and you had to learn how a telephone works and what it can and can't do—I think that's going to exist.
That means prompting is going to exist. The idea that you write something to a model or you say something to a model and then you get the response back, and if it's not what you like, maybe you iterate a little bit—that's going to exist. That's fundamentally how the technology works.
But the idea of it being a discipline that you have to train to do, we've already seen that trajectory; it's already gotten easier. I look for people who know how a language model works and who know that one of the core things that's been a necessary component of working at Cohere is that you can't think the technology is magic.
You can't think this is like we're doing spells. You have to know how a language model works, how it's trained, and what that means for it—what emergent capabilities happen and which ones don't. You can't think, “Oh yeah, I just ask the digital god to do my work and then it does.” That's not what the technology is, and thinking that will not help you build it or use it.
9. Lessons from a $600M Fundraise
If we hone in a little bit more on you, you led or were a large part of the latest fundraise. When you think about the fundraising journey, how was that journey, and are there any big lessons from it? Was it $600 million?
No. So, fundraising at Cohere—lots of people are involved in it. I by no means led the efforts, but I was involved in it. I enjoyed talking to people about the tech and what we're building.
I quite like talking to VCs, pension funds, and people. I love Cohere and what we build, and I also like talking, so I like answering questions.
Are the questions very similar?
Between them?
Yeah.
Yeah, actually, that's an interesting question. I do think, look, the industry is a lot more mature. Two years ago, or three years ago, when we were fundraising, a lot of the questions were, “What is this? How does that work?” We spent more time explaining stuff.
People mostly know how it works and know what it does. Now we can say, “Here's what we're doing for our customers specifically. Here's how RBC is using it. Here's what we're doing with Fujitsu. Here's what LG is doing with it.” We can talk about those things specifically, and that's more interesting.
How much of the $600 million will be spent on compute?
There are 3 components that go into making language models, right? There's talent—people, smart engineers, and researchers—there's compute, and there's data. The importance of those has shifted, and the spend on those has shifted over time.
We train very efficiently. We train efficient models. Our Command A model, the Command A Reasoning model, which we just released, and the Command A Vision model, which we just released, are all trained to fit on 2 GPUs. That's a really important part of our business strategy.
It turns out that if you talk to a lot of companies that want to deploy models into production, they're bottlenecked on deployment because they don't have enough GPUs. 2 GPUs turns out to be a sweet spot between performance and cost, and actually how many GPUs they have access to.
That means we train very efficiently as well. We've spent orders of magnitude less on creating foundational models than some of the other foundational model companies out there—truly orders of magnitude less. I'm very proud of the efficiency of the team and what they've done with the resources they have.
We think about efficiency a lot for ourselves and for our customers, and those 2 things are related. How much of our funding goes to compute shifts over the years, but a lot, you know.
How has it shifted over the years?
When we first started Cohere, one of the very first things we did—because we had no funding, or we did have funding but a very small amount of it—was spend next to nothing on compute. We showed that you could train a model by having a bit of a GPU over here, a bit of a GPU over there, and a bit of a GPU over here, and you could link them together.
We published a few papers on training models with the scraps of GPUs in data centers. That was what we started with. We showed that you could do that. It's very slow, and it's much easier to just rent a big data center and train the model there.
10. How do Cohere compete with OpenAI and Anthropic’s billions?
The question that everyone asks is, how do you compete against competitors who have billions and billions of dollars? Do you hate that question, and how do you respond to it?
No, I don't hate that question. I think that's a fine question. We've announced funding rounds, and you can see that they're smaller than some of the other funding rounds out there.
We're pretty singularly focused in a way that the other companies that build foundational models are not. We don't have a consumer app. We're not trying to get anybody to spend $200 a month on something for their personal lives. We're singularly focused on working with enterprises and businesses and making sure that they get to production with AI.
I'm constantly telling people, “Not AGI—ROI. ROI, not AGI.” It turns out there's a lot of work that still needs to get done there, and there are a lot of companies that have tried to go to production with AI.
Do you just think then that OpenAI and Anthropic will just cede enterprise?
Right now, both of those companies are pretty cool. They've both made good consumer products. I think where this technology adds the most value is in work, for personal reasons. That's where I see this technology being the most useful. I don't know if they'll start working on that.
I know that making models that work in that environment is pretty different from making a model that works in a consumer environment. In a consumer environment, you can make the biggest model possible. You can have complicated switches to tell you to go to this model or that model because you're just hosting it on a huge amount of GPUs.
You can be losing a ton of money on every inference call, but you know you're getting users, and something like that works. The types of models you have to build to succeed there are different.
I know the work you need to do on the interface—North, which is our agentic framework, is privately deployable and customizable for knowledge workers within an enterprise—is pretty different. It looks pretty different from some of the consumer applications.
A big one is that our model doesn't generate images. Nobody in a workforce is really wanting to generate images as part of their work, for the most part. But as a consumer, it's very fun and very cool to be like, “Oh, here, give me a picture of this or something.”
The types of models we train are different, and the interfaces we make are different. I don't know if they'll be interested in that at some point. We stay focused on talking to customers and adding value.
How do you price?
That's entirely dependent on what the customer wants to do with us. We do have some customers where we make a custom model for them and give them that model.
Do you have forward-deployed engineers?
We do. We have forward-deployed engineers, and they're a crucial component of how we get a company up and running and into production with us.
Do you think everyone will have forward-deployed engineers in a future AI world, in a way that Palantir has glamorized?
Yeah, I think forward-deployed engineers are a good idea, right? You're selling technology to somebody. It makes sense to have some engineers who come and help them get it set up and work with them to make sure it's actually delivering value. I think that's a good idea. I don't know if that's true for every business.
Does FDE not just allow for poor technology?
No, no. I think there's this idea sometimes that you can just make the thing and, for every business, it'll work perfectly and require no engagement. That's the way some technology works. That's the way a lot of consumer technology works. That's not the way a lot of enterprise technology works.
You're selling things to people that have to be matched to the way their business is set up. Having engineers go along with it and say, “Cool, here's the model. Here's what we can do to make sure that it's perfect for you in your specific use case,” is helpful.
11. Do Enterprise Companies Trade at Lower Multiples?
Given that you sell to enterprises, I'm an enterprise investor, and I love enterprise. Revenue quality is much higher and much stickier, but growth is slower because you're working with large enterprises. Do you think you had, and have, an enterprise discount applied to your valuation because revenue growth is slower with enterprise?
It's an interesting question. Maybe, yeah. I'm very proud of what we've done and what we've created.
What was the price on the last round? It was public. I think it was like $6.7 billion or something.
Yeah, $6.8 billion.
$6.8 billion.
Yeah. These are all staggering numbers. These are all numbers that are impossible for an individual to conceive of. As a single individual, this is well beyond the realm of what you can reasonably engage with in your life.
For a regular person who grew up working as a cook, my first job was burgers. These are all crazy numbers.
Do you care about money?
And everybody’s motivated by money.
With being motivated by money, you’re an incredibly attractive asset, and it’s been a very strategically important thing for large players to do. Have you had M&A offers across the journey?
Oh, yeah. We have at times.
Yeah. How’s the decision-making gone there? I always want to be in the room. I always picture it kind of like thundery nights, with people coming together while it’s raining outside.
No, no, no, no. Yeah, we have. I mean, look, we’ve been a company for 5 years.
Not tempting?
We’re all really interested in building a generational company. A lot of the things you’ll hear Aidan say are about building a generational company. We’re all really interested in building something that outlasts us and goes beyond our involvement in it. That’s really exciting.
Why do you want that? I know it sounds strange. Why do you want something that outlasts you?
Oh, yeah. That’s a great question.
Remember earlier when I was like, “Sometimes I get asked philosophical questions and then we deviate too far off the thing”? This is one of those questions.
No, it’s cool. I love this.
This is why people say, “AI could just replace you as an interviewer, Harry.” It’s like, no, it couldn’t, because it doesn’t have the ambiguity to go off on the “but why does that actually matter?” tangent.
Well, then, if you think what you just said, why do you not think that exists for all jobs?
Oh, because I think what I do is a very disciplined art honed over 10 years, compared to a social media manager who’s 24, coming out of university, writing, “We’re thrilled to publish our latest report.” Everybody thinks that what they do is honed and trained. Then why is mine paid millions and theirs not? Because society places strategically more value on mine, if we’re being blunt, and I’ll take this out because very few people can do it.
There’s labor that’s easier to work at, that you can learn faster, or get up to speed quicker. There are things that take a really long time to do, and the only way you’re going to be able to do that job is if you spend a really long time doing it. There are some things that are harder, and some that are more agent-like, that have more agency. Our economy is decent at figuring that out and compensating people based on the investment they had to make and the skills they had to have to get there.
But I don’t think it’s perfect. I think everybody thinks, and should think, that the work they do is a skill that they learned. A lot of it is.
Even some of the hardest jobs—the hardest days I ever had at work were working at the grill, making breakfast and burgers for people when it was super chaotic and stressful, and there was no air conditioning because it had broken. I had to run across the street to buy extra potatoes because we hadn’t prepared it right. That was challenging and rewarding work.
I don’t think just because I was getting paid minimum wage at the time means it wasn’t valuable, but it is definitively less valuable.
It’s definitively paid less. Adam Smith’s invisible hand would suggest that it is just definitively less valuable that you went and got more potatoes, which meant more chips for someone who probably didn’t need more chips in a grill that was a single-consumption model.
Yeah.
I mean, I’m really sorry.
No, no, no. I understand.
Versus the work that your team do, which has an impact on thousands of employees in some of the biggest companies in the world. Exactly, which has downstream impact on millions of Fujitsu users. It is legitimately, definitively less impactful and less valuable work.
Yeah. Again, as with a few times over this, there’s one extreme that says the invisible hand is absolutely accurate, and whatever you’re paid is exactly as much value as you’re creating and exactly as much value as you’re worth. Then there’s another side that’s like, “Everything is the same. Who knows? Who has any idea? It’s all the same.” There is some middle ground.
So, going back, why does it matter that you have a generational company?
After we detoured into Adam Smith, one thing that comes to mind is, “Look upon my works in despair.” Nothing—Ozymandias, and the idea of people obsessing over their legacy and building some statue to their grandeur. One day, it will also fall, right? One day all that will be left are two legs in the desert. That’s true, and that’s true regardless of what you build. At some point, that’s true.
When I think about building Cohere, what excites me about building a generational company is that I mean a timescale of generations. I don’t mean my generations. I mean the idea of building something that is there for a long time. It’s rewarding, and I think it’s inherently human. I think we all like to think about what we’re building, how long it’s going to be there, and whether that’s a work of art, an actual building, a philosophy, or an idea. The idea of building something or participating in the construction of something that is bigger than you is rewarding.
Totally.
And is fundamentally human, even though at some point, yes, it will be two feet in the desert. Both those things are true. It’s rewarding and exciting, ultimately, as with all things.
What’s the biggest disagreement that you and Aidan have had?
We had some disagreements about API design. There was a brief moment before RLHF when we were talking about, “We should make an endpoint for summarization, an endpoint for entity extraction,” or something. I think we disagreed about that.
We disagreed on some low-level stuff, but beyond that, I’ve had the privilege of getting to work with both Aidan and Ivan, the other co-founders. I have a huge amount of respect for them. We definitely disagree and argue about the little things—about how to run the business, whether we should do this, or whether we should make that policy—but there hasn’t been any—
I was chatting to Aravind from Perplexity a month or so ago. I actually chatted to him on stage, in front of 4,000 people or whatever it was; it wasn’t a private conversation. I said to him, “It feels to me like you, Sam, and Dario—all leaders of foundational model companies—you’re basically just presidents who sit on top of the machine shouting views,” because we’re in a shouting-views world, and then everyone in the machine does the work.
He was like, “That’s exactly what we’re all doing. Me, Sam, Dario—you name the leader—we just have to be in front of every camera, doing every interview, basically espousing the views of the organization constantly, because it is so important to be front and center and relevant today.” Do you feel that Cohere is telling your story enough publicly?
But they’re all consumer companies, right? They all fundamentally make money via subscriptions from consumers.
And Anthropic, quasi—I mean, you’d say the majority is enterprise.
I would say a lot of it—most of what it is—is API calls from coders, right? Maybe it’s, I think, 50% Cursor or something, and now they’re competing with Cursor, but a lot of it still comes down to an individual decision of a consumer.
I understand their motivation. If you’re selling to consumers, you want to be telling a story. Consumers are really interested in that story right now, so that makes sense for them. That’s not what we’re doing. You can’t spend $200 a month on Cohere as a person; we don’t have that offering.
It’s not as important for us. Could we be doing a better job telling what we’re doing? When I’m asked to come on here, I’m excited to come. I’m excited to talk to you. I like talking about Cohere, and I think it’s important. Do I think it’s the most important thing? No. I think building a product is the most important thing. I think solving problems for our customers is the most important thing. I think making a better model for them is the most important thing. Telling our story—
You think that’s more important than discussing Adam Smith’s invisible hand with me?
I do.
How dare you, though? I did enjoy it. But, yeah, I do.
12. Should countries fund their own models? Is model sovereignty the future?
That’s hilarious. One area that we haven’t covered, which is interesting, is sovereignty. We’re sitting in London now, and Mistral’s in Paris. I get in so much trouble these days, Nick, because I just—my mother says that I have a spurious butt, but I just call it kind of freedom of speech. We all say that for Mistral, it’s the Europe play, and that’s why it’s funded and continues to be funded. Do you think that we will see sovereign models and usage because of geography?
I think this technology is a lot like infrastructure, right? Having a language model that speaks the language of your country is like building infrastructure for the people of your country.
So I think that's broadly a good idea. I think the past 20 years, or longer, of technological history has been very defined by Silicon Valley. It's been very defined by California and America, and I think a lot of people are not very happy about that. I think a lot of people are rightfully upset with some of the developments, right?
I used to be a real technological optimist. I used to love the way technology was built and be like, “Oh, it's so exciting.” I wouldn't describe myself as a technological optimist over the past 10 years.
Wow. Why? What happened to change that?
Oh, well, wait, sorry. Let me answer that first. Let me get back to the sovereign thing before I go off on this tangent.
I think there's a lot of people who are interested in building that infrastructure within their country and having the technology for their economies. Using a model that is built by China or built within America might not set your country and your economy up as well as having a model that understands the context, is built in that language and in that dialect, and has the cultural fluency needed to empower the people of the country.
I think that's a good idea. What that ends up looking like, I'm not exactly sure. Geopolitics obviously influences a lot.
Do you think geopolitics has influenced customer decisions around sovereignty of models in the discussions that you see?
I think us being Canadian is an asset, right? I think that's helpful for people.
Do Canadian companies want to buy from you more?
I think companies around the world are interested in talking to us, and in part that's because we're Canadian. Look, over the past few years, America has shown that they're willing to turn off access to tech based on political reasons, right? We've seen that the connection between American tech and the American government is less clear as time goes on.
What does that mean? Yeah, I'm a Brit. It means Trump influences US tech companies.
It seems to be, yeah. It seems to be, right? I mean, even last week they announced that they're taking a 10% stake in Intel. That's an interesting development. I'm not an economist, and I don't know if that's good for the country or not, but it is an interesting development, right?
I think there's a lot of companies in Canada and around the world that are interested in working with non-American tech companies, and I would say that's been an asset for us.
Do you think governments should fund sovereign models? Is it a European imperative for us to have Mistral as an asset for Europe?
I think it's a good idea for countries to have infrastructure within their countries. I think it's a good idea for people to have power plants in the country. I like that Canada has several nuclear power plants and several hydropower plants. That's great. I think language models are not that dissimilar from infrastructure.
Do you think our primary input device will still be a phone in 5 years' time?
I do think language—I know language is going to be a more important part of it, right? Fundamentally, the way we should be interacting with computers is using language for the majority of it, not all of it. There are times when language is actually not the best way of interacting with a computer. It's much better to have a graphical user interface if you're doing something.
Last year there was the Rabbit R1 and the Humane AI Pin, and neither of them really got it right. But I think there's something cool there about, “Hey, how do we use a language model to work with a computer better?” I haven't seen it done right yet, and I don't know if it will.
Going back to the technological optimism thing, I was really excited when Google Glass came out. I thought that was really cool.
Yeah. So was I.
Yeah. And then I had it as a profile picture.
Yeah, okay.
Yeah. And then I got on a bus one time and somebody was wearing Google Glass, and they were delighted. Suddenly, everybody saw it immediately. They clocked it immediately.
I was really excited about VR for a while, and then I realized I actually don't want to strap a computer to my face. I'm not interested in being disengaged from the world more. I want to be engaged in the world more than I am. I don't want more things removing me, and I don't think many people do.
Do you just worry about the state of the world? It's so messed up in terms of depression, loneliness, eating disorders, and a focus on materiality for young people. The number-one job that any young person wants to have is being an influencer.
Yeah, there are things that you're talking about that I do worry about. I do worry about the dissolution of community. What I was talking about earlier is that I want to be engaged more in the world. I want the technology that I use to connect me to the world better. I don't want it to disconnect me from the world.
I think a lot of people feel that. I think a lot of people are looking for ways of connecting to technology more, right? I play music a lot. One of the reasons I play music is because it's immediate. It connects you to the people who you're playing with and the people who are listening. A lot of the people who come to listen to the music are there to connect in the moment. That's a really important part of it.
I think a lot of people are feeling that because they're experiencing what you're describing statistically in their personal lives. I also know that worry of, “Oh no, the world these days is so bad, and things are going in the wrong direction, and the kids these days are so weird. Oh no, back then, things used to be better,” is historically ubiquitous, and everybody has always thought that.
Going back to Greek philosophers bemoaning the prevalence of writing because it was going to make people not use their memories anymore, and saying, “Oh no, the kids these days don't understand honor,” or going back to people bemoaning the spread of newspapers because they were all sitting on the bus reading newspapers as opposed to sitting on the bus talking to each other, I think there are 2 mutually exclusive things.
One, yes, I'm worried about all the stuff you're talking about, and I think technology and people who make technology need to think very hard about whether their technology is helping with that or hurting it. Two, everybody has always thought that the time that they're alive is the time when it's the worst, and so you have to hold both of those conflicting views in your mind.
We'll do a quick-fire, but I'll send you off a brilliant song. It's not really a song; it's like a commencement speech by Baz Luhrmann, and it's called “Everybody's Free (To Wear Sunscreen).”
Okay.
It basically says exactly this, which is: Every generation always looks back and goes, “Ah, prices are so high today, and kids are so rude today, and it was better in my time.” It's this kind of continuous pattern of life in terms of looking back and thinking it was better than the one we have today.
Yeah. I mean, look, talking about intrinsically human things like that also seems to be intrinsically human: wanting to be part of something bigger, wanting to build something that lasts, and human beings thinking things were better when they were young.
Okay, we're going to do a quick-fire. If you were Sam Altman today, what would you be doing that he's not doing?
I don't think Sam Altman has done a disservice to the world by talking about how close AGI is. I think he has made several predictions now that are wrong and that were obviously wrong at the time he made them.
Which one is most prescient?
Oh, that AI is going to kill the whole world in 2 years. He's made allusions to things like that. He did a world tour where he spoke to every major leader the world over to tell them, “Hey, this technology is going to pose an existential threat.” I think that was academically disingenuous and did a disservice to the technology he loves.
Do you not see a correlation between the words that one uses with regard to the future of AGI and AI and their requirements for funding? Do you see what I mean? Dario and Zuck, for a long time, do not need funding, and they are much more balanced and neutral. Other people who do need funding have to be much more provocative and out there because they need your dollars.
Yeah. I don't know if that was the strategy. The correlation you're pointing out exists. I would say that we're a venture-capital-funded company, and we need funding.
And we don't say that.
What worries me, though, actually, is that even the rhetoric from your Demis and your Zuck has changed. Even their aggression towards the changes that are coming has flipped, which does make me worry.
Makes you worry because for what reason?
For the reason that we're actually far closer than we think to very material shifts in labor patterns and workforce behaviors, when even Zuck, who does not need the money from anyone, or Demis, who doesn't need it from anyone, is saying, “The changes are real.”
Yeah. Well, there are some real changes, right? I'm not—I don't want to deny that this technology is fundamentally transformative, in the same way the personal computer was fundamentally transformative, as were the Industrial Revolution, steam engines, and the printing press. Those are all big technologies.
I really think that there's tons of legitimate things to talk about, and I'm glad people are talking about them. There's also a whole lot of illegitimate things that people are spending their time talking about.
What is your founder ritual after closing each round?
Who told you about that?
Jordan.
Okay, that's funny. We go to McDonald's.
You go to McDonald's?
We go to McDonald's. Yeah.
And have the same meal?
Yeah. I normally get 2 Junior Chickens.
Wow.
Yeah. I don't remember why that started, but literally after every round, the 3 of us go to McDonald's.
How many rounds have you done?
Oh, 5—4. 5? 4. Yeah.
Wow. I love that. That's awesome. Go to McDonald's.
What's the worst thing that can happen with regulation towards AI?
Out of an erroneous understanding of the technology, thinking that what we're building is digital gods, which large language models are not. But if you think that's what they're building, then you could think, “Okay, cool, we need to come up with benchmarks around existential threats.”
13. Why has Sam Altman actually done a disservice to AI?
I think there are times when looking at those benchmarks—fixation on particular benchmarks, which can be gamed and can be trained either to do way better on or way worse on—is not helpful for establishing how the technology can be used and misused. So I think the worst thing a regulation could do would be to say, “Hey, we're going to pick this random benchmark we think represents AGI, and we're going to shut down any development on it.” I think that would be a misplay.
Do you think China will produce leading models in the next 2 years that continue to beat US models?
No, I'm not sure. They haven't yet, right? They've made good models—definitely good models—but I don't think they've made models that are beating the other models out there.
You're not worried by China and its model capabilities?
When I look at the cadence of China, in a week about a month ago, there were 7 new model providers that released 7 models.
Yeah.
They were pretty good.
Yeah.
And I was like, “Wow.”
Wow.
Yeah. I mean, China is a very interesting country. But am I worried about it? I would say no. I don't think I'm worried about it. I think they're going to keep building models, and those models will be useful. They'll be particularly good at the things that they've trained them on, but that doesn't cause me a bunch of concern.
What's your boldest prediction for LLMs in 2026?
In 2026, you'll be able to open up a computer, log into North or whatever application you're using, and say, “File my expenses.” Then the model will figure out what the expense policy is and where the photos are, and do all of that for you.
That's my boldest prediction. I know that's not very bold. I know, in some ways, that's like, “Oh, that sounds like it's not too far off.” And yet getting it to actually work and be a thing you can rely on is still not in every company. Most people don't have the experience I just described, and I think that becoming a ubiquitous way of using a computer is crazy.
I think my CFO would love you. We're always getting expense chasing, so I totally get that.
If you weren't at Cohere, which AI company would you bet your career on?
Google's great. Google DeepMind is building cool stuff. That's exciting.
Have there been any tools that you've added to your workflow that meaningfully improve your productivity? For me, it's Wispr Flow.
Oh, Wispr. Yeah. I mean, certainly North, our own product. But I use Cursor. Cursor is a great coding application.
Do you mandate Cursor across the whole team?
No, not at all.
Does anyone use Windsurf or Devin?
I think so. I'm not sure. I think people have. We don't mandate one. I know lots of people rely on models.
Are you price-sensitive towards Cursor's costs increasing?
No, but I live a life of privilege, right? So, no, I'm not, as a person.
But when the cost goes up 10x and you have 100 engineers—
Oh, as a business. Yes.
As a business, obviously. How does that shake out? Do we see this kind of—
Well, for us? Yeah. I mean, we make our own models. So if we wanted to use our own model and plug that into an extension of VS Code, that would be something we could do if that was—
Would the quality of output be similar right now?
No, no, no, no. Cursor's built a good product. They've done really good UX work, and it's cool. But if it were to go up 1,000-fold, then of course we'd start thinking about it.
Will we see a trillion-dollar AI company outside of the US in the next decade?
Yeah, maybe north of the border.
Okay, final one. What's one trait you love that has contributed to a lot of your success, but that you're also quite wary of?
I think I'm quite curious and contrarian, and that is both an asset and a hindrance. There are times when it's very helpful to be super interested in something, learn about it, and think, “Everybody thinks this, and they're totally wrong.” That's super helpful.
That's why, when Aidan was like, “Hey, do you want to found a company on language models back in 2019?” I was like, “Yeah, absolutely.” That was not a view that was widespread. Not at all. So, yeah, I'm curious and contrarian.
But there are other times when the whole world has definitely been right and I've been wrong, and I've been like, “Oh, yeah, I was super excited about this—”
When were you most wrong?
When I was very young, I was very much a technological optimist, as mentioned, but I thought all metrics of human improvement were going to continue to go up. Life expectancy would go up, income inequality would go down, happiness would go up. I thought the path towards human endeavors and success was monotonic, and that's not true. There are times when that was something I was really wrong about.
Another thing I was wrong about was the data efficiency of reinforcement learning from human feedback, to give a real technical answer. This was back in 2020, and I remember thinking, “No, you can't make a small dataset of feedback from people and make a model better.” That was a technological misstep.
Nick, this has been an interview of many twists and turns. That's the joy of doing what I do, though, actually—the natural conversation that's inspired. So thank you so much for agreeing to partake in such a wide-ranging conversation.
Thanks for having me.