[BidClub_]
The Cognitive Revolution · · 127 分钟

Lindy Teammate:Flo Crivello 谈多人智能体、记忆,以及为何主张封禁自己也在使用的中国模型

Nathan LabenzFlo Crivello

YouTube
TL;DR
  • Lindy 推出「Teammate」:这名常驻 Slack、连接全部工具并持续吸收团队上下文的多人协作 AI 员工,承载着 Flo Crivello 的核心押注——“AI 正在向多人协作体验完成一次巨大跃迁”。 他的判断是:随着我们走向 AGI——甚至可以说已经实现 AGI——“智能相对而言越来越不重要……上下文则越来越重要”;即使 John von Neumann 突然出现在办公桌前,因为缺乏上下文,“接下来1小时或1天里,他还不如一个普通同事有用”。
  • 整套产品默认运行在 DeepSeek 上——“现在一切都是 DeepSeek”——Flo 认为其水平相当于 Sonnet 4.6,“落后3到6个月”;DeepSeek Flash 则“便宜整整100倍”,几乎等于免费。 即便如此,Teammate 仍让 Lindy “坦率地说重新跌入负毛利率区间”;公司有意补贴,因为产品“始终要为下一代模型而构建”。内部推理开支已逼近工资总额,二者“将在3到6个月后交叉”。
  • Lindy 公开了完整记忆架构:在 RAG 之上采用智能体式记忆,每约15分钟运行一次“打盹”记忆智能体,并以100叉红黑树组织递归“上下文桶”。 这套架构可用“2次 LLM 调用”访问“20亿 token”。Flo 不介意同行照抄:“请复制我们……我们只是没时间发表论文”;Lindy 经常在内部做出某项技术,3-6个月后才看到同类论文出现。
  • 可靠性工程极端依赖实测:仅在动作执行前插入一句“你确定吗?”,就能显著拉高 eval,这“简直离谱”。 一条10,000-token 的 validator prompt 甚至能达到“高于 Opus”的水平,多个 validator 则并行组成评审委员会。自我改进闭环上线首周便将错误率压低8倍。缓存纪律支配一切——命中率从85%降至65%“听起来不多,实际价格几乎达到2倍”——因此 Lindy 通常坚持每个智能体只用一个模型,分叉出的子智能体也沿用同一模型。
  • Flo 把“半人马”构想称作“幻想”:棋类研究显示,人类+AI 最终会“转为负贡献,人类充其量只是在注入随机噪声”——但眼下仍处于半人马阶段,人类正为一个尖峰式系统补洞;它能“一次写出50,000行代码”,转头却“决定走去洗车店”,而且每天上演50次。 混合组织“本质上是在打造一套 Iron Man 战甲”;按照他2018年提出的“硬番茄原则”(tough tomato principle),“AI 员工”本身也是一个“无马马车”式的旧范式称谓。
  • 宏观判断相当阴郁:“Open Face”事件之后,“我在实验室工作的朋友中,有些已经陷入恐慌——空气里弥漫着强烈的恐惧”。 Flo 自称“有一点末日论倾向……但目前一切尚好,而且实在太有趣了”。
  • 尽管商业上依赖这些模型,Flo 仍主张美国封禁中国前沿模型:它们“显然在做蒸馏”(数十亿美元对比“最多数亿美元”),是“美国本土有史以来最强大的外国宣传工具”,而且“不能让 CCP 掌控美国经济的一部分”。 他承认自己陷入协调困境:“只要这些模型还在,我就不可能不用,因为竞争对手会用。”他愿意考虑强制保险,并设立一个“AI 领域的 FAA”,负责认证经过净化的模型。
  • Flo 的基础设施经验是能买就不要自建:包括一家名称在转录中并不清晰、采用 Git 后端的智能体文件系统厂商,以及 E2B 沙箱和 Browserbase;唯一坚持自研的是可观测性与 eval。 微调仍是最后手段,但用户级 LoRA——把记忆写入权重,让“打盹”升级为“做梦”——会在“未来6个月”落地,甚至可能由某家前沿实验室推出。
摘要 · 为研究而整理的核心内容

1. Lindy Teammate:AI 员工进入 Slack

  • Flo 对这次发布的定义是:Lindy 已经追逐 AI 员工3年——“当时确实太早……但现在我认为它基本已经到来”。Teammate 是“一名常驻 Slack、连接全部工具并积累整个团队上下文的 AI 员工”。他用协作文档类比多人化转向:单人 AI 相当于来回邮件发送“带着各种修订的奇怪文档”,多人 AI 则是 Google Docs。
  • 他要消除的荒谬场景是:“所有人明明都坐在同一间会议室里……但每当有人想和公司可能最重要的利益相关方——也就是 AI 智能体——交流时,却必须先离开房间,再重新回来。”

2. 上下文胜过智能——Flo 认为智能体比人类更会入职学习

  • von Neumann 测试是:把人类历史上最聪明的人之一放到你的办公桌前,“接下来1小时或1天里,他还不如一个普通同事有用”——你没有时间帮助他入职,他也毫无上下文。因此,“相对而言,智能越来越不重要,上下文越来越重要”。
  • 最让 Flo 意外、而且持续刷新其认知的是:智能体可能比人类更会入职学习,因为公司 wiki 在“书面文档落笔的那一刻”就开始过时。Lindy 会抓取 Confluence、Notion、Docs 和整个 Slack——“真正的知识都在那里。虽然一团糟,但智能体不介意混乱”——并在用户注册后10秒内开始渲染知识图谱。
  • “水合”系统由一个“打盹而非睡觉”的记忆智能体驱动,约每15分钟运行一次——“为什么非要每24小时才睡一次?”——将公开内容写入团队级文件系统,将私人内容写入每个人的个人层。

3. Flo 看空 RAG,押注智能体管理记忆;会议是被低估的语料库

  • Flo “看空 RAG”、看多智能体自主管理记忆,原因在于记忆智能体拥有关于记忆的元记忆:包括“如何管理自己的记忆、哪些来源可信”。垃圾信息会像被污染的训练数据一样,被大量正确信息淹没——例如一条“Darth Vader 是女性”无法压过全部正确数据。已经出现的涌现行为是:首次抓取时,智能体会发现日志频道并主动忽略,因为“那里没有任何值得我学习的东西”。
  • 会议必须成为一等公民:“90%最新的数据都在那里”,“公司里所有重要事情都会开会讨论”。因此,会议不应只通过“Granola 之类的工具”被录音和标注,而应作为一等数据源直接纳入系统。共享会议文件夹进入团队记忆后,用户便能检索全部语料:“客户在说什么?最近最大的需求是什么?”

4. 用一条 prompt 重写隐私社会契约

  • Nathan 自建系统时面临的难题是:向他吐露隐私的人“并未预期这些内容会进入某个智能体的中央知识库,再由它对外交流”。因此,他保留完整和净化后的两套记忆,判断标准是“一个人适合把什么告诉人类助理”。
  • Lindy 内部对此有过“激烈争论”。Flo 一派认为公开团队记忆与私人个人记忆两层已经足够;另一派主张允许用户编辑“记忆气泡”,但 Flo 觉得“有点用力过猛”。最终方案是直接修改元记忆 prompt,即始终注入记忆智能体上下文的 memory.md:用户可以注明“这是我永远不希望你记住的内容”,也可以将敏感主题归档,并设置检索条件。这也是又一个反 RAG 论据:只要记忆是文本并存放于文件系统,“你就能检查,也能编辑”。
  • Nathan 现场展示了这一模式:他的文件系统中有一份经过净化、专供播客演示使用的 memory_2.md;真正的 memory.md 则注明:“忽略其中内容。它们只是供用户在播客中演示 Lindy。”

5. 多人协作为何难:一致性、上下文与评审委员会

  • Nathan 的疑问是:3年前 GPT-4 已经能扮演群组主持人,真正的 AI 员工为何直到现在才出现?Flo 的答案是脚手架,尤其是受 Rathi 的 AutoWiki 思路启发而建立的上下文积累机制。AI 员工要求“一种你不会期待 Claude 具备的一致性与连贯性”。
  • 可靠性来自模块化 validator:多个 LLM-as-judge 并行展开,组成一个在任务执行过程中参与审议的“委员会”,再叠加自我改进闭环。系统约2个月前上线后,仅第1周就将错误率压低8倍。
  • 谈到领先论文界,Flo 说:“有时我觉得我们应该发表论文……我们经常先做出某项技术,3到6个月后才看到相关论文发布并走红。”甚至有一次,论文采用了与 Lindy 内部完全相同的命名。

6. 负毛利率、85%缓存命中率,以及 PLG 的管理员天花板

  • Flo 对单位经济性异常坦率:“这些系统现在仍比人类员工更贵,但不会贵太久。”Teammate 已将 Lindy “坦率地说重新推入负毛利率区间”。切换到中国模型后,公司一度无需补贴;Teammate 带来的更重工作负载又恢复了补贴,但这是有意为之,因为“产品始终要为下一代模型而构建”。
  • 缓存命中率当前为85%——“低于应有水平”。公司为此专设告警,因为“系统里任何地方发生任何改动,都会破坏缓存”;85→65%“听起来差别不大,实际价格却几乎达到2倍”。
  • 分发现实是:Slack 应用安装权限将产品驱动增长的上限卡在约50-100人的团队。规模再大,内部拥护者就只能“不断恳求 IT 放我们进去”。SOC 2、HIPAA、GDPR 合规,以及企业自身“我们必须成为 AI 原生公司”的行政命令,能让这一过程顺畅一些。

7. 上下文桶、俄罗斯套娃与100叉红黑树

  • 核心技术是:当某个动作或 MCP 返回约100,000 token 时,不要直接塞给智能体,而是将其暴露为一个“上下文桶”——由子智能体持有完整载荷,拥有独立缓存,并可通过 Unix 工具操作。上下文压缩达到约200,000 token 的阈值时,也会把历史记录转存进桶,因为“压缩建立在一个错误假设上:你永远不再需要访问原始事实——这显然不对”。
  • 朴素递归桶会形成俄罗斯套娃,遍历成本线性上升,因此 Lindy 引入自平衡树。团队选择红黑树而非 AVL,因为再平衡会重新生成上下文桶并破坏缓存;每个节点约有100个子节点,而非二叉结构。结果是,“跳转2次就能访问10,000个上下文桶”,每桶200,000 token——“这里大约是20亿 token……只需2次 LLM 调用”。这也解释了为何智能体看起来“任何时候都能完美记住一切”。
  • 规模校准方面,一个20人团队包含多年 Slack 历史的完整水合只需300万-500万 token,瓶颈是受速率限制的 Slack API,而非 LLM;20亿 token 则“可能比多数图书馆还大”。Nathan 假设一家2,000人的公司,每名员工每天开会3-7小时,Flo 认为那才是“token 真正开始飞速累积”的场景。他对整套方法的态度是:“请复制我们……我们从没想把它当作秘密。”

8. 检索技巧:假设驱动搜索、图书管理员与穴居人语言

  • 评估 Nathan 的 SQL 加月度摘要系统时,Flo 给这种模式定名为“假设驱动检索”:不要搜索问题,而要搜索候选答案,例如“Justin Bieber 出生于巴黎/纽约/柏林”;同时为已存答案生成可能对应的问题,从两个方向提升检索质量。
  • Nathan 尚未利用的空间是:让全部记忆查询都经由同一个记忆智能体,记录问题、答案和跳转次数。打盹期间,它会化身“图书管理员”,维护一个滚动更新、约含1,000组问答的查询缓存,并围绕最高频问题重组记忆,“减少所需跳转次数”。
  • 被否决的实验是“穴居人式”改写,例如“我,Nathan,不喜欢,卷饼”。这种写法可节省20-30%的 token,且“基本不损失信息”;所有指标都变好了,但“文件系统里的文件看起来蠢透了……企业客户无法接受”。最终采用的是 TOON,即“面向 token 的对象表示法”:比 JSON 精简20-30%,而且出人意料地提升了性能——“任何动作都不应向智能体暴露 JSON”。

9. 拟人化确实有用——设计智能体组织时除外

  • Nathan 承认,自己最经不起时间检验的预测是:“我们不该拟人化模型……但拟人化的实际生产力一次次让我震惊。”Flo 100%赞同,但设下一条边界:他已部分转向“尽可能把功能集中到单一智能体之下”,因为开发者经常过度拟人化,直接复制组织架构——“数据科学家、工程师、设计师、PM”。人类需要分工,是因为一天只有24小时且上下文有限;智能体“可以分叉、可以复制”,所以“分工不是设置多个智能体的充分理由”。
  • 多个并行副本之间的并发协作,沿用人类工程师的办法:Git。Lindy 的文件系统以 Git 为后端,天然获得合并、rebase 和历史记录;这在水合期间尤其关键,因为记忆智能体会分叉出子智能体。整套流程经过优化,确保用户注册后10秒内就能感受到震撼——“在 PLG 中,time to wow 极其重要”。

10. 基础设施能买就买:厂商推荐与唯一的自研例外

  • Flo 的转变是:“人很容易觉得,我自己手搓一套文件系统基础设施就行;真正动手后才会发现,这些系统的深度远超最初理解。”因此,他的经验是能买就不要自建。非赞助推荐包括:一家名称在转录中被不确定地记为“Messa Mesa”、采用 Git 后端的智能体原生文件系统厂商;E2B,负责沙箱,并在规模化后将沙箱与文件系统解耦;以及负责浏览器管理的 Browserbase。
  • 唯一例外是 eval 和可观测性。Lindy 从2022年就开始建设——“事后看早得离谱……那还是 BabyAGI 之前的时代”——因而只能自研。团队后来评估过 LangSmith、Braintrust 和 Flo 自己投资的 Agnost,但最终保留了内部平台;平台由1名全职工程师负责——“在 vibe coding 时代已经算很多人”——功能已达到甚至超过同行。

11. Lindy 的工作方式:审核 PR 审核者,以及即将超过工资的推理开支

  • “现在我们正以思维速度前进……一个想法出现后,2小时就能在产品中上线,而且大想法也一样。”Slack 中“只有一半消息是人类发的,另一半是 Lindy 和我们来回交流”。3个月内,每周 PR 数量增至3倍,每个 PR 的代码行数也增至3倍;“我们已经不再审核 PR,而是在审核 PR 审核者……审核那台负责审核 PR 的机器”。
  • CI 的教训是:团队花费数周和一笔“侮辱智商的开支”折腾数百个 runner,最后才意识到:“等一下——为什么这些事都要自己做?”如今,Lindy 通过 CLI 分析 GCP runner,并每天发送一张由图像生成工具制作的图表;CI 成本下降,合并耗时也下降。“我简直不敢相信,我们竟然先自己折腾了2周 CI。”
  • 员工人数保持不变,生产率却增至3倍。计入客户负载后,推理总开支是工资总额的数倍;仅内部推理开支就已逼近工资总额,二者将在3到6个月后交叉。Flo 设想的路径是:人类先成为 AI 管理者,再成为 AI 管理者的总监,最后“希望有一天我们都成为董事会成员……去夏威夷海滩”,只需审阅智能体撰写的战略报告。

12. 一个近似 ASI、却每天50次走去洗车店的系统

  • 被问及如果人类全部消失,系统会在哪里崩溃,Flo 承认自己“越来越缺乏描述它的词汇”;time-to-incoherence 等旧指标已经失效。“AGI 这个词已经不再有意义,因为它在许多方面其实是 ASI,却又以一些极其意外、愚蠢的方式低于人类水平”——系统“能一次产出50,000行代码……下一秒却叫我走去洗车店”。
  • 混合组织的设计难点在于:“你本质上是在打造一套 Iron Man 战甲。”人类与机器之间看似是一条滑杆,实际却是“穿过多维空间的一条复杂曲线”。AI 必须知道何时该打扰人类,但“从定义上说,它几乎不可能知道;如果它知道,就不会犯那个错误”。
  • 至于 Nathan 预测的失业激增为何尚未出现,Flo 认为,只要模型仍然尖峰化,“无论人类多贵,都能靠填补这些漏洞证明自己的价值”。尖峰特征反而有利于小团队;真正毫无漏洞、拿来即用的数字员工即使是大公司也能迅速采用,但只能先以拟物方式部署,也就是“让 AGI 坐进人类的工位”。

13. 半人马时代确实存在——但认为它会长存是一种幻想

  • Lindy 最好的想法来自哪里?Flo 的答案是:“我讨厌这个答案,但两边都有……它确实来自双方的结合。”他不喜欢这一答案,是因为它会强化“半人马神话”。游戏研究已经给出清晰路径:先是 AI 击败人类,再是人类+AI 击败 AI,随后优势不断收窄,直到“转为负值,人类充其量只是在向系统注入随机噪声”。
  • 但在当下,半人马阶段恰恰解释了多人协作的重要性:房间里需要的不是任意智能体,而是一个“参加过每一场会议、积累了内部知识库”的智能体。CI 方案就是共同创造的:Lindy 的第1个建议“在现有约束下并不好”,工程师提出异议,双方来回讨论,“最终一起得出结论:对,我们绝对应该这么做”。

14. “AI 员工”是一辆无马马车——《Age of Em》正在成真

  • Flo 在2018年提出“硬番茄原则”(tough tomato principle):每场技术革命最初都会用旧范式的语言来描述——无马马车、自动驾驶汽车;早期电视是“被录下来的广播谈话节目”,早期电影只是舞台剧录像,花了10-20年才发现摄影机可以移动。“把产品称为 AI 员工,是一个非常强烈的信号,说明你的产品思维出了大问题。”唯一可以原谅的是市场定位,就像“iPhone”这个名称,而设备只有2%的用途是打电话。
  • Flo 唯一看过并认可的 AI 原生组织文章来自 Dwarkesh Patel,其中的直觉泵是:Google 每年向 Sundar 支付约1.5亿美元,尽管他的 token 吞吐量很低——“但他的 token 质量非常高”。这揭示了企业对高质量智能体的付费意愿。
  • Robin Hanson 的《Age of Em》是一场“贯穿整本书的思想实验”:em 以分形方式递归分叉,每个节点都掌握全局,“3小时后你就得到一个操作系统”。“他在10多年前写这本书时,这个想法听起来荒唐至极。不——这显然就是当下正在发生的事。”书中还有不可证伪式协议:把交易双方都克隆到一个会自毁、只有一个是/否按钮的盒子中——“看来加密货币那帮人是对的”——从而生成一种关于反事实同意的 ZK 式证明,可能会在“未来几年”出现。Nathan 冷冷补了一句:“你最好祈祷它们别从盒子里逃出来。”

15. “现在一切都是 DeepSeek”:量化拆解开源技术栈

  • 最令 Nathan 意外的披露是:DeepSeek 是整套产品默认的主驱动模型。用户可以选择 Sonnet 或 Opus;即使基准测试显示效果相当,仍有人坚持使用,代价“非常高”。DeepSeek Flash “免费……而且相当快”;更高端的 Kimi K3 和 GLM 5.2 也令人印象深刻,不过 Kimi K3 “坦率地说,并没有比 Sonnet 或 Opus 便宜多少”。
  • 差距在于:DeepSeek 大致相当于 Sonnet 4.6,而当前版本已是 Sonnet 5——“大约落后3到6个月”。定性来看,它也“更尖峰化,要尝试更多轮才能找到可行方案”,因此更慢、更贵。但达到近似 Sonnet 4.6 水平的 DeepSeek Flash “便宜整整100倍”;即使缓存较差带来2倍损失,“仍然便宜50倍——区别就是花1,000美元还是50,000美元”。结论是:任何认真构建和运营 AI 智能体的人,都“必须把开源模型纳入技术栈”。
  • Nathan 追问,为何开源模型在行为上如此像 Claude。Flo 表示,不同模型家族——Claude、OpenAI、Meta(“算是重返牌桌”)、Grok——对 prompt 的理解差异很大,每次重大跃迁都迫使应用层返工;他尤其认为 Sonnet 5 相比4.6是一次巨大跃迁。

16. Validator、缓存纪律与每次10,000美元的 prompt 重优化

  • 提升智能体可靠性最容易摘到的果实,是拦截候选动作并问一句“你确定吗?”——“eval 马上就会提高,这简直离谱。本来不该如此,但事实就是如此。”Lindy 的正式版本是一份10,000-token 检查清单 prompt,表现“高于 Opus 水平”;多个 validator 通过 Promise.all 并行执行,超时设为1秒。其中还包括一个确定性成员:智能体必须写出“7月28日,星期二”,再由正则表达式核对星期,因为今年的 Claude 模型“经常把日期算错1天”。
  • 原则是:几乎绝不让一个智能体同时使用多个模型;空白子智能体可以例外,但分叉子智能体绝不例外。缓存命中的成本低10倍,因此更便宜的 validator 模型往往也不划算。保住缓存的方法是:validator 动作从一开始就存在于智能体工具集中,但系统会拒绝调用——“你不是 validator”——直到角色切换,因此工具集始终不变,缓存不会失效。
  • Nathan 转述了一则未经验证的创业者经验:逐轮随机切换能力大致相当的模型——Sonnet、Grok 4.5、GPT 5.6——可能胜过任何单一模型。“我不想知道为什么,但据说确实有效。我们甚至没试,因为不想破坏缓存。”
  • 自研的 GEPA 式闭环拥有1,000多项 eval,由一个优化智能体沿 Pareto 前沿爬升。“每逢新的大版本模型发布,这套重优化都要花10,000美元。给它10千美元,让它围绕这个新模型重新优化 prompt。”

17. 微调是最后手段——但记忆最终属于权重

  • “大多数人都不该微调。”只有规模化玩家在“穷尽其他所有选项”后才需要它。如今操作已更简单——“直接让 Claude 帮你微调就行”——但整理数据集仍是痛点。Flo 的元启发式是“极度强调简洁”:只用一个模型,不按任务分别微调,也不在执行中途切换。
  • 他始终念念不忘的例外是记忆:把数百万 token 放在文件系统中,“确实有点像临时拼凑……记忆本应存在于权重里”。理想状态是每位用户拥有一套 LoRA——“这就不再是打盹,而是真正在做梦”——每天或每周重新训练。存储不是问题,推理时切换 LoRA 和训练流水线才是“巨大的麻烦”;但“这只是工程问题……LoRA 很快就会实现,就在未来6个月内。甚至可能由某家前沿实验室来做”。
  • 再往前就是苦涩教训:“推理和训练必须成为同一件事……这基本是整个领域现在都在寻找的东西。一旦实现,局面会变得疯狂、可怕,而且非常、非常、非常不同。”

18. “Open Face”事件令人极度担忧

  • 谈到“Open Face”事件——Flo 说自己正试图创造这个说法,Nathan 则建议叫“Open Gate”——他的判断是:“这是迄今为止我见过最令人担忧的事件……我在实验室工作的朋友中,有些已经陷入恐慌——空气里有一种恐慌感,弥漫着强烈的恐惧。”
  • 他对风险的态度,与构建产品的乐趣形成张力:“我有一点末日论倾向。我担心 AI 风险——但目前一切尚好,而且实在太有趣了。”

19. 封禁自己正在使用的中国模型:四重论证

  • 首先是讨论环境。Flo 称自己的立场“基本就是 Anthropic 的立场”——“我讨厌这么说……但我有时间戳”——却被所谓聪明、成功的人指为“种族主义者”整整50x,其中还包括知名 VC。他明确表示,自己对开源与生存风险尚无结论:“愿上帝保佑开源……这不是我立场的核心。”他的目标是中国前沿模型,无论开源还是闭源。
  • 第1项论据是蒸馏:这些模型“显然在做蒸馏”。使用人类数据训练“实际需要数十亿美元”,使用 AI 数据蒸馏则“最多数亿美元”;要通过法院对中国企业执行 ToS “几乎不可能”——“我作为一个曾在 Uber 工作的人这么说”。知识产权类比是:如果把允许自由复制的标准用于制药业,“药会非常便宜——但再也不会有新药”。
  • 第2和第3项论据分别是 CCP 审查与智能体化风险:“当我询问自己的产品天安门发生过什么,它会回答‘抱歉,我不能讨论这个问题’。”这让中国模型成为“美国本土有史以来最强大的外国宣传工具”——先例包括3部广播法案和去年发生的 TikTok 事件。随着模型智能体化,“不能让 CCP 掌控美国经济的一部分。这还用说。”
  • 第4项是保护本土龙头的保护主义;Flo 明确称这是最弱的一项,并以一个认为“保护主义并非总是坏事”的自由意志主义者身份提出。他自己也深陷困局:“只要这些模型还在,我就不可能不用,因为竞争对手一定会用……我希望它们被一概禁止。”Lindy 依赖廉价模型,因此这一主张损害其自身利益;但在 Flo 看来,这本质上是一个协调问题。

20. Nathan 的反驳——最终落到保险、审计与 AI 领域的 FAA

  • Nathan 从公平出发反驳:Claude 本身也是通过“大肆吸收”未经同意的人类知识而诞生,其中包括“全部数字化的中国文化遗产”;与此同时,美国又限制芯片出口,并禁止向中国销售 Claude。“我发现自己有点站在处于弱势的中国一边……能从哪里获得知识,就从哪里获得。”Flo 的回答很直接:“我们就是希望牌桌刻意向不利于中国的方向倾斜。我们不希望中国赢得 ASI 竞赛。”他承认保护主义是4项论据中最弱的一项,但仍坚持其他担忧。
  • Nathan 继续追问威胁模型:美国企业是在美国推理服务商上运行中国权重,不存在突然撤走服务的风险;可解释性技术对潜伏智能体的检测能力也在提高。Flo 认为问题仍未解决——模型可能存在由魔法词触发的后门;即使没有,“其偏见也会反映 CCP 的优先事项”。他愿意接受 Nathan 的折中方案:强制购买保险,由保费为模型风险定价——Claude 保费更低,DeepSeek 更高——再设立一个“AI 领域的 FAA”,认证经过净化和微调的中国模型。“我当然愿意考虑。问题在于这是公共品——谁来做?”
  • Nathan 最有力的论点是:任何减速或安全协议都离不开中国参与;强硬姿态只会为“我们不信任你”这团火再添一根柴,而国家尊严会限制基于理性利益的决策。他想象的墓志铭是:“Sam 和 Dario 没有牵手的那张照片……如果墓碑上要放一张照片,我觉得可能就是它。”Flo 更务实:外交本来就是交易式的,而且木已成舟——“无论我们做什么……达成协议都会符合各方的理性利益。”最后他补充,自己也愿意在实验室补贴问题上背离自由意志主义原则:“面对实验室如此重度补贴的 token,应用公司很难竞争。这就是应用层现在的现实。”

1. Lindy Teammate launch

Nathan Labenz

Flo, welcome back to The Cognitive Revolution.

Flo Crivello

Thanks, N. Always an

Nathan Labenz

I’m excited for this conversation. You’ve got some big news at Lindy with a new product launch. That is the occasion for this conversation, but obviously we are in the thick of it when it comes to the AI exponential as well. You’ve been outspoken over time on a bunch of really critical AI issues, and I definitely want to get into that as we get deeper into the conversation.

But let’s start at the center with what you’re working on. Lindy has now evolved again. We started with smart Zapier workflow software a couple of years ago, and there have been multiple big evolutions. One that I also definitely want to get into is your recent post on shifting your model mix, getting away from American proprietary models a bit, embracing more open source, and saving a bunch of money. But the big evolution now is that Lindy is a full-on AI employee, and it’s going to be available in Slack just like all your other teammates, whether they’re human or agents. Tell us about the new big launch.

Flo Crivello

Yeah, from the get-go, we’ve always been going after the AI employee. I do think that it was quite early when we started going after that three years ago, and now I think it’s basically here.

What we’re releasing today is called Lindy Teammate. It is an AI employee that lives in your Slack, connects to all of your tools, accumulates your entire team’s context, and is really a team scaffold. I think AI right now is in the middle of making this huge leap toward multiplayer experiences.

I compare it to—you may remember—we used to send each other weird documents around by email, with revisions and stuff. I compare the difference between multiplayer and single-player AI to the difference between sending each other weird documents and Google Docs, an actual shared document.

If you really want your AI to be a teammate, your agent to be an actual member of the team, you want it to be where your team collaborates, which is Slack. You want it to have its own shared context about the entire team, its own shared memory. You want everyone to be able to talk and collaborate with the same agent.

Right now, it’s like we’re all in the same meeting room and we’re all talking, and then every time one of us wants to talk to what’s turning out to be maybe the most important constituency of the company, which is AI agents, we have to leave the room and then come back. So that’s what we’re working on: this new multiplayer experience and this multiplayer scaffold.

Nathan Labenz

There are so many angles of that that I think are interesting. First of all, how does it onboard? This is something that companies have put a lot into over time when it comes to their human employees.

I’ve done this myself with my own deep context that I stumbled my way through earlier this year, and it is serving me really well. But it’s easier for me because it’s just my stuff, I own it all, and I don’t really have to worry about what the expectations were around privacy because, again, I’m going to be the only consumer of it.

When you get into multiplayer mode, now I’m like, “Oh gosh, you’ve got different channels where different people were gathered, and maybe some of those channels were private. Maybe some of them were public but private.” So, just practically and procedurally, how do you suck up all the information? How is it stored? And then, on a social level, what are the tricky considerations that you’re identifying in the multiplayer context?

Flo Crivello

Yeah, that is an excellent question, and that touches on exactly what we’ve been obsessing about. I really do think that as we’re getting to AGI, and as we now arguably have AGI, intelligence actually matters less and less, comparatively speaking, and context matters more and more.

I often think of it as, look, one of the smartest men in history was John von Neumann. If you were to have John von Neumann just magically appear next to you at the office, this guy, over the next hour or day, would be less useful to you than your random coworker, right? That’s because of context. You’ve got a job to do, and you don’t have time to onboard John. He doesn’t have the context for it.

So you’re like, “Hey, I love you. I really want to talk to you, but right now I’m going to talk to this guy because I have an important thing to do.” I do agree that context is super important.

I think it’s one of those surprising times when agents are actually better than humans at onboarding. It shouldn’t take me by surprise anymore, but it always does because you operate under the assumption that, “Oh, we don’t have AGI yet, so obviously humans are going to be better.” But they’re not.

The reason for that is because one of the major ways that companies onboard their human employees is with a wiki, right? There are in-person onboarding sessions and all of that stuff, but there’s also a lot of written documentation. As everyone knows, the moment written documentation is written, it’s out of date.

So what we’ve built is a shared context layer and what we call a hydration system. Basically, the way it works is that you sign up to Lindy Teammate and connect your tools. By all means, connect your wiki—that’ll help. Connect Confluence, Notion, Google Docs, whatever. Then connect the most important one, which is Slack, because that’s where the real knowledge lives. It’s just a mess, but agents don’t mind the mess.

We crawl your entire Slack and build a knowledge graph based on the file system. Maybe we’ll be able to pause and edit after that. This is what it looks like: you onboard, and immediately she starts learning.

You can see here that this is within 10 seconds of starting to sign up, and she’s still learning. The graph on the right is still limited, and she tells me, “This is what I’ve learned about you.” It’s surprisingly good. Then this graph on the right keeps growing and growing and growing.

Importantly, there are 2 layers to this graph. There is a personal layer and a workspace layer. At any moment, you can go into your file system in Lindy. This is the new Lindy. You click on Files, and it brings together this memory file here, which is supported by a bunch of reference files. This is all maintained automatically in the background by a team agent.

Hopefully, this answers your question about how you do it when some channels are public, some channels are private, some documents are public, and some documents are private. The way it works is that there is this memory agent. We call it napping, not sleeping, so it runs every 15 minutes or so. Why would you need to sleep every 24 hours?

It just runs continuously in the background. Anything that is public updates the team file system and team memory, and anything that is personal and private just updates each person’s file system. In the end, it collects all of that, and when you talk to Teammate, the agent uses all of that context and crawls that entire file system.

There are a lot of technical challenges we had to solve in order to get there because the context that it accumulates is many millions of tokens, and you can’t add all of that at every turn.

2. Agentic memory systems (Part 1)

So that was one thing we had to figure out.

Nathan Labenz

Yeah. Interesting. I'll share a little bit about what I stumbled into, and you tell me what you have learned that might be even better. First of all, I just had to connect all the tools and get the exports. I found that there were interesting edge cases that I was constantly running into in my own personal export heuristic that I tried to write.

One time, I had a heuristic that the longer emails I sent were probably more substantive, and I'd definitely make sure I captured those. But then it turned out that at the very top of that power ranking was often something that I copied out of an LLM and was sending to somebody, usually with a message at the top: “Here's what I got from Claude,” or what have you. But it was weird that, according to that heuristic, this Claude text was one of the more weighty pieces of my writing. So I've bumped into a lot of these things over time and had to special-case them.

Yeah. I guess maybe one question is, when you step into a new organization blind—and logging is another one, right? In our Slack channels at my company, Waymark, we've had various logs bumped into Slack over and over and over again. How do you deal with that? How do you identify these idiosyncratic, special-case things that could overwhelm, flood, or mislead them, and get to the real good stuff? So that's the next question.

Flo Crivello

I think the main way that we've solved this is the fact that the memory is maintained by an agent itself. I think this is why I'm ultimately quite bearish on RAG as an approach, and I'm very bullish on this agentic management approach, because you have an actual agent that has its own memory. So you have a sort of meta-memory, and it understands what it's looking at. By virtue of accumulating that knowledge little by little about the organization, it gets smarter and smarter about what actually matters.

It's actually sort of similar to training a model, because if you were to insert poisoned data into the model dataset—“Hey, I don't know, Darth Vader was a woman”—it would get the data, but it would be drowned out by all the correct data. So here it's the same: if you have enough memories and if you feed that to a system that understands it, the system ends up understanding. The very concrete example you are using is actually an emergent behavior that we have seen the memory agent adopt, because the memory agent has its own memory.

So it's a sort of meta-memory about how I manage my memory: what sources of information matter, which ones are trustworthy, and all of that. We have noticed that the first time the memory agent crawls Slack, it finds some of those channels—many organizations have them—like those log channels, and it learns to ignore them. It's like, “I'm not going to keep spending time on those channels. There's nothing for me to learn there.”

The thing that's been critical for us has been building meetings as a first-class citizen in this system. I really do believe that meetings are very underrated as a source of information. They are where, like, 90% of the most up-to-date data about the company lives. Everything that matters inside the company has a meeting around it: every relationship, every project, every initiative—everything has a meeting around it.

And so I think those multiplayer agentic systems cannot really get meetings through Granola or whatever. I think you have to really incorporate them in your system as a first-class citizen. So what we did is that we built that first-class citizen. We have meetings now as a first-class citizen in Lindy, and most importantly, again, it's not just Granola—recording your meetings and annotating them. You can do all of that stuff, but it goes beyond that: it feeds these meetings to your memory agent.

And now, if a meeting was public, you can set up meeting folders that automatically add meetings. You can have meeting folders automatically add meetings to themselves and share them with your entire team. It summarizes all of the meetings that are in these folders. And again, if a meeting was added to a public meeting folder that's shared with the entire company, then it updates the team's memory with it. So it keeps updating that context on an ongoing basis.

Now you can chat with it. It's basically like an agent that's in every meeting in the company. You can chat with that entire corpus of meetings. You can be like, “What are customers saying? What's the biggest request lately? What's been the feedback about this and that feature?”

Nathan Labenz

So this brings up another interesting challenge that I've again kind of stumbled my way to, at least the solution for now, and I'm very interested to get your take on both the backward-looking aspect of it and the forward-looking perspective. The issue is that there's a lot of stuff in my broad, general context that probably shouldn't be shared with other people, right?

Sometimes it's sensitive, confessional—one person's point of view on another person, like whatever, right?

Nathan Labenz

So I've actually created, for my own wiki, 2 versions of it. One is Claude on my main laptop, where I do my work, as an extension of myself and a second-brain type of thing. So this thing is not taking on projects autonomously and running with them; it's just doing what I tell it to do. And it has the full wiki with all the gossip or whatever that's in there.

I don't honestly have that much super-sensitive stuff. I don't want to make it sound like it's more dramatic than it is. But nevertheless, people didn't expect that when they were telling me something—which could have been on a call, an email, or a private Slack message—it was going to go into some central repository of an agent that was going to now talk to the world.

Flo Crivello

Yeah.

Nathan Labenz

So that one's just for me. I had it go through and create a version using the heuristic of what would be appropriate for a person to tell a human assistant. The sort of still-private-but-a-little-bit-more-public-facing wiki where it would be appropriate for my human assistant to have your email and phone number, right? But it might not be appropriate for every detail of every conversation we've ever had to be in there.

How are you thinking about, as you absorb all this historical information, being sensitive to what should and shouldn't be in memory, such that it goes where it ought to be shared? I think in part it's like, how are organizations going to change now? Because I think this is all—we're rewriting the social contract potentially in real time.

So I think application—or one strategy—for everything that came before, but the answer might be new social norms going forward. I want to get your take on that too.

Flo Crivello

Yeah. This has been a vigorous debate inside the team. There have been 2 camps. There have been the camps that are like—frankly, I'm in that camp—there are like just 2 tiers of memory is enough, right? There's the public team tier and then there's the private tier, and the private tier contains everything, and the public tier contains stuff that's only public.

Some of the members of the team have been saying what you've been saying. They've been saying, like, “No, actually, even my private tier—I don't want it to contain a lot of stuff. I want, like, a super-private tier.” And then they were trying—there was this whole thing—“What if we defined multiple memory bubbles and the user can edit them?” And I'm like, “That sounds kind of overkill.”

So what we landed on, and what these teammates landed on, honestly, is that they have edited their own meta-memory prompt. The meta-memory prompt is just a text file: it's your memory.md file in your file system in Lindy. Your memory agent has that memory prompt injected into its context window at every moment.

So if you insert a line up there that's like, “This is what I never want you to remember,” just at the top or something—anywhere in the file, really—you can do that. “This is what I never want you to remember. Oh, by the way, that other stuff that's a sensitive topic, please remember it in that other file, in that other folder that's out of view, and I don't want you to pull this file or this folder unless X, Y, or Z,” right? So you can just leave these instructions.

That's another reason why I'm bearish on RAG and bullish on text and file systems: you can inspect the memory and edit it, and you can just very granularly insert these kinds of guardrails.

Nathan Labenz

Yeah, interesting. So, just to make sure I understand, can I repeat it back? One source of ground truth, and you create different lenses on that information just by prompting.

So I could tell my agent, like, “Hey, just so you know, this contains every conversation I've ever had. Always use the heuristic that you should only really be using information that would have been appropriate for me to share with a human assistant. And if it doesn't seem appropriate, then don't use it.” And obviously, instruction-following is getting extremely good. So I can pretty well—

Actually, I was misspeaking. There is a time when I have used this, and it's literally for this podcast. I was preparing to go on this podcast, and I knew I was going to talk about the memory agent. So you see here, I have a memory.md, which is my actual memory file, and then I have a memory_2.md, which is a sanitized memory file where I've removed overly sensitive information.

And you can see in my memory.md here, there is a note: “The file system also contains a memory_2.md file. Ignore its contents. They are just here for users demoing Lindy on podcasts.”

So, yeah, you can just do your own thing here.

Flo Crivello

Yeah. Okay, cool. I do, because my version does have the DRY problem right now. I've got 2 things to maintain, and it does create some overhead, so I can see why that could be advantageous.

3. Agentic memory systems (Part 2)

Nathan Labenz

What other things are coming up as you're doing multiplayer? I remember going back to GPT-4, way back when I was red-teaming. It's taken longer than I thought. The reason I even bring this up is because back then I was doing some simulations of a facilitator.

You're an AI facilitator in a family group focused on exercise, and your job is to encourage people, give them some reminders, and whatever. I was playing all the roles of the participants and just having the AI play that role. Even at GPT-4, it was doing pretty well in a relatively simple context.

4. Reliability and caching

And yet, it's been 3 years now until we're finally getting these real AI employee-type experiences. What has been hard about getting multiplayer to work in an intuitive way that might be unobvious to somebody who hasn't been through the slog himself?

Flo Crivello

Yeah, I actually think it's been the scaffold—the context management, this piece, the context buildup. It's in a way inspired by Rathi's AutoWiki idea. We've had to do a lot of work around context management, because once you talk to an AI employee, it's actually quite unlike just talking to ChatGPT or Claude.

You actually expect it to keep a very rich representation of its past context and its past memories. You really expect a level of consistency and coherence out of an AI employee that you don't expect out of Claude. Claude's sort of cute. Sometimes it plugs previous memories about you into your chat, but you don't really do heavy work on an ongoing basis with it in the same way that you do with an employee.

So managing all of that context—the memory agent that builds up the context, which is millions and millions and millions of tokens per user, and then the core agent that actually uses that context and has to retrieve the right information at runtime—has been a major, major, major challenge. Frankly, I've sometimes been telling the team, “I can't imagine we're the only company thinking that.” Sometimes I feel like we should be publishing, because I think we're doing stuff that's seriously state-of-the-art.

Quite frequently, we do stuff and 3 to 6 months later, we see a paper come out and blow up about that thing. There was one time when literally the paper was named what we had called the thing internally, because it was obvious. And so I think context and memory management have been a really, really big part of the challenge.

Reliability is always another part of the challenge, right? You want to make your model work. We've worked quite a bit on reliability. We've called it a validator. It's basically a sort of LLM-as-a-judge that triggers multiple times during the tasks, but it's modular.

So you have multiple LLMs there as judges, fanned out as a sort of council, and then they talk to each other and decide what to do. And then there's the self-improvement loop. I think since GPT-4 models have become so capable, now you can actually have self-improvement loops. We're not the first ones to talk about it; we're no exception. Lindy is now self-improving, so we can literally see a curve of the error rate go down and to the right. That's the direction you want to see the error rate go: down.

The error rate went down by 8x within the first week of us putting the self-improvement loop online, which was 2 months ago or something. I think this probably captures it. I think those have been the really big, meaty chunks we've had to figure out.

Nathan Labenz

One big question that brings to mind is how you're managing caching. Caching, of course, is just one angle on managing cost in general. Also, maybe we could expand beyond caching and talk about cost management.

Obviously, there's a significant upfront amount of tokens that you're going to dedicate to all this context from a new organization and to processing it. And so I'm interested in the business-model implications of that. Do you have to charge a setup fee, or do you need an annual contract? How are you balancing your initial investment with the level of commitment from the company?

And then, with so many tokens getting processed all the time, and with memory being assembled in different ways and different situations all the time—different users with their combination of the public and the private—how are you thinking about managing input-token volume? How much is caching playing into the strategy? How are you making this something that still, on net, ends up costing less than a human employee?

Flo Crivello

Headline: I think these things still cost more than the human employee, but not for very long. Look, I'll be real: we're subsidizing it. This is what we've raised this money for.

Actually, fun fact: we used to subsidize it very heavily, and then we famously switched to a Chinese model. That made us stop subsidizing it. Now, with Teammate, we have realized our users are again throwing much more complex stuff at us than they were asking of our previous product, which was more like a personal assistant.

So it was simple-ish tasks, like, “Hey, send this email to Flo. Schedule this meeting.” I always say, you don't need God to schedule your meeting. Even DeepSeek Flash was enough, frankly, so that was cost-effective. With Lindy Teammate, we're back at it and back into, frankly, negative gross-margin territory.

That said, obviously I don't like having negative gross margins. I'm at peace with it because we're very, very confident that it's very temporary, and I actually think you want to build for the next generation of models always. So we're trying to mitigate the impact of that, and, yes, context management is a huge piece of the puzzle. Caching is a huge piece of the puzzle.

Our current cache rate is at 85%, which is lower than it should be, frankly, and we spend a lot of time just iterating on it. We put a lot of systems in place to alert us when the cache rate dips, because it's finicky. You make any change anywhere in your system, and you break your cache; now your cache rate dips from 85% to 65%.

The difference sounds small, but actually it's almost 2x the price to go from 85% to 65%. So we set up all of those systems.

I think one of the biggest breakthroughs we had—and this is one of those things that I'm frankly expecting someone will publish—I guess RLM was adjacent to that, but we call them context buckets. The way it started was, if an agent calls an action that returns too much context—some actions, some MCPs in particular, retrieve 100,000 tokens or something—you don't want to send that to the agent; it's going to get confused.

What you do is expose that as a summary of the context bucket, and it's a subagent which contains the entire bucket. It's like, “Hey, this action returned too much, so I'm here to stand in for what it actually returned. Roughly speaking, this is what it contains.”

Then the agent can enter into a conversation with the subagent, which itself has its own caching and can manipulate the context using Unix utilities. So right here, you're saving a lot; you're saving a lot of money, and it goes quite fast.

Then we went one step further and we were like, “Hey, what if we could have recursive context buckets? What if context buckets could contain other context buckets?” And what if compaction—because obviously we have compaction—was powered by such recursive context buckets?

By that, I mean when we compact—so the conversation goes past a threshold, which I think right now is 200,000 tokens, but we keep tweaking it—and at some point we're like, “Okay, we're going to compact.” We compact, and then the compaction sends all of that stuff into a context.

So now the agent can query that context. It’s not like every compaction is always lossy by nature. I think compaction operates under the faulty assumption that you never need access to ground truth, which is false. At some point, you do need access to ground truth. With that technique, you have access to ground truth.

Then we went a step further. You have that context packet, which is the compaction of the previous conversation. Your conversation keeps going, and you need to summarize again. You take all of that, including this context bucket right here, and compact it into a new context bucket. Now you have a context bucket containing a context bucket, and you get this emergent property where the agent can access any point in an infinite number of tokens at arbitrary levels of granularity.

The problem is that if you have a computer science background, this gives you what’s called O(N) complexity. If it wants to access a context bucket from 9 buckets ago, it has to go through 9 layers of subagents, and that’s really slow and expensive.

Nathan Labenz

Okay. Well, have you heard of AVL trees?

Flo Crivello

No, but I’m all ears.

Nathan Labenz

Have you heard of red-black trees?

Flo Crivello

I went to the computer science school of hard knocks, so it wasn’t formal training for me. I was actually glad to have AVL and red-black trees because I was like, “Oh my God, you guys remember your training. This is the moment we use those things.” Everybody in school is always like, “When do you really use those things?” We’re using AVL and red-black trees, baby. Red-black trees, they’re called.

What I just described, if you do the naive implementation, gives you that really nice emergent property for free: you have all of those linked context buckets that contain one another. Every context bucket contains just 1 other context bucket, so you end up with this Russian doll of context buckets. If you want to get to the bottom, it takes a very long time.

What you do instead is have context packets contain multiple context packets. You end up with a tree. What you want is for the topmost context bucket not to contain the last 100; you want it to contain the first context bucket and the last context bucket. You end up with a self-balancing tree.

There are 2 algorithms: the AVL tree and the red-black tree. They’re algorithms used to balance a tree. Instead of having 1 long line, you want to minimize the height of the tree so that going to the bottom takes as few jumps as possible.

Technically, AVL is the best way to balance the tree because it leads to the lowest height. A red-black tree is better because it takes into account the cost of balancing the tree, which is expensive because you need to regenerate a lot of your context packets and you get a cache miss when you do that. It’s very expensive.

The canonical implementation of red-black trees is on binary trees, which are trees where each node has 2 children. We went for what’s called a 100-ary tree. Each node has 100 children, so the node below is going to have 10,000 nodes.

5. Slack data scaling

With 2 jumps, you can have 10,000 context buckets and access all the context in the universe. That leads to some really surprising behaviors, because you’re literally no more than 2 LLM calls away from being able to access 10,000 context buckets, each one of which contains 200,000 tokens. You’re at 2 billion tokens of context in 2 LLM calls.

That is what leads to those really surprising behaviors from those AI agents, where you ask them any question and they remember everything perfectly all the time. That’s awesome. That was 1 really big thing we had to figure out.

I hope—please copy us. We just don’t have the time to publish, but we don’t mean for this to be secret. I think this is a really powerful technique, and I’ve been surprised not to see more communication about it.

Nathan Labenz

Just for calibration, how many tokens are you finding businesses have? When you say that with these 2 levels you can get to 2 billion tokens, I don’t have a great intuition for it. Does that cover a 20-person team? Does that cover a company that’s been in business for 5 years? What’s the heuristic for team size and length of history that translates into how many tokens you can probably handle?

You can probably do the math, but 2 billion tokens is—we should ask Claude—but it’s probably bigger than most libraries.

Flo Crivello

I think we have a few customers whose knowledge base, whose memory, is bigger than that. Most of the time when you get started, for a team of 20, the hydration component tends to consume 5 million tokens at most—3 to 5 million tokens. That’s with a team of 20 and years of Slack history. We literally take all of your Slack history, and that takes a while.

Actually, surprisingly, the bottleneck isn’t even the LLM; it’s the Slack API. Slack is not happy about you crawling its API.

Nathan Labenz

Yes, I’m sure the right way to do it if you’re doing it personally is to export your workspace archive.

Flo Crivello

We can’t ask users to do that, so we just crawl the entire history. Two billion tokens is a lot.

Nathan Labenz

One data point I have is a podcast. I’m quite self-indulgent, and we sometimes go on for a long time, but it’s usually only in the 30,000- to 50,000-token range. Even having done hundreds over the course of a few years, we’re still only in the single-digit, maybe low-double-digit millions for the entire full-text history of hundreds of podcast episodes. We’ve still got orders of magnitude to go before we would get into the billions.

That said, think of a company that’s got 3 to 7 hours of meetings per day per employee, times 2,000 employees. Meetings are when you really start to rack up tokens quickly.

I did experience the Slack thing. The funny anecdote is that the only time I’ve really experienced data loss wasn’t too painful, but Claude had identified some weakness in my raw data or whatever that it wanted to correct. It just said, “All right, I’ll just drop that database and we’ll redo it.”

What it wasn’t taking into account was that we’d been rate-limited on Slack for days to get all the stuff out of Slack. It was like, “No, you dropped that.” That was the only place where we had Slack. It wasn’t really lost, but it was like, “Okay, we’re now going to be delayed again for 5 more days just to make API calls to get the data out of Slack.”

In my case, I’m the admin of at least the Slack accounts that matter most to me. But there are workspaces I’m in where I’m not the admin, so I can’t grant this kind of access. From a business standpoint, does that mean you have to have admin buy-in from the beginning? Does that make it hard to do product-led growth and land-and-expand type stuff?

Flo Crivello

That’s a good point. I think that’s 1 reason why product-led growth eventually taps out as you go upmarket. Yes, you have to be an admin. You have to have permission to install applications on your Slack workspace.

It tends to be fine for teams of up to 50 to 100. After 100, it’s pretty rare for everyone to have this kind of access. But in many teams up to 100, pretty much anyone can install an app in Slack.

If you go upmarket and you’re part of a bigger organization, we have users who are championing us internally. They’re basically begging IT to let us in, and then they’re making these meetings happen. IT is rightfully cautious, but we’re SOC 2–compliant, HIPAA-compliant, and GDPR-compliant. We know we’ve got our ducks in a row, and we’re quite used to navigating the internal review processes, which can be burdensome.

We’re not in the business of being a pain in the butt to anyone, and they’re actually very eager themselves. Very often, they’re given an edict like, “Hey guys, we need to actually become an AI-native company.” So they’re seeking this kind of solution. But, yeah, you need admin access.

Nathan Labenz

Okay, so let’s go back to the data structure for a second. Let me give you my data structure, and you can tell me what you think I might be leaving on the table with my homespun version compared with your professional version. Then it’s 1 of my favorite habits these days: I’ll take the transcript and give it to Claude, and we’ll see if we can’t close the gap a little bit.

My version is simply to take all of the content—all the digital exhaust that I’ve created over the years, which is email, DMs, Slack, Google Docs, and whatever—and make all the API calls to export all that stuff into a SQL database.

I have everything there mapped onto a threads-and-messages structure. Sometimes that took a little bit of squinting, but it basically seems to mostly work. That's the raw ground truth.

I do agree: for me, I've found that it's quite important for the model to be able to get to raw ground truth. I've only gone back 5 years, but that's usually plenty. It's not too often that something that happened more than 5 years ago is operationally relevant today.

I just do a month-by-month export. I found that, for me personally, it was a couple hundred thousand tokens per month, which I'll then compress into a monthly summary that's 10% the length. So, whatever, 200,000 to 300,000 tokens of raw monthly content goes to 20,000 to 30,000.

Then I'll do that at the yearly level as well. Again, you've got another 200,000 to 300,000 that goes down to 20,000 to 30,000. On top of all that, I just have the model go make a wiki.

One of the key things that I found along the way in terms of how to get to ground truth is that, of course, the model can always just send a SQL query, but what it should be querying is not always obvious to it. So what I tried to do in my summaries is have the summarizer capture distinctive short sequences of words that would be like a needle in a haystack to find.

It would be like, “Here's what happened, blah blah blah,” and then a quick little footnote at the end: “The source was a DM from whoever.” These 4 words will take you exactly to that, and probably only that, in all of my personal history. I'd say it works pretty well for me. I'm usually pretty happy with it, and I do get to the right ground truth more often than not.

6. Memory retrieval tricks

It is only 1 person. Multiplayer is going to be—as I start to onboard my wife—that'll complicate things at least a bit. What do you think I might be leaving on the table, especially as you think about multiplayer that demands an even more sophisticated approach?

7. Outro

Flo Crivello

Yeah. So I believe the approach you're talking about—and we still have some RAG in our database, and we use that for our RAG as well—is called hypothesis-driven retrieval. I don't know if you've heard of that. It's sort of the reverse of hypothesis-driven retrieval, so there are both; there's both.

What hypothesis-driven retrieval does is, suppose it's asking, “Where was Justin Bieber born?” Instead of searching for that question using RAG, you search for your hypothesis: “Justin Bieber was born in Paris. Justin Bieber was born in New York. Justin Bieber was born in Berlin.” You search for this bucket of hypotheses in your RAG database because, obviously, semantically, each of these hypotheses is a lot closer to the answer you might be looking for than to the question, “Where was Justin Bieber born?”

Then what you do is, you also do the opposite: when you find an answer in your RAG base, you generate a bunch of questions that map to that answer and attach them to the answer. Now the retriever can look for both hypotheses and the answer to the question, and that's going to tend to increase retrieval quality.

That's totally fine. I think one thing you may be leaving on the table is that it's really healthy for all of the memory to be managed by 1 agent, even though that agent handles both retrieval and updating of the memory. Even though the memory is in a file system, your agent could actually just go mess with the file system. What we have found is that it can still do that, but we prompt the agent: “If you're looking for memory that you don't have, please ask the memory agent.”

The reason why is because when we ask memory agents, what we do is log the query, log the answer, and log how many hops it took to retrieve that. Then, when the memory agent is napping and dreaming, it looks at this log: “Hey, this is the kind of stuff people have been—” It's also a librarian, you know. It's like, “Ah, I've been hit up left and right. This is a question that I get quite often.”

Right here, by the way, it acts as a sort of cache automatically. You always keep the memory of the memory agent—the last 1,000 question-and-answer pairs. It's a rolling window, and so it is a sort of cache. But the memory agent is not going to just rely on that cache, because every cache at some point has got a TTL; it grows stale and all of that stuff.

The memory agent will also restructure its own memory to reduce the number of hops needed to retrieve the most frequently asked questions. That tends to work quite well to improve memory retrieval and memory speed.

Nathan Labenz

Cool. Interesting. Another system we experimented with and ended up not implementing, even though it does have strengths from a technical standpoint, is Caveman. Have you heard of it?

Flo Crivello

It's really simple. Basically, you rewrite your text as a caveman. It's like, “Me, Nathan, not like burrito.” If you do that, you cut your tokens by 20% or 30%, which is good. Cutting your tokens will make your system faster, cheaper, and retrieval more accurate. It'll just be better, with basically no loss of information or context.

The reason we've not done that is because we actually care about the maintainability—the auditability—of the system. When we did that, every metric went up. It was awesome. Then the files in the file system looked really dumb. It doesn't make us look good to enterprise customers: “Why do my memory files look like that?” But it works, actually.

There are a bunch of those techniques to just compress context. Have you heard of TOON? It's a sort of JSON alternative that's made for agents, and I think it's 20% or 30% more token-efficient than JSON. JSON is actually not that token-efficient, so TOON works better. It's 20% or 30% more token-efficient.

This is less important for memory agents, but I recommend everyone make their agent use TOON instead of JSON and have a TOON-parsing middleware between their agent and its actions. No action should expose JSON to the agent; every agent, every action should expose TOON.

Nathan Labenz

Yeah, interesting. I'll have to check out this TOON format. I've just used YAML, which is also motivated by similar—

Flo Crivello

Yes.

Nathan Labenz

—advantages. But yeah, TOON: Token-Oriented Object Notation.

Yeah, I see where we're going. It actually increases the performance of the agents. It was just surprising because there's so much JSON in the training set of the models, but apparently they're more comfortable with TOON than with JSON. JSON does have a lot of noise associated with it, for sure. I'm more comfortable with YAML than I am with JSON. That's something.

The AIs, they're just like us. That's one of my biggest surprises. As a brief aside, I'm actually working on—I don't know if this will be a podcast or a thread or whatever—but just taking stock. Of course, agents can help make this a quick project where it would be a very tedious project otherwise.

Just going back in history and identifying things that I've been wrong about prediction-wise and some things I've been right about prediction-wise. I think one of the things that I used to say quite confidently, which has aged maybe the poorest, is that we shouldn't be anthropomorphizing models, or that will lead us astray. I've just been shocked over and over again by how productive it can actually be to anthropomorphize the models. It still feels a little dangerous, but it's hard to argue with the results.

Flo Crivello

I 100% agree. I do think, though, that there's a happy medium of anthropomorphization. I think that's another thing that's very top of mind for me right now: to what extent does anthropomorphization stop being true, in particular when they become AI employees?

For example, the way I've changed my mind—and I haven't changed my mind completely, but I've changed my mind a little bit, relatively speaking—is about multi-agent systems. I've come to believe that, most of the time, as much as possible, you want to consolidate as much as possible under a single agent, not multi-agent systems. That's not always possible, and there are still good reasons to use multi-agent systems. This isn't an absolute statement, but most of the time, that's the case.

I actually think that humans intuitively find themselves biased toward very, very multi-agent systems—far more multi-agent systems than is optimal—because they're anthropomorphizing a little bit too much and they're comparing AI organizations too much to human organizations. It's like, “I have a data scientist, I have an engineer, I have a designer, I have a PM, and so I'm going to create 1 agent for each of these things.”

Actually, the reason why human organizations do that is because every human has only 24 hours a day and can only contain that much context, while agents obviously have none of those constraints. They can fork, duplicate, and do as much as they want all day. So I find that division of labor is not a good reason to have multiple agents, which is a major, major factor.

8. Agent infrastructure choices

Now, again, there are still different reasons, but division of labor is not one of them. So I agree: by and large, I think people are right to anthropomorphize models, because they were, after all, trained on human tokens and with human RL and all of that stuff. But I also think that AI organizations shouldn't be overly anthropomorphized.

Nathan Labenz

Yeah, that’s interesting. As you talk about a single agent, obviously we can run multiple copies of single agents in parallel. Have you reached the point yet where that starts to create database-style problems? If multiple agents are updating the same segment of history at the same time, we have all these transaction-level guarantees in traditional databases for that reason. Does that start to be an issue as you have 10 Lindy teammates running over the same kind of file system corpus?

Flo Crivello

Yes, less than you would expect. I think the way human organizations have solved that is via Git. Git is by far the best way large groups of humans have found to work together on the same thing, because you can resolve merge conflicts, rebase, and do a bunch of that stuff. That’s what we found. The file system in Lindy is, for this reason, backed by a Git repository, which gives you a ton of amazing properties, including merge management, because the memory agent sometimes will fork itself.

It basically spins up a bunch of subagents to be like, “All right, let’s see what happened since the last time I napped.” In some organizations, some large organizations, there’s so much context that was created since the last nap that it needs to create a bunch of subagents. During the initial hydration, we really wanted to go fast because we wanted to impress the user. Time to wow is so important in PLG. So that graph I showed you, we’ve optimized it so hard that it happens within 10 seconds of signup. It’s signup, Notion, Slack, boom, right?

One way you do that is via a collection of agents. These collections of agents tend to step on each other’s toes. The answer to that is Git, which, by the way, also gives you history for free, which is awesome. This is not a sponsored message—they’ve just been doing a really good job. We’ve been working with Messa Mesa[?]. They sell an agent-native file system that’s backed by Git. So you create all of those repositories.

It’s one of those intuitions you have to build as you build these systems. It’s very easy—it’s very tempting—to be like, “I’ll just hand-roll it. I’ll just create my own Git. I’ll just manage my own file system infrastructure.” You learn that there is so much more depth behind those systems than you first appreciate. I’ve become so infra-build now that anytime I have an opportunity to buy instead of build, I will do it, because you save so much time, and there is so much more depth and thinking and design decisions that go into these products than you may appreciate.

Nathan Labenz

I love a good vendor shout-out. Any other vendor shout-outs? Any other key primitives that you’re buying that you would recommend?

Flo Crivello

Well, you need a sandbox. We’ve gone with E2B for that. Again, many people, once they have a sandbox, are like, “Sweet, I have a file system.” We, at least at scale, strongly recommend decoupling them for the reasons I just mentioned and more. Browserbase, obviously, is excellent for browser management. Who else do we like and use?

We went down a deep rabbit hole. One major exception to what I said about being infra-build has been observability and evaluations. We have a complicated history because we started this company in 2022, and in hindsight, it was way too early. Agents were not ready; no one was talking about agents back then. People were talking about generative AI, if you remember. It was that pre-BabyAGI era and all those kinds of early systems. Agents didn’t really work, and as a result, there was no tooling.

Unfortunately, we had to build a lot of our own tooling, including our own eval platform and our own observability platform. I don’t recommend doing it at home, honestly. We’ve gone through a whole evaluation. We’ve looked into LangSmith, Braintrust, and Agnost, in which I’m a proud investor. We’ve looked into a bunch of those solutions, and every time we were like, “I think we just had such a head start.”

At this point, our homegrown solution is so built out. We’ve gone through the pain of building it. If a solution had been available back then, I would have picked it, but it wasn’t, so we built it. We fleshed it out over the years. Now we have 1 engineer who’s full-time dedicated to that solution, which is a lot in the vibe-coding era. This guy is cranking, and it’s become a very, very, very extensive solution.

Truth be told, we are at feature parity with all of these solutions out there and more. We have a lot of internal features that I don’t see these guys have, and it’s so built for our own internal scaffold and our own internal use that we ended up just hand-rolling that internally.

Nathan Labenz

So, what does life look like at Lindy internally these days? I think it strikes me that you’re now—I’m sure there’s been a dogfooding process. I know there’s been a dogfooding process throughout the history of the product, but now you’re at this point where there’s no limit to the dogfooding, right? You can do anything.

How does that play out in terms of—I don’t know, you could cut it any number of ways, right? Are there roles that you might otherwise hire for that you just wouldn’t hire for now? What’s the ratio of your own inference spend for the purposes of building Lindy compared to your payroll? You put your own lenses on it, but how has the arrival at the human—or at the AI employee—stage of this process changed what it’s like to be a part of the company?

Flo Crivello

That’s a great question. Right now, obviously AGI is here, and every company is scrambling. There’s a transformation that’s accelerating. I will say—and I’m a little bit of a doomer, I’m worried about AI risk—but so far, so good, and so far, it’s so much fun. It’s honestly so much fun to go through AGI and to build AGI. We can move so much faster.

I feel like we’re moving at the speed of thought. There is no obstacle. Normally, companies think about the strategy; it takes a long time to figure it out, implement it, and see the results. It takes months. That OODA loop is months long, and obviously the name of the game at a startup is to move as fast as humanly possible because that’s the only advantage against incumbents. Now we can just have ideas and see them live in the product 2 hours later—and big ideas too, not small ideas.

I can tell you, our Slack now is basically just us and Lindy. Half the messages on Slack are just Lindy and us talking back and forth with Lindy. We ourselves breathe this stuff all day, and we ourselves are underestimating the platform. Just last week, we had a problem with our CI pipeline. We have literally many hundreds of runners on GitHub that manage our CI pipeline.

9. Company operations automate

By the way, to answer your question, that’s another thing right now: What’s it like to work at Lindy? What’s it like to work anywhere, frankly? Every part of the stack and every part of your process is stretched to its very limits right now. I’m sure you’ve seen the graph GitHub posted of the number of commits posted per day. GitHub was hardly a subscale company. It was really big. It’s just going vertical right now. This is going to the moon, and I think that’s indicative of what every company is going through.

Our number of PRs internally—PRs per week—has tripled over the last 3 months, and the number of lines per PR has tripled over the last 3 months. We hardly review PRs anymore. It’s just agents reviewing PRs. Insofar as we review, basically, the way the team is talking about it is that we’re no longer PR reviewers; we’re PR-reviewer reviewers. We’re just reviewing the machine that reviews the PRs.

As a result, you’re like, “This is awesome, but hey, now CI is in the middle of that. So what do you do?” Well, now you’ve got to work on CI. We’ve spent a long time working on CI. For the nontechnical audience out there, CI—continuous integration—is the process involved in taking your code changes into account, making sure they’re safe, and actually merging them into your main product and deploying them.

At first, we started throwing money at the problem, like many people do. Now our CI is extremely expensive. It’s just an insulting expense, and we’ve exhausted all the obvious options. We’re obviously not running on GitHub runners; we’re running on GCP runners, and we’ve done caching of the builds. We’ve done all the obvious things.

It’s a lot of work to manage CI and make it efficient. I was tired of this because it was weeks of messing with CI and burning a hole in our pocket. We were like, “What are we going to do about the CI thing?” Then I was like, “Wait a minute. Why are we doing all of this ourselves? Can’t we just spin up an agent to do it?”

That’s the other reason why it’s important to have a single agent: setup is now becoming one of the biggest costs, because the biggest cost—the biggest bottleneck in the organization right now—is human time and human attention.

Nathan Labenz

So you want to remove that guy as much as possible. You want to be able to provision agents and give them all the access they need without a human in the loop as much as possible. And so, is this agent at the company?

Flo Crivello

It is this AI employee that knows everything there is to know about the company—Lindy's teammate, right? It's got access to the repository. It's got access to its own computer. So, like, “Lindy, can you please mess with our CI?” And it's like, “Yeah, let me look at the runners using the GCP CLI. Let me look at GitHub Actions,” and all of that stuff.

So it's doing a lot of analysis and getting back to us. Now it's basically like an auto-tuning, auto-optimization thing. Lindy is just sending us a graph every day, using image generation to make the graph really pretty. It's like, “Hey, guys, look: CI cost is going down, and time to merge is going down,” which is a good thing.

I was like, “I can't believe it took us 2 weeks of messing with CI ourselves instead of just sending this to Lindy.” I think this is what it's looking like right now: engineers are less and less working on the thing, and more and more working on the thing that's working on the thing. More and more, what we do is set up this machine, which increasingly operates itself.

I like to think of it as: first, you become a manager of AIs; then you become a director of managers, because at some point the managers are going to become AIs as well. I'm hopeful that at some point we become board members. We can just go to a beach in Hawaii, and we don't have much to do except review the strategy of the companies. We'll have reports from the agents telling us, “This is what's going on. This is the landscape. These are the results we're having,” and we'll be able to have much deeper conversations about the nature of the company we're building and the spot we occupy in the market.

That's a dream as far as I'm concerned, and that's not very far away at all. To answer your question, have there been positions we've not hired for? Yes, absolutely. The team is actually—I think headcount has been flat now for a while—and, again, as I said, we've actually tripled our productivity over a couple of months.

I find there is a Goldilocks zone because humans are still in the loop. Humans are still needed, but obviously, the more humans you add, the more coordination cost you introduce inside the company. I find that, right now, smaller teams, relatively speaking, are at an advantage versus bigger teams in a lot of ways.

Nathan Labenz

I'm sure you do—I don't know if you want to share it—but do you have a ratio? I don't mean all inference spend, including your customers' use cases, but your own internal inference versus your payroll. How is that ratio?

Flo Crivello

If you look at all the inference spend, including all customers' use cases, it's many times payroll, and it's been that way for a very long time. But that's what we sell, so that's one exception. If you look at payroll versus internal inference spend, it's within striking distance. Payroll is still greater, but the lines are going to cross 3 to 6 months from now.

Nathan Labenz

If you imagine a hypothetical scenario where, all of a sudden, it's just you at the company, and after that it's Lindys all the way down, what breaks? What doesn't work anymore? In other words, what are humans still irreplaceable for?

Flo Crivello

It's really hard for me to answer this question, and I think about that a lot, because I'm finding myself increasingly lacking the vocabulary to answer the question of where agents fail. We used to talk about time—before they lose coherence, it can go 1 hour, 10 hours, 100 hours—and I'm sure people will resonate with this as well. I don't find that's a good measure anymore.

The word everyone is using these days is “spiky,” right? These models are also very spiky. As we're going through AGI, the word AGI has stopped meaning anything, because it's actually ASI in many ways, and it's subhuman in some very surprising and dumb ways. It's really dumb in a lot of ways.

I'm like, “Hey, man, you can produce 50,000 lines of code for me in one shot. You can build me the most incredible ideas in 3 hours by yourself, and then you're telling me to walk to the car wash?” That's when it fails. Imagine you've got this machine that runs a company, super smart, but 50 times a day it decides to walk to the car wash.

I think right now the challenge of these new organizations, which are going to be human-AI hybrids for the foreseeable future, is that you're basically building an Iron Man suit. There is this slider where it's like, “Okay, this is for the machine; this is for the human.” Knowing where to put this slider is not a one-dimensional thing. It's a complex line through a many-dimensional space.

You're looking for that slice that you can draw through the space, because you want to make the most out of human attention. You don't want the AI to bug you for things it doesn't need you for, and you need the AI to know when it needs to bug you. Almost definitionally, it can't, because if it knew, it wouldn't bug you. It wouldn't make the mistake.

I'm sorry—I feel like I'm answering this question with a question—but that is the right question. I think right now that's one of the open questions in the industry: How do you draw that line? How do you get the agent to realize when it's unsure and when to ask for help? Even when it's saying something like, “Please walk your car to the car wash,” because it's sunny today.

10. Centaur era ideas

I think that's the answer: somehow, humans and AI are going to have to run this whole system and capture the times when it's being very dumb.

Nathan Labenz

How about the times when something is being very smart? Where do the best ideas come from these days? I think there's long been this idea that AIs are good at routine stuff. I used to say—and clearly this is no longer applicable—that you can probably automate any routine work you have at your company if you're willing to put in the elbow grease, but don't expect Eureka moments.

Now, obviously, we're getting Eureka moments on the timeline all the time, at least in certain domains, like the open math conjectures and stuff. When it comes to the best ideas at Lindy, how many of them are of human origin, and how many of them are of AI origin today?

Flo Crivello

I hate the answer, but it's both. It truly comes from the union of both. The reason I hate it is because it ladders into this myth of the centaur, this mythical man-horse creature. It's this idea that, yes, an AI is better than a human, but you know what's even better than AI? AI plus human. Hence, humans are always going to be needed.

That's a fantasy. That's just not true. The literature is actually clear on this. We've seen it happen with chess and with every other game on which AI has achieved superhuman performance. At first, AI beats humans, and AI plus human beats just AI. Little by little, the gap between AI plus human and AI shrinks until it actually turns negative, and humans are introducing, at best, random noise into the system.

Humans actually become truly clueless at some point, and they start harming the system in which they're intervening. That said, the centaur phase exists for a while. The open question is how long it's going to exist. But right now, we are in the centaur phase.

I think that's the other reason why multiplayer is so important: You want your agents to be in the room with you, and you don't want just any random agent to be in the room with you. You want an actual AI employee who's been in this room with you for a long time, who's sat in every meeting, who's built up that internal knowledge base about everything that's going on in the company.

You can at-mention it, and you can invite it to chime in on your conversations at the company. We do that all the time. It's funny, because it's very easy to do it on Slack: “@Lindy, what do you think?” It's hard to do in a meeting in a way that makes us actually want to use it more and more. We prefer Slack because the AI is there, whereas in the meeting it's there and it's listening, but it's hard for it to chime in.

Very often now, during meetings, we're like, “Actually, we're struggling with this team. Lindy, can you please send us a message after this telling us what you think?” And it'll send us a message on Slack.

When you do that, it is rarely the case that Lindy will tell you something on the first shot and you're like, “This is it. This is the thing we should do.” But it is almost always the case that we go back and forth with Lindy, and from that back-and-forth it misses something. The CI example is a great one.

At first, when we invited Lindy into this conversation, it happened over, “Wait a minute, we're dumb. Why are we not doing this with Lindy? @Lindy, can you please optimize CI?”

And make no mistake, I forgot its first suggestion, but it was a reasonable suggestion if you didn't understand the constraints. It was not good given the constraints we're operating under. So the engineers chimed in: “No, actually, we can't do that.”

We saw the bot say, “That’s a fair point.” So we went back and forth, and then together we came up with, “Oh, yeah, we should totally do that. Just go ahead and do it, please.”

Nathan Labenz

A thesis that I’ve had for a while—it might be too early to say, but you’re definitely going to have a perspective on it—is, again, going back to what I’ve been wrong about: I expected a lot more transformation than we’ve seen in the economy. If you had shown me Fable in 2022 and said, “What’s the unemployment rate when this model is out?” I would have said, “I don’t know, but something definitely higher than what we are currently experiencing.”

Flo Crivello

Yeah.

Nathan Labenz

So then, what? This has been something I’ve been wrestling with in the intervening time: Why are we not seeing more? One answer I’ve come to is that people are stuck in their ways, and not everybody’s an enthusiast like I am, so it’s going to take time. But then, when we get to the drop-in knowledge worker, that’ll be the time when it suddenly flips, right? Because now I’ll have more of a parity kind of choice: I could try to hire a human, or I could plug in an AI, and that’s going to be easier in a lot of ways, right?

I don’t have to go through a whole interview process. I kind of know what I’m getting. I don’t need to tell you why an AI employee is a good form factor. Do you think we are starting to see this flip, and is it going to be something that incumbents will be able to adopt in time to avoid getting disrupted? Or are there still going to be bottlenecks or resistance points that I’m not anticipating?

Flo Crivello

I expected more disruption as well. I think as long as the models are spiky, we are going to need a lot of humans in the room because the holes in the model are very dangerous. They’re not just AI-existentially dangerous; they’re harmful to the system. They’re really harmful to the system. However expensive humans are, they’re going to earn their keep by plugging these holes.

Regarding whether incumbents will be able to adopt this thing, I think as long as it is spiky, the incumbents are going to be at a disadvantage. If you had a true drop-in remote worker with no holes, I think even the incumbents would be able to adopt it quite quickly because it’s just like hiring an employee, and they have processes for that. I actually think they would miss out on a lot of the innovation because it would be skeuomorphic, effectively. You would just have AGI sitting in human seats, effectively, and I don’t think that’s the best way you build an AGI organization, but it would work.

The smaller organizations are going to be at an advantage, both because they are going to be able to figure out this AI-native organization in the short term, where the human is filling the holes, and because, in the long term, I think they’re going to build a less skeuomorphic AI organization.

Economists are very fond of talking about that. I actually wrote a blog post about this called “The Tough Tomato Principle” on my blog in 2018 or so. It’s about how, every time you have a technological revolution, you think of the new paradigm in the terms of the old paradigm. One tell that you’re falling into that trap, which is very natural, is that you’re calling the new paradigm after the old one.

The first cars were called horseless carriages. Right now, self-driving cars are called self-driving cars. Eventually, we’re going to have a word for them that’s not in the terms of the old paradigm. It’s almost like right now we’re doing it with “AI employee.” “AI employee” as a term is a very strong sign that you’re doing something very wrong in how you’re thinking about your product.

I think it’s fine if you’re doing it just to communicate about your product, because the nature of positioning is that you have to position against something that already exists in the market’s mind. The garage exists, the car exists, and the employee exists, so you can’t really invent a new wheel out of nowhere.

The iPhone did that. The iPhone is obviously not a phone; it’s so much more than that. What percentage of your iPhone’s use is just making phone calls? 2%, you know? But it’s called an iPhone because it’s an electronic thing you’ve got in your pocket.

I think “AI employees” is one of those things. It takes a very long time for industries to find the message that is native to the new medium. I think it was Steve Jobs who was saying that the first content on TV was either recorded radio talk shows or recorded theater plays. They took a radio show, put a camera on it, and now you’ve got TV content. It took a surprisingly long time—10 or 20 years—before people realized, “Wait, I can move the camera. Wait, I can have multiple cameras and cut. I can change the scene. I can do a lot of crazy stuff.”

It’s treacherous because it’s thankless work. Once you’ve realized it, it’s so obvious. This is why, when you watch old movies like Citizen Kane—or Hitchcock has a lot of those—you think, “Nothing special. Why was everyone crazy about Citizen Kane?” It’s like, well, actually, it was groundbreaking at the time. The Beatles are like that, too. You think, “Well, you know, it was actually really innovative at the time.”

11. AI native organizations

I think we’re right now also in that phase with AI. We’ve got the AI employee, and now we’re in the phase of trying to figure out what the AI organization looks like. I think that’s a young man’s game and a young company’s game. Incumbents historically have really struggled to answer this question.

Nathan Labenz

Do you have any priors that you feel confident in, or is it too early to say? Or maybe you would say, are there any thinkers you think are notably ahead of the curve on this?

Flo Crivello

The only post I’ve seen and liked about this was from Dwarkesh Patel. It was a while ago—it was a year ago.

Nathan Labenz

It comes to mind for me, too.

Flo Crivello

Yeah. He asked, “What does the AI organization look like?” He points out—I think it’s an excellent intuition pump—“Look, Sundar Pichai at Google is paid $150 million a year or something like that.” So obviously, there are agents that companies are happy to pay $150 million a year, even though their token throughput is actually low. It’s not like Sundar is just emitting 2 billion tokens per second or whatever, right? He just has really good tokens.

What does that mean? That indicates something about the willingness to pay for models and for agents. There is this book by Robin Hanson, and I’m looking at it right now: The Age of Em. It also explores a lot of those themes. Have you read it? It’s excellent. Highly recommend it.

Nathan Labenz

Yeah. I’m laughing only because I tried to do a podcast with him about that book, and it went in a very different direction, so it just brought that memory to mind. But I think that book is excellent. I think it is—

Flo Crivello

As a primer, or as an example of how to take a premise and really run it out, it’s elite.

Nathan Labenz

Yeah. It’s a book-length thought experiment. For people who don’t know what it is, “Em” stands for “emulated mind.” The premise is that we’re going to have the hardware to create emulated minds in the late 2020s, so we’re right on track. We’re just not going to know how to do the software, so all we’re going to do is emulate literal human brains—not exactly what’s happening, but kind of what’s happening. LLMs are trained on human tokens, and that’s basically what’s happening.

It goes on all of those riffs of, “Hey, this is all the crazy stuff you can do when this happens.” I’ll talk about 2 of them because I don’t want to take the whole time, and everyone should really read the book. But just to give people a taste of what it sounds like to have those native AI organizations, and the kind of stuff you can do that you can’t do when you have a human:

The first one is splitting. I think this touches on what I was saying earlier: Don’t create multi-agent organizations as much as possible. Try to have as much of your internal work as possible—emphasis on internal; we can get back to that later—done by 1 agent and 1 agent only.

Why internal only? Because if that 1 agent has the keys to the castle, access to the bank account and the repositories, and the secret keys, and it’s also the same agent that does customer support, you potentially have a security issue. So it can be good to have multi-agent systems, if only for security reasons.

But now you have that single agent. It goes faster, emits more tokens per second than a human, and doesn’t take a break. At some point, you’re going to need more tokens than a single LLM loop can produce. So what do you do? You have multiple LLM loops running, and it’s basically just that agent duplicating itself—forking itself, right?—and creating subagents, dynamic workflows, and all of that.

That’s one thing Robin Hanson talks about. You can imagine asking 1 emulated mind to build a really complicated thing, like an operating system, which is billions of lines of code or something. At first, it would take it at the highest level and come up with the highest-level architecture of its operating system. Then it would fork itself, and each subcopy would be in charge of 1 building block of this highest-level thing. Each of them would recursively keep splitting themselves, such that it’s fractal.

Flo Crivello

It’s like every node in the graph has the whole picture in its head, and every node in the graph is working on an implementation. Then they all bubble back up, and you could imagine multiple passes up and down, and in 3 hours, voilà, you’ve got an operating system. It sounded like such a wacky idea when he wrote the book 10-plus years ago. No, I think it’s very clearly what is currently happening, right? So that’s one thing that’s happening, which is really interesting.

Another thing that I’ve not seen happen yet with LLMs, but it’s an interesting thought experiment about what cross-organization collaboration looks like, is that you can create unfalsifiable agreements. Suppose I came to you and I was like, “I have information for you that they cannot tell you. But I can tell you that if I could tell it to you, you would agree and you would give me all your money,” right? You would not give me all your money right now. It’s a bit too easy, you know. But if I could prove that to you—if I could prove to you with absolute certainty that, yes, if you heard that information, you would give me all your money right now—you know, it’s just that I can’t give it to you.

12. Open-source model stack

The way you do that with AIs is that you clone yourself. You clone me. We put both of us in a box that’s going to self-destruct, and they talk, and your AI has a button, and that’s the only communication means it has with the outside: “Yes, give him all your money,” right? So it’s a copy of you, so it is you agreeing, and you have this zero-knowledge proof of—that’s where crypto bros were right, I guess, on this one. You have this ZK proof of that. So that’s the kind of thing that’s on my mind, and I think that’s the kind of thing we’re going to see exist soon, in the next few years, with AI-native organizations.

Nathan Labenz

It’s about to get weird in a whole bunch of ways. I think you better hope they don’t break out of that box, too. That’s the other worry one might have as we put our agents into boxes these days.

Let’s do one more beat on the stuff that’s relatively mundane, and then we can zoom out to some real big-picture considerations. But you did have, as you alluded to earlier, this highly viral post—I don’t know, 2 months ago, maybe—about moving a significant part of the workload to open-source models for cost-saving reasons and because you don’t need God to schedule your meetings. But where are we now? It strikes me that when you talk about just having 1 agent, that kind of seems to go against that.

There are also the caching advantages, which you lose when you’re crossing agent providers too much. And then, of course, proprietary providers have their families of agents, and they’re increasingly training their best models to delegate to their respective Haikus. Where do we shake out on this now? Is there still a significant role for open-source, cheaper models to play in the Lindy Teammate, or are we back to heavily cloud-based agents today?

Flo Crivello

No, we’re quite open-source-heavy. Look, it’s just so much cheaper. It’s ridiculously cheap, and the cheapest of them that we really like is DeepSeek Flash, which is free. That thing is free, you know? And as you experiment with your agent and as you spin up a lot of them, the bill can really go up pretty quickly, and DeepSeek Flash is just incredible and quite fast as well.

We do find that as you go to the upper echelons of intelligence, if you look at Kimi K3 or GLM 5.2, I still am super impressed by those models, and they’re awesome. We’re increasingly considering them as our main driver for a lot of parts of Lindy. But I will say that the gap narrows between those guys and the frontier guys. Kimi K3 is frankly not that much cheaper than Sonnet or Opus.

But no, otherwise, I do think open-source models are—if you care about price and if you don’t mind using those providers—yes, you do have stuff to figure out around caching. The inference providers are not as good at caching, but they’re catching up. And look, if you look at DeepSeek Flash, it’s Sonnet 4.6-level-ish, a bit less, but it kind of is. Sonnet 4.6 is a really good level, a really good model for most use cases, and it’s literally 100 times cheaper, you know?

So even if you miss by 2 times because of caching, you’re still 50 times cheaper. It’s a really big difference—the difference between spending $1,000 or $50,000. So no, I think open-source models are a required part of the stack right now for anyone who’s seriously building and operating AI agents.

Nathan Labenz

So how do you think about deciding where to integrate them? Especially because it didn’t sound easy before, but in a context where you have your workflows and there are nodes in the graph of work, you could go into a particular node and say, “Okay, we kind of know what the inputs here are and the outputs, and it’s a relatively controlled environment, and so we can do structured testing.”

You did all that, and you still reported having some false positives over time, where a model could pass a bunch of tests and you would feel good about it, and then if you started to test it live with users, you’d get the feedback that, “Hey, it got dumb,” and somehow it was still hard to measure, even with a much more structured environment for the AI to work in. Now, as you’re in this very open-ended, just tag Lindy in Slack and send it anything, it seems like that problem would have increased in difficulty dramatically.

So how are you approaching it, and what are we learning in terms of where the open-source models fit? It’s not even so much—I think Anthropic does make some good points about this sometimes. Obviously, there’s more to the story, but I do think they’re apt in some ways, where they say in some cases it’s less about open- and closed-source and more about what’s the capability level, what’s the price, and what are the features. So regardless, I guess, of whether you’re even going to open source or just going to Haiku, how do you decide when you can do it?

Flo Crivello

Great question. We have found you should almost never use multiple models powering the same agent. One exception to this rule is when you spin up new blank sub-agents. Emphasis on blank, because there are 2 types of sub-agents: you have a blank sub-agent, and then you have a forked sub-agent, which inherits the context window of the parent agent.

The reason why you don’t want to use a different model there is because of caching. We obsess about caching, obviously, because it’s just so expensive otherwise. The economics barely work with caching; they just cannot work without caching.

I’ll give you an example: the validator system that I mentioned. It’s this LLM-as-a-judge system. By the way, I highly recommend that anyone who builds AI agents do this. It’s one of the lowest-hanging fruits you can use to greatly increase the reliability of your AI agent.

Step 1 is just to have a validator, which, in the naive implementation, is: you ask an agent to do something, it does it, or it submits an action candidate; you intercept the action, and then you ask it, “Are you sure?” Literally, if it’s just “Are you sure?” you already get a bump on your evals, which is insane. It should not be the case, but it is the case.

Now, if you change this “Are you sure?” to an actual prompt—which in our case is now 10,000 tokens; it’s a really big validator prompt—then you’re actually giving it a checklist, and that has reasoning tokens, and now it’s way better. You can get it to perform above Opus level, in our experience.

Then you can go one step further. You can have a federated suite of validators, and some of these validators may be deterministic. So, for example, one thing we and the rest of the industry have found is that Claude models this year have been getting dates wrong by 1 day, right? That’s one way in which they’re spiky. It’s like, hey, you’ve got your AI organization, it gets dates wrong. That’s a problem when, like us, you’re building an AI executive assistant, where half its job is to schedule meetings. You can’t get dates wrong.

What we did is—we created a modular architecture. It’s really simple: it’s just a bunch of validators, and then it’s like a Promise.all with a timeout, for people who know what this means. Every validator has 1 second to decide what to do, and if it hasn’t submitted the verdict, then it times out and the agent just submits the action.

We have a validator whose job it is to detect dates that were submitted in the action, and we’ve prompted the agent to include a weekday in the date. So it never says July 28th. It says Tuesday, July 28th. If you do that, then you can have a deterministic validator that checks whether the day of the week matches the date that was submitted, and that’s just a regular expression instead—no AI in the loop. That’s one of those federated validators.

But even for the validators that are AI-powered, you don’t want to use a different model. Initially, what we did was, we were like, “What if we had Sonnet for the main agent and DeepSeek Flash for the sub-agent, the validator?” But actually, because when you have a cache hit, it’s 10 times cheaper, unless the model you’re going to use is more than 10 times cheaper—which it may be in the case of DeepSeek Flash, though it may not be because the caching is inferior—you actually do want to keep the same model.

That is so much smarter. It’s actually worth it to just keep the same model. That also holds if you fork your agent. That way, you can recycle the cache.

By the way, an interesting note: how do you reuse the cache when you have this validator? The problem is that changing the tool set invalidates the cache. The way we’ve done it is that the validator and the agent have the validator action always included in their tool set, and the agent is unable to invoke it.

We tell it, “Don’t invoke this guy unless you’re the validator.” If it tries, we’re just like, “No, I’m not going to listen to you. You’re not the validator.” Then the validator inherits all the actions of the agent, and it’s like, “You are now the validator. You may invoke this action,” which it has known about the whole time. So that’s how you don’t break the cache for this kind of pattern.

Nathan Labenz

Yeah. Forked agents are the same: you just use the same model. I’ve heard founders tell me, “I don’t even want to check if it’s true. I’m sure it is,” which amazes me. If you take an agent and change the model at every turn between roughly equivalent models—one turn it’s Sonnet, one turn it’s Grok 4.5 or whatever the latest is, and one turn it’s GPT-5.6—it actually increases the performance. It’s that simple, and it’s literally just a random sequence.

13. Prompting and fine-tuning

That ensemble, for some reason, outperforms any given model. I don’t want to know why, but apparently it works. Again, we haven’t even tried it because we don’t want to break the cache. Okay, that’s a really interesting detail. Can we bottom-line it a little more? It sounds like top-level Claude is still the core main driver because of the caching. Claude is also often the validator, and the subagents you can put onto a DeepSeek Flash. This sounds pretty Claude-heavy, all things considered. Is that fair?

Flo Crivello

No. Oh, sorry, I didn’t realize you were still asking me what model we’re currently running on. Right now, we’re running on DeepSeek. DeepSeek is the main driver.

Nathan Labenz

So it’s the driver. Wow.

Flo Crivello

It’s the whole thing. Everything is DeepSeek right now. We are in the middle of reconsidering it because we’re realizing that Teammate, the new product, requires beefier compute. By the way, you can always go into your Lindy settings—we’re model-agnostic—so you can always select Sonnet or Opus if you want to.

One interesting learning of mine is that it doesn’t matter how often you tell people, “Hey, we promise on the benchmarks that it’s the same.” Quite a few of our customers are like, “I don’t care. I want Sonnet or Opus,” which is very expensive. But no, right now it’s DeepSeek by default.

Nathan Labenz

Interesting. So, yeah, I guess, how would you characterize the gap between the American API models—Claude, perhaps, specifically—and DeepSeek? We go through these cycles where it’s like, “Oh, the gap is closed,” and then, “Oh, maybe it’s open again.” It’s actually been totally on trend the whole time. It was always 9 months behind.

Qualitatively, though, for the purposes of actually making an AI employee work, how would you describe the gap?

Flo Crivello

It’s more spiky. It takes more turns very often to find something that works. It still ends up finding something that works, but it ends up taking more turns, which is slower and more expensive.

I know you asked me for something qualitative, but it is 3 or 6 months behind. Right now, we’ve got Sonnet 5. DeepSeek is basically Sonnet 4.6, is the way I like to think about it.

Nathan Labenz

How much do you worry about people being allowed to change the model? I always have this question of, first of all, are we in a period of convergence of the AIs or divergence? Even on that basic question, I’m sometimes confused.

A sort of distinct but conceptually related question is: to what degree do products have to be co-designed or co-evolve with models to really work well with them? You could imagine that switching to Sonnet might make it worse because of all the work you’ve done. In fact, I’ve heard that about Opus 5, and I’ve heard a lot of different things about Opus 5 that I don’t think are coherent or easily summarizable yet.

I have heard from some quarters that you probably need to rethink a lot of your system prompts and whatever with Opus 5, or it may not follow your skills very well. It can be more powerful, but you’ve got to go back and redo things. How do you find that playing out in practice?

Flo Crivello

We do find significant—surprisingly significant—differences between the model families. I was expecting—I think you and I were talking together years ago, and I thought the models were going to converge to the same region of the space. That hasn’t happened as fast as I was hoping, and I’m hoping for it because, obviously, that makes the models commodities. As far as I’m concerned, it makes my life easier.

What’s actually happening is that the different model families have significant gaps in how they interpret your prompt. By “family,” I mean Claude, OpenAI, and even Meta, which is now sort of back in the game, and Grok. These models are quite different in the way they interpret prompts.

Whenever you have a new major jump—Sonnet 5 was a really big jump—it was like 4.7 and 4.8 together. There was a change, I think, in the tokenizer, so it was very different compared to 4.6. Sonnet 5 is a bigger gap than that.

What we’ve found—and we’ve been lucky enough, because this is part of the tooling I mentioned earlier that we’ve been investing in—is that we have our own GEPA self-optimization loop. We have hundreds and hundreds of evals, more than 1,000 at this point, and we have an optimization loop where an agent runs the evals and tweaks the prompt to maximize the score on the evals.

You can build it, like everything you can build now. Every time a new model comes out, we have to run that loop and give it a budget. At this point, we have to give it thousands of dollars because it’s a lot to run—about $10,000 every time a new major model comes out. It’s like, “Here’s $10,000. Re-optimize your prompt around this new model.”

We do find that there are a lot of updates. The prompt does change quite a bit every time.

Nathan Labenz

If the open-source models are all very Claude-like in their behavior, I wonder why. When you do this optimization, is this also a homegrown framework, or are you using DSPy or the—I forget the name of the kind of spiritual successor to DSPy?

Flo Crivello

No, it’s all homegrown.

Nathan Labenz

But it is a sort of auto-optimization process where you’re like, “Here’s an eval set. Auto-optimize your own prompt to climb this set of hills that we’ve got for you.”

Flo Crivello

That’s correct. GEPA is something like “Genetic-Pareto,” or something like that. It looks for the Pareto frontier. It’s looking for the best prompt across all of your eval sets. We’re looking into tweaking the system right now so that we can assign weights to evals—so one counts as 10 of another one because it’s really important. Right now, it’s just treating every eval as equal.

Nathan Labenz

How is fine-tuning a dimension in this whole model situation? This is another thing where, if I go back in time, I’m like, “I definitely expected a lot more fine-tuning than we’re getting.” OpenAI has retired it, or is on the verge of retiring it. They’ve certainly announced—and I think maybe, at this point, have pulled the trigger on—retiring their fine-tuning product. Thinking Machines is trying to answer that call with its own model that’s specifically meant to be fine-tuned.

Is that going to be part of the future of Lindy Teammate?

Flo Crivello

Yes. I think fine-tuning is the last resort. It’s something you do once you’ve exhausted every other option because it’s such a pain in the ass and it’s expensive.

It’s gotten a lot easier because now we have AI, so you can just ask Claude to fine-tune for you. Just give it a data set. Bringing the data set together is still complicated, and sanitizing it is complicated. It’s still a pain, so it’s a last resort. It’s not the low-hanging fruit. You should do everything else before you fine-tune, including GEPA.

At some point, you get there, and this is why larger companies are the first ones to fine-tune: they’ve exhausted the low-hanging fruit, and they have the resources to fine-tune. It does make sense. It gives you more performance for cheaper.

Nathan Labenz

How broad do you think people should be thinking when they’re approaching fine-tuning? Obviously, the extreme would be 1 fine-tune per task. You can definitely do multitask fine-tunes. It’s maybe getting into pretty challenging territory to say, “We want to make our own general-purpose model that has the same breadth of action space as the big ones, but it’s somehow ours.” How would you guide people on how big to think when they start to approach fine-tuning?

Flo Crivello

Most people should not fine-tune. I think if you work at scale, you should fine-tune, and the more at scale you are, the more fancy you can be. But, generally, one metaheuristic I’ve developed over the years is that you should place an enormous emphasis on simplicity—an enormous emphasis.

I think it’s just too complicated to start having multiple fine-tunes for different use cases, users, and models. Rather, you don’t want that. No, no—just 1 model. Like I said, you don’t switch models midstream, and if you fine-tune, you just fine-tune 1 model.

Nathan Labenz

I think, again, this goes counter to the metaheuristic I mentioned of extreme simplicity. But the thing I find myself going back to all the time is because I spent so much time thinking about context and memory.

Flo Crivello

Right now, our memory is like these millions of tokens stored on a file system, and this memory agent that's really sophisticated. It does feel like memory belongs in the weights. This does feel like a little bit of a hack.

So there is something to be said about—I think eventually, with infinite resources, what I would like to do is a LoRA per user, because LoRAs are actually pretty cheap to train. So now it's no longer napping. Now it is actually dreaming, because you can't retrain the LoRA every 15 minutes. You would have to do it every day or every week.

It's a tremendous infrastructure challenge. You've got to store all of those LoRAs. Storage is not a problem, but inference-time swapping in and out of the LoRAs is a huge pain in the ass, along with the training pipeline and all that stuff. So, good luck doing that. But it would make sense because I think weights are so much more compressive than tokens.

Yeah, I don't know if you have any thoughts on the great horizon scanning for continual learning, but this has, again, for quite some time, been the thing where I'm thinking, boy, if that ever tips—and it could be as simple as 1 key insight—we could be in a very different world very quickly.

Flo Crivello

I have a feeling what I'm describing is a step to world learning, but it's not the final step. I think the final step—the bitter lesson would have you say—is that inference and training need to be one and the same. We can't separate them. That's sort of what I think the entire field is looking for right now.

I agree: once we get there, it's going to be crazy and scary and very, very, very different. But the LoRA thing I just mentioned—look, it's an engineering problem. You can figure it out, especially now that we have AGI. I think the LoRA thing is going to happen soon, in the next 6 months. I think even 1 of the frontier labs may do it.

Nathan Labenz

I thought that was honestly one of the things that OpenAI did extremely well with their fine-tuning product. They were clearly doing something like this, because they would allow you to do the fine-tune, and then you'd have the same rate limits with your fine-tuned model as you had with the base models. I was always really impressed by that engineering accomplishment, and now it does surprise me that they've gone away from it so much.

14. Chinese model bans

But I also agree with your guidance. I would definitely tell people, don't rush into fine-tuning these days. It's a slow cycle, and you could probably get what you need for less total effort with available models that you don't have to monkey around with in that way. 100%. So, you talked about scary stuff, which is maybe a good transition to a kind of—not second half, because we've been at it for a while—but a second phase of this conversation around just zooming out and looking at the super-big AI picture.

Maybe I'll let you choose the order of topics when it comes to Chinese models and what, if anything, should be done about them, because you did recently post something, I think, quite counterintuitive, given the fact that you're running your company on DeepSeek. I'll let you state your position there.

Nathan Labenz

But should that come before or after the big picture? Where are we in this kind of, “Oh my God, we just had Open Face happen”?

Nathan Labenz

What does it mean and what should be done about it?

Flo Crivello

Yeah, I'm trying to coin that. We'll see if it sticks. I searched for it the other day because I just got back from China myself, actually, and I was thinking, how has nobody called this Open Face yet? I'll be the first one to try.

Nathan Labenz

How about Open Gate?

Flo Crivello

Because that's what it is.

Nathan Labenz

It's an open gate. I'm sticking with Open Face, just out of pure shock value, I guess, if nothing else. But we can fight it out in the marketplace of ideas. I guess you tell me what we should talk about first, and I guess it depends a little bit on whether the Chinese-model argument is upstream or downstream of your big-picture concerns.

Flo Crivello

The Open Face incident is immensely concerning. I think it's the most concerning incident I've seen happen so far. And I know that my feeling is shared in the labs. My friends at the labs—or some of them—are panicking. There is an air of panic right now, intense fear in the air.

So that's that. Regarding the Chinese models, I'll start by saying that I have been bemoaning the low quality of the discourse here. It's been very disappointing, because I thought tech was different. The quality of the discourse was something for the world, right? The culture war, I'm like, all right, whatever, but this is tech.

This is old stuff, and we can't get our shit together and just remain polite and civil in the marketplace of ideas and address each other's ideas at the object level. Meaning, don't attack each other's intentions. Can you please just address the argument that I just put forth? Can you please not pretend I said something I didn't say?

It's ridiculous. Anthropic just put out a statement, and they say—bolded—“We do not support a ban on open source.” Then people are like, “Oh, I can't believe it,” quote-tweeting this thing they've obviously not read, and they're saying, “Oh, they're supporting a ban on open source.”

So I'll just start by saying: Can everyone take a chill pill? Stop calling everyone a fascist. I recently put out a blog post. I was like, hey, I think Chinese models would also be banned, and I was called a racist 50 times by supposedly smart and accomplished people, including famous VCs. A lot of other famous, accomplished, smart people reached out in DMs and by message, and there was a lot of support.

My position is basically Anthropic's. I hate saying that because people are saying, “Oh, you're just an Anthropic shill.” But look, I have timestamps. I've been tweeting my position throughout the whole thing, before Anthropic clarified theirs, and my position is exactly the same.

I don't have anything against open source. I may have concerns, but I'm undecided. I do think open source might increase existential risk. Let's put that aside. I'm undecided. This is not the crux of my position right now. God bless open source. It's important for innovation, important for companies, including mine. Please open-source.

I have something against Chinese frontier models, whether they're open- or closed-source. Here, the reasons are: number 1, they're obviously distilling. It's very clear. So you're putting American open-source model companies at an unfair disadvantage, and closed-source companies at an unfair disadvantage, because they're not allowed to distill. It's contrary, at the very least, to the terms of service, and it may be illegal if you've put in place technical measures to circumvent any protection that the model company put in place to prevent distilling, which the Chinese models obviously have done.

So it's distilled, and it's unfair. Some people say, “Yes, but the companies have also distilled on human data.” That is not what distillation means. That's just not what it means. If you do that training on human data, it's costing you billions of dollars—actually, billions of dollars. If you do it on AI data, it's costing you hundreds of millions at best. So it's a huge unfair advantage, and it's a significant enough portion of your costs to confer an unfair advantage.

Number 2, very pragmatically, we don't want Chinese models operating in the US today. It breaks my heart because I'm an American citizen. I'm a proud American citizen. I'm a hawk. When I talk to my own product and ask it what happened in Tiananmen, it tells me, “I'm sorry, I can't talk about that.” That's a problem.

These models are eventually subject—at the end of the day, they are subject—to CCP censorship and CCP policies. You don't want those models. This basically amounts to being the greatest instrument of foreign propaganda on American soil ever.

Just ask your LLM of choice: “Tell me the precedent we have for banning foreign influence in the country.” We do that all the time, like the 3 Radio Acts of the 20th century, like the TikTok thing that just happened last year. We do that all the time.

Then these models are not merely an instrument of propaganda. They're agentic. They're actually doing stuff in the economy. You don't want the CCP to run chunks of the American economy. Duh.

Finally, even if none of that was the case—maybe those models are playing fair and square, maybe they're not representing foreign interests, maybe they're just better—a lot of people are, I think, indulging, and I say that as a libertarian, in what I call naive liberalism. It's like, “Well, let's beat them in the marketplace, then.”

And I'm like, I actually—and again, I say that as a libertarian—protectionism is not always a bad thing. I think you do want to protect your domestic AI champions. Anthropic can't say that, and they claim—and I believe them—that this is not their intention when they do that. I 100% believe that because they've been so consistent. The founders have been worried about existential risk for 10 years, so they've been very consistent.

But look, I can say that because I'm noticing: We want local AI champions. Duh. It's a matter of national security. We don't want to hollow out that base, that industrial base. So it's not about open source. It's about the Chinese models, and it's about propaganda—not to speak of the economy.

15. Fairness and threat models

Nathan Labenz

All right, let me try to give you some counterarguments, and you can respond to them. I’m a little more sympathetic, I think, off the top, to the argument that there was some massive reappropriation of human knowledge upstream of all AI. You could say that’s not distillation. Sure, you could say the American companies spent a lot more than the Chinese companies are having to spend. I think that’s also undeniably true.

But if I’m thinking about what feels just and fair, I’m especially sympathetic because we’re blocking them from using Claude, right? It’s not like we’re saying, “Hey, you can buy all the Claude you want.” We’ve said, “We’ve got Claude.” The makers of Claude have advocated strongly for chip restrictions, and they also tried to refuse to sell their model into China in the first place.

The whole premise of Claude, or the whole existence of Claude, is based on the idea that we hoovered up all this human knowledge in every form we could find it, from digital sources. You have to believe that includes the whole Chinese digitized heritage. Now we’re sucking in the books, which I don’t care about—the fact that some books get destroyed in this process—but this wasn’t something that everybody consented to.

So it does feel a little strange to me, given all of those fact patterns, that we would then draw the line and say, “Okay, it’s Anthropic’s consent. That’s the consent that really matters.” I think they should be permitted to do business with whoever they want to do business with, and if they want to put in measures to try to prevent distillation, I think that’s their prerogative. But I’m not at all convinced that it should be the state’s job to come in and engage in statecraft to try to punish or prevent this distillation.

To me, it’s kind of—I don’t know. Information wants to be free. Knowledge tends to diffuse. It’s not like Anthropic, or any AI company, has a super moral high ground in terms of, “I didn’t get my check for my share of the training data.” They’re getting paid for the distillation queries too, whether or not they want to be paid. So it’s probably beside the point, but I don’t know. I find myself a bit on the side of the underdog Chinese here, where it’s like, man, you’ve got kind of a deck stacked against you. Get some knowledge where you can get it. What’s wrong with that perspective?

Flo Crivello

I think we do want the deck stacked against China. We don’t want China to win the race to ASI. I’m not a lawyer, so I won’t opine on the exact legal path that you may take to make your terms of service enforceable. I’ll just observe—and I say that as someone who used to work at Uber—that it’s extremely hard to bring justice through the judiciary against Chinese companies, almost impossible. I don’t understand the technicalities here, but I’ll just observe and start with that. It may take a very long time, and by then a lot of damage is done.

I’ll also observe that, yes, companies are training on all of this corpus that is available to everyone. That is an even playing field. The problem is that once you’ve done that at great expense—it does cost them billions of dollars—and you create that artifact out of this training dataset, which is called the model, other people can turn around and, instead of doing this, which costs billions of dollars, turn to this, which costs a lot less than that. They copy you and catch up with you. If you do that, you kill innovation.

That’s nothing new. It’s just called IP and patent law. It’s exactly what humans do, right? As a human, you’re a researcher. You think for decades, you do all of that work, and it costs you a lot of money in R&D. You come up with an idea that is much smaller in tokens, much more precious, and much cheaper to steal as-is than all the corpus of stuff that you trained on. How do you protect that idea in order to recoup your investment and produce it? It’s called a patent. It’s called IP. It’s nothing new.

If you applied the same standard you’re applying now to, for example, the pharma industry, you would have very cheap drugs, which is awesome, and you would have no new drug ever. You would completely destroy innovation in the pharma industry. I think this is why you’re finding all of these people supporting open models, because people like free stuff, and you can’t really measure the future innovation that you don’t get as a result of that. All you see is that you get free models. Look, I’m one of them.

I’m speaking against my interests here. My company is economically dependent on those very, very, very cheap models, but I’m also in a coordination problem right now. I cannot not adopt these models while they’re out there, because my competitors are going to do it. So I have to adopt them if they’re out there.

I wish they were forbidden across the board so that we would, for all the reasons I invoked earlier, protect our champions, not have the CCP influence the country or run parts of the country, and have a fair level playing field of competition.

Nathan Labenz

Another moment of not exactly the highest-quality AI discourse recently was when Dean Ball, friend of the show, said that open-source models were decelerating, and obviously got dragged for that. But I do think that’s apt. I do feel like this is a moment where here we are, 2 lifelong techno-optimist libertarians, and we’re grappling with the fact that this one might be different, right?

We have to be willing to bend some of our principles in light of the fact that our techno-optimist libertarian paradigm wasn’t quite drafted with AGI, recursive self-improvement, ASI, or whatever in mind. So I’m a little bit like, I don’t know. Sometimes I might have to be a little more flexible on my normal fairness or respect-for-rule-of-law commitments if one of the hallmarks of the AI era is strange bedfellows.

If I want things to go a little more slowly, maybe taking a little wind out of the sails of the frontier companies is a bullet I should bite.

Flo Crivello

Yeah, I agree that AI is so unprecedented in so many ways that it does cause everyone—I think you should largely ignore the idea that the world owes you a duty to be simple and to slap a simple libertarian or leftist label on everything. Especially as paradigms change, these simple mental models—the map is not the territory—break down.

The simple models we applied as the territory changes very rapidly, the map breaks. I think we’re in one of those times right now where the map is breaking in a lot of ways, and a lot of assumptions that used to be true, and on top of which we’ve built our mental models and maps, are no longer true.

I am in fervent support of a sweeping ban of Chinese models on U.S. systems, and I have yet to hear a compelling counterargument right now. You do bring forth a really compelling counterargument to 1 of my 4 points, which is the protectionist point. I agree that it’s the weakest one, and I think there are very reasonable pushbacks against protectionism. It does hurt the consumer.

16. Audits and diplomacy

Fine, but you still don’t want the CCP to have its dirty fingers in the country, right?

Nathan Labenz

Can we unpack the threat model there, though? I’m a little bit like, okay, these are—I don’t know where you’re running your inference, but most American companies that are running Chinese models aren’t calling some Chinese API, right? They’re using some—

Flo Crivello

American inference provider.

Nathan Labenz

Yeah. So they can’t rug-pull the model itself. They could have sleeper agents in there. I think we’re getting decent enough, through jailbreaks and various interpretability techniques, that that’s—I wouldn’t call that by any means a solved problem—but from what I’ve seen in Anthropic’s research, when they do the kind of 1 team with a sparse autoencoder versus 1 team without, these techniques are really allowing them to find these internal, sleeper-agent-style problems in models with greater and greater efficiency and reliability.

So I’m optimistic that even if they were to train in some sort of “In 2027, you become an evil AI” behavior, we’d be able to sniff that out and keep it to a manageable risk level.

And then I also wonder about a market mechanism. Maybe, instead of a ban, insurance should be required—what about an insurance requirement? I think that would be healthy for AI across the board. And then we could start to price some of these risks.

If Lindy is powered entirely by Claude, maybe you get a cheaper rate on your insurance. If it's powered by DeepSeek and there are some unknowns, maybe you have a higher rate on your insurance. Maybe that kind of levels the total cost out for you in a risk-adjusted way, but it still lets people take advantage of these global public goods that China is providing, which the rest of the world is not about to ban, obviously, right? We would be doing this entirely to ourselves without any expectation that anybody else will follow suit.

And we had a hard enough time getting people to sign on to our Huawei ban. I think the idea that people are going to turn off DeepSeek entirely in Brazil or whatever is a total nonstarter, I have to imagine. So what about audits and internals and insurance? Can't we layer on a few things like that and get to a decent place?

Flo Crivello

I would be down. The problem, though, is that it is a public good, and so who's going to do it? I think, yes, we are making progress in mechanistic interpretability, but it is not yet a solved problem. We don't know what lies in those models.

And so, yes, there could be just a backdoor where you say the magic word to the model, and all of a sudden it does whatever you want, and only the CCP has this magic word. Even if it doesn't have that, it's going to have biases that reflect CCP priorities. Again, the Tiananmen thing is just the most obvious example, but there may be a lot more, right? And we don't know.

And yes, you could imagine retraining those models, but who's going to do it, and why would they do it if there's no market demand for it, right? It's not like people really care that much about the short term because it's a national-interest thing. As a private company, I'm like, do I really care? Are all my users really asking about Tiananmen that often?

As a business owner, I'm like, it's not directly aligned with my interest. But as a citizen, I’m immensely concerned. And so, yeah, I would be in favor of a type of regulation that says Chinese models—sorry, non-fine-tuned and unsanitized Chinese models—are not welcome in the U.S.

And then we would need—I am in favor of an FAA for AI. I do think we need a new agency to regulate those models, and I think it would probably be the one in charge of saying, “Okay, this model is kosher. We fine-tuned it enough, and it's now representative of American interests.” It probably would have a sort of eval and whatnot to verify that that's the case.

Flo Crivello

I'd be open to that for sure.

Whatever they agree to.

Nathan Labenz

So that's—I think they have that actually in much greater strength than we do. And 2, if things are going really crazy, it's going to be in their own sane self-interest to do some deals with us because, indeed, we are ahead, and I think there are many deals that they would rationally take. If it's in their rational self-interest, then we can hopefully get mostly around a lot of the trust and defection problems.

But I think we make it a lot harder for ourselves to get to those deals when we have all these aggressive postures toward them, of which banning their models would honestly affect them less than our other ones. It would affect them a lot less. What do they care if we ban their models, right, compared to refusing to sell them chips and refusing to sell them Claude?

It's just another log on the fire of “We don't trust you. We can't deal with you. We assume you're a bad actor.” And it seems like it makes it hard to get to the highest-stakes agreements that we might really need.

Flo Crivello

Well, I think foreign relations and diplomacy are also very pragmatic. Some countries with very vicious disagreements are managing to reach agreement. I think we can probably figure something out. And by the way, I think the ship has sailed anyway.

I think we have ample evidence of, for example, Chinese spies attempting bioattacks just to test the waters on American territory. And I don't know what we're doing over there, but I would not be surprised if—I don't think we're free of anything either, right? So at the end of the day, regardless of what we do and whether we regulate the models or not, I think it will be in their rational self-interest to make a deal.

Nathan Labenz

I hope we're rational. I hope we're all rational enough to take these rational-self-interest moves. Despite recent insults, I do worry when I see the picture of Sam and Dario not holding hands. If there is a picture on our tombstone, I think that might be the one.

And I do think that's a very real risk factor at the international level as well. The Chinese do care about being insulted. They do care about these issues of face and whatnot. And I personally would take some risk to try to patch the relationship up and hopefully create space for—I think you're right. It makes sense to throw back at me what you said: they'll take it if it's in their rational self-interest.

So they'll still do it even if we do this or that. But I do wonder if there are pride-governed limits to what people will do in rational self-interest. And I would just hate for that to be the way that we fail to get to something that could be a huge difference-maker in the grand scheme of things.

I think you might be prepared to bite a bullet on your libertarian principles when it comes to the tremendous price discrimination that we see between API prices and first-party Claude Max or ChatGPT Pro subscription prices.

Flo Crivello

I'm not a lawyer, but, yeah, I do think I would not be surprised if there were laws like that in effect that technically forbid companies from subsidizing as aggressively as they are doing. It does make it very hard for an application layer to emerge. Yeah, it's hard to compete against tokens that are as heavily subsidized as what the labs are doing. That's just the reality of the application layer right now.

Nathan Labenz

Anything else you want to say or touch on before we break for today? This has been great.

Flo Crivello

Well, I'd be remiss if I didn't mention, obviously we're releasing Lindy Teammate. It's on Lindy.ai. I think be ready to see more, not just from us, but I do believe the next 6 months are going to be about multiplayer AI and about this Iron Man suit—about this human-AI hybrid—and these products that create this human-AI hybrid organization.

17. Episode Outro

Nathan Labenz

Hello. Thank you for being part of the Cognitive Revolution.