哪些行业能在 AI 时代存活、新的 AI 基准,以及 2026 年递归学习时间线 | #218
Peter Diamandis × Matthew Fitzpatrick × Salim Ismail × Dave Blundin × Dr. Alexander Wissner-Gross
AI 会率先冲击文档密集型行业,而不是同时抹平所有行业。 Fitzpatrick 将媒体、法律服务和业务流程外包列入爆发区,而石油天然气、房地产仍保留大部分底层职能;竞争关键在于“初创公司能否在大公司造出技术前先拿到分发”,尤其是银行业——其大量应用系统已有 20 多年历史。
企业 AI 缺的不是又一个宽泛的模型分数,而是任务级基准。 公开编码基准显示,模型在 3 年里多项指标提升了 50% 到 100%,但企业需要的是“某项具体任务上的准确率或人类等效水平”——从理赔处理到产权保险都如此;即使达到“80%准确率的超聪明部署”,生产风险仍然过高。讨论因此指向数千个窄领域基准,其所有权可能为被忽视的行业带来曝光度和议价能力。
2026 年的实操打法是跟着价值走,挑 2到3 个用例,并让运营负责人而不是技术部门负责。 Fitzpatrick 建议约 1 个月内做出可运行原型,将其绑定到 CSAT、库存周转天数、缺货率或每通电话成本等指标,并追问:“你愿不愿意拿全年奖金下注,保证部署的用例有效?”对于第一个项目,他倾向于向第三方供应商发出按结果付费的 RFP。
对大多数企业而言,一开始就追求完全自治是错误的架构。 据报道,Klarna 曾称其 AI 每月处理 230万通电话,完成 700 名全职客服的工作,并可每年节省 4000万美元,但在 8到12 个月后逆转方向;Fitzpatrick 的结论是,让智能体处理常规工作,把复杂退款、回写源系统以及希望联系真人的客户升级给人工。“人在回路中将是功能,而不是缺陷。”
通用模型可能让模型构建商品化,但公司上下文、定制评估和数据安全仍是企业问题。 Fitzpatrick 预计企业会定制前沿模型,而不是自行预训练模型,让可替换的模型位于公司专属文档、工作流和基准之下。关键是区分真正的专有信息——例如 Jane Street 的交易数据——与普通后台数据;有必要时使用本地部署或小语言模型,而不是把所有数据一视同仁。
最强的部署都从窄范围的数据整合开始,并且已经带来可量化的运营改善。 Invisible 整合了 750 张 SwissGear 表格,将整体库存覆盖率提高约 30%,并在几个月内让获得可靠预测的 SKU 数量翻倍;它还为 Charlotte Hornets 构建球员移动模型,为 Lifespan MD 搭建符合 HIPAA 的控制塔,并为美国海军提供水下无人机决策系统。反复出现的模式是:具体数据、边界明确的决策,以及可观察的结果。
本期的核心分歧在于,递归自我改进是否会很快消除对专业人类反馈的需求。 Wissner-Gross 给出保守的“最多 2到3 年”外部上限,认为 AI 研究员将达到或超过人类模型研究员;Fitzpatrick 则认为,隐性专业知识、公司特有工作、新模态、新语言、机器人和多步推理会持续制造新的评估需求。他不排除有用的训练前沿在 10到15 年后耗尽,但认为人类专业知识和人在回路中的工作仍将在很长时间内重要。
2026 年的架构判断是多智能体、多模态和重模拟,政府流程自动化可能是最大受益者之一。 Fitzpatrick 预计由 LLM 编排任务专属智能体,增加音频、视频和图像交互,并通过“镜像世界”或 RL gym 在真实部署前测试系统;引用的估算显示,AI 辅助审批可能将能源和数据中心建设周期缩短 50%,许可、福利审批和合规周期则可能缩短 70%。
1. AI 冲击将集中在“文档就是产品”的行业
Fitzpatrick 不接受所有行业会受到同等冲击这一前提。媒体、法律服务和业务流程外包产生大量知识工作和文档,这些活动最直接暴露在冲击之下;而 Wissner-Gross 谨慎收窄了自己的判断:“按我们目前理解的知识工作”已经走到尽头,但不一定意味着知识工作者或整家公司会消失。
石油天然气和房地产明显不同,因为它们的核心职能仍然存在。Fitzpatrick 认为,决定买哪套公寓或哪栋办公楼,整体上仍会像 5到6 年前那样运作,即便 AI 改善了辅助分析;企业应找出真正能够改变的运营环节,而不是宣布整个公司“AI-first”。
竞争关键在于“初创公司能否在大公司造出技术前先拿到分发”。银行业体现了这种张力:其大量应用系统已有 20 多年历史,而 Revolut 等新金融科技公司可以用不同方式搭建系统,不必背负同等的现代化改造包袱。Fitzpatrick 没有判断哪一方会胜出。
能力是另一个独立于行业暴露度的约束。一家 50 人的公司可能连 CTO 都没有,成熟的 IT 团队也未必具备相关技能;即使会 Python,也可能存在能力缺口。Fitzpatrick 的建议很直接:无法招聘或培养相关能力的公司,就通过合作伙伴租用这些能力,不要假设每家公司都能内部自建。
2. 可量化基线决定采用能否安全推进
抵押贷款承保之所以取得进展,是因为银行可以用统计上有效的基线回测决策,并检查最终的信贷决策是否有效且没有构成红线歧视。联络中心也有类似的清晰指标——通话时长、每通电话成本和 CSAT——按理应成为特别适合 AI 的领域。
法律工作分为咨询和商品化生产两部分。Fitzpatrick 预计,一笔大型并购交易仍需要高端法律顾问,但常规 NDA 和标准化文件会大幅压缩;Diamandis 指出,风投文件反复带着 5万美元法律费用上限,只有“大约 8 个旋钮”,却不知为何总是做到“49,999.99美元”,尽管彼此几乎完全相同。
Klarna 成为了反面案例。据报道,该公司曾称其 AI 每月处理 230万通电话,替代 700 名全职客服的工作,每年可节省 4000万美元;在成为智能体成功案例的旗舰故事约 8到12 个月后,公司却宣布回归人工客服中心。
Fitzpatrick 提出的是推测,并非内部消息:一些客户就是要求和另一个人沟通,而退款及其他非一线事项则需要向源系统进行复杂回写。他困惑的是架构——合理设计本来就应该是不断变化的智能体与人工组合,而不是“全人工、全智能体、再回到全人工”。
3. 企业 AI 需要价值驱动的操作系统,而不是战略文件
面对董事会“你的 AI 计划是什么?”这一问题,CEO 应该“跟着价值走”。Fitzpatrick 会选择 2到3 个足以实质性推动业务的杠杆——客户服务、FP&A 预测、库存管理或数字营销——然后只把其中 1到2 个带入真实试点,而不是鼓励无重点的试验。
生成式 AI 颠倒了过去的机器学习部署路径。原型大约 1 个月就能运行,而不是先花几个月搭建,但可靠性要靠之后密集的测试和验证来形成。战略演示文稿不是进展,决定性问题是:“你愿不愿意拿全年奖金下注,保证部署的用例有效?”
对于第一次实施,Fitzpatrick 建议向第三方发出按结果付费的 RFP。内部团队可能缺乏经验,也无法接受“有效才拿钱”这样的约束;与结果挂钩的合同既转移部分执行风险,也迫使所有人事先定义成功标准。
Fitzpatrick 援引一份 MIT 报告称,只有 5% 的企业模型进入了生产环境,并指出组织本身是失败的核心原因之一。应让最好的运营负责人、而非技术组织负责人牵头,并分配一个业务 KPI:服务场景看 CSAT 或通话时长,预测场景看库存天数和缺货率。否则,“百花齐放”最终只会变成无人负责的科学项目。
4. 窄领域基准将成为企业 AI 的控制层
广泛的公开基准,尤其是编码基准,仍然是衡量整体进展的有用指标;按 Fitzpatrick 的解读,模型在 3 年里大多数可观测维度提升了 50% 到 100%。但它们的局限在于相关性:企业需要的不是抽象智能,而是“某项具体任务上的准确率或人类等效水平”。
联络中心基准应把 AI 智能体与公司自己的专家客服在代表性通话上进行比较。理赔处理需要自己的人工等效数据集。“80%准确率的超聪明部署”仍可能风险不可接受,因此每条工作流都需要围绕实际错误、阈值和升级条件构建定制评估。
Diamandis 看到了创业机会:既懂 AI 又懂产权保险等被忽视领域的从业者,可以在其他人之前定义并传播该领域的基准。按他的说法,对无人认领的评估类别建立可信的所有权,可能让一个人“一夜成名”,因为创建基准可能比创建后训练模型更难。
Invisible 已经围绕单项任务构建客户专属基准。通用销售智能体不能像传统 SaaS 那样直接购买;它必须学习公司的产品、知识库、销售方法和“说话方式”,再通过本地评估检验定制后的行为是否真的有效。
5. 通用模型赢得底层,企业上下文留在本地
Wissner-Gross 提到那个“臭名昭著的 BloombergGPT 时刻”:Bloomberg 拥有专有金融数据,希望获得领域优势,但据报道,前沿实验室的通用模型在几个月内就超过了该项目。他质疑,如果通用模型持续吸收专业能力,内部数据和后训练是否还有持久价值。
Fitzpatrick 区分了构建语言模型与添加公司上下文。他不认为单个机构会自行预训练计算密集型 LLM,而是预计它们会利用私有文档、偏好的输出形式和工作流来定制领先模型,同时搭建一个企业层,使更新的基础模型日后能够被“直接替换接入”。
律所偏好的并购文件说明了公共模型无法知道的部分。相关信息仍需要本地定制、后训练或评估。Ismail 将论点进一步推进:关于“我们是怎么做事的?”这类商业秘密知识,可能成为公司的最有价值优势,从而增加内部数据与更广泛 AI 世界之间的保护需求。
Fitzpatrick 反对把每一个字节都视为圣物。银行、医院和交易公司可能会将敏感信息留在本地部署环境,或使用小语言模型,但 Jane Street 的交易数据与其后台预测数据并非同等程度的专有信息。可行的政策应根据实际竞争或监管敏感度对数据分类,而不是拒绝所有外部模型。
6. 数据就绪意味着组装“最低可用真相”
建立在分散客户和产品记录上的智能体“按定义就会出问题”。因此 Fitzpatrick 每个用例都从输入开始,但他不主张耗费 5 年完善整个企业数据湖——许多大型组织已经走过这条路,却仍没能让每个数据仓库都准确、可访问且相互一致。
信贷承保可能需要 5到6 个核心类别:贷款本身、市场环境、公司财务状况、信贷担保以及相关变量。它不需要商业银行的每一条记录。实际问题是:“针对这个具体用例,我需要哪些数据?”然后进行有针对性的修复。
生成式 AI 还提升了那些从未进入记录系统的信息价值:视频、图像、自由文本和其他非结构化文件。这些资产过去并不是人们试图掌握的对象,因此第一步是明确任务,并让相关数据具备可用性。
Blundin 的 Vestmark 案例进一步说明了这种区别:对账记录展示账户最终如何完成对账,却没有记录员工为解决问题做了什么。AI 助手可以先观察并加速工作流,再生成自动化所需的人类反馈或调优数据。另一位银行 CIO 则面对 300 个受保护的客户数据库,每个数据库都由不同的产品孤岛负责。
7. 当决策和数据边界清晰,领域部署才能奏效
对 Charlotte Hornets,Invisible 基于来自多个大学和国际比赛场馆的单点视频,微调了计算机视觉模型。传统统计可以捕捉得分、篮板或正负值;该模型则衡量球员移动、空间以及谁在创造空间,跨越不一致的镜头角度,为选秀评估者提供交易型统计数据无法揭示的特征和球员适配度证据。
Lifespan MD 从数据架构开始,而不是自动诊断。Invisible 的 Neuron 平台整合出符合 HIPAA 的多租户患者、医生和诊所绩效视图,使用户能够提出“35到50 岁男性最常使用哪些长寿测试”等问题。患者数据仍保留在各个诊所,临床医生和中央运营人员只能获得与其职责相匹配的访问权限。
Fitzpatrick 认为,临床决策比行政事务更模糊。美国人均医疗支出约为 1.3万到1.4万美元,德国或加拿大约为 2500到3000美元,其中约 30% 到 40% 流向行政环节。更近的机会是消除排班和文书工作,同时让医生“获得更强的能力”。
其他项目扩大了这一模式,但没有改变其本质:Invisible 与 SAIC Vantour 和美国海军合作,研究围绕传感器丰富的水下无人机群进行决策;在 SwissGear,它整合了 750 张表格,将整体库存覆盖率提高约 30%,并在几个月内让获得可靠预测的 SKU 数量翻倍。库存项目的目标是同时减少缺货和过量库存。
8. 人类反馈正变得更专业,而不是简单消失
Invisible 的 Meridial 业务负责训练模型,企业业务则构建定制应用。Wissner-Gross 追问:如果 AI 研究员能够构建模型、数据集和基准,是否会消除人类机器学习贡献者市场的必要性;Fitzpatrick 回应称,RLHF 即将消失的预言已经持续了 5 年,却始终没有与部署现实吻合。
工作正从商品化的“猫狗、猫狗标注”,转向博士和硕士级别的评估、受控 RL 环境、模拟和 RL gym。Fitzpatrick 表示,研究和运营经验都支持合成数据与人类数据结合,尤其是在多步推理场景中,因为幻觉可能沿着推理链不断累积。
他最典型的例子刻意保持窄范围:一个研究 17 世纪法国建筑演变、并且用法语工作的模型,仍然需要合格的人类来验证。同样的逻辑适用于新的法律数据集:律师助理或并购律师必须判断生成的文件是否真正达到专家工作的可比水平。
Wissner-Gross 的反驳针对的是效率,而不是否认反馈价值。强化微调可能比大规模标注劳动力需要更少的人类工时,而 RL 环境一旦建成就可以扩展。Fitzpatrick 承认形式会演化,但指出后训练反馈只占总算力成本的一小部分,却是最有价值的输入之一。
9. 递归自我改进可能跑赢企业吸收速度
Wissner-Gross 给出“2到3 年”这一保守外部上限,认为 AI 研究员将达到或超过构建机器学习模型的人类研究员。他预计递归自我改进能力将在 2026 年加速,而企业仍会“以蜗牛般的速度”前进,使实施、变革管理和工作流瓶颈成为硬约束。
Fitzpatrick 不排除机器可能在“10到15 年后”耗尽有用的训练前沿,但认为短期内不存在供给不足:新语言、新模态、机器人任务、法律细分领域和公司专属工作流,会持续创造更窄的评估需求。通用能力不会凭空制造私人先例,也不会制造只存在于资深人士头脑中的隐性专业知识。
销售是他反驳完全通用化的例子。最优秀销售人员的打法通常没有文档记录,而一个充斥着 500 家邮件型 SDR 供应商的市场,可能让真正的互动变得更加稀缺和有价值。“人与人接触的元素会变得越来越重要”,尤其是在信任、例外处理或真实性决定结果的场景中。
Wissner-Gross 将商业机会定义为能力与采用之间的缺口:Invisible 这样的公司可以成为连接前沿实验室与企业的“润滑剂”。在他的联络中心测试中,80% 的人更喜欢 AI,而不满意的 20% 可能把系统“折磨到彻底崩溃”,从而为路由、评估、数据和异常处理创造有利可图的业务。
10. AI 原生挑战者会重构流程,而不是自动化岗位盒子
Ismail 的 Canon 思维实验以功能流程取代部门思维:营销、零售销售、注册、墨水补充、预测性维修、追加销售和会计,成为由 AI 管理的一体化打印机生命周期。产品自行报告状态并触发下一步行动,最终可能让人类“90%退出回路”。
他把当下逐个员工自动化的做法称为“电视上的收音机”——就像把广播员放到电视里照读旧稿,而不是适应新的媒介。AI 原生公司从目标功能出发,重建整个运营系统;传统企业则倾向于把助手贴到既有岗位和交接环节上。
Diamandis 再次回到传统企业的分发优势。他的方案是邀请全球 AI 创业者展示他们将如何颠覆公司,资助最好的 5 个项目,给予数据和访问权限,然后收购有效项目或取得多数股权,把它们变成新公司。目标是先“边缘创新”,再让边缘取代传统中心。
Ismail 将 Apple 的小型秘密团队视为组织范本,称大约有 18 个团队在研究不同产业,并耐心迭代,直到进入时机成熟。大型运营商可能难以攻击自身已经优化的核心业务,但它们对相邻行业的理解,可以支持面向邻近市场的 AI 原生创业项目。
11. 多智能体团队、多模态和镜像世界定义 2026 年
Fitzpatrick 对架构的第一项判断是多智能体团队。企业不会让一个智能体做出所有决策,而是把任务专属智能体训练到高准确率,再置于一个负责编排整体逻辑的 LLM 之下。联络中心就是天然案例。
第二项判断是多模态跃迁。音频、视频和图像将成为人们与模型互动方式中更大的组成部分,让交互超越历史上以文本为主的界面。此前篮球、医疗和无人机的案例说明,即使软件界面不是多模态,企业数据表面早已是多模态的。
第三项判断是“镜像世界”或 RL gym:通过模拟环境和数字孪生,让编码、联络中心或制造系统在接触生产环境前执行任务并接受测试。对于现实错误代价高昂或具有危险性的系统,模拟提供了可重复的测试方式。
数字人也是 2026 年可能出现的发展方向。Fitzpatrick 表示,用他的公开言论训练一个数字人并不困难,并提到体育相关数字人项目已经在推进;他预计,人们可能更愿意通过数字人与熟悉的人物交谈,而不是使用通用聊天机器人,让合成人格更自然地成为交互的一部分。
12. 人类工作将转向实体存在、真实性和新类别
被问到“最后留下的专家”时,Fitzpatrick 给出的是大类,而非 3 个精确职业。地震勘探、钻井现场运营、房地产选择、实体技工以及围绕人与人互动展开的工作,会存续更久,因为它们的功能不只是搜索文档。近期冲击仍集中在业务流程外包、法律服务和媒体。
他不认为任务被冲击就等于就业下降。媒体的经济模式和渠道发生了变化,但 Substack、Medium、博客和其他形式创造了更多媒体创业者。Fitzpatrick 援引估算称,美国每届高中毕业生中约有 25% 最终进入一个在其高中时期尚不存在的领域,而美国就业中已有 20% 属于数字生态系统岗位。
Wissner-Gross 给出 3 个相互竞争的最终职业候选:政治家,因为他们制定法律;最伟大的物理学家和数学家,因为他们代表人类智力工作的顶峰;或者那些客户要求真实人类作为交易对手的职业。Diamandis 将第三种假设概括为:“品味定义者将占据主导。”
政府可能带来最清晰的社会回报。Fitzpatrick 援引一项研究称,AI 辅助审批可能将能源和数据中心建设周期缩短 50%;OECD 的估算则显示,许可、福利审批和合规周期可能缩短 70%。他认为,围绕支出和基础设施部署进行项目管理、压缩建设周期,是简单且价值很高的应用。
Most of the public focus today has been on the large public benchmarks for things like coding. I think the problem is, though, Matt Fitzpatrick, CEO of Invisible Technologies.
It sounds like your position is that we need thousands of new narrow benchmarks to capture perhaps every labor category and every industry vertical. That is an interesting second part of this, which is:
We're going to see the largest disruption ever in 2026 from companies that don't make this change.
There are many sectors where the structure of what the industry does is going to change. If you think about knowledge work—the production of large amounts of documentation—these technologies are very disruptive.
What are you seeing most companies get wrong on their mission to implement AI?
You've got 2 different challenges.
In today's episode, we're going to be discussing why all companies need to become AI companies in 2026, how they do that, and what happens if they don't. We'll discuss whether big legacy companies can even make such a dramatic change and how they can best do it.
We'll go over some fun and meaningful AI use cases. I think they're going to get you excited about what you can do, and we'll dive into some predictions from our guest for 2026.
Today, joining us is a friend of the pod, Matt Fitzpatrick, who spent more than a decade at McKinsey, rising to the position of global head of QuantumBlack Labs. I love that name, QuantumBlack. It's so cool. He led the firm's AI software development R&D and global AI products.
A year ago, Matt joined as the CEO of Invisible Technologies, a company started by a brilliant friend of mine, Francis Pedraza. For those of you who don't know Invisible, the company is a modular AI software platform that uses AI training and provides AI training for most of the large language-model providers out there. It builds custom workflows and agents for enterprises. The company anchors its work in creating clean data and human-in-the-loop delivery to ensure measurable business results.
Matt, welcome. Good to have you here. Hey, Matt.
Thank you for having me. We've got DB, AWG, and Salim. Almost happy holidays, guys. It feels like we're on this pod every other day. I think we should just move into a large podcast house, and we will have fully documented the singularity.
I'm really looking forward to hearing from Matt because on Thursday we have to do our predictions for next year, and Matt is going to give us a ton of insight today. One of my predictions out of the gate is that enterprises are going to move super stupidly slowly compared to AI capabilities. Matt is the world-leading expert on the intersection between AI and enterprise, so I cannot wait for this.
You cannot cheat this way, Dave, and use Matt's predictions as yours.
I can't?
No.
I'll be there. But everybody listening to this pod will know that.
All right. Take good notes, if nothing else. Are you guys ready to jump in?
Matt, I'm going to kick it off with a broad question. Within the past year, we've heard from every company out there and every CEO that we're going to be pivoting to become an AI company. Salim, in the last pod, you said something like, "We're going to see the largest disruption ever in 2026 from companies that don't make this change." And I think, Alex, the term you used is that they're going to be cooked if they don't.
I said knowledge work is cooked. Not knowledge workers, not companies. Knowledge work as we currently know it.
So you don't think that companies are going to be cooked if they don't make the transition to AI?
I think we're going to see many more companies over time, and many more smaller companies as well.
We're going to dive into that.
In an earlier episode, we pointed out that when you thought you were at product-market fit and scaling a SaaS company, you're toast because everything needs to be rethought now given AI. This is now applying to big companies as well.
Matt, the question to kick this off is: Can every company truly become an AI company, and how? Which companies and industries do you think need to disrupt themselves now before they become basically irrelevant? It's a softball question to kick us all off here.
Peter, save the hardballs for me, Matt.
Always. I think your second question relates to your first question in some ways. I don't think, based on all the data that has come out on this so far, that all industries are going to be impacted equally by this.
There are some sectors and areas where you're going to see materially different impacts. I think areas like media, legal services, and business process outsourcing are sectors where the structure of what the industry does is going to change.
If you think about knowledge work—the production of large amounts of documentation—these technologies are very disruptive. Where I think the hype has been a bit overblown is in sectors like oil and gas or real estate. The function of what they do is going to stay pretty consistent.
I think most of the good analytics of how job dynamics will change over the next couple of years will get at this. The decision about which apartment building or which office building to buy is going to function pretty similarly to what it did 5 or 6 years ago.
I think the question is which parts of your business can really change with AI. It's not all of them, and some sectors will be more or less affected.
The second part of your question—can everyone actually become an AI company?—is also an interesting one. There aren't that many people who know how to build these sorts of models or deploy these sorts of models well.
One of the big challenges is whether you have the expertise in-house to do this. How do you think about adjusting the operating functions of your company to do it? Is it the same team you have in your IT function doing it now?
Particularly, Peter, I know the set of folks that you and I have spoken with in the past—small businesses. If you're a 50-person company, it's hard to deploy a lot of this stuff at scale if you don't even have a CTO in-house.
I think there's a mix of whether your industry is going to fundamentally change and what the actual core competencies your company has to implement it are. I think what you're going to end up finding is—
So, do you end up bringing a chief AI officer into your company? Are you going to bring that capability in, or are you basically renting it?
Part of the other thing that's going on right now, which we've talked about on the pod a lot, is that your competition isn't really the large multinational. It's the AI-native startup that came out of nowhere and reinvented itself from the ground up as an AI-first company, right?
Like, down and dirty, Matt: Which happens first? I can get a mortgage by talking to an AI and get it done in under an hour, or we're walking on Mars with our own 2 feet? Which of those 2 things is going to happen in the real world first?
The way I've heard the question asked is: Do the startups get distribution before the big companies build the technology? I do think that will be the tension in a lot of ways.
I think there's a lot of big, established companies that are going to figure out how to do this really well. If you take a sector like legal services, I do think the big law firms will figure out how to use a lot of this over time.
I think there are sectors where banking is a really interesting one to look at right now. If you look at the age of the application footprints in banking, most of the tech that exists in banking is north of 20 years old.
You do have a bunch of very fast-moving newer fintechs that are approaching it in different ways, companies like Revolut. I don't know how that plays out, but I do think that becomes the question in a lot of ways: Which moves faster, the emerging entrants or the modernization of the existing companies?
Peter, to hit on what you were asking as a second part of that: Do you buy or rent? I think that's something you've got to be really honest with yourself about as a company.
The idea that everyone can buy, or everyone can hire people to do this, is challenging. The challenge of trying to adapt an existing IT function to do this is that many of the skill sets people hire for—even things like knowing Python—have gaps in them.
I think the answer that most companies I've seen who don't have the resources in-house are coming to, through a very direct push, is that they're finding ways to rent or buy this externally and partner with folks that can allow them to do it.
I think legal and accounting are really cool case studies, and I know you know more about this—your McKinsey time and QuantumBlack. You're the guy understanding and parsing all of this. But they're really cool because they can be replaced by a startup, like Harvey.
Dave, two things I'd say about that. I think one challenge of implementing GenAI in the enterprise setting is having a statistically validatable baseline to compare against.
As an example, if you take something like mortgage underwriting, it has made huge progress in a very positive way. The percentage of mortgage underwriting that's now done by a very guardrailed and very effective set of algorithms developed by the banks is pretty high because they can back-test and say, "This is a correct credit decision that has no redlining or anything else."
But if you think about a domain like this, the reason contact centers have been one of the use cases where we've seen a lot of adoption is that you do have a clear baseline. You have time per call, CSAT, cost per call—you have a set of metrics you can compare against.
Something like, "Let me generate an investment memo," which is different in format at every firm—it could be 10 pages versus 40 pages, and the content is different—has made it harder for folks to build baselines. I do think that's why legal services are interesting: there are certain areas of legal where those baselines are clear. You can look at what documents are really good for, like an ISDA agreement.
Where I think you're going to see this in a lot of different segments is that the high end of that market still persists in a really differentiated way. If you're doing a large M&A transaction, you're still going to want a really good lawyer's advice.
Where it changes, I think, is in the more basic, "Produce an NDA" type of work. I think that's going to be one of the shifts again: really good human guidance is going to persist forever. It's the basic commodity information that right now a lot of people are paid probably excessive amounts of money to produce.
Yeah, well, the NDA is pretty extreme. But I'll tell you, the venture fundings that we do—we do tons of these every year—the term sheets always say that the company we're investing in will bear the cost of the legal work, capped at $50,000. The documents are freaking identical every single time. There are about 8 knobs, and you could store all the combinations on the smallest thumb drive in the world.
I'm like, how is this $50,000? It always runs up to $49,999.99. It's like, wow, what a miracle.
That to me feels like it would be on the mid-to-hard end of the scale, yet it's still so doable. An NDA is a no-brainer. Mortgages are no-brainers.
I completely agree. I think what's been interesting, though, is how slow the actual adoption curve has been in contact centers. Contact centers should have had generally measurable CSAT scores, and people don't really like most contact-center interactions. The general customer feedback you get is pretty unhappy, and that's been true for a decade.
Yeah, I guess technology would have—let's talk about the whole Klarna thing, actually. I know you're an expert on this. The Klarna thing has been really interesting to watch. Wait, tell us the story. What is the Klarna thing?
Well, I was not involved in Klarna, but I can say at least what I know from reading about it and what my hypothesis would be.
Basically, Klarna announced that they were going to move entirely toward a fully end-to-end agentic contact center. By the way, the interesting thing was that at that time, they were the most frequently cited example of agentic success in deployments. Then, about 8 to 12 months later, they announced they were rolling the whole thing back and moving entirely back to human contact-center agents.
I found the entire evolution interesting because, if you think about how these systems should be defined and deployed, a multi-agent system should have an orchestration of the types of calls. You'd have a set of validations on which calls could go well or badly, and you'd have some sense of where you need escalations to human agents.
You would never want to move to doing everything agentically. This is a theme in this whole area: you're never going to want to do everything agentically. You're going to want humans in the loop in almost every industry and on almost any topic.
If these models are trained on precedent data, you can train them really well to continue that logic. But you're going to want humans for things where you don't really have precedent data, or where you need them to work through complex situations for which you don't have enough historical information.
I found the entire structure of how the change happened quite confusing because you would always want to keep a contact center as a mix of humans and agents, and then evolve the mix based on the topics. The whole movement from all humans to all agents and back to all humans was confusing, I think.
Salim, you are presented with a question.
I just wanted to give out some details here. The Klarna situation was that they rolled out an AI to handle customer-service calls, and the claim was that in the first month it did the work of 700 full-time agents, handled 2.3 million calls a month, and was projected to save them $40 million a year.
They were proudly saying this was month 1, and it was only ever going to get better from there. When I saw that, I thought, "Okay, if I were doing that, this sounds like a PR exercise more than anything real, because you'd never put that out in the first month. You'd wait a couple of months to see what exactly happened."
Matt, you may be able to give a little more color on why they rolled it back in the end. Did they find the hard cases were too many? Was the exception handling too much? Or was it a cultural backlash? What was it exactly that had them undo the whole thing?
I don't know, in the sense that I haven't worked with Klarna. But you hear a variety of different pieces of feedback on why folks have struggled in contact centers.
One reason is that there are cases in which humans just want to talk to another human. So, some of the PR around saying, "We're moving to only agents," has its challenges.
Two, a lot of the challenges—and where contact centers are most sensitive—is with non-first-line call-resolution topics. It's not something like, "Check your balance." It might be something like, "Process a refund," which is pretty complex. You have to write back to the source systems.
It was surprising to me how quickly they rolled that out, and I wonder how well it was able to deal with some of the more complex functionality in that example.
Right. You go from level 1 to levels 2 and 3 very quickly on those support calls, and then you do not want an AI dealing with you.
Can we get back to the main question here? 2026 is coming up. If you're listening to this in 2026, it's here now. Here's the question: you're a medium-sized or large-sized company, and your board of directors has just said to the CEO or CTO, "Guys, what's your AI plan? What are you doing?"
We're seeing that over and over again. What is their first reaction typically, and what should they do? I want to get some of the fundamentals here because I want to serve our listener base in that fashion.
If you're that CEO, you've got 2 different challenges. One is, what are the things I should focus on? And two is, who should do them? Do I have those skills in-house?
The first thing I'd start with is making sure you answer the first question. I do think this is a question of following the value. I'd go down a list. I would not start with letting 1,000 flowers bloom. I would start by identifying 2 or 3 things that, if you do them well, materially move the needle for your business.
Maybe it's customer service. Maybe it's forecasting in your FP&A function. Maybe it's inventory management. There are definitely 2 or 3 things that almost any business on Earth, even a small company, has. Digital marketing is probably another one that you see pretty frequently.
You focus on 1 or 2 of those and make sure you get to a pilot stage in those areas.
Meaning, not a strategy document. I do think the one thing that anyone who's spent real time in this space will tell you is, if you take the paradigm of how machine learning is deployed—where you spend months and months building something, and then it works, and you can underwrite statistically that it works—this is kind of the exact opposite paradigm. You can get a prototype up and running in a month, but you have to do a lot of testing and validation to make sure you can trust it. And so, it is really a function of making sure you can get something up and running, and testing and validating.
You know, Peter, the question I always ask is: would you bet your annual bonus that whatever use case you deploy works? And that's a complicated thing. If it's, let's say, generating a claims-processing review, and you have to do 10,000 of them, most companies don't know how to say whether that works or it doesn't. And so, just to summarize, make sure you have a list of the 2 or 3 things that move the needle. Make sure you get to a proof of concept in one of them. And I would probably do that first use case as an RFP to a third-party vendor that gets compensated based on results.
Yeah. And I say that very specifically because I think if you do it in-house, the odds are the in-house team has not had a lot of experience with this. And so, you also can't hold them accountable in the same way: you get paid if it works. And so, I do think tying it to outcomes limits your risk. I mean, that's still the business model for Invisible, right? You're paid by money saved.
Correct. We do outcomes. Yeah, outcomes in various ways.
Yeah. Alex, I want to bring you into the game here. Much appreciated. So, maybe just as a preliminary matter, for full disclosure, I have no financial interest in Matt's company, Invisible. I do have a number of questions, though. First question, maybe pulling the thread on testing. One of the things that we talk about here on the pod all the time is benchmarks, the importance of benchmarking. I'm curious, given that—
We talk about that constantly, Alex. That is all we talk about. We talk about nothing else. That is all we talk about. Oh, wait. Maybe that's you. Okay.
Given that's all we talk about, as Dave just mentioned, and given that Invisible is also in the business of training so many models, what benchmarks do you think most need to be brought into existence in the world? What's most missing? What are the top 3 benchmarks you'd like to see summoned into existence?
Yeah, look, I think—and you've seen a bunch of these start to get publicized in the past couple of months—but most of the public focus to date has been on the large public benchmarks for things like coding. And I think those are very useful as metrics for whether the models are improving broadly. That is why you've been able to see, by any standard, if you look at our last 3 years, the models have improved 50% to 100% on most dimensions that you can look at.
I think the problem is, though, if you think about enterprises or small businesses, your benchmark for most cases is not a broad-based, accurate cognitive benchmark. It's accuracy or human equivalence on a specific task. And so, what I think you're going to see more and more need for is custom evals on highly specific topics. If you go back to the contact-center example, the benchmark you'd want to build if you're going to roll this out for a contact center is a series of expert agents that are in your contact center, how they perform, and then how the AI agents perform similarly. The same applies to claims processing. Basically, most businesses are going to have to get comfortable with doing what's called an eval or a custom benchmark for the tasks they're trying to modernize.
Because an 80% accurate, very smart deployment is not enough—there's still too much risk in that rollout framework. And so, I think the way that we think about benchmarking will evolve from broad-based benchmarks to hyper-specific benchmarks.
I freaking love that because I can immediately see 10,000 listeners right now who just found a calling in life based on what you said. All this benchmarking within any of these domains is really, really hard to figure out unless you know the space. Take title insurance. What's the benchmark for successful AI in title insurance? Somebody in that industry listening to this pod right now is going to be like, “You know what? I was an early adopter of AI, and I know this space inside and out. That's my benchmark to own.”
And if you declare yourself the owner of it and then broadcast the benchmark, the evidence so far is you become an instant star. Nobody's grabbing topic ownership in all these areas, and if you just get there first, you become an instant star.
I completely agree with that. In this era of post-training as a commodity, if you own the benchmark, often the benchmark is the hard part, and you can leverage existing resources to post-train an off-the-shelf model. I am curious, though, Matt, maybe following up on this. It sounds like your position is that we need thousands of new narrow benchmarks to capture maybe every labor category and every industry vertical. Assuming that's correct, is that something that Invisible is working on, can be working on, or should be working on?
Yeah, we do spend quite a bit of time working on that. In fact, a lot of the time, what we're building is customer-specific benchmarks for an individual task. That is a lot of what we think about: how to test equivalence for a given task.
And I think one of the things that folks have not fully realized is, let's say you take a really high-performing LLM and you want to tailor it to your individual context. That process of actually fine-tuning it off of your data.
I think one of the challenges is that people were hoping this would be a SaaS buyers' paradigm, meaning I could just buy something off the shelf that would solve everything I needed. So, I wanted to buy a sales agent; I wouldn't have to do anything. I could just take in a sales agent that would sell well. And the reality is that's pretty hard to do. You need to actually train it up on your specific knowledge corpus and your information.
The way we would think about it is: you take the LLM, or you take an agent that's been trained for sales, and then you fine-tune it off of your specific company information, your products, the way you sell, and your way of speaking. Then you have to build an eval or a benchmark against that to say whether this is performing well or not on that task.
Well, a quick follow-up question, if I may, because there was the sort of infamous BloombergGPT moment, where Bloomberg was sort of in quasi-competition with the frontier labs. They had a wide variety of internal proprietary data sets. Their original plan—this is now sort of an infamous episode from 1 to 2 years ago—was to offer their own proprietary frontier model, basically, but, critically, pre-trained and/or post-trained off of their internal data sets.
The plan was to achieve superb performance in the financial domain because they had all the data, or a lot of data, that were not broadly available to the general public. But what actually happened is the generalist models offered by the frontier labs, which were trained basically off of the internet and more or less publicly available data sets, within a few months leapfrogged BloombergGPT.
And so, I guess the moral of that parable, in my mind, is: how far do you think we can really get with proprietary data sets and proprietary benchmarks before the generalist models completely wipe the floor with them?
Sorry, to clarify, I'm saying you use an LLM. The process I'm describing of actually fine-tuning a large language model for your specific context is basically adding more context. Most of the LLMs offer a paradigm where you can do this, where you can add your knowledge corpus and train it to be more specific to your individual context.
I don't think you'll see individual institutions building their own LLMs. I think that's a very compute-intensive, very difficult thing to do. I think you'll see them tailoring the large language models to their context.
Sure. To be clear, I wasn't asking whether you think every institution is going to get into the business of pre-training its models. I was rather asking whether you think post-training—which is inclusive of supervised fine-tuning, reinforcement fine-tuning, and a variety of other post-training—has a long-term future.
Or will, maybe in 1 to 2 years, we just use a pre-trained plus post-trained generalist model off the shelf and not need any internal benchmarks or any internal data sets for post-training?
Well, I think there are clearly going to be use cases where you are going to need the context of the individual company, right? If you take the law-firm example, there are documents that a company has on how they want their future-state documents for M&A agreements to look, right? And the LLMs are not going to have that information.
So, at some point, you are going to have to see the post-training layer happening at the enterprise. And what we're seeing more and more is there are ways to design that layer so that, as new models evolve, you can drop those in. We are seeing more and more folks experiment with that.
So, they’re using all the new technology that’s being rolled out.
I think, in fact, what’s going to happen is that, over time, that edge in data is going to be the most valuable part of any company. Is that trade-secret type of “How do we do things?” At some point, it may leak into the public models.
Like, if you used OpenAI, right? If you use any of the frontier models connected to them—I remember we were talking to Replika, et cetera—people are using it, and then the data is going straight into the cloud, right? That’s kind of dangerous. They’re going to have to solve that layer in a very powerful way.
That’s one of my predictions to forecast, et cetera: we’re going to need to see a layer of protection between company data and the broader AI world.
Matt, I want to make this a little more tangible. I know you can’t talk about the work you’ve done with the hyperscalers, but you’ve identified, I think, 5 or 6 cases where you can speak publicly about it. If you don’t mind, maybe we can toss a few of those in and talk about them as concrete examples.
Since Alex made his no-financial-involvement statement, I will say I’m a proud advisor and am conflicted, in a positive fashion, supporting what Matt and Francis are doing. Do you want to pick one of those? I loved the example on the basketball court. Can you speak to that one?
Yeah, sure. We worked with the Charlotte Hornets on fine-tuning custom computer-vision models for draft preparation. In their case, they wanted to look at the spatial movement patterns of players on a very broad scale, across single-point cameras at a whole host of different college universities and international locations.
We fine-tuned a custom computer-vision model to specifically look at the movement patterns they were interested in before the draft. That was a big part of their draft evaluation.
In English, you basically took the video and were able to use models to evaluate every player based on the video, to see how well they performed in every different—
I’m not a sports guy, so it’s the—
Yeah, that’s becoming clear here, actually.
Yeah, yeah. It says the salt economy. [Laughter]
Sure. If you take typical NBA stats, there are things like points, rebounds, and what’s called plus-minus, which is one ratio that’s often used. It’s the amount you score versus the amount you give up when you’re in the game.
But those are mostly transactional stats. What they don’t look at is the movement patterns of the players: who creates space and where people are positioned at any point in time. That’s actually a lot of the most interesting data.
If you go back to some of the original baseball analytics that Billy Beane did for the A’s, it’s the movement patterns of players and who is in the best spacing, right? There are companies that do this in very consistent formats, such as on the same court. What we’ve been able to do is perform that analysis over many different camera angles and many different stadiums, very quickly.
That uses custom computer-vision models. We’re effectively able to take a single-point camera and understand the movement patterns of players in many different environments.
And how do the Hornets use this? For team selection? Player selection?
Draft selection—to understand which players fit the characteristics they were looking for.
Fascinating. It’s a complicated problem, too, because chemistry between players matters. It’s not just about finding the best player; the chemistry between players matters, too. It gets infinitely complex, and it’s a cool little case study.
Gavin Baker was saying recently that, in fantasy football leagues all over the country—which I used to love before I ran out of time—
Now you have an agent doing it for you and having fun.
That’s exactly the point.
We’re now obsoleting human sports leagues, replacing them with robot sports leagues and esports.
Yes. Very 21st century, not 20th century.
That’s right. Great fantasy-football players are losing all over the place because the AI agent is tracking a huge amount of more detailed data. If you look at the video footage, somebody might be making it up and down the court very slowly. Nobody’s going to notice that, but the AI will notice it in a heartbeat. Then that just goes into the great model. It’s really a cool little case study.
Since you asked a little bit about how a traditional business might be thinking about doing this, I’ll give a slightly different example, which is Lifespan MD. Peter, I think this one will resonate with you in particular. It’s a concierge—
I know Chris, who runs it.
Yeah. Lifespan MD is a concierge-medicine business. You can think of it as having a network of practices both internationally and in the United States, all of which have very different sets of data on their patients and practice information.
The thing I always start with in any AI use case is that you have to get the data right. Before you can even start with AI, you have to make sure that you have the structured and unstructured data together that you want.
The first thing we’re doing for them on our data platform, Neuron, is creating a HIPAA-compliant, multitenant cloud instance where we bring together all the patient and provider data that’s of interest. We start to bring a 360-degree view of both the patient and the practice.
You can start to think of things like which longevity-focused tests male patients between 35 and 50 are using most frequently. You can start to think about patient outcomes that are really interesting. If you want to understand practice performance, or where you have certain patients who are not compliant or not as interactive, it’s effectively a control tower to understand everything that’s going on across that footprint of practices.
Then I think the area where generative AI has become more important for that is chat agents, where people can ask questions—knowledge-management systems that allow them to interrogate and ask questions of all the key data from all of those practices.
One of the key challenges is that, obviously, in health care, you have to be extremely careful about which data is stored locally at the practice versus how it’s brought centrally. The HIPAA-compliant, multitenant cloud is one of the key components of that. It makes sure that no patient data leaves the premises of the individual practices, while doctors can access certain things and certain practice metrics are organized centrally.
I heard the coolest thing this week. It’s a quality-assurance company that has invented “Talk to Your Defect.” It’s just the coolest concept. The defect actually has a personality, and you can ask it questions about itself, like, “Where did you originate?”
I can totally imagine what you just said in health care being “Talk to Your Illness.” You have a conversation with it: “Where did you come from? How do I treat you? Are you getting better or worse if I do this thing?” It’s talking back to you with a personality. It’s the coolest idea ever, isn’t it?
I think it’s amazing. It’s one thing with the defect. It’s a little awkward when you say, “Here’s the bacteria you’re talking to.”
Well, I just mean the defect is real. “Talk to Your Illness” maybe gets a little weird. I don’t know what voice you would give it. A Voldemort voice or something.
Tell me, how do I kill you? How do I dispatch you? [Laughter]
Well, Dave, one thing I’d note there, too, is that I think there’s a question—and I get asked this often—of how sectors evolve. Peter, you asked earlier how sectors evolve. I think the question of whether the decision-making around individual patient care changes with generative AI is much murkier.
I think the easier place to start, and where it would be very interesting, is that the United States, as an example, spends about $13,000 to $14,000 per patient per capita on health care, compared to $2,500 to $3,000 per capita in, say, Germany or Canada. Something like 30% to 40% of that is administrative cost, and that is not an administrative cost that anyone wants to bear.
This is something where I actually think the idea Lifespan MD is pursuing is not to change the standard of care, but actually to make the physician even more empowered—to take all of the really painful administration and scheduling and make that the part they don’t have to deal with anymore. AI should do a huge amount of damage in those areas.
Exactly. What are you seeing most companies get wrong on their mission to implement AI?
I think it’s a couple of different things. The first one is a lack of focus on data as the starting point. I think the challenge is that if you tried to build an AI agent on fragmented customer and product data, it’s going to break by definition.
You have to be in a place where the data you’re going to feed into the models is clear and working. That’s been one major challenge.
Do you think most companies—as a whole, in the medium and large size—have clean data? How long does it take a company to get its data into a format and to a level of fidelity that’s useful? Is this a hard lift or an easy lift?
It depends. If you take the paradigm of, “I’m going to put everything in a data lake and get everything right,” that can take 5 years.
And the reality is that most big companies have spent half a decade trying to get all their major data schemas in order. But I think if you start with the question, “What data do I need for this specific use case?”—let’s take credit underwriting. To do that well, you need a set of data around the credit itself and the market. You probably have 5 to 6 core data variables you need: the core financials of the business, the security of the credit, and all those kinds of core pieces.
But you don’t need every piece of data across the entire commercial bank to be right. You need the core elements for that use case. And so I think companies that are focused on the exact data they need to get right have done pretty well.
But I do think that trying to get all the data right—I mean, you’ve also seen the enterprise for a long time, Peter. If you asked any Fortune 1000 company to look at their full data repository and determine how much of it is accurate, working, clear, and accessible right now, very few companies have that.
So I do think it’s important to be very tactical about what data you need. The other thing I think, for generative AI in particular, is that a lot of the most important data is non-system-of-record, unstructured data. It’s things like images, videos, and text files. Those are just not things that people have tried to master historically. And so I think the first step in this is saying, “What is the thing I’m trying to solve, and how do I make sure I have that data ready?”
Yep. One thing I see a lot of—I had a long board meeting this morning with a company that’s very AI-forward in portfolio accounting, a company called Vestmark. And the data, for account reconciliation, for example, is abundant. But it doesn’t tell you what the person actually does. It just tells you how it was reconciled.
So now the path to success is first the AI assistant, which helps accelerate you through the day, but it also knows what you’re actually doing. Then that accumulates, and then that becomes the RLHF for the training or tuning data. Because what you’re trying to do is, “What are you doing, guys?” And that’s not really represented in the data.
But a lot of times you go talk to a bank or an insurance company, and they’re like, “Our data is our advantage. Go ahead, bomb it into the neural net and train it.” You’re like, “I don’t even know what that means. I’m just going to throw all the terabytes of spreadsheet data in and see what happens? That’s going to go Clippy on you?”
Well, you have all sorts of other issues as well. I was talking to the CIO of one of the biggest banks in the world, and they have 300 different customer databases. Three hundred: one for mortgages, one for loans, one for this. The mortgage people don’t want to tell the loans people about their customer data, so they guard it jealously. It’s a total disaster for the poor CIO. Fascinating. Alex—
I think these are all very interesting points. I’d like to, if I may be so bold, jump up several levels and speak a little bit more about the business model of Invisible. My understanding—correct me if I’m wrong, Matt—is that there’s an element of the business, I think it’s called Meridial, that serves as a sort of marketplace for ML freelancers, if I understand correctly.
And I’m curious. I think, in my mind, one of the many elephants in the room in this conversation is that we’re arguably on the edge of recursive self-improvement. All of the frontier labs, more or less, I think would agree with the assertion that we’re nearing the point where you could have an AI researcher, where you just turn over computer resources to that AI researcher, and the AI researcher does as good, if not a better, job than the human AI researchers who work for the frontier labs.
If that is indeed the case, surely one of the several elephants in this room—but given limited time, let’s focus on this one—is that the need for a marketplace of ML freelance researchers to train models evaporates entirely as we start to reach the point where AI researchers can build custom models off of custom datasets and custom benchmarks for each client. Doesn’t it evaporate entirely?
Yeah. So, as you said, we have 2 sides of our business. One side, Meridial, is where we train all the large language models. Then, on the enterprise side, we build basic custom applications for enterprises.
Look, I think there has been a 5-year evolution where folks have consistently said that, at some point, you will not need reinforcement learning from human feedback to validate and test models. And I think the challenge of that logic is a couple of different things.
One, the spectrum of expertise—if you take language, multimodality, extreme expertise on things like computational biology—and then the fact that a lot of these are reasoning tasks, you do need human feedback on almost every different sort of agent you want to roll out. There’s a whole host of studies on this showing that pairing synthetic and human data together is stronger, but you do need human feedback in some form.
And so I think the nature of RLHF is changing. I think you’re moving more toward things like RL gyms, controlled environments, and simulations. I think you’re starting to see much more of the expert work being done by PhDs and master’s-level researchers. So it’s less what I’d call commodity “cat, dog, cat, dog” labeling.
But if you say tomorrow you’re going to train a model to figure out different evolutions in 17th-century French architecture in French, you are going to want RLHF to do that, to validate it. And I think you’re seeing that over and over: as the models move more and more into very specific areas, there is more and more RLHF needed for them.
That’s interesting. Maybe I’ll share my intuition, and then I’d be curious to hear what you’re seeing in your version of the ground truth. My intuition, my impression, is that we’re seeing greater and greater data efficiency.
And pardon me, I mean, RLHF was obviously very fashionable over the past 3 years. Maybe it went through peak fashion, if you will, and then we saw the rise of reinforcement fine-tuning mechanisms that are far more data-efficient and maybe even more human-time-efficient.
If you have to just build an RL environment, arguably that’s, per human hour involved, probably a lot more time-efficient than staffing out to some so-called developing-country folks to, as you say, do “cat, dog, cat, dog” supervised fine-tuning or some other RLHF-type mechanism.
Surely I’m projecting. My intuition is you’d see more data efficiency, not less, and therefore the amount of time, effort, and money expended on RLHF—or any sort of mechanism—even if we buy your assertion that we’re seeing hyper-parochialization of lots of different tasks and each of them is going to need artisanal annotation, surely there is a competing force, which is increasing data efficiency from algorithmic efficiencies like reinforcement fine-tuning. What are you seeing?
Yeah. People have been arguing that for 5 years, but I think at least what I’ve seen on the ground is that, given the accuracy that you want, if you think about a reasoning task that involves a several-step leap and you think about the risk of hallucinations, it is more useful to have human feedback involved in that in some form, all right?
And so I don’t think that means—if you think about it, in some ways, RLHF happens after all the pretraining compute cost—it’s a pretty small percentage of the total cost in training. And it is some of the most valuable feedback.
As you see more and more specific agents being trained for specific tasks, take legal services as an example. If you get a new legal-services dataset, which is interesting, and you want to train a model off of that, you’re going to want to see some sort of comparable equivalent, whether it’s an associate or an M&A lawyer equivalent, where you actually test if it works.
Now, is it possible that at some point, 10 to 15 years from now, you run out of things to train on? Possibly. But actually, if you take the number of languages and modalities, robotics is probably the next frontier of this in some ways. RL gyms, contact centers—there’s a lot.
We are, as a company, a full believer in—I talked about it on the enterprise side, too—that human-in-the-loop is going to be a feature, not a bug, for a long, long time. And I think the entire red herring of the enterprise, for example, is that autonomous agents will do all of this with no human in the loop. I actually think you’re going to need more and more humans at every step.
Alex, you’re saying that the level of intelligence of these agents, as we pass through AGI and get to ASI, is such that they’ll figure it all out as well as any human and replace that human in the loop. What’s your timing on that?
That was exactly my question, Peter. So my timeline, if I had to spitball—of course, this is not the predictions episode, so don’t hold me to it. Hold me to my predictions in the next episode—is approximately 2 to 3 years as a conservative outer bound for some element of recursive self-improvement, where we get an AI researcher that’s as good, if not stronger, than the human researchers for building ML models, as a conservative outer bound.
Now, 10 to 15 years? 2 to 3 years max. That’s the outer, outer edge. But I also believe Matt’s totally right that 2026 is going to be the year of recursive self-improvement, with capabilities growing at crazy exponential rates and corporations moving at a snail’s pace compared to what they could be doing.
And it’s all going to be stuck, bottlenecked, log-jammed, and it’s going to frustrate the hell out of Google and OpenAI.
And companies like Invisible are the lubricant that's going to actually get it from point A to point B. But that Clippy use case is a really good one. In our tests for contact centers, 80% of people massively prefer the AI. But the 20% who don't like it would rather torture the whole thing to death, make it better, or repeal the entire thing. There are probably 8 ways to fix that quickly.
Yeah, but it's not going to come from Google, and it's not going to come from OpenAI. It's going to involve data that isn't in the natural data set. If you told me 2 years ago that everyone in the world would know what RLHF stands for, and that there would be 3 people who are multibillionaires from building RLHF companies walking around saying, “That's not even a thing,” I would have laughed. Oh, wait—now it's not only a thing; it's massive in scale.
There'll be new terminology in 2026 for many of these other bottlenecks: the AI can do it, but for whatever reason, the bank isn't doing it, and the contact center isn't doing it. Those bottlenecks are going to be so lucrative for companies like Invisible to just plow through.
I can't answer the specific question of whether your workforce is going to involve the distributed workforce that you just described. What was it called, Alex? Or Matt? It's called Meridial. Meridial. Yeah, so there is a really healthy debate over whether Meridial is a key part of this, whether a network of even more agents is a key part of this, or whether 2026 is the transition year between the 2. It's going to be a really interesting footrace between those 2 different approaches.
But I think that's it. Dave, I think you put your finger on it. That really is what I'm asking, which I think is a distinct question from whether there's value in supervised fine-tuning or reinforcement learning with human feedback going forward. Of course there is. What I'm really asking is how much of that can come from AI bootstrapping it in the near future versus needing human inputs.
What I'm saying is, think about a balance between generalizability and hyper-specificity. I agree with you on generalizability. I don't actually think RLHF is important even now for that. But where it gets more complicated is when you want to train off specific tasks. So let's take the insurance-claim example that I mentioned earlier. You're going to generate a 10-page insurance claim, and you could apply this to any enterprise use case and many consumer use cases.
In that world, an LLM is producing an outcome and is fine-tuned off a specific company's data, but you need a way to actually say, at that point, does this produce an output comparable to what a human doing this task before was doing? When I mentioned custom benchmarks earlier, that's the process by which you do that. You actually do need human-equivalence testing. You need a human to provide a comparable data set and say, “This looks good,” or, “It doesn't.”
You just don't have precedent data to train that off of in any of these LLMs because the human input is not there. Now, again, that's going to keep going down to more and more specific tasks. If you take legal services, take it by language, take it by topic, take it by document type, there's human feedback required for all of that.
I don't want to put too fine a point on it, but I want to make sure that those in this episode who want to drink the bitter pill with the bitter glass of water for The Bitter Lesson are drinking it. I'm curious, Matt, to understand how you see this. Surely there's a wave of generalism that, over time, maybe we can finesse what the appropriate time scale is. It sounds like maybe your timelines are a little bit longer than mine, but would you at least agree with the premise that, over time, even specialized skills end up getting subsumed by generalist models? Or do you think that's just never going to happen? Will we always—or by “always,” I mean on time scales of 10 to 15 years, which is a pretty long time scale—just have generalist models that are always specially fine-tuned?
I don't think all expertise—all specialized expertise—is going to go away. Again, if you think about a lot of the information that specific experts have, there's no training data available for that. It's stuff that sits in people's heads; it's experience.
I'm aware of many of the narratives that human expertise becomes less important. Again, we're a company that actually thinks the human-touch elements become more and more important. Take sales, for example. Many of the best-selling patterns, and many of the people who have done that best—there is no information you can train on from what they do. They live in human interaction.
In a world where there are 500 companies selling email-based SDRs, I think human beings become more important in that world. I don't actually think specialization goes away. I think the shift is that expertise becomes more and more important in many different areas. I think the human loop stays really important.
But if you take a contact center—and Alex, I understand the theory of what you're saying—but we're 4 to 5 years into this, and if you look at the number of US contact centers that have migrated to using agents, it's a pretty small percentage.
Can I ask you the Jane Street question? It's really burning a hole in my pocket, too. It's really clear that stock picking is moving to AI at warp speed.
And the reason is that there are no barriers. You're just placing a trade that's already automated, so—
And that's the bellwether to me. It's a great benchmark. More money. More money.
Yeah, and also almost all of the volume on public-equities markets has long since been dominated by algos. So this happened decades ago. It started with rapid trading, so the quants were already there. Now that it's moving to fundamental analysis, it's the same mindset. That's one of the reasons it's taking off.
Like Peter said, you're making more money. Okay, let's just keep going. There's nobody who's saying, “But I'm going to lose my job.” It's like, “No, we'll just pay you more. Let's just go.” So it's a really interesting bellwether. But within that world, they're struggling because the data is so proprietary. Mhm.
And it's looking more and more likely that these self-improving, massive foundation models are going to get to superhuman IQ this year—this year being 2026. The context window is getting massive, and the recursive chain-of-thought reasoning is getting really good. So you can actually feed it data without having to retrain it and have it achieve the job.
If I take that mindset from Jane Street and move it over—now I'm a mechanic and I'm trying to fix a car, diagnose what's wrong with it, and I have audio and sensor data—great, easy use case. But am I going to put that data into the LLM API and transmit it to OpenAI, where they can accumulate it, and then, if they decide later they want to be a garage, they can have all my data? Or am I going to run some kind of walled-off model?
A garage mechanic's maybe not the best example. That's why I chose Jane Street, because they're never going to take their proprietary data and give it to OpenAI. But in the middle ground, you have banks, insurance companies, hospitals. How are they going to deal with this? It's easy now. Sometime in 2026, it becomes easy, but the data is proprietary. That's my only reason for having a competitive advantage. I don't want to give it over to the API.
Yeah, look, I think you're seeing that there are definitely sectors, many of which you just named—banking and health care—where people are deciding to keep their data on-premise, or they're using things like small language models for those sorts of reasons. I think you may continue to see that as a trend.
I think one mistake folks often make is that not all data is proprietary. Take the Jane Street case: maybe their trading data is proprietary, but their back-office forecasting data might not be, and their back-office finance data might not be. I think one thing is being clear about the data that you need to keep proprietary and around which you do want to take more security measures, and then what data you say, “Look, I'm going to be very careful as a company, but this is data that isn't as proprietary.”
I think that sort of balance is similar to what we discussed with contact centers. The idea of “I will not give anything to the LLM, but I'll keep it all in-house” doesn't make sense either. But I do think that's a paradigm you're seeing more and more.
Yeah. So I want to change tack a bit, if that's okay. I actually do agree that we'll automate, but I think we'll automate in a way that's different from this discussion. Let me give an example.
Let's say I'm Canon and I'm selling home printers. Right now, I have a bunch of people doing marketing, content development, and brand management; salespeople to sell to Best Buy and so on; online folks; post-purchase staff getting the customer to try to register the dang printer; repair-support and technical staff; and accounting folks in the company.
You could get all your job functions managed by AI, right? So you've got pockets of people doing different functions across the board.
If I was going to build an AI-native printer sales company, then I might think about having all of those things automated completely with AI. Then you're not human-centric, but function-centric across those. The printer could report when it's running out of ink, and you ship it a new thing. It tells you when there's a problem with it or a problem coming up, and you alert your repair staff, saying, “Hey, this guy, maybe we can upsell him a printer.”
You essentially automate all the functionality with AI, and you leave the human 90% out of the loop almost completely because you've automated the core functionality. What I'm seeing right now is what I used to call “radio over TV.” When you first had television, we took radio announcers and put them on TV to read radio scripts. We didn't adapt for the medium.
I think what I'm seeing right now is we're automating what human beings are doing at each of those functions, but surely, over time, we're going to automate the functional flow and then get rid of the human beings completely. AI-native, AI-first, right? Not to mention getting rid of the printers. Well, that's a separate question. I'm just using that example.
Who's going to be doing any of the printing? Let's leave that part aside just for the moment. I think you're absolutely right, Salim. This is where a young AI-native company reimagines an entire field and has zero legacy and zero friction in coming forward. The question, as Matt said at the beginning, is: Do they have the distribution?
This is where a large company—Canon, in this case—should actually be investing in entrepreneurs. One of the things you and I talk about a lot of times is, if I'm a large company and I don't know what to do, I would basically hold a competition and ask young AI entrepreneurs around the world to come forward: How would you disrupt my company? Give me a pitch. Then I would pick the best 5 of them and fund them.
I would say, “We're going to fund you to disrupt us, and then we're going to give you access to our data, to everything we have. Ultimately, we're going to buy you or buy a majority stake in you, and we're going to make you our new company.” This is the innovation on the edge, the displacement of the core, et cetera—whatever you want to call it.
You're a medium-sized or large-sized company. I'm not going to focus on the startup right now. What do you do in 2026? You're going to have to do something. You're going to have pressure from your board, from your shareholders, from Alex.
From just competition.
So you've got to do something. What I heard you say so far, Matt, is: Number 1, you've got to get clean data. You need to make sure you understand what your data situation is. Number 2, you should pick 2 or 3 areas—call them benchmarks—where you're going to run experiments. It's not a proposal or an idea. It's actually: Run it. Actually do it—run an experiment to see how it works.
Then pour money on the things that do work, and have an expanding, increasing circumference around the company's major revenue engines. How do you think about that? Walk us through a few more steps.
One of the things that has been a topic of conversation here is, given all the improvements in the models and what Salim was walking through about the potential to clean-sheet and design a company from scratch, why has that been so much harder? There was this MIT report that came out saying that 5% of enterprise models right now make it to production, right?
I think there's a starting question: Given all this tech excitement, why has that been so much harder? It's not the technical challenges that we talked about. It's the data and the focus on which priorities to look at.
I think the other 2 big ones are the organizational structure through which you pursue those initiatives. Particularly, the advice I give everyone is: Do not locate this in your technology organization. Take your best operator—your best ops person—give them an operational KPI, and track it to that. Make sure it's a really clear operational KPI.
We talked a bunch about contact centers. You should have an operational person there lead it around CSAT score, time per call, or whatever the core metrics are that you're looking at. That should be your guide. If you want to take something like inventory forecasting, you should do it around inventory days, stockouts, and all those kinds of metrics.
If you have a clear sense of which operational person is leading it, how they're marshalling resources around it, and you have a clear KPI, you're going to make progress if you focus on a couple of different things. I think the failure mode has been that you let a thousand flowers bloom, none of them have an operational metric, and you end up with a science-project dynamic.
Yes, exactly. That's exactly right. If you walk in, a thousand flowers bloom. You walk in and say, “I am going to give you a million genius-level people for free. Do something.” It fails.
It's like, “Here's a million people for free, and they're all geniuses.” It fails for that same reason. It's like, “I didn't think of an idea, so I said, ‘A thousand flowers, just go bloom.’ I couldn't think of anything, so maybe you will.” How's that going to work? I've seen that. You're exactly right. It's just so sad.
We go even further. We basically say: Not just take the operator and put them outside the organization and let them build something from scratch at the edge. Otherwise, you're getting encumbered by all the internal rules and bureaucracies, and that gets slowed down a huge amount. Then it fails for legacy reasons.
Yeah, it's not the company skunk works; it's the Apple Macintosh team. Apple is actually a master at this. What Apple will do is form a small team that's very disruptive. They'll put them at the edge of the company, keep them secret and stealth, and say to them, “Go disrupt another industry.” Whether it's watches, retail, or whatever.
At last count, I think they have 18 teams looking at different industries to think about. When they think it's ready to disrupt, they go into it and patiently iterate. The Apple Watch, for example.
This is the model I think we're going to see many other companies take on, where you do this. If you think of any operational company, the insights they have on all sorts of adjacent industries are incredible. It's very hard to disrupt their own industry because they're probably pretty optimized for it, unless you come with the AI startup, but they can really disrupt a lot of the edge cases and a lot of the industries around them. So I expect them to launch AI-native startups that go into adjacent industries and attack some of their neighbors.
Nice. We worked with SAIC Vantour and the U.S. Navy on building intelligence for an underwater drone swarm of unmanned underwater vehicles. Think of it this way: If you have a series of drones and enormous numbers of sensors on each of those drones, you need to understand the movement patterns of those different drones.
In each case, you see an object underwater. What do you do? Do you engage? Do you step back? Do you move with other drones? That whole movement pattern and decisioning for underwater unmanned vehicles is what we worked on: fine-tuning a model to do that, training it, and looking at all the movement-pattern data.
Again, this is one of those interesting things about drones: They are autonomous, and so thinking about how those movement patterns evolve in complex environments is very, very tricky to do. But you also have lots and lots of interesting sensor data to do that.
I think one that anchors more on the human decisioning side is SwissGear—like Swiss Army, the luggage brand. Similarly, I actually think this is, Peter, one that a lot of folks in the audience may relate to in some form. They had an enormous mix of different data tables around products, customers, et cetera, that they couldn't really bring together for inventory forecasting.
We used our data platform, Neuron, to bring together 750 tables really quickly and then optimize the forecasting to look at both minimizing stockouts and optimizing which inventory to hold. If you get inventory forecasting right, it's probably one of the major issues for most big and small businesses: You minimize lost revenue, and you make sure that you don't hold lots of excess inventory.
It's one of the hardest things to do, particularly if you have a 6- to 8-month order cycle time. And so that was something we partnered with them on, and I think it was a great outcome. We ended up expanding their overall inventory coverage by about 30% and basically 2x-ing the number of SKUs with a reliable prediction. And again, that was done in a couple of months.
All right. So later this week, my Moonshots mates and I are recording our 2026 predictions. We'll have Emad back, and we'll be talking each of us will provide two predictions for 2026. We'll have our top 10 from the Moonshots podcast. It's going to be fun. Uh it's going to be a battle. Uh we're going to ask our listeners to vote on which predictions they like best. I mean, of course they're all going to vote for Alex's, but hey. Uh Matt, uh talk to us about what you see coming in 2026.
Yeah, I think I'll call out a couple, and we've just done a bunch of research on our 2026 predictions. So I won't say all of them, but I'll call out a couple.
I think one of the first ones I would anchor on is multi-agent teams. I think one of the challenges—and it's inherent in a lot of what we've discussed here—is that if you're a large enterprise or medium-sized company implementing a use case, you won't necessarily have one decisioning agent that does everything. You'll train task-specific agents for individual tasks, usually orchestrated by an LLM. What that allows you to do is pinpoint the accuracy on those specific tasks, and then use the broader logic set of the LLM to make sure they all work together properly.
I think that's been an architecture that's been discussed pretty broadly for a while, but I think we're just starting to see the green shoots of more and more folks having success with that. Contact centers are a good example, so I think that's a big one that I would call out.
I think the second one I'll call out is the multimodal leap. More and more, video, images, and audio are going to become a bigger and bigger part of how people engage with these models. Audio is probably one of the most interesting. I do think the way you'll be able to speak to them, interact with them, and visualize them is going to be a really interesting moment for 2026. And I don't think that will all be text-based like it has been historically.
Mhm. And then maybe one other thing to talk about. No, I was going to ask Alex for feedback. Go ahead, but finish up, Matt.
Yeah, so the third one I'll call out, because we've talked about it a couple of times on this episode, is what we call either the mirror world or RL gyms. I don't actually think that's a well-understood concept for many folks in the audience, but think of that as creating simulated environments or digital twins for tasks you might want to test. Maybe that's a coding environment; maybe that's a contact center, as we've used that a couple of times. But it allows you to simulate a series of function calls, tasks, or environments by which, if you're going to train a model or a task, you can test how it's going to work in a manufacturing environment before you roll it out to your actual physical world. And I think that's more and more an interesting topic for both model builders and the enterprise.
I want to go around and maybe ask some final questions of Matt. Alex, do you want to kick us off?
Yeah. I think the most interesting crux of what we're discussing here is: What is the future of human expertise? For that matter, does human expertise have a future? And assuming it does, what's the half-life of the value of human expertise?
To put that question to Matt, what do you think, of all the forms of human expertise, all of the labor categories and job roles that exist in the economy today, will be the last 3 of those job roles or forms of expertise to disappear or ultimately succumb to AI? What are the last 3 to survive?
Last expert standing. Okay.
That's right. I'll go back to where I started the episode. I think a lot of the commentary on mass shifts ignores the actual function of jobs in society today. So let's take sectors, for example: oil and gas. A lot of the functional expertise—geoscience, if you look at seismic engineers, people on oil and gas sites drilling—that is a human function. Real estate is another example. You can go down a whole list of different areas.
I think there are sectors where you're going to see more disruption in the near term. I call out a couple of them: BPOs and legal services. I think media is a fast-changing area. But I'm also not exactly sure that those lead to negative—meaning, have negative employment consequences.
If you take media, it's a really interesting one. 5, 6, 8 years ago, I think media as a category really struggled in a lot of ways, with paid media as an example. And you've actually now seen, in the last couple of years, Substack, Medium, and all these blogs become much more interesting. You have way more media entrepreneurs. And so you've changed the function of society, and where the money is coming from changes, but it has not changed total employment.
And look, I understand a lot of the skepticism that says AI is going to radically change everything, but I think if you look at American society for the last 100 years, it's something like 25% of every high school class goes into a field that did not exist when they were in high school. And the reason that persists is people go into the working world understanding the tools they have, thinking about what they can create from that.
One of my favorite statistics, which I saw The Wall Street Journal report a couple of weeks ago, is that 20% of U.S. employment right now is in digital ecosystem jobs. And something like 9% of U.S. citizens are full-time social media influencers. It's mind-boggling to me.
But again, this is the changing nature of work, and I think that pattern will persist. I think the core of what will change is that the process of looking up information across multiple systems and documents is going to become less valuable. But I think all the jobs that involve human interaction and physical work—
Optimistic.
I was actually on another panel with someone who runs a recruiting company. They were saying that job profile, I think, will 2x, 3x, 4x over the next couple of years. And so that will have pretty interesting implications for the education system and everything else. But I don't think it will all be displacement; I think we'll see an evolution.
I mean, the humanoid robot electrician and plumber. Alex, very quickly, what are your 3 last-standing human roles here? Or your last 3 standing?
So I'll present multiple competing hypotheses.
Briefly.
One hypothesis is that it's the politician, because they have to make the laws. Another hypothesis is that it's the greatest intellects: the physicists or mathematicians. Even though, as we talk on the pod, math and the sciences are all getting solved, on the one hand, they're still perhaps—to the extent that that represents the culmination of human intellectual accomplishment—maybe the greatest intellects will be the last to be automated.
There's another school of thought that says, "No, it's the roles that involve the greatest need for human authenticity." Because even though it's not actually a capabilities question, people nonetheless demand human contact, or something to that effect. And so it's going to be the highest-touch job roles, where people just want to know that there's a human counterparty on the other side of the interaction. So that's a set of 3 hypotheses.
Tastemakers will dominate. That's authenticity. Bucket number 3. Saylor said that word for word, actually. Yeah, we had that enjoyable sunset conversation on his boat. Salim, do you want to go next with a closing question for Matt?
I think you covered some of the industries that are going after it. You guys have done some government work. Where in government functionality do you see the biggest opportunity for AI automation, efficiency, et cetera?
Yeah, everywhere. Look, I actually think this could be one of the really positive trends for society. I saw a study recently that AI-assisted permitting could cut energy and data center project implementation timelines by 50%.
Think about housing. One of the biggest challenges right now for housing development in the U.S. is NIMBY regulations and how complex it is to build housing because of the myriad of different regulations and zoning constraints by location. The OECD came out with a report that AI could shrink public-sector process-cycle timelines by 70% on licensing, benefits approvals, and compliance.
To me, the simplest thing that AI can do is project management and timelines related to all spending and infrastructure deployments.
This would be a really positive thing for society, in my mind. Amazing. Good question, Salim. Dave, why don't you close us out on the questions here?
I've got so many, but I'll pick the best. First, Matt, how many hours of video footage will there be of you 1 year from today compared to 1 year ago? Because I know we saw each other in Riyadh a few weeks ago.
And I know that you are the thought leader in this whole bottleneck of AI getting into the enterprise. It feels like what we're doing right now. The footage of you that's out there right now is all this CNBC, Bloomberg-type, 5-minute format. But here we're getting your real thoughts. It's just so much better. How many hours can we count on 1 year from today?
Well, look, I think as of 12 months ago, I had done almost no interviews of any kind. So this job has been fun on that front. What I enjoy about the podcast format is that it does allow you to talk about some more complex topics. I particularly find podcasts like this really interesting. So hopefully many more in the year to come.
Well, I'm hoping for at least a 10X on that. And then my follow-up question is: the avatar version of you that's also out there talking—is that a 2026 thing, do you think, or when?
Yeah, it probably happens in 2026. I don't think it would be that hard to train an avatar off of my public statements. So I think that'll be interesting. We are actually working in the sports space on the topic of avatar training. And I think it is an interesting space where you could imagine a lot of different areas where, rather than a chatbot interaction, people want to speak to people they know via an avatar. I actually think that will become a more natural part of society and a pretty interesting one, actually.
I totally agree. I just think the timeline could be as soon as 2 months, as far as I'm concerned.
What makes you think it's not an avatar we're speaking to right now, Dave?
That's a good question.
That seems very human, actually. I don't know. The best ones are.
The orbs behind you kind of give it away.
Yeah, they are pretty strange. That's not real.
Matt, where do people find you? Where do people find Invisible? Who should go to Invisible to check out what you do and how you do it?
Sure. So we have 7 offices now: New York, San Francisco, Austin, Texas, D.C., London, Poland, and Paris. I'm probably the easiest to find. We have an office right off Union Square, which is where I'm at least half the time when I'm not on the road.
In terms of who should come to us, and from the listener base in particular, any mid-cap or enterprise company that knows there is potential in its business, that knows AI can transform it in a positive way, and is struggling to bring all the pieces together. I think that is the main thing I would say. There is no doubt: the technology Alex is asking about has made an enormous step change over the last couple of years. The hard thing is actually the change management, the operationalization, the metric tracking, and the evaluation.
It's kind of bringing together—I think it's the difference between the... Our founder, Francis, has an idea: you have all the components to build a cake, but you don't have a cake. What we do is we actually bake the cake in the end. We build you something that works. We make AI work, and we use all the modern tools to do that.
Amazing. And the website?
invisible.tech.ai.
All right. Thank you, Matt. Salim, Dave, AWG, I'm going to see you guys in a couple of days for our 2026 predictions. Make them brilliant. It's going to be fun. All right.
To benchmark for tracking benchmarks.
All right. No, that's not the one I'm going to talk about. Okay.
All right, guys. Have a great day.