Matt Fitzpatrick:谁将赢得数据标注竞赛,以及为何 Al 需要前线部署工程师
- 核心判断:企业采用 AI 要走“10年,而不是2年”。过去两年,模型在公开基准上的表现提升了40–60%,KPMG 称60%的消费者每周使用 genAI;但 MIT 发现,只有5%的 genAI 部署以任何形式真正运转,Gartner 预计到2027年40%的企业项目会被取消。Fitzpatrick 的诊断是:瓶颈在数据基础设施、工作流重构、责任归属,以及“最重要的,信任”——银行会先按模型风险管理的方式验证 genAI,其他企业则会在此后五到六年跟进。
- “仅靠开箱即用的 SaaS 绝对做不到这件事。”Invisible 不收前期费用——免费提供8周解决方案冲刺,直到用户验收测试通过才收费,前线部署工程师也完全免费(450人、8个办公室)。对 SaaS 投资者而言,结构性结论是:“一旦你在 SaaS 场景下不得不引入 FDE,经济模型会瞬间崩溃。”Fitzpatrick 认为,开箱即用的软件“某种程度上一直都是谎言”——终局是超个性化软件,而不是一个个软件盒子。
- 内部自建正在败下阵来:MIT 数据显示,外部驱动的构建效果是内部构建的2x。一家电商零售商就是典型案例:它花2500万美元做退货客服 agent,自建评估体系,以通话解决速度和情绪作为指标,结果一个虚构“退款200万美元”的回答也能拿满分;几个月后,公司关停项目,退回确定性流程。Fitzpatrick 的解法是:选择3到4个项目,由运营负责人牵头——“不要把它放在技术部门”——并按实际效果付费。
- 行业最大的误称是:合成数据不会取代人类反馈。合成数据适用于数学等具有基础真值的领域;跨45种语言、多步骤推理和多模态场景仍处于“第一局”,而达到法律行业标准的数据掌握在大型律所内部,不在公开语料库中。与 ML 不同,genAI 需要人类参与统计验证——“未来几十年,你都需要人在环路中。”
- 数据标注行业将洗牌为3–5家,而非赢家通吃。集中度是结构性的,因为真正构建 LLM 的公司并不多;但“人们愿意为好数据付费”,因为一周的坏数据就能烧掉巨量算力,所谓完全不在乎价格是“夸张了”。护城河是 Helmer 所说的机构记忆——一条数字化生产线,能在24小时内调度2.6万名筛选过的专家——而且他坚持,那些大数字是收入,不是 GMV。
- 资本策略已经反转:Invisible 成立9年来只募集了700万美元一级资本;如今已募得1.3亿美元,今年不会盈利——“这是有史以来最适合增长的环境……我希望我们永远不要进入收获阶段。”下一 frontier 是物理世界数据:FDE 把 Starlink 终端部署到农场,为畜群安全训练计算机视觉模型;再往后是机器人。
- 对 agent 的怀疑可以量化:Salesforce AI 研究显示,开箱即用的 agent 单轮准确率为58%,多轮准确率为33%;AWS 报告称,70%的“agent”其实是传统脚本。若管理一支4亿美元基金,他的逆向配置是:不要只看开箱即用的 agent,而要押注直接服务客户需求的 AI 原生企业——他认为,YC 最近一届正是靠这一路径,收入做到了此前任何一届的2x。
1. 企业采用 AI 仍然滞后
- Fitzpatrick 同时站在市场两端:Invisible 通过 RLHF 为所有大型语言模型训练数据,也在自己的平台上部署企业用例。他对过去两年“认知失调”的概括是:模型在公开基准上的表现提升了40–60%,消费者采用呈指数增长(KPMG:60%的消费者每周使用 genAI),但 MIT 报告称只有5%的 genAI 部署以任何形式真正运转,Gartner 预计到2027年40%的企业项目会被取消。
- 问题出在模型之外的一切:“支撑这些模型的数据基础设施、工作流重构……以及最重要的,信任。还要有可观测性。”基于自己在银行业构建信贷模型10年的经历,他预计企业会复制银行所称的模型风险管理:测试、训练、验证——“整个流程还处在第一局。”
- 他的时间判断是:“这要走10年,而不是2年。”银行和医疗行业会先完成测试与验证阶段,“之后其他企业会在接下来的五到六年里跟上。”
- Stebbings 亲身经历进一步说明了这一点:他在全球最大银行之一演讲后给团队发消息说“他们完蛋了”,因为 CTO 以数据、安全和权限为由,笑谈他的现成工具。他也承认:“你刚才说的这些,我都听进去了。”
2. 内部自建 AI 项目失败
- 据 Fitzpatrick 转述,MIT 报告中一个被埋没的数据是:外部驱动的构建效果是内部团队构建的2x。他观察到过去10年的模式是:企业先买15个应用,云计算兴起后再套上定制封装,而“genAI 把这一切放大了5倍”——内部团队拿到巨额预算,却没有供应商必须遵守的交付物、时间表、ROI 和里程碑纪律。
- 他谨慎地描述人才问题:“真正知道如何把这件事做好的人并不多。”而这些人才都在 AI 初创公司和大科技公司里,因此企业若从第一性原理自行推演,确实面临真实风险。
- 他见过最典型的案例是一家电商零售商:花2500万美元构建退货 agent,并自建评估体系,把通话解决速度和情绪混在一起衡量。“问题在于,如果 agent 产生幻觉,说‘这里有200万美元’,它确实能很快解决问题,客户也会很开心。”几个月后,公司关停项目,退回确定性流程——“这完全不让我意外。”
- 他预计未来两年的修正来自 CFO 职能:设定 ROI、指标和里程碑。具体打法是,挑出3到4件真正推动业务的事,安排最好的4名运营负责人,“不要把它放在技术部门”,并把付款与结果挂钩。负责人“不需要技术能力特别强”,只要具备管理任何供应商时都需要的同一套肌肉记忆。
3. 前线部署工程胜出
- 买方的现实是:每周有250家供应商上门推销,听起来都差不多——有客户甚至一开会就问:“你和这周向我推销过的另外250个人有什么不同?”而且很多产品根本不好用:Salesforce AI 研究测试开箱即用的 agent,单轮工作流准确率只有58%,多轮只有33%——“这意味着它们其实不能工作。”
- Invisible 的回答是:“我们其实什么都不卖。”公司见客户时会说,免费做8周,证明技术确实可行。当被问到能否在没有强力前线部署机制的情况下服务企业时,他回答:“我不认为可以……仅靠开箱即用的 SaaS 绝对做不到这件事。”公司已进一步押注这一模式:450人分布在8个办公室,FDE 不收费,区别于竞争对手:“我少花钱雇销售,多花钱雇前线部署工程师。我的算式就这么简单。”
- 他对 FDE 的定义很关键:做好 FDE,就是在大约3个月内“执行一个非常具体的工作流构建”,既不是解决方案工程,也不是 Accenture 式的3年期项目。对创业者而言,判断标准很简单:知识库(公开披露文件、医疗参考资料)不需要 FDE;“如果你想改变工作流,就需要 FDE”,而采用和工作流嵌入才是最难的部分。
- 他最尊敬的公司是 Palantir:“他们比科技市场其他人早10年意识到,前线部署工程定制会很重要……这是一次非常反文化的跃迁”,因为当时没有人愿意涉足技术服务。
4. AI 软件需要定制化
- 创始人 Francis 的创办原则是:“如果什么都有一个应用,为什么什么都不好用?”过去20年的 Accenture 模式是:买50个应用,两年花2亿美元把它们层层拼接起来,最后得到“没有可用数据、没有连接关系”的沉积层。Duck Creek 等保险科技公司则借助势头和推动其发展的系统集成商,取得了很好的成绩。
- 他对 SaaS 投资者的异端观点是:看看大型上市软件公司收入中究竟有多少来自服务。“开箱即用的软件某种程度上一直都是谎言……只是把它包装起来。”genAI 进一步打破这一模式,因为每个客户构建的东西都高度具体:“它不是一个盒子。”付款发生在用户验收测试时,也就是项目开始两到3个月后——“机器学习进入企业已经很久了,一直就是这样的路径。”
- Lifespan MD 是一个验证案例:这家医疗 concierge 服务商的数据分散在 EHR、CRM、ERP 和各种笔记中。Invisible 的5个平台——负责数据的 Neuron、构建 agent 的 Axon、构建流程的 Atomic、专家市场,以及负责评估的 Synapse——在2到3个月内完成统一,“Accenture 要花两年”。随后再叠加对话式查询,例如“哪些人用过 peptides,年龄在36到50岁之间,结果如何”,以及定制排班 agent。他的判断是:“你会从开箱即用的 SaaS,走向利用单个客户特定数据的高度超个性化。”
5. 数据标注高度专业化
- 对于外界把 Invisible 与人才市场归为一类,他重新定义了业务:“你必须能在24小时内找到一名牛津大学天体物理学博士,把他放进一条数字化生产线,4天后生成完美且经过统计验证的数据,再与另一家的数据正面对比。”Invisible 每年接触130万名专家,最终保留2.6万名筛选过、必须在24小时内开始工作并产出完美数据的专家。
- 用 Hamilton Helmer 的话说,护城河是机构记忆。他最喜欢的例子是丰田生产体系:丰田可以公开描述这套体系,但没有人能复制。“它就是一条数字化生产线,与汽车工厂没有区别”,同时还沉淀了5年数据,记录谁擅长什么任务。
- 这个行业的变化很大:5年前还是在 Google Sheets 上运行的“猫狗低端商品标注”;如今则是在一天内找到“会说法语的17世纪法国建筑专业建筑师”来完成验证。专业化和拆分成极其细分的供给池确实存在——Stebbings 引述的那位董事会成员没料到这一点,Fitzpatrick 的回答是:“绝对如此。”
- 对于定价,他说:“我把我们的业务看成 Uber。”这是一个价格发现市场,费率取决于下雨、地点等市场环境。关键区别在于:“你可能给一个非常差的专家支付过高费用,那会彻底浪费所有人的时间。”真正的价值在于知道150美元的专家与130美元的专家有什么差别。供给有限也不是问题:“所需的专业能力每月变化太大,即使你把供给固定下来,3个月后也会变。”
6. 数据市场仍将保持集中
- 对于“两家公司贡献超过50%收入”的格局,他并不反对:“真正构建 LLM 的公司并不多,所以整个空间按定义就会集中。”他的反驳是业务正在多元化:2024年业务明显偏重 AI 训练(人才市场一侧占据“相当大的比例”),但 Invisible 已在过去45天确认了12笔企业交易。
- 谈判的僵局最终取决于质量,而不是议价权:“人们愿意为好数据付费……一周的坏数据会烧掉大量算力。”同台测试很常见:一款多模态音频模型出现后,“我们会在当周与另一家正面对决,最后可能赢,也可能输。”但董事会成员关于价格严重不敏感的说法遭到反驳:“我认为那是夸张……所有玩家之间其实存在相当标准的价格区间。”
- 他对行业洗牌的判断是:“我不认为答案是一家……大多数市场最终都会留下3、4或5家。我甚至不认为会是两家。”不同公司会按任务分工,分别专注编码、专业任务和博士级专家。在企业市场,历史上“一直是 Palantir,其他公司不多”,因此替代者才会引发如此大的兴奋。
- 对于 Stebbings 质疑其收入与 GMV 的区分——他此前几期节目因称其为收入而“挨了不少批评”——Fitzpatrick 回答:“我认为那就是收入。”区别在于 Airbnb 收取统一费率,而这里“取决于项目和专家类型,差异极大”,不存在相对于订单金额的统一费率。
7. 人类反馈仍不可或缺
- 在他看来,行业最大的误称是认为合成数据会在两到3年内接管一切。“从第一性原理看,这其实没有太多道理。”合成数据适用于数学等具有基础真值的领域;但跨音频、视频和45种语言的多步骤推理——例如“带南方口音的印地语、法语和英语计算生物学”——仍是一个排列空间,“我们还处在第一局”。他更明确地说:“人类反馈在未来10年都会很重要。我对此有强烈信念。”
- 法律是他反复举的例子:“世界上大量法律数据都在大型律所手里,甚至不在公开领域”,而公开语料库“多年来已经商品化”。更深层的区别来自他在 QuantumBlack 工作时的经验:ML 模型无需人类干预就能回测到统计有效性;但在 genAI 领域,“未来几十年,你都需要人在环路中。”
- 对于 Stebbings 嘲讽 Gemini 3 和 Opus 4.5 的基准不断变化,他认为模型“毫无疑问”在进步,但它们正转向高度具体的任务,而“按定义就不存在公开基准”——比如模型构建 LBO 模型的能力,或一份面向某家 PE 公司的 IC 备忘录是否达到“99%精确率”。因此投资者应注意:基准进步“几乎与企业采用正交”,后者取决于具体任务上的信任和精度,而不是泛化能力。专业能力会复利增长;他以可能与 SAIC、Vantor 和美国海军合作的案例为例:为水下无人机集群微调模型,让它们执行反应、撤退、报警和接战,正是靠逻辑积累出的转换成本。
- 对初级人才被掏空,他并不担心:5到10年的采用曲线给了行业足够的反应时间,而应届毕业生“是采用率最高的人群之一”,他反而在加大招聘。他的类比是:会计从“计算尺走向 Excel”,员工数量并没有下降——这就是 Jevons 悖论,“实际上会计工作更多了”,而且每个 FP&A 职能的规模“可能都比25年前更大”。
8. 资本推动物理世界扩张
- 资本策略发生了反转:Invisible 成立9年来总共只募集了700万美元一级资本;如今已募得1.3亿美元(最初公布的数字是1亿美元),今年不会盈利。他的逻辑是:“你可以收获资本,也可以投资资本……我认为这是有史以来最适合增长的环境……我希望我们永远不要进入收获阶段。”他经常思考、也最害怕的决定,是是否要向500亿至1000亿美元规模的超大规模公司迈进——这意味着要在每个客户身上投入更多。对于 Mercor 募资20亿美元,他并不羡慕:自己的资金投向企业业务和核心软件平台,“和其他人的重点有些不同。”
- 他目前还没有投资、但希望投资的方向是物理世界交互。一个案例是服务美国最大的农业集团之一,解决畜群安全问题:FDE 走进农场,“把 Starlink 终端放进这些农场,并构建定制计算机视觉模型”,判断何时派兽医。下一站是油井平台;机器人“需要更长时间,但一旦奏效会非常有意思”,而且需要针对具体任务的机器人,而不是通用机器人。
- 如果给他一支4亿美元的基金,他的逆向配置是:模型层能带来回报,agent 层较为复杂,应用层“也很棘手”。对社会而言,真正有意思的问题是:“围绕 AI 创立的新公司,能否比大公司学会采用 AI 更快地获得分发。”他的押注是直接服务客户需求的 AI 原生企业,包括税务和会计领域的 genAI 原生服务,以及物理世界运营商。他认为 YC 最近一届的收入达到了“此前任何一届的2x”。
- 对于70–80%的软件毛利率是否已经成为过去,他首先反问:“首先,70–80%的软件毛利率真的存在过吗?”上市软件公司的估值倍数在两年内从20x降到10x,因为盈利能力暴露出增长放缓;与此同时,“那些一体化单元会非常、非常赚钱”——客户获取更快、构建更快,也没有软件盒子的黏性束缚。
9. 信任胜过虚张声势
- 他接手 Invisible 时,“如果你看遍整个公开互联网,我认为只有一篇文章提到这家公司。”如今他70%的时间都在出差,遵循 Marc Andreessen 的一个观点:“当私下叙事与公开叙事出现分歧时,那就是风险或机会。”比如,假设有人声称一个开箱即用的 agent 什么都能做,而事实并非如此,这就会给竞争对手留下机会。Stebbings 的反问值得保留:“这不就是我们的行业吗?……我们的工作就是先卖出去,之后再交付。”Fitzpatrick 的回答是:“我希望与我们合作的公司知道,只要我说这能奏效,它就会奏效。你只有一次机会做到这一点。”
- 非确定性让“先吹出来再说”变得风险更高:合同签的是“交付50个 agent”,但接下来要问的是,究竟能不能交付这些 agent?它们能不能工作?他引用当天 AWS 的一份报告:70%的“agent”其实是传统脚本编写和自动化——“这就是为什么我根本不把自己定义为 agent 公司……agent 只是工具箱中的一种工具。”
- 他也坦言,12年前在 McKinsey 工作时,行业还叫“数据分析”,自己曾在不知道具体要构建什么的情况下设定愿景,因为相信碎片化数据让企业决策陷入失效——美国70%的软件已经使用超过20年,银行平均93%的技术成本用于维护(Stebbings 补充:有一家银行仅 KYC 就有6500人)。反直觉的经验是:“我没有虚构它……我的整个做法是说,我认为这会奏效,这是我的理由——让我们把它做出来。”人们会信任这种表达;而“一个能解决你所有问题的开箱即用 AI”只会触发怀疑。
10. 战略会变,招聘很关键
- 他最常思考的主题是:“只要找到优秀的人,其他一切都会随之而来。”这不只是招聘,还包括“招聘、留住并让优秀人才进化”,不要拘泥于岗位:“真正优秀的人会承担5到6种不同角色……要招全能型选手。”他的体育类比可能指向 Nick Saban:“他不是靠流程打造 Alabama 橄榄球队的,而是靠招募全国最好的橄榄球运动员。”
- Stebbings 以 Revolut 为例提出反驳:Nik 的理念是,胜负取决于文化,“有边界的残酷能驱动人类”。Fitzpatrick 给出了一个值得注意的让步:“不,我认为这其实是对的”,前提是要规模化一个一致的商业模式。但他补充说:“我们很多工作本质上是研究和探索……这既是研究文化,也是执行文化。”当然,交付和运营“相当多时候都处于战时状态”,而优秀工程师应拿出30%的时间做新项目。
- Invisible 不再把远程办公作为默认选项:公司连续9年完全远程,如今已在纽约、旧金山(原 Pinterest 办公空间)、伦敦、巴黎、波兰、华盛顿特区和奥斯汀设立办公室。生产力呈“指数级”提升;他今年把工程团队扩大了3倍,并发现“绝大多数人都希望线下办公”,尤其是入职时间较短的员工,但公司没有强制出勤。他的平衡方式是:每周7天待命与是否共址是两回事,而每周6天到办公室“过头了,会失去优秀人才”。
- 他放弃了两种管理信念。第一是中央控制:“有点像谬误”——应当压平层级,让一线团队拥有自主权,同时提供一致的工具,就像军队一样。第二是战略本身:“至少在 AI 世界里,战略是一个有些被高估的概念……整个世界每3个月就会变化一次。”应保留核心信念,同时保持30–40%的持续迭代;“现在做5年战略规划没有用”,5年期规划适合文化和机构记忆。最后,他以一组数字表达乐观:AI 数据中心只占数据中心耗电的0.25–0.5%,而数据中心占全球用电的1%(空调占14–20%);美国医疗人均支出为1.4万美元,是德国或加拿大的2.5–3x,每年有25万人死于可避免的医疗错误(Johns Hopkins),这些都是 AI 可以解决的问题;而教育最让他兴奋:Invisible 会根据认知能力和技能评估“大量”从未上过大学的候选人。
MIT just released this report that 5% of GenAI deployments are working in any form. You’ve seen Gartner saying 40% of enterprise projects will likely be canceled by 2027. I think the reason for that is externally driven builds are 2x as effective as internal team builds. I don’t think that discipline exists in the same way in internal builds.
They spent $25 million building an agent, and what ended up happening was, a couple of months later, they shut it down and moved back to a deterministic flow. We don’t actually sell anything. We meet a customer and say, “We will do it for free for 8 weeks and prove to you that it works.” The minute you had to bring in FDEs in a SaaS context, your economics broke instantly, right?
Are there any other big misnomers that you think are pronounced in the industry?
Look, I think the biggest one is just the view that synthetic data will take over and you will not need human feedback. It’s interesting from first principles that actually doesn’t make very much sense if you think through it.
In the AI world, at least, strategy is a somewhat overrated concept. What I mean by that is: Is it ready to go?
Matt, I am so excited for this, dude. I think Invisible is one of the most incredible, but also—I’m sorry to say this—one of the most under-discussed businesses when I look at the incredible achievements that you’ve had over the last few years. Thank you so much for joining me.
Thank you for having me. I really enjoy the show.
1. Career Journey and Leadership
Can you just talk to me about how a 10-year McKinsey veteran becomes CEO of one of the fastest-growing data companies in tech? How does that transition happen?
Yeah. I would say my McKinsey journey was nontraditional. I spent 12 years there. I was a senior partner, and I led a group called QuantumBlack Labs, which is the firm’s global tech development group.
About 10 years ago, McKinsey actually started hiring engineers, and I was a big part of this in a pretty big way. I think we went from—when I started, we had about 100 engineers total in the firm. By the time I left, we had 7,000. I oversaw about a fifth of that group, all the application development, all of the data warehouse infrastructure, and all of the GenAI builds globally.
That journey was really interesting. Over the course of it, I spent a variety of my time competing with other large enterprise AI businesses, and I got to know the founder, Francis, really well. About 3 or 4 years ago, we actually met totally unrelated to work, in a social context. It was basically a forum called Dialogue. I don’t know if you’ve heard it, but you basically talk about different ideas.
I keep getting invited to this. It’s in Hawaii, though.
It’s in many different locations.
Far.
It’s actually a great—I really enjoy it because you actually don’t talk about work at all. You’re not allowed to talk about your job. You spend time talking about history, politics, and technology.
What does everyone from San Francisco do?
They don’t talk about it for 2 days, which is—
A silent retreat.
Exactly. Exactly. But I actually think it’s one of the few events I’ve been to where people are not talking their own book. They’re not trying to convince you of anything, and you just really—I’ve made a bunch of really good adult friendships out of that.
Francis and I got to know each other from that 4 years ago. There had been another CEO in the 2 years before I joined who was actually based in Australia, interestingly. When the business got to a certain scale, it was just time to have a US-based CEO who could help take the business to the next level.
Francis approached me and pretty directly said, “Do you want to be our CEO?” That was kind of how it happened.
Was it a no-brainer?
Look, I think when you walk away from a really stable job that you really enjoy, that’s always difficult. McKinsey, actually, the sliver of McKinsey that I was doing, I found to be one of the most intellectual day-to-day jobs ever. I was working with all of the Fortune 1000 on every different AI topic daily, and particularly in the early machine-learning days, around 10 years ago, I think we built some really interesting stuff.
But, yeah, it was kind of a no-brainer in some ways. When you think about it, I think this is the most interesting time to run a company on a topic that has probably never existed in our lifetime. Maybe the 2000s, but to run a company in AI right now is fascinating.
There’s the rate at which you can build, the people you can recruit, and the interest of customers in this topic. I felt like I’d spent 10 years learning one topic, and now I had a chance to run a business and build it the way I wanted to build it on that topic. That’s just something you can’t pass up.
Even though I walked away from a fair amount, I’m much more excited about building something for the next 2 decades out of this.
When we think about decision-making frameworks, I always have one, which is: Find someone who you respect and admire. For me, it’s Pat Grady, who’s the head of Sequoia. I’ve known him for about 10 years. He’s a great father, investor, and husband—3 things that I care about.
Whenever I have a tough decision, I’m like, “What would Pat do?” Most of the time, I get to the answer by asking that question in that framework. If I were to ask you, what do you ask yourself? How do you find direction when struggling with a decision?
I’m not a particularly materialistic person. When I was coming out of college, for example, everyone was focused on going into large finance jobs, which, at that time—pre-financial crisis, obviously—were where a lot of that was. I think a lot of what I think about is doing work day-to-day that I really enjoy with people I really enjoy, and then building something.
I really enjoyed the decade I spent building at McKinsey. I think that was an incredibly interesting experience, to stand up something of that scale within an existing institution.
I do think about it. I read a ton. I read a ton about everything from military history to current entrepreneurs to enterprise executives I really admire. I have a small group of people whose opinions I ask for pretty regularly.
Probably the most telling piece of advice was that my girlfriend and my main mentor, when I asked them, within 2 minutes, were like, “Absolutely do this.” My main mentor is a guy named Sesh Khanna, who had been a senior partner at McKinsey for a long time and is on the board of a whole variety of different companies today.
I remember we got lunch. I walked him through the opportunity and said, “Listen, it’s a big risk.” He goes, “The only risk is if you don’t take this, and the amount of regret you’ll have from not giving it a go.”
I totally agree with that one. I was once given advice that whatever you think you should do, hold that close, and then let your girlfriend tell you what you should do.
And that’s why you still have a relationship. It’s a great piece of advice. That was from someone who’s been married for 40 years, and so it’s worked well for him.
We were chatting before, and I said, “Listen, where do we have to go?” I always think that the best conversations are led by passion. The first one that you said was that there’s a gap, or a chasm, between model performance and adoption. When we break that down, can you explain to me what you meant by that and how we see that in action?
Yeah. Let me set the context. I’ll go into more detail later, but Invisible is an interesting business in that we both train all the large language models with reinforcement learning from human feedback, and we are, at our core, a modular software platform where, in an enterprise context, we deploy all different enterprise use cases.
I think the cognitive dissonance that has occurred over the last couple of years is that model performance has increased exponentially. I don’t think anyone would doubt that. If you look at all the public benchmarks, models have increased 40% to 60% in performance over the last 2 years, and consumer adoption has also been exponential.
KPMG just released that 60% of consumers use AI weekly now, but the enterprise has not. MIT just released this report that 5% of AI deployments are working in any form. You’ve seen Gartner saying 40% of enterprise projects will likely be canceled by 2027.
I think the reason for that is deployment in the enterprise is a lot more than just the models themselves. It’s the data infrastructure to support those models, the redesign of workflows, and the process of figuring out which operational leader takes accountability for that. Most importantly, it’s trust and observability—all the things that I spent a decade building, things like credit models in banking.
In those cases, you need to go through model risk management, testing, training, and validation. I think that whole process is in the first inning in the enterprise. I think it’s going to take a decade, not 2 years.
I do think that is the core mission that we think a lot about. I actually think the evolution of AI deployment will be much like what the model builders have done over the last couple of years. You’ll see banks and healthcare firms start to do the same sort of testing and validation over this period, and then the rest of the enterprise will follow over the next 5 or 6 years after that. That’s the journey that we’re focused on.
2. The Single Biggest Barriers to Enterprises Adopting AI
So I was speaking at one of the largest banks in the world. It's an absolute joke that they get a university dropout like me to speak at their largest retreats. I find it very fun.
I left and messaged the team, and I just said, “Oh my God, they’re toast.” They’re toast because I told them about the amazing tool they should implement internally. The CTO laughed at me. He was like, “Dude, there’s no way that we can ever adopt your off-the-shelf search engine optimization for your LLM tool because of data, because of security, because of permissions.”
I was like, “Wow, everything that you just said there, I listened to.” But that was once you got in the door. Are enterprises even open for business? You see Goldman Sachs developing a huge amount of its own tools. Are they open for AI business?
Yeah, it’s a great question. I think it depends a bit on the sector. There are sectors like banking that are very focused on building this internally. I think that is a reality.
Do you think the internal build will work for them?
So it’s interesting. If you look at the MIT report, which is the one I mentioned that says 5% of models are making it to production right now, they actually cite a statistic that externally driven builds are 2x as effective as internal team builds.
I actually think there’s an interesting 10-year pattern here. Ten years ago, everyone bought software, right? Your tech team did not try to build anything. You started to buy, and you often bought way too many apps, but you bought 15 different apps, and that was what the technology team did.
Then, with the advent of cloud, you started to have a world where the technology functions started to think about building things. Maybe they started to have more custom applications that wrapped around that. I think GenAI has 5x’ed that, where now an internal team is given this enormous budget and told, “Go have at it.”
I think that’s complicated because when you hire any vendor of any kind to build something, you’re pretty disciplined about what they’re delivering, on what timeline, what the ROI is, what the milestones are, and how it works. I don’t think that discipline exists in the same way in internal builds. I also think the talent levels that internal teams often have are challenging.
3. It is BS That Enterprises Can Adopt AI Without Forward-Deployed Engineers
When you say the internal team builds are challenging, there are some things that you can’t say but I can. The perception from external or general tech crowds is that internal teams for—I don’t know, you name your boring large enterprise—are just really low quality. You’re not getting the top-tier AI engineers; you’re not getting top-tier devs. Is that true?
Look, I think the amount of talent that knows how to do this well is not large. That finite group mostly works in AI startups of various forms and large tech companies.
So I do think there’s real risk to the process of figuring this out from first principles in enterprises. That’s part of the cycle that we’re going through right now. A lot of internal groups have gone through the process of saying, “We must do this all internally,” but the reality is that this is an open-architecture ecosystem.
You’re going to adopt things like MCP and all the new technology and the new voice agent that comes out. You actually want a modular, open architecture where you can use all the best technology available and figure out how to link it all together. I think the desire to shape that all internally has been challenged.
I’ll give you one of the more interesting examples I can discuss. I was talking to an e-commerce retailer that had built an agent to handle its returns process. They spent $25 million building this agent, and at the end of it, after I met them, I said, “How did you define whether this agent worked or not?”
They were like, “Well, we built our own eval tool.” It’s not a joke. We basically analyzed a mix of speed of call resolution and sentiment.
The problem with that is, what if the agent hallucinates and says, “Here’s $2 million”? That actually gets resolved quickly, and the person’s happy. So they built this entire system from first principles, and what ended up happening was that a couple of months later, they shut it down and moved back to a deterministic flow. That’s not surprising to me at all.
I do think that’s a little bit of the adoption curve we’re in. Over the next 2 years, you’re going to see the CFO function put different guardrails on how this stuff is built and say, “What is the ROI? What are you investing in? What’s the metric? What’s the return?” That will change the adoption curve.
Right now, there have been a lot of science projects. I think that is realistic.
Okay, we have hundreds of thousands of listeners, and many of them are CEOs. If you are a CEO thinking about whether your CFO is equipped to buy and manage in this new environment, what should you be thinking about? Do we have the right CFO talent pool to manage this new environment?
Yeah. One misconception is that the leader has to be highly technical to make that decision, and I would actually argue they don’t at all. They just need the same muscle memory they’ve used in the past, which is to go through it.
What do you need to get a GenAI initiative working? You need good data that you can work off of for that specific initiative, clear milestones and outputs, and clear line ownership of the initiative. Then, probably most importantly, you want to anchor it in milestones and outcomes where you pay as it works.
The other interesting context for a lot of this is what I would call the Accenture paradigm of the last 20 years. A lot of times, if you think about the wrapper that’s been around software for the last 20 years, our founder, Francis Pedraza, has a founding principle of Invisible: “If there’s an app for everything, how come nothing works?”
It’s an interesting concept, because what ended up happening is you bought 50 apps, had Accenture come in, and paid them $200 million over 2 years to try and layer them all together. Often, you ended up a couple of years in with no working data and no linkages between them.
Those layers of sediment have been how the tech paradigm worked in the enterprise for the last 20 years. I think what’s different now is that if you’re thinking about a specific GenAI initiative, like a contact center, you don’t need to operate that way.
You can think about the operational metrics you want in your contact center. You want to think about call resolution, call performance, cost per call, routing logic, and so on. You can then look at both internal teams and a set of vendors who will deliver those metrics, make an evaluation, and, if the vendor doesn’t work, fire them.
I actually think there’s a very clear way to get ROI in this. Figure out the list of 3 to 4 things that move the needle for your business. Focus on those 3 to 4 things. Don’t spend money on 1,000 science projects.
Take your best 4 operational leaders and put them on those 4 things. Don’t locate it in the tech function. That’s the main advice I give people: your GenAI initiative should be led by the business.
That could be your head of the call center, or it could be your head of operations, but each of those people should have clear operational KPIs. They’ll get the stuff working, and there are a bunch of companies that have. It’s just a very different approach from saying, “I’m building GenAI.”
It’s really interesting that you said not to invest in a bunch of science projects, but to do 3 to 4 initiatives. Let’s do 3 to 4 initiatives again. I’ve put on that CEO hat. A contact center is just a big one that is homogeneous across everything.
Matt, there are so many players in the contact center space. I’m a CEO. I’m not a Silicon Valley guy. How am I meant to understand whether we go for Sierra, Decagon, Zendesk of old, Intercom, or any of the other players that we’ve seen in the space? How do you advise the bigger CEOs on buying in a wave of new innovation?
I think this is the other big challenge of GenAI adoption. You’re an average CTO or COO, and you’ve got 250 vendors a week pitching you. All of them sound pretty similar. In fact, I was with a customer yesterday who literally started the meeting by saying, “How are you different from the other 250 people that have pitched me this week?”
This is the dynamic: We have an oversaturation of companies that all sound relatively similar relative to agents. To make your question even more pointed, a lot of them don’t work.
You’ve got a fair number of the enterprise agent companies. Salesforce AI Research released this report that, if you test a lot of the out-of-the-box agents on single-turn and multi-turn workflows, they’re about 58% accurate on single-turn and 33% accurate on multi-turn workflows, which means they don’t really work.
You’ve got this challenge of 250 agents, 250 companies a week pitching you. You don’t really know how to select them, and you’re worried you’re going to pick someone who’s effectively a charlatan and it won’t work. The more you have a market where there’s a lot of excitement, the more you do have that risk.
I think the simplest advice I give—and, by the way, this is how we sell, quote-unquote—is to start with proofs of concept. Start with what we call solution sprints. Don’t pay a dollar until you prove the tech works.
We don’t actually sell anything. We meet a customer and say, “We will do it for free for 8 weeks and prove to you the tech works.” And that’s a very simple way.
If your tech works, you'll show it.
It is and it's not. But let me give you an example of how one of our deployments works, because I think it's fair enough if the answer is that it takes you 2 years to build anything. I'll give you an example.
Our AI software platform is effectively 5 modular components. Neuron, which is our data platform, brings together structured and unstructured data. Axon is our AI agent builder, and Atomic is effectively a process builder—we can build any custom software workflow. Then we have a Meridian expert marketplace, which has 1.3 million experts a year on any topic you can imagine that we bring into those workflows, and then Synapse, which is our evaluation platform on all of it.
We can take those 5 things and configure them to almost any different enterprise context. Just as an example, we serve food and beverage, the public sector, asset management, agriculture, sports, oil and gas, and a whole host of other different sectors using that same modular architecture. I think we end up scaling pretty materially once we show the tech works.
We're working with a company called Lifespan MD, which is a concierge medicine business across the US and internationally. What we're doing for them is building them an entire tech backbone. They have an enormous amount of fragmented data across EHRs, CRM, ERP systems, notes, and everything else. All of their data sits in a pretty fragmented format, so we're using Neuron to bring all that data together.
We do that very, very fast. If Accenture would take 2 years, we can usually do it in 2 to 3 months. We're then, on the back of that, building a lot of different intelligence and reporting so they can look at things like patient journeys over time, labs, genomics data. I don't know how much you use the Oura Ring or anything else like that, but they want to look at wearables and how all that content is looking. So they have a lot of detail on what any patient is doing at one time.
On top of that, we layer things like the ability to interrogate data and ask lots of different questions: “Let me look at who's used peptides, male, between 36 and 50, and what have been the results?” We're using Axon to build all that and to fine-tune the model to do that. Then we also build lots of specific custom agents for things like scheduling.
What you get at the end of that is a transformed, tech-enabled business with all of those different components. It does take us a little while to stand up, but once that is there, it's effectively hyperpersonalized software. That is my view on where this whole industry goes: You move from out-of-the-box SaaS to much more hyperpersonalization using the specific data of an individual customer, and that is what we do.
Do you think you can work with enterprise today with generative AI and AI implementation without an intense, fully forward-deployed engineering mechanism?
I don't think you can. We've doubled down. A huge part of what we do is forward-deployed engineering. We now have 8 offices in 8 different cities and 450 people. We're fully focused on forward-deployed engineering, and I can tell you from a decade in my prior life, you just cannot do this with out-of-the-box SaaS. It does not work.
What do the economics of FDEs look like? Obviously, Palantir made it the sexiest thing ever. I love the way tech crowds work, where we all just get super excited by an acronym. It's like, “This is the coolest thing.” But what do the economics look like?
One thing I'll say is that forward-deployed engineering has come to mean a lot of different things. A lot of forward-deployed engineering across the broader market is more like solutions engineering, where the people answer your questions and show up at your office.
I think forward-deployed engineering done well is executing a very specific workflow build. You're effectively configuring a set of core platforms to build something hyper-specific for that customer. One of the questions is how good your platform is, because, for example, you could argue Accenture is forward-deployed engineering, right? But that build may take 3 years.
In our case, I think we've built modularity and a lot of the new software workflow development workflows into what we do, so usually our forward-deployed engineering motions are about 3 months. We'll come on board, customize everything to the hyper-specific way a customer wants it, and then hand it over, with something built on a basis that works. It does require ongoing fine-tuning.
That's the other big difference that people should acknowledge, right? You can't fine-tune a model in an enterprise context and just leave it for 4 years and hope it continues to work. I could give you 100 examples, but take healthcare: GLP-1s launch. You do need to fine-tune the model for the new context of the market. We do view it that way, but—
I'm very naive, so forgive me on this. Do they pay additional for FDEs to come? Do you pay additional in terms of ongoing maintenance, just on the economics of it?
For many of our competitors, they do charge. We do not charge anything for FDEs.
Why not?
I think it goes back to my general premise that the best way to differentiate in this market is to prove that your tech works. The way that we do this is we say, “You will pay when the software is up and running.” We're able to do a lot with 1- to 2-person, small FDE teams.
Once that's stood up and running, then we do have ongoing software. The paradigm that we're evolving from is that, over the last 20 years, you had the system-of-record layer, which was where a lot of the value was. What we're building is hyperpersonalized systems-of-agility layers, kind of what sits atop that.
I think the Accenture paradigm is what people are afraid of, and it's very hard to convince somebody you're going to pay time and materials until it gets working. So I spend less on sellers and more on forward-deployed engineers. That's my simple math.
I always think the biggest mistake that people make is that they don't put the hat on of their customer. I think the reason the show has been successful is because I put the hat on of different customers. A lot of the customers that we have are startup founders who create amazing products, and everyone wants to sell into enterprise. It's where the money is.
If I'm a startup founder thinking, “Do we need FDEs? How do we do FDEs? How do we move into an FDE model?” what would you say to them that they should know if they're thinking about starting that model or potentially needing that model, knowing all that you know?
I think it depends a lot on the nature of the business and what you're trying to build. If you're trying to build a knowledge management system of public filings for finance, for example, you don't need FDEs, because what you're building there is a repository of information that people can access. You have similar things in healthcare, for example. If you're trying to change workflows, you do need FDEs.
4. Best Capital Allocation Decision? What did Matt Learn from it?
I think that's the simple paradigm difference in my mind: If you're building something where the hardest part is getting adoption and workflow embedding, and you need to actually change the way a company works, then yes, forward-deployed engineers are the only way to do it. It's interesting. There aren't that many people who have expertise doing that, so it's a hard thing to train and learn. But I do think it is the only way to get the enterprise working.
You've said several times, “Don't pay until you prove that it works,” and you said earlier, “Pay as it works.” That's not the SaaS business that we've been trained on, Matt. I'm a SaaS investor. How does the pricing model of the future look in this very new environment?
Let me step back for a second. I think an interesting thing is that if you look at the economics of SaaS and enterprise 5 to 10 years ago, and then look at any large public enterprise software business and how much of its revenues actually come from services, I think you could argue that out-of-the-box software has always been a lie to some degree. It's a weird thing to say, but they always had a ton of configuration and just dressed it up to some degree.
I think SaaS was even more challenging than that because, often, in the unit economics of SaaS, you're selling at a much smaller cost per customer. The SaaS business that worked was actually about selling something where the out-of-the-box setup was quick enough that you could make it work with the sales team, where you didn't have to do lots of configuration. The minute you had to bring in FDEs in a SaaS context, your economics broke instantly, right?
On the enterprise side, the way people made it work was— that's why Accenture grew so much, that's why Cognizant grew so much, and that's why TCS grew so much. I'll give you an example. If you take insurtechs, every one of the major insurtechs, like Duck Creek, has a set of core data schemas, a series of analytical logic, and a front end.
The ones that did really well had momentum and push from the SIs that got them going. Their economics were geared by having somebody else do all your services around what you did, and you got something up and standing at the end that worked.
I think the challenge with GenAI is that that motion doesn't really work, because what ends up being built at the end of the day is something that is hyper-specific to that customer. If you actually think about the nature of fine-tuning an LLM or creating a knowledge-management system, it's not a box. It is something that uses a lot of different, consistent tooling, but it has to be customized.
5. Are AI Talent Marketplaces Dead? What is the best model?
The way we do that is we stand it up, get it working, and at the end of it, usually 2 to 3 months in, the payment happens when we pass user acceptance testing and validation and it works. Here's the other thing I'll say: we use SaaS as a paradigm because that's how software has worked. But machine learning has been around in the enterprise for 10 years. I was building machine-learning models 10 years ago, and that's always been a motion that looked like this. So what's happening now is we're starting to realize that the GenAI adoption paradigm in the enterprise works the same way that ML does.
Totally get that. When we look at the different products that you have today, the Expert Platform is one that gets a lot of attention. How much of the business today is the Expert Platform? I find companies are lumped into categories because it's easier, and you have your Mercor, your Surge, your Invisible, and you're all put in this category of, “Are you all just talent marketplaces?” No one wants to be a talent marketplace, it seems. How much of your revenue is the talent marketplace, and why does no one want to be a talent marketplace?
Yeah, let me think about that in a couple of different ways. I actually think the AI training space has many different players that do have many different business models within it. There are 4 to 5, but actually, they're all quite different. I think of us much more as an AI training platform than just a talent marketplace.
Meaning, we have 1.3 million experts that come through the marketplace. Here's the simplistic question I think AI training asks: You have to be able to source any expert in the world with 24 hours' notice. You have to be able to source a PhD in astrophysics from Oxford, put them into a digital assembly line, and 4 days later generate perfect, statistically validated data that will be compared head-to-head to somebody else's data and make sure that that is perfect at the end. That is an incredibly difficult thing to do.
A lot of what I saw when I took over Invisible was that this motion was incredibly applicable to the next phase of the enterprise as well, which is the fine-tuning motions, the training, and the ability to statistically validate for an enterprise use case like claims processing. It's the same motion. I actually think AI training will be used next in banking and healthcare, and then after that in many other enterprise contexts.
The historical business I took over in 2024 was pretty materially weighted to the AI training side of the house. But I came in with a thesis that enterprise would be a huge source of growth, and I think as you see next year evolve, we've confirmed 12 enterprise deals in the last 45 days. We see pretty good momentum on that side of the business, and I think that's where we will evolve: to doing both.
I think the 5 core platforms we have allow us to serve a whole host of different end markets, and I do think that's very different from the other AI training players you mentioned. I think we're the only player that spans that broad-based view in the same way.
Can I push on the talent marketplace side? How much of the business is that today?
I won't say an exact number, but it was a pretty material percentage of 2024.
Okay, got you. So it's a pretty material percentage. The one thing that's also striking is the concentration of revenue to a couple of core players. When you look at other providers, it's about 2 players that make up more than 50% of revenues for pretty much every provider. Is that the same for you? How do you think about what that revenue makeup will be, given the enterprise diversification that you're talking about?
Yeah. I do think this is a space where there are not that many players that are actually building LLMs. By definition, the whole space has concentration. I would not disagree with that.
One of the really interesting things for us on the enterprise side is that we have materially more diversification now in the number of customers we serve across a whole different range of topics. I also think you're seeing more early-stage model builders that are building hyper-specific topics. That's the other part of where we see expansion in the total customer base.
When you come to negotiations with a client, given the revenue concentration, how do you play that staring contest? Essentially, they go, “We know that we are one of your core customers, and we will squeeze you on price,” and you go, “I know I'm one of your core data providers. I will stand firm.” How do you handle that negotiation? It is a staring contest of sorts.
I think people are willing to pay for good data. That's my simple, firm belief. If you think about the importance of these models and the cost of compute, that is actually a huge chunk of the cost space. If you think about one week of bad data, it burns a lot of compute.
I think what we've seen—the reason it's been the same 4 to 5 players in this market for a couple of years now—is that it's really hard to do well. People are willing to pay for good data. We have a very collaborative dynamic with all of our customers on that front.
When you provide a service that's helpful, people are willing to pay for it. If you provide a service that doesn't work, people don't pay for it. The interesting thing I would say on that front is that, a lot of the time, the discussion topics anchor around proven value. We'll get a topic that comes in, a multimodal audio model, for example, and we'll go head-to-head with somebody on that that week. At the end of it, we win or we lose. If you win and your data is way better, people are willing to pay for that.
Totally get you. I had a chat last night with a board member of another company in the space, and he said 2 things that really stood out to me. He said, “I'm just drastically shocked at the lack of price sensitivity for the core customers. They're willing to pay pretty much anything.” Is that the case, or is that a bit of an exaggeration?
I think that's an exaggeration. I think it's an exaggeration. I think that there is a fair price. If you think about classic economics, people are willing to pay a fair price for good data. I don't think we operate in a model of trying to give anything unreasonable. I think there's actually fairly standard price bounds across all the players here.
Is data commoditized? When I think about pricing power, I'm a massive fan of Hamilton Helmer's 7 Powers. Amazing book. When you think about pricing premiums, you get that through not being a commodity, through owning the supply of a rare asset.
Is there commoditization of data, and are we in a race to the bottom on the pricing of that data? Or do you own the supply of vetted workflow data for surgeons in Oklahoma?
Yeah. Let me take that. I'll actually start with the market context, and then I'll use 7 Powers. It is a great book, and I'll use one of his frameworks for that.
I think the market context that is somewhat misunderstood here is the way that human data becomes more and more important over the next decade. The reason for that is if you think about the different types of things you could train off of, synthetic data gets mentioned a lot. Most of the time, synthetic data is useful for things like ground-truth information, like math, where there is a clear output that is right or wrong.
Now let's take all of the different reasoning tasks, like a multistep reasoning task. I mean, even a simple one, like what movie would I select based on these 5 preferences?
Exactly. Well, then let's take that question and add into it audio, video, multimodal language, and the ability to do it in 45 language contexts. The ability to think about computational biology in Hindi versus French versus English versus English with a Southern accent—that paradigm is actually incredibly hard to train on, and we're still in the first inning of a lot of those permutations of complexity, is what I would say.
For a multistage reasoning test that requires a PhD in multiple different languages, human feedback is going to be important for that for the next decade. I have a strong belief in that, and when I chose to take this job, that was one of my core convictions: the enterprise is going to need that too.
If you take legal services, for example, a lot of the way you're going to need to validate that is with legal expertise. There's no corpus of information you can train from. So I would start with the idea that I think the market tailwind for the next 10 years is actually in the first inning, because there's the LLMs, then there's the more sophisticated enterprises, and then there's everyone else that needs to train, validate, and move to fine-tuning.
Again, contrasting, there's the pretraining and LLM work, but then to fine-tune a model to a specific context, most companies don't even know what that is in the enterprise yet. We're in the first inning of that whole process.
I think market demand is going to continue to grow pretty materially for a decade or more. The Hamilton Helmer framework is an interesting one. My favorite example is that he talks a little about what he calls institutional memory. He mentions the Toyota Production System as an example, right? Toyota would literally say to people, “This is exactly how our factories are set up,” and nobody could replicate it.
I think the interesting thing about this space, and why you’ve had a consistent set of folks doing it for a while, is that they have to go through the process every week of spinning up. We have 1.3 million active agents, or experts, that come into the pool in any given week. We have 26,000 of those that we’ve selected who have to start within 24 hours and produce perfect data. Think about the challenge of scaling an organization that can do that for 5 years at really high quality and consistently turn and evolve with the different permutations of the market and new ideas of training. It’s really hard to do.
I think that was what got me most excited when I took the Invisible job: the question of whether you can make AI work in a really complicated context. Very few companies know how to do that on the enterprise side or on the training side, for that matter. I thought that was a really unique institutional-memory context. It is a digital assembly line, no different than an auto factory, and I think that is a hard thing to replicate.
6. How Does the Data Labelling Market Shake Out: Who Wins/ Who Loses
The other really interesting area that this board member said to me was that he very much agreed with you. He said exactly the same words as you in terms of the first inning of data, in terms of just how much the market size will increase. He said the other thing that I really didn’t understand when I made the investment was the specialization of data and how we are moving into the acquisition of these insanely niche data-supply pools, where it’s not like cat, hedge, zebra crossing—zebra crossings, what do you guys call it, a pedestrian pathway or something?
I did not see the specialization in the unbundling. Is that something that you see too, in terms of these very micro-niche, specialized data requirements?
Absolutely. I think 5 years ago this space was what I would call cat-and-dog commodity labeling. I think there were a lot of Google Sheets in that era. You’ve seen some comments on that. This sector has evolved the same way most technology sectors do: it started with Google Sheets and cat-and-dog labeling, and it’s evolved to real digital assembly lines, huge velocity of expertise, and incredibly specific expertise.
Here’s a funny example. We have to be able to validate an architectural expert on 17th-century French architecture who speaks French. That is a complex thing to do on 24 hours’ notice, right? The ability to source, assess, and validate is important. One advantage for us is that, because we have 5 years of data on who’s been good at what task, there’s real institutional memory in how you do that selection and assessment. I think that’s one of the core advantages we have from that.
How important is pay? I think a couple of other providers have said that bluntly: it’s about how much you pay. You pay more than the others, you’ll get the good talent.
Look, I think of our business like Uber. What I mean by that is we source talent at the price at which people will do the work that is asked of them. The same way, if you’re standing on a street corner, your question is, “Can I find a ride that will pick you up at this moment within 3 minutes?” That matters. It’s a different price if it’s raining, and it’s a different price if you’re in Rio de Janeiro versus London. The price depends on the market context and the specific place you are.
I think extra pay is the same dynamic. A lot of what we’re doing is what I call price discovery. The nuance I would add to what you’re saying is that you can overpay a really bad expert, and that is a total waste of everyone’s time. What I think our customers appreciate is that we can tell you, between a $150 expert and a $130 expert, the difference in expertise—you—
Do you think you have control of a finite supply of data providers? If you look at 7 Powers, one of them is acquiring finite supply.
I don’t. I actually don’t think finite supply matters. What I mean by that is that the expertise needed varies so much month to month that, if you tried to build a world where you bottled up whatever supply it is, it would change in 3 months. We actually relish that concept.
The dynamic, again, is why I would use Uber and Lyft—you could use Airbnb and VRBO in the same context. I don’t think people—I don’t think experts go on 5 platforms, right? I think, actually, what you want is a 2-way marketplace where you need enough demand for people to be interested, and you need enough expertise—many experts. I think the reason we get 1.3 million inbounds is because of that supply-demand balance.
I don’t think this moves to a world—and I actually would never say it moved to a world—where there is 1 player coming out of this. I think there are benefits to everyone in having numerous players that do AI training. It’s a question of being 1 of the players that has that balance.
You said there about the switching of preference: “3 months ago it was this; now you want something completely different.” Switching costs are another consideration for data providers in this way. Are there inherent barriers to switching? Is there any loyalty?
If you’ve learned how to do a certain data task really well, there’s incredible value in that. That’s what I like about the way—and let’s take the enterprise context again, because I do think it’s a good one.
We’re doing a lot of fine-tuning on some pretty interesting topics. One example is work we did with [likely SAIC], Vantor, and the U.S. Navy on fine-tuning a model for underwater drone swarms. The question in that context, if you think about—
Niche.
Very niche. This is why I use it as an example to answer your question. If you think about it in that context, you’ve got a bunch of underwater unmanned vehicles, and they’re taking in all the drone and sensor data from the interaction patterns of those vehicles.
What they want to know is: an object is in the water near them. What do they do? Do they react? Do they pull back? Do they alert another drone? Do they engage? What are the protocols for that?
Fine-tuning a model to take in all that complex sensor data, fine-tune it, train it, and build a decision-making framework for those drones involves a lot of logic. I think that’s why it’s been a great partnership with [likely SAIC] and Vantor, because we built logic around how to do that.
There is real sustainability and expertise that you build up. The way I think about our enterprise motion, for example, is that every sector is led by somebody with deep, deep sector expertise. We do build real logic around those topics, and I think the same is true for multimodal video and audio. It’s true for legal.
I actually think a lot of the training work, even on the model-builder side, is moving in that direction. One interesting view I have is that people talk a lot about public benchmarks. That tends to be one question you get a lot: “Are we reaching a point where models are not improving?” I think about it very differently. The models are now all moving toward hyper-specific things where there isn’t a public benchmark for them, by definition. They’re moving to more specific tasks that are very different and not something you can publicly benchmark in the same way. That’s where we see more and more model improvement every day, both from model builders and enterprises on these specific tasks.
You said something about the benchmarks. I’m just so interested: Gemini 3 killed it—it’s the best ever—and then yesterday Claude Opus 4.5 killed it—it’s the best ever. Next week Sam’s going to release one. Does it matter? Are we in a world of such transience and flux where, really, we should detach ourselves from these blunt updates that last for days?
Look, I think the benchmarks are a useful framework for society to gauge progress on this topic, and it’s a very often-discussed topic. People want a way to answer the question of how the models are improving, and I can tell you unequivocally that the answer is yes. By every measure you look at, they are improving.
They’re not only improving on the benchmarks, but also on specific tasks, like research for investments, for example. You can see the models are much better at doing certain tasks. I think what you’re seeing start to happen is that people—and we’re doing this as well—are building very specific work-based benchmarks to calibrate certain things, like how well the model does at building an LBO model, for example. You’re going to see more and more benchmarks cited.
The complexity then becomes that, if you move from 5 main benchmarks, like SWE-bench and others, to 600 benchmarks, you kind of lose track of who’s doing well on which things. My interesting view on that would be that I’m not sure benchmark progress is what determines enterprise adoption.
What I mean by that is, if you take the fact that the models have improved exponentially over the last couple of years and say consumer adoption has been massive—KPMG had this report that 60% of consumers use this on a weekly basis.
The adoption curve in enterprise is not going to be a question of generalizability. It's going to be a question of hyperspecific performance on a specific task. There isn't actually a benchmark for that. If I take an investment summary document for a private equity firm, for example, there's no benchmark to say, “Firm 1, this is how you write investment committee memos. Does this generate something that looks, with 99% precision, like something you would roll out?” There's no benchmark to do that.
That's where I see the adoption curve: the fine-tuning and inference layer of actually testing that, getting into a place where that firm could say, “This looks good. I'm okay with this. I've tested it.” Machine learning has a context—I don't know if you've heard the banks do this thing called model risk management—where they actually do a whole host of validation and testing on things like redlining before they roll a model out. That's what the enterprise is going to have to do.
It's not that model improvement doesn't matter. I actually think the benchmarks are a good way to get some sense of model improvement, but they're almost orthogonal to enterprise uptake. I think enterprise uptake depends on trust and precision on specific tasks at 99% accuracy, not generalizability.
If those specific tasks are removed in the way that you said—like summary documents for investments—often, they're done by more junior people in the earlier stages of their careers, when they're building and scaling those skills. Do you think we'll have a talent pipeline problem if we remove a lot of those junior roles, which we're seeing in certain cases already and I think we'll continue to see, where we won't actually have the graduation pathways that lead to the leaders we have today because we've removed those junior roles?
I don't actually think so. I think one of the challenges is that the adoption curve of this stuff is going to take a lot longer than people expect. I said this to you earlier: I think in enterprise, this is a 5- to 10-year adoption journey, not a 1- to 2-year journey. I think you have a dynamic where people will have a lot of time to react and think about what's useful in addition to that.
I actually find that a lot of the people coming out of college right now are some of the highest adopters of this and the most useful for these kinds of tools. We're hiring more and more people of that profile, not less. I think the adoption curve and the usage curve of that group of people mean that certain tasks will not be done, but there will be many more.
I'll give an example: accounting. If you worked at a bank or any accounting firm in the 1980s, this is absurd to think about, but you literally calculated revenue and financial statements with a slide rule. People would literally sit there and generate a financial statement manually on paper with a slide rule. That was how people did accounting.
Then Excel comes around, and that becomes the main tool everyone uses to do accounting. In theory, you'd have fewer accountants because you went from manual generation with slide rules to Excel, which makes it much easier to do that. You look today, and we have about the exact same number of accountants, and about the same number of junior accountants. What's happened is that way more people do way more sophisticated accounting scenarios with the tools they have.
It's this old idea of Jevons' paradox: you increase consumption with advanced technology, and so the number of accountants didn't go down. You actually had way more accounting. In fact, every FP&A function is probably larger now than it was 25 years ago because the work people do is more sophisticated.
I totally get that. I do want to go back to what we said about market composition and how we see the different players. Is this a market where, as you said with Uber and Lyft, there are 1 or 2 players that take the dominant market share, and then there's everyone else? Is it a cloud market where it's much more evenly distributed? How do you project that out over, say, a 10-year horizon?
I think in both AI training and enterprise, I don't think the answer is one player. Interestingly, in enterprise historically, there's probably been Palantir and not many others. I think that's why you've seen more people wanting alternative options to that. I think that's part of the reason you've seen so much excitement around enterprise AI recently.
I think most of these markets end up with 3, 4, or 5 players. I don't actually think it's even 2. I think that choice in consumer markets tends to allow this to happen, and that's a good thing. You'll have some specialization on certain topics. Maybe some will be better at coding, some better at specialist tasks, and some better at PhD-level work, but I think it'll stay with a fair amount of choice.
When you look at the landscape, who do you most respect, and what do you learn from them?
I would say Palantir is the company I probably respect the most in enterprise AI.
It's really interesting. You see them as a competitor more than Surge, Mercor, Turing, or any of the others?
I think they're all competitors in different ways to different parts of our business. I call out Palantir because I think they realized 10 years before the rest of the tech market that forward-deployed engineering and customization would be important. That was a very countercultural leap at the time.
I spent a lot of time running forward-deployed engineering teams, and most of what I saw was players like Accenture. What was called tech services back then was not a place that anyone wanted to play in. Palantir spent a decade, before anyone realized it was important, building good tech. I have a ton of respect for that and for the culture they built out of it.
On the AI training side, I won't comment on anyone specific. I think all the players in the space are good, and they all do different things well.
7. Are Revenue Numbers for Data Labelling Real Revenue? Or GMV?
There are large revenue numbers thrown out.
Yeah.
Are they revenue? I've done shows before with them, and I got battered, bluntly, when people said, “It's not revenue, Harry, and you can't categorize it as revenue.” Is it GMV, not revenue? Are we playing fast and loose with the truth on revenue versus bookings?
I think it is revenue. The rate you get on every project is different, and the margin you make on every project is different.
Can you help me understand? Sorry, I'm very naive. If I'm acquiring amazing talent and I get paid for that, and then I have to pay them and get my take at the end of that, how is that different from booking on Airbnb, where I get my take from a location but have to pay out to the owner?
Good question. Airbnb has one consistent fee. That's the difference. There's actually a fair amount of variation based on the skill set of the expert. You don't have a consistent rate relative to the booking amount. That's the biggest difference. There's huge variety depending on the project, the expertise type, and the expert type of what you book.
Are there any other big misnomers that you think are pronounced in the industry, where you consistently think, “I wish people would change the way they think about it”?
I think the biggest one is the view that, when I first started this job, the main pushback I always got was that synthetic data would take over and you simply wouldn't need human feedback 2 or 3 years from then. From first principles, that doesn't make very much sense if you think it through. If you think about the diversity of tasks that exist in the world and how long it would take you to get comfortable with the accuracy, it doesn't make any sense.
I'll take legal services because it's a really interesting one. A lot of the legal data in the world exists with big law firms; it doesn't even exist in public. The corpus of publicly available information has been commoditized for years at this point. Most of the logic is incredibly contextual to language, culture, multimodal context, and the information stored in individual companies, as an example.
The only way to actually do the fine-tuning process consistently and get it accurate for any specific context is RLHF. I actually think that in my decade at McKinsey, that was the thing I realized was different about traditional ML models versus GenAI. In machine learning, you can backtest. You can get to a really clearly statistically validated outcome without any human intervention.
On the GenAI side, you're going to need humans in the loop for decades to come. I think that's something that most people are starting to realize. It's always confusing to people when they hear, “Oh, that's how models are trained on the back end. I didn't realize that's how the statistical validation works.” I think that's been an interesting evolution.
You're profitable, correct?
This year, we have started to invest a lot more. I think one of the big differences is that historically, Invisible had only raised $7 million of primary capital in its entire 9-year journey.
We initially announced $100 million; actually, right now we’ve raised $130 million, and so I’m investing very heavily in technology. So we will not be profitable this year.
Good. Can you just take me to that decision? This was going to be my question: that’s a very clear decision to be profitable, and profitability often comes at the expense of growth, naturally. Can you take me through that decision-making for you and how you thought about it?
Yeah, look, to me it was a simple one. If you think about the dynamics of return on capital, you can either harvest capital or invest capital, and your decision to invest depends on the growth you see as a result of that investment.
I think we’re in the greatest environment for growth that has ever existed. I think Invisible is really uniquely positioned to capitalize on that growth, too. I think of our 4 or 5 core platforms and the growth vectors across both AI training and enterprise, and there were just way too many different things I thought were interesting to invest in. It was the clear best use of capital.
I’m trying to build this for the next 10 to 20 years, and I think if you want to build enterprise value for 10 to 20 years, now is the time to invest and build. I hope we never get to the harvest stage, but it’s definitely not now.
Where are you not investing that you want to be investing?
I think the simplest answer is actual physical-world interactions. What I mean by that is, I think a lot of the most interesting data that we don’t even really have access to yet is things that exist in the physical world, which are more complicated to acquire and organize.
I’ll give you an example. We’re serving one of the largest agricultural conglomerates in the U.S. on herd safety—actually monitoring risk factors and determining when you should send a vet for their herd of cows. That whole process relies on us actually sending forward-deployed engineers to farms, dropping Starlink terminals into those farms, and building out custom computer-vision models in those contexts.
I think there are so many different physical-world contexts that become really, really interesting, but it does take cost and capital to build those out. Oil and gas—oil rigs are an interesting one, as an example. I think physical-world interaction patterns are some of the most interesting growth vectors for this, but they do take time and money to invest in. Robotics is another big part of that.
8. How Important is Brand for AI Companies Selling Into Enterprise?
One area of investment that I think is interesting is brand. How do you think about Invisible’s brand today?
What was interesting when I took over was that, if you looked at the entire public internet, I think there was 1 article available. We’ve definitely spent a lot more time this year thinking about it.
Was that a deliberate decision?
I think so, to some degree. Invisible has a culture where we believe in doing great work for customers, and we were not really focused on telling the whole world about that.
Does that become detrimental to the business at some point, though?
Yeah, look, I do think branding matters a lot. My view now is that it’s been very helpful for us to spend time where I spend a lot of time—I spend about 70% of my time on the road. I go to a lot of conferences and things like that, and I think building a brand is really important for trust, awareness, and engagement. I also think how you tell that story is really important.
I’m very much a believer in this. One of my favorite quotes is from Marc Andreessen, who has this idea that when private and public narratives diverge, that is the risk or the opportunity. If you say things you don’t believe to be true, or if everyone’s saying things they don’t believe are true, then what is the actual private narrative?
So I think it’s been very important for me to make—
Can you just help me understand that?
Yeah.
What do you mean by that?
Hypothetically, if I was going around saying we have an out-of-the-box agent that does everything, and that wasn’t actually true, that either creates an opportunity for others or risk for us. That’s how I think about it.
I think what’s been very important for me and how—
Is that not our industry? I’m sorry. I don’t mean to pick a fight with Marc Andreessen, but hello, Marc. Our job is to sell and then deliver later. I’m looking at it thinking, well, I’m—
Well, I guess it’s all a question of degrees. In my mind, I want to say things where the narratives are the same publicly, in terms of what our team thinks and what our customers experience.
I think that’s part of why I’ve focused on saying some of the nuances of what’s not working and not claiming everything works out of the box. That is a different approach, but it’s been core to how we’ve thought about building the brand. We are building this around trust. I want a company we work with to know that if I say this will work, it will work. I think you only get one chance to do that, right?
Do you agree with “fake it till you make it”?
Oh, that’s such an interesting question. I think it depends on what faking it means, right?
One of the things I think is really complicated about GenAI is that it’s nondeterministic. If you’ve never built a machine-learning model to do pricing in industrial manufacturing, you can still understand what data is available, understand how the price is being set today, and get pretty comfortable that what you build, if you say you will build it, will work. I think that is okay.
The challenge of nondeterministic systems is that there is more risk to faking it till you make it. You can go out and say your agent will do anything, and then you actually have to deliver an agent that works, right?
Right.
I think that’s part of the interesting dynamic—you’re asking about accounting dynamics. A lot of the contracts that people will sign right now are like, “I’ll sign for 50 agents to be delivered,” but then the question is: do you deliver the agents? Do they work?
I think that is a different thing than SaaS, to go back to your earlier question. If I deliver a SaaS box, I know it will work. If I deliver an agent in the current world, there was actually a report AWS came out with today. It’s interesting that 70% of agents are actually not even AI agents as you think about them. Most of the agentic processes today are actually traditional script writing and just traditional automation.
I don’t self-identify as an agent company at all. I think we do AI agents, and AI workflows are a core part of what we do, but we do data, training, and fine-tuning. Agents are one tool in the toolkit. A lot of the time, it won’t work.
Did you see the video of the robot going around the house recently? It was the worst thing ever. It took 11 minutes to take a glass out of the dishwasher, and at the end it was like, “And this was controlled by Simon in the back room.”
You’re like, the shittiest robot ever was controlled by some weird dude in your back bedroom. This is so—
Yeah, I did see that. I think robotics is another one that will take longer but will be really interesting when it works. By the way, I think even in that case, you’ll need more task-specific robotics, not just broad-based.
Have you ever faked it till you made it and been caught out? Did you learn anything from it?
When I first started working—it wasn’t even called AI back then. It was called data analytics. This was probably 12 years ago, in my McKinsey days.
The firm gave me a pretty interesting purview to try and explore where I could build out AI offerings across different sectors and customer bases. I don’t think I knew what I was going to build, candidly. The interesting dynamic was that I had a lot of conviction, partly because of some of the things I’d done before, that AI could be really useful for a whole host of things, from inventory forecasting to pricing to credit underwriting.
You intuitively thought about the sources of data and the fact that 70% of the software in America is over 20 years old. Most of that data is massively fragmented and not clean, and so a lot of the decision-making that happens in the enterprise is done in a really fragmented way.
This is what I did know: you took your average person, like your average sales rep making a call, and most of the time they were Googling things to try and figure out what information they had. Not now, but this was 12 years ago. They had very little information on the screen to say customer information or what they might sell. So I had a lot of conviction that would work.
I did not know what would be most interesting. In fact, there were areas I thought would be really interesting, like banking, that were actually much harder to do consistently. It was somewhat—you mentioned earlier, like banks. The average bank spends 93% of its tech costs on maintenance initiatives.
7% goes into building new things. It’s my favorite thing with people that I—I just had one of the CEOs of a big vibe-coding platform on, and he was like, “SaaS is dead. We’re going to build our own products.”
Yeah.
And I’m just like, maintaining, provisioning, updating.
Are you buying?
Yeah. If you've never gone through InfoSec and approval at a bank, banks are slow for very good reason. Banks are much more complicated to do a build like that in, right?
This event that I was at last week was a bank. They have 6,500 people in KYC alone. 6,500 people.
It's a great example. When I was doing that in the early days, partly because there was very little media coverage or interest in it, I was figuring stuff out from first principles. The degree to which I faked it until I made it was that I had to figure out other people I worked with and customers that trusted me enough to allow me to co-iterate and develop stuff with them.
I had to figure out a way to recruit really good people. I actually think if you take any business very simplistically, it's a question of whether you can build trust with customers and co-iterate to develop and make things work, and then whether you can recruit unbelievable people to deliver that. It actually comes down to recruiting in a lot of ways. I think that's the number one thing we focus on.
I think of us as a talent company as much as anything else. You could argue that, not to use a sports analogy, [likely Nick Saban] did not build Alabama football with the process. He built it by recruiting the best football players in the country. I think about that the same way: you have to recruit great people.
To some degree, in the early days of that, 10 or 12 years ago, I was setting a vision, trying to figure stuff out, and iterating on a lot of things. I do think we ended up building a lot of things that really worked, but it took time and iteration as much as anything else. It took iteration and trust.
The counterintuitive thing is that I didn't fake it, and I never told people it would definitely work. My entire approach was to say, "I think this would work. This is my reasoning for why I think it would work, and let's build this." A lot of people were very comfortable with that. If you go in and say, "I have an out-of-the-box AI that solves all your problems," people are pretty skeptical.
I do just want to stay on recruiting because, again, I always think the show is successful because you put on the hat and you're like, "As a startup CEO, one of your biggest jobs is to recruit great people." Having recruited people across different companies, both McKinsey and now, obviously, Invisible, what would you advise startup CEOs in the earlier stages, knowing all you know now, on what it takes to be great at recruiting, acquiring, and retaining great talent?
It's probably the topic I spend an enormous amount of time focused on. It's probably the topic I think about the most because I actually do think if you get amazing people, everything else will follow from that.
So you agree with the moniker of, "Hire great people and let them do the work," because people push back on that?
I think it's not just hire. It's hire, retain, and evolve great people because I actually think you have to give them a platform that they enjoy day to day. The 2 things that I believe, which are somewhat counterintuitive, are that when you recruit a great person, I don't think about role most of the time.
People are very role-focused: "I will hire this person, and they will only do oil and gas," as an example. The reality is that really good people will run 5 or 6 different roles across your company. They'll run 7 or 8 different products. Particularly on the business side, you may have somebody who does everything from delivery to sales to account lead, and you can be comfortable with that if you hire great all-around athletes.
The second thing is that it has to be fun. This is my view on one of the narratives that has gotten a bit lost in the last couple of years: if you have a culture that is brutal to work at, people will leave. They might stay around as long as your stock is high, but they're not going to stay.
You have to create an environment where people really enjoy going to work every day, where they're intellectually challenged, and where they feel like they can unleash creativity. I spend a ton of time thinking about that.
I don't want to argue back, but I want to build great companies myself. I'm trying to with 20VC, and I try to build good cultures. Revolut is a brutal culture to work at—it's famous for it—but Nick has famously always told me, "Culture is winning. Winning is what matters. When people win, they learn more, they earn more, and they grow." That really is culture. Brutality in bounds drives humans. Is that wrong?
No. I think it's actually right. Let me caveat what I said: I think it's also the nature of the business I'm in, being AI. I actually think that's a very true statement if what you're trying to do is scale a relatively consistent business model to do 1 or 2 things.
That is a function of execution and hiring people to go into very specific roles and do very specific things well. Let me caveat my prior comments on that: I think the difference is that a lot of what we do is fundamentally research and exploration.
In the AI world, it is a different dynamic in that you're trying to figure out very specific problems with customers, solve and build really unique technology, and so I think in that world you do have a different cultural dynamic. It is a research culture as much as it is an implementation culture.
Is that difficult, then? We do a show every Thursday, which has blown up, which is incredibly nice for us as a business. Essentially, we have Jason Lemkin and Mario Gabriele, 2 VCs, and we talk about the news. We talked about Sam Altman and war mode. Can you do a war mode in the culture of research and AI, where it's maybe more thoughtful? Does that work?
There are definitely parts of our company where, if you take our delivery and operations team, they're in war mode quite a bit of the time. Again, I'm more describing general, countercultural beliefs I have on how to hire certain sets of great people. I don't think it applies to every single function of the company.
I agree that there are definitely times when you have to be able to push really hard to deliver certain outputs, and I think we do a great job of that. But I also think that ideas like every great engineer should be able to spend 30% of their time on new projects as well as sprinting on the existing ones—it's paradigms like that that are important.
What decision are you scared to make, but you think about it often?
The simplest answer I'd have to that is that growth in this industry relates to the amount of capital you raise. Your earlier question about investment is important. I do think there's a world in which, if you pursue hyperscale growth, it is possible, but you have to invest a lot more to do that.
Every new company, every new customer you onboard, does cost money to do the forward-deployed engineering work. You invest more in your technology. There's an interesting question: do you run a business for consistent, steady growth for 20 years, or do you try to build something that gets to $50 billion to $100 billion and becomes game-changing?
I think we have very much tried to operate in a way where we have a path to profitability and everything else, but we are going to invest in the near term because I think it is a very interesting time to do that.
I know you don't like to name names, but I can. When Mercor raises $2 billion, are you like, "Fuck, we need to raise more money"?
It's interesting. If you look at the players in our space, there have been very different levels of capital raised, and people have had more and less success. I actually think a lot of our investment is in different areas than many of our peer set are focused on.
A lot of it is in things like the enterprise. It's in core software platforms that are maybe a little bit different from what others are focused on. I don't know that I think you can raise a lot of money; the question is where you spend it.
9. Remote Work vs. In-Person Collaboration
Again, I actually think most of the capital we need in the next 5 years is more enterprise-focused. I think we've built something on the IT side that I feel very, very good about.
We were talking about recruiting before I went off on a tangent. You now have offices, despite being a remote company for several years. Does remote not work, bluntly?
We were a fully remote company for 9 years until I took over, and we've now gone largely in person. We do have some people who work remotely, but we now have offices in New York—we took the old Pinterest space in San Francisco—London, Paris, Poland, and Washington, DC. We're just opening Austin, Texas, now.
The interesting thing I've experienced since then is that I do think you really struggle to build culture remotely in the same way. The thing I've experienced since we went remote is a much stronger, positive culture from co-location, which I think means people enjoy their work and get to know their coworkers a lot better as a result.
I think it gives us a lot more depth with customers to be co-located in cities where we spend time with them.
If I take London and Paris, we need to be co-located with the customers there. It can't just be someone on a Zoom screen in New York.
Do you see productivity increase exponentially?
Yes. I think if you take engineering as an example, you can execute engineering tasks remotely, but the process of working through really thorny problems is different. As an example, I've tripled the size of the engineering team this year. What I can tell you is that, interestingly, the vast majority of those people wanted to be in person.
I'm not saying that's true of all engineers, but it was interesting how many people, particularly the younger-tenured people, were saying, "I want to be co-located. I want to work through things." I don't even mandate office attendance; I just have it in those offices, and we have a huge appetite for it. We have 40 people in our London office. I was with many of them last night, and they were all commenting on how many of them come in voluntarily, even on a Friday when they might not need to, because they like being around their peers.
Look, I think I would actually bifurcate 2 separate things, and I don't think they're related. One is the hours you work 7 days a week, depending very flexibly on when client needs exist, and the other is physical co-location. I actually don't think they're related. If something comes up on a Saturday or you're pushing on a new product build, you will work on that Saturday. But if you do that from your home, that's totally fine.
If you took a hypothetical thought experiment and said, over a year, there is a diminishing return from being in the office all the time where you lose flexibility. If we were physically remote 100% of the time, that would not work at all. If we were physically in the office 6 days a week, I think that is overkill and you lose great people. Particularly senior enterprise folks don't want to be in the office on Saturdays.
I think what we've found is a nice balance: people come to the office most days, people really enjoy being with their colleagues, and they work most days, but they can do it from their own home on the weekends. I think that sort of flexibility is good.
Final one before you do a quick-fire. What did you believe about management that you now no longer believe?
I think there are 2 things I would highlight. One is that I think control is a bit of a fallacy, depending on the volume of things you have going on. To the question earlier about hiring great people, if you're serving several—let's say, a couple of years from now, you're serving a couple hundred customers on different topics—you actually need to have consistent values, consistent tooling, and consistent approaches, but you need to empower all those teams at the edge to operate and do what they will.
One of the big focuses I've had over the course of the year is reducing a lot of our hierarchy and making the organization much flatter, so that people at the edge serving customers are empowered to make decisions. They have decision-making frameworks and consistent tooling, but they are empowered. Trying to control that centrally may work in a manufacturing business, but you lose a lot in the latency of decision-making.
There is a lot of interesting military history that would say the same thing. If you look at the function of an army, at some point it moves into people in the field making the decisions. You have to have the training, the strategy, and the recruiting to do that, and then you have to empower your teams to work. I think about a lot of that very similarly.
The second thing I think a lot about is that, in the AI world at least, strategy is a somewhat overrated concept. I was talking to a CEO in the biotech space, and he was saying that strategy is very important to them because every time they make a capital decision, it's a 7-year capital cycle. In that case, strategy makes a lot of sense. But in the AI world, one thing that's been interesting to me is that every 3 months, the entire world changes.
I've just had to get very comfortable with that dynamic. You have to think about your investment life cycle as the core beliefs you have, and then 30% to 40% of things that you iterate constantly based on new technology. There is technology that you're going to build, like a new voice agent, that will become obsolete, and you have to be very comfortable that you're building an interoperable set of frameworks that you can integrate the new technology into.
That has to be a core function of the business. 5-year strategic planning is not a useful exercise right now in a lot of ways. I think you want to think about 5 years in terms of the cultural context you build and the institutional memory, to use the 7 Powers framework. But the actual iteration cycles are much, much faster, and if you don't react to the market and think about things quickly, that does not sustain.
The interesting flip side of that is that enterprise sales cycles, for example, are much longer, so it's not like you can't survive unless you're making a decision. But I do think the big thing is that a lot of the technology being developed changes every 2 to 3 months, and you need to be constantly incorporating that into what you build.
I think final final one, I promised, before a quick fire. You said about always being traveling, and you mentioned your girlfriend earlier. How do you make that work, and what would you advise me? What are the tips and tricks to not have a severely pissed-off girlfriend most of the time?
I think the first thing is to find a great girl who understands that you are really passionate and is supportive of that. I think my girlfriend, Claudia, has been great on that front. I'm very appreciative of that.
But look, it's tough. I'm on the road probably 60% of the time. If you look at my last 4 or 5 weeks: Riyadh, Geneva, Paris, Berlin, London, San Francisco, Boston, Singapore, and now London again.
You enjoy this?
I do, in some ways. I feel very lucky to be building something at this particular time and with a group of people I love working with. This happens to be what I spent my last decade doing, and it happens to suddenly now be what a lot of people want to do, which is great.
I feel very lucky because of that. Every day I wake up and see what else I can do to push that forward. I do live on the road, but I think some of the things I've tried are figuring out things like FaceTime and making sure you keep the cadence of interaction high, because being on the road is tough.
But I also don't think it's forever. I think I'm at that fun stage of trying to take something to where we kind of went from 0 to 1, and now we're trying to go from 1 to N, but we're not yet a fully mature public company or anything like that. I think she's been very understanding throughout that process.
Are you ready for a quick-fire round?
Yeah.
Okay. I think I should name it the discomfort round. OpenAI at $500 billion or Anthropic at $360 billion—which would you rather invest in?
I do not comment on any players in the model-builder space, for obvious reasons.
You can see why it's the discomfort round. What's the most underrated infra company today?
I'm going to say Databricks, which you're going to be like, "Well, they're very well rated," but look, I think their technology is great. In a lot of ways, the most useful foundation for AI is really good data and Databricks infrastructure. When I hear a customer has them, I'm always very happy.
What's the best advice that you've been given that you most frequently go back to?
We talked about this a little bit earlier, but a former CEO that I respect a lot—when I took the role, I asked him for his advice. I said, "What's the best way to think about a team?" And he said, "Look, your job as a CEO is to do 3 things really well: recruit great people, create a culture where they love working together and build great things, and try to make them all extremely rich."
I think it's a funny framework, but it's an interesting way to think about it. That is my responsibility to employees. I want to find great people, help them enjoy each other, and then build something that becomes big and helps all of them achieve their dreams.
What's one widely held belief about AI that you think is completely wrong?
That out-of-the-box agents will solve everything with the push of a button. I think that is the biggest misconception now. Many people were hoping the adoption curve would be, "I buy something, I push it into my business, and it takes a whole process and fixes it." I think they're realizing it requires training, fine-tuning, and a whole host of process redesign and business ownership.
You are me today. You have a new $400 million fund.
Yeah.
And you're a partner in the fund with me. Where should we be investing where most people are not? Because everyone is investing in agents out of the box.
Well, I think it's an interesting question because a lot of the reason people are investing in agents out of the box is that they're trying to apply a SaaS paradigm of what's worked historically to AI, which is challenging. I think the model-building layer is clearly producing amazing returns. I think the AI-agent layer is more complicated.
Now, where I think that's also complicated is that the application layer is tricky too. You hear a lot of commentary on this: many of the applications may or may not work. They're not really getting full workflow embedding; they're more of a nice-to-have in a workflow context.
My counterintuitive take would be that one interesting question of the paradigm now is whether new companies built around AI get distribution faster than big companies figure out how to adopt AI. I think that's the interesting paradigm for our society. Some of the most interesting new businesses are actual businesses using AI in the physical world that are AI-native and will be highly disruptive.
You mentioned Revolut and banking, for example, or you could go into loan servicing. There are many different areas where people are standing up new businesses. One of the most interesting stats I've heard recently is that, if you look at Y Combinator's recent class, I think it's the largest—it's 2x the revenue of any prior class. Many of those are businesses that are actually serving a customer need, not selling that customer software, if that makes sense.
From an adoption standpoint, one way to do this is to bet on AI agents, which are more of a SaaS paradigm and will sell stuff to customers. The other way to think about it is: what are the business models that will change because of this? I think there's a whole host of genAI-native services businesses, like tax accountancies, that are really interesting examples of that.
Again, you're a partner with me in the fund. Do we just get used to a world of lower margins? Is that how this business plays out? Is the world of 70–80% software margins over?
First of all, I'd challenge whether 70–80% software margins actually ever existed. What I mean by that is that there's the gross margin, and then, if you look at profitability in public software multiples, it's fascinating. In the last 2 years, you've seen public software multiples go from 20x to 10x, partly because of growth changes and partly because, as they've tried to move toward profitability, their growth slows materially.
What you realize is—I actually would take the flip side of this—the integrated units will be very, very profitable because of the way they grow. They'll be able to acquire customers faster and build them things that are good faster. They won't have the box stickiness, but they'll also—I would argue—a lot of those software companies, below the line, were not that profitable.
10. What Does No-One Know About the Future of AI That Everyone Should Know
When you look forward to the next 10 years, the final one: what are you most excited about? For me, my mother's got MS. I look at potential advancements in MS, drug discovery, and treatment pathways. What are you most excited for? I like to end on a tone of optimism.
I think that despite some of my realism on enterprise adoption, I actually am an AI optimist. I think the current narrative on some of the risks is far outweighed by some of the benefits. To give a couple of examples, including healthcare, if you take energy as an example, there's a lot of question around data-center implications for energy. But you do the math: right now, data centers are about 1% of total global electricity usage. AI data centers are 0.25 to 0.5% of that, so actually really small. Cooling—electric air conditioners—is 14 to 20% of global electricity usage. AI has so many different ways, with grid optimization and cooling, where the World Economic Forum just came out and said this: it's going to be massively net positive from an environmental-impact standpoint.
I think energy is one area where, if you think about all the energy needs we're going to have and the investment now going into clean energy because of all this, we'll actually be in a much better place 10 years from now. I think healthcare is another interesting one. If you look at U.S. healthcare, we spend $14,000 per capita per year on patients in the U.S. That's a rough spend. That's 2.5 to 3x what Germany and Canada spend, as an example.
If you then break down the context of that, roughly 9% of that is administrative, something like 25% of it is waste, and then actual cost of care is really challenging. Johns Hopkins just released this stat that 250,000 deaths a year happen because of avoidable errors. You see things in AI like 20% better identification of breast cancer risk, for example. I think healthcare is another area where the cost framework for healthcare has not been good over the last 20 years, and the cost-of-care improvements will be really material if AI works well.
I think the one I'm probably most excited about is education. If you're a kid growing up in any socioeconomically disadvantaged city in the world, your ability to learn about any topic on Earth incredibly quickly is better now than it has ever been at any point in history. You can take any topic on Earth and, with just an internet connection, learn. You can go through and pick your topic.
One of the reasons that's particularly important is that the educational system we've had for the last 50 years doesn't really work. We have massive K-through-12 challenges with STEM topics in the U.S., for example. We have huge learning gaps, largely by sociodemographic context, and most of our educational system is based around teaching people biology, English, and history, rather than teaching them basic things like FICO scores or how to code.
To add to all of that, the college system has created the student-debt crisis, where way too many people are going to colleges that are not worth going to and taking on enormous amounts of debt to do it. I think the way our educational system will function will shift materially. We're a talent-assessment company, and an enormous number of people we bring in did not go to college. We assess them on cognitive aptitude and skill.
The really positive note I would leave on is that I think the way people learn, the topics they learn, the way we look at resumes, and how to screen and assess people will move in a really positive direction. I think it will be a very different one from what we've had for the last 100 years.
Absolutely thrilled to hear that there is value in non-uni or dropouts. As a dropout myself, this has been so much fun to do, Matt. Thank you so much for being so flexible with the topic type. You've been fantastic, dude.
Thank you for having me.