Turing CEO Jonathan Siddharth:谁将赢得数据标注,以及为什么99%的知识工作会消失
- Jonathan Siddharth 的核心判断是:「数据标注公司时代已经结束,现在进入研究加速器时代。」 背后有3个变化:数据从简单走向复杂(从「写一个 Python 程序给数字排序」变成用 Kotlin、Swift 和 Next.js「写一个连接医生与患者的B2B市场应用」),训练目标从通过测试变成完成真实工作,聊天机器人变成需要 RL 环境的智能体;聊天机器人使用 SFT/RLHF,智能体则加入强化学习。Turing 正在为8家前沿实验室中的7家训练「超级智能」,如今则为每个角色、职能和行业的每条工作流构建「一个商业迷你世界模型」——「这相当于30万亿美元的知识工作」。
- 对于泡沫问题,他的答案是否定的:「我看不到 AI 泡沫……GPT-5 真他妈强,我们只是已经习惯了魔法。」 真正可交易的想法是模型能力过剩:「模型具备完成X的能力,但我们从模型中得到的是X减去delta」——价值的释放依靠脚手架、评测和部署,而不只是新一轮预训练。MIT「95%的试点失败」数据反映的是成长阵痛,不是不可逾越的高墙。
- 关于收入与 GMV 之争,Turing 将自身收入描述为「gap」或传统收入数字,并强调:「这不是 SaaS ARR……这是另一种生物。」 这些是靠表现持续下来的周期性实验室项目。集中度像 Nvidia 一样是设计使然:实验室为韧性起见,只与「少数几家」可信且隔离的合作伙伴合作;Harry 举例称,Nvidia 在5万亿美元市值时,39%的收入来自2个客户。Scale 的收购让 Turing「被需求淹没」。
- Harry 最强的反驳是:企业「距离采用 Slack 和 Notion 都还差得远,更别说构建定制模型」,这要等20年,不是10年。 Siddharth 的判断是:后台自动化很慢,前台会先动,因为「说服人们使用一项技术来多赚钱,比说服他们用来省钱容易得多」;OpenAI 的 GDPval 显示,在真实的单步任务上,最好的模型(很可能是 Claude 4 Opus 第一,GPT-5 也相当不错)约50%的时间能产出「与人类专家无法区分」的工作。
- 「我们所知的 SaaS,我认为已经结束了,彻底结束了。」 3条致命路径是:企业自己构建定制应用,基础模型对应用层「引发音爆」,以及随着环境式 AI 使用 MCP 和工具调用,GUI 的使用将消失。Harry 强烈反对:企业运行着80-100个 SaaS 产品,长尾公司「连 Wix 和 Squarespace 都用不明白」,而垂直化仍有防御力(在专利软件上,「Sam 不会去那里」)。Siddharth 的回应是审计自己的投资组合:如今的初创公司可能使用更少的 SaaS 应用,也雇用更少的人。
- 循环交易背后的 AGI 博弈论是:谁赢得超级智能,谁「很可能赢下搜索……消费设备……操作系统……云……社交网络」。 因此,Zuck 花费1000亿美元(12-18个月的自由现金流),去对冲2万亿至3万亿美元市值下行风险,是理性的。双方都同意:「必须下注。」
- 如果在自己的领域下注,他「可能会投机器人或具身 AI」。 垂直数据获取还处于第一局,但已不再是无人区;机器人数据则「完全开放」。他相信慢速起飞,而非快速起飞——这对世界更好,因为 AGI 不像自动驾驶那样,最后1%未解决就会失去实用性;AGI 每提升1个百分点,都会释放增量价值。主权模型也将到来:Harry 不认为德国医疗体系会运行在美国模型上,Siddharth 也认同,各国政府需要由本国人生成训练数据。
1. 数据标注已死——「现在进入研究加速器时代」
- Siddharth 一开始就重新定义 Turing:它不是人才市场,而是「我们在训练超级智能」。超级智能需要研究(由实验室内部完成)、算力(「我们要感谢 Jensen」)和数据——「Turing 掌管数据这一支柱」,服务8家前沿实验室中的7家。
- 第一个变化是数据从简单走向复杂:几年前,编程数据是「写一个 Python 程序给数字排序」;如今则是「写一个连接医生与患者的B2B市场应用」——Android 用 Kotlin,iOS 用 Swift,网页用 Next.js。「这已经不是低技能或中等技能承包商能够生成的数据。每个领域都需要专家型人类。」
- 第二个变化是从通过测试到完成工作——「重点不再是让 AI 通过律师资格考试,而是 AI 能不能做律师的工作」——无论是隐私律师、合规律师还是律师助理。第三个变化是从聊天机器人到智能体,数据因此彻底改变:聊天机器人使用 SFT 加 RLHF,智能体则在 RL 环境中加入强化学习。于是有了这句核心判断:「实验室希望与一个主动型合作伙伴合作,由它思考哪些数据可能有帮助。」
2. RL 环境是覆盖30万亿美元工作的迷你世界模型
- SDR 的例子是这样设计的:用合成数据库克隆 LinkedIn、Salesforce 和 ZoomInfo,提示智能体「准备与这个人通话……通话结束后更新 Salesforce」,再由验证器检查是否完成,同时让智能体尝试不同轨迹和工具调用。真正的技术在于设计课程——太简单或太难,「模型都学不到太多」。「这与 AlphaZero 掌握围棋时使用的技术非常相似。」
- 规模想象是一张4维矩阵——每个行业×每项职能×每个角色×每条工作流:「我们正在为每个行业的每项职能、每个角色和每条工作流创建 RL 环境。这相当于30万亿美元的知识工作。」
- 一位竞争对手的董事会成员告诉 Harry:「我们所有人都犯的最大错误,是没意识到垂直化数据获取还处于第一局。」Siddharth 回应:「完全正确。这就是第一局」——而整个 RL 环境体系诞生至今只有约12个月,起点是「o1 在12月发布,DeepSeek 在1月上线」。再过1年,「它可能已经完全不同」,这正是他说市场奖励研究基因的原因。
3. 企业现实:小型本地部署模型与「第一公里苦活」
- 保险核保的例子是:输入非结构化医疗数据,输出风险等级。你不需要一个万亿参数的世界模型——一个5亿至100亿参数、部署在本地、用10年专有核保判断微调的模型,更快、更准确,也能让数据远离前沿实验室。这是 Turing 与 Disney、Pepsi、BlackRock、Fiserv 和 Johnson & Johnson 开展的前线部署业务,规模小于实验室业务,但「增长相当快」;目前按时间收费,「我不认为这是正确的方式」,未来会转向价值定价。
- 第一公里的苦活在现实中是这样的:「我们的数据一团糟。它们分散在各个孤岛里……有些数据在 Bob 手里的一个文件中,而 Bob 已经不在这里工作了。」之后还要做评测,设计类似 Cursor 的部分自治界面,并培训人类适应新的工作流。
- 部署采用「人机并行系统」:人和 AI 做同一份工作,由经理比较输出;人的错误用于训练人,智能体的错误则成为下一轮微调的数据。Harry 问:「如果智能体是对的、人是错的,为什么不直接解雇人?」Siddharth 回答,要持续跟踪精确率和召回率——「你不会因为一次错误就解雇他们。」Harry 说:「有点狠。」
4. 采用之争:既有企业衰退与前台切入口
- Harry 的完整反驳是:企业「距离采用 Slack 和 Notion 都还差得远,更别说构建定制模型……也许要20年,但不可能在10年内实现」。Siddharth 的反击是:如果竞争对手用「1/100的人数」运营,同时还能提供更好的保险定价,「他们会把你的午餐吃掉」。Siddharth 将其延伸为「既有企业10年至20年的衰退……价值从老牌企业转移到初创公司——这就是我们投资的原因」。
- Siddharth 的判断是:后台自动化会很慢,前台会先动,尤其是金融服务,因为「说服人们使用一项技术来多赚钱,比说服他们用来省钱容易得多」。他引用 OpenAI 的 Mark Chen:金融服务是标普500的最前沿,但即便如此,「仍落后于最先进水平约2年」。
- Harry 转述 Rory O'Driscoll 的测试:AI 的价值取决于预算是否从人力转移给 AI 技术。Siddharth 认为,这种转移已经发生在客服、文案和 SEO 中——这些都是「失败风险较低的领域」。
- 他用 GDPval 支撑长期判断:OpenAI 的研究覆盖9个垂直领域和44种职业,参与者完成真实交付物;他回忆称,「约50%的时间里,最好的模型产出的工作与人类专家无法区分」。排名第一的模型很可能是 Claude 4 Opus(「要为 OpenAI 点赞」,因为它主动标出竞争对手),GPT-5 也相当不错——不过测试的是单步任务。「我们正在稳步走向 AI 最终自动化所有类型的知识工作。」
5. 收入是「另一种生物」——GAAP、隔离墙与 Nvidia 级集中度
- 对于 GMV 与收入之争,他不愿点名,但划出了自己的界线:「我们以 gap 收入的方式理解收入」,也就是传统收入数字。「这不是 SaaS ARR……这是另一种生物」——这些是有始有终的周期性实验室项目,需求「非常非常多」,但前提是「只要你一直做得好」。
- 信任是运营约束:项目在不同实验室之间、甚至同一实验室的不同团队之间都设置隔离墙——他的类比是 Foxconn 的生产楼层,一层生产 iPhone,另一层生产 Pixel。出于韧性考虑,实验室会刻意保留「少数几家」合作伙伴:「我们知道 Scale 投资发生时是什么情况。」
- 关于 Scale 的收购:「我们突然被大量需求淹没」,Turing 也「显著加码了多模态」,这是 Scale 从自动标注起家后最强的能力。Siddharth 最尊敬的竞争对手是 Alex Wang,因为他「很有先见之明地看到了数据的重要性」。
- 对于7家客户带来的集中度风险,Siddharth 说:「我认为我们和 Nvidia 处于同一条船上。」Harry 则补充,Nvidia 有39%的收入来自2个客户,约50%来自4个客户,「而这是一家5万亿美元的公司」。仅 Stargate 每年就在算力上投资1000亿美元,下一轮需求将来自主权模型:Harry 说,「我不认为德国医疗体系有任何可能与美国模型供应商合作」,Siddharth 表示认同,各国政府都希望由本国人生成 SFT 和 RL 数据。
6. 「我看不到 AI 泡沫」——能力过剩与被迫下注1000亿美元
- 他的立场十分明确:「这些模型今天已经极其强大……GPT-5 真他妈强。我想我们只是已经习惯了魔法」——而且「它们只会越来越好」。
- 他的核心机制是模型能力过剩:「模型具备完成X的能力,但我们从模型中得到的是X减去delta」;通过智能体脚手架、上下文工程和工具访问来弥合差距。Harry 每个周末花12小时、为每期节目挑选15-20个片段的现场演示,正是微调脚手架可以完成的工作。Harry 说:「如果你他妈能让它跑起来,老兄,我会给你很多钱。」
- MIT「95%的试点失败」的数据反映的是成长阵痛,而不是对技术的证伪:数据非结构化、没有脚手架、没有评测,也没有部分自治工作流;他引用 Karpathy 解释 Cursor 为什么有效。有些角色可以直接跳到完全自治,比如客服,因为互联网 tokens 已经够用;有些则不行:一家公司的融资方式可能与另一家不同,因此需要微调。
- 对于循环交易是否预示泡沫,如果你属于 AGI 阵营——他明确属于——那么超级智能的赢家「很可能赢下搜索……消费设备……操作系统……云……社交网络」。Harry 算了 Zuck 的账:花1000亿美元(12-18个月自由现金流)与所有人一起失败,没问题;不花钱、让别人赢,可能损失2万亿至3万亿美元市值。两人异口同声:「必须下注。」
7. 「我们所知的 SaaS 已经结束」——Harry 最强的反击
- 3条致命路径是:LLM 让应用构建变得极其简单,因此企业会自己构建软件;如果智能体模型充分接入企业数据库,企业可能被基础模型公司「音爆」,「中间不再需要任何东西」;GUI 也可能消失——「GUI 是为人类使用键盘和鼠标的世界设计的……人类可以把时间用在比点击更有价值的事情上」。环境式 AI 可能改用 MCP 和工具调用。
- Harry 展开反对:平均一家公司运行80-100个 SaaS 产品,没有人会在内部构建并维护这一切;长尾中的「每一家水管供应商、律所、会计师事务所……连 Wix 和 Squarespace 都用不明白」;垂直化也能提供防御力——在专利创作 AI 上,「Sam 不会去那里」。Siddharth 的经验性回应是:统计 Harry 自己的投资组合在 ChatGPT 出现前后的 SaaS 使用情况——「我的假设是,如今的公司使用更少的 SaaS 应用,也雇用更少的人。」
- 他认为其中一个护城河是数据驱动的反馈循环。PageRank 的配方 Google、Yahoo 和 Microsoft 都知道;Google 胜出,是因为用户偏好产生了具有代表性的查询与点击流——「这是一个高质量梯度,告诉你下一步该往哪个方向走」。企业领域仍然「完全开放」:先部署,再发现模型在哪里失效,最后生成数据填补缺口。
8. 自动化之后:智能即 API、更多工程师与数字替身
- 知识工作完全自动化会带来3个结果:100倍杠杆(Elon 也许能运营600家公司);创业潮——治疗师出身的创始人可以招募「一个营销 GPT、一个软件工程师 GPT、一个产品经理 GPT」,而不是融资几十万美元;以及「100万朵花将盛开」,创业机会远不止伦敦和帕洛阿尔托。
- Harry 的判断更阴暗:英国劳动人口中有650万人没有工作,而且「我们严重高估了普通人的智力」——这难道不会扩大鸿沟?Siddharth 则反过来认为,超级智能就是「智能即 API」;如果每月20美元就能获得,它将提供专家级智能,而昂贵的人类专家才是真正扩大鸿沟的因素。而且没有什么海滩生活:「我们是工具制造者……会在越来越高的抽象层级上解决问题」——治愈疾病、逆转衰老、「奔向星辰」。
- 10年后软件工程师会更多,而不是更少——定义将扩展到任何交付能解决真实问题的软件的人,比如他那位在 Stanford 的肿瘤学家,自己构建了一款居家诊断应用。在无限软件中发现目标,可以通过智能体之间互相对话解决,类似《Her》:Harry 可能同时与「100万名创业者对话」。
- 硬件会跟上:一款始终在线、带摄像头和耳机的可穿戴设备,会在耳边低语提示。「Harry 似乎没那么感兴趣……但我们谈到 AR 时,他立刻来了精神。」手机会退居次要位置:「手机应用是手机上最无趣的部分。」
9. 10年路线图——以及他改变看法的事情
- 市场结构将是「少数赢家」,而不是一家独大——实验室希望合作伙伴之间保持韧性与价格竞争,市场也会奖励研究深度,因为范式每年都在翻转。如果在自己的领域部署资本,他「可能会投机器人或具身 AI」:垂直数据仍处于第一局,但已不再是无人区(Turing 正在扩大规模);机器人数据则「完全开放」,家用机器人与工厂机器人需要不同的数据。
- 快问快答中的判断都保留了限定性:慢速起飞,而不是快速起飞——这反而是好事,因为 AGI 不像自动驾驶,最后1%未解决就会让产品失去实用性;「每提升1个百分点,都会释放增量价值」。关于中国,他所在的前沿圈子「不会低估」中国:DeepSeek、Kimi K2、Qwen 都是「最先进水平」。前沿模型「保留一部分技术不开放有一定价值」;企业会在5亿至100亿参数的小模型区间里混用开源与闭源模型。
- 他改变的看法是:过去他相信应该招聘「一支强大的高管团队」,然后放手让他们做;如今则转向 Elon 式的地面真相——走进工厂,追问「为什么 Model 3 的这扇门要用3颗螺栓,而不是2颗」。他也承认:「Turing 创业初期……我可能潜意识里想要被人喜欢。」
- 他最不受欢迎的决定,是从分布式团队转向中心辐射式架构——旧金山、帕洛阿尔托,以及即将开设的伦敦办公室;「其中一些人离开了」。Harry 分享说,他的母亲患有 MS,相信 MS 药物研发会出现突破。Siddharth 说,最让他兴奋的是「将 AI 研究本身自动化」,形成自我改进闭环。结尾的画面来自《钢铁侠》的智能无人机战衣:「今天 Harry 可能有100个想法,但 Harry 也许只能把其中2个做好。我喜欢这样一个未来:Harry 能完成剩下的98个。」
Harry, I've been so looking forward to this. Thank you so much for joining me in person. It's such a treat to do it in person while you're in London.
Thank you for having me, Jonathan.
So, I think of a talent marketplace as something that's basically matching talent to something. Maybe it's an opportunity. Turing is not a talent marketplace. What we do at Turing is train superintelligence. We work with 7 of the 8 frontier labs.
To get to superintelligence, you need research, compute, and data. Research, the labs do in-house—with OpenAI, Anthropic, DeepMind, et cetera. For compute, we have Jensen Huang to thank, and maybe NVIDIA as well. On the data side, Turing powers the data pillar.
1. Redefining “Talent Marketplaces” Today
Now, I want to start with a little bit of definitions, because everyone thinks they're talent marketplaces, and then everyone pushes back on talent marketplaces. How do you describe it, and why are we not dealing with talent marketplaces anymore?
On the data side, there's been a significant shift in the last couple of years. A few years back, the models weren't quite smart enough, and as the models have gotten increasingly smarter, the data needed to improve them has become harder to generate. This is because more sophisticated data is required to improve the models. It's vertically specific people, tasks, and workflows that aren't so obvious, like cat pictures.
That's correct. That's correct.
There's a shift in the data from simple to complex. Take coding, for example. A few years ago, the kind of data set a contractor could generate might look like, "Hey, write a Python program to sort some numbers." Today, the data that's generated might be, "Write a B2B marketplace app that connects doctors with patients, and write it for Android with Kotlin and Java, write it for iOS with Swift, and write it on the web with Next.js or something," right? That's the complexity.
So, there's a shift from simple to complex. It's no longer the kind of data that low-skilled, medium-skilled contractors can generate. You need expert humans in every domain.
Yeah.
The second shift is we've gone from teaching AI to take tests and pass tests to teaching AI to do real work. It's less about having AI pass the bar; it's more about, can AI do the job of a lawyer? Can it do the job of a privacy lawyer, a compliance lawyer, or a paralegal? It's about having AI be good at doing economically valuable work. So, that's a shift.
The third shift is we've gone from chatbots to agents, right? We started off with ChatGPT, where you're asking questions and getting answers, which is great, but now it's about the models becoming agentic, where they can execute complex, multistep workflows in a real-world business setting. The type of data you need for that is totally different.
2. Data, Compute, Algorithms: What is Most Abundant?
How is that different? That's so interesting. In the transition from chatbots to agents, how does the data required change with that transition?
When you're training a chatbot, you'd usually do a lot of SFT and RLHF. With SFT, you're giving the model input prompts and output completions. You teach the model to imitate experts. With RLHF, you're basically teaching the model to produce responses that a human would tend to prefer. RLHF is used to train what's called a reward model, and then the model is trying to produce completions that give it a high reward.
With agents—and let's define an agent; different people define agents in different ways—I would define an agent as something that's capable of taking action in the real world or in the physical world, something that's executing a multistep workflow and calling different functions. The agent could be operating a computer or making backend API calls to actually do stuff, right? You might have an agent to file your taxes. You might have an agent to prepare your monthly financials.
To train an agent, you would also want to teach the model how to do tool use. So, you teach the model how to call other functions and how to use other applications to be more leveraged. Today, the dominant paradigm is reinforcement learning. Oftentimes, these agents are trained through reinforcement learning, where you'd build what's called an RL environment, which is like a mini world model for business.
Yeah.
In that RL environment, you have input prompts and output verifiers, and you'd have the full system state tracked along with the data model.
Let me give you an example. Imagine a workflow for a salesperson that an SDR would go through, where, before a sales call, the salesperson has to research the prospect, look up Salesforce to see whether somebody from the team has already spoken with this human, and maybe, if needed, look up this person's contact information, perhaps using ZoomInfo or something like that, to reach out to them. This required this human to use 3 different tools: LinkedIn, Salesforce, and ZoomInfo, right?
In an RL environment setup, you'd create a mini world model with clones of these applications that are created with a fake database and synthetic data. The prompt might be, "Hey, prepare for a call with this person and then, after the call is done, update Salesforce." Let's say that's the prompt, and you have what's called a verifier to check whether the agent completed the task.
This is where I think AI is kind of beautiful and somewhat magical: you set up the agents in this environment, and the agent is going to try different trajectories and different tool calls to try to complete the task.
You would set this up so that the curriculum is optimally defined. The curriculum is the set of tasks that you have this agent do. If it's too easy and the agent completes everything, the model doesn't learn much. If it's too difficult, the model doesn't learn much. Ideally, you'd want the right mix, where the model is getting positive and negative feedback.
It's very similar to the technique that AlphaZero used in mastering Go, when the model played against itself. So, this is another—it's kind of like a form of synthetic data because the agent is trying different approaches by itself—but it's humans. In this case, at Turing, we create these RL environments at massive scale for every workflow you can think of, across every function, across every industry.
So, you create the RL environments that create the data that then allow the models to train to have further use cases like that?
Correct.
Correct.
We create RL environments for every industry you can think of: retail, healthcare, and life sciences. Imagine this 4-dimensional matrix, where the first dimension is every industry—financial services, retail, healthcare, podcasting maybe is one of the dimensions.
Please.
The second dimension could be every function: software engineering, marketing, sales, finance, et cetera. The third dimension could be every role in that org chart. Let's say, in sales, an SDR was a role. The fourth dimension is a workflow that a human goes through in that role.
You can think of every role a human has as a composite of workflows, right? We are creating RL environments for every workflow, for every role, in every function, in every industry. That's $30 trillion of knowledge work.
Is it possible to have that breadth and quality?
Yes.
How? Time and lots of money? Because I was speaking to, candidly, one of your competitor's board members the other day in prep for this, and he said, "The big thing that we all got wrong was we are so in innings one of the acquisition of verticalized data. There is so much room to run in the data acquisition for dental, for SDRs, for product managers—you name whatever function you want."
Do you see us very much in innings one of the data acquisition for these very specific, vertically focused workflows?
Absolutely. It's innings one, and I believe in slow takeoff. I'm sorry to pour cold water on all the AI doomers that might be listening to this, but we are not in a rapid takeoff scenario. I believe in slow, steady takeoff for AGI and eventually superintelligence.
So, we're still in innings one. It's going to take a while before we get all of this data into the models.
When we think about the breadth that we go after and your specializing in RL environments, just so I understand the marketplace that we sit in, because there's Mercor, there's Surge—how do you differ from those, for people who are wondering, "Hang on a minute, I thought they were all one?"
Turing is a fundamentally different animal. What we do is we're training superintelligence for all these frontier labs. To get to superintelligence requires research, compute, and data.
The data needs have significantly changed. It's more complex data rather than simple data. It's more real-world data—data that touches how real humans do knowledge work. You need data to train these agentic systems, right?
What the labs need in a partner in this new world is somebody that has research DNA, who could be a proactive research partner for them, because these paradigms keep changing. Last year at this time, we were not talking about reinforcement learning at all, but then 2 things happened.
o1 dropped in December. DeepSeek launched in January. And now it’s all about RL environments. It’s not just imitation learning; it’s also reinforcement learning. So the labs need a data partner that’s more research-oriented.
Second, the labs need a data partner that also touches the real world. At Turing, we don’t just generate data to train the models for the frontier labs. We also work with enterprises. We work with Disney, Pepsi, BlackRock, Fiserv, and Johnson & Johnson to build fine-tuned custom models to solve real-world enterprise problems for those enterprises.
So this is like FDEs that you send in to go and build custom models?
Correct. So we touch reality. We know where the models break in the real world.
How much of the business is that—FDEs deploying custom models—versus more horizontal?
The horizontal business is bigger, but this business is also fast-growing.
And, third, you need a platform with the world’s smartest humans on it, as well as experts in different domains, so that you can build these RL environments. You need that platform to be really good at sourcing talent, vetting talent, matching talent, and generating data.
I think the era of data-labeling companies is over. Turing is a research accelerator, and it’s now the era of research accelerators. The labs want to work with a proactive partner that can think about what types of data are likely to be helpful for these models and can make recommendations to them.
Why would you need a custom model? When you look at a lot of the customers that you mentioned, for the ones where you have FDEs who go in and build custom models, what is the reasoning around that? And is that a temporary moment in time, or is that a permanent requirement from them for a certain reason?
I think it’s a permanent requirement. I’ll give you an example. Let’s pick an insurance company. For insurance companies, 2 really important problems they have to solve are underwriting and claims processing. Let’s pick underwriting, for example.
With underwriting, the problem statement is that you might get multiple types of unstructured medical data. It could be somebody taking a picture of their medical history on their smartphone, or it could be some OCR data from somebody’s medical history, or data in PDFs, et cetera. A human has to look at that person’s medical information and then decide: Is this person high-risk, medium-risk, or low-risk? What medical conditions do they have? Do they have cardiovascular issues? Do they have renal issues? And how do you price insurance for somebody like this? Do you even take them on as a client if you’re an insurance company?
Right now, this is a problem that an LLM can solve really well with a human-in-the-loop system. You may not need a trillion-parameter world model to do a task like this. In fact, there’s lots of research that shows a smaller language model will actually be faster and more accurate at a task like this than a giant world model.
The insurance company also may not want its data to go back to a frontier model. Oftentimes, in these cases, what we would do is work with a frontier lab, take one of their smaller models—maybe something in the half-a-billion-parameter to 10-billion-parameter regime—and have an AI system built that’s on-prem with the customer, fine-tuned on that enterprise’s proprietary data.
This insurance company might have data from the last decade of humans making judgments. You want to make use of that data, but you don’t want to help other competing insurance companies with your own data. Normally, for these cases, you would have a smaller, fine-tuned model that’s trained on your proprietary data, distilling your proprietary human knowledge into the models.
That human underwriter who’s been doing this job has a lot of institutional knowledge in their brain.
Sure. You’d want to distill that into the LLMs.
And this human that I mentioned who’s doing this job of underwriting might be using other internal tools inside this insurance company. You might want to automate those tool calls in the agent. So you’d basically build a version of almost like a ChatGPT agent that’s smaller and fine-tuned for that specific workflow and use case.
I do think this will become more popular across the board, and it’ll be a big market for the frontier labs. If you want a general-purpose assistant, I think you need a trillion-parameter model, right? One that can answer anything to be a universal assistant.
Is that a good business for you when you think about taking someone else’s model, retrofitting it to a business, and doing a lot of custom work with your own engineering teams in those businesses? Is that a good business?
I think it would be a good business. We’re still early, but it’s growing pretty fast.
We have this unique vantage point because, by generating data for all the frontier labs, we get to see a glimpse of the future before it arrives. The glimpse of the future that I see is that all knowledge work is going to be automated.
If a human’s job involves looking at a computer, analyzing what’s on the screen, using different tools, and using a keyboard and a mouse, it’s going to be automated. It’s only a matter of time. These computer-use agents are going to keep improving over the next decade. And that’s $30 trillion of digital knowledge work.
3. The Biggest Challenges Enterprises Have with AI Adoption
My question to you is this: I spend a lot of time with very large companies, mostly when I speak to them. The thing that astounds me is that we hugely underestimate the pace of AI progression in terms of technological capabilities, but consistently, what I see is the laughable state of internal data and internal processes.
I mean, Jonathan, these guys are so far off adopting Slack and Notion, let alone building custom models and embracing the latest AI tools. I respectfully push back on all knowledge work being automated—maybe in 20 years, but not in a 10-year time frame. Am I wrong?
What do you think is the biggest constraint or obstacle?
The inability for them to try and implement new tooling.
But what if the cost is too high? If they didn’t do that—if the hypothetical insurance company that I told you about had a competitor that could operate with one-hundredth the headcount—
While delivering a better experience to their customers by pricing insurance deals better, making more money from insurance premiums, and having fewer claims payouts, they’ll get their lunch eaten.
I think they will. I think you’ll see this transfer of value from an old incumbent that can’t adopt new tools to a startup company—hence why we invest—which is eating their lunch. Absolutely.
Ah, I see. So your theory is that the incumbents won’t adapt and it’ll just be a forest fire?
100%. We will be on a 10- to 20-year decline of incumbents that are unable and unwilling to adopt new tools because of data, because of permissioning, and because of internal buying processes.
That’s an interesting point, Harry.
I mean, I go to a European bank, and you will be astounded by how bad it is internally. Respectfully, the poor quality of the processes for buying technology is just abhorrent.
I have a hypothesis. My hypothesis is that companies will be very slow with back-office automation. But in the front office—for example, I speak with financial services clients in New York, some of the biggest companies—I speak with a suite of these companies, and they are extremely interested in applying AI to help them make better investment decisions because it directly translates into helping them make more money.
I’ve found it’s a lot easier to convince people to use a piece of technology to make more money than to save money.
Yeah, I agree with that.
And in financial services, it’s pretty brutal, right? It’s kind of an efficient market. If there is alpha to be found in how you can allocate capital better, make investment decisions better, figure out what opportunities to invest in, or price deals better, you’ll get killed if you’re not at the bleeding edge.
I’ve heard Mark Chen, the head of research at OpenAI, say this about how financial services is usually at the bleeding edge among all the other industries in the S&P 500. But even they are about 2 years behind, usually, relative to the state of the art.
So I agree with you that in back-office automation, it’ll probably be very slow, and it’ll probably be the upstarts that do things well. I think the change management will be too slow. But I’m optimistic about the front office, especially in financial services, life sciences, and pharma, where, if you can accelerate the time to discover a drug or to get to a molecule—if you can help somebody win in the main thing that they care about in their industry—I think they’ll adopt it faster.
4. Why Will 99% of Knowledge Work Will be Gone in 10 Years
I’m always told by a dear friend, Rory O’Driscoll at Scale—I don’t know if you know Rory, but he’s a fantastic investor—and he always says to me, “Listen, value generation from AI is fundamentally dependent on one simple question: Will we see the transfer of budget from human labor to AI technology?”
If we see that transfer of budget, oh my God, that’s the $30 trillion that you said it is. And if we don’t, we operate in maybe a slightly larger software technology budget world, but by no means a world where we can have the valuations and the money that we have going in.
When you look at that, are there any areas truly today where you’re like, “We have seen the full transition from human labor budgets to AI technology budgets”?
I think the transfer is pretty high in areas like customer support, copywriting, and SEO—some of these marketing-related areas, as you would expect. The transfer is faster in these low-risk-to-fail areas, where it’s relatively easy.
I’d encourage your listeners to look up GDPval, which is this paper by OpenAI where they measured the impact of today’s AI models in automating all types of economically valuable work.
It’s a lovely piece of research. I’d encourage everybody to read it. They did this study where they looked at, I think, 9 verticals and 44 occupations. They took a very diverse sampling of different types of knowledge work, everything from financial services to real estate, healthcare, and law.
They took very specific occupations, and in those specific occupations they took real tasks where a real deliverable has to be produced. Imagine an engineer, like a civil engineer, creating a blueprint for a building they’re about to build, or somebody who’s on the set of a movie studio coming up with a schedule for how you organize your crews. That’s real work. For coding, you can imagine a real-world software engineering project.
They saw that today’s models were quite good at achieving parity with the best human experts in that field. What I noticed was that about 50% of the time in GDPval, the best models were producing work that was indistinguishable from a human expert, which is remarkable. I say “we”—I don’t know which side I’m on. Am I on the side of the AIs, or are the humans?
I think you’re on the side of the AIs.
Oh, really? I mean—
From that positioning, that would infer so.
Yes. But Turing is a blurry line, right? Maybe—as the CEO of Turing, passing the Turing test is about not being able to tell the difference.
What I noticed was that about 50% of the time in GDPval, the best models were producing work that was indistinguishable from a human expert, which is remarkable. Kudos to OpenAI; they also flagged that the number 1 model was Claude Opus 4, although GPT-5 was quite good also.
This was for relatively simple tasks. Imagine a task requiring a single step, whereas in the real world, if I give you a certain project to do, you won’t just go off and do it. You might ask for clarifying information, do other things to acquire more context, or brainstorm with other humans to complete that task. You would do it in a sequence of steps, so there’s more room to go.
But I think we are well on our way to AI eventually automating all types of knowledge work.
What happens in that world? If AI automates all types of knowledge work, what happens then?
Three things will happen. First, I think we will all have the potential to be 100 times more productive. Today, I’m able to run 1 company. Elon can run maybe 5 companies. But in a world where I’m 100 times more productive, maybe I’m able to run 100 companies. Elon maybe runs 600 companies.
I think every human will just be so much more leveraged. The nature of a job itself could change. Today, we are accustomed to the idea of 1 person doing 1 job, but people could be doing multiple jobs at the same time. People could be running different companies at the same time.
The second implication, I think, is that it’s going to be wonderful for entrepreneurship. You, Harry, are going to be very happy, because today, for a lot of ideas, founders are intelligence-constrained. I think of being capital-constrained as a form of being intelligence-constrained.
For example, if you pick a therapist who wants to start a mental health startup, today that founder would have to raise at least a few hundred thousand, if not a few million, to recruit some software engineers, maybe a marketing person or a growth person, and perhaps a product manager. But in a future where AGI exists, this person will recruit a marketing GPT, a software engineer GPT, and a PM GPT, and get off the ground for a lot less capital.
A million flowers will bloom. Lots and lots of nontechnical founders will start companies. We’ll see a broader distribution of founders than just those who live in London or Palo Alto, who are connected to these pools of capital and might start companies, which I think is wonderful for the world.
Do you think we will? What I mean by that is, there are 6.5 million people today in the UK’s working population who actively do not work because of an inability to work. I’m not going to get into the analysis around that because I’ll get in trouble for it.
I think we grossly overestimate the intelligence of the general population. I know that sounds incredibly arrogant, but most people—actually, a lot of people—just don’t want to work and are not at the level of recruiting GPT assistants. Do you not worry that it will widen the chasm between those that have and those that haven’t?
I’m an optimist, and I think the opposite will happen, because what we’re really doing when we’re training superintelligence is basically training intelligence as an API. What’s the alternative to that? It’s hiring a human to provide you with that intelligence, and that human is quite expensive, right? If anything, that creates an even broader gap between the haves and the have-nots.
Whereas, for $20 a month, if you had access to the smartest experts in coding, STEM, sales, and marketing, I feel like more people would be able to start companies and produce actually valuable work.
I believe that when we have access to superintelligence—we are not all going to chill out on a beach somewhere and contemplate what to do next. I firmly believe this. We humans are tool builders and problem solvers. We’ll solve problems at higher and higher levels of abstraction.
I feel like, in a world where we have AGI, we’ll just solve much more exciting problems. Maybe we’ll cure diseases, reverse aging, or go to the stars. There are all sorts of fun things we’ll do. I don’t think we’ll be bored.
I’m glad. I often hear about UBI and how we’re going to sit and write poetry. I’m like, I think that might be a little bit challenging.
5. How Will Data-Driven Feedback Loops Replace Technology as the Moat
When technology is not the moat, what is the moat? I had the founder of Base44 on, and he said 99% of code in the next year will be written by AI. Technology is no longer the moat. What is the moat in that world?
I think one moat will be data-driven feedback loops. For example, one reason Google had such a great lead in search for a while was these data-driven feedback loops that come from people using your product and generating data that gives you, the algorithm developer, a high-quality gradient for which direction to step in.
PageRank—the importance of PageRank and the recipe for ranking search results—was well known among Google, Yahoo, Microsoft, and a few others. Obviously, people move around these companies all the time, but the advantage Google had was that, because everybody preferred Google and liked that search engine, you saw a much more representative set of queries.
You had data from the clickstream of what results people were clicking on, and that helps your algorithms improve at a much faster rate. I think data-driven feedback loops will be key for all types of enterprise applications as well.
Today, OpenAI and ChatGPT have a good data-driven feedback loop in enterprises. Again, I think it’s wide open. Whoever is deploying the right custom, fine-tuned models and agents for specific workflows, roles, functions, or companies—if you get in first and solve a customer’s problem really well, you start getting that flywheel going.
You will discover first where the models don’t work well, and you will use that data to work with a company like Turing to generate additional data to plug that gap. Then you will improve. This is what I mean by why it’s important for the models to touch reality.
I feel like the models have touched reality in consumer. We haven’t yet touched reality in enterprise, and the only way we’ll improve is through deployment.
And that deployment is fundamentally predicated on handholding, correct?
Yes, handholding. I feel like there’s still a lot of first-mile schlep and last-mile schlep that needs to be handled.
What does that mean, first-mile and last-mile schlep?
When I say first-mile schlep, I mean, for that underwriting copilot example I gave you for that insurance company, I painted a pretty rosy picture of how you take this model and fine-tune it on your proprietary underwriting data. In the real world, it doesn’t work that way.
Let’s say I’m talking to the CEO of this insurance company or the CTO of this insurance company. They’ll say, “Our data is a mess. It’s in silos. It’s super fragmented. Some of the data is in spreadsheets. Some of the data is in a file that Bob has, and Bob doesn’t work here anymore.” The data is all over the place.
You first have to acquire the data, convert the unstructured data into structured data, and put it into a format to fine-tune LLMs. You might want to set up good infrastructure for evals. You’d want to create good evals for the models or agents. You might want to build a workflow designed for partial autonomy.
This human underwriter who’s about to use this model to evaluate these medical histories might need a Cursor-like interface so they can work alongside the AI to do their job. You also need to train the humans in these new workflows.
You want to make sure you’re collecting data the right way. The way, for example, we do deployments is with a tandem system, where you’d have a human and an AI doing the same job for a period of time, and a manager can see the output of both.
If the agent is right and the human is wrong, you train the human. If the human is right and the agent is wrong, you’ve created a data point to fine-tune the next iteration of the agent. The agent is steadily improving over time.
If the agent is right and the human is wrong, why don’t you just fire the human?
It depends on what frequency, right? You track things like precision and recall. You’d want to analyze this over a period of time. You wouldn’t fire them over a single mistake.
Bit harsh.
Yes.
What’s the margin on that business?
It varies. We’re also in the early innings of figuring out how to price that. Today, we do it in a relatively simple way where we’re just building these things for time. I don’t think that’s the right way to do it. We’ll switch to a more value-oriented pricing model at some point.
Right now, we’re just laser-focused on the frontier AI labs. Enterprises for us are a longer-term play.
6. Is Revenue BS in Data Labelling? Are Players Calling GMV Revenue?
When you look at revenue numbers in this space, a lot of people shout back, “They’re not revenue numbers; they’re GMV.” Given our understanding now that, Jonathan, there’s no talent acquisition from your business—it’s all an RL environment-creation business—when you look at the other announcements from alternative providers, can you help me understand: are they revenue or are they GMV? Is there mislabeling being done here?
I don’t want to comment on other companies, but I’m an investor, Jonathan. Essentially, I’m trying to understand and learn from you how I should weight revenue in today’s AI world versus the previous historical world. Should I be impressed by these revenue numbers, or should I not?
I think it depends on the type of revenue. Obviously, these are not SaaS ARR numbers, right? These are not those types of revenues. This is a different beast, and I think this requires thinking from first principles.
The revenue here is recurring in the sense that, oftentimes, when you’re working with a lab on helping the models improve in some area—and I’ll speak to Turing; I don’t want to speak to other companies—when we’re helping a lab, let’s say, improve their models for coding, multimodality, or tool use, or working on RL environments for automating all types of professional knowledge work, it’s usually a recurring project where projects will start and projects will end. As long as you’re doing a good job, there’s lots and lots of demand.
But you have to consistently keep doing a good job. It’s also important to be a trustworthy partner to the labs. We take secrecy very seriously. We make sure that our projects are all firewalled between labs, oftentimes even with teams within the labs. Sometimes, that’s the level of secrecy that you would need.
I’m reminded a little bit of how I’ve been told Foxconn operates. I don’t know anything about that, but I’ve been told that they have different floors where maybe on one floor the iPhone is getting made, and on another floor, maybe a Pixel phone is getting made. Obviously, you have to firewall all of that.
Of the 8 largest providers, do they not spend with all of you?
They spend with a handful of companies. They do that to have some level of resilience. I imagine there are some price benefits to having more than 1 person they could work with, but I think the resilience piece is important.
We know what happened when the Scale investment happened. Again, the labs did benefit from having other partners that they could work with. I would say it’s a small handful—trusted companies. Of course, there’s probably a giant pool of smaller startups, but it’s a small handful of big companies in the space.
Which one do you worry about most?
This is just a big, big market that’s growing super fast. I’m excited about all the companies in the space. I feel like different companies come into this world with a different DNA.
Which leader do you most respect? Sam Altman, Elon Musk—of the data providers, Jonathan? Of the data providers. I’m pushing you, dude. I’m going to get a name.
I have a lot of respect for Alexandr Wang from Scale AI. I feel like Alex and Scale were prescient in seeing the importance of data. I admired how, having started in autonomous labeling, they navigated the ups and downs.
I really like the way he operates as well. I feel like there are certain elements of leadership that I think I share with him. I think he did a great job for Scale.
How did Scale being acquired impact Turing’s business?
We just got flooded with a lot of demand. We’ve also amped up significantly in multimodality. Multimodality was something I think Scale was quite strong in.
Multimodality is teaching the models to operate well with not just text but audio, video, images, et cetera. From the outside in, I’ve heard that because of their roots in autonomous labeling, they were quite good in multimodal work. It was good primarily from just increasing demand, and I feel like they were the company that had been working in the space the longest.
Do they have a business left? Again, I mentioned Rory O’Driscoll. I think he said on a show with me recently that there’s this kind of carcass or husk left behind. But what if everyone benefited from their being bought? They can’t be doing that well.
I don’t know enough about their business.
Do you pay attention to competitors?
I pay attention to competitors in terms of the things that they do well and when there are any significant learning opportunities from them that could help us serve our customers better.
Do you worry about revenue concentration? You said about 8 of the biggest labs. Say, if you look at OpenAI, they have—I don’t know, whatever it is—100 million. You all know these numbers much better than me, but say 100 million paying customers. I’m just taking 10% of a billion people, but give or take 100 million, whatever.
Then you look at a business like yours here, where there are 7 core customers. How do we feel about revenue concentration? The last time I checked, I was told that Nvidia has 39% of its revenue coming from 2 clients, right? Roughly 50% was 4 clients.
Which is extraordinary.
Extraordinary, right, for a $5 trillion company—
Wow, that’s nuts.
Yeah. You think about Apple as not even comparable in market cap size, but what is that—like 3 billion customers? I mean, wow, what a comparison.
Yes. So, in this market, I don’t worry that much about revenue concentration. I think we’re in the same boat as Nvidia, in that there will be lots and lots of spend from these big 8 companies. I expect this market to— I mean, look at the scale of the spend.
Stargate is like a $100 billion-a-year investment in compute, and there’s going to be significant amounts of spending on compute, energy, and data. It’s a little weird to have this level of concentration, but things could change. I think it’s also possible that governments spend even more.
I think it would make sense for governments to build their own internal versions of some of these models, which would require proprietary data, again, to be collected.
Do you not think they’ll have to? What I mean by that is, we’ll see sovereignty of models. I do not think there’s any way you’ll have the German healthcare system working with American model providers. Sorry to say.
I think you’re right. I think it’ll be necessary.
In that world, do you provide that FTE service to governments?
Yes. I think the work that we’re doing, not just in training superintelligence but in deploying superintelligence, is with that goal in mind for these governments. I imagine we’ll help them not just with fine-tuning their models with data and evals, but also helping them with the first-mile schlep and the last-mile schlep to make these systems actually work.
You might want to have full control over what type of human data is going into these models. If you’re the German government, presumably you want German nationals to be the ones contributing data, whether it’s for SFT or reinforcement learning.
7. Are We in an AI Bubble?
I’m worried that we are potentially not going to see AI deliver the immediate revenues that we’ve promised, and we’re going to go through a kind of cooling period—which everyone suggests and thinks we’re going to go through in the next 6 to 18 months—in which, as I said, it doesn’t hit the revenues that we said it would, and the AI bubble kind of deflates slowly.
To what extent do you think that’s possible, or will we see this continuing gradual increase, as we touched on there?
I don’t see an AI bubble. I feel like these models are incredibly powerful today. GPT-5 is fucking awesome. I don’t know what people were talking about. I know there was some chatter.
I think we’ve just gotten used to magic. I feel like these models are incredibly powerful today, and they’re the worst they’ll ever be. They’re only going to keep improving. I say that about the Gemini Pro models, the Grok models, and the Claude models. These models are amazing.
There’s a very significant model-capability overhang. By that, what I mean is, the models are capable of X, but what we are getting out of the models is X minus delta. With the right agentic scaffold around these models—in terms of the right system prompts, the right user prompts, and giving the models access to the right context…
Teaching the models how to acquire additional context, teaching them how to use the right internal tools—there is a significant amount of capability that can be unlocked with today's models. For example, Harry, I imagine that when you do an interview with somebody, one of the things you probably do is apply your secret sauce to pull out the right clips from the interviews: what to highlight, what the catchphrases are, and what will drive more engagement. That can be done by a model with the right agentic scaffold, fine-tuned on all the work that you've done in the past.
Every weekend, I go through every single show and pick out 15 to 20 clips per show, and then I make notes on each one.
Yeah. Are you saying, Harry, that you want to use Turing?
If you could make it work, dude, I'd pay you a lot of money.
Yeah, maybe we should partner.
That'd be great. Seriously, every weekend I spend probably 3 hours per show. Three—definitely 12 hours a weekend—doing that.
So I think there is this model capability overhang where the full potential of the model has not been unlocked by humans yet. No, I don't think there's an AI bubble. I think there are some growing pains.
What are the growing pains? I think everybody keeps citing that MIT report about how 95% of pilots fail.
Because we are in the business of deploying AGI in enterprises, I can tell you why I think that happens. One of the growing pains is that, step 1, most enterprises need to do some work to structure their data in the right way. Again, that first-mile schlep has to be done. Second, you should surround the model with the right agentic scaffold that I just described: the right prompting, the right context engineering, and the right internal tool calls.
All of those have to be distilled into the models. You need really good evals. You also need a workflow designed for partial autonomy. Andrej Karpathy articulated this first when talking about why Cursor works so well, because it's not designed for full autonomy. It's designed today for partial autonomy, for humans to collaborate with the AI to do that specific task. So that Cursor for X needs to be built for every role, for every workflow, to help humans work more easily with the models.
Does every role need to go through that pathway of Cursor for X before it goes to full autonomy, or are there some roles, like customer service, where it just goes to full autonomy?
I think for some roles where you can see that the models are quite good at matching humans, we don't need that intermediate step. There are certain roles where, by virtue of how the models are trained—pre-trained with tokens on the internet and then, of course, with talent from research accelerators like Turing that's fine-tuning the models—the tokens from the internet give them sufficient intelligence to do the job well. Customer support is an example.
But if you pick other roles—for example, if you picked the role of an AI researcher, or you picked the role of a lawyer specializing in venture-capital financing—it's possible there aren't enough of those tokens on the internet. So the models will be relatively weak there out of the box. Also, the way one firm does financing might look different from another's. Maybe they have their own way.
You might want to fine-tune them on your own proprietary data and distill the proprietary intelligence of humans working there. So, for those things, you may need to do some fine-tuning; the models may not work very well out of the box.
A lot of people suggest the circular deals between some of the large providers suggest the strains in the ecosystem or the bubble-like tendencies. Do you think that's fair or not?
I've categorized the world into 2 classes. Class 1 is those that believe in AGI. Let's call it the AGI-pilled group that believes we are on the path to getting to AGI, and let's define AGI as an AI system capable of at least matching humans in almost all types of intellectual knowledge work.
Correct. Right.
Then there is another category of people that don't believe this will happen and think we'll hit a wall, right? In the past, there have been other AI paradigms where we did hit a wall. For the camp that believes in AGI—and I believe in AGI, unsurprisingly, because I love AI and it's been my passion for the last 20 years—if you believe that, the grand prize is so amazing.
If you've solved intelligence, you've solved all of humanity's grandest problems, from curing diseases to potentially pausing aging, to interstellar travel, to energy. All of our problems are intelligence-constrained, right? The prize is so large. Whoever wins the superintelligence race will probably win search, will probably win consumer devices, will probably win operating systems, will probably win cloud, and will probably win business productivity software. It's like the prize is so massive that it's worth placing big forward bets in these areas, because the cost of not winning is too high.
Whoever wins AGI would also probably win social networking. So you can see why the big 8 are excited about it, because you're playing for everything. It's like whoever wins this could be responsible for that $30 trillion of knowledge work.
Well, if you're Zuck, you spend $100 billion on it. If you lose or fail, likely everyone else will fail, in which case you're behind just like they are, and you've lost $100 billion, which isn't a huge amount of your free cash flow—maybe 12 to 18 months of free cash flow. If you don't spend that $100 billion and someone else does and wins, you lose $2–3 trillion of market cap. Correct?
You have to play.
You have to play.
Imagine if somebody built a more engaging social network. Social networks have just 1 unit, which is attention. We all only have maybe 4 to 5 hours a day to spend on an app. If there was a more engaging app, then, yeah, those are high stakes.
8. Why is SaaS Dead in a World of AI?
Do you agree, then, with the notion that if you believe in AGI, you cannot be investing in SaaS apps?
SaaS as we know it, I think, is over. I feel like quite a few SaaS apps were built at a time when software was relatively hard and complex to build. Imagine if you were building some customer-support software, some customer-support bot. To build a company like that, you would have had to hire some Stanford PhDs in NLP. You'd collect data for 6 months. You'd use a support vector machine or a neural network that would kind of work, and then you'd deploy it and grind away for a while.
There's a significant amount of capital that needs to be invested to get an app like that to work well. So it made sense for many companies not to bother doing that if it wasn't their core business. They could just use some third-party SaaS app.
Now many of these AI applications are incredibly easy to build on top of these LLMs. I feel like most companies will start building custom software super easily. We help companies build some of these custom apps, and the bar to create many of these apps has come down significantly. So that's risk number 1: companies do it themselves.
Risk number 2 is that you get Sonic-boomed by the foundation-model companies.
Sonic-boomed, meaning—
So today—
They move into the apps layer and just create it themselves.
Yes. I mean, it could happen, right? The models are becoming agentic. You've seen many of these agents. It's fundamentally about computer-use agents. If the models get better and better, it's possible the model is all you need.
Imagine if you wanted the model to—let's say I'm doing some HR thing—"Hey, update my medical benefits information. We've just had a new daughter. We want to update my medical information." If the model is agentic and sufficiently integrated into the database of the organization, you don't need anything else in the middle. So that's the second risk: the models becoming more agentic.
The third, and I worry about this a lot, is that I feel like a lot of our software was designed to be used by humans—humans navigating a GUI and clicking around and doing things. I think that's going to go away with multimodality. Again, this is why I think of 4 pillars to superintelligence: multimodality, reasoning, tool use, and coding.
Multimodality is important because we humans interact in natural language. We talk, there's video, all of that. I think the future might look like some type of ambient AI that you talk to, that will just go and do things. Maybe it'll use the GUI of the current SaaS application as an intermediate step, or it'll use MCP and use tool calls and get what it needs.
The GUI was designed for a world where humans were using a keyboard and a mouse and clicking around and doing things. I think humans can do better things with their time than click around.
Actually, one big change I have is I never actually type emails anymore. I use Wispr Flow, and its transcription is so good that I don't ever type emails. Ever.
Yeah.
Now the only trouble is everyone knows what I'm saying in my emails.
Yeah. The Wispr Flow founders interned at Turing back in the day.
No way.
Yeah.
Oh, wow. How do I feel about it? No, I don't agree. Why? Because the average company today has between 80 and 100 different SaaS products that they engage with. So, one, just the multitude of how many they'd have to create. Number one. Number two, maintaining them. You think they're going to maintain 80 to 100? Oh, my God. You're going to have teams and teams of people doing maintenance, updates, and debugging. I don't think so either.
And that is for the technology-savvy. Let's talk about every plumbing provider, law firm, accounting firm, and restaurant that can barely use Wix and Squarespace, let alone build out their own CRM system and POS system. Not a freaking chance. Then we move to foundation model companies moving into very vertically specific elements. You know, we're in a business that does AI for patent creation, updates, and collaboration. Sam is not going there. Sam has health, solving cancer, and energy utilization. I don't think Sam's touching patent creation and updating.
And so I think the more verticalized you go, the more defensibility you have. For those reasons, I think SaaS has life. Mine is a very biased perspective because it's my job. Is that all wrong again? You're the master here, Jonathan. VCs are literally middlemen.
I'd say, Harry, you have an interesting data set because you invest in a ton of startups.
Yeah. So I would be curious, looking at your sample of startups that you've invested in, to tally how many SaaS apps they use today at every stage and see if that has changed post-ChatGPT. My hypothesis is that today's companies use fewer SaaS apps and have fewer people.
Do you think we have more or fewer software engineers in 10 years?
More.
Help me understand that, then.
I think the definition of a software engineer will change. A Stanford doctor who's in oncology, who has an idea for some cancer-detection-type app, will be able to create a very simple version of an app that somebody could check by themselves and use as a home diagnosis.
I think there'll be more software engineers because if you define a software engineer as somebody who's capable of building a software product to solve a real problem, that pool of builders is going to expand way beyond people who've graduated with a 4-year computer science degree.
So we have more software engineers creating more software, and the problem then becomes discovery. How do we solve the discovery problem in a world of infinite software? You might have an agent for yourself that's talking to other agents on the internet. Have you seen Her, the movie?
Yeah. I really think we'll have—you might have, like, there's probably an agent for Harry that's just tasked with discovering what pieces of software Harry should be using. That might be talking to—I might have an agent that is maybe talking about what types of marketing things that I've built or that I have. I think we'll just be in a world where we have digital surrogates of ourselves that are communicating with other digital surrogates.
You might have an agent of yourself that's discovering startups to invest in. While you and I are chatting, there's probably—I don't know—you might be having a million conversations with entrepreneurs from all over the world. Today, you're constrained by space and time, but in the future you'll only be compute- and data-constrained.
9. Will the Phone be the Primary User Interface to an AI World?
Do we lose the phone as the interface to this world? We obviously see Sam and Jony Ive—there are rumors of pendants and some hardware devices. I'm not asking you to comment on that. I'm just saying, does the phone still remain the primary interface and design device?
We'll have some type of device that we'll carry, that's always on and processing multimodal tokens. For example, as I'm talking to you, if I were to envision my perfect device, it would be something that has cameras. So maybe it's a wearable, like glasses, or something that I'm having on me that's processing visual input, because I want to be able to read your body language.
I might have an AirPod-like thing in my ear that's whispering to me. It might say, "Jonathan, as you were talking about multimodality, Harry seemed less interested. His body cues suggest that he was losing interest, but when we were talking about AR, he perked up." Those types of feedback and cues, I think, would be good.
So I envision a device that—I think of it in terms of sensors and actuators. In terms of sensors, obviously it has to be listening to stuff; it has to be seeing stuff. But in terms of actuators, it'll probably also be speaking in my ear. Ideally, it should be something that you can talk to and have it do things later. For example, I might say, "Remind me to follow up with Harry on that idea for using Turing to automate clip generation." So it has to remember that and come back later.
I do think there'll be all sorts of new devices. Glasses and hearing devices like these AirPod-type devices seem obvious. There could be—do you remember this device called the Meeting Owl?
No.
During the COVID era, one of the tools that spiked was basically a speakerphone for having better distributed-team Zoom meetings. When somebody's talking, it would focus on them with a camera, and it was also a decent speaker.
Okay.
I can imagine devices like that that people have. It's hard to predict, but the thing that I almost feel confident about is that the phone will look so different. When we think of our smartphone, it's basically a computer with a phone app in it, right? The phone app is the least interesting part of the phone.
I think even for an AI device, it'll probably have some phone app in it, but everything else, I feel, will be magical. I feel like I would definitely benefit from a device that's constantly listening to everything that I'm listening to, constantly processing all the video and audio input that I'm processing, and something that's paging things to memory. Maybe it'll write things down and be able to look them up later. I see it almost like an extension of my brain.
Before we move into the quickfire round, I do just have to ask: what does your market and the data-provisioning market look like in 10 years? I always try and think about market composition and dynamics. Is it a winner-take-all? Is it very fragmented? Is it 3 or 4? What does that look like?
The market will reward players with research depth because the pace of AI research is so rapid. All these RL environments have spiked in the last 12 months after o1 came out in December and DeepSeek came out in January. So now, in addition to imitation learning, we are in this reinforcement-learning regime. 1 year ago, 1 year later, it could be something totally different.
I think the market will reward a company with research DNA, and it'll reward a company that can move fast and adapt very quickly.
Do you think this is a monopoly market, or do you think there will be many winners?
I think there'll be a few winners. A few, because I do think for the labs it helps them to have a few partners for resiliency. I imagine also for price competitiveness, there'll be a few winners in the realm of robotics and embodied AI.
We are still very early at Turing. We are scaling up on the robotic side as well in terms of data that we generate, but there's so much data that's missing that the models need to see that they haven't seen yet. I can totally imagine some newer companies also coming up that don't exist today.
If you were to invest in companies in your space, where would you invest?
Probably in robotics or embodied AI. The vertical stuff—we are scaling up pretty massively in generating data for different verticals. So I don't see that as a big white space, but I think everybody is relatively early with robotics. Robotics is such a vast realm that there could be interesting things to do there.
One way I see the space, Harry, is to think of it like these dimensions. The first dimension is the type of intelligence that you're baking into the models. That could be coding, STEM, functional expertise like sales, marketing, and software engineering, or vertical expertise like healthcare, legal, and finance.
I do a cross-product of that with the modality: audio, video, image, and computer use. That's multimodality, the second dimension. The third dimension is multilinguality, like different languages. The fourth dimension is different learning paradigms, like imitation learning, reinforcement learning, and pretraining, which is unsupervised learning.
All of those may require different platforms to be built. We've had to adapt our platform for imitation learning, reinforcement learning, and multimodality. So I feel like in this matrix, there are all sorts of new opportunities that could emerge.
And I only listed the digital intelligence. I didn't talk about physical intelligence. So I think robotics is wide open. The kind of data that a robot that's in someone's home needs is totally different from a robot that's doing things in a factory, and humanoid versus non-humanoid robots.
10. Quick-Fire Round
I could talk to you all day. I do want to move into a quick-fire answer. I’ll say a short statement, and you give me your immediate thoughts. What’s one widely held belief about AI that you think is wrong?
I don’t think we’ll see rapid takeoff. I think we’ll see incremental, continuous improvement in AI. I actually think this is good for the world because, if what we believe happens—which is that all types of digital knowledge work get automated—I think humanity needs time to prepare its workflows. I think we could use the extra time to upskill humans, rethink education, and make sure there isn’t massive job displacement.
I also think that, in the steady, continuous improvement in AI models, there’ll be value realized every step of the way, unlike self-driving cars. I feel like people have this wrong model for AI that comes from self-driving cars, where you get it 99% of the way accurate, and if you can’t solve the last 1%, they’re not useful. AGI is not like that.
I think when we automate the job of an underwriter, a claims processor, or a CEO, there’s incremental value unlocked for every percentage improvement as the model becomes more reliable. So, I believe in slow and steady takeoff, and that’s actually going to be great for the world.
You mentioned DeepSeek a couple of times. Do you think we underestimate China?
It depends on who you ask. The folks that I work closely with don’t underestimate China. I think it’s very impressive, the progress that they’ve made in open source with DeepSeek, Kimi K2, and Qwen. These models are state-of-the-art. So, no, I don’t think—at least among the frontier AI circles that I’m in—I think there’s a clear realization of how close they are.
The world seems to be moving to closed models. Is that good or bad?
I think it depends on the application. Firstly, in enterprises, it’s often a mix of closed and open models. We do see demand from enterprises that want either. The closed models are often easier to get started with, but there are some cases where enterprises prefer open models for cost and customizability.
I’m talking about the small language model regime, between 500 million and 10 billion parameters. I worry a little about frontier models. I feel like, for frontier models, there is some value in keeping some of the technology closed, just because of how powerful they are. I feel like the US labs are extremely responsible and safety-conscious in how they think about training and deploying these models.
You mentioned reading Elon’s book earlier. Elon is often chastised for his lack of care around some of the training elements. Do you think he is? And do you think he’ll actually benefit from not having that guardrail?
I think Elon also cares a lot about humanity. At least, if you read his book, one of the things I recall reading is that his motivation for getting into AI was that he wanted an AI that was speciesist and loved humanity. That was one of his reasons for getting into it.
Everything I see about the Grok team, I feel like their goals are much like those of any of the frontier labs: quite noble, in terms of having this powerful AI that can help humanity understand the universe and solve some of our biggest problems.
What did you believe that you now no longer believe?
I used to believe that, to build an enduring, valuable company, you hire a strong executive team and operate with a lot of leverage. Basically, hire strong people and get out of the way. I used to believe that.
Now I believe you hire great people and work really closely with them, their directs, and their directs, and their directs, and get as close to the ground as you can, where ground truth usually exists with the customers. The people closest to the customers are the engineers writing code and the salespeople talking to your customers.
So now I believe in being—basically, I used to, for lack of a better word, follow the org chart a little bit. This was also part of one of my learnings from Elon’s biography: he was so hands-on. He would be walking the factory floor and asking an engineer why this door in the Model 3 has 3 bolts instead of maybe 2, right?
It is a different way to operate, where you’re in the details of the most important things that matter, completely working in a flat structure and operating as close to the ground truth as you can.
Generally, I feel like in the early days of starting Turing, I may have had a subconscious desire to be liked. I think I must have had that. Now I don’t care about being liked; I just think about doing things that would solve our customers’ problems the best.
What was the most unpopular decision you’ve taken with Turing?
The most unpopular decision I’ve taken with Turing is switching from a distributed team to a hub-and-spoke model. We are now working from an office in San Francisco, and we’ve recently opened an office in Palo Alto. We’re going to be opening an office in London as well.
Very exciting.
I mean, for some people, that wasn’t very popular.
Until you fired them.
Some of them left, and, yeah.
We like in-person. We’re big fans of in-person here. Final one: when you look forward to the next decade, what are you most excited for? For me, my mother has MS. I think that we’ll have some pretty groundbreaking breakthroughs in MS drug discovery that we haven’t had forever. That excites me.
I’m excited about AI making new discoveries and automating AI research itself, to get to a point where AI is in some self-improvement loop, so that we could get to superintelligence faster. Automating AI research and getting AI to the point of making new breakthrough discoveries—that excites me.
I’ve always been fascinated by AI as this exoskeleton that makes you a lot more productive. Have you watched the Iron Man movies?
Yeah.
Right. So, in the early Iron Man movies, he’s wearing the suit, and the suit is obviously giving him superpowers, right? In the later ones, the suit is “agentic,” where he has these drone suits—an army of his suits that go off and do things.
I’m excited about a future like that, where every human on the planet has access to agentic AIs that help them amplify their fullest potential. Today, Harry might have 100 ideas, but Harry is able to do maybe 2 of them really well. I like a future where Harry can do the remaining 98. And I like that for the 7 billion humans on Earth.
I like that too, for my weekends’ sake, to be honest. Jonathan, I love conversations that are very natural and free-flowing. You can tell that I don’t really pay much attention to the schedule, but you’ve been fantastic. Thank you so much for joining me.
Thank you, Harry, for having me.