Mercor 产品负责人谈前沿实验室带来的收入集中度
- 开源不必蚕食前沿数据业务。 它反而抬高底线。Osvald 为 Mercor 的核心辩护是:数据在模型性能前沿最有价值;每家实验室都会购买评测集和训练集,补上当前能力的缺口;“开源模型只意味着,没人会购买 K3 已经能做的东西。”
- 他不接受“90%的企业工作流都能自动化”这种表述。 理由有二:Mercor 的 Apex 基准显示,顶尖模型在长周期工作流上的得分只有约50%;此外,还有一整类尚未被任何人尝试的潜在需求——比如部署一个运行数月、每周只需检查一次的采购代理。对于法律论证或医疗建议,“用百分比来框定完全不对,我们需要更多考虑持续且无上限的回报”。
- Mercor 的 token 开支已经超过薪资。 Benioff 每年在 Anthropic 上花费3亿美元,相当于开发者薪资的约3.8%;Osvald 认为这一比例在宏观层面会升至3%以上,并确认 Mercor 已经在 token 上花的钱多于薪资:“100%听起来合理。”企业目前还没有 ROI 问题,只是处于“探索和试验阶段”,同时在部分环节“逐步收紧开支”。
- 关于收入集中度,以及“这不算收入”的旁观评论:“我们每周结束时,银行里的钱都多了很多……现金流简直惊人。” 业务甚至来不及花钱来满足需求。解决集中度的办法,是通过自助式项目和 AI 项目经理向中小客户下沉,因为实验室项目需要白手套服务,运营上极其繁重。
- RL 环境是增长最快的数据类型。 一方面是应用的高保真模拟——一个“表现得和 Salesforce 一模一样”的 mock;另一方面是包含数百乃至数千个文件的丰富初始状态。他改变看法最大的一点也是环境:原以为环境无法规模化,“但后来它做到了”。他对未来3年的判断是机器人/物理世界数据,行业进展可能更像“Waymo 无人出租车/Cruise 时刻”,而不是 ChatGPT 时刻。
- AI 时代的产品组织正在倒置。 随着工程不再是瓶颈,PM 与工程师的比例会上升;团队开始争相缩小产品表面积;Mercor 正逐步从 Figma 转向云端设计;“技能问题几乎已经消失,所以现在全看判断力。”他的警告是:“不要把决策权……交给模型,否则你会失去这种能力,最后你会得精神病。”
- 竞争格局方面,对手是一群“自己做标注的创始人组成的家庭作坊”——由 VC 补贴、价格“完全错位”、深受实验室喜爱,但无法扩张10倍。 其他竞争者往往很快照抄 Mercor;Surge 则被形容为“有点离谱……非常神秘”。至于前沿实验室会不会吞掉应用层,Google Plus 和 Threads 已经说明,大公司“会输给那些对自己的市场高度聚焦的公司”。
1. 开源抬高底线,不必吞噬前沿数据业务
- Harry 以迁移到开源模型的论点开场,手里拿着 Kimi 的新模型:开源会不会蚕食 Mercor?Osvald 的回答构成了整期节目的主轴:数据在模型性能前沿最有价值。客户购买评测集和训练集,是为了补足当前能力的缺口;因此,“开源模型只意味着,没人会购买 K3 已经能做的东西。”
- 其中的传导机制是:开源模型“只是抬高了人们感兴趣的能力底线”。只要客户仍然想要新的能力,业务就会继续增长;前沿以下的商品化,并不影响资金真正流向的地方。
2. “90%的企业工作流”测量的是错误对象
- Osvald 不相信,无论是开源模型还是前沿模型,都能处理90%的企业工作流。这类计算可能建立在既有需求之上,但还有一批尚未被任何人尝试的潜在需求,最常见的是长周期任务:“设置一个采购代理,把整个采购团队连续数月完全自动化。你可能每周只检查一次。”
- 他愿意认可的数字是:在 Mercor 的 Apex 基准上,顶尖模型在长周期工作流上的得分约为50%。
- 他更强调的是,满足性工作流和无上限工作流不同。前者如更新 CRM——“你其实很难把它做得更好”;后者则包括法律论证,以及某种程度上的医疗建议,因为这类任务总有继续改进的空间。在这些场景里,“用百分比来框定完全不对,我们需要更多考虑持续且无上限的回报。”
3. 企业担心实验室介入核心工作,每家公司或许都会拥有自己的模型
- 对于 Alex Karp 所说的企业不愿把数据交给前沿实验室,关键取决于工作流的核心程度。HR 和采购相对不敏感,企业更愿意把这类工作流放在专有模型上;真正敏感的是能形成差异化的工作——比如律所撰写的“实际备忘录”,以及给客户提供的建议。Harry 认为其中存在反差:敏感数据可能交给开源、很可能是中国的模型,HR 反而放到闭源模型上。对此的反驳是,开放权重意味着推理“可以发生在多个地方……你拥有更多控制权”。
- 对于 Harry 归因于 Filecoin 创始人 Lynn Qual 的“每家公司都需要专门模型”论点,Osvald 回答:“我认同。”他也承认,这一判断对 Mercor 有自利性,因为每个专门模型都需要企业特定的评测和训练数据。市场规模“取决于客户能从专门模型中获得多少价值”,而那些能够被 ROI 证明合理的场景会随时间增加。
4. 没有 ROI 问题,Mercor 的 token 开支已经超过薪资
- “我不认为现在存在 ROI 问题。我认为我们正处于探索和试验阶段。”在 token 价格和模型性能变化过快、计算结果尚未稳定之际,企业会更有容忍度和耐心;与此同时,部分环节的开支也在“逐步收紧”。
- 他的预算框架是:工程师使用编码代理的开支“可能仍然在带来复利式收益”,因此不能算作销售成本。若客服代理消耗的 token 成本超过它服务客户带来的收入,那就“肯定处于糟糕的位置”。
- 针对 Benioff 每年向 Anthropic 支付3亿美元、约相当于开发者薪资3.8%的说法,Osvald 希望更准确地核算每个 token 带来的结果,并区分不同团队的消费结构。但在宏观层面,“这个比例会随着时间推移升至3%以上”。当 Harry 透露 Brendan 曾告诉节目这一比例会达到100%,且 Mercor 已经在 token 上花费超过薪资时,Osvald 回答:“是的,我们确实如此……100%听起来合理。”自他加入以来,公司员工数增长超过10倍,收入也同步增长——“我们甚至来不及花钱来满足手头所有需求。”
5. 服务业务弥合落地缺口,聚焦型公司可以击败大型实验室
- 对于服务业务热潮——包括 Microsoft 的服务部门和 Palantir——他自称的“热辣观点”是:短期内这就是未来。如何部署和评测代理的知识,目前“集中在旧金山的一群人手里”,之后会向各个行业扩散。最终它可能成为类似软件工程的一种职能,“也许会是一个持续10年的变化”。针对 Factory 的 Matt 所说服务只是“糟糕产品的借口”,他的判断是:问题在人才供给,而不是产品本身。
- 优秀工程师愿意做前线部署吗?那些沟通能力强、能直达问题根源的人愿意;而“这类工程师最终也很适合成为创始人”,所以 Mercor 不断有人离职创业。“失去一个人去创办公司,远好过失去一个人去接受另一份工作。”
- 对于 Harry 担心 Anthropic 会成为 Legora 的真正竞争对手,Osvald 让他看看 Google 和 Microsoft:“你还记得 Google Plus 吧?……它后来什么也没做成。”大公司会不断分散业务,但“会输给那些对自己的市场高度聚焦的公司”。
6. AI 时代的产品组织:缩小表面积、增加 PM,判断力成为核心工作
- 工程更快,反而让产品管理更难,而不是更容易:人们因为做得到,就提交几千行代码的 PR,于是团队“始终在努力简化产品表面积”。他最大的遗憾也指向同一个问题——标注平台曾试图“满足所有需求”,最终形成数百个异构项目,“管理起来简直是一团混乱”,本应更早设置护栏。判断需求是否会持续存在,是领导层基于与实验室持续接触所做的判断——“归根结底,这有点像猜测。”
- 组织上的结果是,PM 与工程师的比例会更高,因为工程不再那么受限,而理解用户工作流、理解什么能推动收入“现在成了瓶颈”。专家市场和 Studio 标注/评测平台这两块产品,都以4或5人为一个 pod,PM 与工程师的比例还会继续上升。
- 伟大 PM 所需的能力已经发生两次变化:工具变少了,“甚至 Figma,我们也越来越多地转向云端设计”;所有人都在提升到业务影响层面。“技能问题几乎已经消失了。所以现在全看判断力:我做的事情是不是能带来最大的商业价值?”
7. 招聘资深人才、测试判断力,并拒绝把大脑交出去
- 招聘现在更偏向处于黄金期的资深候选人——“25到35岁”——他们能更快理解什么驱动收入。初级人才是否会使用工具的能力变得“没那么重要”。Take-home 已经缩减为一次 AI 熟练度检查,之后通过白板面试考察实验、统计和系统设计,因为借助 AI 工具,“很容易把大量思考外包出去”,而他需要的是不会“只是复述 Claude 输出内容”的人。
- 他的“脑腐”规则,划清了判断与执行的边界:模型“会让你以为它做的是对的,但你仍然必须对它保持偏执……不要把决策权,也就是你真正的工作,交给模型,因为你会失去这种能力,最后你会得精神病”。他的类比是,手机“有点像会把你的大脑烧坏,把它变成一团糊状物”;关键是弄清楚边界在哪里。
- 他在糟糕的招聘中最难识别的是:“面试过程中真的很难评估一个人的主动性和主人翁意识。”有才华但讨厌的人尚且可以容忍——“性格可以改变……但很难让一个人真正上心。”
8. 业务机制:现金流惊人、收入集中,并有意向下沉市场推进
- 针对“这不算收入”的评论,他的回应是:“我们每周结束时,银行里的钱都多了很多……现金流简直惊人。”他在其他地方见过“有趣的金融工程”,但 Mercor 不是那种情况。
- 关于前沿实验室带来的收入集中,公司的方向是向中小客户下沉:提供自助式人工数据项目和 AI 项目经理。原因在于实验室项目需要白手套服务,运营强度极高:持续挖掘边缘案例、让客户与专家之间“以惊人的速度完成对齐”,并对每一个数据点都保持“极度偏执”,确保其完美。企业数量远多于实验室,而这一转向已经降低了收入集中度。
- 供给侧的秘诀是专家体验:准时、优厚、透明地支付报酬,带来转介绍;同时配备 sourcing 团队,应对尖峰式技能需求。公司并不是靠超过竞争对手的“疯狂奖金”取胜,而是让专家看见未来的工作和技能成长。利润率是在项目结束后才确定的,如今合成数据和自动质控项目的成本,大致由专家报酬与 LLM 开支各占一半。
- 实验室会谈判,但这套合作关系令人羡慕:评测集代表客户真正想要的东西,训练集则不断推动模型爬坡。因此,“只要他们在数据上的花费低于将要获得的收入,人们就会想把这个杠杆越拧越紧。”
9. 环境是前沿数据类型,家庭作坊无法规模化,机器人是未来3年的判断
- 环境是增长最快的数据类型,Mercor 认为自己处于这一品类的领先位置:一方面是应用模拟,另一方面是包含“数百个文件、数千个文件”的丰富初始状态,让训练数据“更接近模型部署时真正看到的东西”——比如一个“表现得和 Salesforce 一模一样”的 mock。它现在之所以难,和过去的 SFT 及偏好排序一样;“实验室正在摸索。最终……企业也能做。”这也是他改变看法最大的一件事:他原以为环境无法规模化,但“我们不断推进,最后它真的奏效了”。
- 竞争格局是“一群自己做标注的创始人组成的家庭作坊”:这些聪明的前技术创始人经营着“实验室喜欢的、由 VC 补贴的工作”,因为价格“完全错位”。一旦客户要求10倍吞吐量,模式就会失效。其他竞争者经常照抄——“我们写一篇博客,通常一周后就有人写一篇一模一样的博客”;Surge 则被形容为“有点离谱……非常神秘”。网络安全是另一个增长方向:对抗式、“AlphaGo 式”的场景,回报没有上限,“目标会一直移动”。
- 3年后尚未出现的收入曲线,来自机器人所需的真实世界物理数据。相较于 GenAI 和自动驾驶,这仍是一个早期市场。针对 Harry 对“15分钟取水”演示的吐槽,Osvald 提到 Cruise 走向 Waymo 的先例:“现在我打 Waymo 的次数比 Uber 多。”因此,行业拐点可能更像“Waymo 无人出租车/Cruise 时刻”,而不是 ChatGPT 时刻。
- 2000亿美元的乐观情景,是一家技术驱动的服务公司,其中“评测就是 PRD,同时也是提升性能的优化目标”。只要更好的模型对经济仍有价值,评测和训练需求就会持续存在;此外,还会叠加不断增长的代理部署企业业务。
核验说明
- 原始字幕提到“Kimi”,并分别出现“K3”;无法确认“Kimi K3”是否作为一个完整名称被说出。
- 原始字幕将2000亿美元乐观情景中的公司识别为“Brex”;结合上下文看可能指 Mercor,但仅凭原始字幕无法确认。
Osvald, it is so good to have you on the show, dude. I've heard so many good things from Brendan. Thank you so much for making this happen, man.
1. Does Open-Source Cannibalize Mercor's Core Business?
Thanks for having me. Super excited.
Dude, I am seeing open everywhere. Everyone is claiming that we'll see a mass migration from closed frontier models to open models. Kimi very recently came out with that new model, and I didn't really know: does open source cannibalize Mercor's core business?
I wouldn't say that improvements in open-source models cannibalize our core business, because data is most valuable on the frontier of model performance. Each of our customers has their own unique goals and is purchasing eval and training datasets to fill gaps in current model capabilities. Open models just raise the floor of what people are interested in. As long as customers still have new capabilities they want to get better at, our business continues to grow. Open-source models just mean that nobody's buying anything that Kimi can already do.
2. Why 90% of Enterprise Workflows Can't Be Done With Open Models
So if 90% of enterprise workflows can be done with open models—which more and more people say they can—and that 10% is really where you serve your customers and provide data, I'm naive: does that not make it harder and harder to make huge amounts of revenue if that 10% at the frontier moves further and further away?
I'm not convinced that 90% of enterprise workflows can be handled by open models or frontier models right now. We think that these calculations might be based on existing demand or things that come to mind when current model users are thinking about what models could do.
But there's a whole category of latent demand that people aren't even trying to address with models yet. Most commonly, we think these are long-horizon tasks, like setting up a procurement agent to fully automate your procurement team for months on end. You only check on it maybe once a week. We think that's not even captured in these calculations when someone says enterprise workflows are being handled because nobody's trying to do these things yet. The market for data to support those use cases is growing, and that's where we see a lot of the leaders moving toward.
Okay, so we see a lot of leaders moving there and seeing new capabilities that they never thought existed. But then we have Alex Karp, in what I thought was a rather sedate performance. Normally, he jumps up and down much more, but it was still rather energetic. He talked about the incredible skepticism we see from large enterprises toward data and sharing data with frontier-model providers. To what extent do you see skepticism and fear from large enterprises in working with frontier-model companies?
We see it depending on the specific workflow and how core it is to the business. Things that are just general things that every company needs to do, like HR and procurement, can be less sensitive, and enterprises are more open to putting these workflows on proprietary models.
It's the core work that the company is doing that's vital to its business, that differentiates it from competitors, where we see more sensitivity. You can imagine this being the actual legal services that a law firm provides. What are the actual memos that it's writing? What is the advice that it's giving to its clients?
Am I the only one who sees the irony in this? We put the sensitive data on open-source, most likely Chinese models, and we put the HR and procurement data on the closed model. Am I a [__]?
Well, it depends on where you run the open models, right? Whether or not that's a bad idea. The beauty of open-weights models is that the inference can happen in multiple places. You could make mistakes using them, but you have more control.
When you look at that dispersion, what do you think is inaccurate? You said you didn't really believe the 90/10. What do you believe is a more accurate representation?
In our APEX benchmarks, we're getting closer to around 50% of long-horizon workflows. Top models are scoring around that much. But I think there's a class of workflows that are just sufficiency-based, where you do it and it's done and you're good. This is something like updating a CRM; you couldn't really get much better at it.
Then there's a class of workflows that we shouldn't even be thinking about in binary terms: Can the models do it or not? These can be things like legal arguments or, to an extent, medical advice, where you could always get better. In those cases, I think the percentage framing is just totally off, and we need to be thinking more about continuous, uncapped rewards.
When we think about the fact that it could be better, I had Lynn Qual, the founder of Filecoin, on the show the other day, and she was like, “Exactly. That is why we'll have specialized models for every single company.” It could be better depending entirely on the company: one company wants to focus on growth, one on margin, and another, if we're in Europe, wants to focus on work-life balance.
So you need individual, specialized models for every company. Do you buy that we'll have specialized models for every company, or is that a little bit self-serving toward Filecoin?
I buy it. I think it's also self-serving toward Mercor, in that we think every specialized model will need enterprise-specific eval and training data to show the model how to perform in its setting.
I think the diversity and the market for this depend on the value that customers can get from the specialized models. There will be cases where the ROI is really justified, and I think those cases will increase over time. But we certainly believe in this future.
3. Do We Have an Enterprise AI ROI Problem?
Switching back, Alex Karp's second point in that show was ROI questionability. You mentioned the word ROI there, which made me think of it. Is that very present, and do enterprises maybe have questions about the ROI they're getting? Do you think we have an enterprise ROI problem with AI today?
I don't think there's an ROI problem right now. I think we're in a period of exploration and experimentation, where there's more tolerance and more patience to get that ROI calculation right. There's a lot of different projections around where token prices will go and where performance will go, and right now we're starting to see some amount of tightening of the screws on spend here and there.
But I think the paradigm we're in is still, “Let's see what happens,” because things are moving so quickly that the ROI calculation might still shift too dramatically.
There are 2 ways I want to go on this. I'll take the first way. We saw Aaron from ClickHouse say that he 6×'d his spend, and that's what they need to do because they need to be at the frontier. Then you see Uber and Microsoft, and some forms of—I think it was Grok or X, or one of Elon's companies—put budgets on a per-user basis.
4. Balancing Token Spend vs Performance
What do you think is the right way to be navigating this cycle? If I'm a founder listening, what would your advice be on how I should think about optimizing the balance between performance and budget?
It totally depends on the use case. I've mostly worked at hypergrowth companies where growth matters at all costs, right? There's a willingness to spend for growth as long as the unit economics are fine.
When you're looking at coding-agent spend for your software engineers, that's not always COGS for your work. If that's really high, that could still be giving you compounding gains. If you're looking at a customer-service agent that has massive token spend and the revenue you're getting from the customers being served is way lower than the token spend, then you're definitely in a bad position.
5. Salesforce Spends $300M on Anthropic
In my experience, it's just been these growth-stage companies, and I think for a lot of founders, considering your token spend—if it's for growth, if it's for improving the efficiency of your headcount—that's just what you need to do to service large amounts of demand when you're starting up.
Mr. Benioff from Salesforce said that he spends $300 million a year on Anthropic, which works out to about 3.8% of developer salaries if you average the salaries. Do you think that is the going rate moving forward? Do you think that will be 20%? Or do you think it'll be 100%? Or will it be way less?
I hope that we can move toward a future of better accounting for the outcomes being driven by token spend. Even here, I think that in a company like Salesforce—a company of that size—you should certainly have different spend profiles depending on what the team is doing.
Again, here you have teams that might be more like solutions engineering or forward-deployed, where you have to think in terms of unit economics, and teams doing R&D where you can have more tolerance for spend. So, I think at large companies, you have to consider which parts of your organization are doing what and how much tolerance you should have in different areas. I think, at a macro level, the percentage will increase over time to more than 3%.
Huh. Do you want to hear something funny? Brendan said on the show that it would hit 100%, and he said that you already spend more today than you do on salaries.
Yeah. Yeah, we do. And 100% sounds reasonable. As I said, I've only worked at hypergrowth companies, and that's what Mercor is and continues to be, more so every day as the growth just accelerates.
For us, it makes sense because the demand that we have is so high. The company has grown more than 10x in headcount since I joined. The revenue has also commensurately increased. We're just in a race nonstop to service our insatiable customer demand. So, for us, it makes sense because we can't spend money fast enough to service all of the demand that we have.
6. AI Makes the PM Role Harder, Not Easier
Dude, do we just build 10x more products more quickly? Help me understand. Do we have smaller engineering and product teams? Do we just build much more than we ever used to? How do you think about that?
I think this paradigm makes the job of product management a lot harder because we're trying not to build 10x more product surface area. It makes things incredibly chaotic. We have moments in time where product surface area rapidly expands because people think, “Oh, I can make all these features really quickly. This is like—I could just push these multi-thousand-line PRs.”
But we're constantly in this battle to try to simplify our product surface area and find the interactions and workflows that are most scalable. The trend that we see is that, as a product team, we're constantly fighting to reduce surface area and simplify things.
We also see a higher ratio of PMs to engineers because engineering is less bottlenecked. There's much more work to be done in understanding the workflows and needs of users, and what products actually drive the most revenue becomes the bottleneck in servicing more demand for us.
If we think about the pre-AI era, how has what it takes to be a great PM changed for this new world?
There are 2 major changes. One is that you don't really need to learn as many tools anymore. You just have to be able to use coding agents; a couple of tools will do everything you need.
Even Figma—we're moving away from it in favor of cloud design more and more. So, less tool diversity for us.
The other is that everyone needs to uplevel a lot and think about business impact much more. I think all work is starting to look higher-level, so the minutiae and details get sorted out way faster. All the PMs at Mercor have to think much more about, “Is what I'm focusing my time on the right thing?” I can do things very quickly now. Skill issues have almost gone away. So now it's all about judgment: am I doing what is going to drive the most business value?
7. The Biggest Product Mistake
Dude, I have to ask. You said that a core part of the job is retaining simplicity and deciding what to do versus what not to do. What did you do in product that, with the benefit of hindsight, you wish you hadn't done? And what did you learn?
One interesting thing that happened this year was that our annotation platform served a lot of different workflows. The demand for human data is so large and heterogeneous, and our delivery team is so good at delivering projects and selling projects, that we supported too many workflows for human data projects.
We built a tool that was extremely flexible in supporting all sorts of different research experiments that customers might want to do. The shape of data has changed a lot since it started with InstructGPT for GenAI—from supervised fine-tuning to preference ranking to all these environment-type projects. There are a lot of multimodal projects that have totally different formats, and your annotation tool needs to support these different workflows.
Customers will ask for all sorts of things. We tried to serve every ask. We made a tool that was maximally flexible. We had hundreds of different projects running on it. That's just chaos to manage.
What we needed to do sooner was put guardrails on the type of services that we support and work closer with our operations team to say, “Hey, here are the best practices.” Customers are going to ask for everything. We can do it, but should we do it? If there's no enduring demand for certain workflows, maybe it's not worth the investment.
Putting guardrails in place and narrowing down the services that we support was something we should have done a lot sooner, and we did it recently.
How do you determine enduring demand?
This is what makes Mercor a hypergrowth company: we're incredibly tapped into the market and the ecosystem. It's really a judgment from leadership, I think. It's very hard to say what data will look like in a year or 2.
The best way to figure it out is to stay in constant touch with leaders from a diverse set of labs and constantly validate hypotheses. I think Brendan does it very well. I think our operations team does it very well. But ultimately, it's kind of a guess.
Which lab has the most advanced and sophisticated data team?
I can't speak too much to customer details.
[laughter]
They're all super good. Everyone is sophisticated. Everyone blows me away in different ways.
That is such an unfair question. Okay, I totally agree. The other question to ask is, which has the worst team?
[laughter]
My question to you is, you mentioned another element, which actually didn't shock me, but I thought it was interesting: the movement away from Figma. Can you talk to me about that? I hear more and more companies doing the same. As a product leader today, how do you think about that, and what was the thinking there?
The team can do whatever is best for them, and this is a trend I've just observed amongst almost everybody. Cloud design has done a great job. People really like using it. It's easy to use, and we've just had a natural movement towards it.
It's also been easier not to have too many tools and not manage too many licenses. Because Claude is making all these other great features, people just gravitate towards it. Then it's a bit less friction to have the procurement team issue licenses for Figma for every single person.
We were talking about the ROI earlier for enterprises, and we're seeing Microsoft set up a services department. We're obviously seeing Palantir skyrocket, and services becoming an increasing part of everyone's business. Is that the future of AI enterprise deployment? How do you think about the incredible rise of services in deployment?
Yeah, I have a bit of a hot take here. I think it's the future in the short term, as knowledge of how to use AI gets disseminated throughout industry.
We have basically a concentration of a bunch of people in San Francisco who really know how to deploy agents, evaluate agents, and be AI-first in engineering and in other areas. That knowledge just isn't out there yet. Eventually it will be, and maybe you won't need teams to go and set things up—to set up AI agents for every enterprise—and it'll become more of a job function, similar to software engineering.
And so, in the short term, it enables deployment. In the long term, products become more and more sophisticated, to the point that they're able to set themselves up. Because Matt from Factory said to me, “You know what? Fuck this. Services are just an excuse for a crap product.”
I think that it's a knowledge-dissemination problem. That's one way to look at it. The other way is, why not hire someone to just do this agent deployment at your own company?
I just don't think the skill is out there yet. I don't think there's enough—I don't think the talent is available for every enterprise to have its own expertise in it at this point in time. But that'll change over the long run. This is, I think, maybe a decade-long change.
The question is, do good engineers really want to be FDEs, though?
[laughter]
There are a lot of different types of good engineers. There are a lot of ways to be a good engineer. One way to be a good engineer is to be a great communicator, cut through to the source of a problem, and simplify. I think those engineers are great fits for FDEs.
I think those engineers are also great fits to eventually become founders. That's a different profile of person who's incredibly valuable. That's what a lot of people are looking for when they're looking for FDEs. It's also a profile that we look for generally, which is why we have so many alumni go off and start companies.
Do you like that? I spoke to Brendan about this, but is it a good thing to have some sort of a core mafia? Because you also want to retain talent.
I'm proud that, of the people I work closest with on my teams, I've only had attrition to founding. We've had quite a bit of it. It's a lot better to lose someone to starting a company than to taking another job.
It's interesting from a personal level because I like these people. I wish the best for them, and I really enjoy seeing it. It is tough, though. It makes the job of management a lot harder because we just have so many high-agency people who are very ambitious, and it's difficult.
But I like it, and I'd rather be in an environment like this than one where everyone's a little soft and saying, “Oh, I don't want to work.”
[Laughter] No wonder you left Europe. How has hiring changed in a post-AI new world? When you look at the people that you add to your team today, especially in product, what do you ask or look for today that you didn't before?
8. Why RL Environments Are the Fastest-Growing Data Type Right Now
I think, touching on the earlier point of everybody needing to up-level and think closer to business impact, we've biased towards more senior hires who are better at understanding what drives the business forward and really grokking how we operate, how we make more revenue, how we deliver better services to our customers, and how we keep our customers happy. I've found that more senior candidates just get that a lot faster.
All these things like, “Can you use the tool?” and all these other, more junior things are becoming less relevant. The hiring for us is biased towards more senior candidates.
Do you worry that you're just falling for the classic—I’m so sorry to be the fast-growth founder-mode person—which is that your VCs come in and say, “Oh, you need to hire this person from Facebook,” and you get the seasoned operator who fits exactly that rubric? And it never works. It never works.
We're not quite doing that. Seasoned is a spectrum, right? I'm not saying we're hiring people who are formerly in executive positions. We are treating everything as an executive search, but we want to find someone who's at the sweet spot. They're still hungry, they've done the job that we want them to do for a few years, and they're right at the point of really hitting their prime.
When do you think people hit that prime?
[Sighs] I think 25 to 35.
Oh, I just turned 30. I'm bang in the middle. Perfect.
Good timing for you.
Perfect timing for me. In terms of the questions and what we look for in the take-home assignments, has that changed?
We've moved away from take-home assignments. We do 1 take-home assignment, which is, “Can you use an agent? You're on your own for a bit of time. Go use an agent, produce this artifact for me, and we'll look at it.” We do that once so that we know the person is AI-fluent.
Then we move towards a lot of whiteboarding because we want to avoid relying on that. We'll do 1 round where we find out if the person is just familiar with AI tools.
Don't laugh. Okay, so we do that. I'm familiar with AI tools, and now you're coming to my room. We've got a whiteboard. What do you want to see? What would impress you?
We care a lot about being able to set up good experiments, understanding statistics, having good judgment, and systems design as well. The reason is these are just skills that are so easy to assess.
Sorry, I'm so sorry to interrupt you. Good experiments and systems design—it feels quite wordy. What does that actually mean?
We ask people a lot. I don't want to give away too much about our interview process, but we need to run a lot of experiments as a product team. We need to make sure that our team knows how to run a good experiment that actually reveals information and isn't just totally fudged.
With AI tools, it's very easy to offload a lot of thinking and a lot of judgment. We want to make sure that people still have the ability to have good judgment, know what they're doing, and not just regurgitate what comes out of Claude.
That's so interesting. I completely agree with you. I have it with my team: we do scripts for content and for Reels. I do all the questions myself. I would never use AI, and I'm very concerned about it because, to me, you lose the muscle.
Can I ask you, how do you retain thinking, thought, and creativity when so many people are so freaking hooked already?
I think it's like phones. They fry your brain and turn it into goop. But I love it—I do a lot of stuff on my phone. I use my phone all the time. You just have to learn personally where that boundary is: when is a good time to scroll through Reels, and when is a bad time?
For work, I think that boundary is between judgment and decision-making and execution. I was very careful never to delegate judgment or decision-making to models because they make you think that they're doing the right thing, but you have to be paranoid with them still. You still have to double-check everything.
That's what I tell my team: don't delegate your decision-making—your actual job—to a model, because you're going to lose that ability, and then you're going to get psychosis.
Totally agree with that. When you look at the experiments that you've run, does the data correlate to the outcome? I often think in investing, sometimes I do no work and no diligence and make loads of money. [Snorts] And sometimes I do lots, make terrible investments, and lose all the money. Do the inputs correlate to the outputs?
It varies because we run a lot of experiments. Sometimes they do, and sometimes they don't. We want to get more experiments that actually show good results and move the business forward, and that's really the job of the team: to find the right experiments to run and make the narrative around, “Hey, these changes to our product have impacted the business in a positive way.”
That's a lot of the core job right now. Week to week and month to month, we get different results, but we try to trend in the right direction over time. People start to learn the dynamics of the product and the dynamics of the users better and better to improve over time.
9. How Mercor's Product Teams Are Structured
When you think about product and engineering running the experiments that you mentioned, how do you structure the teams today? What does that meeting look like? Do you have a weekly product team meeting? What is the right way to approach the cadence of product team meetings and how to run them today?
We break it down into a few different groups that do experimentation. We have 2 major product areas where this is most relevant: our marketplace, which matches experts to jobs, and our annotation and eval platform.
The annotation platform is where experts log in to do annotation for eval or training datasets, where our operations team also logs in to run those projects, and where our customers log in to see their data and run evals. The annotation platform is called Studio; the other one, call it the Marketplace.
These 2 groups are self-contained in trying to make their individual product offering better. We have 2 main modes of engagement within human data: talent-only, which is when we just send experts to our customers and they'll run the project.
This is like a lab needs a doctor, a lawyer, or whatever, and they're like, “We're just going to use them.” Thanks for finding the best person for the job. You'll need to pay them and performance-manage them, but we'll run the project.
Then there's a managed-service project, where we give our customers data. For the talent-only model, we just use the Marketplace. For the managed service, we use the Marketplace to send people to our annotation platform, and then we'll run the project and give them the whole dataset.
Totally. Are they 2 separate product teams?
They are 2 separate product teams.
How big are the product teams?
Around 2 to 3 per product area, with data scientists and a design team as well. There are data scientists dedicated to each, and a design team that flexes between them—just a few, yeah.
So you have pods of 4 or 5.
That's fair, yeah.
Gosh, yeah, totally. Okay, that makes absolute sense. Will those ratios change over time, do you think, between PMs and design, or will that stay the same?
I think that the ratio of PMs to engineers will change over time to have fewer engineers per PM as engineering velocity increases with better coding agents. We will be bottlenecked by understanding business needs and user needs, and that's more of a PM job.
We need to be very careful as a hyper-growth company to grow the teams in lockstep because headcount increased more than 10x in the last year. We just want to be careful not to grow 1 faster than the other. The trend will be to have a higher PM-to-engineer ratio, though.
When we talk about the good experiments and making sure that we're running a really tight process, what does that look like in terms of the meetings? Do you have a weekly product team meeting? What is the right way to approach the cadence of product team meetings and how to run them today?
We break it down into a few different groups. Within these product areas, we'll have a whole-product-area weekly with product and engineering and a lot of other stakeholders as well.
This one is just for everyone. It's broken down into pods. As I mentioned, within that product area, there might be, let's say, 3 product managers. They'll all have a pod of these parts of the product that we can naturally segment work into.
Our marketplace, for example, has an expert-facing side and a hiring-manager-facing side. These are naturally 2 distinct pods. There are some other pods within here as well, like managing the expert experience and making sure that everyone has great customer support, that there are never any issues with working for Mercor. Each of these pods will do their own sprint planning.
They'll come together in the weekly product-area meeting, and we try to keep it efficient while maintaining a lot of visibility between the pods because they all need to have their roadmaps well aligned. We need to have weekly meetings to maintain accountability, right? We do them on Friday, later in the day, to make sure no one's leaving early for the weekend.
Love it. What do you not do in your product meetings that you should do to make them better?
It varies by product area. The challenges in the marketplace versus the annotation platform are a bit different. The main challenge as we grow quickly is having the right amount of communication and feedback from other teams.
Our marketplace and our Studio team need to get information from each other, right? There are cases where something's wrong in one and it's affecting what's popping up in the other. If something's wrong with one product, it's affecting the expert experience when they're on the other one somehow.
That communication channel just explodes very quickly because the headcount and the teams have grown so quickly. We need to do more cross-product-area collaboration. Keeping it efficient is just really hard as the team grows because the nodes just keep moving around and there are more of them.
What has been the secret to scaling supply on the marketplace side so efficiently? That's hard. How have you guys done that so well?
10. The Three Secrets to Scaling Supply on the Marketplace
I probably put it down to 3 things. The first one is a great expert experience. Experts get paid on time, they get paid well, and they get paid transparently. Everybody involved in the expert experience cares deeply about whether or not experts are having any challenges and whether the work is dignified, well paid, and fairly paid.
That is a requirement for a great referral program because nobody's going to refer their friends or colleagues to some kind of job that sucks, right? Everybody caring about the expert experience drives a great referral program. Additionally, having a great sourcing team that's able to find people in every corner of the world with very specific skills helps us fill the gaps when we have spiky demand for a specific skill set.
Are people as short-sighted as just wanting to be paid the most? I've heard that Mercor pays the most. Is that the secret?
It's not—I wouldn't say it's short-sightedness, because we want to retain the top experts as well, right? You might get paid a lot on 1 project, and I know there are a lot of other competitors in the space who will do some crazy bonus payouts and stuff for short-term sprints. That doesn't get you to come back as much as a great experience with a lot of work, visibility into what future work is coming up, and the feeling of, “I'm growing my skill set.”
I have the ability to pick between a few different jobs. I'm doing interesting work. I have great communication from the people running the project. It's really hard to sign up for online work and then just get hit with this 100-page instruction document. It's a very foreign kind of job.
That's part of the experience as well. Knowing that you're going to get paid highly for a long time for something that you can do for a long time is what keeps people interested.
Have you seen your margin improve over time, or are margins relatively fixed given the complexity?
Margins are an interesting thing in this business. We try to think about, as a product team, how we deliver the best value for our customers. That is independent of how we price the project.
There are cases where you could have automatic quality control and synthetic-data improvements to make the delivery better. There are situations where you could think about the staffing on the project to change the cost of the service.
As a product team, we want to make sure that we can deliver the best value to our customers and have the best experience for our experts. Margins are decided after the fact based on consideration of costs. Now, for a lot of these projects, the costs are driven equally by paying experts and LLM spend on things like synthetic data and automatic quality control.
Does the “It's not revenue. It's not revenue” shouting from the crowd, throwing peanuts, annoy you? Is there anything that hasn't been said that you think people are just not getting?
It doesn't annoy me, no, because we end every week with millions more in the bank, right? It's funny how, at other companies, I've seen interesting financial engineering and accounting, and people can have all these different metrics. But we end every week with so much more money in the bank. The business is very healthy, and we can't spend money fast enough.
What people want to call it is up to them, but the cash flow is insane.
11. Mercor Has High Revenue Concentration
Does it matter that you have such high revenue concentration? The frontier-model providers are your biggest customers by far. Some would say, “Woof, that's a lot of concentration.” How do you think about that?
I can answer this from how it affects the product team. We would love to move downmarket. Our biggest challenge is moving downmarket so that every single enterprise can efficiently run human-data projects for eval and training. That'll diversify our revenue for sure, because there are many more enterprises than there are labs.
That's a harder product to build. The direction that we're taking our products and the company is to be able to self-serve projects very efficiently and have AI project managers, so that it's a lot easier to do this work for smaller customers. Running a human-data project for a lab is incredibly hard. It's a white-glove service that requires a lot of people on the operations team.
As we make that more efficient with better products and better processes, we can do smaller projects that are more heterogeneous for more customers. It's the direction we've been heading in, which has reduced concentration, and it's the direction that we'll continue to head in as every enterprise begins to have human-data work for its proprietary use cases.
What's so hard about it? Making it really simple, explaining it—what is the challenge with not dumbing it down, but democratizing it?
Running a human-data project is just hard. There's so much information that needs to be transmitted from the customer, as the end users of our customers, to experts. All the edge cases matter, right?
People will try to write a guideline that says, “Here's how you make a data point,” but the experts will have some edge case that gets bubbled up, and what you do on that edge case matters a lot. The process of making a human-data project is basically continually surfacing these edge cases, which requires insanely fast alignment between customers, maybe their customers, maybe other experts in the field, and the experts who are doing the annotation.
It also requires a huge amount of paranoia from the operations team to make sure that every data point is perfect, that it fits whatever guidelines the customers have, that the projects are running on time, and that all the bottlenecks are removed. It's just an operationally intense process because it necessarily deals with edge cases and things that haven't been seen before and are outside of model capabilities.
The data types also change very frequently. We've moved from supervised fine-tuning to preference ranking to rubric-based annotation and now RL environments across a whole bunch of different modalities. There's a lot of complexity within each project and then between projects.
I would boil it down to those 2 things: the need for paranoia and the need for very crisp communication that make it challenging.
What data type is not hugely in demand today that you think will be hugely in demand next year?
The data type that's growing the fastest for us is environments. You might have seen a lot about these RL environments on Twitter. It's kind of a hype term. Every company has a different definition for it.
We are certainly the leader in the category and view it as basically simulations of apps that you might want your agent to use, and also a very rich start state, which we call the world, that is representative of all the data you might have on your machine, like your laptop. Then we have tasks that train agents how to use those tools to accomplish something that's useful.
It's a bit of a complicated annotation process because the agent has to interact with this simulated world. We have to make that start state, which can be hundreds of files, thousands of files. The shift here is that the data that the models—the agents—are being evaluated and trained on looks a lot closer to what they see in deployment.
Right? So, if you want to learn how to use something like Salesforce, you need a pretty high-fidelity mock that acts exactly like Salesforce in your evaluation and training. It's complicated to get this set up, just like years ago preference ranking was really hard to get set up. SFT was really hard to get set up when InstructGPT first came out. So, this is the frontier right now. Labs are figuring it out; new labs are figuring it out. Eventually, it'll get so smooth that enterprises can do it, too.
Are labs price-sensitive on data acquisition?
By data acquisition?
Well, when they go on a project with you, are they price-sensitive? Are they haggling, going, “Oh, well, Edwin at Surge gave me a 10% discount. Can I have that?” Or are they like, “Just give me the [__] data”?
Well, there's always the aspect of negotiation and the procurement team trying to get a better deal. But we've chosen a great business where our work directly affects the business outcomes of our customers, right?
We have a great setup where, if you're making an eval set like any other lab, you're evaluating something that your customers want to do. If you could just do it better, right? You would make more revenue. If they're buying a training set, they're now hill-climbing that eval set that they've said represents what their customers want to do.
12. Why Data Projects for Enterprises Are So Operationally Intense
As long as the amount of money they're spending on data is less than the revenue that they're going to get, they're happy to crank the lever. People want to crank it harder and harder because spend on data directly translates to more revenue for our customers.
Do you think we'll have an unbundled data-provider world? I'm a venture investor, and I see so many people who are like, “Oh, we're like Mercor, but for domestic robotics.” You're like, “Okay, cool. Good. I get it.” Do you think we'll see this kind of specialized data-provider world where niches have thousands of players?
To an extent, we're already in this world. I wouldn't say it's always that successful for the small players, though.
How I would describe it is that we're facing what looks like a cottage industry of founders doing annotation themselves. You have all of these small startups where, as the skill bar for annotation gets higher and higher as models get better, the founders are actually just making the data.
Labs love this because it's totally mispriced, right? Someone raises a bunch of money, they have loads of cash to blow, and they go to these labs and say, “I need to get your business. Please let me work for you.” They're smart people. They're founders. They're formerly great technical employees. But they're running the projects themselves and doing the annotation themselves.
This is VC-subsidized work that labs love. The problem is scaling it beyond a few data points or beyond what 1 founder or a few full-time employees can do. This is the position that we're in: we're having to compete against basically founder-led annotation, where some of them are even running it as cash-flow businesses and just taking the profits home themselves.
It doesn't scale, though, and our customers know this. It won't scale when you want to 10x the throughput and 10x the amount of projects. But it is indicative of the direction the field's heading in: we need higher-skilled experts. We need the best people in the world to be doing this annotation.
Don't laugh. I have a bit of an ego, and so I like to feel like a special snowflake. What I mean by that is, I would be like, “Oh, when Meta or OpenAI, or you name your large company, is buying data from multiple people, it feels like you're being promiscuous and cheating on me.” Do you mind? Do you monitor budget and the percentage of budget that gets spent with you versus another provider?
Of course, we do a lot of competitive intelligence. Our customers like us, so they'll often share information with us. But everybody just wants models to get better, right?
We're happy to have this kind of competitive pressure that tells us where to go. If someone else is able to do something better than us, we'd love to hear about it and then do it better than them, right? It's healthy to have vendor bake-offs. It pushes us to make our services better.
We do stay on top of it because we want to deliver better services to our customers. We want to know who's doing better than us, and then we want to surpass them. So, it's a totally healthy thing to happen as long as Mercor's winning.
You said models are getting better. We said frontier earlier. I'm an investor in Legora, and everyone's like, “Oh, your real competition is actually Anthropic.” I'm like, if Anthropic goes after legal and wins over Cooley and Goodwin, something's gone very wrong with the world, because they should be solving cancer and climate change.
To what extent am I right, and how do I balance between Anthropic coming for Legora and Figma, and Anthropic also working on the frontier problems that humanity faces today?
I would look to precedents from other big tech companies that have had a lot of different efforts, like Google and Microsoft, which coincidentally also try to solve climate change and cancer, but it's not their main business. They have their hands in a lot of different areas, but competitors still emerge.
You remember Google+, right?
Yeah.
That didn't go anywhere, right? It probably freaked some people out when it happened. You probably remember Threads. I don't know the current state of Threads.
Apparently, 400 million users, according to their marketing team. [laughter]
That's very interesting. I won't comment too much on that because I—
I'd love to see the engagement. [laughter]
I genuinely don't know anything about this.
[gasp]
But I think if you look to precedents here, large companies often try to make new bets and diversify, but they lose to companies that have intense focus on their market.
We'll see how it plays out, but I would wonder if there's anything to learn from history, with Google and Microsoft having many business units and many efforts, but a core business that has driven all of their revenue.
You know, I love Brendan. I remember texting him when there was the hack. It's tough when there's a hack because you don't know what to say, but, “I'm here for you,” you know, and thumbs-up. And I felt like it's such a VC thing because you're like, “I'm here for you. Good luck.” [__]—how helpful that is.
My question to you: how did that change your mindset and approach to product? It's a really hard thing to go through. I remember you were under intense pressure and stress, and I seriously am sorry about that, because it's horrible to go through. How did it change your product mindset?
I'm not an expert in security, but we hired a lot of experts in security and I listen to them. That's the main change: just larger investment and learning from the experts that we've brought in-house.
Are we entering a golden age for cyber? What I mean by that is, we're seeing a huge amount of AI-generated code, which in a lot of cases has holes. We're seeing Lovable and Replit, and you name it, produce a huge amount of output. The threat is going to increase much more significantly than we're anticipating.
Most likely, yes. Where we see it the most is—it's an interesting data type because it's competitive, and you can have these AlphaGo-type situations for cyber offense and defense, where you can have uncapped rewards and performance and the field's constantly moving.
We love this kind of stuff because it's like a game from a data perspective. We see very rapidly increasing demand for cyber-defensive capabilities via data and very interesting data types.
Can you help me understand? What data types do people want around security that they maybe didn't want before there was this explosion in demand?
13. Why Cybersecurity Data Will Never Hit the 90% Sufficiency Ceiling
I have to be careful not to reveal too much about customer work. The category is growing very quickly, and the nature of a lot of security work is that it's adversarial, right?
It's not this sufficiency-style work like, “Update a CRM and then you're good.” There's a constant cat-and-mouse game between the offensive capabilities and the defensive capabilities. To your point earlier about the 90% of enterprise workflows that can already be completed, there's never going to be that 90% for security because the goalposts are always going to move.
Cyber as a category is growing, and the nature of the data types is much more uncapped, evolving, and adversarial in terms of where the goalposts are.
Can you help me out here? You're Estonian by heritage.
14. Hiring in SF: Brutal Talent War & What Mercor Looks For
I say to European founders, SF is the worst place to start a company. It is impossible to acquire talent, impossible to afford it, and impossible to retain it. Is the talent war in SF as brutal as it seems?
Yeah, it's pretty brutal. It is very difficult to hire, and it is difficult to retain. I think it's harder than before, but it's easy when you're on a rocket ship, right? It's always easy when you're on a rocket ship to get someone. It's hard to make the right decisions about who you want to hire.
When you've made a bad hire, what did you not see that you wish you'd seen?
It's really hard to assess agency and ownership in the interview process.
I'm super freaking talented. I'm a bit of an asshole. I'm not a total asshole, but I'm a bit of a douche. Are you okay with that?
If you're super talented, yeah. The company culture here is high agency, high performance, and high ownership. Personalities can change. You can learn how to work with people better. We care about growth, and we want to hire people who give a shit.
That's a lot harder to coach into someone than smoothing things out with your colleagues, making sure that we have happy hours, and making sure people all get along. That kind of thing is easy to work out. You can have a couple of assholes; they get drinks together a few times, and then you smooth it out. It's really hard to make someone give a shit.
Yeah, also, if you hire multiple assholes, they can just hang out together. It's fine. It's a group hug.
We don't hire a lot of assholes.
No, no, I can also be like that. Happy hours, really?
We had a great off-site recently with our annotation team. We went to Tofino in Canada. It's on the west coast of Canada, and it's the only place you can surf. Everyone did surfing lessons, and we went to a floating sauna. It was a great time.
I thought it was great for the team and a great use of money, and everybody loved it. I think doing these outdoor activities where people are active is good.
Are you ready for a quick-fire round, dude?
What have you changed your mind on most in the last 12 months?
Honestly, I think it's probably the environment market. When we were starting it off last year, it was so complicated to do these deliveries, and it was so hard to get it to work that I thought it wasn't going to work out. I thought it wasn't going to scale, but then it did. I was pretty surprised.
What changed?
The demand was very high, and we got it to work. We just had to try a lot of different things to get environments to actually improve model performance. We kept going at it, and it ended up working.
I'm your little brother, and I'm studying computer science at university today. You sit me down and say, “Little brother, you should know this.” What should I know?
Get a real internship as soon as possible, because whatever you learn in school is probably going to be updated quickly.
Interesting. Why should I get a real internship? I know that sounds stupid, but should I start my own company? Should I join a fast-growing company? Should I join a super-established company where there are, you know, adults in the room, so to speak?
Maybe I'm biased, but join a fast-growing company in San Francisco. It doesn't need to have adults in the room, but somewhere on the frontier that's indicative of where the field is going. Somewhere a bit larger than 10 people, not super early, just to filter out the companies that might not go anywhere.
Would you say that you're too late for me?
No. No, we still act like a startup.
How many people do you have?
Maybe 500. It's a cult. Culturally, we're a startup. We're paranoid, we're in the office all the time, and we're fast-moving. We want to hold on to that as long as possible.
[Laughter.] I love it. That's amazing. Totally. Absolutely. [Laughter.]
Which competitor do you respect most, and why?
I don't think about competitors too much. They're all kind of—even in that, they're all behind us. It's a bit of an odd answer, but we really try not to think about them as much as we try to think about our customers.
I respect our customers a lot. I love the work that they're doing. We stay on top of what competitors are doing, but every time I look at one of their websites, they're just doing something we did a week or a month ago. They write a blog; we write a blog, and someone else writes a blog a week later that's the exact same thing. We make an update to our website, and someone else makes an update to their website that's the exact same thing. So, I spend a lot more time—
What about Surge?
It's happened before. They're a bit out there. Honestly, I don't spend that much time thinking about them because I spend more time thinking about customers. We've seen it. They're a bit out there in that they don't copy us as much, and they do seem a bit different from others in the field. It's hard to say why. They're very secretive.
Yeah. Are you kidding me? Yes, absolutely. I totally get that. Can you please paint the bull case for how [likely Mercor] is a $200 billion company?
It looks like we sell services. We're basically a tech-enabled services company. Our services are incredibly valuable in driving revenue gains for our customers, primarily through better model capabilities. Evals and training data are the primary bottlenecks in model performance right now.
If every enterprise needs to have specialized proprietary models, even if the capabilities start to saturate, the evals serve as the PRD for exactly what you want, but also the optimization objective for better performance. As long as better models are valuable to the economy, there will be demand for eval sets and training sets.
If we can make that process faster and faster, we can serve a growing demand for human data for evals and training. We also have a growing agent-deployment enterprise arm as well.
What line of revenue do you not have today that you think will be very significant in 3 years' time?
I think that real-world, physical data is going to grow significantly over the next 3 years. Robotics is an interesting area for us. The data market for robotics is nascent relative to GenAI and autonomous vehicles as well, and we think that's going to grow a lot.
Do you scale supply ahead of demand?
At times, we retain exceptional talent to do work that might be valuable in the future. We can do off-the-shelf data creation to make use of supply when demand is low, and then resell that data later. In that case, we do. Otherwise, we don't.
What's the best piece of advice you've ever been given?
I got a lot of advice to join small companies, join startups, and move to San Francisco. I grew up in Canada, and I went to school in Toronto. I followed that advice. I think it was great.
I've loved living out here, and I like small companies. I like fast-growing companies. It's been super fun and great for my career.
Final one for you. What are you most excited about that you don't think enough people are talking about?
Probably the same answer as before: the 3-years-out opportunity in robotics. I think there's a lot of discussion around robotics on Twitter.
I'm sorry, dear. Can you just help me out here? This is where I get in trouble. It's Friday afternoon. It's past 6:00. Fuck it, I can say what I want. I don't get it, okay?
Whenever you watch a robotics talk, you see this demo of a terribly moving robot around a home. And after watching it take 1 water out of a fridge in 15 minutes, it goes, “And Brandon was in the other room all along.”
And you're like, “Are you fucking kidding me? I had this absolute spaco in my kitchen for 15 minutes getting a water, and Brandon was in my bathroom doing it? That's where we're at?” What am I not seeing? How do you help me get excited?
Yeah, I think if you go back a decade or so, self-driving cars had people in them all the time. You would see Cruise driving around San Francisco, and there was a person in it for years—for years, right? But now I take Waymo more than I take Uber.
I'm thrilled for you. Welcome to London. We still have these people in cars. [Laughter.] I love it, but it's in 1 city. It can't deal with very ambiguous data. It's pretty irrelevant.
It's made leaps and bounds in the past decade, at least in San Francisco and in Austin and Phoenix. It's tough because, yeah, I guess the distribution is unequal, but it's an incredible service here in San Francisco, and people here use it a lot.
Technically, it works. There might be regulatory challenges or other challenges with scaling, but—
Do you think we'll hit a ChatGPT moment with robotics that will cause an inflection in usage and adoption?
I think so.
Yeah, I think so. But I think it might play out similarly to driverless cars, where it's really hard to scale physical things as opposed to software. It might be more of a Waymo, robotaxi, Cruise-type moment than a ChatGPT moment. But I think the progress will be there, yeah.
Dude, you have been fantastic. Thank you so much for putting up with this incredibly wayward, poorly structured conversation, which was brilliant. I so appreciate you putting up with it.
Thanks for having me. Yeah, it was super fun.