Mercor CEO:为什么应用层公司没有护城河,以及招聘 AI 研究员的成本
- Mercor CEO 正面回应黑客传闻:事件确实发生过——攻击者利用“一群编码代理”获取了系统访问权限——但营收停滞的说法不实。 Mercor 在过去60天新增净 ARR 3亿美元,第一时间聘请 Mandiant,并将安全列为公司第7项价值观。他表示,Twitter 上的叙事还包括“一个非常有影响力、投资了多家竞争对手的人”发帖,声称中国方面访问了数据,但这完全不实。
- 这些是实打实的收入,不是 GMV:客户购买的是任务(例如支付1000美元完成一项模型改进),毛利率为30-40%,而 Mercor 覆盖完整链条——专家招募、平台、AI 项目管理和质量检查。 这项业务“非常赚钱”,账上现金超过5亿美元,现金多于公司历史累计融资额;自2025年秋季、约4亿美元年化收入、100亿美元融资轮以来,业务规模已“接近增长4倍”。
- 核心投资判断是:未来12个月,OpenAI/Anthropic 上游的基础设施公司将跑赢下游的应用层公司。 原因在于“模型就是产品”,而应用层的防御性越来越难建立。2025年是让模型生成一次 PR 的一年;“2026年要解决的是,如何让模型端到端克隆 Slack”——这些能力将在12个月内进入模型。SaaS 能否存活的试金石是网络效应(Salesforce 的集成、Slack Connect、Carta);没有网络效应的公司将面临极大困难。
- 未来5年内,平均企业在 token 上的支出将超过人力成本——Mercor 已经如此: “现在,我们为内部代理支付的 token 费用,已经超过员工人力成本。” Benioff 每年向 Anthropic 支出的3亿美元,仅相当于 Salesforce 开发人员薪酬的约3.8%,说明这一转变仍处于早期。每家《财富》500强公司都需要一个按工作流建立的评估“记录系统”;这会让 API 层商品化(零切换成本,每2个月出现一个新的前沿模型),而粘性将留在工作流中。
- 这位嘉宾认为,5年后 OpenAI/Anthropic 中至少有一家估值将超过10万亿美元。 这是他的态度转变:他过去怀疑实验室能否保持定价权,但“这些公司的收入爬坡速度”让他相信它们会成为全球最有价值的公司。然而,他同时预计,5年后大多数推理任务将运行在开源、微调或蒸馏模型上,而非前沿模型;按工作流进行评估往往是“价格性能的10倍杠杆”。Nvidia 可能因多芯片格局失去垄断,但即使只占“遥遥领先的全球最大市场”30-40%的份额,仍会是全球最有价值的公司。
- 训练代理是“历史上创造速度最快的职业类别”: Mercor 每天向其500万+人才网络支付300万美元,预计12个月内大致增长至3倍,甚至4倍。在 Mercor 的 Apex 基准上,前沿模型目前得分约40%,而仅仅12个月前 o1 只有1%。
- 定价弹性方面: Harry 提到 Nebius 提价30%却对需求毫无影响;Mercor“有能力一夜之间将需求翻倍”,但受限于产能——不过嘉宾认为,定价必须在短期优化与竞争之间取得平衡,因为“高利润率会招来竞争”。
- 人才市场已经失衡: 一名候选人拿到了 Meta 超级智能团队每年2000万美元的现金报价,顶尖 AI 研究员的成本是“每年数千万美元的股票”,需求与供给之比达到10:1。他认为欧洲已经因人才网络效应输掉了模型竞赛,应该接受这一现实——实验室只需“在法国招聘1万人,教模型法国法律”,主权论据就会被削弱。
1. 黑客事件被还原:过去60天新增净 ARR 3亿美元
- Harry 先抛出传闻:公司被黑,收入自此停滞。CEO 的回应是:“事件确实发生过,其他部分全是假的。” Mercor 立即聘请 Mandiant 及其他安全公司,主动与客户沟通,此后“与所有前沿实验室的合作关系都进一步扩大,并在过去60天新增净 ARR 3亿美元”。公司还将安全列为第7项价值观,“确保它真正融入文化”。
- Twitter 上的风波比现实更严重:他提到“一个非常有影响力、投资了多家竞争对手的人,发帖称我们的全部数据都被中国访问了,但这完全不实”;律师建议不要明确反击。他从危机中得出的结论是:这“绝对算不上公司历史上压力最大的时刻”,深厚的客户关系和完整的内部信息,比 X 上的“回音室”更重要。
2. 攻击者动用了代理群:AI 网络安全的黄金时代将至
- 入侵方式本身揭示了下一个市场机会:“攻击者使用了一群编码代理,帮助其获取系统访问权限。”人类攻击者只能按人的速度审查代码;代理群则能“非常彻底地审查整个代码库”,让攻击者的推进速度大幅提升。
- 嘉宾的判断是:“AI 网络安全工程工具将迎来巨大繁荣。”客户正专注于提升模型的网络防御能力,目标是打造“能够保护每一家企业的最佳 AI 安全工程师”;而这一轮又一轮的事件,“才刚刚开始”。
3. 客户与传闻:OpenAI“比以往更强”,Meta 暂停合作,没有 Micro 1 offer
- Mercor 是否因黑客事件失去 OpenAI 和 Meta?“假的。我们与 OpenAI 的关系比以往任何时候都更强。”Meta 是例外——“目前双方关系仍处于暂停状态”。CEO 表示,Meta 还有其他事项在推进,并提到 Scale 的收购可能是原因之一:Meta 自然会更多与 Scale 合作。此后,其他所有前沿实验室都扩大了与 Mercor 的合作。他直接否认 Harry 关于 Handshake 收入飙升源于 Meta 将预算从 Mercor 转走的推测:“不是这样”,但不愿展开。
- 关于 Micro 1 以“数百万美元签约包”挖人的说法:一名员工对外发送消息,提出50万美元签约奖金,部分收件人参加了第一次会谈——“我们没有向任何 Micro 1 的人发出过一份 offer。”媒体却把这些消息包装成正式法律 offer。
- 关于 Amazon 以130亿美元收购的传闻:“那是假的。”如果出价300亿美元,他会不会出售?“不会……我可以拿着几十亿美元现金离开,但这不是驱动我的东西。”他的使命是回答“人类如何融入经济”,而如果公司不是独立主体,“我们实现这一愿景的概率不会那么高”。
4. 收入真实存在,毛利率30-40%,业务几乎从未真正烧钱
- 针对“这只是 GMV”的质疑,客户购买的是端到端任务——“他们会支付1000美元,购买一项能够改进模型的任务”——Mercor 从专家招募、AI 项目经理到自动化质量检查全部负责,毛利率为30-40%。“我们和 Uber 依靠司机网络驱动业务一样,依靠人才网络,但人才网络不是最终产品。”收入“远高于”公开披露的约10亿美元。
- 垂直整合会产生复利效应:下游质量信号反过来改善上游专家招募,而数据价值呈幂律分布——“在1万个任务组成的数据集中,前2000个任务会创造大部分价值”,因此质量本身会带来定价权。
- 财务状况方面:“种子轮之后我们烧了50万美元,此后基本一直盈利……我们的现金比历史累计融资额还多。”公司账上现金超过5亿美元,刻意为市场回归理性做准备:在融资狂热期,任何人都可以承受负利润;“等市场回到现实……那时才会出现整合期”。
5. 训练代理是历史上创造速度最快的职业类别
- 面对裁员潮(Intuit 1.6万人、Meta 8000人、ClickUp 裁员22%),嘉宾的回答是“劳动总量谬误”:过去250年生产率提高25倍,“相当于自动化了某个人约96%的工作”,结果是创造了更多岗位,而不是更少。Harry 的反驳在于速度:过去的革命需要几十年;如今“用 Nano Banana Pro,我几乎可以一夜之间让媒体公司的所有设计师失去工作”。嘉宾承认,岗位替代将“非常显著”,但认为如今的经济也更擅长创造新的职业类别。
- Mercor 就是样本:公司每天向人才网络支付“超过300万美元”——这正是“历史上创造速度最快的职业类别”;他预计未来12个月大致增长3倍,“甚至可能4倍”。在 Apex 基准(Mercor 衡量顾问、银行家、律师和工程师 AI 生产力的指数)上,前沿模型目前得分约40%;12个月前,o1 的得分还只有1%。
- 未来5年的新职业是代理训练师。“所有知识工作都在向训练代理汇聚,因为做一次在结构上更高效。”客服代表不再重复回答数百张工单,而是训练一个代理。人类仍能贡献的部分是隐性知识:“人脑中存在海量模型无法从其他地方获取的上下文”;而随着模型推理能力提升,数据清洗本身会交给模型完成。
6. 横向聚合胜过细分数据供应商:实验室想要一个灵活的合作伙伴
- 针对外界关于数据供应会走向拆分的判断(外科医生戴摄像头出售医疗数据),嘉宾表示:“我们为律师构建的数据形态,往往与为医生构建的数据形态非常相似。” Mercor 拥有500万人才网络,且成员会推荐朋友,因此寻找边际上的医生并不困难;实验室更希望与一个具备横向能力的供应商合作,而不是面对“100个不同领域、需要分别培训的供应商”。
- 实验室真正需要的是全覆盖能力:“能够把所有你可以输入 Google Workspace 的东西,以及在经济中的每个职业类别里你希望从另一端得到的所有东西,全部覆盖。”同时,他们需要既是专家、又是 ChatGPT 或 Claude 的“高级用户”,能够找出模型出错的地方。
7. 直升机、法拉利、军舰:2年内从2300万美元估值走到100亿美元
- 按融资轮次回看:2023年9月种子轮时,年化收入约100万美元;General Catalyst 在36小时内发出 term sheet,投后估值2300万美元、融资230万美元。A 轮时,Benchmark 的 Victor 在第二次会面被拒后追问:“你坐过直升机吗?”最终公司以约250万美元收入拿到2.5亿美元投后估值。B 轮时,Felicis 把创始人带上私人飞机,在拉斯维加斯 F1 赛道开法拉利——收入2000万美元时估值20亿美元,增长100倍。随后在2025年9月或10月,公司以约4亿美元年化收入达到100亿美元估值,约25倍;“此后业务又接近增长4倍。”下一轮“估值可能高得多”,但公司盈利,因此并不着急。Harry 按交通工具统计:“D 轮现在需要一艘军舰。”
- 对这些疯狂估值的辩护是:连续6个月实现50%的月环比增长,随后又维持了12个多月——“我当时预计年末收入年化达到5000万美元,次年年底达到5亿美元……而我们超过了预测。”事后回看最不舒服的一轮是 B 轮:“在收入250万美元时做到100倍,和在收入2000万美元时做到100倍,完全是两回事。”
8. 模型就是产品:应用层防御性越来越难建立
- 撑起本期节目的核心判断是:未来12个月,“Anthropic 和 OpenAI 上游的基础设施公司,将明显优于下游的应用层公司”。第一个原因是,“过去2年,所有人越来越清楚地认识到,模型就是产品”;端到端训练的模型胜过所有拼装出来的抽象层,包括拖拽式代理构建器。第二个原因是软件复刻速度:“2025年要解决的是,如何让模型在代码库里生成一次 PR;2026年要解决的是,如何让模型端到端克隆 Slack。”他预计这些能力将在12个月内进入模型;Claude CoWork 要增加医疗和法律领域的能力,并不是很大的跨越。
- Harry 以 Legalese 为例反驳:深度的律师专属工作流、GTM 和 CS 团队都构成壁垒——“防御性就在这里。请你反驳。”嘉宾回应称,真正的护城河不是售前 GTM,而是“前置部署式服务”。一个每年支付100万美元 SaaS 费用的精明客户,“只需告诉 Claude 把它复制出来”;但一个经过公司隐性知识训练的代理,“极其差异化,也很难复制”。这正是 Sequoia 所说的:“服务就是新软件。”
- 判断传统 SaaS 能否存活的试金石是网络效应。Salesforce 的集成市场、Slack Connect、Carta 的跨公司关系图谱,都建立在真实护城河上,能够将产品迭代速度提升10倍。“没有网络效应的公司将陷入极大困境……这就是判断一家公司是否会变得一文不值的试金石。”
- 服务自动化正在 Mercor 内部发生:一个 AI 项目经理刚刚端到端完成了第一个项目——招募专家、回答问题、用自己的编码工具搭建标注工具——这些工作过去由约100-150人的交付组织完成。“专家们向 AI 项目经理汇报时,体验都非常好。”“我们正在实时看到,服务正在被自动化。”
9. Token 支出超过工资单:评估体系让 API 层商品化
- 这个头条数据是嘉宾随口说出的:“现在,我们为内部代理支付的 token 费用,已经超过员工人力成本。”他对未来5年的判断是:“平均企业在算力上的支出将超过人力成本。”相比之下,Benioff 目前每年向 Anthropic 支出的3亿美元,只相当于 Salesforce 开发人员薪酬的约3.8%。尽管效率提升,token 成本仍在上涨,这是 Jevons 悖论的一个“有趣案例”。
- 企业将使用的机制,是按工作流建立的评估体系,作为“记录系统”。Mercor 为每个内部代理运行一套评估:已完成500万+次面试的面试代理、候选人排序、会计和欺诈检测等,都据此在价格性能的 Pareto 前沿上选择模型。企业会利用这些评估“让模型层商品化”,追求零切换成本的完全竞争。API 层会商品化——每2个月出现一个新的前沿模型,按评估得分即可热插拔;但工作流粘性仍然存在:“我有所有这些流程在 Claude Code 里运行,可能不会花时间把它们迁走。”
- 工作流评估通过将任务蒸馏到开源模型上,“往往是价格性能的10倍杠杆”。因此他的预测是两段式的:OpenAI 和 Anthropic 都是“不可思议的投资”,未来需求还会增长4—5个数量级;但“5年后,大多数推理任务将使用开源、定制微调或蒸馏模型,而不是前沿模型”。估值方面:“我完全可以想象其中一家成为10万亿美元公司……至少有一家市值超过10万亿美元。”在快速问答中,他说自己已经改变看法:实验室的收入爬坡,让他从怀疑其定价权转变为“确信它们将成为全球最有价值的公司”。
- 与其直接买 Nvidia?“不是个疯狂的想法。”但多芯片格局正在形成:Cerebras 已经在执行,Etched 也在推进,实验室还在自研芯片,因此 Nvidia 的垄断可能消退。“即使它们在遥遥领先的全球最大市场中只拿到30%或40%的份额,也仍然会是全球最有价值的公司。”Harry 提到 Nebius 提价30%却没有需求影响;Mercor“有能力一夜之间让需求翻倍”,但没有相应产能。嘉宾认为,定价必须在赢下这10年与招来竞争之间取得平衡,因为“高利润率会招来竞争”。
10. 2000万美元现金 offer、输掉竞赛的欧洲,以及取消美国底层50%人口的所得税
- 人才市场的一个真实案例是:嘉宾正在招聘的一名候选人,手里有一个“每年2000万美元现金”的 offer,来自 TBD——Meta 的超级智能团队,股票但具备流动性。顶尖研究员的成本是“每年数千万美元的股票”,需求与供给之比达到10:1。他预计最顶尖人才的竞价还会继续升级,但受实验室训练的人才供给会在“第99百分位”趋于正常。
- 关于欧洲(Harry 说,Mistral 的排名“像欧洲歌唱大赛一样——大概排在最末”),欧洲输掉模型竞赛,根源在于人才网络效应:法国最聪明的研究员聚集到 OpenAI、Anthropic 和 DeepMind,形成“美国所拥有的、最大的经济优势之一,也是最大的地缘政治优势之一”。他的建议是接受现实,保留部分后训练和应用能力,不要“激进地与 Anthropic 正面竞争”。主权论据同样有限:“实验室只会在法国招聘1万人,教模型如何更好地处理法国法律”,迁移学习会完成剩下的工作。
- 他大一时写过、如今被 Bezos 转发的文章主张:取消美国底层50%人口的所得税。这部分税收仅占政府收入约3%,而“经济中最大的正外部性是就业”,但政府恰恰对就业征税。他建议改为征收资本利得税,尤其是短期资本利得税,以及碳税:“在我看来,不对碳征税,却对美国底层一半人口征税,简直不可理喻。”Harry 激烈反驳:“我带着最大的尊重这么说。这就是错的……你跑去一个没有资本利得税的地方,最后税收收入全部消失。”嘉宾同意,任何方案都需要对资本外逃进行“敏感性分析”。
- 快问快答中还有几条值得保留:约一半数据供应商竞争对手“其实只是交易型人才市场”;他最尊重的竞争者是 Surge 的 Edwin,因为对方始终贴近研究;IPO 会在“未来几年”发生,但不是今年或明年;996 传闻不实——“我们从未强制规定工时”,不过他和 Adarsh “从醒来到睡觉一直在工作”。他认为最善良的是 Prod 非营利社区:每周开会、提供营运资金,并带来第一个大客户——“他们没有拿任何股权……如果没有这些人,Mercor 根本不会存在。”
Ready to go? Brendan, it is so good to have you in the studio, dude. Thank you so much for joining me in person.
Super excited to be here. Thanks for having me, Harry.
1. True or False: Mercor lost Meta & OpenAI as a customer with the hack?
I was thinking about how we're going to structure this, and I thought, you know what? There are quite a lot of myths or rumors around Mercor. Given it's our second time, I thought I could break the ice and go straight for them.
Myth number 1 that we're going to tackle is that there was a hack or leak, or whatever terminology you call it, and revenue's been flat. What's really happening with Mercor? True or false?
There was an incident. All of the other parts are false, and we obviously handled it very quickly. We were in touch with customers, and we moved incredibly fast, engaging Mandiant and a bunch of other security consulting firms. The company's been crushing it ever since.
We've expanded our relationships with all of the frontier labs and added 300 million in net new ARR in the last 60 days.
300 million in 60 days? Fuck me.
It's been pretty crazy, yeah. Keeping us busy.
I'm sorry, I just have to ask: where were you when you found out about the hack, and what did you do?
It was a Saturday, so I was in the office, and I was talking with our engineering team. The initial thing, of course, was figuring out how we were communicating this to customers and trying to be very proactive about understanding exactly what happened, what was accessed, et cetera.
Then it was about how we communicated this to the experts and just moved to contain it—moving quickly on the comms—and then, from there, making sure that we put all of the right things in place so that it never happens again.
There's a brilliant poem by the poet Rudyard Kipling that essentially says you have to keep your head when all about you are losing theirs. That is a time when everyone is losing theirs.
Definitely.
I don't by any means want to be patronizing. We're both young; you're younger than me. What do you do to stay calm when that is an “oh, fuck” moment?
Throughout the lifetime of the business, I have been through a lot of very stressful moments. That was definitely stressful, but it definitely wasn't close to the most stressful one.
I mean, seriously. There have been plenty of times when I'm freaking out about making sure we got something right with a customer or whatever it is. But I think part of it is that there was this broad perception on Twitter that was much more exaggerated than what actually happened within the business.
Having as thorough an understanding as possible of what actually happened, and having really strong relationships with customers, gave us a lot of confidence that we would get through it and be on the other side even stronger.
We used to have 6 values as a company, but we added a 7th value of security to make sure it's very ingrained in the culture. I think it's that confidence that we know what's going on, and that there's sort of this echo chamber on X that we need to hedge against a little bit.
Do you pay attention to it? Do founders need to pay attention to it?
Definitely. I think founders need to pay attention to it. We had an all-hands with the company where we just laid out, “Here's exactly what's happening. Here's the trajectory of the business.” I think that was very helpful to the entire team.
It was definitely annoying that there were all of these people saying things that didn't actually happen, and we couldn't quite speak out against them too explicitly. Otherwise, there's going to be the Twitter mob circulating, along with all these recommendations from lawyers, et cetera.
The whole thing is that there are often a lot of people with economic incentives behind the scenes who will absolutely trounce you and be very negative because they're aligned to a competitor.
We're in a YC company that's been through a lot of shit in the last few days, and their competitor has a lot of people behind them through various different means. The alignment is not obvious, but it really sounds out on Twitter.
That is exactly what happened. I can even think of one person who's very prominent, who's invested in multiple competitors, and who made this tweet about how all of our data was getting accessed by China when it was totally untrue.
You mentioned adding security as a 7th pillar there. We've seen so many hacks. It's almost become normalized, as awful as that sounds. Are we about to enter a golden age of cyber, given the new threats awakened by AI?
2. Are We Entering a Golden Age of Cyber?
I think so. We're even seeing this on the customer side, where our customers are obviously very focused on how we improve the model's cyber-defense capabilities so that we can have the best AI security engineer, able to defend every enterprise from all of these attacks.
In our incident, it was the attacker that used a swarm of coding agents to help get access to the system, as is happening in a lot of these attacks. I think there's going to be an enormous boom in AI security engineering tools and various forms of defense that are able to help protect companies against all of the increasing waves of cyber incidents that are just getting started.
Can I just be very naive and dumb here? How do swarms of coding agents make for such dangerous and malicious actors? Why does that actually work?
When a normal attacker is trying to find vulnerabilities, they can only review so much code and go through a certain portion of it at a human speed, bound by the number of people on their team. Whereas, when they're using swarms of agents, they're able to be very exhaustive in reviewing the entire codebase, looking at the entire front end, and examining all the different things they've accessed.
That has allowed a lot of these attackers to move much more quickly. We've been exploring various collaborations with customers and how we can strengthen their cyber-defense capabilities to hedge against exactly this type of attack as well.
Got you. In terms of those various customers, true or false: you lost OpenAI and Meta as customers in the hack?
False. Our relationship with OpenAI is stronger than ever. Obviously, I can't speak too much to specific customer relationships, though.
Can I push on Meta?
Of course. I think that Meta—currently, the relationship is still paused. Every other one of the frontier labs has grown their relationship with us since, and the company has been crushing it, but they're the only one that is—
And it would be paused just because of the security?
There are other things happening there. Obviously, I think that Meta's a unique customer because of the Scale acquisition, and so naturally they're going to work with Scale more. But I don't want to speak too much to the specifics of a customer.
Because I thought when you saw Handshake's revenue go parabolically up, it was just Meta shifting spend from you to them. Is that not true?
That's not true.
Interesting. What is that, then?
I probably shouldn't speak too granularly to that, but, yeah.
Totally cool. Okay, but so we have—
I'll speak to everything except customers.
But we have not lost OpenAI.
Got you.
Cool.
Stronger than ever.
Because I got told by many of your customers before the show that they definitely have great relationships with you.
Good. Thank you.
You're welcome. I've read this article. You've been trying to poach micro1 team members with signing packages in the millions.
We have not extended a single offer to someone from micro1.
So, no millions?
No millions.
Why does that come about? I read this article.
The reason for the article was that someone on our team sent an outbound message to some people at micro1, saying that we were hiring a variety of people with these very high signing bonuses. I think one of them said $500,000 as a potential signing bonus.
They took first meetings, but we didn't move forward with offers to anyone. Obviously, the way that gets framed to the press is, “These are offers that are going out,” when there's a giant distinction between one of our employees sending a message to one of their employees and actually sending out a legal offer letter.
Love it. Press is a wonderful thing, huh?
Totally.
Okay, next myth, but I'm enjoying this. This should be a new show: MythBusters. You might get uncomfortable with this one.
I heard a rumor that Amazon tried to acquire you for $13 billion. True or false?
That one is false. I obviously can't speak too much to other acquisition-type stuff, so I'll reserve any comments on future acquisition questions there.
Would you sell for $30 billion?
No, I wouldn't. Ultimately, we've gotten a lot of acquisition interest, and I could walk away with billions of dollars in cash. The thing is, that's just not what motivates me.
3. AI, Jobs & Layoffs: How Do Humans Fit Into the New Economy?
I'm very motivated by how we solve this incredibly important problem in the world of how humans fit into the economy, and I feel like we have the opportunity to build a legendary company in creating this new category of work. Our probability of executing on that vision wouldn't be as high if we weren't an independent company.
How humans fit into the economy.
Mhm.
Fascinating. When we look at the news, we see Intuit lays off 16,000, Meta lays off 8,000 at 4:00 a.m., LinkedIn 1,000, Coinbase, and so on. ClickUp now has 22% going. It's hard for people to see how humans are going to fit into that new economy.
Totally. I think, to some extent, I share that concern. I believe there's certainly going to be many more jobs in 10 years than there are today, but there's also going to be a lot of job displacement along the way.
Amidst all of these layoffs, I think the most important question is understanding which jobs AI is able to do and which jobs AI is not able to do. We're building a ton of initiatives, such as the AI Productivity Index, or Apex, that are becoming the industry standard in answering that question and measuring across all of the different popular job categories that people are talking about, ranging from consultants to investment bankers to lawyers to software engineers. What are the actual tasks within those jobs that AI can automate, and what are the tasks that it can't?
With the greatest respect, does that not change so quickly? When you saw Andrej Karpathy talk about how he uses coding agents, it was like, “Oh, I use it for 20% of the work.” Then it's like, “Oh, it does 80%, and I do the final 20%,” within a 6-month period.
Definitely. Another example of that is on Apex: the frontier model right now is at about 40%, and 12 months ago, the frontier model was o1, which was scoring 1%. That's been the progress of the last 12 months, and obviously we expect it to continue and be fairly significant.
I think the key thing is that everyone underestimates the elasticity of demand for increased productivity in the economy. Ultimately, over the last 250 years, we've increased productivity by 25x, equivalent to automating about 96% of someone's job.
During every technology revolution, ranging from the agricultural revolution to the Industrial Revolution to the computer revolution, people feared that there would be this enormous job displacement because of the lump-of-labor fallacy, where people assumed that there was a fixed amount of things that had to be done. When we made people more productive, that would all of a sudden mean that there were fewer jobs.
Yet, 250 years later, there are more jobs than ever before. It's because we have no shortage of problems to solve as a society, right? We still need to solve climate change, cure cancer, and do all of these other new things.
I buy that completely. What I don't buy is the speed of transition. When you look at the Industrial Revolution and the agricultural revolution, it took multidecade cycles to implement and train new technologies to do what humans did.
Yeah.
Now, with Nano Banana Pro, I can get rid of all the designers in my media company pretty much overnight.
The thing I agree with you about is displacement. I agree there's going to be a very significant amount of displacement, but I also think that the economy is becoming much more effective at creating new job categories and allocating new labor.
A great example is what we're doing in that: we're now paying out over $3 million a day in the fastest job category ever created in history. I expect that's going to continue growing exponentially from here.
I think there are going to be so many new job categories created across everything within AI, such as training agents for deployed engineering and building data centers, all the way to all of the problems that we otherwise wouldn't have been able to address as a society. How do we build solutions to climate change? How do we have more people working on rockets to explore space, et cetera?
Totally get you. You said $3 million per day paid out. What is that in 12 months' time?
In 12 months' time, that's probably about triple that.
$9 million?
Mhm.
Do you think you're being ambitious enough?
Maybe it's quadruple that.
We have internal projections that are always much more aggressive than our external projections, but we almost doubled our projections last year.
What new role will we have in 5 years that does not exist today?
One of the largest things that people underestimate, both in the context of AI labs as well as within the enterprise, is how significant of a job category it is going to be to train agents. What we're seeing is that all knowledge work is converging on training agents because it is structurally more efficient to do something once.
Instead of having a customer support representative who is redundantly responding to hundreds of tickets, they're going to train an agent how to do that once. Instead of having a lawyer who is redundantly doing dozens of similar redlines on commercial contracts, they're going to train an agent how to automate that.
Even when you're playing around with Claude, you see that there are so many repetitive workflows for how you prepare for a meeting or draft emails or whatever it is, where it's just much more efficient for you to train the agent how to do that activity so that you can amortize that over the entire useful life cycle rather than doing it redundantly yourself.
I think there's going to be this enormous paradigm shift as agents enter the workforce and everyone begins to manage them.
Can I ask you, when we think about enterprise adoption, I think one of the biggest problems we have is data structures and data cleanliness.
Mhm.
I interviewed a guest the other day, and they said we'll have data cleaner as one of the most important jobs in the next 5 years. Is data structure and data cleanliness the biggest barrier to enterprise adoption?
I agree in part. Certainly, the models need to have access to data to perform their jobs effectively, but the caveat is that they'll be able to clean the data themselves fairly effectively as reasoning capabilities go up.
The thing that humans will need to contribute is all of the tacit knowledge within the organization that isn't written down. I found that when I try to get agents to do all of these workflows throughout my career, there's just an enormous amount of context that lives in people's heads that the agents need to have access to in order to perform effectively.
So much of that is going to be the new job of employees: How do we codify all of this knowledge? How do we train agents so that they're able to perform these tasks effectively across every function in the organization?
I'm sorry for digging down, but you said reasoning capabilities will allow enterprises to clean data more efficiently. Why?
The reason is that if a model is able to, for example, read through every message written in Slack over the last 6 months, the model can presumably structure a table of all the different customer conversations that happened in the CRM.
I don't expect humans to be doing that type of work—how do we structure data, how do we classify it, et cetera? But I do think that humans will do the things that models inherently can't do, such as tacit knowledge.
When we look at the market for being a data provider to some of the largest models in the world, it's such a large market that you're seeing the unbundling of it into such verticals. I met a real-world medical data provider serving them the other day. Basically, you have surgeons with video cameras on, and they record all the real-world data.
Do we see the mass unbundling of the data provider market? Is that how it plays out?
It's interesting. We're doing a ton of data collection in the physical world as well, especially across skill domains where you have electricians, mechanics, and scientists strapping cameras to their heads to record things.
I think there's always going to be some degree of value in niche vendors that are able to go really deep in a specific vertical, but what we're finding is that there's enormous value to aggregation and economies of scale. When we have this talent network of over 5 million people who are able to refer their friends, it's just so much easier for us to find the marginal doctor because we have that enormous talent network that can refer us to their friends.
Even more importantly, the data schemas that we would build for a lawyer are often very similar to the kinds of data schemas that we would build for a doctor. All of the tooling that we build is very, very cross-applicable.
And that’s the way that most labs have been scaling out their data quite horizontally. For that reason, we are finding that the labs tend to prefer partnering with a very horizontally capable vendor that is able to flex across all of the different verticals and scale extremely quickly, rather than working with 100 different vendors that they have to train for the same data schema in 100 different domains.
Do you think we’ll go through a period of consolidation? There are a huge amount of them, where you’ll actually end up buying the medical data provider because it’s a really important part of medical data. Do you think you will have that period of consolidation?
I think there will. In most markets, when the markets are so frothy and anyone can get funding and run negative margins, of course there’s going to be this proliferation of companies that pop up. When markets come back to earth and there are natural corrections, that’s when there are periods of consolidation. We view having over $500 million in cash and a super-profitable business as a significant asset, allowing us to be prepared for when there is a market correction and to make sure that we consolidate market share.
You’re profitable today?
Mhm. Very profitable.
How long have you been profitable for?
4. Rejecting a $30B Acquisition
We’ve never really burned cash. We burned $500,000 after our seed round, and then from there we’ve pretty much been profitable ever since. We have more cash than we’ve ever raised, and it’s just because the business has grown so quickly that we obviously try to redeploy capital as fast as we can to invest in growth. But the business has grown so fast that we haven’t been able to redeploy capital commensurate with that.
Can I ask you a myth-buster, which is: after we had Adarsh on the show the first time, people were like, “Oh, the revenue’s not real revenue. It’s like GMV.” When we understand your revenue, what’s the revenue, say?
I can’t share the exact revenue number, but it’s dramatically higher than whatever has been posted publicly.
Let’s give a ballpark, just because my simple numbers—I’m genuinely not asking for a billion. It’s just an easy number.
More than that, but yeah.
Okay. Let’s say a billion because it’s easy for my brain. We have a billion. Is that like sales for Airbnb, and then they get 20% of that?
The revenue has a 30 to 40% gross margin, but the key distinction—and why it’s not GMV but is revenue—is that the experts are actually only one part of the broader value chain that we deliver to customers. When a customer comes to us, they’re generally buying tasks where they would say, “Hey, they’ll pay $1,000 for this task that delivers model improvement.”
Then we do the end-to-end process associated with it: How do we find the experts? How do we hire the experts? How do we build the platform that the experts work on so they can do the work? How do we have our AI project manager manage the experts to automate all of the coordination involved in producing this data? How do we have automated quality checks, et cetera, to produce the end product of the task that we’re delivering for our customer?
That’s the large distinction. We’re powered by a talent network in the same way that Uber is powered by a driver network, but that’s not the end product, in the same way as some of those marketplace businesses.
What’s so interesting for me, and you can tell me if this is right or not, is that you’ve seen the evolution of this business from, “Hey, we provide raw data back to the largest models in the world.” That was how it started.
Mhm.
And now it’s end-to-end. We provide it fully, then we send it to you, we make sure everything’s ready, and it’s full-stack.
Exactly. Very vertically integrated.
So many parts of the downstream signal inform the upstream signal. We can use the quality checks on how high-caliber each of the individual data points is to understand exactly what types of experts we should be onboarding to achieve the data that drives the most model improvement.
There’s often this very power-law nature of data that drives model improvement, and out of a data set of 10,000 tasks, the top 2,000 tasks will create the majority of the value. It allows vendors that are extremely high-quality to be super differentiated in terms of pricing power, because quality is the X factor that becomes dramatically more valuable than any other dimension.
What task is super high-value? Is it medical, financial modeling, that kind of thing?
It corresponds extremely closely to economic value. If you go through the top 5 demands that we serve, it would be software engineering, finance, medicine, law, consulting, et cetera, and the super-long-horizon tasks within those.
I think we’re moving away from the paradigm of, “How do we get an investment banker to prepare a financial model?” and moving towards the paradigm of, “How do we get a banker who can talk with 5 different colleagues, wait to hear back on their responses, and prepare an entire slide deck with a deliverable that includes the financial model and the analysis in a multi-week-long project?”
Those are the kinds of tasks that we need to be building to push the frontier of research and evaluation, so that those are the capabilities that people are able to use in the models in 6 to 12 months.
Can I ask which segment we’re underserved in?
In terms of model capabilities?
In terms of, we don’t have enough medical data, we don’t have enough financial modeling data. Is there a segment where you think, “You know what? If we were to acquire a company in this space to plug a hole in our data supply”?
Yeah. I would say maybe I’ll give it from Mercor’s perspective, and then I’ll give it from the labs’ perspective.
We tend to be now so good at mobilizing experts that we’re able to access pretty much any domain. There are always going to be some degree of niche pockets of oncologists or whatever it is that have a particular background, but generally we can fill those fairly quickly. It’s more about people who are very acclimated to the frontier of AI, because it’s the people who both have the expertise in oncology and are power users of ChatGPT or Claude who are able to find where the model makes mistakes and help the model learn from those mistakes.
That’s from the Mercor perspective. From the perspective of the labs, it seems like it’s all-encompassing. The barrier to automating everything that you can do in, say, Google Workspace is how we cover the full distribution of all of the context—messages, Slacks, slides, Excel sheets—and all of the tasks, prompts, and outputs that correspond to everything that you do in your job.
That applies to every individual and every domain throughout the economy. There’s this enormous mobilization of hundreds of thousands, and soon millions, of people to build out the full distribution of everything that you could pass into Google Workspace and everything that you could want out on the other side in every job category throughout the economy.
Can I ask you, before we dive into a tweet that you did which slightly terrified me, to be quite honest? You said 30 to 40% is how we think about our revenues from that?
Generally, yeah.
Okay. So if we take the rounds that we’ve raised, which round felt most uncomfortably high?
Good question. I’ll talk through the valuation of each and the revenue of each.
5. The Fundraising Story: Helicopters, Ferraris & $10B Valuation
Did Founders Fund not fly you in the chopper?
That was Series A. Our seed round was in September of 2023. We were at, call it, $1 million in revenue run rate, or just shy of that. I initially didn’t want to raise because I wanted to bootstrap the company, but Adarsh’s condition on dropping out was that we needed to raise money.
We met General Catalyst at 8:00 a.m. on a Sunday morning. They gave us a term sheet within 36 hours for $2.3 million at a $23 million post-money valuation.
Pretty good?
That was pretty reasonable in terms of price at the time.
Max and Hemant?
This was Max and Nico. At our Series A, the business hadn’t grown that much from the seed to the Series A, but we found that we had a key differentiation in the market. We met Victor when we were at, call it, $1.5 million in revenue run rate in May of 2024, and Victor got super excited.
Initially, I refused to take a second meeting, but then he said, “Oh, have you ever been in a helicopter?” Peter took us on the helicopter flight, and then Benchmark really wanted to work with us. By the time they gave us a term sheet, we were at, call it, $2.5 million in revenue, and they gave us a $250 million post-money valuation.
Uncomfortable, because that’s a big jump—from $23 million post to $250 million?
Keep in mind, at the time this sounds crazy because we were at $2.5 million in revenue, but I was projecting $50 million in revenue run rate by the end of the year and $500 million by the end of the next year. It felt like a bargain.
Do you know, do all founders project that way?
But we beat the projections.
It just doesn’t happen often.
Yeah, yeah, yeah. Then, 4 months later, we met Felicis. We never made a slide deck or took investor meetings, and so Felicis sent us an email saying, “Hey, we know your co-founder Surya really likes Ferraris, so do you want to go racing Ferraris?”
I replied and said, “You caught my eye. Tell me more.” They said, “We’ll meet at the airport in Hayward and go on Aydin’s private jet to Las Vegas to race Ferraris around the F1 track.”
[laughter] I was like, “We’re available in 3 weeks on a Sunday.” So we do this: we race Ferraris. We’re at $20 million in revenue, and they ask us what valuation we think makes the most sense. I say $1–2 billion, so they give us a term sheet at a $2 billion valuation. At the time, that’s 100 times revenue, and everyone thinks that’s a high valuation. Meanwhile, it was an incredible investment.
So, I’m going to be honest: this is when I interviewed Adarsh at that time. At the end, I was like, “Dude, I would love to invest. Please let me invest.” You very kindly let me put a small check in, and I then spoke to several of the biggest investors in the world. No offense, but they chuckled at me: “Dude, that was such a high price. You paid such a high price.”
Well, here’s the thing. We’d been growing 50% month over month for the prior 6 months, and I think what none of them really realized was that it would continue for the subsequent 12-plus months. So that compounded more and more. By September 2025—or, say, October—we were at, call it, $400 million in revenue run rate.
Then Felicis was like, “We want to invest more.” So they gave us a term sheet at a $10 billion valuation. We didn’t really want to spend much time on a financing because the business was growing 50% month over month, and so we were very preoccupied. That was about 25 times, and the business has almost 4x’d since then.
So, really, which one felt most uncomfortable? If you were to choose any.
If I had to choose any, I would say the Series B priced in the most—the furthest ahead of our growth. Or the Series A. I think it was probably the Series B.
The $2 billion.
Because both were 100 times the revenue, but it’s very different to be 100 times the revenue when you’re at $2.5 million in revenue versus $20 million in revenue. That was probably the largest one, but obviously both were great investments in hindsight.
What’s the next round done at?
We’ll see. Probably a much higher valuation. We’re getting a lot of offers at meaningfully higher valuations, but the company is fairly profitable, and so we’re taking our time to see who the right partner is.
We’re also just going through modes of transport, aren’t we? We had the chopper, we had the Ferraris.
That’s a good observation. Exactly. We need a warship now to get the Series D. [laughter] I totally agree. I’ve never been on a warship before, but that’s a lot of fun.
6. Infrastructure Will Win Over Application Layer
There you go. So we’re lining up the warship. [laughter] The next 12 months will be dramatically better for infrastructure companies upstream of Anthropic and OpenAI than for application-layer companies downstream of them. This was your tweet. Why do you believe that?
Mm-hmm.
The reason I believe that is that the application-layer companies’ businesses are not far removed from the foundation model companies’ businesses. It’s not a far leap for Claude CoWork to add capabilities across medical and legal. Obviously, they did it with software engineering, and Claude can do that across finance. So I feel like building defensibility in the software layer on top of the models is going to be incredibly difficult.
Whereas on the infrastructure side of things, it feels like there are meaningful moats getting built. We are compounding enormous network effects in the business and a pretty significant data moat as we build out the inventory for our customers. Compute companies, obviously, are able to build moats through these very long R&D cycles. So I think there are going to be high margins achieved at the infrastructure layer and sustainable, profitable businesses in a way that’s less immediately clear at the application layer.
I mean, you saw Nebius. I don’t know if you saw this, but they increased their pricing by 30%.
I didn’t know.
Across the board.
Wow.
It will have absolutely no impact on demand. Isn’t that absolutely nuts? You increase the price by 30%, with zero impact on demand.
That’s insane.
Do you do pricing elasticity tests? Because if you can double the price and double the business, you maybe can’t double prices. You could double capacity. You could probably increase prices by 30% without much of an impact.
But the other thing you need to consider is that pricing is not merely a question of optimizing for the next 6 months. It’s optimizing for a structure that wins the market over the next decade, right? For that reason, we’re very focused on how we do what’s best for customers, how we do what’s best for experts, and how we build a sustainable business while we’re doing it—but make sure that we’re not leaving oxygen in the market, because high margins invite competition.
Okay. I am an investor in several application-layer companies downstream, like Legalese, which you mentioned there. We see the Legalese versus Harvey battle. I think everyone is actually coming around to the fact that they shouldn’t be fighting each other. They should be wary of Anthropic, to your point.
Totally, but then I look at it and go, there is incredible defensibility. It’s a very deep product specifically suited to the workflows of lawyers. Anthropic would have to build out whole separate product teams and divisions to come after them. They’d have to build out go-to-market teams, customer success teams, and adoption teams. It’s a different freaking company. The defensibility is there. Argue back.
7. Is SaaS Dead? When Network Effects Are the Only True Moat
Maybe I would say 2 things. First is that I think over the last 2 years, everyone has increasingly realized that the model is the product. We can build so many of these different abstractions of trying to stitch together API calls and having all this patchwork logic, where people used to have all these drag-and-drop agent builders. Then they just realized that if we give the model the end goal and train it to accomplish that end goal, it has outperformed every other solution in almost every case that we go after. That bodes incredibly well for those that are training models end to end.
The second thing to consider is that software layers are able to get recreated very quickly now. We’re building out an eval set that measures how effectively agents can build end-to-end SaaS applications. 2025 was the year of, how do you get a model to make a PR in a codebase? 2026 is the year of, how do you get the model to clone Slack end to end? Those capabilities are going to exist in the models in the next 12 months. That means very significant things for companies that are betting on software moat sustaining their businesses.
If we take that extrapolation further, how effectively can we build Slack internally, agent-led entirely? That would very much concur with the idea that SaaS is dead, because if you’re a large company needing maybe small customizations and integrations—say you’re a real estate company and you need very specific integrations to pricing providers—you’d build your own.
I generally agree. I think the caveat is that when those companies have network effects, there’s probably a significant moat that isn’t being priced in fully. For example, Salesforce has tons of companies that are building integrations on top of their platform. That creates this almost marketplace and network effect around it. Slack has Slack Connect, right?
I think Carta is another great example of this whole network effect of the people that use it and want to use the same platform across all of their companies. The companies that have network effects will be able to, in some ways, generate more value because they can iterate 10 times faster while leveraging those network effects to create more value for their customers and therefore build more valuable products, charge more money, and increase revenue.
The companies that don’t have network effects are going to struggle very significantly, because there’s not really a defensible moat in the pure software associated with the products that they build. To me, that is the litmus test that determines whether this company is going to become worthless or whether this company is going to gain dramatic value from its ability to 10x product velocity.
You said we’re learning more and more that the models are the product. What if I push back and say the go-to-market is the product? When you’re selling to law firms, it’s about being in the room with your biggest law firms—your Cooleys, your Goodwins, your Wachtells, your Clifford Chances—building the relationship with the buyer, and then the CS and the adoption. It’s actually in the go-to-market, not in the product.
I agree with this in part, but the caveat I would give is that it’s arguably more the forward-deployed motion rather than the go-to-market. The forward-deployed motion is the post-sales; go-to-market is the pre-sales.
Ultimately, say you’re just really good at sales, and then you provide a SaaS product, and you have a savvy customer who’s spending $1 million a year on the SaaS product. They realize they could just tell Claude to copy it, and they’ll get the same exact thing. It feels very difficult to maintain your pricing power even if you’re the best in the world at sales.
Whereas, on the other hand, if you have a great forward-deployed motion where you’re going deep with a customer, you’re training the agents based on all of this tacit knowledge within the company so that they understand how to perform effectively, that feels incredibly differentiated and hard to recreate.
That's also the reason that we see the labs, OpenAI and Anthropic, investing so much in this forward-deployed motion. I think that the Sequoia article “Services Are the New Software” resonated a lot in that the software moats are whittling away, and it's the ability to layer services on top of software to meet the customer where they're at and go the last mile that is creating stronger defensibility.
Do you buy this new sexy category? The venture investors are wonderful people, but this new sexy category of AI-enabled services—is that the future gold mine?
I think in a large way I do. I think the key thing is that you need to make sure that they're actually going to leverage AI. There are a lot of companies that are just building services and not getting a significant competitive advantage from AI and using that. That's the thing you've got to be careful about, but I think it's very rational.
I'll give an example in the context of Mercor. Within this process of turning human time from the talent network into building these super-rich environments that mirror everything that people could do in their jobs, there is a lot of human coordination: How do we answer people's questions? How do we track the KPIs of the project and manage it effectively? How do we build the bespoke tooling for that project?
We have about 100 people, or call it 150 people, in our delivery organization who do that for deployed work, helping to go the last mile for the customer. But now we have an AI project manager that just completed its first project managing that entire thing end-to-end. It's able to hire the experts, answer their questions, build the annotation tool using its coding tools within our platform, and produce the end data type.
The experts all had a really good experience on the project, reporting to the AI project manager that was running it. I think we're seeing in real time that services are getting automated and that this is going to be an extraordinary transformation in the economy.
8. Token Spend on Agents Now Exceeds Employee Headcount
One thing that powers the agents that we use is the tokens that power them. I thought the whole point was that we have increased token efficiency and token costs come down. Token costs are rising for everyone.
Mhm.
Help me understand how you see token costs changing in the next 6 to 18 months, and why.
Well, it's a fascinating case study in Jevons paradox, similar to what we were talking about in the context of making humans more efficient leading to more jobs. When we make models improve by 10x year over year, that has just been causing the total consumption of the models to go up and up and up as the cost per performance goes down.
Insofar as how it's going to develop, this trend is going to continue very, very significantly before we start seeing any leveling off of token consumption within the enterprise. Right now, we're spending more on tokens for our internal agents than we are on employee headcount. I think most businesses are going to look like that in—
You're spending more on tokens for agents than you are on headcount?
Exactly.
Your token spend on agents is more than salaries?
That's correct. It's pretty incredible. The way we manage it is that we have a variety of these key workflows throughout the company where we have an AI project manager, as I was describing, that manages operations.
We have our interview question agent. We've done over 5 million interviews, and it asks all the questions in those interviews. We have our interview ranking, or broader candidate ranking, where it helps to assess all of the candidates and figure out who we should be hiring. We have agents for accounting automation, fraud detection, and so on.
Corresponding to each of these agents, we have an eval that tells us which model is best to use for this given use case and what the Pareto frontier of price-performance is for that specific use case. That eval allows us to make decisions around where we should be allocating our inference spend, what provider we should be using, and so on.
I believe that over time, this is going to develop to look very similar across every Fortune 500 company, where they'll need to have this system of record for evaluating and specifying agent behavior across every workflow in their business. They're going to use that to commoditize the model layer because they want to enable perfect competition for the models, with zero switching costs.
We've been growing extremely quickly with the enterprise, helping them to populate the system of record and building out those evals for each of the use cases that they have throughout their business.
Do you think you will see that commoditization at the model layer, whereby enterprise clients are able to efficiently package the workflows that they do, so it does commoditize the model layer? Because right now, it's not quite commoditized.
Yeah. I think the key distinction is that I think the API layer will get commoditized. You can definitely build stickiness in the workflows that people have on top of those APIs.
For example, I have all of these routines running in Claude Code, and I feel like it would probably be difficult—or at least I wouldn't put in the time—to move those routines over. I have a bunch of similar things running in ChatGPT.
I think there are going to be various ways that people can build stickiness, but for pure API-based products, if we're just spending $10 million a year on a specific workflow, obviously we're going to have an eval for that. Every time a new model comes out, we're going to benchmark it and understand exactly how we should be hot-swapping between models and distilling models.
Why does the API layer get commoditized?
Because the switching costs are zero. When the switching costs are zero and there's a new frontier model every 2 months, that means that we're very quickly going to swap them out.
Ultimately, the decisions that we make boil down to the score on the eval corresponding to that workflow. It's very easy to compare model to model one-for-one in a perfectly hot-swappable way, which is almost the definition of a commodity.
I'm still reeling from your token spend with agents being more than headcount. Marc Benioff said the other day that they spent $300 million on Anthropic, which seemed like a lot of money, but when you break it down, it worked out to be about 3.8% of developer salaries being spent on Anthropic, which is much less than one would think.
Yeah.
What do you think that is in 24 months' time?
For a Salesforce?
Yeah.
I don't know about 24 months' time, but I would bet that in 5 years, the average enterprise spends more on compute than headcount. The reason for that is that the models are just becoming so capable that it seems like there is enormous ROI to being able to have models do something for $100K a year that is going to continue compounding at an exponential rate in a way that human intelligence is not going to.
Humans will still play an important role in the things models can't do, but I expect that the cost of inference and the cost of compute will exceed that. The reason that's so interesting to me is that having an eval for your specific workflow—say we take the case of Salesforce, having an eval for how good a specific model is at code generation in their use case—is often a 10x lever on the price-performance of that model.
They can distill the model, and they can have an open-source model that is performing as well as, if not better than, a frontier model for a dramatically lower cost. As we see this enormous shift towards compute and significant inference spend across every workflow in the enterprise, they're going to need to have evals that act as a source of truth for whether those workflows are being done correctly and whether they're using the right models to accomplish that.
With the greatest of respect, evals today are relatively unhelpful. It's like, how good are you at driving around the corner for the driving test in a very specific way, but actually that's not how it works in the real world. It's not very practical.
That's exactly the problem. We used to have this paradigm of all the academic benchmarks that were totally disconnected from the outcomes that enterprises actually care about. People were building everything ranging from GPQA for PhD-level reasoning to IMO for Olympiad math to humanities last exam for this long tail of academic problems no one really cares about.
Now they're focused on how we get the model to do this end-to-end workflow, coordinating with multiple colleagues for a financial model or slide deck like we're discussing. How do we get the model to build an entire SaaS application end-to-end?
That's why there's this enormous buildout and pushing the frontier of evaluation as a critical research problem for the next frontier of model development.
Okay, the next frontier of model development. If I listen to everything that you just said, I would draw 2 conclusions. One, we should just invest all of our money into OpenAI and Anthropic.
Then the realization dawned on me that the majority of startups, especially on the West Coast, use frontier models to see where they can go and how far they can push them. Then they use open-source, often Chinese models, to get as close to that as possible at a much better cost basis.
In which case, OpenAI and Anthropic are inherently challenged by that much more cost-efficient open-source model. Right or wrong?
I think both are true. There’s going to be many, many orders of magnitude more demand in 5 years than there is today—maybe 4 or 5 orders of magnitude more demand—but there’s also going to be increased competition, with people distilling and having fine-tuned open-source models that accomplish their work closely.
Ultimately, I think OpenAI and Anthropic are incredible investments, and it seems like there’s starting to be consensus around that in a way that there wasn’t just a couple of years ago. At the same time, I think the majority of inference in 5 years is going to use an open-source, custom fine-tuned, or distilled model, rather than a frontier model.
Okay, interesting that you said that. Obviously, incredible investments. Where will they be in 5 years’ time?
Valuation-wise? Revenue-wise?
That’s one—valuation-wise.
Valuation-wise, wow.
If we put them both at $1 trillion today, give or take.
Yeah, this is hard to imagine.
This is one I’ll play back to you in 5 years’ time, and we’ll both look back and go, “Ah, either we were very prescient or just completely wrong.”
I could definitely see 1 of them being a $10 trillion company, maybe even significantly higher. It feels like the opportunity associated with being the frontier model is so large that it will eat up so much of the other demand within the economy, because that also means that when you have the frontier model, you can use that as a teacher model to distill your own models and have the best small models, et cetera.
So I would guess that at least 1 of them is worth more than $10 trillion.
My naive assumption was, when you talk about orders of magnitude more, when you talk about spending more on compute than you will on salaries, why don’t we just put all of our money in NVIDIA? I know it sounds supercilious and glib.
I think it’s not a crazy idea. NVIDIA’s obviously a phenomenal business that will continue to execute super well. The only caveat is that it feels like we’re starting to move toward a multichip future, where obviously Cerebras is executing well. I’m good friends with the Etched guys. Most labs are building in-house chips.
So I would guess that in 5 years, it doesn’t feel like NVIDIA has quite the same monopoly. But that’s okay, because even if they only have 30% or 40% market share in the largest market in the world by far, that is the world’s most valuable company.
Speaking of the world’s most valuable company, you’re seeing this concentration of value toward the top 8 names more than ever before. 84% of the year-to-date rally was driven by the top 10 names. Do you worry about the concentration of value in such a small number of players?
Maybe to some extent. I definitely worry about how we smooth out the benefits to society. How do we ensure that every enterprise and every individual is able to reap the full benefits of AI, rather than just a handful of people in San Francisco?
Ultimately, I also think that there’s some natural dynamic associated with capital allocation, where it’s going to be more valuable to give the compute to Anthropic, where they have the marginal demand and can use that right away, versus a less successful company that might not be able to create the most value with it.
So I think that it’s probably good from a capital-allocation and efficiency standpoint, so long as we’re able to manage the societal implications of increasing inequality.
Speaking of increasing inequality, you wrote an essay—and this is taken from your Twitter—about how we should eliminate income tax for the bottom half of Americans.
Mm-hmm. Well, I believe this very strongly. I actually wrote this essay as a research paper when I was a freshman in college. It was one of the few productive things I did in college.
[laughter]
Essentially, the thesis of this was that the largest positive externality in the economy is jobs. People talk about all of these economic theories of how we have negative externalities, like carbon or smoking or whatever it is. We should tax those.
But on the one hand, the largest positive externality is jobs. Yet on the other hand, the way that most economies structurally collect income is by disincentivizing jobs, both on the income-tax side by taxing individuals, as well as on the payroll-tax side by taxing companies.
9. Competing for Talent When Meta Offers $20M Per Year
As we move toward a world where there’s increased job displacement and increased uncertainty around how many jobs there are going to be, especially for the bottom half of Americans, I think that this is going to become extremely problematic.
And so I would suggest that we move toward a paradigm where we instead focus on taxes on things that aren’t necessarily going to have a negative impact on incentives in the economy. One great example is capital gains, where I’m going to invest money in assets regardless. If there’s a higher capital-gains tax, it’s not like I’m just going to not invest, right?
I think that taxing capital gains, especially short-term capital gains, which I think are probably not as beneficial for the economy as long-term capital gains, would probably be structurally much better than taxing income.
With the greatest of respect, if you increase the tax on capital gains, you will disincentivize those investors from taking risk. Why the fuck should I pay more? I’m already taking a risk. I’m already investing in innovation when other people won’t, when banks won’t, when all the data tells me not to.
Now you want to tax me more for doing that, for taking the risk? Of course you will disincentivize investment.
The thing is, when investors are taking very high risks, it’s generally in an aggregated way, in a portfolio. And so you would tax the gains on the portfolio overall. I know that you don’t like to hear the capital-gains-tax theory, but—
No, no, no, no. I think I say this with the nicest respect: it’s just wrong, because you just move.
But I agree that you need to be careful. The main thing you need to be careful about is whether people would move to other geographies, because obviously that creates problems.
But I—I'm so sorry to be a dick, dude, and you can say I withdraw. That creates problems. Yeah, that’s kind of the whole point. You fuck off to somewhere that doesn’t have capital gains, and then you lose all the tax revenue completely.
I’m sorry, forgive me. We live in the UK, where there’s the Green Party, which is this idea-less movement that’s like, “Oh, increase tax.” Well, yeah, then we leave. And then you have nothing.
I agree. I think that there needs to be sensitivity analysis associated with how the increased amount of taxation causes people to just leave and reduce overall government revenue.
But I think that another way of going about it is also taxing consumption of items that probably aren’t the best. It’s crazy to me that instead of taxing carbon, we tax the bottom half of Americans. Why don’t we tax carbon, right? That’s a very clear negative externality in the economy, at least in the US, that’s not taxed.
I feel like there’s a lot of low-hanging fruit with respect to things that we could tax without damaging incentives in a perverse way or causing people to flee the country, which would be far better than taxing the bottom half of Americans.
The other thing is that it’s only 3% of government revenue. The fact that it’s only 3% feels like a very easy decision for policymakers to make in the grand scheme of the impact that it would have on people.
Would you tax prediction marketplaces? It’s gambling.
I probably would. There’s probably some value in having good prediction marketplaces, allowing people to make effective predictions of the future and hedge things within their lives and investment portfolios, but it’s likely okay to tax them.
The thing on that point around taxing the bottom 50% is that Jeff Bezos retweeted me, which I was ecstatic about.
That’s pretty cool. [laughter]
It was pretty great, yeah.
Who’s the coolest person you’ve met?
I really like Jensen, and I really like Satya. So many incredible people. Obviously, Dario and Sam are incredible.
But if I had to choose 1 person, Jensen’s so cool, right? The jacket, his style—he’s always on point. So I would say Jensen is probably one of the coolest.
The fascinating one I would love to ask you—and you shouldn’t give the answer to this, but I think this is what people have asked of me, which is very hard to answer—is who did you think would be amazing who was surprisingly underwhelming?
Okay.
No, but it’s a really good one.
It is an interesting question, yeah.
And I have met a couple where you’re like, “Wow, that gives me confidence that I can do that, too.”
You know, actually, I will say this one thing, which is that I remember when I went to Georgetown, I didn’t get into Harvard, and I was like, “Wow, the people at Harvard are probably dramatically smarter than me.”
And I went to this nonprofit called Prod, where there was a bunch of kids from Harvard and MIT who were all building startups. They’re very smart, don’t get me wrong, but I do think that most of us have this very equalizing feeling. When you spend more time with them, you realize that they’re just normal people to a significant extent.
Not all of them, but most of them, to a significant extent.
And I think that makes you feel like, when I saw Ethan Thornton from Mach raising $70 million as a 19-year-old, I thought, “Wait, Ethan is a chill guy and a good friend, and maybe I could do something like that one day.” It just gives you this sense of being able to accomplish so much more.
It’s so interesting you said that—that kind of dispersion effect from seeing your friends achieve. I think it’s one thing that’s held Europe back in many ways. You were with some of the largest model providers in the world. How do you feel about Europe’s inability to compete and provide leading models to the world?
When you look at the benchmarks, Mistral might make an entry at, you know, 72. It’s like the Eurovision Song Contest, kind of at the bottom. I love art, I love visuals, and I’m very proud of it as a European, but we haven’t delivered on the model side.
I think that it’s going to be difficult to change because there are just so many strong network effects around talent. I know so many brilliant French researchers who go to work at OpenAI, Anthropic, and DeepMind, because when those labs have the best talent, that’s where they all aggregate. Then that compounds to them having more capital, more compute, more impact, et cetera.
I expect that trend to continue and to be one of the largest economic and geopolitical advantages that the US has.
So if you were Europe today, do you go, “You know what? We’ve lost that model race, but we can still be a dominant energy provider”? If you’re Norway, where I’m from, actually, we do pretty well providing energy. Is that what we just accept?
10. Do Sovereign AI Models Actually Matter?
I would accept that. I think that maybe it’s worth having some post-training capabilities, because there is going to be value to distillation and some of the work that happens after foundation models are built. There’s definitely going to be some value in applications, but I don’t know if I would lean aggressively into, “How do we compete head-to-head with Anthropic?”
Do you buy the sovereignty argument that we need sovereign models because we don’t want our data going to the US or China, or wherever that is?
Maybe in some cases. There is value in localization, and I’ll give an example: Oftentimes, labs will come to us and say they need their models not just to be good at American law, but also to be good at British law, French law, or whatever the jurisdiction is in the world. I think that’s going to be an important last mile in making the models useful in whatever jurisdiction they’re operating in.
That said, the labs are just going to hire 10,000 people in France to teach the models how to be better at French law. I don’t think there’s so much that others are going to be able to do to stop that, because the transfer-learning capabilities from all of the other domains that they’re focusing on are just so powerful.
And when you say hire 10,000 people, the thing that’s just astonishing to me is the wave of cash. I’m sure OpenAI is the same, but I’ve seen it specifically with Anthropic. I mean, insane levels of compensation.
Totally.
How do you compete against that?
It’s definitely one of the things that’s most top of mind, in particular because the market for people founding companies is so hot. We’ve had 3 employees who have founded companies worth in excess of $100 million.
I saw your tweets where you do the Mercor Mafia tweets.
We’re a very young company, and I think that it’s difficult for a variety of reasons. A lot of people probably don’t have a full understanding of just how hard it is to build a company, as you know well, Harry, and how low the probability of success is, and how fortunate we were and how lucky we got along the way.
I think that’s definitely one of the large challenges. There was someone I was hiring the other day who had an offer for $20 million in cash per year from Meta’s superintelligence group.
Meta’s superintelligence group?
Meta’s superintelligence group.
$20 million in cash?
Per year. Or it’s in stock, but liquid.
That’s hard to compete against.
It’s hard to compete against, yeah.
Does that change, or does that just continue to escalate?
I think that it’ll probably continue to escalate for a smaller group of people. But I also suspect that as more people gain knowledge of how these labs operate and what the capabilities are for training a frontier model, there’s going to be more supply in the market for people who have that skill set, and more reasonable pricing.
I expect there to be some craziness that continues, but hopefully the 99th percentile, at least within the market, will balance itself out.
What’s the hardest role to hire for today?
Researchers.
Just because of supply?
Because of supply and demand. It’s just this market where there’s 10 times more demand than there is supply, and that makes it very difficult.
How much does it cost to hire a high-quality AI researcher?
Oftentimes, it would be in the tens of millions in stock per year, for the really good people.
When researchers weren’t paid very much—just 10 or 15 years ago, they were the underpaid but brilliant people in society.
Yeah.
Now I feel like that’s relatively changed.
Yeah.
Is it harder than ever to run the company?
I don’t think so. To give a frame of reference, we were 40 people and had a $50 million revenue run rate at the start of last year. Since then, we’ve grown headcount 7 or 8×, and we’ve increased the broader scale of the business by 25 or 30×.
It’s definitely been very stressful to keep up with the growth along the way, but I think that now we have the supporting functions. We have finance and legal, and we’re building out HR. That brings some sense of stability, where I don’t have to deal with all of these little escalations, and I’m able to spend my time focusing on building great products, research, and time with customers.
That, I think, has made it significantly easier to run the business.
I get a lot of shit for everything I say these days, which is wonderful. The trouble is, I don’t deliberately rage-bait, but people just hate me, which is the worst thing.
11. Does HR Slow Companies Down? Brandon Pushes Back
I tweeted after a show with Adam Foroughi at AppLovin: “No great CEO that I’ve met—and it’s true—loves HR. They slow you down, they implement policy and procedure, and it’s just a pain.” Do you agree with me?
The caveat I’ll give is that I think it’s really important. We definitely had challenges in scaling culture when we went from 40 people to 400 people.
How does that show up?
Quickly. It’s so many things, ranging from making sure that we keep a really high talent bar, to making sure that people are bought into the mission of the company, to even the tactical things of making sure that managers are communicating to their team about their performance review and how they’re doing, so that they’re never surprised by a performance review.
When we have a young team with a lot of first-time managers, that creates culture challenges for people who aren’t used to giving feedback and maintaining all of the values and commitment to the mission of the team.
To some extent, I agree, and I think that some of the big tech companies probably go too far in empowering HR. But I also think that it’s important, and one of the large lessons we’ve had over the last 18 months or so is that it’s critical to get these foundations in place as you scale headcount. Otherwise, it creates problems.
Before the show, we said that after the show with Adarsh, a couple of people thought that 996 was the way Mercor is run. It’s like clock in, clock out. Why is that not true, and how do you think about that?
The reason it’s not true is that we’ve never mandated hours at the company. Obviously, I work extremely hard. Adarsh works extremely hard. We work from when we wake up until we sleep pretty much all the time, aside from maybe working out, but I’m still working and thinking about work during that time.
Most of our leadership team does as well. But at the same time, the majority of my leadership team has kids, and we want them to be able to go home and see their families and all of that.
12. Quick-Fire Round
I think it’s some combination of knowing that building a legendary company requires immense dedication to the mission of the business, while also recognizing that we need to ensure that it’s a sustainable environment for the best people in the world to do their life’s work.
Are you ready for a quick-fire round?
Of course.
Would you like to go public?
Definitely.
When?
In the next few years. I think that all legendary companies eventually go public, and so it’s an important part of the journey and of maturing and having a much larger company than we have today.
But it’s not something we’re rushing to do this year or next year, in part because we dropped out of college less than 3 years ago at this point. It’s still a very young business, and we want to make sure that we properly actualize everything that we’re working on, especially on the enterprise side, before going public.
Do you ever lie in bed at night and just go, “Wow, it’s pretty wild”?
I’m always pinching myself, and I feel extremely grateful for the team and Adarsh and Surya and how all of them made it possible, because I could have imagined a hundred things that would have gone differently, and we’d be in a totally different circumstance.
What have you changed your mind on in the last 12 months?
I used to have some questions around whether the foundation model labs would be the largest businesses in the world because of the exact things you asked about in the context of how much those models are going to be able to maintain pricing power amidst a competitive environment. But I think that as we’ve seen the sheer revenue ramp of these businesses, I’ve gained immense conviction that they will be the most valuable companies in the world.
You can invest in OpenAI or Anthropic. Which one?
Oh, I can’t respond to that. I would choose that. [Laughter]
Who do you not have as an investor in the company yet that you would most like to have?
I really admire Jeff Bezos. I think he’s so disciplined about the culture of Amazon. That’s one of the things that’s always stuck with me. Everyone there just understands the values and is steering in the same direction. He’s such a strategic business leader. I’ve never met him, but I’ve always wanted to.
Which competitor do you most respect and why?
I admire that Edwin from Surge has done a really good job in staying super close to research, and it’s something that we’ve obviously been doing a lot of as well. But I think that’s probably one of the largest things that differentiates both us and Surge: our ability to train models, to hire some of the best researchers in the world. I admire them for their execution on that front.
What percent of data providers are just respectfully transactional talent marketplaces?
In terms of volume or number of competitors?
Number of competitors.
About half.
Half?
Yeah.
What would you most like to change about your role today?
I would say that there’s a decent amount of HR things that get escalated to me, and so we’re looking for a really strong head of people who is able to handle a lot of this.
Final one for you, dude. What’s the kindest thing that anyone’s ever done for you?
One that really stuck with me is—I’ll probably attribute this to the entire Prod community, namely especially a couple of people like Rob Wachen, Ben Spachter, and Richard Dahan. Prod was this nonprofit that got started at MIT and Harvard, and I was sort of a blow-in because I didn’t get into those schools, but I went to Georgetown.
For the first year of the business, they would meet with us every week. Ben became a big customer. Richard would give us tons of money just to float working capital, and Rob gave incredibly valuable advice.
They had nothing in it for them. They took no equity. I tried to give them equity, and they wouldn’t accept it. Mercor wouldn’t exist if it weren’t for any of those individuals, I would say. I think that that is something that I’ll always be grateful for for the rest of my life.
Dude, I have to say I loved having you on the show last time. It was incredible to do this in person. I’m so thrilled with how this conversation went, and you’ve been amazing.
Thanks so much for having me, Harry. Always great to come back.