从岗位替代到 AI 训练师:Brendan Foody 谈 AI 时代的工作
Mercor 已融资1亿美元,营收年化运行率突破1亿美元,同时为顶级 AI 实验室招募了数千人。 公司成立于2023年,最初是综合型人才匹配平台,但人类数据正从众包低、中技能劳动力,转向筛选能够直接与能力前沿研究人员合作的高能力人才。
智能体评测是从亮眼基准测试到具备经济价值的自动化之间的瓶颈。 SWE-bench 拿高分,距离替代一名需要与产品团队协作、选择工具并运用品位的软件工程师还很远。Foody 预计,这将是一个持续数年的、按行业推进的建设过程,因为“评测位于上游”,决定了实验室和应用公司希望获得的能力。
Mercor 有两条会复利增长的循环:人才供给和客户绩效反馈。 其免费的职业工具解决了劳动力市场中典型的50:1供需失衡;结果数据则持续改善平台对谁会表现出色的预测。Foody 认为,不那么显眼的数据飞轮最终可能比市场网络效应更重要。
Foody 预计,知识工作岗位的替代将来得很快、过程很痛苦,也会成为政治问题。 客服和招聘已经是出现岗位替代报告的领域,尽管大部分替代尚未发生;体力劳动以及重视人际互动的岗位,自动化速度应会更慢。他给出的生存法则是保持多面性,因为数学以及很快就会到来的代码能力“会被非常迅速地解决”,而品位和创始人判断所能获得的反馈很稀疏。
强化微调可能只用数百到数千个样本,就能解锁企业级智能体。 与监督微调不同,强化微调只需规定期望结果,并奖励模型自行找到实现结果的方法。Foody 认为,基础模型已经具备所需的推理能力,缺的是企业特定的工具使用知识,以及“这个岗位上的优秀表现是什么”。
创建评测可能成为全球最常见的知识工作岗位,尽管劳动者正在帮助自动化自己。 经济逻辑将劳动力从反复执行任务这一项变动成本,转向一次性定义评测这一项固定成本。只要人类评测仍存在前沿,这项固定成本机会就会持续;如果模型变得超越人类,那么经济中的许多其他环节也可能不再需要人类。
1. 专家评测取代商品化标注,成为稀缺投入
Mercor 的运营快照异常紧凑:公司由3名大学辍学生和 Thiel Fellows 于2023年创立,已融资1亿美元,营收年化运行率突破1亿美元。Foody 表示,公司的 LLM 会筛选简历、进行面试,并且“比人类更准确地预测工作表现”。
Mercor 最初是一家综合型人才匹配公司,服务那些无法通过传统人工招聘获得机会的优秀人才。随着人类数据开始从众包低、中技能劳动者生产“语法勉强正确的句子”,转向筛选最有能力、能够直接与前沿研究人员合作的人才,公司的切入点也随之改变。
需求覆盖咨询、软件工程,也包括业余爱好者和电子游戏。只要有相应评测,强化学习就能提升模型能力,因此 Foody 对市场的概括很简单:“评测位于所有这些能力的上游。”
2. 当绩效呈幂律分布,人才预测才真正值钱
Mercor 已经看到,模型在大多数评测中的表现超过人类招聘经理。Foody 预计,即使出于法律原因仍需由人类完成最终审批,不听取模型建议也可能变得“几乎不理性”。
幂律职业中的经济价值最强:识别出90分位的工程师,或者找到成本只有一半、但表现位于前25%的人,都会创造不成比例的价值。投资比软件工程更符合幂律分布;工厂工作则更加商品化。
文本密集型评估应最先自动化,因为模型可以以超越人类的规模向候选人提问、分析面试记录。激情、说服力和销售能力更多依赖多模态信号;而反复招聘同一岗位,比比较20个人分别从事20种不同工作,更容易清晰归因绩效。
招聘还留下大量未被充分利用的线上证据:GitHub 仓库、大学博客、个人项目和设计作品集都包含大量信号,但招聘经理没有时间逐一审阅。Foody 还提到一些更隐性的信号,例如在西方国家留学的国际候选人往往更擅长协作或沟通;论文或副业项目则能显露内在动机。理想情况下,Mercor 会直接测试工作本身,比如让候选人完成一个缩小版 MVP;如果真实结果需要很长时间才能观察到,就使用代理指标。
3. 可验证性决定哪些人类技能最先贬值
Foody 预计,许多岗位的替代会“非常迅速”地发生,并同时变成痛苦的大众政治问题。客服和招聘已经出现岗位替代报告,但大部分替代尚未发生;他预计更多替代即将到来,尤其是在经济收缩期间企业更加重视效率的背景下。
他预计,更多工作会留在实体世界:机器人数据采集、餐饮服务、治疗,以及其他重视人际互动的岗位。实体自动化的推进应会更慢,因为它缺少虚拟世界中快速且自我强化的改进循环。
他对持久技能的建议是保持多面性:模型一次次掌握人们原以为会长期困难的活动。数学以及很快到来的代码能力进展最快,因为答案可以验证;创始人的品位则更难复制,因为反馈信号稀疏。不过,一旦新领域积累了足够数据、开始进入学习阶段,跨领域推理能力仍然可以迁移。
Foody 可能不会鼓励年幼的孩子走计算机科学路线,但也并非完全反对。他更看重能带来智力刺激的兴趣、通用推理能力、创业式试错、对市场缺口的逆向判断,以及产品品位,而不是押注“会写代码的人在5年后仍然独具价值”。
4. 经济活动需要智能体评测,而不是又一个学术基准
历史上的评测类似零样本学术题目;真实工作则是端到端系统。软件工程师必须理解产品需求、跨团队协作、使用工具,并将相互竞争的优先级转化为产出——从 SWE-bench 表现出色到替代软件工程,中间“还差得很远”。
因此,评测创建需要按行业推进。客服是一个相对容易切入的起点,因为工作人员使用统一界面和有限工具;软件工程则更难,架构、协作和品位意味着,即便在一些原本可验证的领域,完整建设也要持续数年。
Foody 表示,如果评测创建成为全球最常见的知识工作岗位,他“不会感到意外”。现有员工可以把什么叫优秀表现编码下来,市场上的承包商则负责扩大覆盖范围,让经济活动从反复执行任务,转向一次性支付固定成本来教会智能体。
Elad 转述了一位朋友关于 Nyquist 的类比:人类可能知道某个东西比自己更聪明,却无法测量它究竟聪明多少。随后他问,如果一个模型已经比单个医生更强,医生的反馈最终是否反而会让模型退化,并以 Med-PaLM 2 为例。Foody 回应称,模型应学会把专家有价值的知识与专家自身的错误区分开来。
5. RFT 将企业定制转化为结果设计问题
模型可能会帮助提出自己的评测标准,再由人类进行验证,但 Foody 仍预计,领域专家需要为整个过程提供锚定。对于那些只有少数人仍然比模型更强的窄领域问题,前沿专家可能获得极高报酬。
更宽泛的智能体工作流仍为普通专业人士保留空间:传达诊断结果、协调工具和发送后续信息,都不只是回答一个医学问题。这一区分意味着,专家报酬会越来越呈幂律分布,但不会立即排除分布中部的人群。
Foody 还认为,模型很快可能成为更强的管理者:拆解大问题、分配工作,并对人类进行绩效管理。Sarah Guo 将产品缺口描述为:只要获得足够的上下文和目标,助手就能完美地安排工作的优先级、分配任务并确定顺序;Elad 则把理想行为压缩成一句话:“告诉我接下来3分钟该做什么。”
底层推理能力可能已经足够。强化微调通过奖励期望结果,教会模型如何选择工具,以及如何理解企业特定的成功标准;与 SFT 的输入—输出配对不同,Foody 称其“数据效率极高”。Elad 将实际所需规模概括为数百到数千个样本,而不是10亿个 token,Foody 表示认同。
6. Mercor 正在走向人类与智能体共用的单一市场
Mercor 的重点是吸引“全世界最聪明的人”,并预测他们的工作表现。由于普通劳动力市场的供给侧与需求侧比例大约为50:1,免费的模拟面试、职业建议和可分享个人资料,可以为那些从未获得工作机会的候选人创造价值;业务的另一端则为这些工具提供资金。
市场网络效应已经显现,但 Foody 预计,绩效数据飞轮会变得更重要:客户反馈会揭示谁取得了成功、为什么成功,从而改善后续筛选。长期来看,产品可能会将每个问题匹配给由人类劳动者和 AI 智能体组成的某种协作组合。
Foody 表示,早期的共享招聘系统,例如 Hired,可以汇总简历和人脉,但无法大规模记录、进行面试,也无法分析解释工作表现所需的数据。LLM 带来了更完整、更自动化劳动力市场的“为什么是现在”。
他的终局是一个全球统一的劳动力市场。如今,候选人会申请十几个岗位,而旧金山的一家公司只会考虑全球人才中的极小一部分;如果评估成本由软件承担,每名候选人都能进入同一个市场,每家公司也都能从中招聘。
对初创公司而言,Foody 建议早期不妥协地追求人才密度,随后随着组织扩大,转向数据驱动的标准。Elad 转述了硅谷时代的传闻:Google 招人能力很强,但花数年时间才会清理表现不佳者;Facebook 早期人才池良莠不齐,但更擅长淘汰低绩效者。
Hi, listeners, and welcome to No Priors. Today, we're chatting with Brendan Foody, co-founder and CEO of Mercor, the company that recruits people to train AI models. Mercor was founded in 2023 by three college dropouts and Thiel Fellows. Since then, they've raised $100 million, surpassed 100 million in revenue run rate, and are working with the top AI labs. Today, we're talking about where the data for foundation model training will come from next, evaluations for state-of-the-art models, and the future of labor markets. Brendan, welcome to No Priors. Brendan, thanks so much for doing this.
Yeah, thanks for having me. Excited to be here.
You guys have had a wild last 6 months or so.
Mm.
There’s huge traction in the company. Can you talk a little bit about what Mercor does?
Yeah. So at a high level, we train models that predict how well someone will perform on a job better than a human can. Similar to how a human would review a resume, conduct an interview, and decide who to hire, we automate all of those processes with LLMs, and it's so effective, it's used by all of the top AI labs to hire thousands of people that train the next generation of models.
What are the skills and job descriptions that the labs are looking for right now?
1. The Human Data Market
It’s really everything that’s economically valuable, because reinforcement learning is becoming so effective that once you create evals, the models can learn them and improve their capabilities. For everything that we want LLMs to be good at, we need evals for those things.
It ranges from consulting to software engineers, all the way to hobbyists, video games, and everything that you can imagine under the sun. Whatever capabilities you’re seeing foundation model companies invest in, or even application-layer companies invest in, the evals are upstream of all of that.
Are you also helping companies outside of the core foundation models with this type of hiring, or is it mainly focused on AI models right now?
Yeah. When we started the business, it was totally unrelated to human data. We saw that there were phenomenally talented people all around the world who weren’t getting opportunities, and we could apply LLMs to make the process of finding them jobs more efficient.
Then we realized, after meeting a couple of customers in the market, that there was this huge vacuum because of the transition in the human data market. The human data market used to be a crowdsourcing problem: How do you get a bunch of low- and medium-skilled people who are writing barely grammatically correct sentences for the early versions of ChatGPT?
It was transitioning toward this vetting problem: How do you find some of the most capable people in the world who can work directly with researchers to push the frontier of model capabilities? We’ve still kept that core DNA of hiring people for roles, human data and otherwise, and a lot of our customers hire for both.
Do you think all hiring eventually moves to these AI systems assessing people, or at least all knowledge work?
2. Models Predict Performance
I think certainly, because we’re already seeing on most of our evals that models are better than human hiring managers at assessing talent, and it’s still very early innings.
I think we’ll get to a point where it’ll almost be irrational not to listen to the model, right? Maybe for legal reasons, we’ll still have the human pressing the button and making the final sign-off, but we’ll trust the model’s recommendations on who should be doing a given task or job more than we trust humans.
I guess in any field, people say that there are 10X people. There are 10X coders who are way more productive than the average coder. There are 10X physicians or investors, or you name it. Do you see that in terms of the output of your models? In other words, are you able to identify people who are outliers?
Totally. This is one of the most fascinating things. The power-law nature of knowledge work frames the importance of performance prediction. Imagine if you can understand the kinds of engineers on an engineering team who are going to perform in the 90th percentile. Or even if you could say, “I know that this person who costs half as much is going to perform in the top quartile.”
It frames how you think about the value that we create for customers and how you think about the long-term economics of the business. It all ties back to how you measure the customer outcomes and really home in on them.
And is it a power law, or what sort of distribution is it? Because people always talk about human performance—
Yeah.
—as a bell curve. Do you think that’s actually true, or do you think that’s the wrong way to interpret human performance relative to knowledge work?
It’s very industry by industry. For you in investing, it’s the most power-law thing imaginable. The top handful of companies each decade are the ones that matter a disproportionate amount, and it’s the investors who win in those companies who matter.
If you’re hiring factory workers, it’s a much more commoditized skill set, and there’s a lot less difference. Software engineering is somewhere in between. It’s definitely very power law, but I don’t think it’s as power law as the handful of best investors in the world.
Do you have a prediction for where you should expect that models are better at evaluating or identifying talent, either because of the distribution of skill level or because of measurability?
Yeah. The models are really good at everything that you can measure with text. If you can ask questions in an interview and read through the transcript, the models are superhuman at that across many more domains than one would think. It’s more domain-agnostic than I would have initially anticipated.
I think the things where models are going to be slower are multimodal signals, such as understanding how passionate this person is about what they’re working on, or how persuasive they are or how good they are at sales. Those capabilities will come, but they’ll just take a little bit more time. That’s my mental model for thinking about it right now.
Right. If I’m interviewing a candidate from one of our companies, and they’re saying the right words about their motivation level but I don’t believe it, that might be a next-level signal if I have any predictive power here.
Totally.
Right?
Totally, exactly. The other thing is that the models are way better at high-volume processes.
Say you’re assessing 20 people for the same job and you hire those people and see how they perform. It’s very easy to attribute features of each person’s background to how they perform. It’s sort of stack-ranking, where you can understand that this person had this nuance in their interview, or this person had this nuance in their resume, and that was the thing that explained how well they performed on the job.
If those 20 people are performing 20 different jobs, then it’s just a mess of figuring out what is causing what. It’s way more difficult to understand what features are actually driving signal. I think those higher-volume processes will also be automated first.
Is there anything that surprises you about the discovered features in terms of any domain that you’re working on today that identifies amazing talent?
That’s a very good question.
Or maybe in engineering, because it’s relevant for many of our listeners.
3. The Hidden Signals of Talent
Yeah. One of the really interesting things about engineering is that there’s so much signal about a lot of the best engineers online that I don’t think people properly tap into. It’s everything ranging from their GitHub profiles to the personal projects on their websites to the blog posts that they wrote during college.
There’s a lot of signal, but it’s bottlenecked by manual processes. Hiring managers don’t have time to read through all of this stuff. With designers, they don’t have time to consider every proposal or every image from someone’s Dribbble profile before doing their top-of-funnel interviews.
I think one of the things where people are under-indexing on signal the most is the information that can be found online. A lot of the things that can be picked up on during an interview—how passionate this person is, whether this person has the skills required for the job—I think humans are relatively good at. At least, they’re a little bit more adapted to it right now.
Are there hidden signals for other types of domains where there’s less online work? An example of that would be physicians—
Mm-hmm.
—lawyers. There are a lot of other professions where—
Totally. Yeah. There are all sorts of these hidden signals.
One interesting signal we've seen in the past is that people who are based internationally but study abroad in a Western country tend to work much more collaboratively or communicate better with people. These are the kinds of signals that make sense when you look backward and evaluate them but are hard for a human without full context of everything happening in the market to really understand and appreciate.
One of the most important things, as you can imagine, is how intrinsically motivated and passionate people are about a domain. So you look for signals not just on their résumé and in their interviews, but also online, of what indicates that. It pertains not just to whom you hire but also to what those people should be working on, right? Imagine the nuance between hiring a biology PhD to work on biology problems versus hiring the person who wrote their thesis on drug discovery to solve problems and come up with innovative solutions contextual to their thesis. There's just so much inefficiency with the way we do matching and use all of those signals right now.
So you're evaluating people. Are you also evaluating the models relative to the people?
Yeah, of course.
When or what is your view in terms of the proportion of people who will eventually get displaced by these models? In other words, if you can tell the relative performance and you can look at relative output, how do you start thinking about displacement or augmentation, or other aspects like that?
4. The Coming Work Displacement
I think displacement in a lot of roles is going to happen very quickly, and it's going to be very painful—a large political problem. I think we're going to have a big populist movement around this and all the displacement that's going to happen, but one of the most important problems in the economy is figuring out how to respond to that, right? How do we figure out what everyone who's working in customer support or recruiting should be doing in a few years?
How do we reallocate wealth once we approach superintelligence, especially if the value and gains of that are more of a power-law distribution? I spend a lot of time thinking about how that's going to play out.
Mm-hmm.
I think it's really at the heart of—
What do you think happens eventually? X percent of people get displaced from white-collar work. What do you think they do?
I think there's going to be a lot more in the physical world. I think there's also going to be a lot of niche skills—
What does the physical world mean?
It could be everything ranging from people creating robotics data to people who are waiters at restaurants, or who are just therapists because people want human interaction—whatever that looks like. I think automation in the physical world is going to happen a lot slower than what's happening in the digital world, just because of so many of the self-reinforcing gains and the amount of self-improvement that can happen in the virtual world but not the physical one.
Mm-hmm.
Do you have a point of view on what types of skills, knowledge, and reasoning are worth investing in now as a human expecting to stay economically valuable?
5. Skills That Survive Automation
Sam Altman said this thing when someone asked him this, about how people should optimize for just being very versatile and able to learn quickly and change what they do. I think that resonates a lot because there are so many things that one would think the models aren't good at that they get very good at very fast. I almost think you just need to be able to navigate that quickly.
What are the characteristics of those things that you think models will learn the fastest? If you were to say, "Here's a heuristic," what do you think are the components of that?
If it's verifiable. Things like math, or soon code, that are verifiable will get solved very quickly.
So you want a feedback loop or utility function they're optimizing against as a model?
Exactly.
For things that aren't verifiable, like maybe it's your taste in a founder, that's much harder to automate. It's also a very sparse signal because there's just not that much data on it.
This is a pretty fundamental research question right now, but what do you think are the most interesting ideas about verifiability beyond code and math?
I think there are ways that you can have certain autograders or criteria that humans can apply, or that models can apply, and I'm very interested in how that will play out over time. There are obviously a lot of other domains where models will take unstructured data, structure it, figure out how to verify it, and it's very much industry by industry.
I think it's going to be hard for 1 lab to do everything there. There will be more specialization as we progress further and further, and marginal gains in each industry become more challenging.
How much do you believe in a generalization from code- and math-type reasoning and intelligence? If I'm this much better at proof math, does it make me funny eventually? Me being the intelligence.
I generally believe in it, but to a certain extent. You still need a reasonable amount of data for the new domain to kickstart it. But there are going to be a lot of transfer learnings.
I think it's very funny when Sarah does proofs, so I think it all fits.
She's good at proofs?
Spurtle.
Yeah.
I actually think being bad at proofs is funnier. Okay, let's talk about evals, because you're working on the bleeding edge of model capability. There has been this whole sense of what people call an evaluation crisis, where the models are so good and somewhat indistinguishable at the fringe of capability today that we don't know how to test them.
Ignoring all the issues with people gaming the benchmarks, what do you think? What right ideas are there about evaluating models, especially as they become superhuman?
6. Agents Need Better Evals
I think one of the most important things is that a lot of the evals historically have been for zero-shot performance of a model, or a test question. That might be academic. The thing that we actually need to evaluate is economically valuable work, right?
When a software engineer goes to their job, it's so much more than writing a PR. It's coordinating with all of the relevant parties to understand what the product manager wants, how that fits into the priorities of each team, and how that all translates to the end output of the work. I think we're going to see an immense amount of eval creation for agents, and that is the largest barrier to automating most knowledge work in the economy.
Where should people start? That feels not terribly generalizable.
Yeah.
So Sierra has something called τ-bench that I think people are trying, and there are other efforts here, but it is perhaps more specific to a certain function.
Yeah. I think people will need to have these by industry, and they should probably start with tasks that are more homogeneous. For customer-support tickets, I think that's a great example because there's 1 interface that the customer-support agent interacts with. Maybe they call a couple of tools, like accessing the database or reading through the documentation, but it's a relatively homogeneous, uniform task.
I think the things that are going to be more challenging but also, in many cases, more valuable are creating evals for these very, very diverse tasks—all of the things that go into making a good software engineer. That's going to be really hard to do. I think it's going to be a years-long build-out for even some of the verifiable domains, because there's so much that goes into being a good software engineer: Do they have taste for what is the right way to approach a problem, or what are the products that people really enjoy using? I'm really excited for that.
If you were to counsel people with young kids—say, your child is 5 to 10—should their kids learn computer science?
I would probably not push them toward teaching their kids computer science, but I'm not totally against it.
What would you teach them?
I would encourage them to find something that's intellectually stimulating, that they're really passionate about, and where they can learn general reasoning capabilities. Those reasoning capabilities will probably be very valuable and cross-applicable.
I always loved building companies growing up, hustling, and doing small things like that, and I think that is something that could be helpful.
But I am skeptical that the really valuable thing is just people who can code in 5 years. I think it’s much more likely that the people who have these contrarian ideas around what’s missing in markets and have the taste for what features and nuances need to go into solving that problem.
Mm-hmm.
You’ve said taste a few times. Are there signals of taste that you feel like you can discover in any domain?
Yeah, absolutely. I think that oftentimes you just want to see the softer signals of how people think about certain problems. Certain people have intuitions, whether it be the way they approach a problem or, if they’re looking at different products, how they notice nuances. It’s very contextual to the industry, but it’s important to measure.
How do you score it?
What’s the positive feedback loop here?
We’ve done a variety of things, but oftentimes we will give people a problem that as closely as possible mirrors what they would solve on the job, and then we would see how they compare to other people. And so that helps with scoring it.
Yeah, some further thought process is part of that. I know, for example, it’s almost like looking at code reviews or other sort of intermediate work along the way relative to something.
Totally.
Yeah.
Relative to something.
We definitely do. One thing I’ve realized about talent assessment is that a lot of people focus too much on the proxy for what they care about rather than the thing they actually care about. Ideally, you want to measure the thing that you actually care about, so if it’s that person building an MVP of the product, ideally you have an interview that’s a scoped-down version of doing that.
The place where you need to use proxies is when it’s a longer-horizon task, where you just want to structure the proxy to get as much signal as possible. That’s sort of how I think about talent assessment.
Can I ask a scale-of-impact question?
Mm-hmm.
So if I think about the very largest employers today, let’s call it low single-digit millions of employees, right? Or, I don’t know, when you think about contractors and Amazon workers and such. But how many people do you think will end up doing data collection?
I think it’s a huge volume. I think the reason is that it all comes down to creating evals for everything in the economy. I think part of that will be current employees of businesses that are creating evals for that business so that those agents can learn what good looks like. Part of that will be hiring out contractors through a marketplace to help build out those evals. But it would not surprise me if that becomes the most common knowledge-work job in the world.
How long does that last? Effectively, people are being brought on to displace themselves.
This is true.
Is that a 6-month cycle? Is it a 2-year cycle? What is the length of time at which people have relevancy relative to some of these tasks?
There’s always a frontier.
Ah, unless it becomes superhuman, right?
Yeah, unless it becomes superhuman.
In other words, it’s almost like time to superhuman.
But I had an interesting conversation, which is that you don’t even know that you have superintelligence without having evals for everything.
Mm.
Because you need to understand what the human baseline is and what is good, and it’s grounded in this understanding of human behavior.
Yeah, a friend of mine basically believes that—do you know Nyquist’s theorem?
Mm-hmm.
It’s basically that if you’re sampling a signal, you need to be able to sample it at twice the frequency in order to actually extrapolate what it is. Otherwise, you’re not sampling richly enough to know.
And so he views that there’s some version of that for intelligence. You can tell if somebody’s smarter than you, but you don’t know how much smarter because you aren’t capable of sampling rapidly enough to understand it. I always wonder about that in the context of superintelligence or superhuman capabilities, in terms of how smart you can actually be, since it’s hard to bootstrap into the eval.
Well, I think when you take it to the limit and you have superintelligence, what you’re saying makes a lot of sense. But another way I think about it is that if we classify knowledge work in 2 categories, one is solving an end task where it’s a variable cost—you need to do that repeatedly—and the other is creating an eval to teach a model how to solve that task, which is a fixed cost that you incur 1 time.
It does seem structurally more efficient for work to trend away from the variable cost of doing it repeatedly toward this fixed cost of how we build out the evals and processes for models to do this themselves. That said, it all comes down to how fast we’re approaching superintelligence, right? If the models are just getting that good that fast, then sure, I don’t think we would need humans creating evals very much.
But I also don’t think we would need humans in many other parts of the economy. And so you sort of need to be thoughtful about the ratio of that.
Does that create an asymptote in terms of how good these things get, or do they start creating their own evals over time?
I think they’ll play a role in creating their own evals.
Some kind of bootstrap.
Yeah, they might come up with certain criteria for what a good response looks like, and humans validate that criteria. However, I think you often need to ground this in the experts in that particular domain.
Sure.
But I’m just thinking of Med-PaLM 2.
Yeah.
The output of the model was better than the average physician. It was basically a health model that Google built. They would use physician panels to rate outputs of the model versus individual physicians, and the model did better by far than individual physicians.
At some point it should do better than the physician panels, where feedback from the physician panel should make the model worse, right? In other words, if you just RL’d it off individual physicians, the model already was going to get worse. And so there’s a little bit of this question of when human scoring creates worse outcomes because the humans aren’t as good at a task.
Well, I think the models will be able to delineate between the valuable human knowledge and the human knowledge that’s not valuable. Maybe you have doctors who create a bunch of evals for this particular task, and the model realizes, “Wow, I see the mistake that the doctor made on these particular tasks, but I’m going to ignore them. Here are the things that seem insightful, or the things that I can learn.”
Sure.
And the models will use that data and value that data immensely. The other thing I’ll say is that I think it’s easy to look at these evals and the rate of improvement on the evals and just think we’re a lot closer to superintelligence than we are.
Mm-hmm.
But the truth of the matter is, there is a lot between being really good at SWE-bench and replacing software engineering.
Uh-huh.
There are all the coordination problems that we talked about. There’s so much else that goes into that, and I think we’re just going to need a lot of evals for tool use. We’re going to need a lot of evals for agents, and that build-out is going to be a lot longer than a couple-year time horizon.
Mm.
How do you think about incentives for all of these expert knowledge workers?
Mm-hmm.
Because for a great software engineer with taste and architectural understanding, the opportunity cost is a great job at Mercor or another interesting tech company, whereas some of the geo-arbitrage on basic knowledge work does not exist as the skill level increases over time. That’s true in coding, true for physicians, true for finance people—lots of areas where you might want evals and labels.
Totally. I think it’ll definitely become more power-law over time, which means that the best people are going to, of course, make an incredible amount of money.
Do you think it’s more just turning up the dial on what any piece of information is worth from the higher-skilled workers?
Yeah. But you also want the evals at the frontier of what the models can’t do. And so it might be that for a very well-scoped problem, like answering a medical question that someone has, you might need to get the world-class doctor who is one of the handful of people that’s able to be better than the model at that very well-scoped problem.
But for the broader agentic problem of how do we talk about this case in a way that the patient is receptive to? How do we then coordinate with this set of tools to help complete the diagnosis and send whatever emails at X time? I think for those kinds of things, I still expect that the bulk of the bell curve—people who are closer to the mean of the distribution—will be able to contribute for a longer period of time.
What do you think is the biggest shift that nobody's really anticipating that's coming? It could be domain-specific or broader.
Well, I'll answer this in 2 parts, because when you think about “nobody,” it feels like the bulk of the country is not really coming to grips with how fast jobs will be displaced, and that just feels like a big problem, as I said before. I think that we need to stay very proactive as a government, as an economy, et cetera.
Are there certain areas where you're already seeing large-scale job displacement that you don't think is being reported on?
It's definitely being reported on in customer support and recruiting. I think one of the challenges is that a lot of this happens during economic contractions, when people get more efficient and more focused on the bottom line. A lot of it hasn't happened yet, but it's going to happen imminently.
In terms of things that maybe no one, even in San Francisco, is thinking about, another interesting part of that problem is that these agentic evals for non-verifiable domains are significantly under-indexed. Another thing is that people in San Francisco tend not to think critically about the role humans will play in the economy because they're so focused on automating humans, and so I think it's important to think more about that problem.
One thing I've thought about is that, ideally, models should help us figure that out over time, right? What are the things that people are passionate about? What motivates them? Maybe it doesn't need to be an economically valuable thing. Maybe it's just a certain kind of project that they like working on, and I think that people aren't indexed enough on how humans will fit into the economy in 10 years.
You know, one thing that I feel I've really misunderstood, or didn't quite understand the scope of, was the degree to which we effectively had different forms of UBI, or universal basic income, in different sectors of the economy. Government is a clear example—
Totally.
…where there's enormous waste, fraud, grift, et cetera, happening.
Yeah.
Parts of academia, if you just look at the growth of the bureaucracy relative to the actual student body—
Mm-hmm.
…or faculty. Big tech—
Mm-hmm.
If you look at the size, it basically shows that a lot of these things were effectively UBI. To some extent, one could argue that parts of our economy are already experiencing what you're saying, in terms of there being high-paying jobs that may or may not be super productive on a relative basis. And so the question is: Is that something that we actually embrace as a society, given some of these changes in displacement? And if so, where does that economic surplus come from?
Yeah. It's interesting. I think that as we have better analytics around the value of employees, it seems intuitive that these companies will start doing more layoffs, more cuts, et cetera.
Do you think those evals become illegal at some point? Because it feels like that happened a little bit with certain aspects of merit, or merit-based testing, for different disciplines or fields. That happened with the government in the '70s, where they removed it as a criterion. I'm just wondering if that becomes something that, more generally, people may not want to adopt because it exposes things, or do you think it's something that's economically inevitable?
There's definitely going to be pushback, but I think it's inevitable economically because it's hard to regulate and so strongly valuable to companies that they'll move toward it. Should companies adopt that now?
I think it depends on what segments of the economy, because some of these are not economically driven already. They're just not efficient as sectors, but if you look at healthcare or education, everybody's seen this chart that shows a bunch of industries that have some measure of output per dollar spent, and you have increasing spend on healthcare and education with no improved output.
And that's happened for a long time, when there's been an increase in productivity in many other sectors. The answer is that there's no economic pressure, actually.
Yeah.
Sure. It's regulated versus unregulated sectors, effectively, and the regulation is what causes the divorce from economics.
Yes.
Yeah. Also, one thing that I think is very interesting is that a lot of people are in the mindset of AI being really good as an independent contributor, when actually it may soon become much better at being a manager, right? In taking a large problem, breaking it down, and figuring out how to performance-manage people and how they should be doing.
This ties into your point around what we should do with all of those unproductive employees. If we have a ruthlessly rational agent making that decision, it is probably going to be very different from a lot of the decisions that have been made historically.
One of our companies recently asked what I would expect an assistant to do that it doesn't do today. Right? I think the biggest thing is: If I give it enough context and some objectives that I'm trying to achieve—I'm not a particularly organized person, and I have a lot of output, I think, all things relative—but is it perfectly prioritized, tasked out, and sequenced so I'm not bottlenecked on a particular thing? No. And I would absolutely expect that the assistant can do that for me.
Totally.
Yeah.
Well, it goes to the point earlier, right?
Just tell me what to do for the next 3 minutes.
We have these models that are incredibly good at math, right?
Yeah.
You give them a test and they can ace the test.
Mm-hmm.
But they still can't do basic personal-assistant work, right? And I think it goes to show that there's still a lot of research and product to be built out in how we actually bridge the gap with what's economically valuable to complete that end-to-end job that you're willing to pay a human salary for.
Do you think the models are good enough for that and there's just incremental engineering work to make it better?
They are.
Or do you think it's—okay. So we actually have model capabilities that you think would allow us to build certain types of agentic systems, versus we need—
That are proactive, too.
Yeah.
Or—actually, maybe let me put it this way. I think with a small amount of evals for agents in various categories, the base model has all the reasoning capabilities. The reason you still need those evals is that the models need to understand when they should be using tools in certain ways, and they need to understand how to synthesize information from those tools.
But it's not a reasoning problem. It's much more the problem of learning each company's knowledge base and what good looks like in that role. There is going to be some post-training, and I'm very bullish on RFT and everything that's going to mean.
Can you say more about RFT and explain it for our audience?
Yeah. Basically, everyone used to talk about fine-tuning the model in the context of SFT, or supervised fine-tuning, where you would have inputs and outputs for a model, and the model would learn from those input-output pairs.
The main issue in supervised fine-tuning was that customization never really took off because it wasn't very data-efficient. Companies would create a few hundred and eventually try to scale it up to tens of thousands or hundreds of thousands of SFT pairs, but oftentimes wouldn't be able to get a lot of the capabilities that they were looking for.
Whereas in reinforcement fine-tuning, you instead define the outcome that you care about. In Sierra's case, I was talking with them about how they define what a good customer-support response would look like. In our case, we define the key things that you should identify as characteristics of this candidate, whether it is that they're passionate during their interview, that they demonstrate XYZ domain knowledge, or that they worked on this side project that demonstrated that skill.
Then you reward the model for identifying that. You set the solution, and then the model can learn in that environment how to get really good at it. The reason I'm so optimistic about it taking off is that it's profoundly data-efficient, right? And it finally makes sense to customize models at the application layer in the enterprise.
And “profoundly data-efficient” is actually hundreds to thousands of examples—some tenable number for an enterprise or a medium-sized business to think about—versus, I don't know, a billion tokens.
Yeah.
Yeah.
Yeah, exactly. And so it'll be very cool. I think we're going to have these agents that fill all roles that employees currently fill, working alongside employees.
Human employees will help create the evals. I also think that contractors in our marketplace will play a large role in that. It will just be this huge build-out of evals to create custom agents across every enterprise.
What is most important for Mercor to get done in the next year or so?
7. Mercor Scales Its Talent Network
There are 2 things that we focus on as a business, and I think those will be most important for this year as well as for the next 5 years. The first is: How do we get all of the smartest people in the world on our platform? That ties into the supply side of our marketplace and the marketplace network effects, similar to an Uber or Airbnb, because if we have the best candidates, then we're able to give them job opportunities and understand what they're looking for.
The second thing is predicting job performance.
Are you trying to offer anything that isn't comp?
Yeah, we are. One of the things that we realized is that the average labor marketplace has a 50-to-1 ratio of supply side relative to demand side, which means the average person who applies talks to their friend who also applied, and neither of them got jobs. It's almost this structural part of building labor marketplaces.
The way to actually scale up the labor marketplace to have hundreds of millions of the smartest people in the world on the platform is to build all of these free tools, such as AI mock interviews, AI career advice, and shareable profiles for people—all of the things that create the most magical experience possible for consumers—and give that away for free because it's powered by this monetization engine on the other side of the business. That's a very significant focus for us.
I interrupted you.
Yeah.
You were going to talk about what else was important.
Part 2.
Yeah.
It's performance predictions. We get all the data back from our customers about who's doing well and for what reasons, and how we can learn from all of those insights to make better predictions around who we should be hiring in the future.
That's the data flywheel that you would find in many of the most prominent companies in the world. I think that the marketplace network effect is the more obvious one when you look at the business, but I actually believe that the data flywheel will become more important over time based on a lot of the initial results that we're seeing.
How do you view the labor markets evolving over the very long term?
I think that the largest inefficiency in the labor market is fragmentation, and that a candidate, wherever they are in the world, will apply to a dozen jobs, while a company in San Francisco will consider a fraction of a percent of people in the world because it's all constrained by these manual processes for matching. They need to manually review every résumé, conduct every interview, and decide who to hire.
When you're able to solve this matching problem at the cost of software, it makes way for a global, unified labor market that every candidate applies to and every company hires from. I believe that that's not only the largest economic opportunity in the world, but also the most impactful one, insofar as how you can find everyone the job that they're going to be passionate about and successful in.
Would that include AI agents? In other words—
I think—
—the marketplace would be a hybrid of people and agents all competing for labor globally?
I think so, because customers ultimately come with a problem to be solved, right? Ideally, it's some coordination of how those 2 fit together.
Given you spend all your time thinking about how to attract high-skilled candidates and determine their effectiveness, what advice would you have for people who are hiring in startups and scaling companies?
Early on, it's hard to stress the importance of talent density. There's always a trade-off between hiring speed and hiring quality, and for those early employees, you should always index on quality. You need to be patient, and you need to make sure that people are extremely high caliber.
When you're scaling up an organization, you obviously don't want to drop those standards, but people need to be a lot more data-driven around the characteristics of people who actually drive the outcomes they care about. A lot of problems happen when that slips, when it's this vibes-based assessment that doesn't scale very well, where each hiring manager is doing it in a fragmented way and it's hard to enforce those standards across the board.
Being very disciplined around your hiring goals, the characteristics of people you know are actually going to achieve the business outcomes you care about, and how you measure those things is really important.
I find that almost every great company either hires well, like what you're talking about, or fires well, which is your phase 2. But I think often they do one of those things really well early.
For some reason, most people don't seem to get both right early on. I don't know why it is. I think it's almost a founder bias or something like that. Hopefully, over time, they pivot into both.
Google was a good example of an organization that would always hire well but couldn't fire well.
Mm-hmm.
It took them a really long time to clean people out. Years. Literally years.
Interesting.
Facebook, on the other hand, was known for a more mixed early talent pool, but they were very good at removing early people who weren't performing. I always thought that was an interesting dichotomy between the 2.
Those were the rumors in the Valley when each company was—
Yeah.
…tens or low hundreds of people. Now, obviously, they're all very professionalized in terms of how they do both.
They have their UBI.
Yeah, exactly. So I always thought that was interesting.
I think it's—because I mostly think about engineering hiring, go-to-market hiring, and investor hiring—they're all professions that have some timescale of outcomes that isn't an hour. I think you're always looking for a proxy for outcomes in these longer-outcome jobs.
I think there's a really interesting question, very related to evals and assessment: What are the proxies we're going to discover for each of these roles? I think it's a huge shortcut in hiring—hiring well, not necessarily firing well. If you can do references and work trials with engineers, you actually know a lot in the first 5 days or 30 days of whether or not something's going to work out.
Totally.
I'm always looking for proxies for that.
One of the crazy things about the market is that any candidate you do a work trial with has probably done work trials with a lot of other top companies in San Francisco, but you don't have any of the data on that.
Mm-hmm.
Right? Obviously, there are interesting data privacy and centralization questions: Companies want that to be their proprietary knowledge. But I think that market is going to trend toward becoming a lot more efficient over time.
Or even references for people you don't hire. Theoretically, it's beneficial for the top companies to understand the reasons that other companies in different markets aren't hiring specific candidates, et cetera.
What do you think companies that attempted some sort of common, generic evaluation, like Hired in a previous generation, got wrong? The theory of having a common application of some kind or a shared assessment has existed but hasn't worked at scale or at quality.
I think LinkedIn centralizes and aggregates the very first layer of the application process: What are the things that this person has done, and who are they connected to? The challenge historically has been that the rest of the process to facilitate a transaction has not been possible to aggregate and automate.
It wasn't possible to actually record all of these interviews and scalably conduct interviews of everyone. It wasn't possible to get all of this data and analyze it properly to understand what goes into causing someone to perform well. I think there's just this huge “why now” that's enabled by LLMs becoming so capable so quickly.
That makes sense. One of the theories that my partner Mike has is around the scalability of LLMs being able to interrogate humans and the usefulness of that data in a bunch of different domains.
It would be great to see the aggregate of that for hiring.
My cofounders and I are all Thiel Fellows, and we're very passionate about how we could apply LLMs to help identify the next Thiel Fellows.
I often wonder: Imagine if you could have Peter Thiel, as a heuristic, interview everyone in the world when they're 18, right? Maybe he could meticulously spend time determining who is actually going to be good at what job. I think we're approaching that world very quickly. It'll be fun to see how that impacts the labor market, the investing market, and everything else.
That's really cool.
Thanks for doing this, Brendan.
Yeah, this was awesome.
Yeah, thanks for having us.
Thanks for coming.