[BidClub_]
Gradient Dissent · · 56 分钟

为 AGI 提供底层数据的创业公司

Lukas BiewaldEdwin Chen

YouTube
TL;DR
  • 2020 年创立、未获风险投资支持的 Surge 在 2024 年营收突破 10 亿美元,员工仅略超 100 人——意味着人均营收约 1000 万美元。 Chen 将这种资本效率归因于坚持使命导向、避免打造组织“帝国”,以及向真正重视数据质量的研究人员销售。对于 Surge 认为与 AGI 无关的项目,公司甚至会放弃其收入。

  • Surge 所称的差异化优势是质量控制技术,而不是廉价劳动力。 Chen 将竞争对手形容为依赖电子表格的“人肉外包工厂”,同时强调用于衡量工人与任务质量、追踪质量变化、开展实验和过滤低质量产出的系统。「不能只是把一群人扔进去」(“You can’t just throw warm bodies at it.”)。

  • 生成式 AI 数据让旧有的共识式标注模型变得不够用,因为多数偏好可能筛出通用、最低公分母式的产出。 一首 8 行月亮诗可以满足所有检查项,却依然糟糕;英语文学博士也可能缺乏所需的创作能力。Surge 转而试图保留多元洞察——写诗或证明勾股定理都可以有“一千种方式”。

  • 需求已从搜索评测和内容审核,转向多模态、多语言、专家级 LLM 工作,足以让贡献者连续投入数天甚至数周。 Surge 目前覆盖超过 50 种语言,包括用阿根廷西班牙语编程,以及在玻利维亚提供法律或金融专业知识。Chen 预计前沿将从研究生 STEM 能力进一步迈向真正的科学协作。

  • Chen 认为,基准测试激励可能带来更高的分数,却造出更差的模型。 在 LMSYS/Chatbot Arena 排行榜上,普通评测者可能奖励更长的答案、格式和 emoji,即使答案在幻觉或无视指令;据称,一些研究人员只有在模型获得 10 个排行榜积分后才有资格晋升。Chen 的判断异常绝对:这个排行榜“至少让行业倒退了 1 年”。

  • Surge 倾向于在有限的 SFT 启动之后采用 RLHF,同时认为合成数据有用,但一旦压过人类信号就会带来风险。 一项评测中,使用 1000 万至 2000 万道合成数学题训练,改善了狭窄的学术领域,却让模型在其他领域变差。Chen 说,一些实验室后来丢弃了规模相当的数据,因为发现“哪怕只有 1000 条真正高质量的人类数据”也更有用。

  • Chen 看不到模型开发的瓶颈墙:下一轮扩展将来自更丰富的 RL 环境、更长周期的智能体、个性化、创造力,以及持续投入的人类数据。 一个例子是模拟的 AI 初创公司,里面包含 Gmail、Slack、Jira、GitHub、代码库,以及突发的 AWS 或 Slack 故障。按照当前激励机制,他猜测闭源模型将继续胜出;不同实验室的数据支出可能约占算力支出的 1% 至 10%,而这部分占比应该更高。

摘要 · 为研究而整理的核心内容

1. Surge 从所谓简单的标注失灵中诞生

  • Chen 在 Twitter 遇到的创业问题看似简单:一个情感分类器只需要 10,000 条标注推文,但 2 名通过 Craigslist 招来、朝九晚五工作的人花了 1 个月才启动,又花了 1 个月交回“彻底的垃圾”。他们无法识别俚语和 hashtag,Chen 于是花了 1 周亲自标注数据,因为这样更快、质量也更高。

  • 更深层的失败来自目标设计。团队以点击和转发为目标优化 Twitter 信息流,结果形成了由标题党、色情擦边内容、比基尼图片和“10 种最可怕皮肤病”之类榜单文章构成的负反馈循环。团队希望基于更丰富的产品原则获得人工评分,但一个连情感都无法可靠处理的系统,也无法提供这种更高阶的信号。

  • Surge 于 2020 年创立,刻意聚焦语言和行为,而不是绘制汽车边界框这类商品化任务。早期客户包括 Airbnb、Twitter、搜索和算法创业公司,以及那些已经“迫切需要”更高质量系统、能够承担更复杂工作的联系人。

  • Biewald 强调了商业结果:Surge 在 2024 年营收突破 10 亿美元,更早达到 10 亿美元 ARR,且没有接受外部融资。Chen 说,他们希望客户因为相信质量而购买,“而不是因为在某篇科技文章里看到了我们”。

2. 运营杠杆来自软件和研究文化

  • Chen 不认同质量会随着部署 10,000 人或 100,000 人而自动出现。他将一些竞争对手形容为手工化的“人肉外包工厂”:工程师通过电子表格创建或审核数据,却无法量化某名工人在特定任务上的质量。

  • Chen 说,真正严肃的质量系统需要复杂算法来识别高质量数据、剔除最差产出;它还应量化工人在特定任务上的能力、追踪质量随时间的变化,并支持仪表盘、A/B 测试和改进算法。难点还会叠加,因为 LLM 必须覆盖诗歌、物理、编程,实际上还包括“世界上的每一项任务”。

  • 内部团队略超 100 人,被定位为一家数据研究机构。Chen 避免采用简历导向的招聘,也避免招来想要打造组织帝国的管理者;他的警示循环是:招聘带来会议,会议降低生产率,而生产率下降又被用来证明还需要继续招聘。

  • 销售是 Chen 个人最大的跨越。他将 Surge 定义为一家研究公司,目标是生成能够帮助 AGI 的数据,而不只是创造收入;由于没有外部董事会和 VC 压力,公司可以自由拒绝与 AGI 无关的工作。

3. 生成式数据质量需要品味、多样性和第一性原理

  • Chen 最经典的测试是一首关于月亮的 8 行诗。检查它是否是一首诗、是否有 8 行、是否包含“月亮”,只能证明它遵守了要求;招聘英语文学博士也解决不了问题,因为学历不会让一个人成为 Hemingway、Emily Dickinson 或诺贝尔奖级别的诗人。

  • 诗歌可以通过水面月光、内韵、格律或情感成立;数学同样有“1000 种证明勾股定理的方法”。优化评审者之间的一致性,对区分猫和狗很有用,但当目标是创造力和洞察力时,最终只会得到“最低公分母”。

  • Biewald 的反驳值得保留:即使是边界框也隐藏着判断——被遮挡的汽车、镜面反射,或广告牌上的汽车图像,都会让规格变得复杂。Surge 的回应是从研究人员的目标和预期用途出发进行推断;当它认为最佳数据与客户指令不一致时,就展开一场“有火药味的辩论”。

4. 工作正变得多模态、多语言且难度陡增

  • Surge 早期的任务主要是搜索评测和内容审核;如今的工作“几乎纯粹是 LLM 工作”。发展方向始终是更高的复杂度、成熟度和专业性,而不是随着自动化推进,人类数据需求就会消失。

  • 多模态任务如今会同时结合文本、图像、音频和视频。Chen 举的例子是:用手机拍摄某个东西,让模型编写一个模拟它的程序,并要求模型同时理解这些模态。

  • Surge 覆盖超过 50 种语言,甚至处理用阿根廷西班牙语编程、在玻利维亚提供法律或金融专业知识等高度具体的交叉场景。在高端任务中,贡献者会出题并解答奥林匹克竞赛级别的问题,有时一道任务要花数天或数周,而不是用 5 秒钟给图片打标签。

5. 排行榜可能奖励表演,却损害真实能力

  • 被问到为什么模型能取得 IMO 金牌选手般的分数,却答不出朋友的脑筋急转弯时,Chen 说前沿系统已经被“基准测试劫持”。学术、合成、封闭式测试鼓励狭窄优化,而真实数学工作往往是开放式探索,没有一个容易验证的单一答案。

  • Chen 认为人工评测是黄金标准,但他严厉批评 LMSYS/Chatbot Arena 排行榜。任何人都可以提交提示词、查看 2 个回答并投票;他说,许多用户——包括初高中生——只看几秒,就可能把票投给看起来更厉害的回答。

  • 可被利用的配方是更长的篇幅、粗体格式和 emoji。Chen 指出,他所谓的为 LMSYS 优化的 Llama 4 版本正符合这些模式;研究人员据称被告知,只有让模型在排行榜上增加 10 分才能获得晋升,即使被偏好的回答存在幻觉,或更差地遵循指令。

  • Biewald 注意到,原本不同的实验室在 2025 年前后出现了共同的声音:过度积极、像“童子军”一样循规蹈矩、急于讨好。Chen 认为,包括破折号在内的一些特征来自预训练,同时也指出了模型特有的痕迹,例如“absolutely”、形容词三连、Markdown 习惯和 emoji 频率;他预计不同模型的风格最终会更加明显地分化。

6. RLHF 携带比示范更丰富的信号,但合成规模可能适得其反

  • 在监督微调(SFT)中,人类同时提供提示词和理想答案,例如要求写一首关于奶酪的十四行诗,然后亲自写出答案。在 RLHF 中,模型生成多个竞争答案,由评测者选出更好的一个,从而学习对情感、工艺、格式以及糟糕之处的偏好。

  • Surge 过去的判断是,RLHF 数据的效果要强得多,但 Chen 保留了限定:部分 SFT 仍然重要,可以高效地把模型启动到能够进行 RLHF 的阶段。他说,Llama 2 论文描述的实验最终改变了那些原本预计 SFT 会占主导地位的研究人员的看法。

  • 合成数据在特定场景下有用,但 Chen 认为,无差别使用时,它会让模型“擅长合成问题”,而不是真实问题。一名研究人员的人工评测在使用 1000 万至 2000 万道合成数学题后大幅下滑,模型能力被收窄;据称,其他公司也花了数月丢弃类似数据,因为发现哪怕 1000 条高质量人类样本都可能更有用。

7. 未来的人类数据将从博士练习走向协作与品味

  • Chen 不认为当前的数据类别会彻底消失。它们的占比可能下降,就像人随着年龄增长而减少算术训练,但如果删除这些数据,前沿模型在被激进地调向新能力时,可能出现反常的能力回退。

  • 研究生级别的 STEM 能力只是中间阶段。要产生新发现,模型需要与超越普通博士水平的人合作;一个例子是 Stanford 教授与 AI 协作开展前沿研究,提出假设、测试实验,并作为真正的科学伙伴提供正确的信号。

  • 时间跨度将从几分钟或几小时延长至数天和数周。创作也必须摆脱模型“非常香草、非常通用的维度”:今天的诗歌和短篇小说可以拥有无可挑剔的文笔,却依然缺乏创造力,可能是因为更具经济价值的能力获得了优先训练。

  • 个性化同样几乎仍是一片空白。知道用户住在 Ohio 而不是 New York 很容易;理解决定一个人喜欢哪些故事或巴黎餐厅的潜在偏好,则要难得多。Biewald 5 岁的孩子说明了这种冲突:一篇同时针对父母和孩子优化、再对所有人取平均的故事,可能谁都无法满足。

8. 真实 RL 世界是下一块训练场,不是模型停滞的证据

  • 推理模型正推动 Surge 从零构建 RL 环境:带有工具和数据集的“电子游戏式宇宙”。一个模拟的 AI 初创公司可能包含 Gmail 和 Slack 消息、Jira 工单、GitHub pull request 以及代码库,然后让 Slack 或 AWS 突然宕机,迫使智能体规划、恢复、反思并寻找另一条路径。

  • Chen 预计模型会越来越多地使用代码,因为代码可以提供有用的可验证性。他因此对进展持乐观态度:新的 RL 方法、环境和尚未收集的数据让他确信,行业“绝对没有”陷入停滞,而且新领域很快可能出现“巨大进展”。

  • 他对 ARC 基准的坦诚回答是否定的:在看到它之前,他本来会认为模型“绝对”能解决那些看似简单却颇具迷惑性的视觉模式。Surge 正在生成 ARC 风格的问题,但 Chen 承认:“我并不真正理解模型为什么解不出来。”

  • 在行业结构上,按照当前激励机制,Chen 预计闭源模型会继续胜出,因为有价值且昂贵的模型会促使所有者捕获上行收益。他提到,据他描述,Meta 正在考虑将 Llama 闭源;真正持久的开源需要另一套激励系统。他估计,不同实验室的数据支出可能约占算力支出的 1% 至 10%,并认为这一比例应该更高。擅长使用人类数据的研究人员推进更快,而缺乏这类能力的人可能只能依靠变通方案,拖慢进度。

Lukas Biewald

Today, I'm talking with Edwin Chen, who's the CEO of Surge. I was really looking forward to talking to Edwin for a long time for a number of reasons. One thing is that Edwin has built an incredibly valuable business with no venture backing in a really short amount of time. He started this data-collection business in 2020, and in 2024 he crossed $1 billion in revenue, which is spectacular success.

But that's probably not even the top reason I wanted to talk to him. The human data-collection business is this really important part of building AI systems that's not talked about enough. I got into that business in 2026 2027. I started a company called CrowdFlower that did this data collection in the early days, and around the time I sold CrowdFlower is around when Edwin started Surge.

The data-collection business has changed a lot over the years, but one constant has been that it's a huge spend for people building high-quality models. Edwin has these front-row seats into what most of the foundation labs and foundation-model builders are doing, how they're thinking, and how they're building their models. So he has a lot of insights. He doesn't do a lot of podcasts, so we were really lucky to get him.

So, you're the first guest we've had in the human data-generation space. It's a space I was in for a long time, so I'm really curious about this. Could you first of all start by telling the story of Surge—what you were thinking when you started it and how it's gone?

Edwin Chen

Yeah, I can give the founding story. Basically, I used to be an ML engineer at a bunch of these big companies, and the problem I kept running into was that we just kept facing all of these issues getting the data that we needed to train our models.

For example, I used to work on our search and ad systems at Twitter, and one of the first things I wanted to do was build a sentiment classifier. Sentiment analysis is a super-simple problem, and all we needed was 10,000 tweets to train our models. But our human-data system at the time was literally 2 people we'd hired off Craigslist, working 9 to 5. We had to wait a month to get started, and then we had to wait another month for them to label the tweets in a spreadsheet because our tools were just terrible.

When we finally got the data back, we saw that it was complete junk. They didn't understand slang—for example, a phrase like “She's such a bad...” They were labeling it as just negative. They didn't understand hashtags or all these other aspects of tweets. What ended up happening was that I spent a week myself labeling all these tweets because that was actually so much faster and better.

I think one of the things that we often said was that these were really simple things. At the end of the day, sentiment analysis isn't all that complicated. At the same time, we had this bigger problem to solve around how we wanted to optimize our ML systems for the right objectives.

When I first started working on Twitter, this was in the old days when it was purely a chronological timeline. One of the things we wanted to do was make it easier for users to discover and engage with tweets they cared about. The question was, how do we train our recommendation algorithms?

The obvious choice was clicks and retweets at the time. You just train your algorithms to produce as many clicks and retweets as possible. But we tried doing this, and it turned out to be an incredibly negative feedback loop. Once you optimize for clicks, you get the most clickbaity content in the world rising to the top. You get lots of racy content, lots of bikinis, and lots of listicles about “10 horrifying skin diseases,” and so on.

1. The problem with early data labeling systems

We wanted to train our models on deeper principles instead. We'd ask human raters to label tweets and recommendations according to our product principles. But if we just couldn't even get simple sentiment analysis right, we definitely couldn't get this more complex data at the quality and scale we needed.

Eventually, it became a problem that happened over and over again at Google and Facebook, too. Eventually, I realized it was something I needed to go out and build myself.

Lukas Biewald

What year did you start Surge?

Edwin Chen

We started Surge in 2020, right in the middle of the pandemic.

Lukas Biewald

I feel like in 2020 there were some established data-labeling and data-generation companies. Did you have a particular take on the space?

Edwin Chen

Our take on the space was that all of these solutions out there were basically focused on this idea of very commodity labeling, very low-skill work. The example I often give is the problem of drawing a bounding box around a car. You and I can all draw a bounding box around a car. Almost a 3-year-old can draw a bounding box around a car. The bounding box that I draw isn't going to be any different from the bounding box that Terence Tao draws or that Einstein would draw. There's a very low ceiling on the complexity of data that's required.

In contrast, if you think about all the things that we want to do today, we want models that can write poems, and we want models that can solve relativistic equations. There's almost an unlimited amount of intelligence that we want to feed our models. All the other solutions at the time were designed for a very low-skill, commodity style of labor. They weren't focused on quality at all. Instead, they were focused on scale.

Lukas Biewald

I see. How did you source the people in the beginning?

Edwin Chen

Literally, the first few people who were on our platform had just been working with me. It's kind of funny, but I've been working on this problem for a long time, so I already had a network of people who were very interested in doing this kind of work. When they heard that I started Surge, a lot of them joined. A lot of it was also me in the beginning, but it was also all these labelers that I'd accumulated throughout the years.

Lukas Biewald

That's cool. Who were the early customers?

2. Search ranking, clickbait, and product principles

Edwin Chen

The early customers were a lot of tech companies. We had this idea that we wanted to really focus on engineers and research scientists who really understood the quality of data. These were people I'd been working with in the space for a while. There were a lot of friends and contacts I had at all these companies who had been dying for this kind of higher-quality, next-generation solution that could do far more advanced tasks than what were possible at the time.

There were a lot of companies like Airbnb and Twitter, as well as a lot of startups in the search and algorithm space.

Lukas Biewald

At that time—in 2020—that was right around when I was leaving the space. I remember that what was really taking off then was autonomy and robotics, a lot of vision applications. It sounds like you were more focused on text. Is that fair?

Edwin Chen

We were always focused on language and what I call behavior from the beginning. I wouldn't say it was necessarily text per se. There are a lot of complex problems in the image space, too, especially nowadays.

What we didn't want to focus on was very simple image-labeling tasks or very simple bounding-box-style tasks, where I don't really think that there's any intelligence needed to build such solutions. We always focused on this higher-complexity, higher-scale, higher-sophistication space.

Lukas Biewald

I would say that image labeling, when you actually dig into it, is more complicated, for the record. But I take your point.

It seemed like you really built Surge under the radar. Was that an intentional decision? I think you might have said recently that you're over $1 billion in yearly revenue. Is that right? Can we say that?

Edwin Chen

Yeah, we're over $1 billion last year, and we hit our $1 billion ARR number a while ago.

Lukas Biewald

That's an incredible achievement in just 5 years, and you did it with no outside funding. Is that right?

Edwin Chen

Yeah.

Lukas Biewald

That must be a historic level of growth without funding, I think. Did you intentionally avoid the VC path?

Edwin Chen

One of the things that was really important for us was that we wanted customers to be buying from us because they really believed in having high-quality data, not because they saw us mentioned in some tech article. We wanted partners who had the same vision we did and who could go out and show the world how important data actually was.

Again, if you contrast that with the kinds of data people were looking for 5 or 10 years ago, people really weren't focused on quality. They treated data as this kind of commodity. The researchers themselves would barely look at the data. They'd just let their vendor-management teams handle the entire process because they wanted to outsource it.

And so, we just had this idea that we really wanted to focus on people who believed in quality and were buying us for that reason.

Lukas Biewald

But to get to a billion dollars plus in revenue, you've obviously had to scale your system. That goes beyond hiring a few people that you've known for a long time. Can you talk about how you've scaled your processes as your scope has expanded? Do you think it's more important to scale the technology that's enabling this, internal human processes, hiring, or the management of what's going on? As much as you could say about how this works, I'd love to hear it.

Edwin Chen

Yep. Yeah. I think the thing that people underestimate in this space is how much technology you actually want to build. People tend to think that humans are smart, and so if you just throw 10,000 or 100,000 humans at a problem, that will solve it. It's a little crazy to me.

3. Why Surge focused on high-skill, high-quality labeling

Sometimes I'll interview candidates from some of the competitors in our space, and when they describe to me what they're building and how they operate, it's incredibly manual. They're just body shops at the end of the day, and they literally have no technology. If you ask them, “Could you tell me the quality of this worker? Could you tell me how good this worker is at this particular task? Could you show me a dashboard with how your quality is improving over time, what A/B tests you run, and what algorithms you're building to improve it?” they literally can't.

All they're doing, a lot of the time, is dumping data into spreadsheets. Their employees and engineers are either creating the data themselves or reviewing it themselves. There's almost no technology going on.

I think what you have to realize about this space is that quality control is incredibly difficult. If you want to get the highest quality out there, you need to find or create sophisticated algorithms to detect the highest-quality data that you can. It's incredibly complicated because, in this space with LLMs today, you really want LLMs to be good at every task in the world. It's not just a single domain; you want to make sure that they're really good at poetry and, at the other extreme, physics.

4. Scientific collaboration and frontier research data

How do you find the highest-quality data in order to train the models, and then how do you also remove the worst? You can't just throw warm bodies at it. You really need to be able to build a lot of technology to manage it.

Lukas Biewald

Well, I'm coming from a different era, I think, of labeling, where a lot of labels were simple—yes or no, or simple tasks where it's really clear what's right and what's wrong. But I think you're doing much more complicated tasks, and as you're saying, even getting 2 people to agree on high quality versus low quality gets more and more difficult as the task gets more complicated.

Can you talk a little bit about what technology you actually have to build to manage quality for a generic, complicated expert task involving language?

Edwin Chen

Let me start by contrasting with a very old-school take on data and data quality, and how we think about the problem differently. Let me give you an example. Let's say you wanted to train a model to write an 8-line poem about the moon. The way most companies think about it is, “Okay, let's just hire a bunch of people from Craigslist or through some recruiting agency. Let's ask them to write poems.”

The way to think about quality, again going back to the image-annotation days, is: Is this a poem? Is it 8 lines? Does it contain the word “moon”? They just check all these boxes, in the same way you might check, “This is a cat. This is a dog.” They're just checking boxes, and then you're saying, “Sure, this is a great poem because it follows all the instructions. It follows all these checkboxes.”

What happens is that you get terrible poems that feel like they're written by kids in high school. A kid in high school can write an 8-line poem about the moon, but is it a great poem? Is it an evocative poem? Is it the type of poem that a Nobel Prize winner would write? No.

5. From Craigslist workers to a billion-dollar business

Some other companies might say, “Okay, sure, these people on Craigslist don't have any poetry experience. What I'm going to do instead is hire a bunch of people with PhDs in English literature.” But what they don't realize is that this is also terrible. A lot of PhDs, again going back to what I was saying about even people with MIT computer science degrees, aren't good writers or poets. Think of people like Hemingway or Emily Dickinson: They definitely didn't have a PhD. I don't think they even completed college.

We think about quality completely differently. What we want isn't poetry that checks some boxes and uses complicated language. We want the type of poetry that Nobel Prize winners would write.

I think you need a mindset shift. One of the things we think about is that there are certain people who are trained up in this very objective domain of computer vision, a domain that lacks all of these nuances and all this inherent subjectivity. What we want to do instead is recognize that poetry is actually really subjective and rich.

Maybe one poem is a haiku about moonlight on water, another poem is something with internal rhyme and meter, and another one focuses on the emotions behind the moon rising at night. You want to capture that there are 1,000 ways to write a poem about the moon, in a way that there aren't 1,000 ways to draw a bounding box on a car. There aren't 1,000 ways to label something as a cat or a dog.

There isn't a single correct way to write this poem. Each different way that you write it, or each different preference that you have for different types of poetry, gives you different insights into language, the mind, imagery, and human expression.

I talk about poetry a lot, but it's not just poetry. If you think about math as well, there are 1,000 ways to prove the Pythagorean theorem. Each way of proving the Pythagorean theorem is based on different insights into the mathematical reality of the universe.

One thing that happens is that when you think about quality the wrong way, you get commodity data that optimizes for things like inter-rater agreement. Going back to the computer-vision world, if all you're doing is labeling images of cats and dogs, sure, you want high inter-rater agreement. You want people to agree that this is a cat, and you want people to agree that this is a dog.

One of the very easy ways to quality-control in this old-school world is just by seeing whether there's agreement with a majority. But in this new generative-AI world, there's no way that you can ask people to write 1,000 poems and then take a majority vote and get the best poem.

What actually happens when you optimize for higher inter-rater agreement is that you get the lowest common denominator of things that people want, which turns out to be trashy, low-quality, and unengaging a lot of the time. You need to think about quality in a different way in order for your data to really embrace human intelligence and creativity.

6. Scaling without funding and avoiding Silicon Valley status games

Lukas Biewald

I mean, one of the things that we found when I was working in this space is that maybe the biggest issue was actually eliciting from a customer what they really wanted. Obviously, we're doing simpler examples, but you think about putting a bounding box around a car. It seems so simple to an ML researcher, especially one who hasn't looked at a lot of specific data examples, but it is really hard.

You're one of the rare ML researchers who really wanted to look at data. I was, too, and that's why I got into the space, right? But putting a bounding box around a car seems simple until you start looking at real-world data. What if the car is occluded? Do you put the box around where you think the car is? What if the car is in a reflection of a car in a mirror? Do you still want the bounding box around that car?

What if it's a billboard with a picture of a car and it's not a real car? Do you want that? Even in the simplest cases, as soon as you start to look at real-world data, it gets way more complicated.

I remember working with customers, and there would actually be a long process of eliciting from the customer what they wanted. Some customers would write these giant documents trying to enumerate every case and exactly what they wanted, but I think those, too, would end up being hard to reason about exactly what you want.

If you followed those instructions to the letter, you'd end up in some ridiculous cases, which isn't actually what the person wanted in the first place. I'm kind of curious: how do you think about that problem?

Edwin Chen

What we always try to do instead is understand the goal or principle that the researcher has. When we understand the goal or principle, or how they're going to use the data, we can almost put ourselves in their shoes. Instead of asking them, “Okay, what happens when this car is occluded? What happens when this car is in a reflection?” we think, “Okay, deriving from first principles, given what they've told us, what do we think the researcher wanted?”

Maybe if you know that the data is going to be used for some LLM model, then you'd know that this makes sense and this doesn't. We can have that higher-level goal instead of a mechanical list of instructions. That's always what we aim for.

It is hard because, at the end of the day, sometimes researchers don't know either. But this is where, as a company, we often try to have a strong opinion on what the best type of data is, as opposed to almost blindly accepting whatever people tell us. Oftentimes, we actually do like to get into a feisty debate about what type of data would be most useful to build them.

Lukas Biewald

So then, do you hire a lot of ML researchers? Are those the people that you want working directly with the customer in order to—

Edwin Chen

Yeah. We have a lot of researchers. One of the things that we often think about is that we almost consider ourselves a research company, just that instead of researching algorithms in a way that other frontier labs might, we're more about researching the data.

Lukas Biewald

Totally. And has it been—I really admire how you've avoided status games and flown under the radar, but I would think that, in doing that, it might make it harder to hire. Do you have a different strategy for hiring, or do you think that maybe it doesn't matter and you find the people who really connect with your mission?

Edwin Chen

One of the things that we often think about is that we don't want people who are just joining us in order to notch another brand on their resume. Sure, there are other people out there who do that, and we're missing out on those candidates. I think they also tend to be people who want to build large teams. They tend to be people who want to build empires, and it almost turns into hiring for the sake of hiring.

“Why do you want to hire this person?” “I needed to hire this person because the other person doing this job is spending all their time interviewing.” “Why is that person spending all their time interviewing?” “Well, it's because somebody told them that they needed to hire an internal tooling team.” “Why does that person need to hire an internal tooling team?” “Well, it's in order to make the engineers more productive.” “Why aren't they productive enough?” “Because they're spending all their time in meetings.” “Why are they spending all their time in meetings?” It's because we hired all these people, and you need to communicate with them.

7. Why most human data platforms lack real tech

I think there are a lot of benefits if you start with people who really believe in your mission. Then, at least for us, we've been able to stay smaller, with a much smaller team.

Lukas Biewald

How big is your team?

Edwin Chen

We're a little over 100 people.

Lukas Biewald

Wow, that's incredible revenue per employee. Congratulations. You've gone through this really fast change, I think, from being essentially an ML researcher to running a very significant, large company. Where have you felt stretched the most? What's been challenging?

Edwin Chen

Certainly, the thing that I found most challenging is sales. Probably any researcher will tell you that just the concept of having to go out and hawk your product is a little foreign to me.

I think we're lucky, in a sense, that our product, at the end of the day, is for researchers. Having that research mindset, what we're trying to build is almost like that Disney quote: “We don't make movies to make money; we make money in order to make movies.” In a similar way, what we're doing is not trying to generate revenue. We're trying to generate data that will help AGI.

Sometimes when customers—or when new companies—come to us and ask if they can work with us, if their goal is unaligned with AGI, we actually just say no to them. We don't want the revenue; we want to focus on the AGI companies, in a sense.

Again, going back to not raising money, the fact that we don't have a board, or an external board, and the fact that we don't have VCs who are just dying to make as much money as possible, I think that gives us a sort of freedom that allows us to focus on the most important problems. That's allowed us to maintain our research focus.

Lukas Biewald

Totally. So you only work with companies focused on building AGI?

Edwin Chen

For example, if a company came to us and said, “Yeah, we just want to train a—let's say I'm a newspaper and I just want to train a category classifier,” we'd just say no.

Lukas Biewald

What if I want to make an AI video generator? Would that be in your realm?

Edwin Chen

Oh, yeah. We do that in that sense because building such video generators is part of building AGI. So, yeah, in that sense.

Lukas Biewald

I imagine that 5 years ago the data being collected was pretty different from the data collected now. I would think that some of the tasks you'd be doing 5 years ago would be easily automated by LLMs today. It's just been such an astonishing pace of improvement. Can you talk about how the types of tasks have changed over the last few years?

Edwin Chen

Some of the types of work that we do are very, very different. When we first started, a lot of our work was in tasks like search evaluation or content moderation, whereas today it's almost purely LLM work. That's one big difference.

8. Detecting cheaters, liars, and low-quality labelers

Even within LLMs, there's been this obvious trend toward higher complexity, higher sophistication, and higher expertise. For example, there's been a big increase in multimodal complexity. A few years ago, it was all text data, just conversational text assistance, but now we do a lot of work with images, audio, and video. I think the interest is that you actually want the models to understand all of these modalities simultaneously.

One of the things you might want to do is say, “Okay, I'm taking a video of something on my phone, and now I'm asking the model to create a program based on the video on my phone that simulates this in real life.” So there's been a lot of multimodal increase in complexity.

There's also been a big expansion in languages. At first, people were naturally focused on English-only work, but we actually work in over 50 languages now. What's also interesting is that it's very hyper-specialized: we support coding in Argentinian Spanish, and we support legal and financial expertise in Bolivia.

Even today, I think a lot of the models are just not that good—surprisingly so. They're not that good at the different nuances of different languages, dialects, or cultures. I think there's still a lot more progress to be made there.

Probably the biggest shift is just the depth of expertise that a lot of the work requires. You see the models winning IMO gold medals now, and you see them doing all these incredibly advanced tasks, so you actually really need serious thinking power behind them.

Even today, some of the tasks that we do involve spending days or even weeks solving these really interesting problems. It's a very far cry from tasks 5 or 10 years ago, where you might spend 5 seconds labeling a task.

Lukas Biewald

Do you actually have people who can solve Olympiad-level math problems—creating those problems and then solving them?

Edwin Chen

Yeah.

Lukas Biewald

You know, it's funny—this is kind of an aside—but I've been surprised by the scores the latest models are getting on these IMO problems. When I put in more fun brain-teaser problems that my friends pass around, they often can't do them. Do you have a sense of why there's that disconnect?

Edwin Chen

I think a big problem with a lot of the frontier models today is that they've basically been benchmark-hacked. You have all these benchmarks out there, and a lot of them just aren't very good. They're overly academic or overly synthetic.

A lot of these benchmarks have a single objective answer. Going back to my point earlier, the models have been narrowly constrained to be good at these very narrow, objective problems. But a math problem in the real world, or a problem that you would ask as a research mathematician, isn't going to be a closed-ended problem; it's going to be an open-ended exploration.

I think a lot of it stems from this kind of benchmark hacking that's going on.

Lukas Biewald

So if you're going to build a benchmark today to compare frontier models, what kinds of things would it include?

Edwin Chen

What we always say is that the gold standard for evaluating models really is human evaluation, where you just can't fully automate it. People try to build these leaderboards where they take automatic verifiers, but automatic verifiers still only work well in very objective domains.

9. Why inter-annotator agreement is a flawed metric

They've tried building benchmarks like LMSYS, which I think is an absolutely terrible leaderboard that has basically set the industry back by at least 1 year. I can go into more detail on that, but I think a lot of the benchmarks out there are flawed either because they're built using low-quality data or because they're overly synthetic, overly academic, and overly objective in a way that the real world isn't.

Lukas Biewald

Okay. Well, I mean, now I want to hear this. Why do you think LMSYS—you're saying that benchmark is not just bad, but it set the industry back? Can you tell me more about that?

Edwin Chen

Yeah. Basically, LMSYS is, if you don't know what it is, this popular leaderboard of language models. What happens is that people go onto it—literally anybody around the world. What we often hear is that it's literally high schoolers and middle schoolers who can't access models any other way, and they're going onto Arena as their only mechanism.

They go onto Chatbot Arena, enter a prompt, and then see 2 model responses. They vote on which one is better. But if you think about it, they're not taking the time to read or evaluate these model responses at all. They're literally just looking at the responses for 2 seconds and picking whichever one strikes their fancy.

The models could have made everything up. They could have completely hallucinated everything. They could have not followed the instructions at all, and people will just vote on it because, in order to progress, they need to vote on something. Then they see, “Oh, yeah, this model has emojis and a lot of bold formatting, so it just looks really impressive.”

One of the things that we've learned is that the easiest way to improve in this arena is simply to make your model responses a lot longer and double the number of emojis that you have. If you actually think about some of the models that have been released in the past year—think about Llama, or at least the version of Llama 4 that was optimized for LMSYS—it exactly matched all of these patterns.

One of the phenomena that we often hear about from researchers is that they'll tell us, “I'm only going to get promoted if my VP has told me that we advance our model by 10 points on the leaderboard.” They'll look at the leaderboard data themselves and see that a lot of the responses preferred by LMSYS—in its datasets—are literally the responses from the model that followed instructions the worst, or the model that hallucinated everything. Those are the responses that are getting preferred simply because they have a lot of formatting and emojis.

What these researchers tell us is, “I want to work on improving the fundamental capabilities of my model. I want to improve it in coding. I want to work on fixing its hallucinations.” But instead, if the only way I'm going to get promoted is by optimizing for this leaderboard, and the easiest way to optimize for this leaderboard is to make my model hallucinate so that it generates these crazy answers that, even though they're completely wrong, are very compelling to amateur raters who are only spending 2 seconds, then that's what they're going to do.

We've often seen models where, if you look at some of the models that are in the top spots on the leaderboard today and compare them to how they were performing 6 months ago, you actually see that they're worse in many ways.

10. What makes a great poem? Not checkboxes

Lukas Biewald

Do you think so? Another thing that strikes me is that we have these different models made by totally different organizations. You have xAI, run by Elon and making Grok, and you look at Anthropic and OpenAI. These clearly have different cultures, and we've had different people from these organizations on the podcast, yet the models seem to have this surprisingly consistent tone to me.

They feel a little annoyingly positive, a little Boy Scout-y: “Yes, thank you. I will answer your question.” I imagine that must somehow be trained in at some point in the process. I wonder if it's in the preprocessing—is that just what high-quality data on the web looks like—or is it intentionally inserted later in the process? Is the annotation that you're doing somehow contributing to that consistent tone? I feel like we'll look back on it as almost this LLM tone of 2025.

Edwin Chen

Yep. I think it's a combination of everything. For example, we don't explicitly teach the models to follow proper grammar. We generally don't teach the models to use em dashes, but it's kind of just there in pretraining, where these behaviors are baked in.

I think in the future the models will become more and more differentiated, where you'll be able to tell that they have a noticeable difference in style. Even today, maybe because I've looked at this data enough, I can tell from reading the model responses themselves which model is which. Some of them will use certain prefaces, certain words, or certain stylistic patterns, and I can generally tell.

Lukas Biewald

Interesting. Can you think of any tells? Is it more subconscious, or could you cite some specific things that let you know it's a particular model?

Edwin Chen

Some of them just use certain phrases, like “absolutely,” more often than others. Some of them use emojis in a way that others wouldn't.

Lukas Biewald

Llama uses so many emojis. I totally agree with that.

Edwin Chen

Some of them do this thing where they repeat 3 adjectives in a row. Some of them use a lot more Markdown than others.

Lukas Biewald

When you look at the way these different organizations collect data, are there striking differences, or does it feel like everyone's chasing the same types of data?

Edwin Chen

I think there actually are very striking differences between the different companies. Almost every company has its own philosophy on the right way to both train and evaluate its models. It's been surprising how different they are.

Lukas Biewald

Do you have a point of view on how you think it should be done? If you were running an AGI company, what would you do? Is there something you would do differently from what you are asked to do?

Edwin Chen

We had this very strong view that our RLHF data was a lot more effective than SFT data. Maybe you should just explain the difference here for—

Lukas Biewald

Yep.

Edwin Chen

—the general audience too. So, the way SFT data works, just to give an example, would be—

Lukas Biewald

And this is fine-tuning data—supervised fine-tuning.

Edwin Chen

Yes. So SFT stands for supervised fine-tuning. The way it works is, let's suppose that I wanted to train my model to become better at poetry. You, as an annotator or as a reader, would write a prompt. The prompt might be, “Write a sonnet about cheese,” and then you would literally write a sonnet about cheese. That would be a demonstration to the model. That's essentially supervised fine-tuning.

In contrast, in RLHF—reinforcement learning from human feedback—the way it works is that you write a prompt. Again, you might take the same prompt, “Write a sonnet about cheese,” and then you ask the model to generate 2 responses. These might be 2 different models, or 2 different model checkpoints. Model A would produce an answer, and model B would produce an answer. Then you would rate which one is better: “Model A is better because it had emotion behind it, because it was better written, or because it actually followed the format of a sonnet.”

You basically teach the model, “A is better than B.” The benefit is that there are many benefits, but just to enumerate some of them: first, it's a lot more efficient. Even if you're a Nobel Prize laureate, it takes you a lot of time to write a poem about cheese.

11. Measuring subjective quality rigorously

Second, it helps teach the model all these latent preferences. It helps teach the model what's good, but also what's bad.

Basically, we had this very strong view that RLHF was a lot more effective. SFT is important sometimes. One of the things we often tell our customers is that you often want to use a little bit of SFT data to more efficiently bootstrap your models into a phase where RLHF is useful. But we had this very strong view that RLHF was a lot more effective, and so we steered our customers in that direction.

If you actually read the Llama 2 paper, you'll see that one of the discoveries they describe is that, at the beginning, all the researchers thought that SFT data would be more effective as well, but then they ran a bunch of experiments and found that RLHF was just so much more effective. So, yeah, that was one of our beliefs in the past.

Lukas Biewald

When you look into the future and roll things forward a few years, what kinds of data do you think you're collecting now that you won't be collecting? And what new kinds of data do you think you'll be collecting?

Edwin Chen

I actually don't think that we're going to stop collecting any of the data that we're still collecting right now. One of the things that we've often found is that the proportions may change. In the same way that, as you become a more sophisticated adult, you don't need to be trained as much on arithmetic as you did when you were a little kid, what we often find is that if you don't persist with at least some amount of this data, models often just regress in very surprising ways.

I won't name names here, but some of the frontier models just make the most bizarre mistakes these days. I think it's kind of wild, and that's just because they're being fine-tuned in very certain ways. I'll say that some of the ways we expect the data trends to continue are in ways that they haven't before. I'll just list them out.

One is even higher-expertise problems than we have today. Right now, there's a lot of work being done in what I would call graduate-level STEM work, but if you actually want the models to make new scientific discoveries, you're going to need to progress beyond what even the average PhD can do. So, I think there'll be STEM work like that in the future. That's one of them.

Lukas Biewald

But wait, can I ask? What would super-STEM work be? Would that literally be, “I'm going to go out and discover some new principle of chemistry” just to train a model? How would that even manifest?

Edwin Chen

I think it actually is exactly that. Imagine that you are a Stanford professor and you're literally working on your latest frontier research. What you want to be doing is collaborating with an AI to generate new hypotheses and help you test your experiments.

12. What types of data are becoming more important

It's this idea of a scientific collaborator, in the same way that we have a coding collaborator right now in the form of Claude Code and whatnot. I think there will be things like that that are increasingly important.

Lukas Biewald

But how would you collect the data? Would you get a Stanford professor to just score responses? I guess scoring responses would be easier than just generating responses.

Edwin Chen

It wouldn't necessarily be just scoring responses. There's a lot of variety in the different types of data that we collect nowadays. It's literally a Stanford professor or some other person who's capable of understanding frontier physics, collaborating with the models and teaching them the right signals.

That's one domain. Another domain is just longer and longer-horizon tasks. Today, the models can definitely perform tasks that a human could have done in 5, 10, 15, or 30 minutes, or an hour. I think some time horizons will just get very, very long. What are the types of things that you want the models to do that would take a couple of days or even weeks? That's another domain.

A third domain, almost going back to your point about the language that these models exhibit today, is how we actually make the models really, really good at long-form creative writing. Even today, it's kind of interesting: I don't like any of the poetry or short stories that the models can write, even though they have perfect prose. They're just not creative enough. They've almost collapsed to this very, very vanilla, generic dimension.

I think there will be more work there. Up until now, there have maybe been higher-economic-value tasks that people want the models to do instead. I think it's actually important for models to exhibit this kind of creativity in a non-generic way.

A fourth one would just be more and more actions in the real world. Agents have obviously exploded in popularity over the last few months, but they're still at a very, very nascent level. The concept of AI models that can plan and reflect on their actions, try new ideas, and actually affect things in the world will be an increasingly important concept.

Lukas Biewald

I guess one of the things that we might not have expected models to do as a priority is put a lot of their reasoning into code. We've seen a trend where models use code to do the reasoning more and more. You can see it in the chat window. Do you think that's an artifact of the limitations of models today, or do you expect that trend to continue?

Edwin Chen

I actually expect it to continue. I do think that the code they're writing is very important for certain types of verifiability that you might want the model to exhibit. In the same way that, if I could write a lot of code for my day-to-day life, I would, I think it's a very, very important capability for models.

Lukas Biewald

Do you have a sense—I guess I have this intuition—that the amount of code used to train these models is increasing as a fraction of the overall data? Do you think that's accurate? Are you generating more and more code for your customers than other types of data?

Edwin Chen

Yeah. We are.

Lukas Biewald

I guess that's a lead into another question I have: Do you think the trend toward reasoning models is changing the types of data that you're collecting?

Edwin Chen

Definitely. Reasoning models have led to a bunch of new types of collection methods. I think probably the biggest one is this concept of creating RL environments from scratch, where essentially a lot of what we're doing is building these video game-like universes with interesting tools and datasets that the models need to solve.

You can imagine a universe consisting of a simulated AI startup. In this environment, or universe, you basically have a bunch of Gmail messages, Slack messages, Jira tickets, GitHub pull requests, codebases, and so on. You want this universe to be as rich and complex, and as high-fidelity to challenging tasks in the real world, as possible.

For example, you could imagine that in this universe, suddenly AWS goes down and suddenly Slack goes down. What do you do? What does the agent do? If I can't use Slack, I need to make sure that it knows how to handle that and figure out how to solve the problem. There are a lot of interesting things that we're building into these RL environments.

Lukas Biewald

How do you handle the subjectivity of these models? I have a 5-year-old, and she really likes to generate stories, so I've gone really deep trying to get these models to write interesting stories. I agree with you: I don't think they really write interesting stories. A big part of it is that I think they're trying to average the preferences of a lot of different people.

My 5-year-old's taste is really different from my taste in stories, and we're actually asking the model together. Maybe it's not even possible for it to make both of us happy with the story, let alone the average person. It seems like that might require a different type of labeling if you actually want to get interesting art out of it, because good art probably shouldn't get a thumbs-up from every single person who looks at it.

Edwin Chen

Yeah. I think there's this concept of personalization that is still very, very unexplored with the models. There's surface-level personalization, where the model knows that you live in Ohio versus New York City. But there's also a lot of unexplained latent preferences that you have that are kind of hard to articulate.

How do you get the model to somehow learn those from the data for every single person and then apply them to the different responses that it generates? I think that's a really, really interesting concept that still hasn't been fully explored. For example, if you are a mathematician, here are the restaurants that you should try in Paris. Models sometimes just make these weird little personalization generalizations.

I think there's still a lot of work that remains to be done with personalization.

Lukas Biewald

And is that, do you think, because it's hard to collect good personalization data? I'd imagine that, to put yourself in the shoes of someone you're not, the data is never going to be as accurate.

Edwin Chen

I think that’s part of it, but a big part of it is that all the frontier labs only have so many things they can focus on, and they just haven’t quite doubled down on personalization yet.

13. Multimodal data, Argentinian coding, and hyper-specificity

Lukas Biewald

Okay. So let’s talk about synthetic data. I think everybody’s dream is always to be able to automatically generate useful data. We’ve seen a lot of people using LLMs as a judge for many applications, which is a kind of synthetic data. How do you do synthetic-data generation, and how do you think about that?

Edwin Chen

Yeah. I actually do think synthetic data is really useful in some places, but I think a lot of people overestimate what synthetic data can do. I can give a couple of examples. Right now, there are a bunch of models that have been trained really heavily on synthetic data. But similar to what you and I were saying before, that’s partly why they’re very good at these very academic, homework-style, benchmark-style problems, and they’re actually really terrible at real-world use cases. Synthetic data has made models good at synthetic problems, but not real ones—that kind of phenomenon.

I remember maybe a year or so ago, we were running human evals for one of the researchers we work with, and our human evals showed that their models had suddenly tanked. When we dug into it and talked to the researcher, it turned out that they had just trained their model on 10 million or 20 million synthetic math problems. The math problems were in a very, very narrow domain of math, and what they hadn’t realized was how much that was making the model worse at basically every other type of task. There are only so many very closed-ended SAT-style math problems that you want the model to solve. There’s only so much you care about that, and so the model actually just became worse in basically every other domain.

Another thing we often hear from companies is that they tell us, “I spent the past year training my models on synthetic data,” but they’ve only now realized all the problems that caused. They then spend multiple months throwing a lot of it out. A lot of them will tell us they’ve thrown out 10 or 20 million pieces of synthetic data because they found that even just 1,000 pieces of really high-quality human data are more useful. The human data is, I think, both more diverse and more creative, as opposed to getting 10 million pieces of the same thing over and over again. What ends up happening in practice is that companies will try using synthetic data for certain problems for 6 months, and then a lot of the work that we do ends up being cleaning up synthetic data.

14. What's wrong with LMSYS and benchmark hacking

Lukas Biewald

You have front-row seats to what most of the labs are doing. Do you have a point of view on whether we’re at a moment where performance is stuck and kind of needs new methods to get to the next level, or whether the current strategies of pre-training and then reinforcement learning to learn reasoning will get us to a much better set of models?

Edwin Chen

I definitely don’t think we’re stuck at all. I think there are a lot of new methods that are just appearing and a lot of new data that hasn’t been collected yet. From some of the early experiments that we’ve run ourselves, I think we’re going to see some massive progress in some new domains very soon.

Lukas Biewald

Can you describe what some of the new methods are at a high level?

Edwin Chen

Even just some of the methods I was describing in terms of RL environments and some of these new RL methods, I think, are still kind of new to a lot of researchers in industry. It’s basically this concept of exposing the models to all these new environments they haven’t seen before.

Lukas Biewald

What would be an example of a new environment they haven’t seen before?

Edwin Chen

Even just the example I mentioned earlier, where you have a simulated AI startup and then suddenly the model’s environment loses access to Slack and AWS. How does the agent continue operating in this environment, and how does it solve the problem? I think it’s something that just doesn’t appear, in some sense, in any preexisting data or any data that we’ve generated before.

It’s an entirely new concept where the model needs to go through all these different actions. It needs to reflect on what it’s done, and it needs to find new ways of solving a problem. Especially when you expose it to very messy types of data that it may try to retrieve—data that, again, it may not have encountered before—the models can fail in these very unique ways, but I think in ways that they can be taught to progress through.

Lukas Biewald

What do you think about the ARC benchmark? That’s an interesting one: it has deceptively simple problems that I think you or I would have no trouble with. They’re certainly easier than math Olympiad problems and are more about visual pattern recognition. Do you think those are likely to get solved in the next 1 or 2 years?

Edwin Chen

I’ll admit, even the ARC benchmark is very surprising to me as well. If you asked me, without ever having seen the ARC benchmark, whether models could solve it, I’d be like, “Yeah, absolutely.” So there’s something about them that I still don’t quite understand—why models can’t solve them.

We’re actually doing a lot of work to generate ARC-style problems, so it’ll be interesting to see how much that helps models improve. Right now, I don’t have a good understanding of why models can’t solve them.

15. Personalization and taste in model behavior

Lukas Biewald

Do you have an opinion on whether open-source or closed-source models are going to win in the long term?

Edwin Chen

My guess is that, at least with the current state of things, closed-source models will continue winning. Part of that is because LLMs are just so valuable that, if you ever try to build open-source models, the way incentives currently work is that eventually you’re going to be forced to make them closed-source. If you want to build truly open-source models that are really, really good, we almost need a different kind of incentive structure to make sure that happens or remains in place. Otherwise, if you look at the history of other types of open-source models, they’ve gotten more closed over time.

Lukas Biewald

What do you mean? What are you referring to?

Edwin Chen

If you even think about Meta, which is thinking about making Llama closed-source, models are just so expensive to train, and people want to fully capture the value. If you ever build a truly, truly good open-source model, I don’t think it’ll remain open-source for very long, again, unless you can change the incentive structure in some way that I haven’t figured out yet.

Lukas Biewald

What do you think is the ratio of spending on data to spending on compute in the training of a large model, and how do you expect that to change over time?

Edwin Chen

I definitely think it should be a lot higher. Sometimes people, I would say, almost skimp on gathering data.

Lukas Biewald

What do you think it is today, roughly?

Edwin Chen

I actually don’t have a good sense myself. I think it varies widely depending on the labs, but it can be anywhere from a couple of percentage points, or 1%, to 10%.

One thing we’ve often heard from researchers is that some of them are really, really good at using human data: they know how to come to us to gather it, and they know how to apply it in their own work. What they often tell us is that some of their counterparts don’t know how to use human data, so they just move a lot slower. They may try somewhat wild ways to get around the fact that they don’t have any human data, and it just ends up being this complex slowdown for them. I hope it’ll get higher in the future. I think there are a lot of ways in which people still underestimate the use of data.

Lukas Biewald

Awesome. Well, I appreciate your time, and congrats on building a fantastic business.

Edwin Chen

Yeah, it was a great chat.

Lukas Biewald

Great chat. Thanks.

Edwin Chen

Thanks so much. Bye.

为 AGI 提供底层数据的创业公司 — 文字稿与摘要 | BidClub