Lukas Biewald
Today, I'm talking with Edwin Chen, who's the CEO of Surge. I was really looking forward to talking to Edwin for a long time for a number of reasons. One thing is that Edwin has built an incredibly valuable business with no venture backing in a really short amount of time. He started this data-collection business in 2020, and in 2024 he crossed $1 billion in revenue, which is spectacular success.
But that's probably not even the top reason I wanted to talk to him. The human data-collection business is this really important part of building AI systems that's not talked about enough. I got into that business in 2026 2027. I started a company called CrowdFlower that did this data collection in the early days, and around the time I sold CrowdFlower is around when Edwin started Surge.
The data-collection business has changed a lot over the years, but one constant has been that it's a huge spend for people building high-quality models. Edwin has these front-row seats into what most of the foundation labs and foundation-model builders are doing, how they're thinking, and how they're building their models. So he has a lot of insights. He doesn't do a lot of podcasts, so we were really lucky to get him.
So, you're the first guest we've had in the human data-generation space. It's a space I was in for a long time, so I'm really curious about this. Could you first of all start by telling the story of Surge—what you were thinking when you started it and how it's gone?
Edwin Chen
Yeah, I can give the founding story. Basically, I used to be an ML engineer at a bunch of these big companies, and the problem I kept running into was that we just kept facing all of these issues getting the data that we needed to train our models.
For example, I used to work on our search and ad systems at Twitter, and one of the first things I wanted to do was build a sentiment classifier. Sentiment analysis is a super-simple problem, and all we needed was 10,000 tweets to train our models. But our human-data system at the time was literally 2 people we'd hired off Craigslist, working 9 to 5. We had to wait a month to get started, and then we had to wait another month for them to label the tweets in a spreadsheet because our tools were just terrible.
When we finally got the data back, we saw that it was complete junk. They didn't understand slang—for example, a phrase like “She's such a bad...” They were labeling it as just negative. They didn't understand hashtags or all these other aspects of tweets. What ended up happening was that I spent a week myself labeling all these tweets because that was actually so much faster and better.
I think one of the things that we often said was that these were really simple things. At the end of the day, sentiment analysis isn't all that complicated. At the same time, we had this bigger problem to solve around how we wanted to optimize our ML systems for the right objectives.
When I first started working on Twitter, this was in the old days when it was purely a chronological timeline. One of the things we wanted to do was make it easier for users to discover and engage with tweets they cared about. The question was, how do we train our recommendation algorithms?
The obvious choice was clicks and retweets at the time. You just train your algorithms to produce as many clicks and retweets as possible. But we tried doing this, and it turned out to be an incredibly negative feedback loop. Once you optimize for clicks, you get the most clickbaity content in the world rising to the top. You get lots of racy content, lots of bikinis, and lots of listicles about “10 horrifying skin diseases,” and so on.
1. The problem with early data labeling systems
We wanted to train our models on deeper principles instead. We'd ask human raters to label tweets and recommendations according to our product principles. But if we just couldn't even get simple sentiment analysis right, we definitely couldn't get this more complex data at the quality and scale we needed.
Eventually, it became a problem that happened over and over again at Google and Facebook, too. Eventually, I realized it was something I needed to go out and build myself.
Lukas Biewald
What year did you start Surge?
Edwin Chen
We started Surge in 2020, right in the middle of the pandemic.
Lukas Biewald
I feel like in 2020 there were some established data-labeling and data-generation companies. Did you have a particular take on the space?
Edwin Chen
Our take on the space was that all of these solutions out there were basically focused on this idea of very commodity labeling, very low-skill work. The example I often give is the problem of drawing a bounding box around a car. You and I can all draw a bounding box around a car. Almost a 3-year-old can draw a bounding box around a car. The bounding box that I draw isn't going to be any different from the bounding box that Terence Tao draws or that Einstein would draw. There's a very low ceiling on the complexity of data that's required.
In contrast, if you think about all the things that we want to do today, we want models that can write poems, and we want models that can solve relativistic equations. There's almost an unlimited amount of intelligence that we want to feed our models. All the other solutions at the time were designed for a very low-skill, commodity style of labor. They weren't focused on quality at all. Instead, they were focused on scale.
Lukas Biewald
I see. How did you source the people in the beginning?
Edwin Chen
Literally, the first few people who were on our platform had just been working with me. It's kind of funny, but I've been working on this problem for a long time, so I already had a network of people who were very interested in doing this kind of work. When they heard that I started Surge, a lot of them joined. A lot of it was also me in the beginning, but it was also all these labelers that I'd accumulated throughout the years.
Lukas Biewald
That's cool. Who were the early customers?
2. Search ranking, clickbait, and product principles
Edwin Chen
The early customers were a lot of tech companies. We had this idea that we wanted to really focus on engineers and research scientists who really understood the quality of data. These were people I'd been working with in the space for a while. There were a lot of friends and contacts I had at all these companies who had been dying for this kind of higher-quality, next-generation solution that could do far more advanced tasks than what were possible at the time.
There were a lot of companies like Airbnb and Twitter, as well as a lot of startups in the search and algorithm space.
Lukas Biewald
At that time—in 2020—that was right around when I was leaving the space. I remember that what was really taking off then was autonomy and robotics, a lot of vision applications. It sounds like you were more focused on text. Is that fair?
Edwin Chen
We were always focused on language and what I call behavior from the beginning. I wouldn't say it was necessarily text per se. There are a lot of complex problems in the image space, too, especially nowadays.
What we didn't want to focus on was very simple image-labeling tasks or very simple bounding-box-style tasks, where I don't really think that there's any intelligence needed to build such solutions. We always focused on this higher-complexity, higher-scale, higher-sophistication space.
Lukas Biewald
I would say that image labeling, when you actually dig into it, is more complicated, for the record. But I take your point.
It seemed like you really built Surge under the radar. Was that an intentional decision? I think you might have said recently that you're over $1 billion in yearly revenue. Is that right? Can we say that?
Edwin Chen
Yeah, we're over $1 billion last year, and we hit our $1 billion ARR number a while ago.
Lukas Biewald
That's an incredible achievement in just 5 years, and you did it with no outside funding. Is that right?
Edwin Chen
Yeah.
Lukas Biewald
That must be a historic level of growth without funding, I think. Did you intentionally avoid the VC path?
Edwin Chen
One of the things that was really important for us was that we wanted customers to be buying from us because they really believed in having high-quality data, not because they saw us mentioned in some tech article. We wanted partners who had the same vision we did and who could go out and show the world how important data actually was.
Again, if you contrast that with the kinds of data people were looking for 5 or 10 years ago, people really weren't focused on quality. They treated data as this kind of commodity. The researchers themselves would barely look at the data. They'd just let their vendor-management teams handle the entire process because they wanted to outsource it.
And so, we just had this idea that we really wanted to focus on people who believed in quality and were buying us for that reason.
Lukas Biewald
But to get to a billion dollars plus in revenue, you've obviously had to scale your system. That goes beyond hiring a few people that you've known for a long time. Can you talk about how you've scaled your processes as your scope has expanded? Do you think it's more important to scale the technology that's enabling this, internal human processes, hiring, or the management of what's going on? As much as you could say about how this works, I'd love to hear it.
Edwin Chen
Yep. Yeah. I think the thing that people underestimate in this space is how much technology you actually want to build. People tend to think that humans are smart, and so if you just throw 10,000 or 100,000 humans at a problem, that will solve it. It's a little crazy to me.
3. Why Surge focused on high-skill, high-quality labeling
Sometimes I'll interview candidates from some of the competitors in our space, and when they describe to me what they're building and how they operate, it's incredibly manual. They're just body shops at the end of the day, and they literally have no technology. If you ask them, “Could you tell me the quality of this worker? Could you tell me how good this worker is at this particular task? Could you show me a dashboard with how your quality is improving over time, what A/B tests you run, and what algorithms you're building to improve it?” they literally can't.
All they're doing, a lot of the time, is dumping data into spreadsheets. Their employees and engineers are either creating the data themselves or reviewing it themselves. There's almost no technology going on.
I think what you have to realize about this space is that quality control is incredibly difficult. If you want to get the highest quality out there, you need to find or create sophisticated algorithms to detect the highest-quality data that you can. It's incredibly complicated because, in this space with LLMs today, you really want LLMs to be good at every task in the world. It's not just a single domain; you want to make sure that they're really good at poetry and, at the other extreme, physics.
4. Scientific collaboration and frontier research data
How do you find the highest-quality data in order to train the models, and then how do you also remove the worst? You can't just throw warm bodies at it. You really need to be able to build a lot of technology to manage it.
Lukas Biewald
Well, I'm coming from a different era, I think, of labeling, where a lot of labels were simple—yes or no, or simple tasks where it's really clear what's right and what's wrong. But I think you're doing much more complicated tasks, and as you're saying, even getting 2 people to agree on high quality versus low quality gets more and more difficult as the task gets more complicated.
Can you talk a little bit about what technology you actually have to build to manage quality for a generic, complicated expert task involving language?
Edwin Chen
Let me start by contrasting with a very old-school take on data and data quality, and how we think about the problem differently. Let me give you an example. Let's say you wanted to train a model to write an 8-line poem about the moon. The way most companies think about it is, “Okay, let's just hire a bunch of people from Craigslist or through some recruiting agency. Let's ask them to write poems.”
The way to think about quality, again going back to the image-annotation days, is: Is this a poem? Is it 8 lines? Does it contain the word “moon”? They just check all these boxes, in the same way you might check, “This is a cat. This is a dog.” They're just checking boxes, and then you're saying, “Sure, this is a great poem because it follows all the instructions. It follows all these checkboxes.”
What happens is that you get terrible poems that feel like they're written by kids in high school. A kid in high school can write an 8-line poem about the moon, but is it a great poem? Is it an evocative poem? Is it the type of poem that a Nobel Prize winner would write? No.
5. From Craigslist workers to a billion-dollar business
Some other companies might say, “Okay, sure, these people on Craigslist don't have any poetry experience. What I'm going to do instead is hire a bunch of people with PhDs in English literature.” But what they don't realize is that this is also terrible. A lot of PhDs, again going back to what I was saying about even people with MIT computer science degrees, aren't good writers or poets. Think of people like Hemingway or Emily Dickinson: They definitely didn't have a PhD. I don't think they even completed college.
We think about quality completely differently. What we want isn't poetry that checks some boxes and uses complicated language. We want the type of poetry that Nobel Prize winners would write.
I think you need a mindset shift. One of the things we think about is that there are certain people who are trained up in this very objective domain of computer vision, a domain that lacks all of these nuances and all this inherent subjectivity. What we want to do instead is recognize that poetry is actually really subjective and rich.
Maybe one poem is a haiku about moonlight on water, another poem is something with internal rhyme and meter, and another one focuses on the emotions behind the moon rising at night. You want to capture that there are 1,000 ways to write a poem about the moon, in a way that there aren't 1,000 ways to draw a bounding box on a car. There aren't 1,000 ways to label something as a cat or a dog.
There isn't a single correct way to write this poem. Each different way that you write it, or each different preference that you have for different types of poetry, gives you different insights into language, the mind, imagery, and human expression.
I talk about poetry a lot, but it's not just poetry. If you think about math as well, there are 1,000 ways to prove the Pythagorean theorem. Each way of proving the Pythagorean theorem is based on different insights into the mathematical reality of the universe.
One thing that happens is that when you think about quality the wrong way, you get commodity data that optimizes for things like inter-rater agreement. Going back to the computer-vision world, if all you're doing is labeling images of cats and dogs, sure, you want high inter-rater agreement. You want people to agree that this is a cat, and you want people to agree that this is a dog.
One of the very easy ways to quality-control in this old-school world is just by seeing whether there's agreement with a majority. But in this new generative-AI world, there's no way that you can ask people to write 1,000 poems and then take a majority vote and get the best poem.
What actually happens when you optimize for higher inter-rater agreement is that you get the lowest common denominator of things that people want, which turns out to be trashy, low-quality, and unengaging a lot of the time. You need to think about quality in a different way in order for your data to really embrace human intelligence and creativity.
6. Scaling without funding and avoiding Silicon Valley status games
Lukas Biewald
I mean, one of the things that we found when I was working in this space is that maybe the biggest issue was actually eliciting from a customer what they really wanted. Obviously, we're doing simpler examples, but you think about putting a bounding box around a car. It seems so simple to an ML researcher, especially one who hasn't looked at a lot of specific data examples, but it is really hard.
You're one of the rare ML researchers who really wanted to look at data. I was, too, and that's why I got into the space, right? But putting a bounding box around a car seems simple until you start looking at real-world data. What if the car is occluded? Do you put the box around where you think the car is? What if the car is in a reflection of a car in a mirror? Do you still want the bounding box around that car?
What if it's a billboard with a picture of a car and it's not a real car? Do you want that? Even in the simplest cases, as soon as you start to look at real-world data, it gets way more complicated.
I remember working with customers, and there would actually be a long process of eliciting from the customer what they wanted. Some customers would write these giant documents trying to enumerate every case and exactly what they wanted, but I think those, too, would end up being hard to reason about exactly what you want.
If you followed those instructions to the letter, you'd end up in some ridiculous cases, which isn't actually what the person wanted in the first place. I'm kind of curious: how do you think about that problem?
Edwin Chen
What we always try to do instead is understand the goal or principle that the researcher has. When we understand the goal or principle, or how they're going to use the data, we can almost put ourselves in their shoes. Instead of asking them, “Okay, what happens when this car is occluded? What happens when this car is in a reflection?” we think, “Okay, deriving from first principles, given what they've told us, what do we think the researcher wanted?”
Maybe if you know that the data is going to be used for some LLM model, then you'd know that this makes sense and this doesn't. We can have that higher-level goal instead of a mechanical list of instructions. That's always what we aim for.
It is hard because, at the end of the day, sometimes researchers don't know either. But this is where, as a company, we often try to have a strong opinion on what the best type of data is, as opposed to almost blindly accepting whatever people tell us. Oftentimes, we actually do like to get into a feisty debate about what type of data would be most useful to build them.
Lukas Biewald
So then, do you hire a lot of ML researchers? Are those the people that you want working directly with the customer in order to—
Edwin Chen
Yeah. We have a lot of researchers. One of the things that we often think about is that we almost consider ourselves a research company, just that instead of researching algorithms in a way that other frontier labs might, we're more about researching the data.
Lukas Biewald
Totally. And has it been—I really admire how you've avoided status games and flown under the radar, but I would think that, in doing that, it might make it harder to hire. Do you have a different strategy for hiring, or do you think that maybe it doesn't matter and you find the people who really connect with your mission?
Edwin Chen
One of the things that we often think about is that we don't want people who are just joining us in order to notch another brand on their resume. Sure, there are other people out there who do that, and we're missing out on those candidates. I think they also tend to be people who want to build large teams. They tend to be people who want to build empires, and it almost turns into hiring for the sake of hiring.
“Why do you want to hire this person?” “I needed to hire this person because the other person doing this job is spending all their time interviewing.” “Why is that person spending all their time interviewing?” “Well, it's because somebody told them that they needed to hire an internal tooling team.” “Why does that person need to hire an internal tooling team?” “Well, it's in order to make the engineers more productive.” “Why aren't they productive enough?” “Because they're spending all their time in meetings.” “Why are they spending all their time in meetings?” It's because we hired all these people, and you need to communicate with them.
7. Why most human data platforms lack real tech
I think there are a lot of benefits if you start with people who really believe in your mission. Then, at least for us, we've been able to stay smaller, with a much smaller team.
Lukas Biewald
How big is your team?
Edwin Chen
We're a little over 100 people.
Lukas Biewald
Wow, that's incredible revenue per employee. Congratulations. You've gone through this really fast change, I think, from being essentially an ML researcher to running a very significant, large company. Where have you felt stretched the most? What's been challenging?
Edwin Chen
Certainly, the thing that I found most challenging is sales. Probably any researcher will tell you that just the concept of having to go out and hawk your product is a little foreign to me.
I think we're lucky, in a sense, that our product, at the end of the day, is for researchers. Having that research mindset, what we're trying to build is almost like that Disney quote: “We don't make movies to make money; we make money in order to make movies.” In a similar way, what we're doing is not trying to generate revenue. We're trying to generate data that will help AGI.
Sometimes when customers—or when new companies—come to us and ask if they can work with us, if their goal is unaligned with AGI, we actually just say no to them. We don't want the revenue; we want to focus on the AGI companies, in a sense.
Again, going back to not raising money, the fact that we don't have a board, or an external board, and the fact that we don't have VCs who are just dying to make as much money as possible, I think that gives us a sort of freedom that allows us to focus on the most important problems. That's allowed us to maintain our research focus.
Lukas Biewald
Totally. So you only work with companies focused on building AGI?
Edwin Chen
For example, if a company came to us and said, “Yeah, we just want to train a—let's say I'm a newspaper and I just want to train a category classifier,” we'd just say no.
Lukas Biewald
What if I want to make an AI video generator? Would that be in your realm?
Edwin Chen
Oh, yeah. We do that in that sense because building such video generators is part of building AGI. So, yeah, in that sense.
Lukas Biewald
I imagine that 5 years ago the data being collected was pretty different from the data collected now. I would think that some of the tasks you'd be doing 5 years ago would be easily automated by LLMs today. It's just been such an astonishing pace of improvement. Can you talk about how the types of tasks have changed over the last few years?
Edwin Chen
Some of the types of work that we do are very, very different. When we first started, a lot of our work was in tasks like search evaluation or content moderation, whereas today it's almost purely LLM work. That's one big difference.
8. Detecting cheaters, liars, and low-quality labelers
Even within LLMs, there's been this obvious trend toward higher complexity, higher sophistication, and higher expertise. For example, there's been a big increase in multimodal complexity. A few years ago, it was all text data, just conversational text assistance, but now we do a lot of work with images, audio, and video. I think the interest is that you actually want the models to understand all of these modalities simultaneously.
One of the things you might want to do is say, “Okay, I'm taking a video of something on my phone, and now I'm asking the model to create a program based on the video on my phone that simulates this in real life.” So there's been a lot of multimodal increase in complexity.
There's also been a big expansion in languages. At first, people were naturally focused on English-only work, but we actually work in over 50 languages now. What's also interesting is that it's very hyper-specialized: we support coding in Argentinian Spanish, and we support legal and financial expertise in Bolivia.
Even today, I think a lot of the models are just not that good—surprisingly so. They're not that good at the different nuances of different languages, dialects, or cultures. I think there's still a lot more progress to be made there.
Probably the biggest shift is just the depth of expertise that a lot of the work requires. You see the models winning IMO gold medals now, and you see them doing all these incredibly advanced tasks, so you actually really need serious thinking power behind them.
Even today, some of the tasks that we do involve spending days or even weeks solving these really interesting problems. It's a very far cry from tasks 5 or 10 years ago, where you might spend 5 seconds labeling a task.
Lukas Biewald
Do you actually have people who can solve Olympiad-level math problems—creating those problems and then solving them?
Edwin Chen
Yeah.
Lukas Biewald
You know, it's funny—this is kind of an aside—but I've been surprised by the scores the latest models are getting on these IMO problems. When I put in more fun brain-teaser problems that my friends pass around, they often can't do them. Do you have a sense of why there's that disconnect?
Edwin Chen
I think a big problem with a lot of the frontier models today is that they've basically been benchmark-hacked. You have all these benchmarks out there, and a lot of them just aren't very good. They're overly academic or overly synthetic.
A lot of these benchmarks have a single objective answer. Going back to my point earlier, the models have been narrowly constrained to be good at these very narrow, objective problems. But a math problem in the real world, or a problem that you would ask as a research mathematician, isn't going to be a closed-ended problem; it's going to be an open-ended exploration.
I think a lot of it stems from this kind of benchmark hacking that's going on.
Lukas Biewald
So if you're going to build a benchmark today to compare frontier models, what kinds of things would it include?
Edwin Chen
What we always say is that the gold standard for evaluating models really is human evaluation, where you just can't fully automate it. People try to build these leaderboards where they take automatic verifiers, but automatic verifiers still only work well in very objective domains.
9. Why inter-annotator agreement is a flawed metric
They've tried building benchmarks like LMSYS, which I think is an absolutely terrible leaderboard that has basically set the industry back by at least 1 year. I can go into more detail on that, but I think a lot of the benchmarks out there are flawed either because they're built using low-quality data or because they're overly synthetic, overly academic, and overly objective in a way that the real world isn't.
Lukas Biewald
Okay. Well, I mean, now I want to hear this. Why do you think LMSYS—you're saying that benchmark is not just bad, but it set the industry back? Can you tell me more about that?
Edwin Chen
Yeah. Basically, LMSYS is, if you don't know what it is, this popular leaderboard of language models. What happens is that people go onto it—literally anybody around the world. What we often hear is that it's literally high schoolers and middle schoolers who can't access models any other way, and they're going onto Arena as their only mechanism.
They go onto Chatbot Arena, enter a prompt, and then see 2 model responses. They vote on which one is better. But if you think about it, they're not taking the time to read or evaluate these model responses at all. They're literally just looking at the responses for 2 seconds and picking whichever one strikes their fancy.
The models could have made everything up. They could have completely hallucinated everything. They could have not followed the instructions at all, and people will just vote on it because, in order to progress, they need to vote on something. Then they see, “Oh, yeah, this model has emojis and a lot of bold formatting, so it just looks really impressive.”
One of the things that we've learned is that the easiest way to improve in this arena is simply to make your model responses a lot longer and double the number of emojis that you have. If you actually think about some of the models that have been released in the past year—think about Llama, or at least the version of Llama 4 that was optimized for LMSYS—it exactly matched all of these patterns.
One of the phenomena that we often hear about from researchers is that they'll tell us, “I'm only going to get promoted if my VP has told me that we advance our model by 10 points on the leaderboard.” They'll look at the leaderboard data themselves and see that a lot of the responses preferred by LMSYS—in its datasets—are literally the responses from the model that followed instructions the worst, or the model that hallucinated everything. Those are the responses that are getting preferred simply because they have a lot of formatting and emojis.
What these researchers tell us is, “I want to work on improving the fundamental capabilities of my model. I want to improve it in coding. I want to work on fixing its hallucinations.” But instead, if the only way I'm going to get promoted is by optimizing for this leaderboard, and the easiest way to optimize for this leaderboard is to make my model hallucinate so that it generates these crazy answers that, even though they're completely wrong, are very compelling to amateur raters who are only spending 2 seconds, then that's what they're going to do.
We've often seen models where, if you look at some of the models that are in the top spots on the leaderboard today and compare them to how they were performing 6 months ago, you actually see that they're worse in many ways.
10. What makes a great poem? Not checkboxes
Lukas Biewald
Do you think so? Another thing that strikes me is that we have these different models made by totally different organizations. You have xAI, run by Elon and making Grok, and you look at Anthropic and OpenAI. These clearly have different cultures, and we've had different people from these organizations on the podcast, yet the models seem to have this surprisingly consistent tone to me.
They feel a little annoyingly positive, a little Boy Scout-y: “Yes, thank you. I will answer your question.” I imagine that must somehow be trained in at some point in the process. I wonder if it's in the preprocessing—is that just what high-quality data on the web looks like—or is it intentionally inserted later in the process? Is the annotation that you're doing somehow contributing to that consistent tone? I feel like we'll look back on it as almost this LLM tone of 2025.
Edwin Chen
Yep. I think it's a combination of everything. For example, we don't explicitly teach the models to follow proper grammar. We generally don't teach the models to use em dashes, but it's kind of just there in pretraining, where these behaviors are baked in.
I think in the future the models will become more and more differentiated, where you'll be able to tell that they have a noticeable difference in style. Even today, maybe because I've looked at this data enough, I can tell from reading the model responses themselves which model is which. Some of them will use certain prefaces, certain words, or certain stylistic patterns, and I can generally tell.
Lukas Biewald
Interesting. Can you think of any tells? Is it more subconscious, or could you cite some specific things that let you know it's a particular model?
Edwin Chen
Some of them just use certain phrases, like “absolutely,” more often than others. Some of them use emojis in a way that others wouldn't.
Lukas Biewald
Llama uses so many emojis. I totally agree with that.
Edwin Chen
Some of them do this thing where they repeat 3 adjectives in a row. Some of them use a lot more Markdown than others.
Lukas Biewald
When you look at the way these different organizations collect data, are there striking differences, or does it feel like everyone's chasing the same types of data?
Edwin Chen
I think there actually are very striking differences between the different companies. Almost every company has its own philosophy on the right way to both train and evaluate its models. It's been surprising how different they are.
Lukas Biewald
Do you have a point of view on how you think it should be done? If you were running an AGI company, what would you do? Is there something you would do differently from what you are asked to do?
Edwin Chen
We had this very strong view that our RLHF data was a lot more effective than SFT data. Maybe you should just explain the difference here for—
Lukas Biewald
Yep.
Edwin Chen
—the general audience too. So, the way SFT data works, just to give an example, would be—
Lukas Biewald
And this is fine-tuning data—supervised fine-tuning.
Edwin Chen
Yes. So SFT stands for supervised fine-tuning. The way it works is, let's suppose that I wanted to train my model to become better at poetry. You, as an annotator or as a reader, would write a prompt. The prompt might be, “Write a sonnet about cheese,” and then you would literally write a sonnet about cheese. That would be a demonstration to the model. That's essentially supervised fine-tuning.
In contrast, in RLHF—reinforcement learning from human feedback—the way it works is that you write a prompt. Again, you might take the same prompt, “Write a sonnet about cheese,” and then you ask the model to generate 2 responses. These might be 2 different models, or 2 different model checkpoints. Model A would produce an answer, and model B would produce an answer. Then you would rate which one is better: “Model A is better because it had emotion behind it, because it was better written, or because it actually followed the format of a sonnet.”
You basically teach the model, “A is better than B.” The benefit is that there are many benefits, but just to enumerate some of them: first, it's a lot more efficient. Even if you're a Nobel Prize laureate, it takes you a lot of time to write a poem about cheese.
11. Measuring subjective quality rigorously
Second, it helps teach the model all these latent preferences. It helps teach the model what's good, but also what's bad.
Basically, we had this very strong view that RLHF was a lot more effective. SFT is important sometimes. One of the things we often tell our customers is that you often want to use a little bit of SFT data to more efficiently bootstrap your models into a phase where RLHF is useful. But we had this very strong view that RLHF was a lot more effective, and so we steered our customers in that direction.
If you actually read the Llama 2 paper, you'll see that one of the discoveries they describe is that, at the beginning, all the researchers thought that SFT data would be more effective as well, but then they ran a bunch of experiments and found that RLHF was just so much more effective. So, yeah, that was one of our beliefs in the past.
Lukas Biewald
When you look into the future and roll things forward a few years, what kinds of data do you think you're collecting now that you won't be collecting? And what new kinds of data do you think you'll be collecting?
Edwin Chen
I actually don't think that we're going to stop collecting any of the data that we're still collecting right now. One of the things that we've often found is that the proportions may change. In the same way that, as you become a more sophisticated adult, you don't need to be trained as much on arithmetic as you did when you were a little kid, what we often find is that if you don't persist with at least some amount of this data, models often just regress in very surprising ways.
I won't name names here, but some of the frontier models just make the most bizarre mistakes these days. I think it's kind of wild, and that's just because they're being fine-tuned in very certain ways. I'll say that some of the ways we expect the data trends to continue are in ways that they haven't before. I'll just list them out.
One is even higher-expertise problems than we have today. Right now, there's a lot of work being done in what I would call graduate-level STEM work, but if you actually want the models to make new scientific discoveries, you're going to need to progress beyond what even the average PhD can do. So, I think there'll be STEM work like that in the future. That's one of them.
Lukas Biewald
But wait, can I ask? What would super-STEM work be? Would that literally be, “I'm going to go out and discover some new principle of chemistry” just to train a model? How would that even manifest?
Edwin Chen
I think it actually is exactly that. Imagine that you are a Stanford professor and you're literally working on your latest frontier research. What you want to be doing is collaborating with an AI to generate new hypotheses and help you test your experiments.
12. What types of data are becoming more important
It's this idea of a scientific collaborator, in the same way that we have a coding collaborator right now in the form of Claude Code and whatnot. I think there will be things like that that are increasingly important.
Lukas Biewald
But how would you collect the data? Would you get a Stanford professor to just score responses? I guess scoring responses would be easier than just generating responses.
Edwin Chen
It wouldn't necessarily be just scoring responses. There's a lot of variety in the different types of data that we collect nowadays. It's literally a Stanford professor or some other person who's capable of understanding frontier physics, collaborating with the models and teaching them the right signals.
That's one domain. Another domain is just longer and longer-horizon tasks. Today, the models can definitely perform tasks that a human could have done in 5, 10, 15, or 30 minutes, or an hour. I think some time horizons will just get very, very long. What are the types of things that you want the models to do that would take a couple of days or even weeks? That's another domain.
A third domain, almost going back to your point about the language that these models exhibit today, is how we actually make the models really, really good at long-form creative writing. Even today, it's kind of interesting: I don't like any of the poetry or short stories that the models can write, even though they have perfect prose. They're just not creative enough. They've almost collapsed to this very, very vanilla, generic dimension.
I think there will be more work there. Up until now, there have maybe been higher-economic-value tasks that people want the models to do instead. I think it's actually important for models to exhibit this kind of creativity in a non-generic way.
A fourth one would just be more and more actions in the real world. Agents have obviously exploded in popularity over the last few months, but they're still at a very, very nascent level. The concept of AI models that can plan and reflect on their actions, try new ideas, and actually affect things in the world will be an increasingly important concept.
Lukas Biewald
I guess one of the things that we might not have expected models to do as a priority is put a lot of their reasoning into code. We've seen a trend where models use code to do the reasoning more and more. You can see it in the chat window. Do you think that's an artifact of the limitations of models today, or do you expect that trend to continue?
Edwin Chen
I actually expect it to continue. I do think that the code they're writing is very important for certain types of verifiability that you might want the model to exhibit. In the same way that, if I could write a lot of code for my day-to-day life, I would, I think it's a very, very important capability for models.
Lukas Biewald
Do you have a sense—I guess I have this intuition—that the amount of code used to train these models is increasing as a fraction of the overall data? Do you think that's accurate? Are you generating more and more code for your customers than other types of data?
Edwin Chen
Yeah. We are.
Lukas Biewald
I guess that's a lead into another question I have: Do you think the trend toward reasoning models is changing the types of data that you're collecting?
Edwin Chen
Definitely. Reasoning models have led to a bunch of new types of collection methods. I think probably the biggest one is this concept of creating RL environments from scratch, where essentially a lot of what we're doing is building these video game-like universes with interesting tools and datasets that the models need to solve.
You can imagine a universe consisting of a simulated AI startup. In this environment, or universe, you basically have a bunch of Gmail messages, Slack messages, Jira tickets, GitHub pull requests, codebases, and so on. You want this universe to be as rich and complex, and as high-fidelity to challenging tasks in the real world, as possible.
For example, you could imagine that in this universe, suddenly AWS goes down and suddenly Slack goes down. What do you do? What does the agent do? If I can't use Slack, I need to make sure that it knows how to handle that and figure out how to solve the problem. There are a lot of interesting things that we're building into these RL environments.
Lukas Biewald
How do you handle the subjectivity of these models? I have a 5-year-old, and she really likes to generate stories, so I've gone really deep trying to get these models to write interesting stories. I agree with you: I don't think they really write interesting stories. A big part of it is that I think they're trying to average the preferences of a lot of different people.
My 5-year-old's taste is really different from my taste in stories, and we're actually asking the model together. Maybe it's not even possible for it to make both of us happy with the story, let alone the average person. It seems like that might require a different type of labeling if you actually want to get interesting art out of it, because good art probably shouldn't get a thumbs-up from every single person who looks at it.
Edwin Chen
Yeah. I think there's this concept of personalization that is still very, very unexplored with the models. There's surface-level personalization, where the model knows that you live in Ohio versus New York City. But there's also a lot of unexplained latent preferences that you have that are kind of hard to articulate.
How do you get the model to somehow learn those from the data for every single person and then apply them to the different responses that it generates? I think that's a really, really interesting concept that still hasn't been fully explored. For example, if you are a mathematician, here are the restaurants that you should try in Paris. Models sometimes just make these weird little personalization generalizations.
I think there's still a lot of work that remains to be done with personalization.
Lukas Biewald
And is that, do you think, because it's hard to collect good personalization data? I'd imagine that, to put yourself in the shoes of someone you're not, the data is never going to be as accurate.
Edwin Chen
I think that’s part of it, but a big part of it is that all the frontier labs only have so many things they can focus on, and they just haven’t quite doubled down on personalization yet.
13. Multimodal data, Argentinian coding, and hyper-specificity
Lukas Biewald
Okay. So let’s talk about synthetic data. I think everybody’s dream is always to be able to automatically generate useful data. We’ve seen a lot of people using LLMs as a judge for many applications, which is a kind of synthetic data. How do you do synthetic-data generation, and how do you think about that?
Edwin Chen
Yeah. I actually do think synthetic data is really useful in some places, but I think a lot of people overestimate what synthetic data can do. I can give a couple of examples. Right now, there are a bunch of models that have been trained really heavily on synthetic data. But similar to what you and I were saying before, that’s partly why they’re very good at these very academic, homework-style, benchmark-style problems, and they’re actually really terrible at real-world use cases. Synthetic data has made models good at synthetic problems, but not real ones—that kind of phenomenon.
I remember maybe a year or so ago, we were running human evals for one of the researchers we work with, and our human evals showed that their models had suddenly tanked. When we dug into it and talked to the researcher, it turned out that they had just trained their model on 10 million or 20 million synthetic math problems. The math problems were in a very, very narrow domain of math, and what they hadn’t realized was how much that was making the model worse at basically every other type of task. There are only so many very closed-ended SAT-style math problems that you want the model to solve. There’s only so much you care about that, and so the model actually just became worse in basically every other domain.
Another thing we often hear from companies is that they tell us, “I spent the past year training my models on synthetic data,” but they’ve only now realized all the problems that caused. They then spend multiple months throwing a lot of it out. A lot of them will tell us they’ve thrown out 10 or 20 million pieces of synthetic data because they found that even just 1,000 pieces of really high-quality human data are more useful. The human data is, I think, both more diverse and more creative, as opposed to getting 10 million pieces of the same thing over and over again. What ends up happening in practice is that companies will try using synthetic data for certain problems for 6 months, and then a lot of the work that we do ends up being cleaning up synthetic data.
14. What's wrong with LMSYS and benchmark hacking
Lukas Biewald
You have front-row seats to what most of the labs are doing. Do you have a point of view on whether we’re at a moment where performance is stuck and kind of needs new methods to get to the next level, or whether the current strategies of pre-training and then reinforcement learning to learn reasoning will get us to a much better set of models?
Edwin Chen
I definitely don’t think we’re stuck at all. I think there are a lot of new methods that are just appearing and a lot of new data that hasn’t been collected yet. From some of the early experiments that we’ve run ourselves, I think we’re going to see some massive progress in some new domains very soon.
Lukas Biewald
Can you describe what some of the new methods are at a high level?
Edwin Chen
Even just some of the methods I was describing in terms of RL environments and some of these new RL methods, I think, are still kind of new to a lot of researchers in industry. It’s basically this concept of exposing the models to all these new environments they haven’t seen before.
Lukas Biewald
What would be an example of a new environment they haven’t seen before?
Edwin Chen
Even just the example I mentioned earlier, where you have a simulated AI startup and then suddenly the model’s environment loses access to Slack and AWS. How does the agent continue operating in this environment, and how does it solve the problem? I think it’s something that just doesn’t appear, in some sense, in any preexisting data or any data that we’ve generated before.
It’s an entirely new concept where the model needs to go through all these different actions. It needs to reflect on what it’s done, and it needs to find new ways of solving a problem. Especially when you expose it to very messy types of data that it may try to retrieve—data that, again, it may not have encountered before—the models can fail in these very unique ways, but I think in ways that they can be taught to progress through.
Lukas Biewald
What do you think about the ARC benchmark? That’s an interesting one: it has deceptively simple problems that I think you or I would have no trouble with. They’re certainly easier than math Olympiad problems and are more about visual pattern recognition. Do you think those are likely to get solved in the next 1 or 2 years?
Edwin Chen
I’ll admit, even the ARC benchmark is very surprising to me as well. If you asked me, without ever having seen the ARC benchmark, whether models could solve it, I’d be like, “Yeah, absolutely.” So there’s something about them that I still don’t quite understand—why models can’t solve them.
We’re actually doing a lot of work to generate ARC-style problems, so it’ll be interesting to see how much that helps models improve. Right now, I don’t have a good understanding of why models can’t solve them.
15. Personalization and taste in model behavior
Lukas Biewald
Do you have an opinion on whether open-source or closed-source models are going to win in the long term?
Edwin Chen
My guess is that, at least with the current state of things, closed-source models will continue winning. Part of that is because LLMs are just so valuable that, if you ever try to build open-source models, the way incentives currently work is that eventually you’re going to be forced to make them closed-source. If you want to build truly open-source models that are really, really good, we almost need a different kind of incentive structure to make sure that happens or remains in place. Otherwise, if you look at the history of other types of open-source models, they’ve gotten more closed over time.
Lukas Biewald
What do you mean? What are you referring to?
Edwin Chen
If you even think about Meta, which is thinking about making Llama closed-source, models are just so expensive to train, and people want to fully capture the value. If you ever build a truly, truly good open-source model, I don’t think it’ll remain open-source for very long, again, unless you can change the incentive structure in some way that I haven’t figured out yet.
Lukas Biewald
What do you think is the ratio of spending on data to spending on compute in the training of a large model, and how do you expect that to change over time?
Edwin Chen
I definitely think it should be a lot higher. Sometimes people, I would say, almost skimp on gathering data.
Lukas Biewald
What do you think it is today, roughly?
Edwin Chen
I actually don’t have a good sense myself. I think it varies widely depending on the labs, but it can be anywhere from a couple of percentage points, or 1%, to 10%.
One thing we’ve often heard from researchers is that some of them are really, really good at using human data: they know how to come to us to gather it, and they know how to apply it in their own work. What they often tell us is that some of their counterparts don’t know how to use human data, so they just move a lot slower. They may try somewhat wild ways to get around the fact that they don’t have any human data, and it just ends up being this complex slowdown for them. I hope it’ll get higher in the future. I think there are a lot of ways in which people still underestimate the use of data.
Lukas Biewald
Awesome. Well, I appreciate your time, and congrats on building a fantastic business.
Edwin Chen
Yeah, it was a great chat.
Lukas Biewald
Great chat. Thanks.
Edwin Chen
Thanks so much. Bye.