[BidClub_]
No Priors · · 33 min

No Priors Ep. 124 | With SurgeAI Founder and CEO Edwin Chen

Sarah GuoElad GilEdwin Chen

YouTube
TL;DR
  • Surge says it surpassed $1 billion in revenue last year with a little over 100 employees, five years after launching in 2020, serving clients including Google, OpenAI, and Anthropic. Chen credits being profitable from the start: capital was not Surge’s bottleneck, so he avoided giving up control for fundraising as social proof and argues founders should “go out and build whatever you're dreaming of” before raising.

  • Surge's stated moat is not labor supply but a technology-mediated quality system for data whose ceiling rises with generative complexity. Hemingway and a second-grader can draw essentially the same bounding box, but poetry, proofs, code, and games admit outcomes with radically different quality. Without systems that measure those differences, vendors are merely “scaling up mediocrity.”

  • Surge's work spans SFT and preference labels, verifiers, failure-mode analysis, and rich RL environments that simulate work over long horizons. Chen's salesperson world spans Salesforce, Gmail, Slack, spreadsheets, documents, presentations, calendars, and even a car accident that changes meeting travel. He sees “no ceiling” on useful realism and doubts one terminal reward can capture a complicated trajectory.

  • Chen is categorical that human feedback “will never run out,” even if models become superhuman, because models still need external objectives and synthetic volume routinely fails quality tests. Customers may spend six months generating 10 million-20 million synthetic items only to find “99% of it just wasn't useful”; he says 1,000 highly curated human examples can outperform 10 million synthetic ones.

  • LMArena can reward clickbait rather than capability, creating an incentive to make answers longer, more formatted, and more emoji-heavy. Five-to-10-second raters often do not check factuality or instruction-following, and Chen says researchers have accepted regressions in both to improve leaderboard rank. His alternative is expensive but direct: careful human evaluation with fact-checking, instruction verification, and taste.

  • Chen expects a plural frontier-model market rather than commodity convergence, and picks xAI as the underdog most likely to catch OpenAI, Anthropic, and DeepMind. Anthropic's coding and enterprise focus, OpenAI's consumer orientation, and Grok's different behavioral boundaries create distinct products; eventual AGI “may encompass this all,” but companies can only sustain so many priorities at once.

  • The Meta–Scale deal increased Surge's visibility among legacy Scale users, while public model evaluation is its next strategic expansion. Chen says those users had not known about the under-the-radar company, and Surge wants to expose benchmark-driven failure modes as frontier labs publish less. The larger thesis: independent quality measurement becomes more valuable as model providers optimize against increasingly gameable public signals.

Digest · the substance, structured for research

1. Bootstrapping preserved control because capital was not Surge's constraint

  • Chen traces Surge to the data bottleneck he repeatedly encountered doing ML at Google, Facebook, and Twitter: teams could barely source what they needed for a simple classifier, let alone the “next generation AI systems” they imagined. Founded in 2020, Surge reached its fifth anniversary with just over 100 employees and, Chen says, more than $1 billion in revenue during the prior year. Sarah says it serves clients including Google, OpenAI, and Anthropic.

  • Bootstrapping was initially practical: Surge was “profitable from the start” and did not need money, while fundraising meant surrendering control. Chen's sharper complaint is cultural—too many founders optimize for a $10 million round and TechCrunch headline before identifying a problem. His first-principles default is to build first, then raise if an actual financial constraint appears.

  • Elad's pushback—worth keeping: Silicon Valley may raise too readily, but founders elsewhere often underuse venture capital when it could unlock scale; Sarah adds that unknown founders may need external validation to recruit. Chen distinguishes founders who need subsistence capital from those with savings, then challenges assumed headcount: data scientists tune 2%-5%, while early startups need 10x-100x swings, and founders—not an early PM—should own product.

2. Generative data turns quality from a checkbox into the product

  • At the simplest level, “our product is our data”: for coding models, Surge supplies SFT solutions and unit tests, preference judgments between code or explanations, and verifiers such as whether a web app contains a login button or triggers the intended action. Evaluation accompanies training data through comparative model insights, loss patterns, and failure modes delivered back to labs.

  • Chen separates this from competitors he calls “body shops,” whose deliverable is “warm bodies” rather than measured output. His analogy illustrates the quality ceiling: Hemingway and a second-grader can draw essentially the same bounding box around a car, but Hemingway can write a much better poem—generative AI has “almost an unlimited ceiling” on quality.

  • Asked how Surge can evaluate enough evaluators, Chen compares its system to Google Search or YouTube ranking millions of pages and videos. Surge collects signals about annotators, their work, and their activity, then feeds them into internally built ML systems to measure quality.

  • The eight-line moon poem exposes why credentials and checklists fail: “is it a poem,” eight lines, and the word moon can all pass while the writing remains high-school-level. English-literature PhDs are not automatically poets, just as some MIT CS graduates are poor coders. Quality must embrace haiku, internal rhyme, emotion, and “a thousand ways” to prove the Pythagorean theorem—or the vendor is “scaling up mediocrity.”

3. Humans and models must co-produce data as RL worlds grow realistic

  • Rising model baselines do not remove humans; they change the interface through “scalable oversight,” which Chen defines as humans and models producing data better than either can alone. An SFT story that once began from a blank page can now start with a model's generic skeleton, leaving the person to perform substantial creative editing instead of low-value “cruft.”

  • RL environments make that collaboration materially harder. Chen's salesperson simulation spans Salesforce, Gmail leads, Slack conversations, Excel tracking, Google Docs, PowerPoint, and a calendar, then adds evolving time and external events—a car accident should alter travel to a customer meeting. Thousands of messages and hundreds of emails must remain realistic, interesting, and mutually consistent.

  • On realism, Chen's answer is categorical: “there's no ceiling.” More diversity, richness, and longer time horizons give models more to learn from; synthetically generating an environment is insufficient if the resulting world lacks coherent events, creativity, or the tools needed to represent an entire job.

  • His five-to-10-year demand forecast is “all of the above,” not an RL-only substitution for other data. Environments can create extremely long, rich trajectories, but a single terminal reward may not capture all the work involved in achieving a complicated goal. Chen expects expert reasoning, environments, verifiers, and multiple rewards to remain complementary.

4. Synthetic scale and public leaderboards can optimize the wrong objective

  • Chen is categorical that human feedback “will never run out.” Customers sometimes arrive after six months and 10 million-20 million synthetic items, only to conclude that “99% of it just wasn't useful” and search for the small usable slice. His deliberately extreme comparison: 1,000 highly curated human examples can be more valuable than 10 million synthetic points.

  • The need is not merely for better examples but for an external objective. Chen cites a top frontier model that, in roughly 10% of his uses, inserts random Hindi or Russian characters into otherwise unrelated answers about subjects such as Donald Trump or Barack Obama. The model is not self-consistent enough to flag the error, so a person must still say, “this is wrong.”

  • His indictment of LMArena is that raters spend five-to-10 seconds choosing whichever response “looks better,” rewarding formatting, bolding, emojis, and length without checking factuality or instruction-following. Training to that signal becomes the model equivalent of clickbait; Chen says the easiest way to improve rank is simply to make responses longer.

  • The alternative is slow, proper human evaluation: fact-check the answer, verify every instruction, and use people with taste to judge writing. Sarah's cross-domain analogy sharpens the risk: protein evolution can select bizarre, unanticipated activities against a narrow catalytic signal, just as model training drives toward a local maximum shaped by whatever feedback designers chose.

5. A plural model market raises the value of independent evaluation

  • On the Meta–Scale deal, Chen says the effect was beneficial for Surge, which he describes as already the largest player: legacy teams that had not heard of the under-the-radar company became aware of it. His broader concern is category damage—low-quality vendors burn labs on human data, pushing them toward slower methods with poor objectives and ultimately slowing model progress.

  • Asked which underdog could catch OpenAI, Anthropic, and DeepMind, Chen chooses xAI because it is “very hungry and mission-oriented.” He expects more frontier providers rather than commodity convergence: Anthropic has excelled at coding and enterprise, OpenAI carries a consumer orientation through ChatGPT, and Grok has different boundaries around what it will say and build.

  • The mechanism is organizational focus and model personality: companies can prioritize only so many principles at once, producing different skills and behavior. Chen already switches among models depending on the task and expects that habit to deepen across personal and professional life, even while conceding that eventual AGI “may encompass this all.”

  • Surge's next strategic expansion is public research and external model evaluation, partly because frontier labs publish less. Chen says researchers have explicitly told him their VPs accept worse factuality and instruction-following if LMArena rank rises; six months of longer, flashier answers can therefore equal “zero progress.” Meanwhile, IFEval rewards contrived tasks such as capitalizing five letters in every mention of Abraham Lincoln rather than demonstrating real-world usefulness.

Sarah Guo

Hi listeners, welcome back to No Priors. Today Elad and I are here with Edwin Chen, the founder and CEO of Surge, the bootstrapped human data startup that surpassed a billion in revenue last year and serves top-tier clients like Google, OpenAI, and Anthropic. We talk about what high-quality human data means, the role of humans as models become superhuman, benchmark hacking, why he believes in a diversity of frontier models, the Scale–Meta M&A deal, and why there's no ceiling on environment quality for RL or the simulated worlds that labs want to train agents in.

Edwin, thanks for joining us.

Edwin Chen

Great seeing you guys today.

Sarah Guo

Surge has been really under the radar until just about now. Can you give us a little bit of color on the scale of the company and what the original founding thesis was?

Edwin Chen

We hit over $1 billion in revenue last year. We're kind of the biggest human data player in the space, and we're a little over 100 people. Our original thesis was that we really believed in the power of human data to advance AI, and we had a really big focus from the start on making sure that we had the highest-quality data possible.

Sarah Guo

Can you give people some context for how long you've been around and how you got going? I think, again, you've accomplished an enormous amount in a short period of time, and you've been very quiet about some of the things you've been doing. It would be great to get a little bit of history—when you started, how you got started, and how long you've been around.

Edwin Chen

We've been around for 5 years. I think we just hit our 5-year anniversary, so we started in 2020. Before that, I used to work at Google, Facebook, and Twitter. One of the reasons we started Surge was that I used to work on ML at a bunch of these big companies, and the problem I kept running into over and over again was that it was really impossible to get the data we needed to train our models. It was this big blocker that we faced over and over again.

There was so much more that we wanted to do, even just the basic things. We struggled so hard to get the data. It was really the big blocker. But simultaneously, there were all these more futuristic things that we wanted to build.

If we thought about the next generation of AI systems, if we could barely get the data we needed at the time just to build a simple image classifier, how would we ever advance beyond that? That really was the biggest problem. I can go into more of that, but that was what we faced.

Sarah Guo

You guys are also known for having bootstrapped the company, versus raising a lot of external venture money. Do you want to talk about that choice—going profitable early and then scaling off of that?

Edwin Chen

In terms of why we didn't raise, a big part of it was obviously just that we didn't need the money. I think we were very lucky to be profitable from the start, so we didn't need the money. It always felt weird to give up control.

One of the things I've always hated about Silicon Valley is that you see so many people raising for the sake of raising. I often see that a lot of founders I know don't have some big dream of building a product that solves a problem they really believe in. If you talk to a bunch of YC founders, or whoever it is, what is their goal? It really is to tell all their friends that they raised $10 million and to show their parents they got a headline on TechCrunch. That is their goal.

I think of my friends at Google. They often tell me, "I've been at Google or Facebook for 10 years, and I want to start a company." I'm like, "Okay, so what problem do you want to solve?" And they don't know. They're like, "I just want to start something new. I'm bored."

It's weird because they can pay their own salaries for a couple of months. Again, they've been at Google and Facebook for 10 years; they're not fresh out of school. They can pay their own salaries, but the first thing they think about is going out and raising money.

They might try talking to some users and building an MVP, but they do it in a throwaway manner where the only reason they do it is to check off a box on a startup accelerator application. Then they'll just play around with these random product ideas, and they happen to get a little bit of traction so that the VC DMs them.

They spend all their time tweeting and go to these VC dinners, and it's all so that they can show the world they raised a big amount of money. Raising immediately always felt silly to me. Everybody's default is to immediately raise.

But if you thought about it from first principles, if you didn't know how Silicon Valley worked or that raising was a thing, why would you do that? What is money really going to solve for 90% of these startups where the founders are lucky to have some savings?

I really think your first instinct should be to go out and build whatever you're dreaming of. If you run into financial problems, think about raising money then, but don't waste all this effort and time up front when you don't even know what to do with it.

Elad Gil

Yeah, it's funny. I feel like I'm one of the few investors who actually tries to talk people out of fundraising often, right? I had a conversation today where the founder was talking about doing a raise, and I'm like, "Why? You don't have to. You can maintain control."

On the flip side, I would actually argue that outside of Silicon Valley, too few people raise venture capital when the money can actually help them scale. So I feel like in Silicon Valley there's too much, and outside of Silicon Valley there's too little. It's this interesting spread of different models that stick.

Edwin, what would you say to founders who feel like there's some external validation necessary to hire a team or scale their team? This is a very common complaint or rationale for raising more capital.

Edwin Chen

I think about it in a couple of ways. I guess it depends on what you mean by external validation. In my mind, I often think about things from a first-principles perspective: Are you trying to build a startup that's actually going to change the world? Do you have this big thing that you're dreaming of? If you have this big thing you're dreaming of, why do you care?

Maybe the way to think about it is in Sarah's context. If you're, say, a YC founder, you haven't been at Google, Meta, or Twitter; you don't have this network of engineers; you're a complete unknown; you haven't worked with very many people; and you're straight out of school.

Sarah Guo

How do you then attract that talent? To your point, you can tell a story about how you're going to build things or what you're going to do, but it is a harder obstacle to convince others to join you or come on board, or to have money to pay them, if you don't have a long work history.

Edwin Chen

I would differentiate between 2 things. One is: Do you need the money? First of all, there's a difference between people who are literally fresh out of school, or maybe have never gone to school in the first place, so they don't have any savings and literally need some money in order to live. Then there are others who don't necessarily need money because they've been working at Google or Facebook for 10 years—or 5 years, whatever it is—and have some savings.

I would say one question is: Do you really need to go out and hire all these people? One thing I often see—I'm curious what you guys see—is founders will tell me, "I'm trying to think about the first few hires I'm going to make. I'm going to hire a PM. I'm going to hire a data scientist. These are 2 of my first 5 to 10 hires."

I'm like, "What?" This is just wild to me. I would never hire a data scientist as one of the first few people in a company. I say this because I used to be a data scientist. Data scientists are great when you want to optimize your product by 2% or 5%, but that's definitely not what you want to be doing when you start a company.

You're trying to swing for 10x or 100x changes, not worry and nitpick about small percentage points that are just noise anyway. Similarly, product managers are great when your company gets big enough, but at the beginning you should be thinking for yourself about what product you want to build.

Your engineers should be hands-on; they should be having great ideas as well. Product management is this weird conception that big companies have when your engineers don't have time to be in the weeds on the details and drive things themselves. It's not a role that you come up with beforehand.

Elad Gil

So I guess with the initial Surge AI team, it sounds like you had a small, tight engineering team. You started building the product and were bootstrapping off of revenue. At this point, you’re at over $1 billion in revenue, which is amazing. How do you think about the future of how you want to shape the organization, how big you want to get, and the different products you’re launching and introducing? What do you view as the future of Surge AI, and how is all of that going to evolve?

Sarah Guo

Before we do that, can you just explain, at whatever level of detail makes sense here, what the $1 billion of revenue is? Maybe how the product supports the company, what your data is, and who your humans are, because I think there’s just very little visibility into all of that.

Edwin Chen

In terms of what our product is, at the end of the day, our product is our data. We literally deliver data to companies, and that is what they use to train and evaluate their models. Imagine when you’re one of those frontier labs and you want to improve your model’s coding abilities. What we will do on our end is gather a lot of coding data.

This coding data may come in different forms. It may be SFT data, where we are literally writing out coding solutions, or unit tests: these are the tests that a good piece of code must pass. Maybe it’s preference data, where it’s, “Here are 2 pieces of code,” or, “Here are 2 coding explanations. Which one’s better?” These might be verifiers: “Here’s a web app that I created. I want to make sure that in the top-right hand of the screen there’s a login button,” or, “I want to make sure that when you click this button, something else happens.”

There are a bunch of different forms that this data may take, but at the end of the day, what we’re doing is delivering data that will help the models improve on these capabilities. Very related to that is the notion of evaluating the models. You also want to know: Is this a good coding model? Is it better than this other one? What are the errors on which this model is weak, and is this model worse? What insights can we get from that?

In addition to the data, oftentimes we’re delivering insights to our customers. We’re delivering loss patterns and failure modes. There may be a lot of other things related to the data, but at the end, it’s this universe of applications, or just this universe around the data, that we deliver. That is our product.

Sarah Guo

Yeah, and maybe going back to Elad’s question, maybe “product” isn’t actually the right word here. What’s repeatable about the company, or what are the core capabilities that you guys have that you would say your competitors fail to meet the mark on?

Edwin Chen

The way we think about a company, and the way we differentiate from others, is that a lot of other companies in this space are essentially just body shops. What they are delivering is not data; they are literally just delivering warm bodies to companies. At the end of the day, they don’t have any technology.

One of our fundamental beliefs is that quality is the most important thing. Is this high-quality data? Is this a good coding solution? Is this a good unit test? Is this mathematical problem solved correctly? Is this a great poem?

Basically, a lot of companies in this space—just based on how things have worked out historically—have treated quality and data as commodities. One of the ways we often think about it is: Imagine you were trying to draw a bounding box around a car. Sarah, you and I are probably going to draw the same bounding box. Ask Hemingway and ask a second grader; at the end of the day, we’re all going to draw the same bounding box. There’s not much difference in what we can do.

There’s a very low ceiling on the bar of quality. But take something like writing poetry. I suck at writing poetry. Hemingway is definitely going to write a much better poem than I am. Or imagine a VC pitch deck: you’re going to create a much better pitch deck than I will. There’s almost an unlimited ceiling in this generative AI world on the type of quality that you can build.

The way we think of our product is that we have a platform and actual technology that we’re using to measure the quality that our workers or annotators are generating.

Sarah Guo

If you don’t have that technology, if you don’t have any way of measuring it, is the measurement through human evaluation? Is it through model-based evaluation? I’m a little bit curious how you create that feedback loop, since to some extent it’s a question of how you have enough evaluators to evaluate the output relative to the people generating the output. Or do you use models? How do you approach it?

Edwin Chen

I think 1 analogy that we often make is to think about something like Google Search or YouTube. You have millions of search results, millions of web pages, and millions of videos. How do you evaluate the quality of these videos? Is this a high-quality video? Is this a high-quality web page? Is it informative, or is it really spammy?

The way you do this is that you gather many signals. You gather page-dependent signals, user-dependent signals, and activity-based signals, and all of these feed into a giant machine-learning system at the end of the day.

In the same way, we gather all these signals about our annotators, the work that they’re performing, and their activity on the site. We feed it into a lot of these different systems. We basically have a machine-learning team in-house that builds many of these algorithms to measure all of this.

Sarah Guo

What is changing or breaking as you are scaling increasingly sophisticated annotations? If model-quality baselines are going up every couple of months, then the expectation is that they exceed what might have been a random human at some point—as you said, something that can draw a bounding box—across all these different fields where we have models better than the 90th percentile at some point.

Edwin Chen

This is actually something that we do a lot of internal research on ourselves as well. There’s basically this field of AI alignment called scalable oversight, which is the question of how you have models and humans working together, hand in hand, to produce data that is better than either one can achieve on their own.

Even today, something like writing an SFT story from scratch—sure, a couple of years ago we might have written that story completely from scratch ourselves. Today, it’s just not very efficient. You might start with a story that a model created and then edit it, perhaps in a very substantial way.

Maybe the core of it is very vanilla and generic, but there’s just so much craft that is inefficient for a human to do and doesn’t really benefit from the human creativity and ingenuity that we’re trying to add to the response. You can just start with this bare-bones structure that you’re basically layering on top of.

Again, there are more sophisticated ways of thinking about scalable oversight, but this question of how you build the right interfaces, how you build the right tools, and how you combine people with AI in the right ways to make them more efficient is something that we build a lot of technology for.

Sarah Guo

A lot of the discussion in terms of what human data the labs want has moved to RL environments and reward models in recent months. What is hard about this, or what are you guys working on here?

Edwin Chen

We do a lot of work building our environments. I think 1 of the things that people really underestimate is how complicated it is. You can’t just synthetically generate it. For example, you need a lot of tools because these are massive environments that people want.

Sarah Guo

Can you give an example, just to make it more real?

Edwin Chen

Imagine you are a salesperson. When you are a salesperson, you need to be interacting with Salesforce, getting leads through Gmail, talking to customers in Slack, creating Excel sheets to track your leads, writing Google Docs, and making PowerPoint presentations to present things to customers.

Basically, these are very rich environments that are literally simulating your entire world as a salesperson. Everything in the future is now on your desktop as well. Maybe you have a calendar invite to meet a customer, and then you want to simulate a car accident happening and getting notified of that, so you need to leave a little bit earlier. All these things are things that we actually want to model in these very rich RL environments.

The question is how you generate all the data that goes into this. You’re going to need to generate thousands of Slack messages and hundreds of emails. You need to make sure that these are all consistent with each other. You need to make sure that, going back to my core example, time is evolving in these environments and certain external events happen.

How do you do all of this in a way that's actually interesting and creative, but also realistic and not incongruent with each other? There's just a lot of thought that needs to go into these environments to make sure that they're, again, rich, creative environments that the models can learn interesting things from. So, yeah, you basically need a lot of tools and a balance of sophistication for great UX.

Elad Gil

Is there any intuition for how real or how complex is enough, or is there just no ceiling on the realism that's useful here or the complexity of the environment that's useful here?

Edwin Chen

I think there's no ceiling. At the end of the day, you just want as much diversity and richness as you can get, because the more richness you have, the more the models can learn from—the longer the time horizons, the more that the models can learn and improve on. So I think there's almost an unlimited ceiling here.

Elad Gil

If you were to make a 5- or 10-year bet on what scales most in terms of demand from people training AI models and types of data, is it RL environments, traces of expert reasoning, or what other areas do you think there's going to be a really large demand for?

Edwin Chen

I think it will be all of the above. I don't think RL environments alone will suffice, just because—it depends on how you think about it—but RL environments often have very, very rich trajectories that are very, very long. So it's almost inconceivable that a single reward—I mean, I think even today we often think about things in terms of multiple rewards, not just a single reward—but a single reward just may not be rich enough to capture all the work that goes into the model solving some very, very complicated goal. So I think it'll probably be a combination of all those.

Elad Gil

If you assume eventually some form of superhuman performance across different model types relative to human experts, how do you think about the role of humans relative to data and data generation versus synthetic data or other approaches? At what point does human input sort of run out as a useful point of either feedback or data generation?

Edwin Chen

I think human feedback will never run out, and that's for a couple of reasons. Even if I think about the landscape today, I think people often overestimate the role of synthetic data. Personally, I think synthetic data is very, very useful. We use it a ton ourselves in order to supplement what the humans do. Again, like I said earlier, there's a lot of craft that simply isn't worth a human's time.

What we often find, for example, is that customers will come to us and say, “For the past 6 months, I've been experimenting with synthetic data. I've gathered 10 to 20 million pieces of synthetic data.” Actually, we finally realized that 99% of it just wasn't useful. And so we're trying to find right now—we're trying to curate the 5% that is useful—but we are literally going to throw out 9 million of it.

Oftentimes, you'll find out that even 1,000 pieces of high-quality human data—highly curated, really, really high-quality human data—is actually more valuable than those 10 million points. That's one thing I'll say. Another thing I'll say is that sometimes you need an external signal to the models. The models just think so differently from humans that you always need to make sure that they're aligned with the actual objectives that you want.

Let me give 2 examples. One example is kind of funny: sometimes, if you use one of the frontier models—one of the top models, or one of the models everybody thinks is one of the top ones—maybe 10% of the time when I use it, it'll just output random Hindi characters and random Russian characters into one of my responses. I'll be like, “Tell me about Donald Trump. Tell me about Barack Obama,” and just in the middle of it, it'll output Hindi and Russian. It's like, “What is this?” The model just isn't self-consistent enough to be aware of this. It's almost like you need an external human to tell the model that this is wrong.

One of the things I think is a giant plague on AI is LMArena. I'll skip the details for now, but people will often train their models on the wrong objectives. The mental model that you should have of LMArena is that people are writing prompts, they'll get 2 responses, and they'll spend 5 or 10 seconds looking at their responses and just pick whichever one looks better to them. So they're not evaluating whether or not the model hallucinated. They're not evaluating factual accuracy or whether it followed any instructions. They're literally just vibing with the model: “Okay, yeah, this one seemed better because it had a bunch of formatting, it had a bunch of emojis, it just looks more impressive.”

And people will train on basically LMArena's subjective evaluations, and they won't realize all the consequences of it. Again, the model itself doesn't know what its objective is. It's like you almost need an external quality signal in order to tell it what the right objective should be. And if you don't have that, then the model will just go in all these crazy directions. Again, you may have seen some of the results with LMArena before, but it'll just go in all these crazy directions that kind of mean you need these external validators.

Sarah Guo

This also happens when you do different forms of protein evolution or things like that, where you select a protein against a catalytic function or something else, and you just randomize it and have a giant library of them. You end up with the same thing, where you have these really weird activities that you didn't anticipate actually happening.

Sometimes I think of model training as almost this odd evolutionary landscape that you're effectively evolving and selecting against, and you're shaping the model into that local maximum or something. It's this really interesting output of anything where you're effectively evolving against a feedback signal, and depending on what that feedback signal is, you just end up with these odd results. It's interesting to see how it transfers across domains.

These, of course, as you said, 5-second-reaction academic benchmarks or even nonacademic industrial benchmarks are easily hacked, or are not the right gauge of performance against any given task. They are very popular. What is the alternative for somebody who's trying to choose the right model or understand model capability?

Edwin Chen

The alternative that I think all the frontier labs view as the gold standard is basically human evaluation. Proper human evaluation, where you're actually taking the time to look at the response: you're going to fact-check it, you're going to see whether or not it followed all the instructions, and you have good taste, so you know whether or not the model has good writing quality.

This concept of doing all that and spending all the time to do that, as opposed to just vibing for 5 seconds, I think is really, really important, because if you don't do this, you're basically just training your models on the analog of clickbait. So I think it actually is really, really important for model progress.

Sarah Guo

If it's not LMArena, how should people actually evaluate model capability for any given task?

Edwin Chen

What all the frontier labs find is that human evals really are the gold standard. You really need to take a lot of time to fact-check these responses and verify whether they followed the instructions. You need people with good taste to evaluate the writing quality, and so on and so on. And if you don't do this, you're basically training your models on the analog of clickbait. I think that really, really harms model progress.

Sarah Guo

Is there work that Surge AI is doing in this domain of trying to standardize human eval or make it more transparent to end consumers of the API or even users?

Edwin Chen

Internally, we do a lot of work today with all the frontier labs to help them understand their models. Again, we're constantly evaluating them, constantly surfacing loss areas for them to improve on, and so on.

Right now, a lot of this is internal, but one of the things we actually want to do is external forms of this as well, where we're helping educate people on the different capabilities of all these models: these models are better at coding, these models are better at instruction following, and these models are actually hallucinating a lot, so don't trust them as much. We do want to start a lot of external work to help educate the broader landscape on this.

Sarah Guo

If we can zoom out and talk just about the larger competitive landscape and what happens with frontier models over time, what does a Meta–Scale AI deal mean for you guys, or what do you make of it?

Edwin Chen

I think it's kind of interesting. We were already the number-one player in the space. It's been beneficial because there were still some legacy teams using Scale; they just didn't know about us because we were still pretty under the radar.

I think it's been beneficial because one of the things we've always believed is that sometimes, when you use these low-quality data solutions, people get burned on human data, and then they have this negative experience, so they don't want to use human data again.

And so, to try these other methods—which are honestly just a lot slower and don't come with the right objectives—I think it just harms model progress overall. The more we can get all these frontier labs using high-quality data, I think it actually really, really is beneficial for the industry as a whole. So I think overall it was a good thing to happen.

Sarah Guo

If you were to make a bet that an underdog catches up to OpenAI, Anthropic, and DeepMind, who would it be?

Edwin Chen

I would bet on xAI. I think they're just very hungry and mission-oriented in a way that gives them a lot of really unique advantages.

Elad Gil

I guess maybe another broader question is: Do you think there are 3 competitive frontier models, 10 competitive frontier models, or a couple of years from now? And are any of those open source?

Edwin Chen

Yeah. I actually see more and more frontier models opening up over time because I don't think that the models will be commodities. One of the things that has actually been surprising over the past couple of years is that you see all of their models have their own focuses that give them unique strengths.

For example, I think Anthropic has been really, really amazing at coding and enterprise, and OpenAI has this big consumer focus because of ChatGPT. I actually really love its model's personality. Grok, you know, just has a different set of things it's willing to say and build.

It's almost like every company has its own set of principles that they care about. Some will just never do one thing. Others are totally willing to do it. Models will have so many different facets to their personality and so many different facets to the types of skills that they will be good at.

Sure, eventually AGI will maybe encompass all this, but in the meantime, you just need to focus. There are only so many areas of focus that you can have as a company, and so I think that will lead to different strengths for all the model providers.

Today, we already see a lot of people, including me, switching between all the different models depending on what we're doing. In the future, I think that will happen even more as people use more and more models for different aspects of their lives, both their personal and professional lives.

Sarah Guo

Going back to something Elad mentioned, where should we expect to see Surge investing over time? What do you think you guys will do a few years from now that you don't do today?

Edwin Chen

Again, I think I'm really excited about this more public research push that we're starting to have. I think it is really interesting that, for obvious reasons, a lot of the frontier labs just aren't publishing anymore. As a result, I think the industry has almost fallen into a trap that I worry about.

Maybe to dig into some of the things I said earlier, there are some negative incentives in the industry and some concerning trends that we've seen. Going back to LLMs, a lot of researchers will tell us that their VPs make them focus on increasing their rank on LMArena.

I've had researchers explicitly tell me that they're okay with making their models worse at factuality and worse at following instructions, as long as it improves their ranking. Again, that literally happens because the people ranking these things on LMArena don't care whether the models are good at following instructions. They don't care whether the models are emitting factual responses.

What they care about is: Did this model emit a lot of emojis? Did it emit a lot of bold words? Did it have really long responses? Because that's just going to look more impressive to them. One of the things that we found is that the easiest way to improve your rank on LMArena is literally to make your model's responses longer.

What happens is that there are a lot of companies trying to improve their leaderboard rank. They'll see progress for 6 months because all they're doing is unwillingly making their model responses longer and adding more emojis, and they don't realize that all they're doing is training their models to produce better clickbait.

They might finally realize 6 months or a year later—you may have seen some of these things in the industry—that they spent the past 6 months making zero progress. In a similar way, besides LLMs, you have all these academic benchmarks, and they're completely divorced from the real world.

A lot of teams are focused on improving these SAT-style scores instead of real-world progress. I'll give an example: There's a benchmark called IFEval, which stands for Instruction-Following Evaluation. If you look at IFEval, some of the instructions it's trying to check whether our models can follow are like, “Hey, can you write an essay about Abraham Lincoln? Every time you mention the words Abraham Lincoln, make sure that 5 of the letters are capitalized and all the other letters are uncapitalized.”

It's like, what is this? Sometimes we'll get customers telling us, “We really, really need to improve our score on IFEval.” What this means is that you have all these companies and researchers who, instead of focusing on real-world progress, are just optimizing for these silly SAT-style benchmarks.

One of the things that we really want to do is think about ways to educate the industry, think about ways of publishing on our own, and think about ways of steering the industry into hopefully a better direction. I think that's one big thing that we're really excited about, and it could be really big in the next 5 years.

Elad Gil

Sarah brought up earlier that everybody wants high-quality data. What does that mean? How do you think about that? Can you tell us a little bit more about your thoughts on that?

Edwin Chen

Let's say you wanted to train a model to write an 8-line poem about the moon. The way most companies think about it is, “Let's just hire a bunch of people from Craigslist or through some recruiting agency, and ask them to write poems.”

Then the way they think about quality is: Is this a poem? Is it 8 lines? Does it contain the word moon? If so, I hit these 3 checkboxes. So, sure, this is a great poem because it follows all these instructions.

But the reality is you get these terrible poems. Sure, it's 8 lines and has the word moon, but they feel like they're written by kids in high school. Other companies will say, “These people on Craigslist don't have any poetry experience, so what I'm going to do instead is hire a bunch of people with PhDs in English literature.”

But this is also terrible. A lot of PhDs are actually not good writers or poets. Think of Hemingway or Emily Dickinson: They definitely didn't have a PhD. I don't think they even completed college.

One of the things I will say is that I went to MIT. I think, Elad, you went there too. A lot of people I knew from MIT who graduated with a CS degree are terrible coders.

So we think about quality completely differently. What we want isn't poetry that checks the 3 boxes and uses some complicated language. We want the type of poetry that Nobel Prize laureates would write.

We want to recognize that poetry is actually really subjective and rich. Maybe one poem is a haiku about moonlight on water, another poem has a lot of internal rhyme and meter, and another one focuses on the emotions behind the moon rising at night.

You actually want to capture that there are thousands of ways to write a poem about the moon. There isn't a single correct way, and each one gives you all these different insights into language, imagery, and poetry.

It's not just poetry; it's math. There are probably 1,000 ways to prove the Pythagorean theorem. The difference is that when you think about quality the wrong way, you get commodity data that optimizes for things like inter-rater agreement and checking boxes off some list.

One of the things that we try to teach all of our customers is that high-quality data actually really embraces human intelligence and creativity. When you train models on this richer data, they don't just learn to follow instructions. They really learn all these deeper patterns about all the stuff that makes language and the world really compelling and meaningful.

A lot of companies just throw humans at the problem and think that you can get good data that way. But you really need to think about quality from first principles and what it means. You need a lot of technology to identify that these are amazing poems, these are creative math problems, and these are games and web apps that are beautiful and fun to play, while these others are terrible to use.

You really need to build a lot of technology and think about quality in the right way. Otherwise, you're basically just scaling up mediocrity.

Sarah Guo

That sounds very domain-specific. So, in every domain, do you build a lens for what quality looks like along with your partners?

Edwin Chen

Yeah, I think we have holistic quality principles, but oftentimes there are differences by domain. So it’s a combination of both.

Elad Gil

I think we got all the core topics. Nice work on podcast number 2, Edwin, and thanks for doing this. Congrats on all the progress with the business.

Sarah Guo

Yeah, no, thanks so much for joining us.

Edwin Chen

Yeah, it was great meeting you guys.

Sarah Guo

Find us on Twitter at no prior pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way, you get a new episode every week. And sign up for emails or find transcripts for every episode at no-bers.com.

No Priors Ep. 124 | With SurgeAI Founder and CEO Edwin Chen | BidClub