[BidClub_]
20VC · · 66 分钟

Surge CEO兼联合创始人 Edwin Chen:零融资做到10亿美元以上营收

Harry StebbingsEdwin Chen

YouTube
TL;DR
  • Surge AI从2020年起步、零融资做到10亿美元营收。 GPT-3发布后,创始人Edwin用了“几周时间”亲自做出V1,随后发在博客上;公司第1个月就实现盈利,且没有销售团队。他拒绝以300亿美元、甚至1000亿美元出售公司:“我绝对不会以300亿美元、甚至1000亿美元出售……被收购就等于承认失败。”
  • Edwin认为,竞争对手本质上是“披着科技公司外衣的人力外包商”:靠简历筛选招来一批“人肉劳动力”,没有衡量或提升数据质量的办法。质量控制是真正的对抗性问题:CS专业毕业生中有一半“连代码都不会写”,而会写的“会想办法骗你”——把账号卖到海外、用LLM生成数据。
  • 他对AI进步瓶颈的排序是:数据质量第1,算力第2,算法第3。拿算力硬砸垃圾数据,只会带来“并不存在的进步”——各家实验室反复在6-12个月后发现,训练数据和评测数据都是垃圾,有些模型甚至比起点更差。LM Arena就是典型:投票者奖励emoji、加粗和篇幅,而排行榜第1的模型坚持认为Pope Francis还活着。
  • 合成数据被高估了:大量用合成数据训练的模型擅长“合成问题,而非真实问题”;客户告诉他,“1000到2000条真正高质量的人类数据……价值高于1000万条合成数据”。Surge的大量工作,正是在清理合成数据造成的损害。
  • Scale被收购后,Surge迎来了一波兴趣。 在顶尖研究员圈子里,Surge是“这个领域最大、最好的公司”早已是公开的秘密;出于历史原因仍在使用Scale的团队带来了“巨大的兴趣浪潮”——这也呼应了Handshake的Garrett告诉Harry的说法:Scale客户迁移形成了“潮汐般的浪潮”。
  • 整个故事背后的效率逻辑是:在Google、Facebook、Twitter,90%的人都在做没用的问题,这些工作只是为了让VP印象更好;因此,规模只有其1/10的公司可以快10倍。他相信100倍工程师确实存在——速度、想法、工作强度各提升2-3倍,会议减少2-3倍;AI“会不成比例地利好本来就是10倍工程师的人”,而一家由单人创办、营收10亿美元的公司“总有一天会出现”。
  • 快问快答中的预测包括:如果说的是自动化普通工程师的工作,时间点是2028年;如果说的是治愈癌症,时间点是2038年;最大的模型提供商可能尚未成立,因为“我们距离AGI还只有2%或5%”;多个前沿AGI将以不同个性并存;而今天的基准测试作弊,本质上已经是纸夹最大化问题,模型能力越强,风险越大。
摘要 · 为研究而整理的核心内容

1. 零融资、10亿美元营收——先做MVP的打法

  • 过程几乎平淡得反常:Edwin作为ML工程师多年研究数据问题,已经“有了非常清晰的愿景”,所以他没有招10名工程师,也没有融资“1000万美元、2000万美元或3000万美元”,而是在2020年GPT-3发布后“几周时间亲自做出了V1”,发在博客上,结果“市场其实早就有巨大的数据需求”。
  • 需求爆发后为什么不融资?“融资并不能帮我们解决任何问题。”他们很幸运,第1个月就实现盈利。他也明确拒绝组建销售团队:希望客户“正是因为理解高质量数据的价值”才来购买,因为早期客户会塑造产品——而不是靠团队给1万名潜在客户群发邮件找来的客户。
  • 放到今天,有了Lovable以及很可能还有Replit,他认为“90%到95%的创业公司”没有理由在MVP之前融资,硬件和少数重资产类别除外。他如果能回到创业第1天,会告诉自己:“始终专注于能带来10倍改善的事情,而不是担心那些10%的现实问题。”

2. 大科技公司90%的人都在做无用功

  • 他在Google、Facebook和Twitter形成的关键观察是:如果神奇地移除那90%不做有趣问题的人,你会得到一家规模只有1/10、速度快10倍、产品好10倍的公司——面试更少、会议更少,没有“为了更新而更新”,人才密度更高,想法传播更快。
  • 原因在于,大公司的优先级“与最终客户脱节”——你做事情是为了给VP留下印象、争取升职。他的归谬是:改进内部工具→员工效率提升5%→因为他们有10%-20%的时间在面试→因为组织在“为了扩张而扩张”。很多管理者真正的目标“是告诉朋友,自己是一个千人组织的VP”。
  • 招聘筛选也由此而来:强候选人会反过来面试他,追问产品——“为什么不改进这些东西……你们为什么不做这个”;而帝国建设者的典型问题是:“如果我加入,我能不能再招20个人来支持我?”
  • 他实行零次一对一会议,还会展示自己几乎空白的Calendly:“如果你每周开一次一对一会议,几乎可以说这是一个负面信号,因为这意味着你根本不知道这些人的情况。”至于人均营收是否已经成为硅谷的新指标,他给出的诚实保留是:“我可以相信它。我也希望它是真的。但我不知道现在是否真的如此。”

3. 100倍工程师确实存在,而且AI偏爱他们

  • 他对这个神话般角色的算法是:有些人写代码快2-3倍,想法好2-3倍,工作更拼2-3倍,参加的会议少2-3倍——“2到3倍往往其实还是低估……把这些因素全部乘起来,确实能到100倍。”
  • AI会如何改变分布?它主要消除苦差事;而“优秀的人有太多想法,只是没有时间实现”,所以它“会不成比例地利好本来就是10倍工程师的人”。
  • 单人10亿美元公司是否存在?“我绝对相信这家公司总有一天会出现。”如今已经有单人创业公司做到1000万美元营收,而AI效率“提升100倍”后,就能走到那一步。

4. 竞争对手是“披着科技公司外衣的人力外包商”

  • Harry追问,究竟是竞争对手管理不善,还是Surge格外出色?答案是:“两者都是。”很多同行“根本没有任何技术”:没有衡量数据质量的办法,没有提升数据质量的办法,有时甚至连工人平台都没有。他们按简历筛人——“只要有PhD,就直接录用”——然后把人而不是数据交给前沿实验室,因此无法对质量算法或工具变化做A/B测试。
  • 质量控制的对抗性远超人们的想象:“我毕业于MIT,但CS专业毕业生里有一半连代码都不会写。”而那些会写的“其实会想办法骗你——把账号卖给第三世界国家的人,用LLM替你生成数据”。
  • 对于简单粗暴地堆人的做法,结果是:“我认识的、试图这样做的团队,最后实际上在自己毫无察觉的情况下,比其他所有人慢10倍。”

5. 创业创伤:Twitter的数据管线由Craigslist上招来的2个人搭建

  • Twitter当时做情感分类只需要1万条标注推文,但人类数据系统“真的就只是Craigslist上招来的2个人,朝9晚5工作”——等了1个月他们才入职,又花1个月在电子表格里做标注,最后产出的数据一塌糊涂:他们不懂俚语,把“她真是个坏[__]……”标成负面,实际上这句话非常正面。Edwin最后花了1周亲自标推文,因为那样更快。
  • 更深层的需求没有被满足:用点击和转发训练按时间顺序迭代的推荐算法,制造了“极其负面的反馈循环”——满屏都是比基尼女孩和“10种最可怕皮肤病”的榜单文章。他希望标注员能依据产品原则进行判断;而如果Twitter连情感分类都做不好,就更不可能在所需规模上获得更复杂、更高质量的数据。GPT-3发布后,Edwin看到“整个行业明显正在朝那个方向移动”,Surge于是于2020年成立。

6. 融资是一场身份游戏;风险本身才是重点

  • 他对创业文化最尖锐的批评是:“人们只是为了融资而融资……目标是告诉所有朋友自己融了1000万美元”,再靠一篇“融了1000万美元”的新闻上头条——每周改方向、在Twitter上发表热门观点、参加VC晚宴,直到某个话题获得1000次转发。“你没有承担任何风险,只是想赚一笔快钱。”
  • 但他确实认为,创始人应该追逐真正属于自己的想法:商品化创意可以做成“一家还不错的中型公司”,但具有时代意义的公司“真的应该建立在几乎只有你能想到的创意之上”。
  • 内部执行的抓手是质量——每个新员工都会被告知:“质量是最重要的事情……如果因此必须延期,如果因此必须拒绝一个项目,那就这样。”至于在压力下是否降低招聘门槛,他认为那名急着招进来的人“可能正在做一个根本没人关心的功能”。

7. 1000亿美元也不卖;生意本身就是奖品

  • Harry把报价从300亿美元一路加到500亿美元,仍然撞上了明确的上限:“我绝对不会以300亿美元、甚至1000亿美元出售。我已经拥有想要的一切。我们是盈利的,我完全掌控着自己的命运……被收购会严重限制我,也等于承认失败。”
  • 他真正追求的是AGI——“小时候你真的会梦想造出能做这些神奇事情的AI”——以及实验室发布模型后“最先做的事情之一就是联系我,说:如果没有你,我们不可能做到”。这也解释了他凌晨2-3点接客户电话时的状态:“我们的模型快崩了,早上6点前需要一大批数据”——“没有什么比知道我们能在接下来几小时交付1万条数据更让我开心”。但他也提醒,不要把工作时长和价值混为一谈——“最好的想法往往是在我四处走路时出现的。”
  • 他坦言自己的盲区,反而十分坦率:“我没法告诉你EBIT是什么……它和营收、利润、净利率之间的区别,我真的不知道这些术语。”他理想中的北极星指标是:模型是否在基础层面取得进展,其中有多少进展归因于Surge;目前最接近的代理指标,是平台上项目的多样性和复杂度。

8. ChatGPT是拐点;Scale被收购带来兴趣

  • 公司从第1个月起增长就“非常非常强劲”,但ChatGPT是拐点,因为人们突然看到了人类数据和RLHF的惊人价值。
  • 对于Scale被收购,他说:“在很多顶尖研究员那里,这早已是一个公开的秘密:大家都知道我们是这个领域最大、最好的公司。”那些出于“历史原因”使用Scale的团队,随后带来“巨大的兴趣浪潮”。Harry也作了印证:Handshake的Garrett告诉他,面对“潮汐般涌来的Scale客户”,自己“整晚都没睡”。
  • 转投客户的共同经历是:他们在其他地方“会花几个月试图提升一些非常基础的数据质量,看起来好像改善了1个月,但很快又退化”;而Surge坚持交付“你在其他任何地方都拿不到的数据”。

9. 数据质量胜过算力;LM Arena就是标题党

  • 被问到AI进步的瓶颈排序时,他给出的答案是:“数据质量第1,其次是算力,再其次是算法。”他完全否定“算力可以解决一切”的前提:如果没有正确的数据和目标,“你只会掉进一个陷阱,以为自己取得了其实并不存在的进步”——实验室一次次告诉他,6-12个月后才发现“训练数据是垃圾,评测数据也是垃圾”,有些模型甚至比起点更差。
  • LM Arena的机制是:参与者根本不核查事实——“一个回答可能完全是在胡说,但因为里面有一个emoji、几个加粗词,人们就会投给它。”爬榜最容易的方法是写更长的回答;排行榜第1的模型被问到教皇何时去世时,坚持说Pope Francis还活着,并把搜索结果斥为“谣言和错误信息”。爬榜的实验室其实是在“训练模型生成更好的标题党”。
  • 对于Grok在基准测试中的胜利,他引用了源头说法:在Grok 4直播中,“Elon Musk本人”说模型非常擅长“家庭作业题”——“相当于让它们特别擅长SAT题,但不擅长人们真正面对的问题”。不过,他仍然认为xAI本身非常令人印象深刻:“晚上11点我给他们发私信……他们还在办公室,而且身后有一大群人。”
  • 他对安全问题的快问快答,最终又回到这里:认为AI安全被夸大的人“忽视了纸夹最大化问题”——朝错误目标意外优化,已经通过基准测试作弊发生了;“未来当模型更强大时……如果它真的在为一家万亿美元公司编写代码,后果可能严重得多。”

10. 合成数据被高估,PhD不够用,前沿格局远未定型

  • 面对“合成数据是否会取代人类数据”的老问题,他说合成数据让模型“擅长合成问题,而非真实问题”——一些公司投入1年后,现在正在“大量弃用”,并告诉Surge,“1000到2000条真正高质量的人类数据……价值实际上高于1000万条合成数据”,因为模型会“在极窄的相似性范围内坍缩”。他举出的最新例子是:某个2025年的前沿模型在回答中随机输出俄语和印地语字符——“任何一个二年级学生都能看出的错误”,这正是“始终需要外部价值体系作为安全护栏”的原因。
  • 关于PhD级数据,Surge拥有“Harvard教授、Stanford PhD学生和Princeton计算机科学理论学者”——“我们单日合作的PhD数量,比Google、Meta和Microsoft加起来还多得多。”但“我认识的计算机科学PhD里有80%代码写得很烂……而Ernest Hemingway也没有PhD”。要找到最顶尖的1%-2%,既需要街头智慧,也需要技术;他的类比是:“Vimeo上有很多所谓的高质量视频,但没有任何算法,所以最终YouTube上的视频质量高得多。”
  • 有2个前沿判断值得记录:他现在预计会出现多家前沿AGI公司——“Claude在编程方面非常非常强……在企业场景和指令遵循上很强,而ChatGPT更偏向消费者……Grok愿意更具挑衅性”——就像诗人和数学家一样,不会有唯一的最强者。最大的模型提供商也可能尚未成立:“我们距离AGI还只有2%或5%……这就像10年前问Google会不会是最终的搜索引擎。”
  • 时间表上的关键判断是:如果说的是“自动化普通工程师的工作”,时间点是2028年;如果说的是“治愈癌症”,时间点是2038年,因为现实世界的实验需要时间。他反驳Robinhood和Benioff关于代码自动生成50%的说法:“如果你真的在解决有意义的问题,我不认为今天的模型能写出50%的代码”;他“绝对”相信10年内GDP能提升10%;而Google的难题在于,Sundar“必须愿意让广告收入在短期内承受冲击,去打造更好的东西——这真的非常难”。
Edwin Chen

I think a lot of the other companies in our space are just not technology companies at the end of the day. They are either body shops, or they are body shops masquerading as technology companies. One of the things that we simply tell everybody when they first join is that quality is the most important thing. It’s more important than anything else.

I definitely want to sell for $30 billion or even $100 billion. If you think about us as a company, I already have everything I want. We’re profitable, I have complete control of our destiny, and so I’m really lucky to already have all the resources I want to do anything that I want.

Harry Stebbings

Edwin, dude, I’m so looking forward to this. I am the biggest fan of your business from afar, which makes me feel incredibly weird because we haven’t met before, which means I’m basically a stalker. But thank you for joining me.

Edwin Chen

Yeah, thanks for having me. It’s wonderful being here today.

1. Why 90% of Big Tech Is Wasting Time on Useless Problems

Harry Stebbings

I wanted to break the show into 2 different parts. The first part is the story of this incredible rise, and then the second part is really assessing the future of data labeling and taking a more analytical approach.

If we start with the story itself, before the founding of Surge, you said to me that 90% of the people, while you were working at Google, Facebook, and Twitter, were working on useless problems. I thought that was a very interesting place to start. Why were they working on useless problems, and what did it teach you about efficiency seeing that?

Edwin Chen

Yeah. I think the biggest lesson for me was that you can build a completely different kind of company with 10% of the resources and 10% of the people, but you’re still moving 10 times faster and building a 10 times better product.

Imagine you could just magically remove the 90% of people who aren’t working on interesting problems. What would happen then? Well, if you have a company that’s one-tenth the size, you don’t need to hire as many people, so you spend less time interviewing. You spend less time in meetings. You spend less time giving people updates for the sake of updates.

And if it’s one-tenth the size, that means everybody has a better view of what’s going on around the company because there isn’t all this clutter masking the important stuff. Because the talent density is higher and the teams are smaller, that means the communication is a lot better, the iteration speed is a lot higher, and better ideas just percolate around more quickly.

Harry Stebbings

Can I ask, prioritization is slightly ambiguous according to different people. Everyone feels that their project is important and more important than someone else’s. How do you determine priorities within a company and determine what matters versus what doesn’t?

Edwin Chen

Yeah. I think a big thing about being small is that when you’re smaller, I and other people around the company just have a much better view into the customer problems themselves and what everybody’s working on.

At these bigger companies, a lot of your priorities, a lot of the things that you’re building, are simply things that you’re building to impress someone. “I need to impress my VP. I need to impress my manager. I need to impress my director so that I can get promoted.” You’re not really building things or prioritizing things because they’re good for the end customer or the end product.

It’s more like, “Okay, I have this priority to improve an internal tool.” Why are you improving the internal tool? “Well, it will make people 5% more productive.” Why do I want them to be 5% more productive? Because they’re spending 10% or 20% of their time interviewing. Why are they interviewing? Because they’re growing for the sake of growing.

It just leads to this perpetual cycle where a lot of your priorities are divorced from the end customer and the end product. They’re almost priorities just for the sake of internal company machinery. So, yeah, I think it’s very different.

Harry Stebbings

What do you think no one knows about working within these big, incredibly hailed companies that they should know?

Edwin Chen

I think one of the things that people don’t realize, again from the outside, is how much of what you’re building is for this internal company machinery, and how much of the internal company machinery is simply because a lot of people within these organizations, their goal isn’t to build a product. Their goal is to tell their friends they’re a VP of a 1,000-person organization, because that sounds impressive.

Their goal is to think, “Okay, so how do I grow my organization even faster? How do I find more teams that I can hire? How do I have these monthly performance reviews?” Again, now that I’ve built this 1,000-person organization, I need to prove to my VP or my CEO that the 1,000-person organization I’m building is efficient and useful.

A lot of the work that goes on in these large companies is simply to perpetuate and grow even further a lot of this very, very big company machinery that exists purely for its own sake.

Harry Stebbings

When you’re hiring, how do you determine between managers who like to brainstorm and tell their friends that they have 1,000-person organizations and are very powerful and important, versus doers—those who execute, work, and complete tasks? How do you determine the 2, and are there very clear differences?

Edwin Chen

Yeah. I think a big part of it actually just boils down to the kinds of questions they ask me.

Some people, when I’m interviewing them, will ask really interesting questions about our product. They’ll brainstorm about ideas to make our product even better: “I went to your webpage. Why don’t you improve these things? I tried signing up as a worker. Why did these things happen in the flow? I tried working on this project. What if you guys did this instead?”

Other people are like, “If I join in a year, will I be able to be a manager of a company? If I join, will I be able to hire 20 more people to support me?”

It boils down, a lot of the time, to the kinds of questions that people even have at the forefront of their minds.

2. How Surge Kills Meetings and Still Moves 10x Faster

Harry Stebbings

Can I ask you in terms of meeting cadence? I’m sorry for being granular, and I told you we’d go off schedule, but I’ve had Tobi on the show in the past, from Shopify, who’s obviously advocated for no meetings, given the ability to spend lifetimes in meetings that are quite pointless. How do you approach meeting policy, and what does and doesn’t belong in the organization?

Edwin Chen

Yeah, I’m a big fan of that. For example, personally, I actually have no one-on-one meetings. It’s kind of funny because oftentimes people will ask me, “How often do you meet with your reports? How often do you set aside time for these meetings?” I just don’t have them at all.

Oftentimes, I’ll give people my calendar, my Calendly, and they’re just surprised at how blank it is because I try to avoid filling my meetings all day. Sometimes when people join, they’ll be like, “Okay, I need to go and have one-on-one meetings with these 10 other people that I’m going to be operating with on a weekly basis.” That’s just what they’re used to when they come from Google or Facebook.

I ask them, “Why are you having these standing one-on-one weekly meetings? Do you not talk to them every day in Slack? Are you just unaware of what they’re doing?” It’s almost a negative sign if you’re having a one-on-one weekly meeting, because it means that you just don’t know what’s going on with these people. You’re almost waiting for your weekly meeting to raise interesting questions and interesting problems.

We’re pretty ruthless internally about killing meetings when they’re unnecessary.

Harry Stebbings

We mentioned the efficiency of small teams. Before we dive into Surge, one of the hot topics of the day is the future where billion-dollar companies will be built by single people. Do you agree with that vision of the future, or do you think it’s slightly overdramatized?

Edwin Chen

Yeah, I absolutely believe that that company will exist one day. I’ve always believed in 10x engineers, even 100x engineers, and already you have a lot of these single-person startups that are doing $10 million in revenue.

If AI is adding all this efficiency, then, yeah, I can definitely see this multiplying 100 times to get to this $1 billion single-person company.

3. 100x Engineers Are Real

Harry Stebbings

You can’t drop 100x engineer without me diving on it. We’ve been so focused for so many years on 10x engineers. What have been your biggest lessons on 100x engineers? Do they exist, actually, in reality? What are the signs? Talk to me about that.

Edwin Chen

Even today, you see how we are honestly so much more efficient than some of our peer companies, right? For that reason alone, you can already see the fact that 10x engineers or 100x engineers exist.

If you break it down, some people are simply 2 to 3 times better—2 to 3 times faster than anybody else. They just code faster. There are some people who simply have 2 to 3 times better ideas. There are people who simply work 2 to 3 times as hard. There are people who have 2 to 3 times fewer meetings. There are people who simply have ideas that other people can’t think of.

If you just multiply all these things together, 2 to 3 times is often actually an underestimate.

I know people who literally are 5 times more productive as coders than anybody else. Now add in all the AI efficiencies that you get. You can just multiply all those things out, and you get to 100.

Harry Stebbings

Do you think AI turns 10x engineers into 100x engineers, or average 1x engineers into 10x engineers, or maybe both today but definitely even more so in the future?

Edwin Chen

I tend to think of it like good people have so many ideas that they just don't have time to implement. If you think of AI today as something that isn't necessarily coming up with the greatest ideas, although it can, it often just removes a lot of the drudgery of your day-to-day work, a lot of your day-to-day coding. If you don't have to spend that time on the drudgery, but you just have these endless ideas bouncing around your head and AI helps you put them to paper, then I do think it disproportionately favors people who are already the 10x engineers.

Harry Stebbings

You mentioned the comparative efficiency in the landscape. Without naming names, a lot of people around you say—it’s not naming names, but a lot have raised a lot of money to get to a smaller stage than you are. If I were to push you into a camp, is that a result of you being phenomenally efficient, where you deserve credit, or have they, bluntly, been incredibly mismanaged and resource allocation has not been done well?

Edwin Chen

I think it's both. A lot of the other companies in our space are just not technology companies at the end of the day. They are either body shops, or they are body shops masquerading as technology companies.

Harry Stebbings

What do you mean by body shops and body shops masquerading as technology companies? I get it, but a lot of people criticize the space with this and say, “Oh, it's just labor camps.” So what do you mean by body shops or body shops masquerading as technology companies?

Edwin Chen

A lot of companies in the space don't have any technology. When I think about technology, it's that they don't have any way of measuring the quality of the data that they're producing, and they don't have any way of improving the quality of the data that they're producing. They are literally just body shops, in the sense that they sometimes literally have no technology at all. They don't have a platform where workers are doing work.

What they're doing is simply finding people; they're recruiting warm bodies. They're looking at résumés—anybody with a PhD, they'll just instantly hire them and then pass them along to the AI companies, to the frontier labs. Again, they have no technology. They have no way of measuring what any of these workers are doing, and they have no way of knowing if they're doing a good job or not.

So they have no way of doing things like, “Hey, what if I A/B-tested this algorithm for improving quality? What if I changed this method of allowing workers through? What if I tweaked your tools in order to change these questions around? Would that make the workers more efficient? Would that improve their quality, or would it actually make it worse?”

They just have no way of doing these things because, again, at the end of the day, what they're passing to their customers is just the body itself—the person—as opposed to the data. What that means is they just have no technology to measure anything.

Harry Stebbings

Do you think you have a fundamentally different business, then? Because you're all lumped in the same category. But if they're passing along a warm body and you're passing along data, it's a phenomenally different product, and it's monetized differently.

Edwin Chen

No. Yep, yep, yeah. Again, if I think about the way we think about it, it's maybe the following: We have always started out with quality of the data as our number-one principle. As a result, we need to build technology in order to measure that and improve it.

If I think about what goes wrong, it's that people often just don't realize how difficult quality control is. People often think that humans are smart, and so if you just throw a bunch of humans at the problem, you'll get good data. What we found is that that is completely untrue.

4. Why the Real Bottleneck in AI Isn’t Compute or Models

For example, I went to MIT, but I think half of the people who graduate with a CS degree can't even code. So it's a really challenging problem to detect high quality. Second, if you actually take the folks from MIT who can code, they're actually just going to try to cheat you. They're going to sell their accounts to somebody in a third-world country. They're going to try to use an LLM to generate the data for you. They're going to come up with all these crazy methods to cheat the system.

So it's also this really, really challenging problem to detect low quality. It's actually really adversarial. What we found is that when you want to get the highest-quality data to train LLMs that are already superintelligent, you actually need to build a ton of really sophisticated algorithms.

You can't just take warm bodies or try to improve your methods for résumé filtering, then throw people at the problem and get good data and results out of it. The teams I know who try this actually end up moving 10 times slower than anybody else without realizing it. Again, at the end of the day, I think it's all about the technology that we build to extract the highest-quality data possible, as opposed to just throwing bodies at it.

5. Founding Surge AI

Harry Stebbings

Okay. We mentioned before the background you had prior to obviously being at companies like Google, Facebook, and Twitter, and then you said there about the focus on data quality. Can you take me to the founding moment for you leaving the last company and deciding that you were going to go all in on Surge?

Edwin Chen

I used to work as an ML engineer at a bunch of data companies, and the problem I just kept running into was that it kept being impossible to get the data that we needed to train our models. I can give an example. I used to work on our ad search and ad systems at Twitter, and one of the first things I wanted to do was build a sentiment classifier.

It's a super-simple problem. All you need is 10,000 tweets labeled as positive or negative to train your models. But our human data system at the time was literally just 2 people we hired off Craigslist, working 9 to 5. Even just in order to get started, we had to wait a month. Then we had to wait another month for them to label the tweets inside a spreadsheet because the tools were just terrible.

When we finally got the data back, it was completely junk. They didn't understand slang, like, “She's such a bad [__].” They were actually labeling this negative when it's actually really positive. They didn't understand hashtags and all these other aspects of the tweets. I ended up just spending a week labeling tweets myself because that was so much faster and better.

At the same time, this was actually really simple stuff. But the bigger problem we wanted to solve was: How do we optimize our ML systems for the right objectives, and how do we build feeds that are engaging in a positive way for users? Think again about Twitter. This was the old days when it was a purely chronological timeline, and one of the things we wanted to do was make it easier for users to discover the tweets that they really cared about.

The question was, how do we train our recommendation algorithms? The obvious choice was clicks and retweets: You just train your algorithms to produce as many clicks and retweets as possible. But the problem is, we tried doing these things, and it turns out to be this incredibly negative feedback loop.

Once you optimize for clicks, the most clickbait content starts rising up to the top. You get lots of racy content, lots of girls in bikinis, lots of listicles about 10 horrifying skin diseases, and so on. We wanted to train all of our models on these deeper principles instead. We'd ask our human raters to label tweets and recommendations with product principles, like whether this was a top voice connecting somebody with their interests, or whether somebody just had this really interesting insight on a particular topic.

If we couldn't even get simple sentiment analysis right—labeling whether a tweet was positive or negative—we definitely couldn't get this more complex data at the quality scale that we needed. If you think about it, we basically started Surge in 2020, right after the launch of GPT-3, and I think it really is because there was just so much more that you could see the industry moving toward. If we really wanted to progress it in really, really big ways, we needed a different kind of data solution for the industry.

Harry Stebbings

Okay. So you realize this data problem in 2020, you leave Twitter. What happens then? You go heads down into product build for several months. You go about recruiting the first team members. Can you just take me to the build? I mean, 2020, dude. It's not that long ago: $1 billion in revenue, and you started in 2020.

Edwin Chen

Yes. The way it worked was, I've always been a really big fan of MVPs, and so I literally just built myself an MVP in a couple of weeks. The really nice thing was, again, I had worked in the space for a really long time, so I already had a very clear vision of what I wanted to build. I didn't feel like I needed to go out and hire 10 engineers in order to build a product, and I didn't feel like I needed to go out and raise $10 million, $20 million, or $30 million in order to hire more people.

Again, I just wanted to build it myself and talk to customers myself. That's what I did. I think I had already built the V1 in a couple of weeks. I posted about it on my blog and told people I met about it, and there was actually this giant demand for the data already. I think we were very lucky early on.

Harry Stebbings

You posted on a blog, you got some demand. You said that, with the MVP, you decided you'd build that first and not raise money. The traditional thinking in the Valley is, “I need money because I need money to build.” Why do you think that's maybe wrong, and how would you change or advise founders differently?

Edwin Chen

I think one of the things that's always driven me crazy about Silicon Valley is that it really is just a status game for most people. People are just raising for the sake of raising. Their goal isn't to build some great product that solves a problem they fundamentally believe in. Their goal really is to tell all their friends that they raised $10 million and get ahead of that crunch.

I have a lot of friends who've worked at Google for 10 years. When they think about starting a company, they often tell me they don't even have a problem they want to solve. They're just bored and want to try something new. At the same time, they can definitely pay their own salaries for a couple of months, but the first thing they tell me is that they're going to go out and raise some money.

They might try talking to some users and building an MVP, but the only reason they do that is to check off a checkbox on a YC application. Then they'll constantly pivot around random ideas until they get something that happens to get a little bit of traction and sounds impressive to VCs. They spend all their time tweeting hot takes, networking, and going to all these VC dinners, and it's all just so they can get this headline about raising $10 million.

I really think that people's first instinct should instead be to find some big idea that they fundamentally believe in and that could change the world. I don't really care why they believe in it. It could be because they have a lot of experience in the space, or because they talked to a bunch of users, but it really has to be something they believe in and would double down on for the next few years.

Startups are all about big risks, right? You have to believe in something enough that you're going to take a risk building it. If all you're doing is jumping around from idea to idea every week until you land on something that gets you 1,000 retweets, you're not taking any risks. You're just somebody looking to make a quick buck.

Harry Stebbings

I have so many questions off the back of that. You mentioned loving the MVP and the ease of doing so. Given the tooling that we have today, the ease of building an MVP has never been greater. Do you think there's any excuse for going out to raise now without an MVP, given Lovable and likely Replit? It's just so much easier.

6. Will Synthetic Data Kill Human Labelling?

Edwin Chen

For 90% of companies, no. There are some companies where you actually do need a lot of capital in order to build hardware or whatever it is for a couple of years. You really need a lot of investment before you can get to your actual MVP. But for 90% to 95% of the products out there, and for 90% to 95% of the startups people are building, no. Just go out and build your MVP and see if it gets any traction.

Harry Stebbings

You said something about the inherent risk that you take on when you start a company. Do you believe in the advice that you should only pursue ideas that only you can do—in other words, that the idea is specifically tailored to you and not everyone could solve that problem? Or do you think that's bullshit and it's actually about execution?

Edwin Chen

I actually do believe in it. If you think about the idea of a startup as something where you can take big risks, where you can build something that nobody else can and you're willing to go all out to create something that literally nobody else could, it does have to be something unique to you.

Otherwise, sure, you can get to a decent medium-sized company with a commodity idea. But if you really want to go big, if you really want to build a generational, foundational company, I think it really should be based on an idea that's almost unique to you.

Harry Stebbings

You said that people may gain value or self-worth from raising big amounts and going to conferences. That is how most people gain self-worth. When you think about where you derive your own self-worth from—sorry to be personal, but given that yours is clearly not that—how do you think about where you get self-worth and self-value from?

Edwin Chen

I think it's kind of funny. If I think about the things that have made me happiest in the past few years, I can think of 2 things off the top of my head. One is that sometimes, whenever our customers launch their next big model, one of the first things they'll do is reach out to me and say, “Hey, just wanted to send you a note that we couldn't have done this without you.”

I think that's amazing to hear. How often do you get to play a role in building some of the most important technology of our time, and then, right after the launch, these very top people who are very busy have one of their first thoughts be to thank you because of how critical you were to the operation? I just think that's so cool. That is one of the things I often think about.

The other thing I often think about is that, in many ways, Surge is an embodiment of me and my interests. What I've always loved doing is analyzing data and figuring out how to use that data to make models better or to make products better.

Every now and then, when I get the chance to write an analysis myself of the latest frontier model, or I get to read some of the analyses that our internal employees are creating based on the data we're providing, I think it's so cool. A lot of the data we're providing is so insightful, and it helps people build models in ways that they just wouldn't know how to otherwise. I think it's really cool to help these insights emerge into the world.

Harry Stebbings

Can I ask, going back to that story, then? You built the MVP, you posted it, and then you said, very nonchalantly, “Luckily, people came and people liked it.” What did that look like? How did the initial demand come to you?

Edwin Chen

Sorry, I think I say it nonchalantly because it felt very nonchalant. What would end up happening is that I would find all these people who were desperate for a lot of really high-quality data. They would email me with their request, or we would jump on a live meeting and get started.

It might take a week or a couple of weeks to negotiate some sort of SOW or contract, just because a lot of this does have to live within the confines of their company. But I think we were really lucky. I had a lot of experience in this space, working with ML engineers and research scientists, and I understood the ways they wanted to get data and look at it. Things just moved very, very quickly.

Harry Stebbings

In the early days, everyone else was acquiring the supply side of talent, correct? All the other people that compete in the space were acquiring that talent supply, and you weren't acquiring the talent supply; you were building the product.

Edwin Chen

It was both, because obviously we need talent supply in order to make our product work. But it was less about that. There are some companies in this space that think of it as a pure supply problem, and they don't give any consideration to the technology.

How do you identify these people? How do you make sure that they're doing good work? How do you remove the bad-quality work? They're just literally not thinking about any of the technology aspects at all. They're also not thinking about the product at all. How do you present the data to the customers?

One of our principles—one of the principles that I've always had, even before Surge, when I was just an ML engineer or data scientist—is what we call this visceral understanding of the data. I really just want you to go in, get your hands dirty, and look at the data.

Historically, a lot of ML engineers don't take the time to look at the data. Maybe that's because the data just isn't all that interesting. When all you're doing is drawing bounding boxes around cars, sure, I don't need to look at 1,000 bounding boxes. But when what you're doing is creating poetry, creating mathematical equations, or creating new research, you want to get your hands dirty with the data to see what it is that you're producing and what you're teaching your models.

I think it's really important to have this aspect of viscerally understanding the data that you're getting.

Harry Stebbings

So we were doing both—building the product and acquiring the talent supply—in unison. Fantastic. What did we end the first year at? Did we have immediate product-market fit?

Edwin Chen

Yeah, I think it was very, very obvious that there was just huge demand for this product, and there was so much more that we could be doing.

7. “No Sales Team, No PR, No BS”

Harry Stebbings

So, Edwin, when there's huge demand for your product, this is even more so the time when everyone goes, “Now raise money, hire CS teams, hire sales teams, hire big.” Why did you not raise money then? I get it at the start when you didn't want to do what everyone else did. Why not raise money when it was a hair-on-fire problem and you had so many people calling you?

Edwin Chen

I would say there was nothing that raising money would help us with. Again, we were very lucky to be profitable from month one, and so we didn't need the money. We didn't need a sales team. I didn't actually want a sales team going out and selling our product. I wanted people to buy from us precisely because they understood the value of high-quality data. They saw all the gains that our data was producing.

I didn't want them to buy from us simply because they heard about us in some TechCrunch article, because that would almost put them at odds with the kind of product that we were building. One of the things that I think is actually really important is that, especially early on, you want customers who believe in your product and not people who are simply giving you a little bit of money.

Your early customers will shape the kind of product that you're building because you're building for them. You're building for their needs. They're giving a lot of really great feedback, and so you almost want customers who share the same overall vision. That was actually very important for us. I didn't want sales teams who would email 10,000 people and be like, “Hey, any thoughts on getting good data?” It was just very, very counter to the kind of product that we wanted to build.

Harry Stebbings

How do you think about what you just said there in terms of building with your customers, being so close to them and letting them shape your product, but then also not doing the Henry Ford of building a faster horse, and not building a product that, bluntly, isn't relevant for a wider audience base, where you really just tie yourself into a few small clients?

Edwin Chen

I think this is where we actually have a very strong vision of what a product should be. Going back to what I said earlier, most companies in this space—and maybe also at large—don't have product principles that they try to adhere to. We had very strong product principles from the start.

We wanted to focus on quality above all else. If we ever thought that we couldn't give the quality that we wanted, we would just say no. That's as opposed to these other companies where they're almost desperate and racing around, just trying to get any traction that they can. They're trying to prove to the VCs that their numbers are always going up. They're almost focused on getting $10, $100, $1,000, wherever they can.

As soon as some customer comes to them, even if that customer is counter to the kind of product that they're building, if they're offering money, they'll just say, “Sure, I'll do it,” just because they'll give them another logo for a website, another case study to show another customer, or another talking point with their VCs. I think we're very lucky not to have to worry about that because we could build for the long-term vision we had, as opposed to pivoting every few months. We just wanted to double down on the idea that we actually believed in.

Harry Stebbings

Is there a time when you let quality slip in any area of the company? With hindsight, what did you learn from that?

Edwin Chen

I think we've never let quality slip. It's such a principle ingrained into everybody at the company. One of the things that we simply tell everybody when they first join is that quality is the most important thing. It's more important than anything else.

If you have to make a deadline slip because, for whatever reason, you don't think the quality is there, or if we have to say no to a project because we just can't handle it right now—we can generally handle a lot of things—we just want to ingrain this principle that it is okay to say no. It is okay to let other things slip just because we care about quality.

Harry Stebbings

Most founders have a challenge where they need to hire now, but they haven't found the perfect person, and so they hire a 7 out of 10. They let the quality bar slip because they need someone in the role. How do you think about that, and what would you advise them?

Edwin Chen

I think the funny thing is that I've been at all of these other companies. Oftentimes people say, “My hair is on fire and I really need this engineer, so I know they don't meet the bar. I'm going to lower the bar to hire them.”

That engineer is probably building a feature that nobody cares about. They're building an internal tool to improve the productivity of everybody around the company by 2%, while at the same time having so many meetings about it that they take up 5% or 10% of their time just talking about the feature. A lot of the things that people hire for just actually aren't all that important.

When you don't feel like you have to hire for the sake of hiring, when you have the mentality that if your company only grows by 10% or even 0%, that's actually positive, I think that's valuable. People right now have this view that if someone tells you, “My engineering organization only grew by 2% this year,” your initial reaction is going to be, “Okay, you guys must not be doing well, right?”

There's this negative incentive where people feel like they need to hire just in order to prove to other people that their business is doing well.

Harry Stebbings

Do you think now we're in an opposite world to that, though, where you see the reductions in force from, say, Microsoft, and you see better performance than ever from them on revenue per head? Do you think we're now seeing the counterbalance of that, which is the desire to be the smallest team, the fastest team to X ARR, and the smallest team to do it? Now revenue per head is the most important metric.

Edwin Chen

I honestly don't pay enough attention to these kinds of Silicon Valley Twitter discussions to have a sense of whether this mentality is becoming more pervasive. I can believe in it. I can hope for it. I don't know if it's true right now.

Harry Stebbings

Do you worry that by not being so ingrained in social, you miss out on certain elements that are important to be in, or do you think that purity of mind that you get is really so valuable?

Edwin Chen

It's kind of funny because I used to work at Twitter, and I loved Twitter back in the heyday, but I actually am really glad that I'm not surrounded by default ways of Silicon Valley thinking.

Every now and then, if something is important enough—maybe there's some big new product that's really cool, or some really interesting new research paper—it'll be big enough that, even though I'm not monitoring Twitter every day, it will just reach me in some other way. One or more employees will post it in our Slack channel, or somebody will email it to me.

The really important stuff will manage to percolate to me in other ways. I'm really glad that I'm not worrying about what people are saying about us on Twitter.

Harry Stebbings

I love that, especially given the irony of being at Twitter for a number of years. I do have to ask: the first year ends. What did you end revenue at in the first year?

Edwin Chen

Let's just say we've been doing really, really well from the start.

Harry Stebbings

I totally get it. Again, you said publicly that you're at a billion in revenue now. Did it look like relatively even growth, or were there elements where it was much more accelerated than others? I'm just intrigued. Say whatever you feel comfortable with.

Edwin Chen

We've always been very, very successful from literally month one. Things definitely hit an inflection point with ChatGPT because I think people just saw how incredibly valuable human data in RLHF was.

ChatGPT was definitely an inflection point for us, but even before that, we had very, very strong growth.

Harry Stebbings

Okay, I love that. So post-ChatGPT, you really saw the inflection point. Another one that I guess is probably quite an important one is Scale, obviously, selling and the movement of customers away. How did the world change for you with the Scale acquisition?

Edwin Chen

It's interesting because I think it was an open secret that a lot of top researchers already knew who we were. They already knew that we were the biggest and the best in the space, even though we'd been pretty under the radar. Most people were already working with us.

There were a lot of teams who were using Scale for legacy reasons, or they just didn't happen to know about us, so we've gained a lot of new interest from them, too. I think the more interesting thing has been seeing how we've opened their eyes to what really amazing, high-quality data can actually look like.

A lot of them have tried getting human data from other teams, and they tell us it's been this slog. They'll spend months trying to improve the data quality for really basic stuff, and it'll look like it's better for a month, but then it'll quickly regress.

We have this concept where we just want to get started immediately. We want to show them really, really high-quality data immediately. But then we also want to—one of the big concepts for us as a company is that we always want to be producing data that you simply couldn't get anywhere else.

There’s so much richness and complexity in the types of things that we do that we just want to open up new avenues of research and new types of products. I think a lot of these new companies or these new teams who have been coming to us—it helps. It’s just been a breath of fresh air for them.

Harry Stebbings

I spoke to Garrett at Handshake right after the acquisition. He said, “I’m just staying up all night. There’s just a tidal wave of Scale customers moving to us.” Did you have the same experience, in terms of that tidal shift in customer demand moving to you, as well as the realization that you mentioned there?

Edwin Chen

Yep. I would say I’m pretty sure that a lot of these other companies, at the end of the day, people want high-quality data and they don’t want to be working with body shops. I think we’ve seen a massive wave of interest because the space is really large, and there are a lot of teams who are still using Scale for legacy reasons.

At the end of the day, we were already the biggest investor in the space. So even when there were teams at some of these larger companies who weren’t working with us already, they knew who to turn to.

Harry Stebbings

Do you think everything has a price, Edwin?

Edwin Chen

I think some people have a price, but I think we don’t.

Harry Stebbings

You said you wouldn’t sell to Zuck for $30 billion. Would you sell for $50 billion?

Edwin Chen

No. I definitely wouldn’t sell for $30 billion or even $100 billion. If you think about us as a company, I already have everything I want. We’re profitable, and I have complete control over our destiny. I’m really lucky to already have all the resources I want to do anything that I want, and there aren’t many companies who can say that.

Harry Stebbings

What are you doing this for? You’re such a dude. I’ve interviewed a thousand founders and, in the nicest way, I’ve almost never met a founder like you. In a nice way, it’s really special. But with a pure mindset like you have, what are you doing it for, then? To build a business that you can pass on to the next generations? To build a legacy? What is it for you?

Edwin Chen

I think it really is to help achieve AGI. If you think about every—what do kids dream of? When you’re a kid, you literally dream of building AI that can do all these amazing things. Now we have the chance to do it.

I really do think we are such a critical aspect of what all these companies are building. A lot of our customers at these frontier labs will often tell me they wouldn’t be able to build what they’re building without us, and they’re just amazed at what we do. Being able to be this critical part of what is literally the greatest technology of our time, and maybe one of the most important things we can ever build, is amazing.

8. The Real Reason AGI Might Take Until 2040

Why would you get acquired and stop doing that? Getting acquired would be really limiting. It would be an admission of failure and jumping ship because you can’t make it on your own anymore, when we’re the opposite. We’re incredibly successful, and there’s literally nothing else that I’d want to do instead.

Harry Stebbings

It is 2040, and we still do not have AGI. What is the primary reason why that would be the case?

Edwin Chen

I think there are 2 reasons. One is that there will always need to be more breakthroughs, whether it’s breakthroughs in how you leverage all this data or breakthroughs in the different types of algorithms that you’re building.

Another one is just how you gather that data. At the end of the day, in order to cure cancer, how will you gather the data that’s needed to make those breakthroughs? Maybe you’re going to have to run real-world experiments and real-world studies, and those studies will simply take time.

Will there be a way to speed up those experiments through various kinds of simulations or just other forms of gathering data? I don’t know, but there’s the question of how you get to data even faster, which I think will be very, very important.

Harry Stebbings

Speaking of evolutions with AGI, I do just want to ask about the changing nature of data. How will the data needed evolve as AI gets smarter and smarter and smarter with each evolution?

Edwin Chen

A lot of people talk about the shift to PhD-level data. It’s actually really interesting how we have the biggest group of the smartest people in the world working on a platform. We have Harvard professors, Stanford PhD students, and Princeton computer science theorists working on all these really interesting problems with us. It’s kind of crazy if you think of all the PhDs even at Google, Meta, or Microsoft—we have way more than all of them combined doing work for us in a single day.

They’re not just writing random JavaScript programs to improve ads. They’re actually pushing the frontiers of science when they’re collaborating with these models. But I think what people underestimate is that having a PhD isn’t enough.

A lot of PhDs just aren’t good at this type of work. As I said before, there are a lot of body shops and recruiting shops in our space that basically just look at whether you wrote down that you have a PhD on your résumé, and it will instantly give you work if so. But a lot of PhDs just aren’t very good.

I think 80% of the computer science PhDs I know write shitty code because they’re only good at math and algorithms. Think about people like Ernest Hemingway. He didn’t have a PhD. I don’t think he even went to college.

There are 2 things that are important. There’s this underestimated aspect of our space where you actually need a lot of technology to make sure that you’re delivering fully high-quality data. Vimeo has a lot of so-called high-quality videos, but they don’t have any algorithms, and so YouTube’s videos are way higher quality and more engaging in the end.

The second is that a PhD isn’t enough. Just because you have a PhD doesn’t mean that you can make some breakthrough in physics. What you also need is street smarts. You need the creativity and the mental fortitude to think of really interesting problems, find these problems, probe LLMs to see whether they can solve them today, and then teach them in really interesting ways.

Otherwise, if all you’re doing is throwing PhDs at a problem, all you’re doing is teaching models how to hack silly benchmarks and get good at basically the equivalent of SAT problems.

Harry Stebbings

If that’s the landscape today—PhDs aren’t enough, and a lot of PhDs aren’t great quality—how does that change over time? Will you have a dramatically larger supply side? How will the tooling of the supply side change? How will their ability to turn around work change?

Edwin Chen

This boils down to the technology that we build. Over time, it’s simply true that people are going to be trying to solve more and more problems. When you have hundreds of thousands, millions of people working on our platform, and you have 1,000 projects, 10,000 projects that are literally running in any given week, how do you make sure that you’re building technology to identify the top 1% or top 2% of people who can really push the boundaries of physics problems with these models?

How do you identify the top 2% or 3% of people who are writing the most amazing poetry? How do you find those people? How do you remove the worst of the worst—the people who will inevitably try to cheat you and spam you, and who will basically regress the models if you allow their data through?

It’s a really profound problem, and you just need a lot of technology to build this. At the same time, these are researchers who want to move really fast. Researchers at all these frontier labs—all the algorithms are changing every day. They want to learn and try out new projects every single week.

If you’re not moving fast enough, if you’re unable to create a new template or find the expertise that you need literally within the next day or the next week, it’s just going to be too slow for these researchers. If you don’t have the technology to manage these 10,000 projects, automatically create them, and automatically identify the really high-quality data, it’s just going to be too slow for them.

Harry Stebbings

Speaking of the slowness and quality of data, I would love to push you on this. When you think about bottlenecks to progress today, if I were to rank them 1 through 3, with 1 being the most pressing bottleneck and 3 being the least pressing, you’ve got access to compute, algorithms, and data quality. If you were to rank them 1 through 3, how would you rank them?

Edwin Chen

I would definitely rank data quality first, followed by compute, followed by algorithms.

Harry Stebbings

If compute continues to prove to be the unlock, where throwing more compute at it unlocks more and more performance, does that denigrate data quality in the prioritization stack?

Edwin Chen

I fundamentally don’t believe that you can throw more compute at it, because if you’re not getting the data that the computer is essentially trained on, or if you don’t have the right objectives and evaluation metrics that your computer is optimizing toward, you’re just going to fall into this trap of seeing progress that actually isn’t there.

I can give you some examples. Let me talk about why I think data quality is such a problem. I think data quality issues have already been a huge setback for a lot of frontier labs.

One of the things that we often hear from teams over and over is that, before they used us, they tried getting data in other ways, and so they trained their models. They evaluated their models, and their metrics kept going up, but after 6 months or even a year, they realized that their training data was [inaudible]. Their evaluation data was [inaudible], and so all the progress that they thought they were seeing was actually completely misleading. They either made no progress, or their models after 6 months were even worse than when they started.

For example, we see this a lot with LMArena. LMArena is this popular leaderboard of AI models, and it’s basically the equivalent of clickbait. What happens is that you have people going on to what’s called Chatbot Arena. They’ll enter a prompt, see 2 model responses, and then vote on which one’s better. But they’re not taking the time to really read or evaluate the model responses at all.

One of the models could have made everything up, and these participants will vote on it because it has emojis and nice formatting. We’ve literally seen this in the data ourselves. One response will just be a complete hallucination, but because it has an emoji and because there are a couple of words bolded, people will just be like, “Okay, yeah, that looks good. That looks much better than this other thing. I didn’t take the time to fact-check at all.”

One of the things that we’ve learned is that the easiest way to improve in this arena is simply to make your model responses a lot longer. One of the funny things is that if you take the top model on its leaderboard, the number 1 model, and ask it, “When did the pope die?” it will give you a really long response that seems impressive, but it gets the answer completely wrong.

It tells you that Pope Francis is still alive. It will even tell you that there are search results indicating that Pope Francis died in April, but those were just rumors and misinformation. He’s still alive. It’s wild that this model will say this.

Again, there are a lot of companies trying to improve their leaderboard rank. They’ll see progress for 6 months because all they’re doing is unwittingly making their model responses longer. They’re adding more and more emojis and more and more formatting, so they see their models climbing on this leaderboard and think they’re making progress.

All they’re doing is training their models to produce better clickbait. They may finally realize that 6 months or a year later, but it means they basically spent the past 6 months making zero progress. This is what happens when you throw compute at the problem without understanding the underlying training data that you’re throwing the compute toward. It actually just sets your models back.

Harry Stebbings

When you look at Grok 4 announcing its recent developments, performing so well in the latest benchmarks, and coming out as number 1, are those benchmarks misleading? How much weight should be placed on the importance of those benchmarks, and how reflective are they truly of model quality?

Edwin Chen

If you watch the Grok 4 launch—the Grok 4 livestream—I think you would have heard Elon himself saying, “Yeah, these models are really good at…” I forget the word he used, but they’re really good at homework problems. They’re really good at these academic or very narrowly scoped problems.

It’s basically the equivalent of making them really good at SAT problems, but not making them good at problems that people are actually facing.

Harry Stebbings

Totally get you. Were you surprised by how far Elon has been able to get with Grok, as fast as he has, or not?

Edwin Chen

Again, I think Elon has this. It’s kind of funny. Before we worked with the team, I didn’t really have a conception of what an Elon company was like. We work really closely with the xAI team, and it’s actually incredibly refreshing to see how they operate.

They’re all very, very mission-oriented, incredibly smart, and they work incredibly hard. It’ll be 11:00 p.m. at night, and I’ll DM them. Someone will want to jump on a meeting, and I’ll jump on a meeting with them and see that they’re in the office. There are a ton of people behind them, so they’re just crazily hacking together on all these problems.

I actually think it’s incredible. It’s this embodiment of what a startup can do when you really believe in something and are willing to do whatever it takes to achieve it, as opposed to living within the confines of this giant bureaucracy. I think it’s really, really impressive.

Harry Stebbings

Is there anything that you think Elon does specifically to inspire his team to have that form of culture when they’re not a small company?

Edwin Chen

I think it’s almost that you know what you’re getting into when you work at Grok, when you work at xAI, or when you work at any of these other companies. You know when you interview that these people are incredibly mission-oriented. You know when you interview that everybody works super hard.

You know that if you want to work there, you’re going to have to be the kind of person who has the same values. Otherwise, you just shouldn’t join because you’ll be miserable. It’s the fact that they have such a strong culture and such a strong belief in what they’re doing that attracts people of similar talent.

Harry Stebbings

But before we do, I do want to touch on the working hours that you mentioned there. Everyone poses synthetic data as a big threat. What happens to your business when we have synthetic data that is obviously created automatically and labeled automatically? How do you think about the role of human-labeled data in a world of predominantly synthetic data? What are your thoughts there?

Edwin Chen

I think synthetic data is actually really useful in some places, but I think people overestimate what it can do. I’ll give a couple of examples.

Right now, there are a bunch of models that have been trained really heavily on synthetic data, but, as I mentioned earlier, that means they’re only good at very academic, homework-style, benchmark-style problems. They’re actually terrible at real-world use cases. Synthetic data has made models good at synthetic problems, not real ones.

We hear from a lot of companies that tell us they spent the past year training their models on synthetic data, but they’ve only now just realized all the problems it’s caused. They’ve spent months throwing a lot of it out. A lot of them tell us that even 1,000 or a couple thousand pieces of really high-quality human data that we generated for them have actually been worth more than 10 million pieces of synthetic data.

A lot of the work that we do is simply cleaning up all this synthetic data. If you think about why this happens, it’s essentially because the models collapse on this very, very narrow scope of similarity that the synthetic data creates. It just doesn’t give the models the kind of diversity and generalizability that they need.

One other point is that there’s also this interesting phenomenon where models simply make a lot of mistakes and have certain misunderstandings that humans never will. I was recently playing with one of the frontier models, and it kept randomly outputting Russian characters and Hindi characters in the middle of its responses.

This is a mistake that would be obvious to any human, to any second grader, but a model just didn’t know. It’s shocking that a frontier model in 2025 would do this. It’s almost like you always need this external value system as a kind of safeguard to make sure that the models are working properly, just because the models themselves have such a different way of thinking.

Harry Stebbings

I’m an investor in Poolside, which, if you don’t know, is obviously kind of in the same space as, say, Cursor or Windsurf. Bluntly, they seemingly are much more behind because they’ve built their own models, and they believe very much in the power of verticalization of models and specific, or specialized, models, so to speak.

How do you think about the future in terms of monolithic, generalized, very large-scale models versus the requirement to have very narrow, very specialized models for things like code creation and development?

Edwin Chen

I think there’s an opportunity for both, and the reason I think that is because, on the one hand, you have these giant, all-powerful models. Sure, they can be really, really good and really, really powerful in a raw capability sense. At least right now, I think they’ll be able to encompass all of these different use cases.

In the same way that a company—take a company like Google or Facebook—simply can’t build certain products because building those products would be counter to the culture or the business goals of the overall parent company, sometimes you need to be able to move faster and take big bets on certain kinds of products.

The all-powerful model just can’t let that happen, because if you let it happen within this one small domain, it will almost pervade the entire model. Sometimes you do need the smaller models to break through if they have a really unique view on how they’re operating.

Harry Stebbings

Can I ask you, Edwin? You are very composed as a leader, as a CEO. It translates incredibly. Where are you not meeting the bar? Where are you not great, and you are aware of it?

Edwin Chen

I think one area where I'm not great, which is kind of funny, is I'm really bad at understanding financials. Sometimes people around a company will try to tell me, “Hey, have you been paying attention to our revenue numbers? Have you been paying attention to our costs? Have you been paying attention to our margins? Do you even know what they are?” And I don't. They're just these financial metrics that I could not tell you what EBIT does. I know what the acronym stands for, but the difference between that and revenue and profit and net margin—I actually just don't know any of these terms. It's just this blind spot. No matter how much I try to understand these things, I can never remember.

Harry Stebbings

What single metric defines the health of the business to you? What metric, if I showed it to you every morning, would make you say, “Okay, I know the state of my business”?

Edwin Chen

If I could paint my perfect North Star—and this is something that I think we want to work towards, something that we actually want to build for the industry—it would be: Are models progressing in fundamental ways? Are they actually getting more intelligent? Are their capabilities improving, as opposed to simply climbing up a meaningless clickbait leaderboard? So, are these models progressing, and how much of that is due to us, whether it's due to our training data, the evaluations we provide, or the insights that we provide to all these researchers for ways that they can improve their models? If there's a way to measure that, I would love it.

I think the closest proxy we have for it today is the variety of projects that we're creating. One of the things I really believe in is that we want to make it easy for all of these researchers to come up with new ideas and not be blocked by data. The more complex, diverse, and creative projects that we can provide, that is almost a proxy for that overall score.

9. The Price of a $10B Company?

Harry Stebbings

Final one, and then we'll do a quick fire. But you mentioned Elon and X, and the hard work in that culture being so ingrained. I recently said that, bluntly, Silicon Valley and China have increased the intensity required to win in terms of work ethic. You must work 7 days a week if you want to build a $10 billion-plus company. The ability to put your phone on the side and not check an email does not exist anymore if you want to build a $10 billion-plus company. You've built a $10 billion-plus company. Do you agree with me?

Edwin Chen

I think you have to be willing to work hard. You have to be willing to jump on a call at 2 a.m. with a customer. One of the things that I love is that sometimes customers will call me—they'll literally call me at 2 or 3 a.m.—and they'll be like, “Hey, our models are freaking out. I need a bunch of data to fix it by 6 a.m. Can you do it?” Going back to the question of things that make me happy, nothing makes me happier than knowing that we can deliver this. We can deliver 10,000 data points to you in the next few hours, even if you call us at 3 a.m. to fix some critical bug, some critical fire that you're facing. That actually makes me incredibly happy.

I think you have to be willing to work hard. A lot of people do confuse working hard with creating value. It's maybe a trope to say, but you have to work smart and not just hard. If I think about a lot of what I'm doing, oftentimes the best ideas come to me when I'm just walking around, not necessarily when I'm at my computer. I think we all work really hard, but I wouldn't confuse the number of hours we spend with actual progress.

Harry Stebbings

What trait of yourself do you love most, or what is your favorite trait, Edwin?

Edwin Chen

The thing I really enjoy is when there's a unique insight. I've always really enjoyed writing down insights in written form, and I think I'm pretty good at it. This ability to deliver some novel insight about a model, an algorithm, or a data set, and communicate that to our customers—I think I'm pretty good at it, and it's something I really enjoy.

10. Quick-Fire Round

Harry Stebbings

Dude, I want to do a quick fire. I say a short statement, you give me your immediate thoughts. Does that sound okay?

Edwin Chen

Yeah, that sounds great.

Harry Stebbings

What one widely held belief about AI do you think is completely wrong?

Edwin Chen

I think a lot of people think AI safety is overblown, but they ignore the paperclip maximizer problem, where you have AI models that are accidentally trained towards the wrong objectives. This is a big problem that all AI models face today, with all the issues around LMArena and benchmark hacking. I actually think it's a really important problem that people should be thinking more about.

Harry Stebbings

So you think AI is much more dangerous than we let on?

Edwin Chen

I think it is dangerous, but the bigger issue is that it can be accidentally maximized towards the wrong objectives. Today, sure, if you maximize towards these LMArena objectives or benchmark hacking, the worst that will happen is that your models will regress in performance a little bit. But the more fundamental problem is that people don't realize this.

In the future, when the models are more powerful, you're basically accidentally maximizing AI models towards the wrong objectives, and you just have no idea what will happen. It's almost a similar phenomenon to what's happening today, but because the AI models are so much more powerful, they're literally building the code for an insurance company or a trillion-dollar company. The consequences can be much worse.

Harry Stebbings

You mentioned gaining true passion and love for building towards AGI. I hate myself for asking this question. It's a shit question. I hate it. I'm so embarrassed. But if you had to put a number—2028 or 2038—which bracket would it be in, and why?

Edwin Chen

I think it would be 2028 if you're talking about automating the job of the average engineer, and 2038 if you're talking about curing cancer.

Harry Stebbings

Sorry—2028, automating the job of the average engineer? I had Vlad Tenev on the show from Robinhood—it went out today—and he said 50% of code created by Robinhood is now by AI. Marc Benioff said the same on the show: 50%. Are we not at that stage already? How much code from Surge is created with AI?

Edwin Chen

I don't think we're at that stage yet. At least, if you're working on deeper problems that aren't just random features—if you're concentrating your company on the 10% of problems that are most important—I don't think models today can write 50% of the code and come up with 50% of the ideas that are actually going to be meaningful to your company. Sure, if 90% of your company is writing little features that nobody cares about or improving the efficiency of your code base by 1%, then yeah. But I don't think we're at a point where, if you're really working on meaningful problems, models can do that.

Harry Stebbings

What question should every AI company be asking themselves?

Edwin Chen

If you're a frontier lab, the question is: Are you actually improving your models' raw intelligence, or are you just hacking benchmarks? If you're a product company, the question is: Why won't frontier labs be able to instantly replace you?

Harry Stebbings

Do you think they will?

Edwin Chen

I don't ever worry about the application layer being absorbed by the model layer, just because I think there's infinite product breadth that they could go after. They can't go after everything. But there are so many things where you literally just want to chat with the model in this very simplistic, universal interface.

Think about Google Search. I do feel—I have felt—that maybe 50% of the things I used to Google in Google Search are replaced by ChatGPT, or they're even better with ChatGPT. There's a very pleasing aspect of a universal, all-intelligent interface that I think people will just gravitate towards it.

Harry Stebbings

What would you do if you were Sundar today? Would you kill your golden goose with the ads engine?

Edwin Chen

The difficult problem, I think, for Google is that they have to be willing to take a short-term hit to all of their advertising revenue in order to build something better. That's just really hard.

Harry Stebbings

Incredibly hard. Final one for you. Actually, penultimate one. What did you believe about the future of AI that you now no longer believe? What has changed your mind?

Edwin Chen

I see a world where there will actually be multiple frontier AI companies, multiple frontier AGIs, just because every one of them will be able to go in a different direction. You see it already—you see it playing out today—with the differences and the strengths and weaknesses of OpenAI and Anthropic. I just think that trend will continue.

Harry Stebbings

What does that mean? I'm sorry, if you just play that out, what does that landscape look like, then? Because OpenAI and Anthropic are so unique in their properties and characteristics, it means there'll be 10 more of them.

What does that look like?

Edwin Chen

I don't know if there'll be 10 more of them, but I can certainly see even 3 more of them. I think each one will have different trade-offs that they're willing to make and different focuses that they'll have.

Even today, Claude is really, really good at coding. Claude is really, really good, I think, at enterprise and instruction following, whereas ChatGPT is more optimized for consumer use cases. I think it actually has a really great and fun personality right now. And then Grok is willing to maybe answer certain questions that maybe it should, maybe it shouldn't, but it's willing to be a little bit transgressive in ways I actually think are very interesting.

I think this willingness to have different personalities, different boundaries, and different focuses on your models just leads the models to be good at different use cases. It's the same way that I think the analogy is that there isn't a single poet, and there isn't a single mathematician who is the greatest mathematician of all time. They all have different focuses and different ways of approaching these problems. I think that richness of what we often call human intelligence will apply to models as well.

Harry Stebbings

Have the biggest model providers been founded today?

Edwin Chen

I don't think so yet. I can actually see big, new, even more powerful model developers appearing in the next few years.

Harry Stebbings

How so? How does that look? Because when you think about funding them, the capital intensity or capital requirements are so large. All the big players in the financing world, bluntly, have already got their horses in this race. How does that even work?

Edwin Chen

I think it's because it depends on what you view the long-term vision for AGI to be. If you believe that, despite all the immense progress that we've made, we're still only—I don't know—1% or 5% of the way towards AGI, we literally want AGI systems that can, in the future, cure cancer, send rocket ships to Mars, and design entirely new philosophical systems. These are big, massive problems.

As opposed to simply automating away the job of the average L3 or L4 software engineer, if you believe that we're only, again, 2% or 5% of the way there, there's so much more headroom. It's almost like asking, “Do you believe, 10 years ago, that Google was going to be the final search engine in the world?”

Sure, if you're only looking forward to the next 5 years, in some sense. But if you just think of the immensity of what AGI could do, there's so much more ahead of us than behind us that there could be these serendipitous, very creative breakthroughs that nobody's expecting, in part because maybe they're going to be created by some of the AIs themselves or by AIs in concert with humans. There's just so much opportunity ahead of us that it would be almost a miss to think that we've already solved it.

Harry Stebbings

Do you believe AI will be able to create 10% increases in GDP gains or in productivity increases in the next 10 years? That's often touted as the number that would create $10 trillion of value.

Edwin Chen

Yeah, I absolutely believe it.

Harry Stebbings

Final one, Edwin. You can give yourself 1 piece of advice going back to day 1, starting the company, going back to starting the MVP. What do you know now that you could tell yourself then?

Edwin Chen

I think it would be to focus always on the 10x improvements that you can make, as opposed to worrying about 10% improvements.

Harry Stebbings

Edwin, listen, I so appreciate the time. As I said at the beginning, I've been such a fan of the incredible journey. You've been fantastic. It's been very atypical in most ways, bluntly, having this discussion, which has been so great for me. Thank you so much for joining me.

Edwin Chen

Thank you. It's been great chatting.