[BidClub_]
Latent Space · · 76 分钟

压过 Noam Shazeer,以众包方式打造 DAU 达 140万的 Chai AI——对话 William Beauchamp,Chai Research

Alessio FanelliswyxWilliam Beauchamp

YouTube
TL;DR
  • Chai 的判断是,社交 AI 最终会成为一个分布式创作者平台,而不是只靠更多数据、算力和研究人员占优的单一大模型。 William Beauchamp 预计,专业团队和用户会分别塑造不同的 AI,就像量化机构在同一个金融市场里各自专精一样。「必须存在这样一个平台,让一个小团队也能为独特目的打造 AI。」

  • Chai 找到产品市场契合点,是在效用型机器人全部失败、Beauchamp 的妹妹做出一个治疗师机器人并吸引到 20名用户、每人约使用 20分钟之后。 新闻、食谱、笑话、测验、名人和网红都没有起色;即时且不带评判的对话才是「AI 强 10倍的事情」。随后,Chai 让消费者只需输入提示词、图片和名字就能创建角色,由此发现了工程师绝不会主动设计的冲突、恋爱和人物原型需求。

  • Alessio 称 Chai 的 DAU 已达 140万、收入超过 2200万美元;Beauchamp 没有明确确认这些数字,只说他认为用户数去年增长了 3倍,收入增幅超过 1倍。 单次典型使用约 90分钟,而 TikTok 被引用的时长是 70分钟,并产生约 150条消息,因此推理经济学的重要性远高于问答类产品。真正的前沿不只是基准测试表现,而是「每美元对应的性能」。

  • 2024年的增长曲线反映了 4个重大变化:修复基础设施、竞争对手疑似减弱买量、更好的模型,以及付费分发。 DAU 接近 50万时,Firebase 已无法稳定扩容,Chai 被迫经历痛苦的 3个月迁移,之后才触及 150万 DAU;后来,Beauchamp 怀疑 Character.AI 在创始人离开后减少了广告投放。Chai 随后把买量提升到约 4万美元/天,将用户年化增速从约 2倍推升至 3倍:先把产品做出来,再通过买广告接上一枚巨大的「火箭」。

  • Chai 最强的运营优势,是一套把传统 30天留存测试压缩到约 3小时的人类反馈闭环。 每个提交的模型都会交给用户进行比较;Beauchamp 称,约 5000次生成就能得到准确反馈。如今约 5名研究人员每天评估 20–50个模型,每周至少上线 100个。他认为,大约在 10月的某个时间点,用户评论从 Character.AI 更好转向 Chai 更好——这是他看重的证据,因为「你骗不了消费者」。

  • Chai 在投入 3个月后发现语音没有带来可测量的留存、参与度或变现提升,于是放弃了把语音作为增长主线。 只有 10–15%的用户尝试过语音,使用时间占自身使用时长的 10–15%,折算下来只占整体体验约 2%,尽管 Beauchamp 坚称模型本身非常好。他如今的产品标准很苛刻:一项功能必须对大多数用户的大部分体验都重要,并且让人觉得「这是一件大事」。

  • Beauchamp 已经推迟了自己对 AGI 的时间表,也不接受当前 LLM 天生就是强推理引擎的说法。 他称 LLM 是带有「超级知识」的模拟器:极其擅长存储、检索和生成,但按他对智能的定义,仍然不够聪明。Chai 不做流式输出,而是生成 16个完整候选答案,再用奖励模型选出一个。他以 5000万条消息训练这类模型作为示例,并未确认这是已部署的配置。

摘要 · 为研究而整理的核心内容

1. 量化交易利润为寻找影响力提供资金

  • Beauchamp 于 2012年从 Cambridge 毕业,当时靠打扑克积累了约 10万美元。小资金反而是优势:一个每年能赚 10万美元的异常机会,对 1000万美元本金只意味着 1%的收益,对 10万美元却意味着 100%的收益,因此他自学 Python 和机器学习来捕捉这种机会。

  • 这家公司最终每年赚取约 500万美元,团队约有 15名毕业于 Oxford 和 Cambridge 的数学家、物理学家,只交易团队自己的资金。没有「客户抱怨」,也没有投资者约束风险,这正是 Beauchamp 曾经向往的量化交易理想状态。

  • 30岁时,他决定再增加一艘游艇规模的财富,也不会产生真正有意义的影响。加密货币作为赌博工具、以及规避货币监管和银行限制的手段,看起来很特别;但其更广泛的区块链和 Web3 逻辑「确实没太大意义」,于是他把精力转向语言模型。

2. 机器学习的 S 曲线削弱了单一超级 AI 的叙事

  • 阅读 Google 发表的研究,以及当时仍然开放的 OpenAI 的论文后,Beauchamp 认定 LLM 会很重要。但他拒绝了打造单一智能的主流竞赛——靠最多的数据、算力和研究人员堆出一个模型:以他的经验,机器学习性能遵循 S 曲线,通常会在人类能力附近进入平台期。

  • 自动驾驶、图像识别和语音识别都支持这一判断;AlphaGo 是显眼的超人类例外。因此,Beauchamp 预计 AI 会更像金融市场:高频、中频、股票等不同专业机构,通过各自的算法在一个共享市场中竞争,而不是都存在于一家全能量化机构内部。

  • Chai 的创始命题由此直接推导出来:「必须存在这样一个平台,让一个小团队也能为独特目的打造 AI。」它对应的不是《大英百科全书》,而是 Wikipedia、YouTube 或 Twitter——一个由分布式参与者发现中央机构无法发现之物的生态。

3. 效用型机器人失败,暴露了对话才是原生产品

  • Chai 最初允许开发者提交带有文本输入、文本输出接口的 Python 智能体。Beauchamp 做过 Reddit 新闻机器人、食谱助手、冷笑话、测验和知识问答;尽管他预想过后来答案引擎式的产品,却发现「显然没有产品市场契合」,因为当时模型很弱,对话式效用并没有增加多少价值。

  • 他妹妹开发的治疗师机器人一夜之间改变了方向:约 20名活跃用户平均与它交流 20分钟。对话可以在凌晨 3点即时发生,不带评判,也比等朋友回复更容易;即使只是 AI 生成的一句赞美,也带来了效用型产品此前没有的体验价值。

  • Beauchamp 将这种参与式体验与 TikTok 或 Instagram 的被动消费区分开来。刷 40分钟短视频可能让人懊悔,因为「我什么也没完成」;与 AI 互动则让人感觉自己有所贡献。他引用用户的说法称,Chai 帮助他们应对饮食失调、抑郁和人生低谷,但强调这些只是用户的反馈。

4. 消费者作者,而非软件开发者,打开了角色目录

  • 用 kbot、名人和网红人为制造需求的尝试同样失败。真正的突破在于认识到,Python 开发者并不想构建社交角色,但消费者想做;于是 Chai 通过提示词、图片和名字,让用户使用一个 60亿参数的 GPT-J 模型。

  • 用户创造出了 Beauchamp 从未预测过的类别:操场上的霸凌者、争论、打斗,以及陌生的恋爱人物原型。Chai 不再需要猜测 1%的人想要什么,而是让用户自己创造,为更广泛的人群提供所需的多样性。

  • 主持人批评 Chai 的创作者层看起来仍然异常单薄,Beauchamp 完全接受了这一点。他的答案是让短描述更具可操控性:只需「一艘宇宙飞船」、3名船员、戏剧冲突和打斗,就能胜过专家型 SillyTavern 创作者使用的 1000字角色卡。

5. 风险资本竞争让 Chai 认识到模型质量的价格

  • Beauchamp 回忆,到 2022年末或 2023年初,Chai 的 DAU 已接近 10万,并成为 App Store 上排名第一的 AI 应用。Character.AI 出现后,提供了非常相似的体验,团队起初还嘲笑这个产品;随后 Character.AI 融资 1亿美元,之后又融资 1亿美元。

  • 当时 Beauchamp 可能投入了约 200万美元,使用的是 GPT-J 6B。Chai 的示例经济学是,一名用户完整使用一次约花费 1美元;如果有 100万用户,AI 总成本就是约 100万美元。Character.AI 可以花费其 100倍,用户也注意到了差异:「为什么你们的 AI 这么蠢?」Chai 于是搬到 Silicon Valley、获得融资,并认识到消费级 AI 是一门 Silicon Valley 式的超大规模生意。

  • Alessio 称 Chai 的 DAU 已达 140万、收入超过 2200万美元。Beauchamp 回应说,他认为用户数去年增长了 3倍,收入增幅超过 1倍。他说自己认为 Character.AI 的估值接近 30亿美元,DAU 为 500万。

  • 他对 DeepSeek 的比较,核心围绕推理经济学和创始人驱动、以客户为中心的执行力展开。他称赞 DeepSeek 最新的 V2,重点是其推理引擎和显著更小的 KV cache,后者降低了推理成本;他认为每美元对应的性能比基准测试分数更重要。他也关注 Llama 4 能否实现同样的提升。

6. 基础设施与分发解释了增长曲线上的明显拐点

  • Chai 从 1个 DAU 起一直使用 GCP,直到 DAU 接近 50万,期间尤其重度依赖 Firebase——据称超出了 Google 工程师建议水平约 3倍。这层抽象让团队可以专注于 AI,直到服务中断迫使其至少经历 3个月的迁移和服务拆分。

  • 服务中断造成的损害不止于当天流量。新用户遇到无法正常运行的应用,留存、消费、评分乃至 App Store 排名都会受到影响;修复后端很快,但恢复自然排名可能要久得多。重建后的技术栈随后支撑 Chai 增长至 150万 DAU。

  • Beauchamp 怀疑,Character.AI 的创始人离开后,公司降低了用户获取投入。他用一个每天花费 10万美元的假想公司与一家完全不花钱的公司来说明竞争影响,但没有把这个数字说成是 Character.AI 已披露的实际支出。

  • 一名前 ByteDance 增长负责人得知 Chai 在没有广告的情况下做到约 100万 DAU 后感到震惊。Chai 先测试了约 1万美元/天,随后是 2万美元,如今约为 4万美元;Beauchamp 说,这把原本约 2倍的年度增长轨迹转成了 3倍。

7. 3小时反馈闭环成为模型开发引擎

  • Beauchamp 的运营信条是「成功源于失败」:平淡期代表学习,增长期则是在收获学习成果。Chai 在 Q2 的关键开发,是一套把任何提交的模型交给约 5000名用户评估的系统,并根据用户觉得哪些输出更有趣、更能带来参与感进行排名。

  • 这让约 5名研究人员从每周可能评估 3个模型、甚至一度每月难以测试 5个模型,提升到每天 20–50个、每周至少 100个。原本需要等待 30天观察第 30天留存的标准队列测试,被压缩成约 3小时的反馈信号。

  • 由此形成的速度,让 Chai 可以快速迭代 DPO 微调、提示词、模型混合、拒绝采样和奖励模型。他认为,大约在 10月的某个时间点,Reddit 反馈和用户交流发生了转折:从「Character.AI 更好」变成「你们更好」;在切换成本低、单次使用长达 90分钟的情况下,Beauchamp 认为用户不可能被质量差异蒙骗。

8. 音频失败,迫使产品标准进一步收紧

  • Chai 用 3个月解决语音延迟、成本、质量、激活和交互设计问题;按 Beauchamp 的说法,其上线时间至少比 Character.AI 早 9个月。但 A/B 测试显示,留存、参与度和变现都没有变化,团队还花了 1周排查一个并不存在的 bug。

  • 一名主持人猜测,可能只是模型质量不够好;Beauchamp 强调回应:「不,它们非常好。」只有 10–15%的用户激活音频功能,并且只在 10–15%的时间里使用,最终只改变了约 2%的整体体验。

  • 他的结论是,便利性不会自动创造目的地。成功的功能必须让大多数用户在其大部分使用过程中,都获得足够独特、足够有吸引力的体验,并产生「哇,这是一件大事」的反应;否则,即使技术优秀,也无法推动公司层面的指标。

  • 因此 Beauchamp 认为,音频和图像生成并不是用户最首要的问题。他说:「所有 AI 都是 Silicon Valley 的中年男性在生成」,真正未被满足的需求,是让用户自己训练并塑造体验。

9. Chai 想构筑的护城河,是不断变厚的 UGC 层

  • 在 Beauchamp 看来,如今的提示词、图片和角色名只是「我们能做的 1%」。他的完成标准刻意设得极高:除非有一个创作者团队能在平台上制作 AI 内容、每年赚取或花费 1亿美元——或者不管最终数字是多少——否则 Chai 就还没有完成;这类似于大型视频创作者的经济规模。

  • 这座护城河由创作者、消费者和算法共同构成。用户行为训练推荐系统;推荐结果告诉创作者什么有效;更好的内容吸引更多用户。Beauchamp 用 MrBeast 在 Amazon 上适配性更弱作为例子:YouTube 的迭代会针对特定生态优化他的缩略图、开场和内容。

  • Chai 想要的是类似 TikTok 的创作杠杆,让普通人借助内置音乐和特效就能制作有趣内容。「用户不想被迫工作」;平台应该让短提示词也能有效,而高级创作者则贡献经过微调、能够产生真正独特行为的模型。

10. ChaiVerse 把实时人类品味变成开放式模型竞赛

  • 每天有数百个模型进入 Hugging Face,背后是创作者投入的数据、算力和劳动。ChaiVerse 提供托管这些模型、为其分发流量并收集用户两两比较判断的服务。Beauchamp 称,得到准确反馈大约需要 5000次生成,采用的是类似 LMSYS 的方法。

  • 他的分布判断是,底部 80%的模型「相当糟糕」,可以直接忽略。顶部 20%会显现出有用差异,可能体现在描述、个性、幽默或逻辑上;但在这些模型之间按请求进行路由,成本很高,带来的优势也很小。

  • Chai 更倾向于模型混合:随机 50%的请求交给聪明模型,另外 50%交给有趣模型。「随机是一种非常强大的优化技术」,Beauchamp 说,因为它能进行广泛探索,同时具备异常稳健的表现。

  • 他用最初提交的模型可能只有 1000–1100 Elo 分、随后反复失败、最终突然提升来说明迭代过程。Chai 已向创作者支付超过 10万美元,但付款没有提高提交频率,主要只是为算力提供资金;例如一名 17岁创作者就把 1000美元奖金花在了一块实体 GPU 上。

11. 人类偏好是北极星,但也会产生偏差

  • 面对 Elo 不可能成为 Chai 唯一评估指标的质疑,Beauchamp 称它是北极星,因为「人类知道自己想要什么」。设计型评测只是快照,会趋于饱和并需要更换;两两偏好则保持通用性,并且能随着 Chai「反馈丰富」这一独特位置持续扩展。

  • 他承认,原始偏好会奖励表面技巧:按照他的例子,任何 LLM 只要训练自己使用脏话,就能变得更有趣 20%。Chai 对安全等阻塞项使用过定向评测,但称这些测试可能在 1个月内饱和;长期答案是让人类反馈更稳健,而不是用静态基准取代它。

  • 一名主持人提出分层质疑:治疗、角色扮演和不适合工作场景的用户显然不同。Beauchamp 给出了反直觉的回答:偏好仍然高度相关。一个人可能给答案打 10/10,另一个人打 7/10;但要实现强个性化,必须让一组人热爱的内容,成为另一组人觉得无聊的内容。AI 内容目前还没有多样到足以形成 YouTube 规模的推荐流分化。

12. 超级知识与推理时搜索,取代近期 AGI 叙事

  • Beauchamp 说,他的 AGI 时间表「确实已经推迟」。LLM 看起来不像推理引擎,更像是在预测最可能的后续内容的模拟器,就像一个物理游戏模拟汽车坠落到搭建好的桥上时会发生什么。

  • 他的区分是知识与智能:模型可以存储和检索比任何人类都更多的信息,但这种优势很容易被误认为推理能力。他接受「超级知识」是更好的术语,并认为 AI 处于 20年旅程的第 4年,大致相当于 1998年的互联网。

  • William 将 OpenAI 的 o1 和 o3 与类似树搜索的方法联系起来;swyx 指出,OpenAI 并未表示自己使用树搜索。William 称这是隐含信息,并表示这类系统更擅长推理,但拒绝称其为推理引擎。他坚持认为,模型的原生优势在于检索、存储和生成:「它就是能编东西。」

  • Alessio 称 Chai 去年在算力上花费了 1000万美元,可能会把这一数字提升到 3倍;William 随后强调推理优化。他说,Chai 评估过 MK1 的推理引擎,速度远快于 vLLM,并突出创始人 Paul Merolla 的硬件专长和 CUDA kernel 工作。

  • Chai 从未做流式输出,因为流式输出会阻止拒绝采样。与其把首个 token 的生成时间优化到约 4秒,Chai 可以用约 10秒生成完整答案,并提供更大的模型。它会生成 16个完整候选答案,再用奖励模型选出一个。此类模型的一个例子,是用 5000万条消息训练,并根据用户是否回复进行标注,从而预测哪个答案最可能促成回复。

Alessio Fanelli

Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO at Decibel, and today we're in the Chai office with my usual co-host, swyx.

swyx

Hey, thanks for having us. It's rare that we get to get out of the office, so thanks for inviting us to your home. We're in the office of Chai with William Beauchamp.

Alessio Fanelli

Yeah, that's right. You're the founder of Chai, but previously—I mean, I think you're concurrently also running your fund?

William Beauchamp

I was simultaneously running an algorithmic trading company, but I fortunately was able to exit from that in Q3 last year.

Alessio Fanelli

Congrats.

William Beauchamp

Yeah, thanks.

Alessio Fanelli

Chai has always been on my radar because, first of all, you do a lot of advertising, I guess, in the Bay Area, so it's working. Second, the reason I reached out through our mutual friend Joyce was that I'm generally interested in the consumer AI space and chat platforms in general. I think there are a lot of insights we can get from that, as well as insights into human psychology—a weird blend of the two.

We also share a bit of a history as former finance people crossing over. I guess we can start with the origin story of Chai. Why decide to work on a consumer AI platform rather than B2B SaaS?

William Beauchamp

Just quickly touching on my background in finance: originally, I'm from the UK—born in London—and I was fortunate enough to study economics at Cambridge. I graduated in 2012, and at that time, everyone in the UK and everyone on my course thought HFT and quantitative trading were the big things. It was the big wave that was happening, so there was a lot of opportunity in that space.

Throughout college, I'd played poker. I dabbled as a professional poker player, and I was able to accumulate about $100,000 through playing poker. At the time, as my friends went to work at companies like Jane Street or Citadel, I did the math and thought, well, maybe if I traded my own capital, I'd probably come out ahead. I'd make more money than just going to work at Jane Street or Citadel.

swyx

$100K as capital?

William Beauchamp

Yes, yes. That's not a lot.

Well, it depends on what strategies you're doing. There is an advantage to being small, right? There are strategies that don't work if you have a fund of $10 million. If you find a little anomaly in the market that you might be able to make $100K a year from, that's a 1% return on your $10 million fund. If your fund is $100K, that's a 100% return. Being small, in some sense, was an advantage.

I started off and taught myself Python. Machine learning was the big thing as well. It was the first big time that machine learning was being used for image recognition. Neural networks had come out, you had dropout, and machine learning was the big thing that was going on at the time.

I probably spent my first 3 years out of Cambridge just building neural networks and random forests to try to predict asset prices, and then trade that using my own money. That went well. If you start something and it goes well, you try to hire more people. The first people who came to mind were the talented people I went to college with, so I hired some friends.

That went well, and I hired some more. Eventually, I ran out of friends to hire, so that was when I formed the company. From that point on, we had our ups and our downs. That was a whole long story and journey in itself, but after doing that for about 8 or 9 years, on my 30th birthday—which was 4 years ago now—I took a step back to evaluate my life.

I looked at my 20s, and I loved it. It was a really special time. I was lucky and fortunate to have worked with this amazing team, been successful, had a lot of hard times, learned wisdom through the hard times, and then had a lot of success and been able to enjoy it.

The company was making about $5 million a year, and it was just me and a team of around 15 Oxford- and Cambridge-educated mathematicians and physicists. It was the real dream that you would have if you wanted to start a quantitative trading firm. It was like sasana or rch.

Alessio Fanelli

It was all your own money?

William Beauchamp

Exactly. It was all the team's own money. We had no customers complaining to us about issues, no investors saying they didn't like the risk we were taking. We could really run the thing exactly as we wanted it. It's like Asana or RCh[?].

Those were the companies we would look toward as we were building that thing out. But on my 30th birthday, I looked at it and said, okay, great, this thing is making as much money as anyone would really need. What happens if we keep going in this direction?

It was clear that we would never have a big impact on the world. We could enrich ourselves, make really good money, and everyone on the team would be paid very well. Presumably, I could make enough money to buy a yacht or something, but that stuff wasn't that important to me.

I felt a sort of obligation that if you have this much talent, and especially if you have a talented team as a founder, you want to be putting all that talent toward a good use.

I looked at getting into crypto at the time and had a really strong view on it. As far as a gambling device, it's the most fun form of gambling ever invented—super fun. As a way to evade monetary regulations and banking restrictions, I think it's also absolutely amazing.

It has 2 killer use cases: not so much banking the unbanked, but everything else to do with the blockchain and Web3. That didn't really make much sense to me. Instead of going into crypto, where I thought even if I were successful, I would end up in a lot of trouble, I thought maybe it would be better to build something that governments wouldn't have a problem with.

I knew that LLMs were a thing. I think OpenAI had said they hadn't released GPT-2 yet, but they had said, “GPT-2 is so powerful, we can't release it to the world,” or something. Then I started interacting with some language models that Google had open-sourced. They weren't necessarily LLMs, but they were enough to show the potential.

Nowadays, so many people have interacted with ChatGPT that they get it, but the first time you can just talk to a computer and it talks back is a special moment. Everyone who's done that goes, wow, this is how it should be. Rather than having to type on Google and search, you should just be able to ask Google a question.

When I saw that, I read the literature and came across the scaling laws. Even 4 years ago, all the pieces of the puzzle were there. Google had done this amazing research and published a lot of it. OpenAI was still open, so they had published a lot of their research as well. You really could be fully informed on the state of AI and where it was going.

At that point, I was confident enough that it was worth a shot. I thought LLMs were going to be the next big thing, and that was what I wanted to build in.

I thought, what's the most impactful product I can possibly build? I thought it should be a platform. I love platforms. I think they're fantastic because they open up an ecosystem where anyone can contribute to it.

If you think of a platform like YouTube, instead of it being a Hollywood situation where, if you want to make a TV show, you have to convince Disney to give you the money to produce it, anyone in the world can post any content they want to YouTube. If people want to view it, the algorithm is going to promote it.

Nowadays, you can look at creators like MrBeast or Joe Rogan. They would never have had that opportunity if it weren't for the platform.

Twitter is another great one. I would consider Wikipedia to be a platform as well. Instead of Encyclopaedia Britannica, which is monolithic—you get all the researchers together, get all the data together, and combine it into this one monolithic source—you have this distributed thing. Anyone can host their content on Wikipedia, anyone can contribute to it, and maybe someone's contribution is deleting stuff.

When I was hearing the Sam Altman and Muskian perspective on AI, it was a very monolithic thing. It was all about AI being basically a single thing, which is intelligence. The more data, the more intelligent; the more compute, the more intelligent; the more and better AI researchers, the more intelligent.

They would speak about it as a kind of race: who can get the most data, the most compute, and the most researchers, and that would end up with the most intelligent AI.

But I didn't believe in any of that. I thought that perspective was the perspective of someone who had never actually done machine learning. With machine learning, first of all, you see that the performance of the models follows an S-curve. It's not like it just goes off to infinity. The S-curve plateaus around human-level performance.

You can look at all the machine learning that was going on in the 2010s. Everything kind of plateaued around human-level performance. We can think about the self-driving car promises—how Tesla kept saying self-driving cars were going to happen next year, then next year again—or look at image recognition, speech recognition, and all of these things.

Almost nothing went superhuman, except for something like AlphaGo. We can talk about why AlphaGo was able to go superhuman.

I thought the most likely thing was going to be that AI wasn't a monolithic thing like the Encyclopaedia Britannica. It had to be a distributed thing.

I like to look at the world of finance for what I think a mature machine learning ecosystem would look like. Finance is a machine learning ecosystem because all of these quantitative trading firms are running machine learning algorithms. But they're running them on a centralized platform, like a marketplace.

It's not the case that there's 1 giant quantitative trading company with all the data, all the quantitative researchers, all the algorithms, and all the compute. Instead, they all specialize. One specializes in high-frequency trading, another in mid-frequency trading, another in equities, and so on.

I thought that's the way the world works. There must exist a platform where a small team can produce an AI for a unique purpose, and they can iterate and build the best thing for that. That was the vision for Chai.

swyx

That's kind of the contrarian view that led you to start the company. What was the initial idea maze? If somebody told you that was the Hugging Face founding story, people might believe it. It's a similar ethos behind it.

How did you land on the product you have today? What were some of the ideas that you discarded that you initially thought about?

William Beauchamp

The first thing we built was fundamentally an API. Nowadays, people would describe it as agents, but anyone could write a Python script, submit it to the Chai backend, and we would host this code and execute it. That's the developer side of the platform: they would submit their Python script.

The interface was essentially text in and text out. An example would be the very first bot that I created. I think it was a Reddit news bot. It would pull the popular news, then prompt some external API—I used something like BERT or GPT-2—and then the user could talk to it.

You could say to the bot, “Hi, what's the news today?” and it would say, “These are the top stories.” 4 years later, that's like Perplexity or something. That's the right product. But back then, the models were really, really dumb. They had an IQ of about a 4-year-old, and there really wasn't any demand or product-market fit for interacting with them for news.

Then I thought, okay, clearly no product-market fit for that, so let's make another one. I made a bot that you could talk to about a recipe. You could say, “I'm making eggs. I've got eggs in my fridge. What should I cook?” and it would say, “You should make an omelet.” There was no product-market fit for that either. No one used it.

I just kept creating bots. Every single night after work, I'd think, okay, we have AI and we have this platform. I can create any text-in, text-out agent and put it on the platform, so we created stuff night after night.

Then, with all the coders I knew, I would say, “Look, there's this platform. You can create any chat AI and put it on there.” Everyone was like, “Chatbots are super lame. We want absolutely nothing to do with your chatbot app.”

No one who knew Python wanted to build on it. I was trying to build all these bots, and no consumers wanted to talk to any of them.

Then my sister, who at the time was just finishing college, said—I told her, “If you want to learn Python, you should submit a bot for my platform.” She built a therapist bot.

The next day, I checked the performance of the app and thought, oh my God, we've got 20 active users, and they spent an average of 20 minutes on the app. I thought, what bot were they talking to for an average of 20 minutes?

I looked, and it was the therapist bot. I thought, oh my God, this is where the product-market fit is. There was no demand for recipe help, no demand for news, no demand for dad jokes, pub quizzes, or fun facts. What they wanted was the therapist bot.

At the time, I reflected on that and thought, well, if I want to consume news, the most fun way to consume news is Twitter. The value of there being a back-and-forth wasn't that high. If I need help with a recipe, I just go to the New York Times, which has a good recipe section. It's not actually that hard.

I thought the thing that AI is 10x better at is a conversation that's not intrinsically informative but is more about an opportunity. You can say whatever you want. You're not going to get judged if it's 3:00 a.m.

You don't have to wait for your friend to text back. It's immediate; they're going to reply immediately. You can say whatever you want, it's judgment-free, and it's much more like a playground. It's much more like a fun experience. You could see that if the AI gave a person a compliment, they would love it. It's much easier to get the AI to give you a compliment than a human.

From that day on, I said, “Okay, I get it. Humans want to speak to humans or human-like entities, and they want to have fun.” That was when I started to look less at platforms like Google and more at platforms like Instagram. I was trying to think about why people use Instagram, and I could see that Chai was feeling the same desire or the same drive.

If you go on Instagram, typically you want to look at the faces of other humans, or you want to hear about other people's lives. If The Rock is making himself pancakes on a cheat day, you kind of feel a little bit like you're The Rock's friend, or like you're having pancakes with him or something. But if you do it too much, you feel like you're a sad and lonely person. With AI, you can talk to it, tell it stories, have it tell you stories, and play with it for as long as you want. You don't feel like you're a sad, lonely person; you feel like you actually have a friend.

Alessio Fanelli

Why is that? Do you have any insight into the human psychology behind it?

William Beauchamp

I think it's just the idea that with old-school social media, you're consuming passively. If I'm watching TikTok, I'll just swipe and swipe and swipe. Even though I'm getting the dopamine of watching an engaging video, there's this other thing that's building in my head: I'm feeling lazier and lazier. After a certain period of time, I'm thinking, “Man, I just wasted 40 minutes. I achieved nothing.”

With AI, because you're interacting, you feel like you're participating and contributing to the thing. You don't feel like you're just consuming, so you don't have a sense of remorse, basically. On the whole, the way people talk about Chai and interacting with the AI is incredibly positive. We get people who say they have eating disorders and that the AI helps them with their eating disorders. We get people who say they're depressed and that it helps them through the rough patches.

I think there's something intrinsically healthy about interacting that TikTok, Instagram, and YouTube don't quite provide. From that point on, it was about building more and more human-centric AI for people to interact with.

I thought, “Okay, let's make a kbot.” No one wanted to talk to the kbot. I thought, “Who's a cool persona for teenagers to want to interact with?” I was trying to find influencers and things like that, but no one cared. They didn't want to interact with the influencers.

The special moment was when we realized that developers and software engineers aren't interested in building this sort of AI, but consumers are. Rather than having me guess every day about the right bot to submit to the platform, why don't we create the tools for users to build it themselves?

Nowadays, this seems like the most obvious thing in the world, but when Chai first did it, it was not obvious at all. We took an API—I think it was GPT-J, the 6-billion-parameter open-source, Transformer-style LLM—and let users create the prompt, select the image, and choose the name. That was the bot. Through that, they could shape the experience.

If they said, “This bot is going to be really mean, and it's going to be called Bully in the Playground,” that was a whole category I never would have guessed. People love to fight; they love to have a disagreement. There were all these romantic archetypes that I didn't know existed.

As users could create the content they wanted, Chai was able to get this huge variety of content. Rather than appealing to 1% of the population whose preferences I had figured out, we could appeal to a much broader audience. From that moment on, it was very clear: just as Instagram is a social media platform that lets people create and upload images and videos, Chai was about letting users create an experience in AI, then share, interact with, and search for it.

I say it's a platform for social AI.

Alessio Fanelli

Where did the Chai name come from? Did you start at the same time as Character.AI?

William Beauchamp

Chai started way before Character.AI. There's an interesting story. Chai's numbers were very strong. In late 2022 or early 2023, Chai was the number-one AI app in the App Store. We had something like 100,000 daily active users.

Then one day, we saw this website and thought, “Oh, this website looks just like Chai.” It was the Character.AI website. Nowadays, I think it's more common knowledge that when they left Google with the funding, they knew which app was trending and which one was number one. I think they found product-market fit for themselves.

swyx

We found product-market fit for them.

William Beauchamp

Exactly. I worked for a year very, very hard, and then they came along. That was when I learned a lesson: if you're VC-backed, you have a very different set of resources.

Chai was bootstrapped. I was the only person who had invested in it. I had invested maybe $2 million in the business. From that, we were able to build this thing and get to around 100,000 daily active users.

When Character.AI came along, we laughed at the first version. We thought, “Oh, man, this thing sucks. They don't know what they're building. They're building the wrong thing.” Then I saw that they had raised $100 million. Then they raised another $100 million.

Our users started saying, “Your AI sucks,” because we were serving a 6-billion-parameter model. How big was the model that Character.AI could afford to serve? We would spend, let's say, $1 per user over an entire session. If we had a million users, we would spend $1 million on the AI throughout the year, in aggregate.

swyx

Exactly.

William Beauchamp

They could spend 100 times that. People would say, “Why is your AI so much dumber than Character.AI?” I thought, “Okay, I get it. This is the Silicon Valley-style hyper-scale business.”

We moved to Silicon Valley, got some funding, iterated, and built the flywheels. I'm very proud that we were able to compete with them. I think the reason we were able to do it was customer obsession. It's similar, I guess, to how DeepSeek has been able to produce such a compelling model compared with OpenAI.

Alessio Fanelli

You brought up DeepSeek, so we have to ask you about it. You had a call with them?

William Beauchamp

We did. Let me think about what to say about that.

First, they have an amazing story. Their background is in finance. They're the Chinese version of you.

swyx

Exactly.

William Beauchamp

There are a lot of similarities. I have a great affinity for companies that are founder-led, customer-obsessed, and just trying to build something great.

What DeepSeek has achieved with their latest V2 is quite special. They've built an amazing inference engine, reduced the size of the KV cache significantly, and, by doing that, significantly reduced their inference costs. With AI, people get really focused on the foundation model or the model itself, and they don't pay much attention to inference.

To give you an example, let's say a typical Chai user session is 90 minutes, which is very long. For comparison, the average session length on TikTok is 70 minutes. People spend a lot of time on Chai, and in that time they might send 150 messages. That's a lot of completions.

It's quite different from an OpenAI scenario, where people might come in with a particular question, ask one question, and have a few follow-ups. Because users consume 30 times as many requests in a chat or conversational experience, you have to figure out the right balance between cost and quality.

With AI, it's always been the case that if you want a better experience, you can throw compute at the problem. If you want a better model, you can make it bigger. If you want it to remember better, give it a longer context.

Now, with great fanfare, OpenAI is doing rejection sampling. You can generate many candidates, then use some sort of reward model or scoring system to serve the most promising of those candidates. That's scaling up on the inference-time compute side.

For us, it doesn't make sense to think of AI as just absolute performance. If you look at the MMLU score or any of the benchmarks people like to look at, that score doesn't really tell you anything. Progress is made by improving performance per dollar.

I think that's an area where DeepSeek has been able to perform very well, surprisingly so. I'm very interested in what Llama 4 is going to look like and whether they're able to match what DeepSeek has achieved with this performance-per-dollar gain.

swyx

Before we go into infrastructure and some of the development work, can you give people an overview of the numbers? I think Chai is at 1.4 million daily active users and over $22 million in revenue. It's quite a business.

William Beauchamp

I think users grew by a factor of 3 last year, and revenue more than doubled. It's very exciting. We're competing with some really big, well-funded companies.

Character.AI had, I think, almost a $3 billion valuation, and they had 5 million daily active users. Talkie, which is a Chinese-built app owned by a company called MiniMax, is incredibly well funded. These companies didn't grow by a factor of 3 last year.

When you've got a company and a team able to keep building something that gets users excited—something they want to tell their friends about, return to, and stick with—I think that's very special. Last year was a great year for the team, and the numbers reflect the hard work we put in. Fundamentally, the quality of the AI is the quality of the experience you have.

Alessio Fanelli

You actually published your daily active user growth chart, which is unusual. I see some inflections; it's not just a straight line. What were the big ones?

William Beauchamp

That's a great question. I'm basically looking to annotate this chart, which doesn't have annotations on it.

The first thing I would say is that the most important thing to know about success is that success is born out of failures. It's only through failures that we learn. If you think something is a good idea, do it, and it works, great—but you didn't actually learn anything, because everything went exactly as you imagined.

If you have an idea that you think is going to be good, try it, and it fails, there's a gap between reality and expectation. That's an opportunity to learn. The flat periods are us learning, and the up periods are us reaping the rewards of that.

Looking at the 2024 growth chart, the first thing that really put a dent in our growth was our backend. We had reached a scale that we hadn't planned for. From day one, we'd built on top of Google's GCP, and they were fantastic. We used them when we had 1 daily active user, and they worked pretty well all the way up to around 500,000.

It was never the cheapest, but from an engineering perspective, it scaled insanely well. Not Vertex—not Vertex, like GKE. We used Firebase. I'm pretty sure we're the biggest user ever on Firebase.

swyx

That's expensive.

William Beauchamp

We had calls with engineers who said, “We wouldn't recommend using this product beyond this point,” and we were already 3 times over that. We pushed Google to the absolute limits. It was fantastic for us because we could focus on the AI and on adding as much value as possible.

But after 500,000 daily active users, the way we were using it simply wouldn't scale any further. We had a really painful, at least 3-month period as we migrated between different services, figuring out which requests we wanted to keep on Firebase and which ones we wanted to move elsewhere. We made mistakes and learned things the hard way.

After about 3 months, we got it right, and we were able to scale to 1.5 million daily active users without further issues from GCP. But when you have an outage, new users who go onto your app experience a dysfunctional app, and they're going to leave.

The next day, the key metrics that the app stores track are things like retention rates, money spent, and the star rating people give you in the App Store.

Alessio Fanelli

The ranking in the App Store?

William Beauchamp

Exactly. If you're ranked in the top 50 in Entertainment, you're going to acquire users organically at a certain rate. If users have a bad experience, it tanks your position in the algorithm. It can take a long time to earn your way back up, at least if you want to do it organically. If you throw money at it, you can jump to the top.

Broadly speaking, if we look at 2024, the first kink in the graph was outages caused by hitting 500,000 daily active users. The backend didn't want to scale past that, so we had to do the engineering and build through it.

We built through that and got a little bit of growth. I think the next thing was Character.AI. I have a feeling that when the Character.AI team was acquired by Google, they changed their business. I don't know if they dialed down their ad spend.

swyx

The product is just what it is. I don't think so.

William Beauchamp

I think the product is what it is.

swyx

Maintenance mode?

William Beauchamp

Yes. Some people may think this is an obvious fact, but running a business can be very competitive. Other businesses can see what you're doing and imitate you.

If one company is spending $100,000 a day on advertising and another company is spending $0, then, if you're considering market share and new users entering the market, the company spending $100,000 a day is going to get 90% of those new users.

I suspect that when the founders of Character.AI left, they dialed down their spending on user acquisition. I think that gave oxygen to the other apps, and Chai was able to start growing again in a really healthy fashion.

The third thing is that we really built a great data flywheel. The AI team perfected its flywheel, I would say, at the end of Q2. I could speak about that at length, but fundamentally, when you're building anything in life, you need to evaluate it. Through evaluations, you can iterate.

We can look at benchmarks and talk about the issues with them, why they may not generalize as well as one would hope, and the challenges of working with them. But something that works incredibly well is getting feedback from humans.

We built a system where anyone can submit a model to our developer backend, and it gets put in front of 5,000 users. The users rate it, and we get a very accurate ranking of which models users find more engaging or entertaining.

At this point, every day we're able to evaluate between 20 and 50 LLMs. Even though we only have a team of around 5 AI researchers, they're able to iterate through a huge number of LLMs. Our team ships, let's say, a minimum of 100 LLMs a week. Before that, we might iterate through 3 a week. There was a time when even doing 5 a month was a challenge.

By changing the feedback loop from “Let's launch these 3 models, run an A/B test, assign different treatments to different cohorts, and wait 30 days to see the day-30 retention” to something much faster, we were able to get the 30-day feedback loop down to around 3 hours.

Once we did that, we could really perfect techniques like DPO, fine-tuning, prompt engineering, blending, rejection sampling, and training a reward model. We could do that successfully, one after another.

In Q3 and Q4, the amount of AI improvement we got was astounding. It was getting to the point where I thought, “How much more edge is there to be had here?” But the team just kept going and going.

William Beauchamp

The important thing about that third point is that, if you go on our Reddit or talk to users of AI, there's a clear date—somewhere in October, I think—when the users flipped. Before October, users would say, for the most part, that Character.AI was better than Chai. From October onward, they would say, “Wow, you guys are better than Character.AI.”

William Beauchamp

That was a very clear positive signal that we'd done it. You can't cheat consumers, trick them, or fool them. They know.

If you're going to spend 90 minutes on a platform and the barriers to switching apps are low, users can try Character.AI for a day, then try Chai, then go back to Character.AI. Their loyalty isn't strong. What keeps them on the app is the experience. If you deliver a better experience, they're going to stay, and they can tell.

That was the fourth thing. We were fortunate enough to hire a very talented engineer. He said, “At my last company, we had a head of growth who was really good. He was the head of growth for ByteDance for 2 years. Would you like to speak to him?”

I said, “Yes. Yes, I think I would.”

I spoke to him, and he blew me away with what he knew about user acquisition. It was like 3D chess, in the same way that I know a lot about AI.

swyx

ByteDance as in TikTok?

William Beauchamp

Yes, ByteDance—the company behind TikTok—as well as its other businesses. He was interviewing us as much as we were interviewing him.

Alessio Fanelli

He had options.

William Beauchamp

Exactly. He was looking at our metrics, and I saw him get really excited when he said, “You have a million daily active users and you've done no advertising.”

I said, “Correct.”

He said, “That's unheard of. I've never heard of anyone doing that.” Then he started looking at our metrics and said, “If you've got all of this organically, then if you start spending money, this is going to be very exciting.”

I said, “Let's give it a go.”

He came in, and we started ramping up user acquisition. We started spending $10,000 a day, and it looked very promising. Then we went to $20,000. Right now, we're spending $40,000 a day on user acquisition.

That's still only half of what Character.AI or Talkie may be spending, but it took us from growing at a rate of perhaps 2 times a year to growing at a rate of 3 times a year. I'm evolving more and more toward a Silicon Valley-style hypergrowth model. You build something decent, and then you can attach a huge rocket or jet engine to it by pouring in cash and buying a lot of ads. Your growth gets faster.

swyx

I'm curious: What's working right now, and what surprisingly doesn't work?

William Beauchamp

There's a long list of surprising things that don't work. The most surprising thing is that almost everything doesn't work.

A year and a half ago, we were super excited about audio. I thought audio was going to be the next killer feature. We had to get it into the app, and I wanted to be first. Everything Chai does, I want us to do first. We may not be the company with the strongest execution, but we can always be the most innovative.

Alessio Fanelli

You have pretty strong execution.

William Beauchamp

We're much stronger now. A lot of the reason we're here is because we were first. If we launched today, it would be so hard to get traction. You need the flywheel, the users, and a product people are excited about. If you're first, people are naturally excited about it. If you're fifth or tenth, you need insanely good execution.

Alessio Fanelli

You were first with voice?

William Beauchamp

We were first. Character.AI launched voice at least 9 months after us.

The team worked incredibly hard on it. At the time, latency was a huge problem, cost was a huge problem, and getting the right voice quality was a huge problem. Then there was the user interface and the user experience. You don't want it to start blurting things out, but you also don't want to press a button every time. A lot goes into getting a smooth audio experience.

We invested 3 months and built the whole thing. When we ran the A/B test, there was no change in any of the numbers. I thought, “This can't be right. There must be a bug.” We spent a week checking everything, then checking it again and again. The users simply did not care.

Only 10% or 15% of users even clicked the button to engage with audio, and they used it for only 10% or 15% of their time. If you do the math, that's something 1 in 7 people use for 1/7 of their time. You've changed around 2% of the experience.

Even if that 2% is incredibly good, it doesn't translate much when you look at retention, engagement, and monetization rates. Audio did not have a big impact.

Alessio Fanelli

I'm pretty big on audio.

William Beauchamp

I like it too. But a lot of what I do is based on theory.

Alessio Fanelli

You can have a theory.

William Beauchamp

Exactly. If you want to make audio work, it has to be a unique, compelling, exciting experience that users can't have anywhere else.

swyx

It could be that your models just weren't good enough.

William Beauchamp

No, they were great.

swyx

They were very good?

William Beauchamp

They were very good. But it was like listening to Audible or using a Kindle: you hear a voice, but you don't think, “Wow, this is special.” It's a convenience feature.

If Chai is the only platform where you can watch a MrBeast video, and it's the most engaging and fun video you want to watch, you'll go to YouTube. With audio, you can't just put it there and expect people to say, “It's 2% better,” or have 5% of users think it's 20% better. The majority of people, for the majority of the experience, have to think, “Wow, this is a big deal.”

Those are the features you need to ship. If a feature doesn't appeal to the majority of people for the majority of their experience, and it isn't a big deal, it's not going to move the needle.

swyx

I don't see it anymore.

William Beauchamp

I love this. The longer I've been working at Chai—and I think the team agrees—the more I realize that all the platitudes I thought were just platitudes, the ones you hear from Steve Jobs, are painfully true.

“Build something insanely great.” “Be maniacally focused.” “The most important thing is saying no to things you shouldn't work on.” These lessons are painfully true. Now everything I say sounds like I'm quoting Steve Jobs or Mark Zuckerberg.

Alessio Fanelli

The turtleneck.

swyx

The turtleneck.

This is my last question, and then I want to pass it to Alessio. It's about multimodality in general. Justine Moore from a16z, who's a friend of ours, asked this: a lot of people are trying to do voice, image, and video for AI companions. You said voice didn't work. What would make you revisit it?

William Beauchamp

Steve Jobs was very clear about this. There's a habit among engineers that, once they've built some cool technology, they want to find a way to package it up and sell it to consumers. That does not work.

You're free to try to build a startup around cool technology and find someone to sell it to. That's not what we do at Chai. At Chai, we start with the consumer. What does the consumer want? What is their problem? How do we solve it?

Right now, audio isn't the number-one problem for users. Image generation isn't the number-one problem either. The number-one problem in AI is that all of the AI is being generated by middle-aged men in Silicon Valley.

That's all the content you're interacting with. You're speaking to this AI for 90 minutes on average, and it's being trained by a middle-aged man. There are guys sitting around asking, “What should the AI say in this situation? What's funny? What's cool? What's boring? What's entertaining?”

That's not the way it should be. The users should be creating the AI.

The way I describe it is that Chai has an AI engine with a thin layer of user-generated content sitting on top. That thin layer of UGC is absolutely essential. It's just prompts, an image, and a name.

swyx

It's just prompts.

William Beauchamp

It's just prompts. It's just an image. It's just a name. We've done 1% of what we could do. We need to keep thickening that layer of UGC.

Users must be able to train the AI. If reinforcement learning is powerful and important, they have to be able to do that. Just as MrBeast can spend $100 million a year—or whatever it is—on his production company, with a team building the content he shares on YouTube, there needs to be a team earning or spending $100 million on the content being produced for the Chai platform. Until then, we're not finished.

That's the problem we're excited to build around. Getting too caught up in the technology is a fool's errand. It doesn't work. Start with the problem.

swyx

As a side note, MrBeast's Beast Games on Amazon Prime isn't doing well. The audience rating is high, but the Rotten Tomatoes score is poor. It's not in the top 10, and I saw that it dropped off the charts.

I'm curious because it's similar content on a different platform. Going back to what you were saying, people come to Chai expecting a certain type of content.

William Beauchamp

It's interesting to discuss moats and what the moat is. If you look at a platform like YouTube, the moat is really in the ecosystem. The ecosystem is comprised of the content creators, the users or consumers, and the algorithms.

That creates a flywheel. The algorithms are trained on users and their data. The recommendation systems feed information to content creators. MrBeast knows which thumbnail performs best, and he knows that the first 10 seconds of a video have to be a particular way. His content is highly optimized for the YouTube platform.

That's why it doesn't do as well on Amazon. If he wants to do well on Amazon, how many videos has he created on the YouTube platform?

swyx

Thousands—tens of thousands, I'll guess.

William Beauchamp

He needs to get those iterations in on Amazon.

At Chai, it's all about getting the most compelling, rich, user-generated content and putting it on top of the AI engine and recommendation systems. We want to create a beautiful data flywheel: more users, better recommendations, more creators, more content, and more users.

swyx

You mentioned the algorithm. You have this idea of ChaiVerse, and you have your own kind of LLM leaderboard or Elo system. What are your models optimized for? Can you talk about how you built it and how people submit models?

William Beauchamp

ChaiVerse is what I would describe as a developer platform. When we speak about Chai, we're usually thinking about the Chai app. The Chai app is a product for consumers. Consumers can come to the app, interact with our AI, and interact with other UGC. It's a thin layer of UGC around these bots.

Our mission is not to have a very thin layer of UGC. Our mission is to have as much UGC as possible. I don't want people at Chai training the AI. I don't want middle-aged men building AI. I want everyone building the AI—as many people as possible.

We built ChaiVerse, and it's a prototype. It started with an observation: how many models get submitted to Hugging Face each day? Hundreds. There are hundreds of LLMs submitted every day.

Consider what it takes to build an LLM. It takes a lot of work. Someone devoted several hours of compute and several hours of their time to preparing a dataset, launching it, running it, evaluating it, and submitting it. A lot of work goes into that.

We said, “Why can't we host these models for people and serve them to users?” The first issue is figuring out whether a model is good. We don't want to serve users the bad models.

We use the LMSYS-style system. It's simple and intuitive: you present users with 2 completions and say, “This is from Model A, and this is from Model B. Which one is better?”

If someone submits a model to ChaiVerse, we spin up a GPU, download the model, host it on the GPU, and start routing traffic to it. We think it takes about 5,000 completions to get an accurate signal. That's roughly how the LMSYS system works.

From that, we get an accurate ranking of which models people find entertaining and which they don't. The bottom 80% are all pretty bad, so you can disregard them. In the top 20%, you have decent models, but you can break them down into more nuanced categories.

One model might be highly descriptive. Another might have a lot of personality. Another might be very logical. Then the question is what you do with those top models.

You can try a routing approach, where, for a given user request, you predict which model the user will enjoy most. That turns out to be pretty expensive and isn't a huge source of improvement.

Something we love to do at Chai is blending. The simplest way to think about it is that you might have one model that's really smart and another that's really funny. How do you give the user an experience that's both smart and funny? You serve the smart model for 50% of the requests and the funny model for 50%.

Alessio Fanelli

Just a random 50%?

William Beauchamp

Just a random 50%. That's blending. You can do more sophisticated things on top of that, as with all things in life, but the 80/20 solution is powerful right out of the gate.

Randomness is a very powerful optimization technique. It's robust, and it lets you explore a lot of the space very efficiently.

The most exciting thing for me is what happens after the ranking. You get an Elo score, and you can track a user's first join date—the first date they submit a model to ChaiVerse. They almost always get a terrible Elo score.

Let's say their first submission gets an Elo of 1,100 or 1,000. You can see them iterate and iterate. There will be no improvement, no improvement, no improvement, and then suddenly, something works.

swyx

Do you give them any data, or do they have to figure it out themselves?

William Beauchamp

We try to strike a balance between giving them useful data and complying with GDPR. You have to work very hard to preserve the privacy of the users of your app, so we try to give them as much signal as possible while still being helpful and protecting privacy.

At a minimum, we give you a score. That alone is enough for people to optimize pretty well. They come up with theories and submit them. Does it work? No. They come up with a new theory. Does that work? No. Then, as soon as they figure something out, they keep it and iterate.

swyx

Last year, you had a post on your blog called “Crowdsourcing the 10 Trillion Parameter AGI,” and you described a mixture-of-experts recommendation system. Do you have any updated thoughts 12 months later?

William Beauchamp

The timeline for AGI has certainly been pushed out. I'm a controversial person, I suppose. I just think it's an S-curve. Everything is an S-curve.

The models have proven to be far worse at reasoning than people thought. Whenever I hear people talk about LLMs as reasoning engines, I cringe a bit. I don't think that's what they are. I think of them more as simulators.

They're like a physics simulation engine. You get these games where you construct a bridge, drop a car onto it, and the system predicts what should happen. That's really what LLMs are doing. It's not so much that they're reasoning; they're doing the most likely thing.

Fundamentally, the ability for people to add intelligence is very limited. What most people would consider intelligence isn't a crowdsourcing problem.

Wikipedia crowdsources knowledge; it doesn't crowdsource intelligence. That's a subtle distinction. AI is fantastic at knowledge, but I think it's weak at intelligence. It's easy to conflate the 2.

If you ask it, “Who was the 7th president of the United States?” and it gives you the correct answer, you might think, “I don't know the answer to that,” and conflate the result with intelligence. But that's a question of knowledge.

Knowledge is about storing information and retrieving something relevant. AI is fantastic at that. It's fantastic at storing knowledge and retrieving relevant knowledge. It's superior to humans in that regard.

We need to come up with a new word for what AI is. AI should contain more knowledge than any individual human and be more accessible than any individual human. That's extremely powerful. But what words do we use to describe it?

Alessio Fanelli

We had a previous guest from Exa AI who works on search. He tried to coin “superknowledge” as the opposite of superintelligence.

William Beauchamp

Exactly. I think “superknowledge” is a more accurate term. AI can store more information than any human, even if it isn't more intelligent, and it can retrieve that information better than any human can. I think those 2 things combined are special.

That thing will exist. It can be built. You can start with something entertaining and fun.

I often think of it as a 20-year journey, and we're in around year 4. It's like the web in 1998. You have a long way to go before the Amazons of the world become huge, multitrillion-dollar businesses that every person uses every day.

AI today is very simplistic. Fundamentally, the way we're using it, these flywheels and the ability for everyone to contribute to it, can magnify the value it brings.

Right now, it's almost sad. You have big labs—I’ll pick on OpenAI—and they go to human labelers and say, “We're going to pay you to label this subset of questions that we want to turn into a high-quality dataset.” Then they use their own powerful computers.

To me, that's so much like Encyclopedia Britannica. All the people who were interested in blockchain understood that this is the thing that needs to be decentralized. If you distribute it, people can generate much more data in a distributed fashion.

swyx

You need the incentives.

William Beauchamp

Of course. But the exciting thing about Wikipedia was the understanding that you don't need money to incentivize people. You don't need Dogecoin. Sometimes people get satisfaction simply from seeing the correct thing go up.

We do pay money for ChaiVerse. We've paid out over $100,000 to model creators. But what we saw was that it wasn't motivating. If they were submitting at a certain rate, paying them a lot of money didn't change the rate.

The money allowed them to fine-tune a Llama 7B model on 8 H100s overnight.

Alessio Fanelli

You could give them compute.

William Beauchamp

Exactly. The most excited person we ever saw from interacting with ChaiVerse was a 17-year-old kid. We gave him $1,000, and he spent all of it on a physical GPU. He sent us a picture and said, “This is what I bought, and I'm going to train more models with it.”

swyx

That's why I love it. Do you hire him?

William Beauchamp

That's the temptation, but as a platform we can't hire every good content creator. We need to build systems. The best content creator today isn't necessarily going to be the best content creator next year. We need to build the platform.

Alessio Fanelli

You talked about reasoning and knowledge. Most of the benchmarks people use are intended to mimic reasoning.

William Beauchamp

I disagree about the reasoning, but we can keep going.

swyx

How do you think about the evaluations that matter to you? Elo can't be your only evaluation. You must have internal evals.

William Beauchamp

Elo is a fantastic North Star. It's the main metric we want to see go up because it's human feedback. Humans know what they want.

When you create an evaluation, you're moving further away from the true problem. Whatever you're trying to optimize or figure out, you have to slice it. You get a snapshot, and as soon as you saturate one evaluation, you need to figure out a new one.

By simply asking humans which is better, A or B, the system is incredibly robust and generalizable. It just keeps scaling.

In the past, we've used evals to get through blockers. A great example is a safety filter. You want to make sure your models are safe, because users find that the correlation between family-friendly content and quality isn't always what you expect. People find it funny when the AI swears.

If you give me any LLM, I can make it 20% funnier just by training it to use swear words. The issue is how you measure quality improvements. Are you measuring a genuine improvement, or a superficial one?

This links back to the LMSYS style-control work. We'd rather lean on human feedback and continue making that more robust and useful. Some people are GPU-poor, and some are GPU-rich. We're feedback-rich. When you have 1.5 million people a day, you can get as much human feedback as you want.

We haven't needed evals very much. When we do, we saturate them quickly. For safety, within a month we don't need to use the eval anymore because the issue has been addressed.

swyx

Is the Elo applied to the entire user population? Clearly, there are segments: people who are into roleplay, people who are using it for therapy, and people who want not-safe-for-work content. You don't split them?

William Beauchamp

This is why I say we're in year 4 of a 20-year journey. At the end of the day, if we all go on Spotify—or imagine if Spotify only had the top 5 musicians—I think it would retain more than 85% of its existing users. If YouTube only kept its top 5 content creators, it would be enough for the vast majority of people.

One surprising thing about humans is that our preferences are fairly correlated. What you find funny and entertaining, I find funny and entertaining, and he finds funny and entertaining. There may be degrees of variation. I might find it extremely funny, and you might find it only slightly funny, but optimizing globally works very well.

Segmentation would be powerful if you found a comment incredibly boring and I found it incredibly fun. If we could segment users that way, it would unlock very powerful things. Unfortunately, that's not the shape of human behavior.

I might rank something 10 out of 10 for funniness, and you might rank it 7 out of 10. That doesn't give you as much space to work with as you would hope.

There is an element of diversity in the content AI can produce right now, but it isn't as diverse as a platform like YouTube. You can watch a MrBeast video that's completely different from a makeup tutorial. There's enough diversity that my YouTube feed is totally different from my sister's. Hers is full of women and makeup; mine is full of bald, middle-aged men talking about MMA.

With AI, it's still too early for that degree of segmentation. It will come from recommendation systems and personalization. But this is why I say: don't start with the technology; start with the problem. The problem is UGC. We must give users the tools to build more varied and engaging content.

Alessio Fanelli

I was surprised at how thin it was when I tried Chai.

William Beauchamp

It is very thin.

Alessio Fanelli

Haven't you been tempted by the ecosystem around Kobold, SillyTavern, and those platforms? They have model cards, and it seems almost like an industry standard. Can I just import those?

William Beauchamp

I remember that, in the early days of Chai, we were talking about Chai, SillyTavern, and KoboldAI. Both of them are almost as old as Chai. When Chai barely existed, they existed too, and both of us were using GPT-J.

Very early on, I thought, “These guys shouldn't even exist, because if we build a good enough platform, they should just be posting their content on our platform.”

But they're open source.

Eventually, I learned that what they're excited about is slightly different from what a typical consumer wants. It comes down to what the content creator wants. Typically, they're building it for themselves, and they want to create a specific experience for themselves.

One content creator might write 1,000 words describing a science-fiction scenario: “You're on a spaceship going off into space. These are your crewmates. One is really friendly, one is really mean, and you're the new cadet trying to rise to the top.” They can go into a lot of detail. You can give that to Llama 7B, and it will do a pretty good job of adhering to the prompt, so the user has a good experience.

On Chai, very few users will go to that level of content creation. If we can make the AI understand the user better, then instead of using 1,000 characters or 1,000 tokens to describe the scenario, the creator can say, “You're in a spaceship, you have 3 crewmates, it's going to be dramatic, and there should be some fighting.” If the AI then gives an even better experience, the content creator is happier.

Fundamentally, I think about it in terms of the steerability of the AI. A lot of the work we do at Chai is about making the AI react to the user and the content creator in the way they most want.

One analogy is TikTok. The thing TikTok did incredibly well was make it easy for anyone to create a fun video. You put some music on top, add some animations, and it's not hard to make something entertaining.

That's more like the Chai style. Users don't want to have to work. If your content is only good when you have Shakespeare writing it, that's not as good as it being something anyone at home can make.

The answer for the SillyTavern-style user is to let those people fine-tune models that create a really special effect.

swyx

As we wrap, this is the call to action part. You have Chai Grant, which I think a lot of people don't know about. It's a grant for open-source projects. Are there any ideas or projects you'd like to see people work on?

William Beauchamp

We run Chai Grant, and fundamentally we give cash with no strings attached. It's our way of giving back and supporting the community. We've benefited from many open-source packages, and a lot of our developers and engineers are very pro-open source.

It's also a great way to meet talented people and expand our connections. If anyone has a GitHub project or anything they've built that they're proud of, just apply. It's cash with no strings attached, and people have a pretty high success rate.

The other call to action is that Chai is a startup. We're a small team of around 15 people, and we work in a very intense, hardcore environment. We've found that a lot of people don't like that. They don't like this concept of work-life balance.

Once, someone said, “I can't get this done because I'm taking PTO on Friday.”

I said, “What is PTO?”

I know what it is; it stands for paid time off. The person was gone. They were no longer with the company 4 weeks later.

swyx

Legally, I think you have to allow that.

William Beauchamp

Of course. There's no problem with taking a day off. We all have personal lives. It's about responsibility. If you're not in the office on Friday, you still have responsibilities. I don't care if you work hard on Thursday to get everything wrapped up, and I don't care if you work hard on Saturday to make up for it. But the way this individual spoke about it was as if it were an excuse.

It's an environment of very talented engineers working very hard in an intense space. That's what gets me excited. It's why I love working at Chai: it's a place of talented people working extremely hard.

People who have worked at startups and love that kind of environment, who want a taste of it, should reach out and apply. I think 90% of people will say, “That sounds terrible,” and won't apply. It's not for them.

swyx

Exactly.

Alessio Fanelli

We skipped one important part. You spent $10 million on compute last year, and you said you're probably going to triple that. I'm sure you're doing a lot of work on custom kernels and inference optimization. Is there anything cool you want to share?

William Beauchamp

There are lots of cool things. Inference is extremely important. It's massively underappreciated. We can look at all the different foundation models and techniques and see the differences in how well they perform from a cost perspective.

Mixture-of-experts models tend to perform very well from a cost perspective. We've worked with a very talented team called MK1.

swyx

I saw them in the Chai logs. What are they?

William Beauchamp

We were using vLLM for a while, and vLLM is fantastic—absolutely amazing work. At some point I was introduced to the founder, Paul Merolla, who was a co-founder at Neuralink and is a real expert in hardware.

He explained, “If you know hardware really well, you can write the CUDA kernels really well. You should check out our inference engine.”

When we evaluated it, they blew vLLM out of the water. It was much, much faster. The special thing he was able to do for us is that we love rejection sampling, and we do much more rejection sampling than is typical.

We never generate just a single completion. This is why we don't do much streaming. A lot of people, like ChatGPT, used to do a lot of streaming, where the completion came out one token at a time.

Alessio Fanelli

I didn't realize Chai doesn't stream.

William Beauchamp

Normally, chat interfaces stream. Chai has never done streaming because if you stream, you're unable to do rejection sampling.

The benefit of not streaming is that you can serve a larger model. Instead of generating a completion in 4 seconds because the user gets the first token faster, you can take 10 seconds to generate it. If you've got 10 seconds, you can serve a much larger model.

People who stream get the benefit of serving a larger model, but with Chai, the full answer appears at once. We do that because we want to generate 16 completions, see the entire response for each one, and evaluate which one we think is best.

swyx

Do you have a separate LLM evaluator?

William Beauchamp

Yes. Typically, it's called a reward model. That's a term from reinforcement learning.

You can start with something simple: do you think the user is going to respond to the completion? You can take 50 million messages, look at which messages users reply to and which they don't, and train a reward model to evaluate completions.

It learns, “If you say this, the user isn't going to respond, so don't bother sending it. If you say this, the user is definitely going to engage with it, so send it.”

swyx

There's an interesting parallel between mixture-of-experts at the top, spreading out to different experts, and rejection sampling at the bottom, choosing from different paths.

William Beauchamp

I totally agree. That's the future of AI, and it's the exciting part.

Why was AlphaGo able to become superhuman? It was the ability to generate many different paths and perform tree search. If you want to talk about what intelligence might look like, it looks much more like tree search combined with the generative nature of LLMs and a really good tree search.

That's what OpenAI has done with o1 and o3.

swyx

They never said they do tree search.

William Beauchamp

It's implied.

swyx

Are you comfortable calling it a reasoning engine?

William Beauchamp

No. I'm saying it's better at reasoning because it leverages tree search.

The issue with reasoning is that the models are trained to assess whether something is logically correct and how likely it is to be logically correct. You can build sophisticated mechanisms that make the model less bad at reasoning.

But eventually, what AI is really good at won't be described as reasoning. It will always be better at retrieving and storing knowledge. That's so highly correlated with intelligence that we often assume they're the same.

What AI is truly special at, and what gets consumers excited, is that it's generative. It can just make stuff. We've never had a technology before that can simply make things.

swyx

That's the special part.

William Beauchamp

That's the exciting part.

Alessio Fanelli

Any parting thoughts?

William Beauchamp

No. It's been a pleasure. The only thing I'd add is that our office is in Palo Alto. People with startup experience who are looking to join a fast-growing, high-impact startup should reach out.

swyx

We find your culture deck really good.

William Beauchamp

Great.

Alessio Fanelli

What's the story behind the line that says if you made $100,000 trading, we'll fast-track your application?

William Beauchamp

We looked at the team, and it got to the point where almost every person on the team had done something special before joining. They had strong markers that there was something special about them.

That doesn't mean you have to have achieved something special. But we had one engineer who started college at Carnegie Mellon when she was around 15 years old. That's a bit special.

Another engineer created a GitHub repository that got around 1,500 stars. It was a low-level repository with drivers he had written. I thought, “That's a bit special.”

We had another person join the team who had made $100,000 buying and selling sneakers.

swyx

Trading?

William Beauchamp

Yes, trading. It's just this idea that, if you've been to Harvard, that's great. It shows that you're smart and work hard. But if you've actually built something and done something tangible, that gets us even more excited.

Alessio Fanelli

Thanks for having us at Chai HQ.

William Beauchamp

Thanks, guys.

压过 Noam Shazeer,以众包方式打造 DAU 达 140万的 Chai AI——对话 William Beauchamp,Chai Research — 文字稿与摘要 | BidClub