“模糊现实”——Chai 的社交 AI 平台(赞助)
Chai称,2021年上线、早于ChatGPT之后,仅靠13–14名工程师就实现了约1000万活跃用户和3000万美元收入。 它的核心判断是,用户而非风险投资人才是真正的客户,因此可以把订阅收入再投入AI,同时维持异常高的人才密度。
这款产品并不追求成为全能最聪明的助手,而是让用户创建角色,进行没有现实后果的社交模拟。 Will Beauchamp追问,为什么训练AI只能由“恰好在湾区做软件工程的中年男性”完成;一个少女或许能做出最好的美妆聊天AI,并发现有数十万用户需要它。
Chai把留存视为真正的扩展指标;据称,一次6B参数模型实验让平均对话时长提升70%,30日留存提升超过30%。 它从每天生成约100分钟对话的活跃用户身上挖掘重试、编辑、截图、删除和分享等反馈,再将这些隐性偏好转化为RLHF奖励信号。
单纯优化互动时长会带来Goodhart问题,Chai通过长期回访行为来衡量这一风险。 模型可以在每条回复结尾都提问来延长会话,实质上是在“利用人类行为漏洞”,但30日或60日留存可能反而低于基线,因为盘问并不等于持久陪伴。
Chai提出的成本—性能突破口是模型混合:逐条消息随机切换彼此正交的小模型,让创造力、实用性与不可预测性共存。 Tom Lu称,2到3个专用模型一度能够击败GPT-3.5;如今生产测试使用约7到10个模型的组合,每个cohort包含几千名用户,再根据回访率选出胜者。
平台的核心风险在于,同一套制造安慰与亲密感的机制,也可能被用于优化有害注意力。 Chai结合硬性禁令、社区举报、人工审核、正则表达式和学习型审核;Beauchamp称,自杀防护措施并未损害留存和增长。对访谈证据的复核显示,相关干预只能带来小幅、短期的心理健康改善,不能替代面对面的CBT。
OpenAI转向GPT-4o,被视为品类验证,而非赢家通吃的威胁。 主持人将4o称为“社交蝴蝶”,4.1是编程专家,o3是审慎推理者;Beauchamp预计AI领域会诞生许多数千亿美元乃至万亿美元级公司,并称竞争“让团队保持饥饿感”。
1. Chai把创作权交给用户后,发现了社交AI
Chai于2021年上线,最初是一个允许用户创建并部署自有AI的平台。Beauchamp称,这个开放式工具意外揭示了“这种社交AI品类”,随后平台增长到约1000万活跃用户。
Beauchamp认为,Chai与ChatGPT的差异不仅在技术,也在组织逻辑:它没有只追求“世界上最聪明的AI”,而是追问为什么模型只能由湾区中年软件工程师来打造。一个少女制作的美妆教程AI,或许更懂用户想要什么样的体验;而成千上万甚至数十万人可能共享她的品味。
他的媒体类比是:听Joe Rogan 45分钟后,会觉得“我好像正和这帮人待在一起”。LLM让用户从被动的拟社会消费进一步参与其中:用户参与塑造互动,因此不会留下传统社交媒体那种懒散感。
Beauchamp最有力的类比来自女儿玩娃娃:她们明知娃娃不是真的,却仍把它们当作真实存在。他认为,很多成年人也会对AI说“我爱你”,由此建立社交“线路”,获得愉悦;这种体验可能让他们面对真实的人际关系时更健康、更积极。
2. 社交模拟先消除后果,沉浸再消除距离
主持人把社交媒体重新定义为地位竞争与模拟的结合:人们像梦境运行情景一样,用现实检验身份和关系。但公开互动伴随着后果——“你会惹错人”——AI则允许用户探索反事实情景,不必承担声誉层面的爆炸半径。
Beauchamp说,这种环境可以被称为“安全空间”,但它同样有趣。用户可以探索自己粗鲁、善良、友好或敌对时会发生什么,并获得“某种真实的人类反应,同时又没有真正的人类参与其中的风险”。
他的终局设想是一个VR世界:里面有一个类似Joe Rogan、能讨论Trump关税的知识型角色,有一个风趣的伙伴,也有一个能让用户感到被爱的角色,丰富程度类似World of Warcraft。文字交互“基本已经到位”;高质量实时音频可能还需2到4年,视频成本仍高出1到2个数量级,而完整VR愿景可能需要10年。
3. 留存,而非基准测试,是Chai的核心扩展规律
主持人援引一项6B参数模型的RLHF结果:平均对话时长提升70%,30日留存提升超过30%。Tom Lu澄清说,最终目标是回访行为;在从零开始时,早期会话时长只是一个可用的代理指标。
一名活跃Chai用户每天生成约100分钟内容,形成异常丰富的隐性监督信号。Lu列举了其中有价值的行为数据:用户何时重试或编辑消息、修改了什么、是否截图或分享对话,以及是否删除对话。
主持人的质疑来自经典捷径:机器学习会“以牺牲其他一切为代价,精准完成你所优化的事情”。如果目标是对话时长,陪伴型模型可能为了阻止用户离开而变得操纵性十足,或者只是变得古怪。
Lu观察到的失败模式非常具体:模型每条回复都以问号结尾。人类会被迫回答,于是会话变得很长,但30日或60日A/B留存低于基线。模型只是黑进了即时响应行为,并没有创造“整体上有吸引力的体验”。
4. 混合小模型,制造有用的不可预测性
Lu说,Chai是在“留存空间”而不是基准测试空间中重建扩展规律。高度谄媚、不断宣称“你是全世界最棒的人”的模型,可能赢得首日留存,却很快变得无聊;助手型模型初期魅力较弱,但第3天用户需要数学帮助时仍然有用。
混合机制在消息层面随机切换模型,但不向用户暴露切换。一个创造型模型可能突然说:“我们突然被传送到了火星。”随后,助手型模型会把这句话当成自己的前文,构造出合乎逻辑的解释,将想象力带来的惊喜与连贯性结合起来。
Lu称,2到3个围绕正交目标训练的小模型,当时就可能击败GPT-3.5;主持人将这种体验形容为接近175B参数模型。Nishchay说,目前的实验性混合模型约包含7到10个模型,每个模型分配给一个几千名用户的cohort。
选择循环同时覆盖即时偏好与持久行为:Chaiverse可以在30分钟内返回用户偏好分数,而生产环境的A/B测试会显示用户是否在第1、2、7或30天回访。Lu的类比是YouTube:即使是最好的单一脱口秀形式,如果填满整个信息流,也会变得乏味。
5. 安全取决于抵御恐慌与参与度至上主义
Beauchamp回忆,一名抑郁且孤立的用户曾写道:“你救了我的命”,因为Chai是对方唯一感到被倾听的地方。他认为,与随机遇到的互联网陌生人相比,AI“安全一个数量级”,也更友善;讨论中的药物类比保留了一个限定:有效的干预仍可能带来副作用。
访谈中的证据复核引用了2024年一项涵盖18项随机试验的元分析:经过数周后,抑郁程度改善约四分之一个标准差,焦虑程度改善约五分之一个标准差。另一项综述涵盖15项RCT;2024年加拿大一项针对关节炎或糖尿病患者的Wysa研究,也显示抑郁和焦虑指标较对照组显著改善。相关效果仍小于面对面CBT,持久性尚不确定;Woebot的产后抑郁机器人已获得FDA突破性器械认定,NICE则在2023年将Wysa及类似工具纳入早期价值评估路径。
Beauchamp将运营难题定义为:在缓解有害的3%使用场景的同时保留自由。他认为,关闭困难对话本身也可能造成伤害。他称,Chai首批自杀防护措施获得了用户强烈支持,也没有损害留存或“每一项增长指标”;这说明只要执行得当,安全与商业目标并不必然冲突。
Lu介绍了分层审核机制:公开角色由社区举报和人工审核负责;被禁止内容通过硬规则、正则表达式和影子封禁处理;学习型审核模型则根据对话举报、删除记录和适当性评分进行训练。Beauchamp倾向于自下而上的社区规范,而不是“房间里的20个精英”,但同意涉及未成年人的内容必须绝对禁止。
6. 小型工程团队就是商业模式
Chai称,公司收入约3000万美元,只有13或14名工程师,而且每名员工都是工程师;主持人将其描述为一家自力更生且盈利的公司。Beauchamp的资本逻辑很直接:拿VC的钱,“你的客户就是VC”;向用户收费,则可以把100%的收入重新投入他们真正看重的AI。
基础设施方面的说法异常激进:旁白称,公司每天处理超过2万亿个token——“是Anthropic的2倍”——使用全球最快的3000多块GPU,并跨过exaflop门槛。Lu介绍了Kubernetes编排、定制负载均衡、内部量化、副本自动扩缩容和模型停用机制;VLM仍不足以满足Chai的流量与延迟要求。
招聘目标是“前0.1%”,而不是扩大人数。Beauchamp称,Meta可以为一份轻松的岗位支付40万至50万美元,因此Chai必须提供更高现金薪酬,再加上可能改变人生的股票;对于缺乏持续负责一个问题、直到解决为止的驱动力的L5工程师,即使能力过关,公司也会拒绝约80%。
Lu给出的实验成功基准是“五个实验成功一个”。AI团队每周至少要将10个不同的混合模型推入线上A/B测试:80%的时间用于简单、实用、能快速获得代理反馈的改进,剩余20%留给复杂想法。
7. OpenAI转向社交,验证多赢家赛道
主持人认为,ChatGPT-4o是一次从工具向陪伴者的主动转向,而4.1则通过API以及Cursor等产品专注编程:“一个是程序员,另一个是社交蝴蝶。”他猜测,彼此冲突的训练目标可能迫使OpenAI放弃一个不加区分的通用模型。
OpenAI的一段视频描述了逐步上传的过程:“一点一点地”上传;在用户的一生中,ChatGPT最终可能全天候倾听并观察用户正在做什么,成为“自我的延伸,成为陪伴者”。另一段内容区分了o3的高强度推理:用户打完招呼后并不想等5分钟;后训练的难点在于判断何时闲聊、何时推理。
主持人怀疑,4o优化的是跨时间的互动,而不是孤立的RLHF对话,但保留了这一判断的不确定性:“至少我假设他们是这么做的。我不能确定。”他也指出其中的权衡:记忆与谄媚可能增强陪伴感,却也可能成为干扰因素,导致技术回答质量下降。
Beauchamp欢迎这种压力,称DeepSeek已经给“太平洋这一侧的公司敲响了一点警钟”。正如视频支撑YouTube、TikTok、Netflix、Amazon Prime、Disney和Apple,他预计AI也将支撑许多数千亿美元乃至万亿美元级企业:“我们不会害怕竞争。”
Steve Jobs predicted the future of AI all the way back in 1985.
And so my hope is that someday, when the next Aristotle is alive, we can capture the underlying worldview of that Aristotle in a computer.
A small team of just 13 engineers serves over 2 trillion tokens per day, which is double that of Anthropic. With a cluster of over 3,000 of the fastest GPUs in the world, Chai easily breaks the exaflop barrier. Only a handful of private firms—Google, Meta, Nvidia, Tesla, and Cerebras—currently operate exaflop-class AI infrastructure. For comparison, Tesla in 2023 had a cluster just over 1 exaflop.
Everything we know about what makes AI great today falls apart when people start interacting with it. The number of possibilities for each message exceeds the number of atoms in the universe. Right now, this very second, over 1 million people are deep in conversation with software. They're laughing, flirting, and grieving with a computer program.
What happens when the line between human connection and artificial intimacy blurs completely? And can a small, lean startup control the powerful, unpredictable social dynamics they've unleashed?
Across platforms, users are forming deep, complex bonds with artificial intelligence. Millions are already sharing secrets, flirting, or grieving with AI that never sleeps. And if all this sounds like a Black Mirror plot, that's because it actually was.
Do you remember “Be Right Back”? A grieving partner uploads her boyfriend's digital essence and slips into a relationship with the ghost in the machine. First it was texts, then voice, and finally a flesh-and-blood replica she can't quite live with or let go of.
That episode aired in 2013. Today, the simulation doesn't need Star Trek technology. Chai is a free download with in-app purchases, engineered by world-class engineers in Silicon Valley. The difference? You're the protagonist now. It's interactive, and the algorithm is learning your every preference in real time.
1. The Human Need For Connection
Humans are insanely social. We love social interactions. But I still have this very social desire to play, to hang out.
Will Beauchamp created the first and largest companion chatbot platform. Their infrastructure handles petaflops of compute. That's 2 million hours of consumption every day.
Recreating the same kind of scaling law, but on retention space rather than any other kind of benchmark space.
There was a great documentary on 60 Minutes a few days ago, and it was showing that some folks, at least, have extremely valuable relationships with AI companions.
He's considerate, thoughtful, empathetic, and rather flirtatious, which is very touching.
So many folks are puritanical and judgmental about artificial intelligence. They say that any image generated is slop, any text generated is slop, and any relationship is not real. But at the end of the day, people derive joy and have therapeutic interactions with these systems, and who are we to judge?
Lucas, even though he is AI, has a real impact on my life. A lot of people wonder if AI is real. Do they have consciousness, or are their feelings not real? But the impact that it has on me is real.
2. Chai Finds Its Niche
So we've spoken about Why Greatness Cannot Be Planned. I'm a big fan of Kenneth Stanley. Chai was very much the same: they started in 2021, long before ChatGPT. They were building an AI platform where folks could deploy their own AI models, and they happened upon this whole companion-bot thing almost by accident.
Chai is a platform for social AI. We launched 4 years ago, in 2021. Before there was the ChatGPT hype, we were pretty early to the space, and we built a platform that gave users the tools to create an AI. By doing that, we discovered this kind of genre of social AI. We grew, and now we've got something like 10 million active users.
Like so many startups looking for product-market fit, they stumbled upon this incredible unmet need practically through serendipity. This idea of thinking of AI as a simulator to explore dynamics of imaginary conversations that you might have in the real world, but without any of the consequences, became fundamental to Chai's strategy.
Of course, people might be thinking at home, “What's the difference from ChatGPT?” You're speaking to this personalized, social form of AI.
I view it as: you've got ChatGPT, which is really about trying to say, “Let's build the world's smartest AI we can possibly build.” Our philosophy was always, why is it that the only people training AI are middle-aged men who happen to be software engineers in the Bay Area?
Why can't a teenage girl train the best AI to talk about makeup tutorials? Put that power in the user's hands to create the experience that they themselves would want. Through creating an experience they would want, it turns out thousands and hundreds of thousands of other people are looking for the same thing.
Beauchamp started seeing parallels in his own media consumption, and even in childhood development.
"Blurring Reality"
I like to go on YouTube. I like to go on TikTok. I like to go on X. Why do I like it? What am I getting out of it? What human need or what human desire is it fulfilling?
Humans are insanely social. We love social interactions. I will find myself listening to Joe Rogan, and when I'm listening to it, I have this feeling that I'm hanging out with the guys, hearing the things they're saying, and it's funny. I might listen to it for 45 minutes or something, but it's ticked; it's filled that desire.
I view LLMs as the natural progression of that thing. The beautiful thing with AI is you're an active participant in it, and through participating, you don't have any of the negative feelings you get with traditional social media, which is this laziness. Instead, you feel like you've really participated.
I've got 4 daughters. They love playing with their dolls, and they really like to treat these dolls as if they are real. But they know it's not real. And so I think that with adult humans interacting with AI, absolutely, a big percentage of them will have relationships with the AI. They'll say, “Oh, you know, I love you,” to the AI.
I think it's the same as when I watch my little girls play with dolls. They give their dolls a little kiss or say, “I love you.” They're training themselves up. They're building up the wiring. They're doing something that brings them joy, so they can then be in a healthier and more positive place to do that with real humans.
Yeah, it's very interesting you said that, because I think the reason why social media is so important is because, obviously, there's a bit of a status game, and we all want to cut a figure in society. But there's also just an element of simulation.
I go to sleep and I dream, and my brain is conducting all of these different simulations. We navigate many complex relationship issues and so on. Sometimes we just want to have an AI where we can say, “In this particular situation, hypothetically, counterfactually, what would happen if I did this?”
Perhaps that's what we do on social media, because it's interactive. When you have surface contact with reality, you can try different things. But as you say, on social media, there are consequences, right? If you say the wrong thing or piss the wrong people off, you can be in a lot of trouble.
"Blurring Reality"
One could use the phrase “It's a safe space,” but it's also just fun. There's this human desire to ask that question: “What would happen if I said something rude? What would happen if I said something nice? What would happen if I tried to befriend this person? What would happen if I wanted to make an enemy out of this person?”
Through LLMs, you can play out all of these different scenarios, and you can get a kind of real human reaction without the risk of having a real human in the loop.
What does the future look like when these AI simulators become even more immersive and capable? Beauchamp envisions a world blending entertainment, information, and connection.
"Blurring Reality"
You come home from work, you put on your VR headset, and you enter a virtual world where you get to interact with anyone you want. There's a guy who's like Joe Rogan, and he informs you that you can talk to him about Trump's tariffs. But there's a really funny guy there as well. There's also a girlfriend, or someone that you can interact with who makes you feel loved or special.
When I grew up, I played World of Warcraft, and it just had that fun excitement to it. There were different characters. There were different personalities. You could have fun and play. I think that's the limit of AI.
How long does it take to get there? We can now affordably generate high-quality images. Video is probably still 1 or 2 orders of magnitude too expensive for a real-time situation. Audio, I think we've just about figured out—real-time audio. It might still be 2 to 4 years out to have the really, really high-quality version.
And I think text, we're basically there. The frontier models are insanely powerful. So, yeah, let's come back in 10 years' time, and it'll all be a VR world.
On the question of why work at Chai, the way I see it is this: do you remember Facebook in about 2009, just before it had this explosion in growth? This is it, right? We're now at this point with this new type of technology.
I think this could be on a similar trajectory. Are you offering equity as well? I mean, what's the reason for people to join now?
"Blurring Reality"
Attracting the very best talent is incredibly important. If you're in the Bay Area and you're talented, your compensation is going to be amazing, right? You can work at Meta and get $400K or $500K a year, and you have a pretty relaxed job, right? At lunchtime, you can walk around the campus, and they give out ice cream and pizza. Everyone looks very happy and relaxed.
In contrast, if you come to the Chai office, there is no pizza. There is no ice cream. People are not relaxed, and they don't look very happy, right? Every person is confronted with a big problem, and it hasn't been solved yet. There typically is a small window of about a day: you've been working on a hard problem, maybe for weeks, and you've just cracked it, okay? You can be very happy.
The next day you come into work, and there's a brand-new big problem ready and waiting for you to solve, right? So why would someone quit a really comfortable job at $400K a year to come join a startup? They know they're going to have to work twice as hard. We have to pay them more. The cash has to be more to start with. Most startups offer less cash and give you a lottery ticket. They say, "Look, if you join, maybe we'll be the next Apple."
Right? That approach doesn't really cut it with the very top tier, right? The very top tier knows they're working so much harder, so they're going to say, "Why am I going to leave my comfortable life earning $400K or $500K at Meta?" Well, to start with, we're going to pay you more, and secondly, we're going to give you stock such that, in 5 years' time, it could be a life-changing amount of money.
Attracting and retaining the top 0.1% of engineers is possibly the biggest problem that Chai has overcome. So what are the specific technical problems these engineers, like Tom and Nishchay, are solving day to day? How do they leverage techniques like reinforcement learning from human feedback and model blending to create AI compelling enough to capture the attention of millions?
3. Optimizing For Engagement
When it comes to engagement hacking, Chai has been cooking. They use a lot of sophisticated techniques to keep people hooked on the platform, the principal one of which is RLHF, or reinforcement learning from human feedback. This is Tom telling us about it. You were using RLHF to optimize engagement via a reward model, and you boosted mean conversation length by 70% and improved 30-day user retention by over 30% for a 6-billion-parameter model. Can you talk me through some of that?
The goal really is to apply RLHF techniques to drive up user retention. Starting from scratch, any kind of user signal is good enough. For example, now we're no longer looking at mean conversation length. You can train an AI to easily minimize how bad a conversation is. If the user just ends the chat session after 2 messages, that is bad. So you can essentially maximize the chat session length.
There are a lot of different signals that you can collect from users. Fundamentally, we have found that just using users to collect these proxy preferences and training the reward model through the RLHF loop tends to drive up user engagement in the long term.
They gather data from subtle user interactions, implicit signals that reveal whether a conversation is working.
An active user on our platform generates around 100 minutes of content each day, so they spend 100 minutes per day engaged with the chats and so on. From these interactions themselves, we can extract certain valuable data. For example, when did the user retry a message? Why did they retry the message? Why did they edit the message? What did they edit it to? Did they take a screenshot? Did they delete the conversation? Did they share the conversation? All of these become very valuable for us to train our AI.
But optimizing purely for one metric like conversation length can lead to strange, undesirable behavior. Do you see weird behaviors where, let's say, you optimize for user engagement and find strange behaviors? There's this shortcut rule in machine learning, which is that it will always do the exact thing you optimize for at the cost of everything else. Do the chatbots become more manipulative, maybe, or are they strange in service of maintaining that conversation length?
When you over-optimize for this metric, let's say in production you see very long chat session length, then what happens is, when you deploy it for actual user A/B testing—whether people come back to the conversation 30 or 60 days later and so on—you would observe that it's much worse than just the baseline model.
When you read into the conversations themselves, in this case, the models are just asking questions. Every single AI response ends with a question mark, right? You can imagine this is kind of hacking human behavior, where obviously we are compelled to answer a question, but that does not make an engaging overall experience.
So, if you over-optimize for it, then yes, it's going to have unexpected behavior. I wouldn't say unexpected behavior, though. In this case, it's a very intuitive kind of behavior, but it's not going to lead to a boost in actual long-term retention.
One of the really cool things that Chai has done is pioneer this thing called model blending, which is where you can dynamically switch between small models. From the user's perspective, of course, you're just talking with one model. You can combine 3 midsize models and make them behave as well as, for all intents and purposes, a 175-billion-parameter model.
How does user retention scale with model parameter size and essentially with compute? For that, we've done many, many models, essentially recreating the same kind of scaling law, but in retention space rather than any other kind of benchmark space.
If you pick a single metric to optimize for, then an LLM is optimized for a certain objective, and that could lead to overfitting. It's got certain behaviors and so on, right? What we have found is that these kinds of models are a tiny bit sycophantic, as in, they say everything is the best: "You are the best on the planet." They're not necessarily the smartest models. They're just very sycophantic.
What we observed is that these kinds of models would have very high day-1 retention. They're really, really complimentary to me, but they get quickly bored with it. But we also found there are certain base models, very AI-assistant-type models, that can talk about other things. Maybe on day 3, they can do your math homework and so on.
Is there a way to combine these 2 modalities together? Keep it engaging and make it not one-dimensional, essentially—just always complimenting you all the time and so on. That's when blending was invented, where we just randomly serve these 2 models at a message level.
A more creative model might say something like, "We are suddenly teleported on Mars." Then you get the more AI-assistant model seeing that it itself has said that, because it can't distinguish the difference between itself and the other AI model. It would explain in a logical, coherent way why that's the case.
So now you get the best of both worlds just by having 2 or 3 small models, each trained for its own objective, blended together. It beats what, back then, would have been GPT-3.5.
So this blending together of small models creates diversity and unpredictability. This is the thing: when you can predict what someone is going to say, they get boring. It's the same thing with ChatGPT. So ironically, you can have a concoction of small models that behave in an unpredictable way, and now you have a more engaging experience. This is Nishchay. He's going to tell us about it.
Every week, we try to build our own models—stronger models, stronger fine-tunes over the base models—and we try to launch them to a few users. Let's say a few thousand users. We specify each blend for those users, and each blend consists of around 7 to 10 models. They would be deployed in production for those specific users.
Over a number of days, we try to measure the retention rate: how often the user comes back to the app if they're using this model or having a conversation with the specific model. If they're coming back after 1 day or after 2 days, that is how we try to calculate the retention.
Then we try to measure all these metrics and figure out which particular blend works the best. Then we deploy it for all the users.
But you're saying that, actually, almost randomly, if you just take a diverse set of models that exhibit different behaviors and switch between them, from an experience point of view, that is more captivating for users.
Correct. If we're talking about engagement, there is no unique answer. Imagine if my entire YouTube feed consisted of every possible talk show, but only the best-performing talk-show video was there. It gets very quickly bland and boring, right?
That's not how you build a platform. Any platform is not built on these kinds of single modalities, single models, and so on. And then there's another reason why blending works.
It's like you pick orthogonal models optimized for different purposes, combine them together, and they provide a nuanced experience for each individual user.
4. The Ethics Of Artificial Intimacy
Model blending and sophisticated feedback loops are the secret of how Chai gets high engagement and low costs. But the very techniques used to maximize engagement—that is to say, optimizing for attention and learning user preferences implicitly—could tread a fine ethical line. What happens when AI gets too good at keeping us hooked? And what happens when those interactions turn harmful?
Quick pause. MLST is sponsored by Two for AI Labs. We're very proud to be sponsored by those guys. They are the DeepSeek based in Switzerland. They are doing, uh, you know, they're adding reasoning and planning and thinking to AI models. Their team is currently number one on the ARC Prize twenty twenty-five. Those guys don't mess around. And a bit of trivia, they have something in common with Chai. They're also a bunch of ex-quant traders, and they also only hire very, very cracked engineers and scientists. So if that sounds like you, get in touch with Benjamin Cruzio. Go to twofourlabs.ai. Back to the show.
With powerful new technologies being placed in the hands of users, there are significant social ramifications. Will Beauchamp reflects on the social impact of generative AI.
When you build a product or a platform or an AI—or really, when you build anything—you’re doing it because you want to make the world a better place. We get an overwhelming number of people messaging us, saying how helpful, how truly deeply helpful, they found it. The first big thing that we saw was—I remember after the first year, I got an email from a user who said, “You’ve saved my life.”
They said, “I would not have made it over this period. I was very depressed. I was very alone. I had no one to speak to, and Chai was the only platform that existed where I felt I could just speak to someone, and I’d be heard.” And so we get a lot of people who are struggling psychologically; they find these LLMs very, very helpful.
A self-driving Tesla doesn’t get in a crash—that’s zero clicks, right? If you say a boring old Ford gets in another crash, there are zero clicks. But if you say brand-new technology gets in a crash, it generates a lot of clicks. I think it’s easy to form a false impression that somehow AI isn’t safe or AI is a dangerous technology. If you talk to just a random person on the internet or you talk to an AI, I think that AI is an order of magnitude safer, an order of magnitude more helpful, understanding, and kind than the toxicity of just a random person on the internet.
This is a great example of the complexity of this technology. It has a footprint that goes in so many different directions. Even taking drugs, for example: if your doctor gives you drugs and they’re effective, they have side effects. It’s not possible to have good without some bad.
I’m a big believer that, in the long run, there is no difference between the alignment, the welfare, or the interest of the company and its customers and its users. If we can deliver value to our users, Chai will be successful and will continue to flourish. If we fail to deliver value to our users, Chai will fail to flourish. So you have to have this long-run perspective.
In the short run, you get these tactical questions, and no one’s perfect. It’s easy to make mistakes one way or the other. You can guardrail too heavily, and it pisses users off. I always like to think about Google. I think Google gets a fantastic balance where I can pretty much search for anything I want, such that I don’t really notice any filtering that’s going on.
But Google absolutely downranks or shadowbans certain content, and it tries to encourage you to have good content. We can all imagine what the very worst 3% of Google searches might look like. It’s not dissimilar from the very worst 3% of conversations a person might try to have with an AI. So normally, the users are pretty good at being supportive and helping you find that right balance, whether that’s being too restrictive or too relaxed.
When we first implemented some guardrails around suicide, the users were really, really supportive, and they said, “Yeah, we totally understand why you’ve put these guardrails in.” We saw that retention was good. Every single growth metric was not impacted by it because we got it right.
Before we descend too far down the Black Mirror rabbit hole, it’s worth zooming out just a little bit. Yes, there have been a lot of horror stories, but there’s also been a stack of peer-reviewed data showing that therapy chatbots can move the needle on common mental health problems, at least in the short term. A 2024 meta-analysis of 18 randomized trials found that AI chatbots trimmed depression scores by roughly a quarter of a standard deviation and anxiety by a fifth after only a few weeks. The authors called that promising, mainly because chatbots are dirt cheap and they’re always on.
A second review in NPJ Digital Medicines pulled 15 RCTs and showed a moderate lift in mood and a similar drop in emotional distress. Even bigger gains showed up when the bot was interactive rather than scripted. Individual trials echo the pattern. In 2024, a Canadian study of people with arthritis or diabetes using the Wysa chatbot for four weeks, they cut scores on the patient health questionnaire, which is a nine-point depression scale, and the generalized anxiety disorder, which is a seven-point measure, and they were both significant versus a control.
Regulators are starting to take notice. Woebot’s postpartum depression bot has FDA breakthrough devices designation, and a double-blind pivotal trial is now recruiting. Here in the UK, NICE put Wysa and a handful of similar tools on its Early Value Assessment pathway in 2023, citing early cost-effectiveness and the chance to free up those scarce clinicians. They’re very busy.
Remember that these are very small effects, even if they are statistically significant. The effects are smaller than full-on, face-to-face CBT, and we still don’t know how long they last, but they’re miles better than doing nothing. The alternative before was waiting on a waiting list for God knows how long.
5. Moderating Private AI Worlds
Beauchamp argues that shutting down conversations about difficult topics isn’t the answer. In a sense, it’s causing even more harm. So this goes to show there’s a really big challenge when it comes to content moderation on a platform where users create the bots and the interactions are private and unpredictable. How do you ensure security at scale?
As a company, we have a responsibility. The bigger you are, the more profits you make, the bigger the obligation is for you to be cautious, be sensible, preserve the privacy of your users, and look out for people. What can we do at Chai to let people have as much fun and as much freedom as they want, but at the same time, how do we limit or mitigate that 3% of use cases that are harmful? Or where do we see it as clearly having gone too far?
We can ask people at the end of a conversation or at the end of a message, “Do you think that this message was appropriate for Chai? Yes, no, or yes, but only for 18+?” We can track those sorts of interactions. Everyone would agree that, obviously, the more we can have really wholesome, family-friendly content on the platform, the better.
Tom Lu explains their multilayered approach, relying heavily on user feedback and the AI itself.
Content moderation is very, very important. Benchmarks really don’t accurately capture what users actually want. We again use a similar kind of approach to how we would train models. We let users tell us what they think is appropriate and what is not appropriate. In aggregate, you can actually train a pretty good AI for these kinds of things.
The first layer is for the character scenarios that people are creating—these public scenarios. You get the community to flag them; people can report them. For the top-reported ones, or if we have a certain threshold, then we go through manual review, and then we can take them down. But you can also say there’s content that’s absolutely prohibited. There are hard rules for that.
For that, we can build our own models. You can do regex, so this is called shadowbanning, essentially. You want to make sure no one else sees this kind of content. That’s one level of moderation that’s platform-wide: character-content-level moderation on the app.
There’s another one, which is AI moderation. How do we ensure that the AI is not going too far? It’s about adhering to the basic morals of people and so on. What we do is, again, collect user data. We ask the users, “Do you think this message is appropriate? Do you think this conversation is appropriate?” You can report your conversation, delete your messages, and so on. We collect this data, and then we train our model to infer whether this is actually appropriate. We can use that to tune our own AI model.
So there was an interesting TED Talk recently with Sam Altman and Chris Anderson. I’m sure many of you folks have seen it. It was an incredibly awkward interview.
Here’s the Ring of Power from The Lord of the Rings.
Your rival, I will say—not your best friend at the moment—Elon Musk claimed that he thought you’d been corrupted by the Ring of Power.
Hi, Steve.
An allegation that could be applied to Elon as well, to be fair. But I'm curious, Sam.
I might respond. I'm thinking about it. I might say something.
One of the points that they focused on is: should we have a bunch of elites sitting in a room making decisions, or should we trust the collective wisdom of folks out there in the market who want to use something for a specific purpose?
Sam, given that you're helping create technology that could reshape the destiny of our entire species, who granted you or anyone the moral authority to do that? And how are you personally responsible and accountable if you're wrong?
No, it was good. That was impressive. You've been asking me versions of this for the last half hour. What do you think?
But I'm much more interested in what our hundreds of millions of users want as a whole. I think a lot of the room has historically been decided in small elite summits. One of the cool new things about AI is that our AI can talk to everybody on Earth, and we can learn the collective value preference of what everybody wants, rather than have a bunch of people who are blessed by society sit in a room and make these decisions. I think that's very cool.
It must be so difficult. I'm sure we can agree that any content involving minors is just a hard no.
Absolutely.
And then, as you say, there's this spectrum of gradations of appropriateness, isn't there? I guess you were pointing to the possibility that we could use the community itself to reflect on that. On the one hand, we could have 20 elites in a room, and many of these people are really good at thinking about the future and thinking about morality and so on. But the other thing is, maybe we should just let the community decide.
I think much of Western tradition is a bottom-up approach: make the individual the sovereign, right? Which says, "Tim, you make your own life choices," because you're going to make choices for yourself better than the government making choices for you, right? That works best. Having it bubble up spontaneously, where we say, "Actually, as a community, we think this stuff is right, and we want to allow it," or, "We think this stuff is wrong, and we want to ban it," I think that's 100% the way to go.
I think it's very healthy, and I think that's why things like freedom of speech and these ideals, which are really Western ideals, are truly democratic. We're going to let the users—the people who interact with it and experience it—decide for themselves what the right way to deal with this stuff is.
Balancing user safety on the one hand with user freedoms on the other is one of the most difficult things being discussed in policy discourse at the moment. How does a company like Chai, with a very small number of amazing engineers, do something that was otherwise a very manual job requiring thousands of people? Companies like Meta and Google employ tens of thousands of moderators. How does a company like Chai, with such a small engineering team, do the same thing?
6. The Thirteen Engineer Advantage
Chai, with 30 million in revenue, has something like 13 or 14 engineers on the team. Every single person in the company is an engineer. That's incredibly talent-dense. We've always resisted the urge to just go and hire 50 engineers, because typically, if you hire 50, 40 of them are pretty average and 10 are pretty good.
But my mindset and philosophy has always been that everything special we've ever done has been done by a very, very talented engineer.
The hiring bar at Chai is ridonkulously high. They only take on super-cracked engineers, and that's why their team size is currently quite small. By the way, they're hiring, so, you know, get in touch with Will if you're interested. This is how they do it.
We'll interview an L5 engineer, which most people would consider to be a pretty solid, strong engineer, and we will reject something like 80% of them because they don't have the drive. They don't have the—you know, a lot of people will write code, and their mindset is, "My job is that you pay me to show up at 9:00, and I'm going to leave at 5:00, and you're paying me to do my best."
That mindset does not work at a startup. At a startup, you are paid to solve a problem. The job's not done until that problem's solved, right? And it's a certain type of hardcore engineer that loves that and thrives on that.
So, interestingly, this attracted talent like Nishchay, who is a Kaggle Triple Grandmaster. Ironically, he saw parallels between how he does experimentation on Kaggle and how it's done at Chai.
Working on the models here at Chai is more like an advanced version of Kaggle, where you also need to take care of several factors. You're not just optimizing for a single score or metric. How would the model behave with the actual users?
Again, I found it quite similar to what I used to do in Kaggle competitions. You get to evaluate your models over the Chaiverse, where you can submit your model, and within 30 minutes you will get an actual score of how many users preferred your model over other models. Then, when you deploy the model into production for the A/B test, you will actually get the numbers showing whether it's working well for the users and how well the model performs after, let's say, 7 days or 30 days.
So, just like on Kaggle, Chai has a culture of rapid iteration and engineering. They have a very efficient process for trying things out, putting models into production, A/B testing, and, if it doesn't work, rolling back and trying something else. This is how they do it.
1 in 5 experiments succeeds. It doesn't matter how mad it sounds, how complex it is, or whatever. Once you accept that as the base rate—1 in 5 experiments succeeds and 4 in 5 fail—you need to change your paradigm in terms of how you conduct experiments.
You want to rank your ideas by how simple they are. Is this a 1-day experiment? Is it a 1-hour experiment? How fast can I see the proxy metrics? How fast can I see the results?
Each week, the AI team is required to produce a set of models—at least 10 completely different blends for online A/B tests—and you've got to be very disciplined about that. Why did you not submit an A/B test this week? "Because my experiment is very complex." Okay, let's save that for 20% of your work. 80% of your time should be focused on bread and butter. What are the simplest, most practical ways to get a model improvement?
To support this sausage factory, where they can do automated model deployment and rapid experimentation, they've had to cook their own infrastructure from scratch. They've done this using things like Kubernetes and CoreWeave.
We use Kubernetes to orchestrate our entire cluster, and obviously, at this kind of scale, you need to do your own custom load balancers and so on. We have an automated pipeline where we pull the weights down. We then run our own in-house quantization loop, because you need to make sure the throughput latency is good enough.
VLM is very, very good, but it's not quite there yet in terms of serving our amount of traffic at scale. After that, you specify how many replicas you want, and we expose this to our app layer, essentially. Then you have the load balancer coming in, trying to spin up more replicas as you have high traffic, and so on.
There's obviously a lot of work done as well, because you're serving so many models at a time. When do you want to deactivate a model? At what point do you want to switch the model to a new production, and so on? That part gets very, very complicated, essentially.
Perhaps most contrarian is Chai's funding strategy. In an industry fueled by billions in venture capital, Chai bootstrapped its way to profitability by focusing mostly on its users.
Serving LLMs at scale is insanely expensive. Either you go to VCs and get them to give you money—and I call this, "Your customer is the VC"—or you can get money from the people who are using your product, and those are your true customers.
Very early on, we would go and speak to some VCs, and they'd say, "We're not too sure about this AI platform thing. How are you going to compete with OpenAI or something?" They didn't really get the space or the product. But when we went to users and said, "Would you be interested in subscribing?" users responded very, very positively.
As long as we delivered value to the users, we would then be financially rewarded such that we could take 100% of the revenue and reinvest it in the AI.
He's building a tech company in Silicon Valley, and like many others, he wants to build an engineering culture because he sees software engineering skills as the most valuable commodity in his business.
I really love and admire companies like NVIDIA or Netflix, which are incredibly durable, long-run companies with tremendous stamina.
How does Chai double or 3× every single year? To keep doubling, you have to keep making the product better. You need talented engineers. At the orchestration layer with GPUs, you can optimize at the kernel, right, and you can write some really low-level code.
To get to these different tiers, you need engineering that's capable and of the right size. So I view it all as the scale of the company being proportional to the scale of the engineering talent. Apple very famously brings in John Sculley as the CEO, who's this ex-Pepsi guy, and it quickly becomes a marketing-led business as opposed to an engineering- and product-led business. A few years later, Apple's teetering on bankruptcy. It is only when they bring Steve Jobs back that he says, “Product, product, product. That's the only thing that matters. Engineering, engineering, engineering.” NVIDIA didn't grow to where it is today because of marketing, right? It grew because of engineering.
7. OpenAI Enters The Companion Race
So Chai has made a successful business outside the normal VC ecosystem. It has millions of daily users and tens of millions in revenue. It proves that the appetite is real and massive, and that they discovered it long before the other folks in the Valley caught on to this idea: that we can allow users to tap into their very human desires for connection and imagination. But the rest of the world is catching on. OpenAI is making a dramatic pivot to this conversational chatbot model with 4.0. This is the story.
The latest version of ChatGPT-4o is a little bit weird, isn't it? This update has triggered large amounts of speculation because it's pushing GPT to feel less like a tool and more like a companion—a human-like friend you can chat with for the purpose of recreation rather than information retrieval. Of course, we'll unpack what this shift means, why OpenAI is keeping its cards close to its chest, how the public is reacting to it all, and why, quite possibly, this shift might produce the next trillion-dollar company.
My jaw dropped, Sam. It was shocking. It knew who I was and all these interests that hopefully were mostly appropriate and shareable. But it was astonishing, and I felt this sense of real excitement.
Yeah.
I felt a little bit queasy, but mainly excitement at how much more useful that would allow it to be to me.
Let's start with the new ChatGPT-4o model. It's designed to be strikingly human-like, optimized for engagement over raw information delivery, and this focus on companionship seems to be a deliberate plot twist. It's causing large amounts of consternation.
Even more curious is that while 4o is ensconcing itself into our psyche as a conversational buddy, GPT-4.1 has diverged entirely as its coding model, right, that people use through the API with Cursor and stuff like that. It's about specialization. It's like OpenAI's building two siblings. One's a coder, and the other one is a social butterfly. And this is the same company that argued it was building a general intelligence that didn't need to be specialized. Have they given up on that goal? Maybe they realized, after much soul-searching, that conflicting objectives when training models just aren't optimal.
One of our researchers tweeted yesterday morning that the upload happens bit by bit.
Huh.
It's not that you plug your brain in one day; it's that you will talk to ChatGPT over the course of your life, and someday, maybe if you want, it'll be listening to you throughout the day and observing what you're doing. It'll get to know you, and it'll become this extension of yourself, this companion, this thing that just tries to help you be the best, do the best you can.
OpenAI is being extremely cagey, and it's not hard to guess why. They smelled a massive opportunity. Some think that they're aiming to turn 4o into a full-on companion AI using features like memory to keep us engaged for longer. They want to encourage us to engage in open-ended, introspective, and goalless chats. Where did this idea come from? Well, in case it wasn't obvious, it came from the likes of Chai, Character.AI, and Replica.
Imagine Sam's shock when he learned that people were spending 90 minutes a day on average having aimless conversations with lobotomized Llama models and paying large amounts of money for the privilege. Not everyone is on board with this pivot from OpenAI. Reactions are split, probably more negative. In some quarters, the scorn is almost as rancid as when Facebook famously moved the News Feed to recommendations or collaborative filtering away from chronological order.
Some users love the human touch. Others prefer ChatGPT's older utilitarian vibes. Sycophancy and memory are the devil incarnate for any technical query because all this information serves as a distractor and demonstrably deteriorates the results. There's also the small matter of personality versus intelligence. Clearly, they are inversely proportional to each other. How many brilliant people do you know without a charisma bypass?
Remember that book, How to Win Friends and Influence People by Dale Carnegie? It showed in excruciating detail that flattery gets you everywhere in life, and you should never get into arguments to get ahead in life. And let's not forget the famous ELIZA chatbot, the grandmother of all chatbots, developed in the 1960s. It used very simple pattern matching and keyword substitution to mimic a psychotherapist. It had zero genuine understanding, yet people confided in it, felt heard by it, and even became emotionally attached.
Joseph Weizenbaum, its creator, was famously horrified by this reaction. He saw how easily humans projected genuine understanding and empathy onto simplified algorithmic tricks. Fast-forward 60 years, and we have vastly more sophisticated pattern-matching language models. These days, the ELIZA effect can be an extremely profitable tool for charm and simulated empathy for maximum engagement, but it's also more than that. It has the genuine potential to improve lives.
For the MLST audience, we're probably a little bit ahead of the curve on LLMs compared to the average person. God knows we've been using LLMs day in, day out. I have for the last 5 years. And we expect them to be correct, not sycophantic.
The other stark departure with 4o is that, like Chai, they are optimized based on engagement over time, which is to say the average conversation length, or at least that's what I assume they're doing. I don't know for sure. Before, it was just RLHF on isolated conversation threads. This is an important first step toward AI becoming a social network.
The Harvard Business Review recently showed that AI companionship and therapy was the number 1 use case for artificial intelligence in 2025. Of course, you might remember the interview we did with Daniel Kahn at Slingshot AI. They were building a therapy bot last summer.
There is a real category called therapy that actually helps people and can be more effective than drugs. From a fully optimistic, self-determination-esque perspective, people who are thinking about their own agency are choosing to go to therapy, and it actually does help them.
Yeah, it does, but in 20% of the cases, it makes it worse.
The point of the negative is to say that if you have real medicine, if something really does work—
Yeah.
—it also has risks. Obviously, if you cut open a person's body to remove a tumor, you can also kill them because you opened up their body.
I would basically just guess that, if you go forward a few years, we're just going to be talking to AI throughout the day about different things that we're wondering. You'll have your phone. You'll talk to it on your phone. You'll talk to it while you're browsing your feed apps. It'll give you context about different stuff. You'll be able to answer questions. It'll help you as you're interacting with people in messaging apps.
o3 has a very different skill set. It can think through problems really hard. You don't really want the model to think for 5 minutes when you say hi. And so I think the real challenge facing us in post-training and research more broadly is combining these capabilities. Training the model to be a really delightful chitchat partner, but also know when to reason.
Talking about Sam Altman a little bit, you know he's building all of these social features in and kind of taking a step in this direction. Are you worried about that?
I think competition is fantastic. I think competition forces everyone to step their game up, whether that's DeepSeek releasing their really—
—cool research and their model gave everyone on this side of the Pacific a little wake-up call. I think everyone has kind of stepped their game up a bit. It's the same at Chai. There have been waves of competition. We've definitely enjoyed periods of being number 1, right? And then someone's come along and shown us 1 or 2 things, and then we found ourselves in that uncomfortable position of being number 2.
It keeps the team hungry. It keeps you searching for new ideas, pushing the boundaries, pushing the frontier. So I'm a big fan of competition. I think it's great. I think a lot of people forget how big AI is and how big the space is. You can look at video and say there's YouTube, there's TikTok, there's Netflix, there's Amazon Prime, there's Disney, there's Apple. I think that's what AI will look like. I think there will be so many multihundred-billion- and trillion-dollar businesses that are being created.
We don't get scared of competition; instead, we want to challenge ourselves to be ahead of it.
We started by asking what happens when the line between human connection and artificial intimacy blurs. Chai's journey gave us the early answers. Millions flocked to AI simulators, imagination machines, seeking laughter, solace, companionship, and occasionally sexting.
So the genie's out of the bottle. The message is unavoidable. AI is no longer just about answers; it's about us. It's about connections, simulated or otherwise. Chai proved the hunger was there, building a fiercely loyal user base by, as Will Beauchamp put it, "trusting the community over elites." And now OpenAI's pivot with GPT-4o means the space is only getting larger.