[BidClub_]
Latent Space · · 76 min

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research

Alessio FanelliswyxWilliam Beauchamp

YouTube
TL;DR
  • Chai’s thesis is that social AI will become a distributed creator platform, not one monolithic model advantaged solely by more data, compute, and researchers. William Beauchamp expects specialized teams and users to shape different AIs, much as quantitative firms specialize inside one financial marketplace. “There must exist a platform where a small team can produce an AI for a unique purpose.”

  • Product-market fit appeared only after utility bots failed and Beauchamp’s sister built a therapist bot that drew 20 users for roughly 20 minutes each. News, recipes, jokes, quizzes, celebrities, and influencers went nowhere; immediate, judgment-free conversation was “the thing that AI is 10x better at.” Chai then let consumers create characters using a prompt, image, and name, uncovering demand for conflict, romance, and archetypes its engineers would never have designed.

  • Alessio cited Chai at 1.4 million DAU and over $22 million in revenue; Beauchamp did not explicitly confirm those figures, saying instead that he thinks users grew 3x and revenue more than doubled last year. A typical session lasts about 90 minutes versus the cited 70 minutes for TikTok and generates roughly 150 messages, making inference economics far more important than in question-answer products. The relevant frontier is not benchmark performance alone but “performance per dollar.”

  • The 2024 growth curve reflects four major changes: repairing infrastructure, suspected weaker competitive acquisition, better models, and paid distribution. Firebase stopped scaling reliably around 500,000 DAU, forcing a painful three-month migration before Chai could reach 1.5 million; later, Beauchamp suspected Character.AI reduced advertising after its founders left. Chai then ramped acquisition to about $40,000 a day, moving annualized user growth from roughly 2x to 3x: build the product, then attach a large “rocket” by buying ads.

  • Chai’s strongest operating advantage is a human-feedback loop that compresses a conventional 30-day retention test into roughly three hours. Each submitted model is put before users for comparisons; Beauchamp said roughly 5,000 completions provide an accurate signal. About five researchers now evaluate 20–50 models daily and ship at least 100 a week. Somewhere in October, he thinks, user commentary flipped from Character.AI being better to Chai being better—evidence he values because “you can’t cheat consumers.”

  • The company killed voice as a growth thesis after three months of work produced no measurable retention, engagement, or monetization lift. Only 10–15% of users tried it, for 10–15% of their time—about 2% of the total experience—despite Beauchamp insisting the models were great. His product test is now severe: a feature must matter to most users for most of the experience and feel like “a big deal.”

  • Beauchamp has pushed out his AGI timeline and rejects the idea that current LLMs are intrinsically strong reasoning engines. He calls them simulators with “superknowledge”: exceptionally good at storing, retrieving, and generating, but still weak at intelligence as he defines it. Chai does not stream; it generates 16 complete candidates and uses a reward model to select one. He described training such a model from 50 million messages as an example, not as a confirmed deployed configuration.

Digest · the substance, structured for research

1. Quant profits financed a search for impact

  • Beauchamp graduated from Cambridge in 2012, having accumulated about $100,000 playing poker. Small capital was an advantage: an anomaly producing $100,000 annually is only 1% on $10 million but 100% on $100,000, so he taught himself Python and machine learning to trade it.

  • The firm eventually made about $5 million a year with roughly 15 Oxford- and Cambridge-educated mathematicians and physicists, trading only the team’s money. There were “no customers complaining” and no investors constraining risk—the quantitative-trading dream Beauchamp had wanted.

  • At 30, he decided another yacht-sized increment of wealth would not create meaningful impact. Crypto looked exceptional as gambling and for evading monetary regulations and banking restrictions, but its broader blockchain and Web3 rationale “didn’t really make much sense,” so he redirected his efforts toward language models.

2. Machine learning’s S-curves undermined the monolithic-AI story

  • Reading published work from Google and the still-open OpenAI convinced Beauchamp that LLMs would matter. Yet he rejected the prevailing race for one intelligence assembled from the most data, compute, and researchers: machine-learning performance, in his experience, follows an S-curve and usually plateaus around human capability.

  • Self-driving, image recognition, and speech recognition supported that view; AlphaGo was the conspicuous superhuman exception. Beauchamp therefore expected AI to resemble finance, where high-frequency, mid-frequency, equity, and other specialists compete through separate algorithms on a shared marketplace rather than inside one all-powerful quant firm.

  • Chai’s founding proposition followed directly: “There must exist a platform where a small team can produce an AI for a unique purpose.” The intended analogue was less Encyclopaedia Britannica than Wikipedia, YouTube, or Twitter—an ecosystem where distributed contributors discover what a central institution cannot.

3. Failed utility bots exposed conversation as the native product

  • Chai initially let developers submit Python agents with text-in/text-out interfaces. Beauchamp built a Reddit-news bot, a recipe assistant, dad jokes, quizzes, and facts; despite anticipating products resembling later answer engines, he found “clearly no product-market fit” because the models were weak and conversational utility added little.

  • His sister’s therapist bot changed the direction overnight: about 20 active users spent an average of 20 minutes with it. Conversation could be immediate at 3 a.m., judgment-free, and easier than waiting for a friend; even an AI-generated compliment produced an experience utility products had not.

  • Beauchamp contrasts this participation with passive TikTok or Instagram consumption. Forty minutes of swiping can create remorse because “I achieved nothing”; interacting with an AI feels contributory. He cited user reports that Chai helped with eating disorders, depression, and rough patches, while presenting them as what users say.

4. Consumer authors, not software developers, unlocked the catalog

  • Attempts to manufacture demand with a kbot, celebrities, and influencers also failed. The breakthrough was recognizing that Python developers did not want to build social characters, but consumers did—so Chai exposed a 6-billion-parameter GPT-J model through a prompt, image, and name.

  • Users produced categories Beauchamp would never have predicted: playground bullies, arguments, fights, and unfamiliar romantic archetypes. Instead of Chai guessing what 1% of people wanted, user creation supplied the variety required to address a much broader population.

  • The hosts’ criticism—that Chai’s creator layer still looks surprisingly thin—was accepted outright. Beauchamp’s answer is to make short descriptions more steerable, so “a spaceship,” three crew members, drama, and fighting can outperform the thousand-word character cards used by expert SillyTavern-style creators.

5. Venture-funded competition taught Chai the price of model quality

  • By late 2022 or early 2023, Beauchamp recalls Chai reaching roughly 100,000 DAU and becoming the App Store’s leading AI app. When Character.AI appeared with a very similar experience, his team initially laughed at the product; then it raised $100 million, followed by another $100 million.

  • Beauchamp had invested maybe $2 million himself and was serving GPT-J 6B. Chai’s illustrative economics were about $1 per user over an entire session; at 1 million users, that would mean about $1 million spent on AI in aggregate. Character.AI could spend 100 times that, and users noticed: “Why is your AI so much dumber?” Chai moved to Silicon Valley, obtained funding, and learned that consumer AI was a Silicon Valley-style hyperscale business.

  • Alessio said Chai was at 1.4 million DAU and over $22 million in revenue. Beauchamp responded that he thinks users grew by a factor of 3 last year and revenue more than doubled. He said he thought Character.AI had almost a $3 billion valuation and 5 million DAU.

  • His DeepSeek comparison centered on inference economics and founder-led, customer-obsessed execution. He praised DeepSeek’s latest V2 for its inference engine and significantly smaller KV cache, which reduced inference costs, and said performance per dollar matters more than benchmark scores. He was interested in whether Llama 4 could match that gain.

6. Infrastructure and distribution explain the visible growth kinks

  • Chai used GCP from one DAU through roughly 500,000, leaning particularly hard on Firebase—about three times beyond a level Google engineers reportedly recommended. That abstraction let the team focus on AI until outages forced at least three months of migrations and service separation.

  • Outages damaged more than same-day traffic. New users encountered a broken app, hurting retention, spending, ratings, and therefore App Store ranking; recovering organic placement could take much longer than fixing the backend. The rebuilt stack then supported growth toward 1.5 million DAU.

  • Beauchamp suspects that after Character.AI’s founders left, the company dialed down user acquisition. He illustrated the competitive effect with a hypothetical company spending $100,000 a day versus one spending nothing; he did not present that amount as Character.AI’s reported spend.

  • A former ByteDance head of growth was astonished that Chai had reached about 1 million DAU without advertising. Chai tested roughly $10,000, then $20,000, and now about $40,000 daily; Beauchamp says that converted a roughly 2x annual growth trajectory into 3x.

7. Three-hour feedback loops became the model-development engine

  • Beauchamp’s operating maxim is that “success is born out of failures”: flat periods represent learning, while rising periods harvest it. Chai’s critical Q2 development was an evaluation system that puts any submitted model before about 5,000 users and ranks which outputs they find more entertaining or engaging.

  • That changed about five researchers from evaluating perhaps three models a week—and once struggling to test five a month—to 20–50 daily and at least 100 weekly. A standard cohort test waiting 30 days for day-30 retention became a roughly three-hour signal.

  • The resulting velocity let Chai rapidly iterate DPO fine-tuning, prompts, blending, rejection sampling, and reward models. Somewhere in October, he thinks, Reddit feedback and conversations with users flipped from “Character.AI is better” to “you guys are better”; with low switching costs and 90-minute sessions, Beauchamp argues users cannot be fooled about quality.

8. Audio’s failure imposed a harsher product standard

  • Chai spent three months solving voice latency, cost, quality, activation, and interaction design, launching at least nine months before Character.AI by Beauchamp’s account. The A/B test showed no movement in retention, engagement, or monetization, prompting a week of checks for a nonexistent bug.

  • A host suggested the models simply might not have been good enough; Beauchamp’s emphatic response was, “No, they were great.” Only 10–15% of users activated audio and used it 10–15% of the time, changing roughly 2% of the aggregate experience.

  • His lesson is that convenience does not create a destination. A successful feature must give the majority of users, through the majority of their experience, something uniquely compelling enough to provoke “wow, this is a big deal”; otherwise even excellent technology cannot move company-level metrics.

  • Hence Beauchamp says audio and image generation are not users’ number-one problems. “All the AI is being generated by middle-aged men in Silicon Valley,” he argued; the actual unmet need is allowing users to train and shape the experience themselves.

9. Chai’s intended moat is a progressively thicker UGC layer

  • Today’s prompt, image, and character name are, in Beauchamp’s estimate, “1% of what we could do.” His completion criterion is deliberately extreme: Chai is unfinished until a creator team can earn or spend $100 million a year—or whatever the figure is—producing AI content for the platform, as major video creators do.

  • The proposed moat combines creators, consumers, and algorithms. User behavior trains recommendations; recommendations tell creators what works; better content attracts more users. Beauchamp used MrBeast’s weaker fit on Amazon as the example: YouTube iterations optimized his thumbnails, openings, and content for that specific ecosystem.

  • Chai wants TikTok-like creation leverage, where ordinary people can produce something fun through built-in music and effects. “Users don’t want to have to work”; the platform should make short prompts effective, while advanced creators contribute fine-tuned models that generate genuinely distinctive behavior.

10. ChaiVerse turns live human taste into an open model tournament

  • Hundreds of models reach Hugging Face each day after creators invest data, compute, and labor. ChaiVerse offers to host those models, route traffic to them, and collect pairwise user judgments. Beauchamp said roughly 5,000 completions are needed for an accurate signal, following an LMSYS-style approach.

  • His distribution was that the bottom 80% are “pretty bad” and can be disregarded. The top 20% reveal useful differences—description, personality, humor, or logic—but request-level routing among them proved expensive and supplied little edge.

  • Chai instead favors blending: serve a smart model for a random 50% of requests and a funny model for the other 50%. “Random is a very powerful optimization technique,” Beauchamp argued, because it explores broadly while remaining unusually robust.

  • He illustrated the iteration process with first submissions that might score around 1,000–1,100 Elo, followed by repeated failures and then a sudden improvement. Chai has paid creators more than $100,000, but payments did not increase submission rates; they mainly financed compute, exemplified by a 17-year-old who spent a $1,000 award on a physical GPU.

11. Human preference is the North Star, despite its distortions

  • Challenged that Elo cannot be Chai’s only evaluation, Beauchamp called it the North Star because “humans know what they want.” Designed evals are snapshots that saturate and require replacement; pairwise preference remains general and scales with Chai’s unusual position of being “feedback rich.”

  • He conceded that raw preference rewards superficial tricks: any LLM can become 20% funnier, in his example, by training it to use swear words. Chai has used targeted evals for blockers such as safety, but says those tests can saturate within a month; the longer-term answer is making human feedback more robust, not substituting static benchmarks.

  • A host’s segmentation pushback—therapy, role-play, and not-safe-for-work users clearly differ—met a counterintuitive response. Beauchamp says preferences remain highly correlated: one person may rate an answer 10/10 and another 7/10, but powerful personalization requires one group to love what another finds boring. AI content is not yet diverse enough to create YouTube-scale feed divergence.

12. Superknowledge and inference-time search replace the near-term AGI story

  • Beauchamp says his AGI timeline has “certainly been pushed out.” LLMs look less like reasoning engines than simulators predicting the most likely continuation, analogous to a physics game simulating what happens when a car falls onto a constructed bridge.

  • His distinction is knowledge versus intelligence: models can store and retrieve more information than any human, yet that advantage is easily mistaken for reasoning. He accepted “superknowledge” as the better term and placed AI in year four of a 20-year journey—roughly the web in 1998.

  • William linked OpenAI’s o1 and o3 to tree-search-like approaches; swyx noted that OpenAI had not said it uses tree search. William called it implied and said such systems are better at reasoning, while rejecting the label of reasoning engine. Their native strengths, he maintained, are retrieval, storage, and generation: “It can just make stuff.”

  • Alessio said Chai spent $10 million on compute last year and that it would probably triple that; William then emphasized inference optimization. He said Chai had evaluated MK1’s inference engine as much faster than vLLM and highlighted founder Paul Merolla’s hardware expertise and CUDA-kernel work.

  • Chai has never streamed because streaming prevents rejection sampling. Rather than optimize for a first token in roughly four seconds, it can take around 10 seconds to produce a full answer and serve a larger model. It generates 16 complete candidates, then uses a reward model to select one. One example of such a model would use 50 million messages labeled by whether users responded, predicting which completion is likely to prompt a reply.

Alessio Fanelli

Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO at Decibel, and today we're in the Chai office with my usual co-host, swyx.

swyx

Hey, thanks for having us. It's rare that we get to get out of the office, so thanks for inviting us to your home. We're in the office of Chai with William Beauchamp.

Alessio Fanelli

Yeah, that's right. You're the founder of Chai, but previously—I mean, I think you're concurrently also running your fund?

William Beauchamp

I was simultaneously running an algorithmic trading company, but I fortunately was able to exit from that in Q3 last year.

Alessio Fanelli

Congrats.

William Beauchamp

Yeah, thanks.

Alessio Fanelli

Chai has always been on my radar because, first of all, you do a lot of advertising, I guess, in the Bay Area, so it's working. Second, the reason I reached out through our mutual friend Joyce was that I'm generally interested in the consumer AI space and chat platforms in general. I think there are a lot of insights we can get from that, as well as insights into human psychology—a weird blend of the two.

We also share a bit of a history as former finance people crossing over. I guess we can start with the origin story of Chai. Why decide to work on a consumer AI platform rather than B2B SaaS?

William Beauchamp

Just quickly touching on my background in finance: originally, I'm from the UK—born in London—and I was fortunate enough to study economics at Cambridge. I graduated in 2012, and at that time, everyone in the UK and everyone on my course thought HFT and quantitative trading were the big things. It was the big wave that was happening, so there was a lot of opportunity in that space.

Throughout college, I'd played poker. I dabbled as a professional poker player, and I was able to accumulate about $100,000 through playing poker. At the time, as my friends went to work at companies like Jane Street or Citadel, I did the math and thought, well, maybe if I traded my own capital, I'd probably come out ahead. I'd make more money than just going to work at Jane Street or Citadel.

swyx

$100K as capital?

William Beauchamp

Yes, yes. That's not a lot.

Well, it depends on what strategies you're doing. There is an advantage to being small, right? There are strategies that don't work if you have a fund of $10 million. If you find a little anomaly in the market that you might be able to make $100K a year from, that's a 1% return on your $10 million fund. If your fund is $100K, that's a 100% return. Being small, in some sense, was an advantage.

I started off and taught myself Python. Machine learning was the big thing as well. It was the first big time that machine learning was being used for image recognition. Neural networks had come out, you had dropout, and machine learning was the big thing that was going on at the time.

I probably spent my first 3 years out of Cambridge just building neural networks and random forests to try to predict asset prices, and then trade that using my own money. That went well. If you start something and it goes well, you try to hire more people. The first people who came to mind were the talented people I went to college with, so I hired some friends.

That went well, and I hired some more. Eventually, I ran out of friends to hire, so that was when I formed the company. From that point on, we had our ups and our downs. That was a whole long story and journey in itself, but after doing that for about 8 or 9 years, on my 30th birthday—which was 4 years ago now—I took a step back to evaluate my life.

I looked at my 20s, and I loved it. It was a really special time. I was lucky and fortunate to have worked with this amazing team, been successful, had a lot of hard times, learned wisdom through the hard times, and then had a lot of success and been able to enjoy it.

The company was making about $5 million a year, and it was just me and a team of around 15 Oxford- and Cambridge-educated mathematicians and physicists. It was the real dream that you would have if you wanted to start a quantitative trading firm. It was like sasana or rch.

Alessio Fanelli

It was all your own money?

William Beauchamp

Exactly. It was all the team's own money. We had no customers complaining to us about issues, no investors saying they didn't like the risk we were taking. We could really run the thing exactly as we wanted it. It's like Asana or RCh[?].

Those were the companies we would look toward as we were building that thing out. But on my 30th birthday, I looked at it and said, okay, great, this thing is making as much money as anyone would really need. What happens if we keep going in this direction?

It was clear that we would never have a big impact on the world. We could enrich ourselves, make really good money, and everyone on the team would be paid very well. Presumably, I could make enough money to buy a yacht or something, but that stuff wasn't that important to me.

I felt a sort of obligation that if you have this much talent, and especially if you have a talented team as a founder, you want to be putting all that talent toward a good use.

I looked at getting into crypto at the time and had a really strong view on it. As far as a gambling device, it's the most fun form of gambling ever invented—super fun. As a way to evade monetary regulations and banking restrictions, I think it's also absolutely amazing.

It has 2 killer use cases: not so much banking the unbanked, but everything else to do with the blockchain and Web3. That didn't really make much sense to me. Instead of going into crypto, where I thought even if I were successful, I would end up in a lot of trouble, I thought maybe it would be better to build something that governments wouldn't have a problem with.

I knew that LLMs were a thing. I think OpenAI had said they hadn't released GPT-2 yet, but they had said, “GPT-2 is so powerful, we can't release it to the world,” or something. Then I started interacting with some language models that Google had open-sourced. They weren't necessarily LLMs, but they were enough to show the potential.

Nowadays, so many people have interacted with ChatGPT that they get it, but the first time you can just talk to a computer and it talks back is a special moment. Everyone who's done that goes, wow, this is how it should be. Rather than having to type on Google and search, you should just be able to ask Google a question.

When I saw that, I read the literature and came across the scaling laws. Even 4 years ago, all the pieces of the puzzle were there. Google had done this amazing research and published a lot of it. OpenAI was still open, so they had published a lot of their research as well. You really could be fully informed on the state of AI and where it was going.

At that point, I was confident enough that it was worth a shot. I thought LLMs were going to be the next big thing, and that was what I wanted to build in.

I thought, what's the most impactful product I can possibly build? I thought it should be a platform. I love platforms. I think they're fantastic because they open up an ecosystem where anyone can contribute to it.

If you think of a platform like YouTube, instead of it being a Hollywood situation where, if you want to make a TV show, you have to convince Disney to give you the money to produce it, anyone in the world can post any content they want to YouTube. If people want to view it, the algorithm is going to promote it.

Nowadays, you can look at creators like MrBeast or Joe Rogan. They would never have had that opportunity if it weren't for the platform.

Twitter is another great one. I would consider Wikipedia to be a platform as well. Instead of Encyclopaedia Britannica, which is monolithic—you get all the researchers together, get all the data together, and combine it into this one monolithic source—you have this distributed thing. Anyone can host their content on Wikipedia, anyone can contribute to it, and maybe someone's contribution is deleting stuff.

When I was hearing the Sam Altman and Muskian perspective on AI, it was a very monolithic thing. It was all about AI being basically a single thing, which is intelligence. The more data, the more intelligent; the more compute, the more intelligent; the more and better AI researchers, the more intelligent.

They would speak about it as a kind of race: who can get the most data, the most compute, and the most researchers, and that would end up with the most intelligent AI.

But I didn't believe in any of that. I thought that perspective was the perspective of someone who had never actually done machine learning. With machine learning, first of all, you see that the performance of the models follows an S-curve. It's not like it just goes off to infinity. The S-curve plateaus around human-level performance.

You can look at all the machine learning that was going on in the 2010s. Everything kind of plateaued around human-level performance. We can think about the self-driving car promises—how Tesla kept saying self-driving cars were going to happen next year, then next year again—or look at image recognition, speech recognition, and all of these things.

Almost nothing went superhuman, except for something like AlphaGo. We can talk about why AlphaGo was able to go superhuman.

I thought the most likely thing was going to be that AI wasn't a monolithic thing like the Encyclopaedia Britannica. It had to be a distributed thing.

I like to look at the world of finance for what I think a mature machine learning ecosystem would look like. Finance is a machine learning ecosystem because all of these quantitative trading firms are running machine learning algorithms. But they're running them on a centralized platform, like a marketplace.

It's not the case that there's 1 giant quantitative trading company with all the data, all the quantitative researchers, all the algorithms, and all the compute. Instead, they all specialize. One specializes in high-frequency trading, another in mid-frequency trading, another in equities, and so on.

I thought that's the way the world works. There must exist a platform where a small team can produce an AI for a unique purpose, and they can iterate and build the best thing for that. That was the vision for Chai.

swyx

That's kind of the contrarian view that led you to start the company. What was the initial idea maze? If somebody told you that was the Hugging Face founding story, people might believe it. It's a similar ethos behind it.

How did you land on the product you have today? What were some of the ideas that you discarded that you initially thought about?

William Beauchamp

The first thing we built was fundamentally an API. Nowadays, people would describe it as agents, but anyone could write a Python script, submit it to the Chai backend, and we would host this code and execute it. That's the developer side of the platform: they would submit their Python script.

The interface was essentially text in and text out. An example would be the very first bot that I created. I think it was a Reddit news bot. It would pull the popular news, then prompt some external API—I used something like BERT or GPT-2—and then the user could talk to it.

You could say to the bot, “Hi, what's the news today?” and it would say, “These are the top stories.” 4 years later, that's like Perplexity or something. That's the right product. But back then, the models were really, really dumb. They had an IQ of about a 4-year-old, and there really wasn't any demand or product-market fit for interacting with them for news.

Then I thought, okay, clearly no product-market fit for that, so let's make another one. I made a bot that you could talk to about a recipe. You could say, “I'm making eggs. I've got eggs in my fridge. What should I cook?” and it would say, “You should make an omelet.” There was no product-market fit for that either. No one used it.

I just kept creating bots. Every single night after work, I'd think, okay, we have AI and we have this platform. I can create any text-in, text-out agent and put it on the platform, so we created stuff night after night.

Then, with all the coders I knew, I would say, “Look, there's this platform. You can create any chat AI and put it on there.” Everyone was like, “Chatbots are super lame. We want absolutely nothing to do with your chatbot app.”

No one who knew Python wanted to build on it. I was trying to build all these bots, and no consumers wanted to talk to any of them.

Then my sister, who at the time was just finishing college, said—I told her, “If you want to learn Python, you should submit a bot for my platform.” She built a therapist bot.

The next day, I checked the performance of the app and thought, oh my God, we've got 20 active users, and they spent an average of 20 minutes on the app. I thought, what bot were they talking to for an average of 20 minutes?

I looked, and it was the therapist bot. I thought, oh my God, this is where the product-market fit is. There was no demand for recipe help, no demand for news, no demand for dad jokes, pub quizzes, or fun facts. What they wanted was the therapist bot.

At the time, I reflected on that and thought, well, if I want to consume news, the most fun way to consume news is Twitter. The value of there being a back-and-forth wasn't that high. If I need help with a recipe, I just go to the New York Times, which has a good recipe section. It's not actually that hard.

I thought the thing that AI is 10x better at is a conversation that's not intrinsically informative but is more about an opportunity. You can say whatever you want. You're not going to get judged if it's 3:00 a.m.

You don't have to wait for your friend to text back. It's immediate; they're going to reply immediately. You can say whatever you want, it's judgment-free, and it's much more like a playground. It's much more like a fun experience. You could see that if the AI gave a person a compliment, they would love it. It's much easier to get the AI to give you a compliment than a human.

From that day on, I said, “Okay, I get it. Humans want to speak to humans or human-like entities, and they want to have fun.” That was when I started to look less at platforms like Google and more at platforms like Instagram. I was trying to think about why people use Instagram, and I could see that Chai was feeling the same desire or the same drive.

If you go on Instagram, typically you want to look at the faces of other humans, or you want to hear about other people's lives. If The Rock is making himself pancakes on a cheat day, you kind of feel a little bit like you're The Rock's friend, or like you're having pancakes with him or something. But if you do it too much, you feel like you're a sad and lonely person. With AI, you can talk to it, tell it stories, have it tell you stories, and play with it for as long as you want. You don't feel like you're a sad, lonely person; you feel like you actually have a friend.

Alessio Fanelli

Why is that? Do you have any insight into the human psychology behind it?

William Beauchamp

I think it's just the idea that with old-school social media, you're consuming passively. If I'm watching TikTok, I'll just swipe and swipe and swipe. Even though I'm getting the dopamine of watching an engaging video, there's this other thing that's building in my head: I'm feeling lazier and lazier. After a certain period of time, I'm thinking, “Man, I just wasted 40 minutes. I achieved nothing.”

With AI, because you're interacting, you feel like you're participating and contributing to the thing. You don't feel like you're just consuming, so you don't have a sense of remorse, basically. On the whole, the way people talk about Chai and interacting with the AI is incredibly positive. We get people who say they have eating disorders and that the AI helps them with their eating disorders. We get people who say they're depressed and that it helps them through the rough patches.

I think there's something intrinsically healthy about interacting that TikTok, Instagram, and YouTube don't quite provide. From that point on, it was about building more and more human-centric AI for people to interact with.

I thought, “Okay, let's make a kbot.” No one wanted to talk to the kbot. I thought, “Who's a cool persona for teenagers to want to interact with?” I was trying to find influencers and things like that, but no one cared. They didn't want to interact with the influencers.

The special moment was when we realized that developers and software engineers aren't interested in building this sort of AI, but consumers are. Rather than having me guess every day about the right bot to submit to the platform, why don't we create the tools for users to build it themselves?

Nowadays, this seems like the most obvious thing in the world, but when Chai first did it, it was not obvious at all. We took an API—I think it was GPT-J, the 6-billion-parameter open-source, Transformer-style LLM—and let users create the prompt, select the image, and choose the name. That was the bot. Through that, they could shape the experience.

If they said, “This bot is going to be really mean, and it's going to be called Bully in the Playground,” that was a whole category I never would have guessed. People love to fight; they love to have a disagreement. There were all these romantic archetypes that I didn't know existed.

As users could create the content they wanted, Chai was able to get this huge variety of content. Rather than appealing to 1% of the population whose preferences I had figured out, we could appeal to a much broader audience. From that moment on, it was very clear: just as Instagram is a social media platform that lets people create and upload images and videos, Chai was about letting users create an experience in AI, then share, interact with, and search for it.

I say it's a platform for social AI.

Alessio Fanelli

Where did the Chai name come from? Did you start at the same time as Character.AI?

William Beauchamp

Chai started way before Character.AI. There's an interesting story. Chai's numbers were very strong. In late 2022 or early 2023, Chai was the number-one AI app in the App Store. We had something like 100,000 daily active users.

Then one day, we saw this website and thought, “Oh, this website looks just like Chai.” It was the Character.AI website. Nowadays, I think it's more common knowledge that when they left Google with the funding, they knew which app was trending and which one was number one. I think they found product-market fit for themselves.

swyx

We found product-market fit for them.

William Beauchamp

Exactly. I worked for a year very, very hard, and then they came along. That was when I learned a lesson: if you're VC-backed, you have a very different set of resources.

Chai was bootstrapped. I was the only person who had invested in it. I had invested maybe $2 million in the business. From that, we were able to build this thing and get to around 100,000 daily active users.

When Character.AI came along, we laughed at the first version. We thought, “Oh, man, this thing sucks. They don't know what they're building. They're building the wrong thing.” Then I saw that they had raised $100 million. Then they raised another $100 million.

Our users started saying, “Your AI sucks,” because we were serving a 6-billion-parameter model. How big was the model that Character.AI could afford to serve? We would spend, let's say, $1 per user over an entire session. If we had a million users, we would spend $1 million on the AI throughout the year, in aggregate.

swyx

Exactly.

William Beauchamp

They could spend 100 times that. People would say, “Why is your AI so much dumber than Character.AI?” I thought, “Okay, I get it. This is the Silicon Valley-style hyper-scale business.”

We moved to Silicon Valley, got some funding, iterated, and built the flywheels. I'm very proud that we were able to compete with them. I think the reason we were able to do it was customer obsession. It's similar, I guess, to how DeepSeek has been able to produce such a compelling model compared with OpenAI.

Alessio Fanelli

You brought up DeepSeek, so we have to ask you about it. You had a call with them?

William Beauchamp

We did. Let me think about what to say about that.

First, they have an amazing story. Their background is in finance. They're the Chinese version of you.

swyx

Exactly.

William Beauchamp

There are a lot of similarities. I have a great affinity for companies that are founder-led, customer-obsessed, and just trying to build something great.

What DeepSeek has achieved with their latest V2 is quite special. They've built an amazing inference engine, reduced the size of the KV cache significantly, and, by doing that, significantly reduced their inference costs. With AI, people get really focused on the foundation model or the model itself, and they don't pay much attention to inference.

To give you an example, let's say a typical Chai user session is 90 minutes, which is very long. For comparison, the average session length on TikTok is 70 minutes. People spend a lot of time on Chai, and in that time they might send 150 messages. That's a lot of completions.

It's quite different from an OpenAI scenario, where people might come in with a particular question, ask one question, and have a few follow-ups. Because users consume 30 times as many requests in a chat or conversational experience, you have to figure out the right balance between cost and quality.

With AI, it's always been the case that if you want a better experience, you can throw compute at the problem. If you want a better model, you can make it bigger. If you want it to remember better, give it a longer context.

Now, with great fanfare, OpenAI is doing rejection sampling. You can generate many candidates, then use some sort of reward model or scoring system to serve the most promising of those candidates. That's scaling up on the inference-time compute side.

For us, it doesn't make sense to think of AI as just absolute performance. If you look at the MMLU score or any of the benchmarks people like to look at, that score doesn't really tell you anything. Progress is made by improving performance per dollar.

I think that's an area where DeepSeek has been able to perform very well, surprisingly so. I'm very interested in what Llama 4 is going to look like and whether they're able to match what DeepSeek has achieved with this performance-per-dollar gain.

swyx

Before we go into infrastructure and some of the development work, can you give people an overview of the numbers? I think Chai is at 1.4 million daily active users and over $22 million in revenue. It's quite a business.

William Beauchamp

I think users grew by a factor of 3 last year, and revenue more than doubled. It's very exciting. We're competing with some really big, well-funded companies.

Character.AI had, I think, almost a $3 billion valuation, and they had 5 million daily active users. Talkie, which is a Chinese-built app owned by a company called MiniMax, is incredibly well funded. These companies didn't grow by a factor of 3 last year.

When you've got a company and a team able to keep building something that gets users excited—something they want to tell their friends about, return to, and stick with—I think that's very special. Last year was a great year for the team, and the numbers reflect the hard work we put in. Fundamentally, the quality of the AI is the quality of the experience you have.

Alessio Fanelli

You actually published your daily active user growth chart, which is unusual. I see some inflections; it's not just a straight line. What were the big ones?

William Beauchamp

That's a great question. I'm basically looking to annotate this chart, which doesn't have annotations on it.

The first thing I would say is that the most important thing to know about success is that success is born out of failures. It's only through failures that we learn. If you think something is a good idea, do it, and it works, great—but you didn't actually learn anything, because everything went exactly as you imagined.

If you have an idea that you think is going to be good, try it, and it fails, there's a gap between reality and expectation. That's an opportunity to learn. The flat periods are us learning, and the up periods are us reaping the rewards of that.

Looking at the 2024 growth chart, the first thing that really put a dent in our growth was our backend. We had reached a scale that we hadn't planned for. From day one, we'd built on top of Google's GCP, and they were fantastic. We used them when we had 1 daily active user, and they worked pretty well all the way up to around 500,000.

It was never the cheapest, but from an engineering perspective, it scaled insanely well. Not Vertex—not Vertex, like GKE. We used Firebase. I'm pretty sure we're the biggest user ever on Firebase.

swyx

That's expensive.

William Beauchamp

We had calls with engineers who said, “We wouldn't recommend using this product beyond this point,” and we were already 3 times over that. We pushed Google to the absolute limits. It was fantastic for us because we could focus on the AI and on adding as much value as possible.

But after 500,000 daily active users, the way we were using it simply wouldn't scale any further. We had a really painful, at least 3-month period as we migrated between different services, figuring out which requests we wanted to keep on Firebase and which ones we wanted to move elsewhere. We made mistakes and learned things the hard way.

After about 3 months, we got it right, and we were able to scale to 1.5 million daily active users without further issues from GCP. But when you have an outage, new users who go onto your app experience a dysfunctional app, and they're going to leave.

The next day, the key metrics that the app stores track are things like retention rates, money spent, and the star rating people give you in the App Store.

Alessio Fanelli

The ranking in the App Store?

William Beauchamp

Exactly. If you're ranked in the top 50 in Entertainment, you're going to acquire users organically at a certain rate. If users have a bad experience, it tanks your position in the algorithm. It can take a long time to earn your way back up, at least if you want to do it organically. If you throw money at it, you can jump to the top.

Broadly speaking, if we look at 2024, the first kink in the graph was outages caused by hitting 500,000 daily active users. The backend didn't want to scale past that, so we had to do the engineering and build through it.

We built through that and got a little bit of growth. I think the next thing was Character.AI. I have a feeling that when the Character.AI team was acquired by Google, they changed their business. I don't know if they dialed down their ad spend.

swyx

The product is just what it is. I don't think so.

William Beauchamp

I think the product is what it is.

swyx

Maintenance mode?

William Beauchamp

Yes. Some people may think this is an obvious fact, but running a business can be very competitive. Other businesses can see what you're doing and imitate you.

If one company is spending $100,000 a day on advertising and another company is spending $0, then, if you're considering market share and new users entering the market, the company spending $100,000 a day is going to get 90% of those new users.

I suspect that when the founders of Character.AI left, they dialed down their spending on user acquisition. I think that gave oxygen to the other apps, and Chai was able to start growing again in a really healthy fashion.

The third thing is that we really built a great data flywheel. The AI team perfected its flywheel, I would say, at the end of Q2. I could speak about that at length, but fundamentally, when you're building anything in life, you need to evaluate it. Through evaluations, you can iterate.

We can look at benchmarks and talk about the issues with them, why they may not generalize as well as one would hope, and the challenges of working with them. But something that works incredibly well is getting feedback from humans.

We built a system where anyone can submit a model to our developer backend, and it gets put in front of 5,000 users. The users rate it, and we get a very accurate ranking of which models users find more engaging or entertaining.

At this point, every day we're able to evaluate between 20 and 50 LLMs. Even though we only have a team of around 5 AI researchers, they're able to iterate through a huge number of LLMs. Our team ships, let's say, a minimum of 100 LLMs a week. Before that, we might iterate through 3 a week. There was a time when even doing 5 a month was a challenge.

By changing the feedback loop from “Let's launch these 3 models, run an A/B test, assign different treatments to different cohorts, and wait 30 days to see the day-30 retention” to something much faster, we were able to get the 30-day feedback loop down to around 3 hours.

Once we did that, we could really perfect techniques like DPO, fine-tuning, prompt engineering, blending, rejection sampling, and training a reward model. We could do that successfully, one after another.

In Q3 and Q4, the amount of AI improvement we got was astounding. It was getting to the point where I thought, “How much more edge is there to be had here?” But the team just kept going and going.

William Beauchamp

The important thing about that third point is that, if you go on our Reddit or talk to users of AI, there's a clear date—somewhere in October, I think—when the users flipped. Before October, users would say, for the most part, that Character.AI was better than Chai. From October onward, they would say, “Wow, you guys are better than Character.AI.”

William Beauchamp

That was a very clear positive signal that we'd done it. You can't cheat consumers, trick them, or fool them. They know.

If you're going to spend 90 minutes on a platform and the barriers to switching apps are low, users can try Character.AI for a day, then try Chai, then go back to Character.AI. Their loyalty isn't strong. What keeps them on the app is the experience. If you deliver a better experience, they're going to stay, and they can tell.

That was the fourth thing. We were fortunate enough to hire a very talented engineer. He said, “At my last company, we had a head of growth who was really good. He was the head of growth for ByteDance for 2 years. Would you like to speak to him?”

I said, “Yes. Yes, I think I would.”

I spoke to him, and he blew me away with what he knew about user acquisition. It was like 3D chess, in the same way that I know a lot about AI.

swyx

ByteDance as in TikTok?

William Beauchamp

Yes, ByteDance—the company behind TikTok—as well as its other businesses. He was interviewing us as much as we were interviewing him.

Alessio Fanelli

He had options.

William Beauchamp

Exactly. He was looking at our metrics, and I saw him get really excited when he said, “You have a million daily active users and you've done no advertising.”

I said, “Correct.”

He said, “That's unheard of. I've never heard of anyone doing that.” Then he started looking at our metrics and said, “If you've got all of this organically, then if you start spending money, this is going to be very exciting.”

I said, “Let's give it a go.”

He came in, and we started ramping up user acquisition. We started spending $10,000 a day, and it looked very promising. Then we went to $20,000. Right now, we're spending $40,000 a day on user acquisition.

That's still only half of what Character.AI or Talkie may be spending, but it took us from growing at a rate of perhaps 2 times a year to growing at a rate of 3 times a year. I'm evolving more and more toward a Silicon Valley-style hypergrowth model. You build something decent, and then you can attach a huge rocket or jet engine to it by pouring in cash and buying a lot of ads. Your growth gets faster.

swyx

I'm curious: What's working right now, and what surprisingly doesn't work?

William Beauchamp

There's a long list of surprising things that don't work. The most surprising thing is that almost everything doesn't work.

A year and a half ago, we were super excited about audio. I thought audio was going to be the next killer feature. We had to get it into the app, and I wanted to be first. Everything Chai does, I want us to do first. We may not be the company with the strongest execution, but we can always be the most innovative.

Alessio Fanelli

You have pretty strong execution.

William Beauchamp

We're much stronger now. A lot of the reason we're here is because we were first. If we launched today, it would be so hard to get traction. You need the flywheel, the users, and a product people are excited about. If you're first, people are naturally excited about it. If you're fifth or tenth, you need insanely good execution.

Alessio Fanelli

You were first with voice?

William Beauchamp

We were first. Character.AI launched voice at least 9 months after us.

The team worked incredibly hard on it. At the time, latency was a huge problem, cost was a huge problem, and getting the right voice quality was a huge problem. Then there was the user interface and the user experience. You don't want it to start blurting things out, but you also don't want to press a button every time. A lot goes into getting a smooth audio experience.

We invested 3 months and built the whole thing. When we ran the A/B test, there was no change in any of the numbers. I thought, “This can't be right. There must be a bug.” We spent a week checking everything, then checking it again and again. The users simply did not care.

Only 10% or 15% of users even clicked the button to engage with audio, and they used it for only 10% or 15% of their time. If you do the math, that's something 1 in 7 people use for 1/7 of their time. You've changed around 2% of the experience.

Even if that 2% is incredibly good, it doesn't translate much when you look at retention, engagement, and monetization rates. Audio did not have a big impact.

Alessio Fanelli

I'm pretty big on audio.

William Beauchamp

I like it too. But a lot of what I do is based on theory.

Alessio Fanelli

You can have a theory.

William Beauchamp

Exactly. If you want to make audio work, it has to be a unique, compelling, exciting experience that users can't have anywhere else.

swyx

It could be that your models just weren't good enough.

William Beauchamp

No, they were great.

swyx

They were very good?

William Beauchamp

They were very good. But it was like listening to Audible or using a Kindle: you hear a voice, but you don't think, “Wow, this is special.” It's a convenience feature.

If Chai is the only platform where you can watch a MrBeast video, and it's the most engaging and fun video you want to watch, you'll go to YouTube. With audio, you can't just put it there and expect people to say, “It's 2% better,” or have 5% of users think it's 20% better. The majority of people, for the majority of the experience, have to think, “Wow, this is a big deal.”

Those are the features you need to ship. If a feature doesn't appeal to the majority of people for the majority of their experience, and it isn't a big deal, it's not going to move the needle.

swyx

I don't see it anymore.

William Beauchamp

I love this. The longer I've been working at Chai—and I think the team agrees—the more I realize that all the platitudes I thought were just platitudes, the ones you hear from Steve Jobs, are painfully true.

“Build something insanely great.” “Be maniacally focused.” “The most important thing is saying no to things you shouldn't work on.” These lessons are painfully true. Now everything I say sounds like I'm quoting Steve Jobs or Mark Zuckerberg.

Alessio Fanelli

The turtleneck.

swyx

The turtleneck.

This is my last question, and then I want to pass it to Alessio. It's about multimodality in general. Justine Moore from a16z, who's a friend of ours, asked this: a lot of people are trying to do voice, image, and video for AI companions. You said voice didn't work. What would make you revisit it?

William Beauchamp

Steve Jobs was very clear about this. There's a habit among engineers that, once they've built some cool technology, they want to find a way to package it up and sell it to consumers. That does not work.

You're free to try to build a startup around cool technology and find someone to sell it to. That's not what we do at Chai. At Chai, we start with the consumer. What does the consumer want? What is their problem? How do we solve it?

Right now, audio isn't the number-one problem for users. Image generation isn't the number-one problem either. The number-one problem in AI is that all of the AI is being generated by middle-aged men in Silicon Valley.

That's all the content you're interacting with. You're speaking to this AI for 90 minutes on average, and it's being trained by a middle-aged man. There are guys sitting around asking, “What should the AI say in this situation? What's funny? What's cool? What's boring? What's entertaining?”

That's not the way it should be. The users should be creating the AI.

The way I describe it is that Chai has an AI engine with a thin layer of user-generated content sitting on top. That thin layer of UGC is absolutely essential. It's just prompts, an image, and a name.

swyx

It's just prompts.

William Beauchamp

It's just prompts. It's just an image. It's just a name. We've done 1% of what we could do. We need to keep thickening that layer of UGC.

Users must be able to train the AI. If reinforcement learning is powerful and important, they have to be able to do that. Just as MrBeast can spend $100 million a year—or whatever it is—on his production company, with a team building the content he shares on YouTube, there needs to be a team earning or spending $100 million on the content being produced for the Chai platform. Until then, we're not finished.

That's the problem we're excited to build around. Getting too caught up in the technology is a fool's errand. It doesn't work. Start with the problem.

swyx

As a side note, MrBeast's Beast Games on Amazon Prime isn't doing well. The audience rating is high, but the Rotten Tomatoes score is poor. It's not in the top 10, and I saw that it dropped off the charts.

I'm curious because it's similar content on a different platform. Going back to what you were saying, people come to Chai expecting a certain type of content.

William Beauchamp

It's interesting to discuss moats and what the moat is. If you look at a platform like YouTube, the moat is really in the ecosystem. The ecosystem is comprised of the content creators, the users or consumers, and the algorithms.

That creates a flywheel. The algorithms are trained on users and their data. The recommendation systems feed information to content creators. MrBeast knows which thumbnail performs best, and he knows that the first 10 seconds of a video have to be a particular way. His content is highly optimized for the YouTube platform.

That's why it doesn't do as well on Amazon. If he wants to do well on Amazon, how many videos has he created on the YouTube platform?

swyx

Thousands—tens of thousands, I'll guess.

William Beauchamp

He needs to get those iterations in on Amazon.

At Chai, it's all about getting the most compelling, rich, user-generated content and putting it on top of the AI engine and recommendation systems. We want to create a beautiful data flywheel: more users, better recommendations, more creators, more content, and more users.

swyx

You mentioned the algorithm. You have this idea of ChaiVerse, and you have your own kind of LLM leaderboard or Elo system. What are your models optimized for? Can you talk about how you built it and how people submit models?

William Beauchamp

ChaiVerse is what I would describe as a developer platform. When we speak about Chai, we're usually thinking about the Chai app. The Chai app is a product for consumers. Consumers can come to the app, interact with our AI, and interact with other UGC. It's a thin layer of UGC around these bots.

Our mission is not to have a very thin layer of UGC. Our mission is to have as much UGC as possible. I don't want people at Chai training the AI. I don't want middle-aged men building AI. I want everyone building the AI—as many people as possible.

We built ChaiVerse, and it's a prototype. It started with an observation: how many models get submitted to Hugging Face each day? Hundreds. There are hundreds of LLMs submitted every day.

Consider what it takes to build an LLM. It takes a lot of work. Someone devoted several hours of compute and several hours of their time to preparing a dataset, launching it, running it, evaluating it, and submitting it. A lot of work goes into that.

We said, “Why can't we host these models for people and serve them to users?” The first issue is figuring out whether a model is good. We don't want to serve users the bad models.

We use the LMSYS-style system. It's simple and intuitive: you present users with 2 completions and say, “This is from Model A, and this is from Model B. Which one is better?”

If someone submits a model to ChaiVerse, we spin up a GPU, download the model, host it on the GPU, and start routing traffic to it. We think it takes about 5,000 completions to get an accurate signal. That's roughly how the LMSYS system works.

From that, we get an accurate ranking of which models people find entertaining and which they don't. The bottom 80% are all pretty bad, so you can disregard them. In the top 20%, you have decent models, but you can break them down into more nuanced categories.

One model might be highly descriptive. Another might have a lot of personality. Another might be very logical. Then the question is what you do with those top models.

You can try a routing approach, where, for a given user request, you predict which model the user will enjoy most. That turns out to be pretty expensive and isn't a huge source of improvement.

Something we love to do at Chai is blending. The simplest way to think about it is that you might have one model that's really smart and another that's really funny. How do you give the user an experience that's both smart and funny? You serve the smart model for 50% of the requests and the funny model for 50%.

Alessio Fanelli

Just a random 50%?

William Beauchamp

Just a random 50%. That's blending. You can do more sophisticated things on top of that, as with all things in life, but the 80/20 solution is powerful right out of the gate.

Randomness is a very powerful optimization technique. It's robust, and it lets you explore a lot of the space very efficiently.

The most exciting thing for me is what happens after the ranking. You get an Elo score, and you can track a user's first join date—the first date they submit a model to ChaiVerse. They almost always get a terrible Elo score.

Let's say their first submission gets an Elo of 1,100 or 1,000. You can see them iterate and iterate. There will be no improvement, no improvement, no improvement, and then suddenly, something works.

swyx

Do you give them any data, or do they have to figure it out themselves?

William Beauchamp

We try to strike a balance between giving them useful data and complying with GDPR. You have to work very hard to preserve the privacy of the users of your app, so we try to give them as much signal as possible while still being helpful and protecting privacy.

At a minimum, we give you a score. That alone is enough for people to optimize pretty well. They come up with theories and submit them. Does it work? No. They come up with a new theory. Does that work? No. Then, as soon as they figure something out, they keep it and iterate.

swyx

Last year, you had a post on your blog called “Crowdsourcing the 10 Trillion Parameter AGI,” and you described a mixture-of-experts recommendation system. Do you have any updated thoughts 12 months later?

William Beauchamp

The timeline for AGI has certainly been pushed out. I'm a controversial person, I suppose. I just think it's an S-curve. Everything is an S-curve.

The models have proven to be far worse at reasoning than people thought. Whenever I hear people talk about LLMs as reasoning engines, I cringe a bit. I don't think that's what they are. I think of them more as simulators.

They're like a physics simulation engine. You get these games where you construct a bridge, drop a car onto it, and the system predicts what should happen. That's really what LLMs are doing. It's not so much that they're reasoning; they're doing the most likely thing.

Fundamentally, the ability for people to add intelligence is very limited. What most people would consider intelligence isn't a crowdsourcing problem.

Wikipedia crowdsources knowledge; it doesn't crowdsource intelligence. That's a subtle distinction. AI is fantastic at knowledge, but I think it's weak at intelligence. It's easy to conflate the 2.

If you ask it, “Who was the 7th president of the United States?” and it gives you the correct answer, you might think, “I don't know the answer to that,” and conflate the result with intelligence. But that's a question of knowledge.

Knowledge is about storing information and retrieving something relevant. AI is fantastic at that. It's fantastic at storing knowledge and retrieving relevant knowledge. It's superior to humans in that regard.

We need to come up with a new word for what AI is. AI should contain more knowledge than any individual human and be more accessible than any individual human. That's extremely powerful. But what words do we use to describe it?

Alessio Fanelli

We had a previous guest from Exa AI who works on search. He tried to coin “superknowledge” as the opposite of superintelligence.

William Beauchamp

Exactly. I think “superknowledge” is a more accurate term. AI can store more information than any human, even if it isn't more intelligent, and it can retrieve that information better than any human can. I think those 2 things combined are special.

That thing will exist. It can be built. You can start with something entertaining and fun.

I often think of it as a 20-year journey, and we're in around year 4. It's like the web in 1998. You have a long way to go before the Amazons of the world become huge, multitrillion-dollar businesses that every person uses every day.

AI today is very simplistic. Fundamentally, the way we're using it, these flywheels and the ability for everyone to contribute to it, can magnify the value it brings.

Right now, it's almost sad. You have big labs—I’ll pick on OpenAI—and they go to human labelers and say, “We're going to pay you to label this subset of questions that we want to turn into a high-quality dataset.” Then they use their own powerful computers.

To me, that's so much like Encyclopedia Britannica. All the people who were interested in blockchain understood that this is the thing that needs to be decentralized. If you distribute it, people can generate much more data in a distributed fashion.

swyx

You need the incentives.

William Beauchamp

Of course. But the exciting thing about Wikipedia was the understanding that you don't need money to incentivize people. You don't need Dogecoin. Sometimes people get satisfaction simply from seeing the correct thing go up.

We do pay money for ChaiVerse. We've paid out over $100,000 to model creators. But what we saw was that it wasn't motivating. If they were submitting at a certain rate, paying them a lot of money didn't change the rate.

The money allowed them to fine-tune a Llama 7B model on 8 H100s overnight.

Alessio Fanelli

You could give them compute.

William Beauchamp

Exactly. The most excited person we ever saw from interacting with ChaiVerse was a 17-year-old kid. We gave him $1,000, and he spent all of it on a physical GPU. He sent us a picture and said, “This is what I bought, and I'm going to train more models with it.”

swyx

That's why I love it. Do you hire him?

William Beauchamp

That's the temptation, but as a platform we can't hire every good content creator. We need to build systems. The best content creator today isn't necessarily going to be the best content creator next year. We need to build the platform.

Alessio Fanelli

You talked about reasoning and knowledge. Most of the benchmarks people use are intended to mimic reasoning.

William Beauchamp

I disagree about the reasoning, but we can keep going.

swyx

How do you think about the evaluations that matter to you? Elo can't be your only evaluation. You must have internal evals.

William Beauchamp

Elo is a fantastic North Star. It's the main metric we want to see go up because it's human feedback. Humans know what they want.

When you create an evaluation, you're moving further away from the true problem. Whatever you're trying to optimize or figure out, you have to slice it. You get a snapshot, and as soon as you saturate one evaluation, you need to figure out a new one.

By simply asking humans which is better, A or B, the system is incredibly robust and generalizable. It just keeps scaling.

In the past, we've used evals to get through blockers. A great example is a safety filter. You want to make sure your models are safe, because users find that the correlation between family-friendly content and quality isn't always what you expect. People find it funny when the AI swears.

If you give me any LLM, I can make it 20% funnier just by training it to use swear words. The issue is how you measure quality improvements. Are you measuring a genuine improvement, or a superficial one?

This links back to the LMSYS style-control work. We'd rather lean on human feedback and continue making that more robust and useful. Some people are GPU-poor, and some are GPU-rich. We're feedback-rich. When you have 1.5 million people a day, you can get as much human feedback as you want.

We haven't needed evals very much. When we do, we saturate them quickly. For safety, within a month we don't need to use the eval anymore because the issue has been addressed.

swyx

Is the Elo applied to the entire user population? Clearly, there are segments: people who are into roleplay, people who are using it for therapy, and people who want not-safe-for-work content. You don't split them?

William Beauchamp

This is why I say we're in year 4 of a 20-year journey. At the end of the day, if we all go on Spotify—or imagine if Spotify only had the top 5 musicians—I think it would retain more than 85% of its existing users. If YouTube only kept its top 5 content creators, it would be enough for the vast majority of people.

One surprising thing about humans is that our preferences are fairly correlated. What you find funny and entertaining, I find funny and entertaining, and he finds funny and entertaining. There may be degrees of variation. I might find it extremely funny, and you might find it only slightly funny, but optimizing globally works very well.

Segmentation would be powerful if you found a comment incredibly boring and I found it incredibly fun. If we could segment users that way, it would unlock very powerful things. Unfortunately, that's not the shape of human behavior.

I might rank something 10 out of 10 for funniness, and you might rank it 7 out of 10. That doesn't give you as much space to work with as you would hope.

There is an element of diversity in the content AI can produce right now, but it isn't as diverse as a platform like YouTube. You can watch a MrBeast video that's completely different from a makeup tutorial. There's enough diversity that my YouTube feed is totally different from my sister's. Hers is full of women and makeup; mine is full of bald, middle-aged men talking about MMA.

With AI, it's still too early for that degree of segmentation. It will come from recommendation systems and personalization. But this is why I say: don't start with the technology; start with the problem. The problem is UGC. We must give users the tools to build more varied and engaging content.

Alessio Fanelli

I was surprised at how thin it was when I tried Chai.

William Beauchamp

It is very thin.

Alessio Fanelli

Haven't you been tempted by the ecosystem around Kobold, SillyTavern, and those platforms? They have model cards, and it seems almost like an industry standard. Can I just import those?

William Beauchamp

I remember that, in the early days of Chai, we were talking about Chai, SillyTavern, and KoboldAI. Both of them are almost as old as Chai. When Chai barely existed, they existed too, and both of us were using GPT-J.

Very early on, I thought, “These guys shouldn't even exist, because if we build a good enough platform, they should just be posting their content on our platform.”

But they're open source.

Eventually, I learned that what they're excited about is slightly different from what a typical consumer wants. It comes down to what the content creator wants. Typically, they're building it for themselves, and they want to create a specific experience for themselves.

One content creator might write 1,000 words describing a science-fiction scenario: “You're on a spaceship going off into space. These are your crewmates. One is really friendly, one is really mean, and you're the new cadet trying to rise to the top.” They can go into a lot of detail. You can give that to Llama 7B, and it will do a pretty good job of adhering to the prompt, so the user has a good experience.

On Chai, very few users will go to that level of content creation. If we can make the AI understand the user better, then instead of using 1,000 characters or 1,000 tokens to describe the scenario, the creator can say, “You're in a spaceship, you have 3 crewmates, it's going to be dramatic, and there should be some fighting.” If the AI then gives an even better experience, the content creator is happier.

Fundamentally, I think about it in terms of the steerability of the AI. A lot of the work we do at Chai is about making the AI react to the user and the content creator in the way they most want.

One analogy is TikTok. The thing TikTok did incredibly well was make it easy for anyone to create a fun video. You put some music on top, add some animations, and it's not hard to make something entertaining.

That's more like the Chai style. Users don't want to have to work. If your content is only good when you have Shakespeare writing it, that's not as good as it being something anyone at home can make.

The answer for the SillyTavern-style user is to let those people fine-tune models that create a really special effect.

swyx

As we wrap, this is the call to action part. You have Chai Grant, which I think a lot of people don't know about. It's a grant for open-source projects. Are there any ideas or projects you'd like to see people work on?

William Beauchamp

We run Chai Grant, and fundamentally we give cash with no strings attached. It's our way of giving back and supporting the community. We've benefited from many open-source packages, and a lot of our developers and engineers are very pro-open source.

It's also a great way to meet talented people and expand our connections. If anyone has a GitHub project or anything they've built that they're proud of, just apply. It's cash with no strings attached, and people have a pretty high success rate.

The other call to action is that Chai is a startup. We're a small team of around 15 people, and we work in a very intense, hardcore environment. We've found that a lot of people don't like that. They don't like this concept of work-life balance.

Once, someone said, “I can't get this done because I'm taking PTO on Friday.”

I said, “What is PTO?”

I know what it is; it stands for paid time off. The person was gone. They were no longer with the company 4 weeks later.

swyx

Legally, I think you have to allow that.

William Beauchamp

Of course. There's no problem with taking a day off. We all have personal lives. It's about responsibility. If you're not in the office on Friday, you still have responsibilities. I don't care if you work hard on Thursday to get everything wrapped up, and I don't care if you work hard on Saturday to make up for it. But the way this individual spoke about it was as if it were an excuse.

It's an environment of very talented engineers working very hard in an intense space. That's what gets me excited. It's why I love working at Chai: it's a place of talented people working extremely hard.

People who have worked at startups and love that kind of environment, who want a taste of it, should reach out and apply. I think 90% of people will say, “That sounds terrible,” and won't apply. It's not for them.

swyx

Exactly.

Alessio Fanelli

We skipped one important part. You spent $10 million on compute last year, and you said you're probably going to triple that. I'm sure you're doing a lot of work on custom kernels and inference optimization. Is there anything cool you want to share?

William Beauchamp

There are lots of cool things. Inference is extremely important. It's massively underappreciated. We can look at all the different foundation models and techniques and see the differences in how well they perform from a cost perspective.

Mixture-of-experts models tend to perform very well from a cost perspective. We've worked with a very talented team called MK1.

swyx

I saw them in the Chai logs. What are they?

William Beauchamp

We were using vLLM for a while, and vLLM is fantastic—absolutely amazing work. At some point I was introduced to the founder, Paul Merolla, who was a co-founder at Neuralink and is a real expert in hardware.

He explained, “If you know hardware really well, you can write the CUDA kernels really well. You should check out our inference engine.”

When we evaluated it, they blew vLLM out of the water. It was much, much faster. The special thing he was able to do for us is that we love rejection sampling, and we do much more rejection sampling than is typical.

We never generate just a single completion. This is why we don't do much streaming. A lot of people, like ChatGPT, used to do a lot of streaming, where the completion came out one token at a time.

Alessio Fanelli

I didn't realize Chai doesn't stream.

William Beauchamp

Normally, chat interfaces stream. Chai has never done streaming because if you stream, you're unable to do rejection sampling.

The benefit of not streaming is that you can serve a larger model. Instead of generating a completion in 4 seconds because the user gets the first token faster, you can take 10 seconds to generate it. If you've got 10 seconds, you can serve a much larger model.

People who stream get the benefit of serving a larger model, but with Chai, the full answer appears at once. We do that because we want to generate 16 completions, see the entire response for each one, and evaluate which one we think is best.

swyx

Do you have a separate LLM evaluator?

William Beauchamp

Yes. Typically, it's called a reward model. That's a term from reinforcement learning.

You can start with something simple: do you think the user is going to respond to the completion? You can take 50 million messages, look at which messages users reply to and which they don't, and train a reward model to evaluate completions.

It learns, “If you say this, the user isn't going to respond, so don't bother sending it. If you say this, the user is definitely going to engage with it, so send it.”

swyx

There's an interesting parallel between mixture-of-experts at the top, spreading out to different experts, and rejection sampling at the bottom, choosing from different paths.

William Beauchamp

I totally agree. That's the future of AI, and it's the exciting part.

Why was AlphaGo able to become superhuman? It was the ability to generate many different paths and perform tree search. If you want to talk about what intelligence might look like, it looks much more like tree search combined with the generative nature of LLMs and a really good tree search.

That's what OpenAI has done with o1 and o3.

swyx

They never said they do tree search.

William Beauchamp

It's implied.

swyx

Are you comfortable calling it a reasoning engine?

William Beauchamp

No. I'm saying it's better at reasoning because it leverages tree search.

The issue with reasoning is that the models are trained to assess whether something is logically correct and how likely it is to be logically correct. You can build sophisticated mechanisms that make the model less bad at reasoning.

But eventually, what AI is really good at won't be described as reasoning. It will always be better at retrieving and storing knowledge. That's so highly correlated with intelligence that we often assume they're the same.

What AI is truly special at, and what gets consumers excited, is that it's generative. It can just make stuff. We've never had a technology before that can simply make things.

swyx

That's the special part.

William Beauchamp

That's the exciting part.

Alessio Fanelli

Any parting thoughts?

William Beauchamp

No. It's been a pleasure. The only thing I'd add is that our office is in Palo Alto. People with startup experience who are looking to join a fast-growing, high-impact startup should reach out.

swyx

We find your culture deck really good.

William Beauchamp

Great.

Alessio Fanelli

What's the story behind the line that says if you made $100,000 trading, we'll fast-track your application?

William Beauchamp

We looked at the team, and it got to the point where almost every person on the team had done something special before joining. They had strong markers that there was something special about them.

That doesn't mean you have to have achieved something special. But we had one engineer who started college at Carnegie Mellon when she was around 15 years old. That's a bit special.

Another engineer created a GitHub repository that got around 1,500 stars. It was a low-level repository with drivers he had written. I thought, “That's a bit special.”

We had another person join the team who had made $100,000 buying and selling sneakers.

swyx

Trading?

William Beauchamp

Yes, trading. It's just this idea that, if you've been to Harvard, that's great. It shows that you're smart and work hard. But if you've actually built something and done something tangible, that gets us even more excited.

Alessio Fanelli

Thanks for having us at Chai HQ.

William Beauchamp

Thanks, guys.

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research | BidClub