Pioneers Insight Method Research Author
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Back to Episodes

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research

Summary

  • Chai’s thesis is that social AI will become a distributed creator platform, not one monolithic model advantaged solely by more data, compute, and researchers. William Beauchamp expects specialized teams and users to shape different AIs, much as quantitative firms specialize inside one financial marketplace. “There must exist a platform where a small team can produce an AI for a unique purpose.”

  • Product-market fit appeared only after utility bots failed and Beauchamp’s sister built a therapist bot that drew 20 users for roughly 20 minutes each. News, recipes, jokes, quizzes, celebrities, and influencers went nowhere; immediate, judgment-free conversation was “the thing that AI is 10x better at.” Chai then let consumers create characters using a prompt, image, and name, uncovering demand for conflict, romance, and archetypes its engineers would never have designed.

  • Alessio cited Chai at 1.4 million DAU and over $22 million in revenue; Beauchamp did not explicitly confirm those figures, saying instead that he thinks users grew 3x and revenue more than doubled last year. A typical session lasts about 90 minutes versus the cited 70 minutes for TikTok and generates roughly 150 messages, making inference economics far more important than in question-answer products. The relevant frontier is not benchmark performance alone but “performance per dollar.”

  • The 2024 growth curve reflects four major changes: repairing infrastructure, suspected weaker competitive acquisition, better models, and paid distribution. Firebase stopped scaling reliably around 500,000 DAU, forcing a painful three-month migration before Chai could reach 1.5 million; later, Beauchamp suspected Character.AI reduced advertising after its founders left. Chai then ramped acquisition to about $40,000 a day, moving annualized user growth from roughly 2x to 3x: build the product, then attach a large “rocket” by buying ads.

  • Chai’s strongest operating advantage is a human-feedback loop that compresses a conventional 30-day retention test into roughly three hours. Each submitted model is put before users for comparisons; Beauchamp said roughly 5,000 completions provide an accurate signal. About five researchers now evaluate 20–50 models daily and ship at least 100 a week. Somewhere in October, he thinks, user commentary flipped from Character.AI being better to Chai being better—evidence he values because “you can’t cheat consumers.”

  • The company killed voice as a growth thesis after three months of work produced no measurable retention, engagement, or monetization lift. Only 10–15% of users tried it, for 10–15% of their time—about 2% of the total experience—despite Beauchamp insisting the models were great. His product test is now severe: a feature must matter to most users for most of the experience and feel like “a big deal.”

  • Beauchamp has pushed out his AGI timeline and rejects the idea that current LLMs are intrinsically strong reasoning engines. He calls them simulators with “superknowledge”: exceptionally good at storing, retrieving, and generating, but still weak at intelligence as he defines it. Chai does not stream; it generates 16 complete candidates and uses a reward model to select one. He described training such a model from 50 million messages as an example, not as a confirmed deployed configuration.

Deep dive

1. Quant profits financed a search for impact

  • Beauchamp graduated from Cambridge in 2012, having accumulated about $100,000 playing poker. Small capital was an advantage: an anomaly producing $100,000 annually is only 1% on $10 million but 100% on $100,000, so he taught himself Python and machine learning to trade it.

  • The firm eventually made about $5 million a year with roughly 15 Oxford- and Cambridge-educated mathematicians and physicists, trading only the team’s money. There were “no customers complaining” and no investors constraining risk—the quantitative-trading dream Beauchamp had wanted.

  • At 30, he decided another yacht-sized increment of wealth would not create meaningful impact. Crypto looked exceptional as gambling and for evading monetary regulations and banking restrictions, but its broader blockchain and Web3 rationale “didn’t really make much sense,” so he redirected his efforts toward language models.

2. Machine learning’s S-curves undermined the monolithic-AI story

  • Reading published work from Google and the still-open OpenAI convinced Beauchamp that LLMs would matter. Yet he rejected the prevailing race for one intelligence assembled from the most data, compute, and researchers: machine-learning performance, in his experience, follows an S-curve and usually plateaus around human capability.

  • Self-driving, image recognition, and speech recognition supported that view; AlphaGo was the conspicuous superhuman exception. Beauchamp therefore expected AI to resemble finance, where high-frequency, mid-frequency, equity, and other specialists compete through separate algorithms on a shared marketplace rather than inside one all-powerful quant firm.

  • Chai’s founding proposition followed directly: “There must exist a platform where a small team can produce an AI for a unique purpose.” The intended analogue was less Encyclopaedia Britannica than Wikipedia, YouTube, or Twitter—an ecosystem where distributed contributors discover what a central institution cannot.

3. Failed utility bots exposed conversation as the native product

  • Chai initially let developers submit Python agents with text-in/text-out interfaces. Beauchamp built a Reddit-news bot, a recipe assistant, dad jokes, quizzes, and facts; despite anticipating products resembling later answer engines, he found “clearly no product-market fit” because the models were weak and conversational utility added little.

  • His sister’s therapist bot changed the direction overnight: about 20 active users spent an average of 20 minutes with it. Conversation could be immediate at 3 a.m., judgment-free, and easier than waiting for a friend; even an AI-generated compliment produced an experience utility products had not.

  • Beauchamp contrasts this participation with passive TikTok or Instagram consumption. Forty minutes of swiping can create remorse because “I achieved nothing”; interacting with an AI feels contributory. He cited user reports that Chai helped with eating disorders, depression, and rough patches, while presenting them as what users say.

4. Consumer authors, not software developers, unlocked the catalog

  • Attempts to manufacture demand with a kbot, celebrities, and influencers also failed. The breakthrough was recognizing that Python developers did not want to build social characters, but consumers did—so Chai exposed a 6-billion-parameter GPT-J model through a prompt, image, and name.

  • Users produced categories Beauchamp would never have predicted: playground bullies, arguments, fights, and unfamiliar romantic archetypes. Instead of Chai guessing what 1% of people wanted, user creation supplied the variety required to address a much broader population.

  • The hosts’ criticism—that Chai’s creator layer still looks surprisingly thin—was accepted outright. Beauchamp’s answer is to make short descriptions more steerable, so “a spaceship,” three crew members, drama, and fighting can outperform the thousand-word character cards used by expert SillyTavern-style creators.

5. Venture-funded competition taught Chai the price of model quality

  • By late 2022 or early 2023, Beauchamp recalls Chai reaching roughly 100,000 DAU and becoming the App Store’s leading AI app. When Character.AI appeared with a very similar experience, his team initially laughed at the product; then it raised $100 million, followed by another $100 million.

  • Beauchamp had invested maybe $2 million himself and was serving GPT-J 6B. Chai’s illustrative economics were about $1 per user over an entire session; at 1 million users, that would mean about $1 million spent on AI in aggregate. Character.AI could spend 100 times that, and users noticed: “Why is your AI so much dumber?” Chai moved to Silicon Valley, obtained funding, and learned that consumer AI was a Silicon Valley-style hyperscale business.

  • Alessio said Chai was at 1.4 million DAU and over $22 million in revenue. Beauchamp responded that he thinks users grew by a factor of 3 last year and revenue more than doubled. He said he thought Character.AI had almost a $3 billion valuation and 5 million DAU.

  • His DeepSeek comparison centered on inference economics and founder-led, customer-obsessed execution. He praised DeepSeek’s latest V2 for its inference engine and significantly smaller KV cache, which reduced inference costs, and said performance per dollar matters more than benchmark scores. He was interested in whether Llama 4 could match that gain.

6. Infrastructure and distribution explain the visible growth kinks

  • Chai used GCP from one DAU through roughly 500,000, leaning particularly hard on Firebase—about three times beyond a level Google engineers reportedly recommended. That abstraction let the team focus on AI until outages forced at least three months of migrations and service separation.

  • Outages damaged more than same-day traffic. New users encountered a broken app, hurting retention, spending, ratings, and therefore App Store ranking; recovering organic placement could take much longer than fixing the backend. The rebuilt stack then supported growth toward 1.5 million DAU.

  • Beauchamp suspects that after Character.AI’s founders left, the company dialed down user acquisition. He illustrated the competitive effect with a hypothetical company spending $100,000 a day versus one spending nothing; he did not present that amount as Character.AI’s reported spend.

  • A former ByteDance head of growth was astonished that Chai had reached about 1 million DAU without advertising. Chai tested roughly $10,000, then $20,000, and now about $40,000 daily; Beauchamp says that converted a roughly 2x annual growth trajectory into 3x.

7. Three-hour feedback loops became the model-development engine

  • Beauchamp’s operating maxim is that “success is born out of failures”: flat periods represent learning, while rising periods harvest it. Chai’s critical Q2 development was an evaluation system that puts any submitted model before about 5,000 users and ranks which outputs they find more entertaining or engaging.

  • That changed about five researchers from evaluating perhaps three models a week—and once struggling to test five a month—to 20–50 daily and at least 100 weekly. A standard cohort test waiting 30 days for day-30 retention became a roughly three-hour signal.

  • The resulting velocity let Chai rapidly iterate DPO fine-tuning, prompts, blending, rejection sampling, and reward models. Somewhere in October, he thinks, Reddit feedback and conversations with users flipped from “Character.AI is better” to “you guys are better”; with low switching costs and 90-minute sessions, Beauchamp argues users cannot be fooled about quality.

8. Audio’s failure imposed a harsher product standard

  • Chai spent three months solving voice latency, cost, quality, activation, and interaction design, launching at least nine months before Character.AI by Beauchamp’s account. The A/B test showed no movement in retention, engagement, or monetization, prompting a week of checks for a nonexistent bug.

  • A host suggested the models simply might not have been good enough; Beauchamp’s emphatic response was, “No, they were great.” Only 10–15% of users activated audio and used it 10–15% of the time, changing roughly 2% of the aggregate experience.

  • His lesson is that convenience does not create a destination. A successful feature must give the majority of users, through the majority of their experience, something uniquely compelling enough to provoke “wow, this is a big deal”; otherwise even excellent technology cannot move company-level metrics.

  • Hence Beauchamp says audio and image generation are not users’ number-one problems. “All the AI is being generated by middle-aged men in Silicon Valley,” he argued; the actual unmet need is allowing users to train and shape the experience themselves.

9. Chai’s intended moat is a progressively thicker UGC layer

  • Today’s prompt, image, and character name are, in Beauchamp’s estimate, “1% of what we could do.” His completion criterion is deliberately extreme: Chai is unfinished until a creator team can earn or spend $100 million a year—or whatever the figure is—producing AI content for the platform, as major video creators do.

  • The proposed moat combines creators, consumers, and algorithms. User behavior trains recommendations; recommendations tell creators what works; better content attracts more users. Beauchamp used MrBeast’s weaker fit on Amazon as the example: YouTube iterations optimized his thumbnails, openings, and content for that specific ecosystem.

  • Chai wants TikTok-like creation leverage, where ordinary people can produce something fun through built-in music and effects. “Users don’t want to have to work”; the platform should make short prompts effective, while advanced creators contribute fine-tuned models that generate genuinely distinctive behavior.

10. ChaiVerse turns live human taste into an open model tournament

  • Hundreds of models reach Hugging Face each day after creators invest data, compute, and labor. ChaiVerse offers to host those models, route traffic to them, and collect pairwise user judgments. Beauchamp said roughly 5,000 completions are needed for an accurate signal, following an LMSYS-style approach.

  • His distribution was that the bottom 80% are “pretty bad” and can be disregarded. The top 20% reveal useful differences—description, personality, humor, or logic—but request-level routing among them proved expensive and supplied little edge.

  • Chai instead favors blending: serve a smart model for a random 50% of requests and a funny model for the other 50%. “Random is a very powerful optimization technique,” Beauchamp argued, because it explores broadly while remaining unusually robust.

  • He illustrated the iteration process with first submissions that might score around 1,000–1,100 Elo, followed by repeated failures and then a sudden improvement. Chai has paid creators more than $100,000, but payments did not increase submission rates; they mainly financed compute, exemplified by a 17-year-old who spent a $1,000 award on a physical GPU.

11. Human preference is the North Star, despite its distortions

  • Challenged that Elo cannot be Chai’s only evaluation, Beauchamp called it the North Star because “humans know what they want.” Designed evals are snapshots that saturate and require replacement; pairwise preference remains general and scales with Chai’s unusual position of being “feedback rich.”

  • He conceded that raw preference rewards superficial tricks: any LLM can become 20% funnier, in his example, by training it to use swear words. Chai has used targeted evals for blockers such as safety, but says those tests can saturate within a month; the longer-term answer is making human feedback more robust, not substituting static benchmarks.

  • A host’s segmentation pushback—therapy, role-play, and not-safe-for-work users clearly differ—met a counterintuitive response. Beauchamp says preferences remain highly correlated: one person may rate an answer 10/10 and another 7/10, but powerful personalization requires one group to love what another finds boring. AI content is not yet diverse enough to create YouTube-scale feed divergence.

12. Superknowledge and inference-time search replace the near-term AGI story

  • Beauchamp says his AGI timeline has “certainly been pushed out.” LLMs look less like reasoning engines than simulators predicting the most likely continuation, analogous to a physics game simulating what happens when a car falls onto a constructed bridge.

  • His distinction is knowledge versus intelligence: models can store and retrieve more information than any human, yet that advantage is easily mistaken for reasoning. He accepted “superknowledge” as the better term and placed AI in year four of a 20-year journey—roughly the web in 1998.

  • William linked OpenAI’s o1 and o3 to tree-search-like approaches; swyx noted that OpenAI had not said it uses tree search. William called it implied and said such systems are better at reasoning, while rejecting the label of reasoning engine. Their native strengths, he maintained, are retrieval, storage, and generation: “It can just make stuff.”

  • Alessio said Chai spent $10 million on compute last year and that it would probably triple that; William then emphasized inference optimization. He said Chai had evaluated MK1’s inference engine as much faster than vLLM and highlighted founder Paul Merolla’s hardware expertise and CUDA-kernel work.

  • Chai has never streamed because streaming prevents rejection sampling. Rather than optimize for a first token in roughly four seconds, it can take around 10 seconds to produce a full answer and serve a larger model. It generates 16 complete candidates, then uses a reward model to select one. One example of such a model would use 50 million messages labeled by whether users responded, predicting which completion is likely to prompt a reply.