Personalized AI Language Education — with Andrew Hsu, Speak
Summary
- Speak’s central wager—that speech and language models would become superhuman within five to 10 years—has matured into a business well over $50 million ARR. Andrew Hsu says “80 to 90% of the tech is here now,” enabling the original 2016 vision: a pure-software tutor that can make learners fluent faster than a human tutor. The company endured four or five painful years without pivoting because that model trajectory remained compelling.
- South Korean product-market fit came from narrowing the product, market, and customer rather than expanding them. In 2018, Speak replaced a free, multi-language catalog with guided English courses for one market and went premium, recognizing that “people don’t want to choose.” Speak is now South Korea’s biggest English app, roughly 6% of the population has tried it, and the company says it is well on the way in Japan and Taiwan.
- Hsu positions Speak as “Gen 3” language learning: functional fluency rather than gamified study or digitized textbooks. Learners rehearse sentence patterns and real situations—such as talking to an Uber driver—until speech becomes spontaneous, “almost like you’re in a gym.” The product’s private setting also removes the psychological cost of making mistakes in front of a human.
- Speak’s technical stack is hybrid, not dependent on a single frontier API. Its custom streaming ASR is trained on extensive non-native speech and powers latency-sensitive drills; Whisper and LLMs handle open-ended speech, semantic feedback, and tutoring. Hsu’s operating model is to “saturate model capability” with products, then repeat when the frontier advances.
- The next scaling engine is AI-generated curriculum governed by human pedagogy and an evolving learner-specific knowledge graph. Speak wants “100x more content,” 10x more languages, and eventually 100x more language pairs, using tutor and curriculum-writing agents while humans review syllabi and lessons. Vocabulary, sentence patterns, and clustered mistakes should ultimately roll into a holistic Speak Score, where 54 versus five is directionally meaningful even without teaching to an exam.
- Real-time voice is close, but unit economics and interaction design—not raw model intelligence—remain the bottlenecks. OpenAI’s real-time API pricing works better when replacing hourly labor than when consumers may converse for hours; a mistake at Speak’s scale could cost “millions of dollars.” Hsu also calls first-audio latency a “vanity metric” unless it includes turn detection, because learners may hesitate for 10 seconds mid-answer.
- Hsu sees real-time translation as more complement than existential threat, while language remains the beachhead for a broader learning platform. He argues that German-to-English translation must wait for the sentence-final verb, and says Speak’s Asian users want direct human connection, not merely a Babel fish; he nevertheless expects translation inside Speak. B2B is already expanding toward communication, management, and hospitality skills, supporting his larger view that AI will reinvent learning even though “real-world inertia is enormous.”
Deep dive
1. Speak survived because its original model thesis never changed
Hsu entered the first Thiel Fellowship class in 2011 at 19, already in graduate school after an accelerated education. The fellowship’s offer—$100,000 for people under 20 to leave school and pursue almost anything—was “life-changing”; it also connected him to his eventual Speak co-founder, who joined the second class.
In 2016, the founders took a year-long sabbatical to study AI, including conversations with Andrej Karpathy as he finished graduate school. Their conviction was unusually specific: within five to 10 years, speech and language models would become superhuman, making a language tutor built from “pure software, pure AI” possible.
The timing was badly wrong before it was right. Four or five early years were painful enough that the founders repeatedly asked, “Why are we doing this?” Yet they never pivoted; when the company replayed its original YC application in Taipei, its claims matched Hsu’s current pitch almost word for word. He now estimates that “80 to 90% of the tech is here.”
2. South Korea turned focus and premium pricing into product-market fit
Speak’s first “red app” offered free packs of content across multiple languages and let users choose what to study. It failed. In 2018, the team tore it down, focused exclusively on South Koreans learning English, created courses and new lesson types, and put the experience on rails: “They’re already using some of their motivation…just to open the app. They don’t want to make another choice.”
Abandoning free access was equally important. Premium pricing filtered for people already motivated to learn English, sidestepping part of the motivation question. No single change was the silver bullet; Hsu attributes the turnaround to accumulated lessons across three or four years.
Korea was not initially inevitable—the team almost selected Taiwan. The choice was somewhat serendipitous. On the ground, Seoul’s academies and “skyscrapers full of classrooms” revealed intense demand. The wager was that winning a market full of human competitor products and people who fundamentally cared about fluency would establish strong, portable PMF: “If we can really make headway and win this market…then we probably have something pretty real and strong PMF.”
Localization supplied credibility. Speak’s first employee, Sungjae, helped refine details down to button wording, and early users were shocked to learn the company was American. Speak is now South Korea’s biggest English app; about 6% of the population has tried it, and the business exceeds $50 million ARR.
3. “Gen 3” language learning trains automatic speech, not test knowledge
Hsu divides the market into Rosetta Stone’s CD-ROM “Gen 1,” mobile-first “Gen 2,” and an AI-native third generation. He praises Duolingo’s gamification but says its closest analogue is a productive mobile game; Speak instead targets functional fluency through tutoring and open-ended practice.
Speak does not organize its core method around vocabulary and grammar. It teaches sentence patterns, then makes users combine and repeat them “almost like you’re in a gym” until speaking becomes spontaneous. Role-plays emphasize situations such as conversing with an Uber driver rather than recalling textbook rules.
“Magic onboarding” is meant to signal this new category immediately: a tutor asks about goals, then a separate LLM converts the answers into abstracted summaries rather than showing the full transcript. The result remains unsettled—install-to-sign-up is lower because speaking is harder than tapping, while trial starts are higher. Hsu explicitly says, “We still don’t know yet.”
4. Custom ASR and frontier models occupy different latency regimes
Before LLMs, Speak built custom speech-recognition models and accumulated large volumes of non-native English audio from daily lessons. That system remains valuable because core recording loops must stream quickly; the company uses its data to fine-tune models and understand its own users rather than treating generic transcription as sufficient.
Whisper’s 2022 arrival was the long-awaited threshold. Four employees tested audio from a beginner Korean learner that none could understand with their eyes closed; Whisper transcribed it correctly. For Hsu, that was not merely an improvement but direct evidence of “superhuman” speech recognition.
Whisper and GPT-3.5 Turbo then expanded Speak beyond listen-and-repeat. The product could interpret spontaneous responses, explain errors, and say that a native speaker would choose a different word or formulation. The simpler pre-Whisper product had already reached several million dollars of ARR in South Korea, but these models turned supplemental speaking practice into a more complete tutor.
Hsu rejects the idea that better foundation models invalidated Speak’s earlier work. Custom ASR still serves fast, constrained interactions; frontier models serve semantic and open-ended ones. The road-map discipline is to build until Speak can “saturate model capability,” then build again after the next advance.
5. AI-generated curricula are a leverage play, not a no-human pipeline
Speak’s in-house “Speak method,” teachers, scripts, and production tools created heavy operational overhead. Expanding now requires “100x more content,” 10x more languages, and eventually 100x more native-language-to-target-language pairs—far beyond what its Los Angeles studio and manual writing process can supply.
The company is building a tutor agent, curriculum-writing agent, and large LLM-based pipeline that scaffolds curricula and writes lessons. Hsu uses “agent” cautiously, but the objective is explicit: keep the organization small while launching into substantially more languages and markets.
Evaluation remains the hard part. Even teaching a new human writer why one subtly different lesson fits the Speak method better than another is difficult to articulate. Speak relies heavily on its content team, is developing model-graded evals, and may eventually reinforcement-fine-tune a curriculum agent on internal data; Hsu calls that work “still pretty early.”
The hosts frame this as job transformation rather than simple deletion: instead of one person authoring two courses, that person might review 50 machine-generated ones. Hsu agrees with the leverage thesis—human review still covers syllabi, curricula, and individual lines, but the same team should eventually launch “100x more courses.”
6. Fluency is jagged, measurable, and grounded in real-world tasks
Hsu’s core example is intentionally practical: can a learner visit Mexico City and order at a street taco stand? Competence there does not imply an ability to discuss family, so “the frontier of fluency is very jagged.” Speak consequently models multiple subscores rather than treating language ability as one uniform ladder.
Speak is developing a domain-specific knowledge graph to organize vocabulary, sentence patterns, and clusters of mistakes accumulated over time. These dimensions should ultimately roll into a Speak Score: 54 means little in isolation, Hsu concedes, but its difference from five is intuitively meaningful and can map to concrete things the learner can do.
The hosts challenge the absence of a recognized exam target. Hsu’s response is that Speak does not “teach for the test”; it may eventually offer test preparation, but current optimization is real-world proficiency. From A0 through B1, learners share a fairly linear conceptual backbone; intermediate and advanced paths diverge more sharply, with the graph modifying that foundation around individual weaknesses.
Speak teaches casual language rather than “textbook English,” although current scale forces pragmatic defaults such as standard American English. For now, each language gets a chosen standard; Hsu expects more sharply differentiated accents and dialects later. Pronunciation is secondary to communicating an idea: users should “literally move your mouth and…make the sounds.” An English-only coach, fine-tuned from wav2vec on Speak’s phonetic data, currently scores single words and is planned for sentences and additional languages.
7. Real-time voice is gated by code-switching, turn detection, and cost
A genuine language tutor must code-switch: an English speaker learning Spanish should move naturally between both languages, even inside one sentence. Few TTS models can pronounce that properly. The hosts suggest routing by detected language, but Hsu notes that sub-word switching defeats simple routing and concatenation because the result no longer sounds human.
Speak received early access to OpenAI’s real-time API, but Hsu clarifies that nothing was yet in production. Pricing fits a customer-support agent replacing hourly labor better than a consumer practicing for hours. Speak is close, but custom WebRTC infrastructure and scale mean an architectural mistake “will cost us millions of dollars.”
The initial application is a three-to-five-minute instructional lesson that alternates listening or viewing with short interactive conversations. The lesson is semi-on-rails, and Speak built scaffolding to switch reliably between those modes; it is intended to augment the existing teacher videos, not immediately replace every lesson.
Conventional request-to-first-audio timing is, in Hsu’s words, “actually like a vanity metric.” The meaningful interval starts when the learner stops speaking, and turn detection can add a second or more. Semantic VAD works for fluent conversation but fails when a learner pauses for 10 seconds mid-sentence, making custom turn detection the dominant perceived-latency problem.
8. Language is the beachhead for a broader AI learning company
Speak now teaches English in 40 more countries, has Spanish and French live, and planned several additional languages for that year. B2B began roughly a year earlier as a side experiment, then “just started working”; Hsu expects it to become meaningful alongside the still predominantly consumer business.
Hsu treats the translation threat as technically and emotionally incomplete. In German, a sentence-final verb means English translation cannot make progress until the whole sentence arrives, so he argues it can never be “truly truly perfect.” More importantly, Asian users want self-improvement and direct connection: “They want to be able to look you in the eye and speak English.” Hsu nevertheless expects Speak eventually to incorporate translation into learning.
Speak for Business could eventually connect to work documents through a Mac app or browser integration, although Hsu calls that “a whole can of worms.” When a host warns that employees may avoid a language tool secretly evaluating management skills, Hsu immediately concedes the products should be separated.
The longer arc extends from language into communication, hospitality, management, and eventually subjects such as math or biology. Hsu’s closing tension is that AI is “probably the most transformative technology we’ve ever built,” while lives outside the Bay Area have changed “pretty close to zero.” His prescription is more application builders and more scaled consumer products; when AI anxiety surfaces, he returns to one controllable variable: “I just focus on our users.”