Pioneers Insight Method Research Author
Back to Pioneers
Andrew Hsu
Innovators 1 Curated Dialogues

Andrew Hsu

Key Views & Dialogues

Personalized AI Language Education — with Andrew Hsu, Speak

  • 🗓️ Date2025-07-11 | 🎙️ Show:Latent Space

Speak’s focused premium English product helped it become South Korea’s biggest English app and exceed $50 million ARR, supporting its Gen 3 fluency thesis. Custom speech models, frontier LLMs, and AI-generated curricula could multiply content and languages, but evaluation, turn detection, real-time costs, and learner-specific pedagogy remain execution risks.

View Dialogue Notes & Key Takeaways
  • Speak’s central wager—that speech and language models would become superhuman within five to 10 years—has matured into a business well over $50 million ARR. Andrew Hsu says “80 to 90% of the tech is here now,” enabling the original 2016 vision: a pure-software tutor that can make learners fluent faster than a human tutor. The company endured four or five painful years without pivoting because that model trajectory remained compelling.

  • South Korean product-market fit came from narrowing the product, market, and customer rather than expanding them. In 2018, Speak replaced a free, multi-language catalog with guided English courses for one market and went premium, recognizing that “people don’t want to choose.” Speak is now South Korea’s biggest English app, roughly 6% of the population has tried it, and the company says it is well on the way in Japan and Taiwan.

  • Hsu positions Speak as “Gen 3” language learning: functional fluency rather than gamified study or digitized textbooks. Learners rehearse sentence patterns and real situations—such as talking to an Uber driver—until speech becomes spontaneous, “almost like you’re in a gym.” The product’s private setting also removes the psychological cost of making mistakes in front of a human.

  • Speak’s technical stack is hybrid, not dependent on a single frontier API. Its custom streaming ASR is trained on extensive non-native speech and powers latency-sensitive drills; Whisper and LLMs handle open-ended speech, semantic feedback, and tutoring. Hsu’s operating model is to “saturate model capability” with products, then repeat when the frontier advances.

  • The next scaling engine is AI-generated curriculum governed by human pedagogy and an evolving learner-specific knowledge graph. Speak wants “100x more content,” 10x more languages, and eventually 100x more language pairs, using tutor and curriculum-writing agents while humans review syllabi and lessons. Vocabulary, sentence patterns, and clustered mistakes should ultimately roll into a holistic Speak Score, where 54 versus five is directionally meaningful even without teaching to an exam.

  • Real-time voice is close, but unit economics and interaction design—not raw model intelligence—remain the bottlenecks. OpenAI’s real-time API pricing works better when replacing hourly labor than when consumers may converse for hours; a mistake at Speak’s scale could cost “millions of dollars.” Hsu also calls first-audio latency a “vanity metric” unless it includes turn detection, because learners may hesitate for 10 seconds mid-answer.

  • Hsu sees real-time translation as more complement than existential threat, while language remains the beachhead for a broader learning platform. He argues that German-to-English translation must wait for the sentence-final verb, and says Speak’s Asian users want direct human connection, not merely a Babel fish; he nevertheless expects translation inside Speak. B2B is already expanding toward communication, management, and hospitality skills, supporting his larger view that AI will reinvent learning even though “real-world inertia is enormous.”

  • 🔗 Original source & video: Personalized AI Language Education — with Andrew Hsu, Speak

Listen to full conversation →