No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski
Summary
ElevenLabs reached $300 million ARR with 350 employees by pairing a five-million-MAU creator product with an enterprise agents business approaching half of revenue. Self-serve subscriptions and creators remain roughly 50% of the mix; several thousand enterprise customers range “from Fortune 500s to some of the fastest-growing AI startups.” The operating signal is breadth without abandoning the original research advantage.
The founding wedge was Poland’s broken dubbing experience, where one flat narrator voices every character, but the thesis expanded into voice as computing’s natural interface. ElevenLabs aimed to carry the “original voice, original emotions, original intonation” across languages, as demonstrated with Lex’s interview with Narendra Modi. Mati’s larger claim is that keyboard-and-screen interaction “feels broken,” making the opportunity far larger than conventional dubbing spend.
Mati’s operating model is research first, product in parallel, with small cross-functional labs formed around concrete problems. A roughly five-person voice lab solved human-sounding narration before ElevenLabs expanded into audiobooks and dubbing; an agent lab then combined speech-to-text, LLMs, text-to-speech, integrations, testing, and monitoring. Product teams see the research roadmap and generally build alternatives only when the relevant breakthrough looks more than three months away.
Customer support is the first major agent category, but voice is moving from reactive ticket resolution into discovery, checkout, education, media, and government services. Meesho’s agent progressed from refunds and tracking to product guidance and potentially checkout; MasterClass lets learners practice a negotiation with Chris Voss; Ukraine is pursuing what Mati calls a “first agentic government.” That expands the economic case from labor savings toward revenue generation and entirely new experiences.
Voice quality is not one scalar benchmark because preference can flip with voice identity, language, audience, and delivery even when underlying model quality differs. One customer requested a voice “as robotic as possible,” while a Japan-and-Korea deployment wanted an excitable voice for younger callers and a calm, slow one for older users. ElevenLabs therefore supplies voice coaching and expects selection eventually to become dynamic for each person and context.
ElevenLabs assumes base models will commoditize and treats research leadership as only a six-to-12-month advantage. Mati says “research is a head start”; durable value must accumulate through distribution, voices, integrations, workflows, and a product layer that connects models to business logic. Open-source narration is already strong, but controllability and real-time orchestration remain differentiated, while a concentrated pool of perhaps 50–100 elite audio researchers—Mati says roughly 10 are probably at ElevenLabs—can still produce architectural leaps.
The near-term roadmap targets controllable multimodal creation and lower-latency, emotionally aware agents while retaining cascaded systems for reliable enterprise work. Scribe v2 is reported at under 150 milliseconds and 93.5% accuracy across the top 30 languages on FLEURS; fused speech-to-speech may become more expressive but can introduce hallucinations and offers less visibility. Sarah estimates that humanlike conversational interaction is still at least a year away but may arrive within a year, with real-time dubbing or translation within two; personalized education—“your own teacher on demand”—is the largest unrealized application.
Deep dive
1. A bad Polish dub revealed a global interface market
Sarah Guo grounds the discussion in scale: ElevenLabs has 350 employees, $300 million ARR, more than five million monthly active creative users, and several thousand enterprise customers. Revenue is “roughly 50/50” between self-serve creative subscriptions and an enterprise business increasingly centered on agents.
Her original investor pushback is worth keeping: creation tools such as ElevenLabs, Midjourney, Suno, and Hunyuan could look like products for a narrow population, while traditional dubbing itself is not enormous. Mati’s conviction came from Poland, where foreign movies routinely gave every male and female character the same flat narrator—“a terrible experience.”
The proposed replacement preserves the performer rather than merely translating words: “original voice, original emotions, original intonation carried across.” ElevenLabs demonstrated that idea on Lex’s interview with Narendra Modi, but Mati extends it to all computing: humans still spend their time at keyboards and screens even though speech is “the most natural interface there is.”
2. Problem-specific labs connect foundational research to products
ElevenLabs initially tried optimizing existing models for narration and dubbing, but their speech was too robotic for people to enjoy. Mati credits his co-founder—whom he has known for 15 years—and the early research team with producing the foundational improvement the product required.
The first “voice lab,” initially about five people, combined researchers, engineers, and operators around one mission. Research came first, followed by a simple usable layer, then progressively fuller workflows for audiobooks, movie narration, and dubbing.
Once human-sounding output worked, an “agent lab” tackled knowledge on demand by orchestrating speech-to-text, LLMs, and text-to-speech. The product requirement quickly expanded beyond latency and accuracy into legacy-system integrations, functions, production deployment, testing, monitoring, and evaluation.
Some labs followed expressed demand: customers wanted music beside generated speech, prompting a fully licensed music model and eventually a broader creative suite incorporating partner image and video models. Dubbing was more conviction-led; Mati expects the “full Babel fish idea” of real-time communication across language barriers to become immense.
3. Voice quality cannot be reduced to one leaderboard
Sarah identifies the buyer problem: most enterprise decision-makers are not machine-learning scientists, while even researchers lack benchmarks for every audio domain. Customers may recognize a convincing clone, but they do not necessarily know how to select or evaluate the best voice.
ElevenLabs responds with a dedicated voice specialist who translates branding and audience requirements into voice selection, including access to iconic talent such as Sir Michael Caine. Briefs vary radically: one major European company requested a voice “as robotic as possible,” while a Japan-and-Korea deployment paired an excitable voice with younger callers and a calm, slower voice with older customers.
Mati’s harder problem is that swapping voices can materially change perceived model comparisons even when technical quality differs. Standard labeling also captures what was said better than “how it was said”—emotion, accent, and delivery—so ElevenLabs had to build that qualitative audio-labeling capability itself.
4. Agents are moving from support into transactions and experiences
Customer support is scaling fastest, but Mati sees a shift from “I have a problem” toward proactive assistance across the customer journey. At Indian e-commerce company Meesho, the agent moved beyond refunds and package tracking to product discovery, gift recommendations, navigation, and potentially checkout; he also cites Square in connection with the shift from voice ordering toward discovery.
Interactive media turns static intellectual property into something conversational. ElevenLabs worked with Epic Games to put Darth Vader into Fortnite, where millions of players could interact with the character live—a template Mati expects to extend to books, games, and other story worlds.
Education is his strongest application thesis. Chess.com can make Hikaru Nakamura, Magnus Carlsen, or the Botez sisters into teachers; MasterClass lets a learner call Chris Voss and practice a negotiation rather than merely watch his lesson. The model changes content from sequential playback into interactive practice.
In Ukraine, ElevenLabs is working with the Ministry of Digital Transformation on what Mati calls a “first agentic government.” The proposed system combines citizen support about benefits, employment, and processes; proactive public information; and personal tutoring. A central digital-transformation function is paired with engineering leaders in individual ministries.
5. Platform breadth—not a point solution—is ElevenLabs’ competitive answer
Sarah maps the enterprise choices: deploy through Palantir or a large consultancy, use a platform such as ElevenLabs or OpenAI, or buy a use-case specialist such as Sierra. Mati’s candid boundary is that for one isolated problem, ElevenLabs is “likely” not the best choice.
The platform fits organizations deploying conversational interfaces across customer support, internal training, sales, and new customer experiences. Customers can use selected components rather than the whole stack, combine them with existing internal investments, and bring in ElevenLabs engineers; international voices, languages, and integrations are another stated advantage.
Against Google and OpenAI, Mati argues that audio needs both focused research and the unglamorous product layer that meets customers where they are. He claims benchmark leadership in text-to-speech, speech-to-text, and orchestration, adding that audio depends less on sheer scale than architectural breakthroughs: “the number of people doesn’t matter,” but exceptional people do.
His concession is explicit: “In the long term models will commoditize,” perhaps in two, three, or four years. Out-of-the-box narration quality is already converging across open-source and commercial models; controllability remains harder. The lasting sequence is research, product, then an ecosystem of distribution, voices, integrations, and workflows.
6. Research buys time while the interface shifts toward tutors and robots
Sarah offers the uncomfortable formulation that technology advantages may last one year or 10, but are not infinitely defensible. Mati agrees: research supplies “six to 12 months of advantage,” letting customers receive a capability earlier while ElevenLabs builds the surrounding product. Teams work alongside internally shared research initiatives and use three months as the rough build-versus-wait threshold.
Current work spans controllable text-to-speech, speech-to-text across almost 100 languages, a fully licensed music model, and combinations of audio with visual models. On agents, Scribe v2 is reported below 150 milliseconds with 93.5% accuracy across the top 30 languages on FLEURS, while a forthcoming orchestration mechanism should reduce end-to-end latency and incorporate emotional context.
Cascaded speech-to-text, LLM, and text-to-speech remains Mati’s choice for reliable enterprise deployments over the next year because every step is inspectable and tools can be called. Fused speech-to-speech may be more expressive but comes with hallucination risk. Sarah—not Mati—estimates that a conversation like theirs is still at least a year away but may arrive within a year, and places real-time dubbing or translation within two.
Mati expects companions to become significant but personally prefers a “Jarvis” super-assistant to a social substitute: something that understands his context, opens the blinds, reports weather, and plays music. His highest-conviction future remains “your own teacher on demand,” balanced by explicit technology-free human interaction; voice then becomes a key interface for the “decade of agents” and subsequent “decade of robots.”