Pioneers Insight Method Research Author
Back to Pioneers
Mati Staniszewski
Researchers 3 Curated Dialogues

Mati Staniszewski

ElevenLabs · Co-Founder & CEO

Frontier Insights

Frontier Thesis: Voice is the next universal AI interface, but raw model advantages decay within 6–12 months. Sustainable defensibility shifts from foundational weights to vertical workflow orchestration, enterprise integration, and reliable telephony deployment.

Strategic Decisions: ElevenLabs decoupled R&D from delivery via autonomous squads operating on rapid shipping cycles. They converted early model lead into structural moats: a two-sided voice marketplace, native infrastructure independence, and a hard pivot from pure creator tools toward high-margin enterprise agent workflows.

Risks & Warnings: Rapid voice commoditization poses severe threats. Long-term survival hinges on second-year enterprise net retention and proving agents can govern end-to-end mission-critical systems, not just conversational support.

Key Views & Dialogues

No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski

  • 🗓️ Date2025-12-11 | 🎙️ Show:No Priors

ElevenLabs has reached $300 million ARR with 350 employees and more than five million monthly active creative users, while enterprise agents approach half of revenue and expand voice AI beyond dubbing into support, discovery, checkout, education, and government services. Research supplies a six-to-12-month head start, but durable value depends on distribution, voices, integrations, and workflows as base models commoditize; Scribe v2’s reported under-150-millisecond latency and 93.5% accuracy point toward near-term real-time translation and more capable agents.

View Dialogue Notes & Key Takeaways
  • ElevenLabs reached $300 million ARR with 350 employees by pairing a five-million-MAU creator product with an enterprise agents business approaching half of revenue. Self-serve subscriptions and creators remain roughly 50% of the mix; several thousand enterprise customers range “from Fortune 500s to some of the fastest-growing AI startups.” The operating signal is breadth without abandoning the original research advantage.

  • The founding wedge was Poland’s broken dubbing experience, where one flat narrator voices every character, but the thesis expanded into voice as computing’s natural interface. ElevenLabs aimed to carry the “original voice, original emotions, original intonation” across languages, as demonstrated with Lex’s interview with Narendra Modi. Mati’s larger claim is that keyboard-and-screen interaction “feels broken,” making the opportunity far larger than conventional dubbing spend.

  • Mati’s operating model is research first, product in parallel, with small cross-functional labs formed around concrete problems. A roughly five-person voice lab solved human-sounding narration before ElevenLabs expanded into audiobooks and dubbing; an agent lab then combined speech-to-text, LLMs, text-to-speech, integrations, testing, and monitoring. Product teams see the research roadmap and generally build alternatives only when the relevant breakthrough looks more than three months away.

  • Customer support is the first major agent category, but voice is moving from reactive ticket resolution into discovery, checkout, education, media, and government services. Meesho’s agent progressed from refunds and tracking to product guidance and potentially checkout; MasterClass lets learners practice a negotiation with Chris Voss; Ukraine is pursuing what Mati calls a “first agentic government.” That expands the economic case from labor savings toward revenue generation and entirely new experiences.

  • Voice quality is not one scalar benchmark because preference can flip with voice identity, language, audience, and delivery even when underlying model quality differs. One customer requested a voice “as robotic as possible,” while a Japan-and-Korea deployment wanted an excitable voice for younger callers and a calm, slow one for older users. ElevenLabs therefore supplies voice coaching and expects selection eventually to become dynamic for each person and context.

  • ElevenLabs assumes base models will commoditize and treats research leadership as only a six-to-12-month advantage. Mati says “research is a head start”; durable value must accumulate through distribution, voices, integrations, workflows, and a product layer that connects models to business logic. Open-source narration is already strong, but controllability and real-time orchestration remain differentiated, while a concentrated pool of perhaps 50–100 elite audio researchers—Mati says roughly 10 are probably at ElevenLabs—can still produce architectural leaps.

  • The near-term roadmap targets controllable multimodal creation and lower-latency, emotionally aware agents while retaining cascaded systems for reliable enterprise work. Scribe v2 is reported at under 150 milliseconds and 93.5% accuracy across the top 30 languages on FLEURS; fused speech-to-speech may become more expressive but can introduce hallucinations and offers less visibility. Sarah estimates that humanlike conversational interaction is still at least a year away but may arrive within a year, with real-time dubbing or translation within two; personalized education—“your own teacher on demand”—is the largest unrealized application.

  • 🔗 Original source & video: No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski

Listen to full conversation →


ElevenLabs CEO: Why Voice is the Next AI Interface

  • 🗓️ Date2025-11-04 | 🎙️ Show:The a16z Show

ElevenLabs combines a research foundation with roughly 20 autonomous product teams, turning voice breadth into a marketplace with nearly 10,000 voices and $10 million returned to contributors. Enterprise demand has pushed the company toward sales, orchestration, telephony, security, and compliance, while a three-month research rule and a ban on foundation-model customers show how shipping speed and strategic discipline are evolving.

View Dialogue Notes & Key Takeaways
  • ElevenLabs keeps research ambitious without letting it become a product bottleneck. After resisting a simple speed control for nine months because the company hoped research would make voices infer pacing automatically, it adopted a three-month rule: beyond that horizon, product teams may bridge the gap however they choose. “We don’t want to become the same as the previous generation of the editing suite” was the rationale for resisting the slider, not an absolute constraint.

  • Shipping velocity comes from roughly 20 autonomous product teams of five to 10 people, backed by a research foundation. That structure tolerates duplicated work and uneven speeds in exchange for unusually high ownership and the ability to pursue creative tools and conversational agents simultaneously. Each new team gets six months to prove itself; the infrastructure team grew from three people at the partnership to 11.

  • Its distributed model is a global-talent thesis, not merely a remote-work policy. ElevenLabs hired globally and unconventionally—including an open-source text-to-speech developer who was working in a call center—then added hubs once headcount exceeded 30. Titles were removed, tenure does not determine hierarchy, and information access is deliberately limited when transparency becomes distraction.

  • The voice marketplace converts model breadth into an ecosystem with measurable creator economics. Nearly 10,000 voices are available and $10 million has been paid back to contributors; one deep Spanish voice found little demand in Spain but became a top-three voice after being offered in other languages and taking off in an English-speaking market because of its deepness. Staniszewski’s preferred posture is to help industry participants “disrupt together rather than just disrupt.”

  • Enterprise expansion required abandoning the conceit that engineers could simply perform sales. The customer-facing mix is now roughly 80% sales and 20% engineering, while the product has expanded from text-to-speech into speech-to-text and orchestration, and enterprise deployments require telephony, evaluation, monitoring, security, and compliance. The production ambition is eventually “four nines or five nines,” though Staniszewski concedes that reliability is difficult in AI.

  • At 350 employees, incentives have become part of product and competitive strategy. Staniszewski calls quotas and commissions “a lagging indicator of strategy”: the company may still grant commission while killing a strategically wrong deal, including a recent request from a foundation-model competitor to license ElevenLabs models for demos. The resulting policy explicitly prohibits sales to foundation-model companies.

  • 🔗 Original source & video: ElevenLabs CEO: Why Voice is the Next AI Interface

Listen to full conversation →


ElevenLabs CEO/Co-Founder, Mati Staniszewski:The Untold Story of Europe’s Fastest Growing AI Startup

  • 🗓️ Date2025-09-08 | 🎙️ Show:20VC

ElevenLabs says it has crossed $200 million in revenue with roughly 250 employees, with enterprise now the majority and the largest contract around $2 million. Its six-to-12-month research lead is being converted into a product and distribution moat spanning voice agents, knowledge bases, enterprise integrations, testing and monitoring. Owned infrastructure breaks even over roughly two years and supports faster experimentation, while commoditization, hardware advances and growth toward 400 employees remain key risks to monitor.

View Dialogue Notes & Key Takeaways
  • ElevenLabs says it has crossed $200 million in revenue with roughly 250 employees, after ending 2023 near $35 million. Mati described about 20 months to $100 million and initially “around 10 months” to $200 million, then corrected that second leg to “a bit longer, 15 months maybe.” Enterprise is now the majority of revenue, the largest contract is around $2 million, and Mati cautioned that growth “can also go quickly down.”

  • The moat is not a permanent model lead but the product and distribution built during a six-to-12-month research advantage. ElevenLabs concentrates scarce voice talent—Mati estimates only 50-100 researchers operate at the highest level—and turns models rapidly into production workflows. His answer to “Why won’t OpenAI do this?” is that it will do something, but ElevenLabs relies on its exceptional team, focus, execution speed and product layer.

  • Voice agents could become a multi-billion-dollar business for ElevenLabs “if we play it right.” The company is moving from speech components into knowledge bases, functions, testing, monitoring and enterprise integrations such as Salesforce, ServiceNow and SIP trunking, with email and WhatsApp potentially extending the platform into omnichannel support. Mati expects routine scheduling and refunds to automate first while high-stakes, domain-heavy work remains human.

  • The company’s financing story moved from speaking with 30-50 pre-seed investors to US investors competing through speed and operational help. ElevenLabs raised $2 million at a $9 million post-money valuation in 2022, then $19 million in 2023 after its launch broke out; a later round was priced at $3.3 billion. Mati says American firms are “playing a different game”: they ask how to bet bigger, test whether they can help before investing and, in his reference checks, generally showed stronger evidence of supporting founders through failure than some European investors.

  • Owning training infrastructure is an economic and strategic bet, not a vanity project. ElevenLabs calculated that continuously training models and moving large datasets through rented infrastructure would make an owned data center break even over roughly two years; ownership now enables faster experimentation and greater control. Mati concedes that hardware innovation could invalidate the equation, while arguing that a new model should initially optimize for “magic” before cost.

  • Mati’s European-company thesis is to build from Europe without building only for Europe. He calls Europe “hard mode,” rejects the idea that Europeans inherently work less hard, and says its ambitious talent is underused. ElevenLabs develops local leaders, supplements them with experienced advisers and operates roughly 20 teams. The tension is that “small and mighty” now means growing from 250 toward 400 people while opening small teams in Brazil, Japan, India and Mexico.

  • Liquidity is being used to support employee risk-taking and a longer-duration bet. ElevenLabs has received acquisition approaches but became a “flat no” after learning from the first processes; nearly every financing includes secondary liquidity or a tender for vested employee shares. In the host’s private-market choice among OpenAI at $300 billion, Anthropic at $170 billion and xAI at $120 billion, Mati would buy OpenAI but declined to name a sale, while saying Anthropic is especially compelling in coding and Google is “definitely in the race,” though he is not “super bullish” on it.

  • 🔗 Original source & video: ElevenLabs CEO/Co-Founder, Mati Staniszewski:The Untold Story of Europe’s Fastest Growing AI Startup

Listen to full conversation →