No PriorsNo Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski
CHAPTERS
- 0:00 – 1:48
ElevenLabs’ mission and product surface area (creative + agents)
Sarah introduces Mati Staniszewski and frames ElevenLabs as a voice-first company aimed at changing human–computer interaction. Mati explains ElevenLabs’ stack: foundational audio models plus products for creators (narration, dubbing, voiceover) and for enterprises (agent experiences).
- •Mission: make interacting with technology more natural through voice
- •Core capabilities: generate speech, understand speech, orchestrate components for interactivity
- •Creative platform: audiobooks, voiceovers, ads/movies, multilingual dubbing
- •Agents platform: customer experience, personal AI, education, immersive media
- 1:48 – 2:46
Hypergrowth snapshot: team, ARR mix, and customer base
Mati quantifies the company’s scale: headcount, global hubs, ARR, and the split between self-serve creators and enterprise. He also describes the breadth of their user base, from millions of monthly active creators to thousands of enterprise customers.
- •~350 employees; remote-first with major hubs (London, New York, Warsaw, SF, Tokyo, Brazil)
- •~$300M ARR with ~50/50 self-serve vs enterprise mix
- •5M+ monthly active users on the creative side
- •Thousands of enterprise customers incl. Fortune 500 and fast-growing AI startups
- 2:46 – 6:52
Why voice is bigger than ‘dubbing’: origin story and the globalization thesis
Sarah probes why ElevenLabs believed voice creation would be a broad market. Mati traces the founding insight to Poland’s ‘single narrator’ dubbing experience and a conviction that high-quality, emotionally faithful voice translation would unlock global content.
- •Polish film dubbing as a ‘broken’ user experience that inspired the company
- •Goal: preserve original speaker identity, emotion, and intonation across languages
- •Example: translating a podcast while keeping the same speakers’ voices
- •Long-term thesis: voice becomes a primary interface across devices and contexts
- 6:52 – 10:11
Sequencing research and product: ‘labs’ organized around problems
Sarah asks how ElevenLabs balances deep research with product execution across multiple markets. Mati explains their approach: start with a concrete problem, build a cross-functional ‘lab,’ ship a simple product layer, then expand into fuller workflows and new labs.
- •Early attempt to adapt existing models failed due to robotic output quality
- •Problem-first org design: small cross-functional labs (research + eng + ops)
- •Voice lab first (recreate voice convincingly), then product layer, then full workflows
- •Second lab: agents—combine STT, LLMs, TTS with low latency and production tooling
- 10:11 – 12:36
Market pull vs vision push: adding music and multimodal creation
Sarah challenges whether customers knew they wanted these capabilities. Mati describes a mix of proactive bets (like dubbing) and reactive expansion driven by customer demand, including launching a music lab and integrating with partner image/video models.
- •Some initiatives are ‘future inevitable’ bets; others are direct customer requests
- •Music lab: fully licensed model driven by users wanting speech + music creation
- •Creative suite expansion: combine speech, music, sound effects; integrate partner image/video
- •Real-time dubbing as a broader ‘Babel Fish’ vision beyond media translation
- 12:36 – 17:03
Voice quality is hard to measure: selection, personalization, and missing benchmarks
Sarah digs into evaluation: how non-ML customers pick ‘good’ voices and how ElevenLabs handles benchmarking. Mati explains enterprise guidance via ‘voice sommelier’ roles, dynamic personalization of voices, and why audio benchmarks lag because preference is highly voice-dependent.
- •Enterprise support includes voice coaching/consultation to match brand and audience
- •Celebrity marketplace expands options for iconic talent voices
- •Customers may want different voices by demographic or context; personalization can be dynamic
- •Benchmarking challenge: switching the voice can dominate perceived quality vs model differences
- 17:03 – 17:54
Describing audio data is still immature: labeling emotion, accent, and delivery
Mati adds a deeper data problem: the industry struggles to label ‘how’ something is said, not just ‘what’ is said. ElevenLabs had to build internal capabilities for richer qualitative annotation (emotion, accent, delivery) to improve models and controllability.
- •Audio needs labels beyond transcription: emotion, accent, delivery style
- •Traditional vendors often lack the skill to label nuanced vocal attributes
- •Internal efforts to interpret and describe qualitative audio characteristics
- •Better labeling supports controllable, expressive generation and improved user experience
- 17:54 – 19:34
Agent platform traction: from reactive support to proactive commerce assistants
Sarah pivots to agent use cases and what’s surprisingly working. Mati highlights customer support as the fast mover, then describes a shift toward proactive, end-to-end shopping and discovery experiences that guide users through selection and checkout.
- •Customer support is the biggest early enterprise driver (partners like Cisco/Twilio/TELUS Digital)
- •Trend: reactive support → proactive, front-of-experience assistants
- •Meesho example: voice widget for discovery, navigation, recommendations, and checkout
- •Square example: voice ordering expanding into broader discovery experiences
- 19:34 – 20:35
Immersive media: interactive IP and real-time character experiences
Mati describes how voice agents enable interactive storytelling and gaming experiences. He cites work with Epic Games that let players talk to Darth Vader inside Fortnite, pointing to a broader shift from static media to interactive, conversational IP.
- •Shift from static media to immersive, interactive content
- •Epic Games/Fortnite: real-time interaction with Darth Vader at massive scale
- •Future pattern: talk to characters from books, films, and other IP
- •Voice makes immersion feel natural and increases engagement
- 20:35 – 21:46
Education as the breakout category: tutors, coaches, and practice-by-calling
Mati argues education will be one of the most transformative uses of voice agents. He shares examples like chess coaching with famous players and interactive MasterClass scenarios where users can call and practice skills with an expert persona.
- •Personal tutor ‘in your headphones’ enables always-available, adaptive learning
- •chess.com example: learn from voices/personas like Hikaru or Magnus
- •MasterClass example: interactive practice with Chris Voss negotiation scenarios
- •Education becomes more experiential: practice, feedback, and iteration via conversation
- 21:46 – 23:23
‘Agentic government’: Ukraine case study for citizen services + proactive outreach
Mati shares a real deployment story: Ukraine’s Ministry of Digital Transformation pursuing an ‘agentic government.’ He describes consolidating citizen support, proactive citizen communication, and education-style services, enabled by strong engineering leadership embedded in ministries.
- •Vision: redesign government services around agents and digital interfaces
- •Use cases: benefits/employment inquiries, travel processes, citizen support via app
- •Proactive outreach: informing citizens about relevant events or actions
- •Operating model: engineering leaders in each ministry coordinating with a central transformation team
- 23:23 – 26:43
Choosing a technology partner: platform vs point solution vs consulting
Sarah asks how enterprises should decide between platforms like ElevenLabs, foundation model vendors, point solutions, or big consultancies. Mati explains when ElevenLabs fits: organizations wanting an open platform across multiple conversational experiences, strong international coverage, and optional forward-deployed engineering help.
- •If you want a single narrow solution, a point vendor may be better
- •ElevenLabs fit: multiple conversational use cases (support, training, sales) on one platform
- •Open platform: use components selectively, integrate with existing enterprise systems
- •Differentiator: multilingual/international breadth and production deployment support
- 26:43 – 29:58
Against the ‘labs will do it’ narrative: why specialization wins in audio
Sarah raises the classic objection: why won’t Google/OpenAI just solve voice? Mati argues top labs optimize for many things and won’t prioritize the specialized product layer and audio-specific research breakthroughs needed for seamless, controllable voice experiences.
- •Delivering value requires strong product: integrations, deployment, monitoring, workflows
- •ElevenLabs differentiates by focusing fully on audio quality + seamlessness + controllability
- •Audio progress depends more on architectural breakthroughs than sheer scale
- •Talent concentration matters: a small number of top audio researchers can drive outsized gains
- 29:58 – 34:25
Open source, commoditization, and defensibility: research head start → product → ecosystem
Sarah and Mati discuss how advantages erode and what persists. Mati expects base models to commoditize over time, making product, fine-tuning layers, workflows, integrations, brand/distribution, and ecosystem the durable moats; research provides a time-bound head start.
- •Expectation: base models become broadly available; differentiation compresses
- •Narration quality is converging; key gap is controllability
- •Interactive orchestration remains harder: low latency, reliability, ‘Turing test’ conversation
- •Defensibility sequence: research head start (6–12 months) + product shipped in parallel + ecosystem (voices/integrations/workflows)
- 34:25 – 36:52
Current R&D roadmap: latency, orchestration, emotion, and multimodal audio
Sarah asks what research they’re doing now and how they decide what to ship. Mati outlines two buckets—creative and agents—covering multilingual STT, controllable TTS, licensed music, multimodal combinations, and next-gen orchestration that adds emotional context; he also contrasts cascaded vs fused speech-to-speech approaches.
- •Creative roadmap: controllable TTS + high-accuracy multilingual STT (~100 languages) + licensed music
- •Multimodal direction: combine audio with visual/video workflows for better delivery
- •Agents roadmap: real-time STT/TTS and improved orchestration; emotion-aware responses
- •Cascaded systems suit enterprise reliability; fused speech-to-speech may win for expressiveness over time
- 36:52 – 41:38
Future: AI companions, Jarvis-like assistants, and how content interaction changes
In closing, Sarah asks about AI companions and broader interaction shifts. Mati predicts companions will be common but is more excited by ‘Jarvis’ super-assistants; he expects a blend of AI tutoring with deliberate human-only time, and sees education as the biggest near-future change in content consumption.
- •Companions will exist, but the bigger unlock is a proactive personal ‘super assistant’
- •Not everything must be personified; some device control can remain utilitarian
- •Robots and embodied agents will strongly drive voice interfaces
- •Education/tutoring via voice will reshape how people learn and interact with knowledge (even via famous teacher personas)