Skip to content
No PriorsNo Priors

No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski

Imagine learning chess from a grand master, or negotiating tactics from an expert FBI hostage negotiator. ElevenLabs’ voice AI technology is making that unlock possible. Sarah Guo sits down with Mati Staniszewski, co-founder of ElevenLabs, to explore how the three-year old company is transforming how humans interact with technology through voice. Mati talks about the technical challenges of building foundational audio models, the strategic thinking between conducting research and deploying products in tandem, and why voice is the ultimate interface for everything from computers to robots to immersive media. They also discuss how the coming revolution of AI personal tutors will shift agentic AI from reactive to proactive support, break down language barriers globally, and even provide the framework for agentic government services. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @elevenlabsio |@matiii Chapters: 00:00 – Mati Staniszewski Introduction 00:46 – 11 Labs: Growth and Scale 02:46 – Voice Technology and Applications 06:52 – Research and Product Development 12:36 – Voice Quality and Customer Preferences 17:54 – Agent Platform and Use Cases 23:21 – Choosing the Right Technology Partner 26:43 – The Role of Foundation Models 29:58 – Open Source Models and Future Trends 32:37 – Research and Development Focus 36:53 – Future of AI Companions and Education 41:37 – Conclusion

Sarah GuohostMati StaniszewskiguestElad Gilhost
Dec 11, 202541mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:48

    ElevenLabs’ mission and product surface area (creative + agents)

    Sarah introduces Mati Staniszewski and frames ElevenLabs as a voice-first company aimed at changing human–computer interaction. Mati explains ElevenLabs’ stack: foundational audio models plus products for creators (narration, dubbing, voiceover) and for enterprises (agent experiences).

    • Mission: make interacting with technology more natural through voice
    • Core capabilities: generate speech, understand speech, orchestrate components for interactivity
    • Creative platform: audiobooks, voiceovers, ads/movies, multilingual dubbing
    • Agents platform: customer experience, personal AI, education, immersive media
  2. 1:48 – 2:46

    Hypergrowth snapshot: team, ARR mix, and customer base

    Mati quantifies the company’s scale: headcount, global hubs, ARR, and the split between self-serve creators and enterprise. He also describes the breadth of their user base, from millions of monthly active creators to thousands of enterprise customers.

    • ~350 employees; remote-first with major hubs (London, New York, Warsaw, SF, Tokyo, Brazil)
    • ~$300M ARR with ~50/50 self-serve vs enterprise mix
    • 5M+ monthly active users on the creative side
    • Thousands of enterprise customers incl. Fortune 500 and fast-growing AI startups
  3. 2:46 – 6:52

    Why voice is bigger than ‘dubbing’: origin story and the globalization thesis

    Sarah probes why ElevenLabs believed voice creation would be a broad market. Mati traces the founding insight to Poland’s ‘single narrator’ dubbing experience and a conviction that high-quality, emotionally faithful voice translation would unlock global content.

    • Polish film dubbing as a ‘broken’ user experience that inspired the company
    • Goal: preserve original speaker identity, emotion, and intonation across languages
    • Example: translating a podcast while keeping the same speakers’ voices
    • Long-term thesis: voice becomes a primary interface across devices and contexts
  4. 6:52 – 10:11

    Sequencing research and product: ‘labs’ organized around problems

    Sarah asks how ElevenLabs balances deep research with product execution across multiple markets. Mati explains their approach: start with a concrete problem, build a cross-functional ‘lab,’ ship a simple product layer, then expand into fuller workflows and new labs.

    • Early attempt to adapt existing models failed due to robotic output quality
    • Problem-first org design: small cross-functional labs (research + eng + ops)
    • Voice lab first (recreate voice convincingly), then product layer, then full workflows
    • Second lab: agents—combine STT, LLMs, TTS with low latency and production tooling
  5. 10:11 – 12:36

    Market pull vs vision push: adding music and multimodal creation

    Sarah challenges whether customers knew they wanted these capabilities. Mati describes a mix of proactive bets (like dubbing) and reactive expansion driven by customer demand, including launching a music lab and integrating with partner image/video models.

    • Some initiatives are ‘future inevitable’ bets; others are direct customer requests
    • Music lab: fully licensed model driven by users wanting speech + music creation
    • Creative suite expansion: combine speech, music, sound effects; integrate partner image/video
    • Real-time dubbing as a broader ‘Babel Fish’ vision beyond media translation
  6. 12:36 – 17:03

    Voice quality is hard to measure: selection, personalization, and missing benchmarks

    Sarah digs into evaluation: how non-ML customers pick ‘good’ voices and how ElevenLabs handles benchmarking. Mati explains enterprise guidance via ‘voice sommelier’ roles, dynamic personalization of voices, and why audio benchmarks lag because preference is highly voice-dependent.

    • Enterprise support includes voice coaching/consultation to match brand and audience
    • Celebrity marketplace expands options for iconic talent voices
    • Customers may want different voices by demographic or context; personalization can be dynamic
    • Benchmarking challenge: switching the voice can dominate perceived quality vs model differences
  7. 17:03 – 17:54

    Describing audio data is still immature: labeling emotion, accent, and delivery

    Mati adds a deeper data problem: the industry struggles to label ‘how’ something is said, not just ‘what’ is said. ElevenLabs had to build internal capabilities for richer qualitative annotation (emotion, accent, delivery) to improve models and controllability.

    • Audio needs labels beyond transcription: emotion, accent, delivery style
    • Traditional vendors often lack the skill to label nuanced vocal attributes
    • Internal efforts to interpret and describe qualitative audio characteristics
    • Better labeling supports controllable, expressive generation and improved user experience
  8. 17:54 – 19:34

    Agent platform traction: from reactive support to proactive commerce assistants

    Sarah pivots to agent use cases and what’s surprisingly working. Mati highlights customer support as the fast mover, then describes a shift toward proactive, end-to-end shopping and discovery experiences that guide users through selection and checkout.

    • Customer support is the biggest early enterprise driver (partners like Cisco/Twilio/TELUS Digital)
    • Trend: reactive support → proactive, front-of-experience assistants
    • Meesho example: voice widget for discovery, navigation, recommendations, and checkout
    • Square example: voice ordering expanding into broader discovery experiences
  9. 19:34 – 20:35

    Immersive media: interactive IP and real-time character experiences

    Mati describes how voice agents enable interactive storytelling and gaming experiences. He cites work with Epic Games that let players talk to Darth Vader inside Fortnite, pointing to a broader shift from static media to interactive, conversational IP.

    • Shift from static media to immersive, interactive content
    • Epic Games/Fortnite: real-time interaction with Darth Vader at massive scale
    • Future pattern: talk to characters from books, films, and other IP
    • Voice makes immersion feel natural and increases engagement
  10. 20:35 – 21:46

    Education as the breakout category: tutors, coaches, and practice-by-calling

    Mati argues education will be one of the most transformative uses of voice agents. He shares examples like chess coaching with famous players and interactive MasterClass scenarios where users can call and practice skills with an expert persona.

    • Personal tutor ‘in your headphones’ enables always-available, adaptive learning
    • chess.com example: learn from voices/personas like Hikaru or Magnus
    • MasterClass example: interactive practice with Chris Voss negotiation scenarios
    • Education becomes more experiential: practice, feedback, and iteration via conversation
  11. 21:46 – 23:23

    ‘Agentic government’: Ukraine case study for citizen services + proactive outreach

    Mati shares a real deployment story: Ukraine’s Ministry of Digital Transformation pursuing an ‘agentic government.’ He describes consolidating citizen support, proactive citizen communication, and education-style services, enabled by strong engineering leadership embedded in ministries.

    • Vision: redesign government services around agents and digital interfaces
    • Use cases: benefits/employment inquiries, travel processes, citizen support via app
    • Proactive outreach: informing citizens about relevant events or actions
    • Operating model: engineering leaders in each ministry coordinating with a central transformation team
  12. 23:23 – 26:43

    Choosing a technology partner: platform vs point solution vs consulting

    Sarah asks how enterprises should decide between platforms like ElevenLabs, foundation model vendors, point solutions, or big consultancies. Mati explains when ElevenLabs fits: organizations wanting an open platform across multiple conversational experiences, strong international coverage, and optional forward-deployed engineering help.

    • If you want a single narrow solution, a point vendor may be better
    • ElevenLabs fit: multiple conversational use cases (support, training, sales) on one platform
    • Open platform: use components selectively, integrate with existing enterprise systems
    • Differentiator: multilingual/international breadth and production deployment support
  13. 26:43 – 29:58

    Against the ‘labs will do it’ narrative: why specialization wins in audio

    Sarah raises the classic objection: why won’t Google/OpenAI just solve voice? Mati argues top labs optimize for many things and won’t prioritize the specialized product layer and audio-specific research breakthroughs needed for seamless, controllable voice experiences.

    • Delivering value requires strong product: integrations, deployment, monitoring, workflows
    • ElevenLabs differentiates by focusing fully on audio quality + seamlessness + controllability
    • Audio progress depends more on architectural breakthroughs than sheer scale
    • Talent concentration matters: a small number of top audio researchers can drive outsized gains
  14. 29:58 – 34:25

    Open source, commoditization, and defensibility: research head start → product → ecosystem

    Sarah and Mati discuss how advantages erode and what persists. Mati expects base models to commoditize over time, making product, fine-tuning layers, workflows, integrations, brand/distribution, and ecosystem the durable moats; research provides a time-bound head start.

    • Expectation: base models become broadly available; differentiation compresses
    • Narration quality is converging; key gap is controllability
    • Interactive orchestration remains harder: low latency, reliability, ‘Turing test’ conversation
    • Defensibility sequence: research head start (6–12 months) + product shipped in parallel + ecosystem (voices/integrations/workflows)
  15. 34:25 – 36:52

    Current R&D roadmap: latency, orchestration, emotion, and multimodal audio

    Sarah asks what research they’re doing now and how they decide what to ship. Mati outlines two buckets—creative and agents—covering multilingual STT, controllable TTS, licensed music, multimodal combinations, and next-gen orchestration that adds emotional context; he also contrasts cascaded vs fused speech-to-speech approaches.

    • Creative roadmap: controllable TTS + high-accuracy multilingual STT (~100 languages) + licensed music
    • Multimodal direction: combine audio with visual/video workflows for better delivery
    • Agents roadmap: real-time STT/TTS and improved orchestration; emotion-aware responses
    • Cascaded systems suit enterprise reliability; fused speech-to-speech may win for expressiveness over time
  16. 36:52 – 41:38

    Future: AI companions, Jarvis-like assistants, and how content interaction changes

    In closing, Sarah asks about AI companions and broader interaction shifts. Mati predicts companions will be common but is more excited by ‘Jarvis’ super-assistants; he expects a blend of AI tutoring with deliberate human-only time, and sees education as the biggest near-future change in content consumption.

    • Companions will exist, but the bigger unlock is a proactive personal ‘super assistant’
    • Not everything must be personified; some device control can remain utilitarian
    • Robots and embodied agents will strongly drive voice interfaces
    • Education/tutoring via voice will reshape how people learn and interact with knowledge (even via famous teacher personas)

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.