CHAPTERS
- 0:00 – 0:47
Cold open: Solving voice nuance at the research level + creator payouts at scale
A quick cold open frames ElevenLabs’ philosophy: avoid “editing-suite” knobs and instead make models infer natural delivery (pace, style) from context. Mati also previews the Voice Marketplace as a mechanism to unlock voice diversity while paying creators for usage.
- •Preference for model-level intelligence over UI sliders/toggles for speech control
- •Need for broad coverage: languages, accents, styles, and voice variety
- •Voice Marketplace concept: creators share voices and earn revenue
- •Scale metrics: ~10,000 voices and ~$10M paid back to the community
- •Overcoming knee-jerk skepticism by showing concrete tech demos
- 0:47 – 2:03
Welcome and product scope expansion: from TTS to agents and licensed music
Jennifer Li introduces Mati Staniszewski and opens with an ElevenLabs-generated audio example. Mati outlines how the company has expanded beyond voices into orchestration for voice agents and a fully licensed music model.
- •ElevenLabs-generated audio used as event/walk-on music
- •Product expansion trajectory: voices → agent orchestration → licensed music
- •Positioning as a broad audio company, not just TTS
- •Emphasis on licensing and commercial usability for generated media
- 2:03 – 2:48
How ElevenLabs ships fast: small, independent teams with high ownership
Mati explains the operating model behind rapid execution: many small product teams with independence and strong accountability. He also describes the top-level product “buckets” spanning creative workflows and conversational agents.
- •~20 product teams, typically 5–10 people each, operate with autonomy
- •High ownership increases delivery speed, despite some duplication risk
- •Two primary pillars: creative platform (narration, voiceover, dubbing) and agents (conversational experiences)
- •Organization designed to keep shipping velocity in a fast-moving AI market
- 2:48 – 4:30
Research roots and early breakthroughs: building contextual, emotional speech
Mati credits cofounder Piotr and the research team for foundational model advances—especially context-aware speech with emotion and accurate voice characteristics. He connects these early research wins to later expansions like STT and music.
- •Cofounder Piotr as research leader and talent magnet in voice AI
- •Goal: TTS that understands context → emotion, intonation, and delivery
- •Capturing voice characteristics: style, age, gender, dialect, etc.
- •Research foundation later extends to speech-to-text and music models
- 4:30 – 6:18
Balancing research vs product: when to ship workarounds vs wait for breakthroughs
The conversation turns to the tension between shipping product features now versus waiting for research to solve problems “the right way.” Mati shares a concrete speed-control example and a pragmatic internal rule of thumb.
- •Initial resistance to adding speed sliders to avoid legacy editing-suite UX
- •Attempted to solve pacing via research/model intelligence, but it took too long
- •Pragmatic threshold: if research will take > ~3 months, product can ship interim solutions
- •Quarterly alignment: near-term research initiatives vs longer-term bets
- 6:18 – 8:10
Remote-first to hub model: meeting talent where they are
Mati describes why ElevenLabs began as a remote-first company and later added hubs as the team scaled. The goal is to access global research/engineering talent while still offering in-person immersion for newer hires.
- •Company started between Warsaw and London; founders are Polish
- •Remote-first enabled hiring top talent globally (Europe, Asia, etc.)
- •Hubs emerged after ~30 people to help onboarding and cultural immersion
- •Hybrid approach: early-career hires encouraged in hubs; experienced remote workers supported
- 8:10 – 10:01
Hiring beyond résumés: unconventional signals and standout global recruits
Mati explains how the team deliberately avoided traditional “LinkedIn-first” hiring and sought real capability signals, like open-source work. He shares a story of hiring an outstanding TTS contributor who was simultaneously working in a call center.
- •De-emphasizing traditional pedigree/background checks as primary filter
- •Valuing open-source contributions and demonstrated technical output
- •Example hire: call-center worker with an impressive open-source TTS model
- •Blending non-traditional hires with experienced “traditional” operators, including in sales
- 10:01 – 10:34
US vs Europe work culture: finding the highly driven pockets
Mati contrasts social/work norms in the US and Europe, noting that Europe can have fewer “work-centric” social interactions. He argues, however, that motivated talent exists—often lacking the right local company environment—making it powerful once concentrated.
- •US culture: more open enthusiasm for talking about work socially
- •Europe: often less work-centric social norm, but strong motivated sub-communities exist
- •ElevenLabs attracts and concentrates those driven pockets in Europe
- •Claim: European team includes some of the most passionate/motivated contributors
- 10:34 – 13:28
Flat org and removing titles: impact, trade-offs, and focus management
Mati outlines the rationale for removing titles and minimizing hierarchy to emphasize impact over tenure. He also covers the real operational challenges: leads must manage cross-team complexity, and transparency can create distraction if not controlled.
- •Titles removed to reinforce impact-based contribution over hierarchy
- •Small teams get ~6 months to prove themselves; rapid growth based on performance
- •A thin layer of leads coordinates across research, creative, agents, and GTM
- •Trade-off: too much Slack transparency can distract; sometimes access must be limited to protect focus
- 13:28 – 14:56
Creative industries and AI: from resistance to adoption through partnership
Jennifer and Mati discuss early creative-industry skepticism and how it shifted toward adoption. Mati emphasizes spending time with creators to understand incentives and identifying where AI genuinely helps versus where humans should stay central.
- •Early resistance in creative fields has softened into broader adoption
- •Approach: deep industry engagement to learn priorities and workflows
- •Learning from high-profile collaborators about where AI fits in production
- •Thesis: disrupt with the industry rather than doing disruption “to” the industry
- 14:56 – 16:03
Voice Marketplace mechanics: diversity of voices + monetization for creators
Mati explains how the marketplace addresses two needs: broad voice coverage for product quality and a fair economic model for voice actors/creators. He shares usage and payout numbers plus a story showing how voices can find unexpected audiences across languages.
- •Marketplace solves for breadth: voices, accents, languages, and styles at scale
- •Creators can publish a voice and earn money when it’s used
- •Scale: ~10,000 voices; ~$10M paid back to the community
- •Cross-language portability: the same voice can perform in dozens of languages
- •Unexpected demand story: a Spanish voice becomes top-performing in English contexts
- 16:03 – 17:56
Licensed music model: 18-month label negotiations and “forcing functions”
Mati describes the long process of securing agreements with music labels (e.g., Merlin, Kobalt and majors) to train/deploy a licensed music model with commercial rights. He highlights negotiation tactics—especially deadlines/forcing functions—and the importance of education to reduce fear.
- •Goal: fully licensed music generation with commercial rights and user protection
- •Negotiations took ~18 months to reach workable agreements
- •Forcing functions (deadlines/triggers) created urgency and enabled progress
- •Finding compromise required aligning label priorities with new tech capabilities
- •Education and demos helped move stakeholders past “AI is bad” knee-jerk reactions
- 17:56 – 19:01
Hiring for complex, high-risk domains: legal strategy and advisory scaffolding
The discussion moves to hiring in domains the team doesn’t yet understand—especially legal and licensing. Mati explains a blended approach: bring in a small number of experienced operators, then supplement with specialized consultants who already know the ecosystem players.
- •For new domains, hire 1–2 experienced insiders who’ve worked with key counterparties
- •Use consultants (e.g., music lawyers) as a bridge to shared language and norms
- •Consultants help navigate stakeholder networks and negotiation patterns
- •Operating principle: build internal ownership, but scaffold with external expertise
- 19:01 – 20:45
Risk-tolerant counsel: avoiding “risk lists” and seeking decision-grade guidance
Mati details painful lessons hiring early legal staff: some profiles over-index on enumerating risks without offering practical lines to draw. The ideal counsel provides both risk identification and actionable recommendations informed by startup and industry precedent.
- •Early legal hires were mismatched; separation was necessary
- •Fortune 500-only experience led to overly conservative, non-actionable guidance
- •Need counsel who can say: here’s the risk, here are options, here’s the line
- •Best counsel acts as a thought partner using real-world precedents from similar companies
- 20:45 – 25:11
From creator-first PLG to enterprise: orchestration, reliability, and long sales cycles
Mati explains how enterprise pull emerged early even though the team initially resisted building a sales function. Enterprise use cases pushed ElevenLabs beyond single-model APIs into full orchestration (STT + LLM + TTS), integrations, deployment, evaluation, and compliance—plus cultural adaptation to longer cycles.
- •Early enterprise inbound conflicted with initial “no salespeople” instinct
- •Learning: pure engineer-led selling didn’t work; hybrid model emerged (mostly sales, some engineering)
- •Enterprise agents require orchestration: STT + LLM + TTS + integrations (telephony, etc.)
- •Operational needs: testing, versioning, evaluation/monitoring, fine-tuning, security/compliance
- •Cultural challenge: aligning teams to tolerate 6–12 month sales cycles
- 25:11 – 27:44
Shipping for enterprise without slowing down: alpha labeling and PMF vs non-PMF teams
Mati outlines how ElevenLabs maintains speed while serving enterprise demands by clearly labeling alpha products and letting customers opt in. Internally, they separate teams working on pre-PMF experiments from post-PMF products that require more rigor and long-term support.
- •Clear external delineation: alpha vs stable releases; customers choose early access
- •Opt-in experimentation enables innovation without compromising expectations
- •Internal separation: pre-PMF teams ship fast; post-PMF teams optimize reliability
- •6-month “prove it” window for experimental teams; products can be killed if they don’t land
- 27:44 – 31:29
Scaling phase realities: incentives, commissions, and turning down a competitor deal
In the final segment, Mati reflects on the shift from passion-driven behavior to incentive-driven behavior as the company scales. He shares how commissions can unintentionally steer strategy—and a concrete case where they refused to license models to a competitor, updated policy, and still honored commissions to reinforce values.
- •As GTM scales, incentive structures strongly shape day-to-day decisions
- •Commissions/quotas can lag strategy unless carefully aligned
- •New norm: explicitly escalate “strategically wrong” deals even if they pay well
- •Example: refused licensing to a foundational-model competitor; later codified policy
- •Cultural takeaway: clarity and explicit rules preserve strategy during scaling
