Skip to content
Nikhil KamathNikhil Kamath

The $11B Bet That Voice Will Replace Everything | Mati Staniszewski x Nikhil Kamath | WTF Online

ElevenLabs is the AI company that makes machines sound human. They dubbed the Lex Friedman–Modi podcast into Hindi. They built an AI Gordon Ramsay that teaches you how to cook while you cook. They automated 60,000 customer support calls for Meesho. The company is worth $11 billion and is racing OpenAI to own the future of how humans talk to machines. The CEO is a 29-year-old from Poland named Mati Staniszewski who started it because every foreign film in Poland had one man dubbing every character. I sat down with him in Davos. We got into why headphones will matter more than phones, why the real opportunity for young entrepreneurs isn't building AI models but going deep into one boring domain — automotive, healthcare, e-commerce — and deploying voice agents better than anyone else. Then it took a turn. We ended up talking about why social media is fundamentally broken, why no foreign algorithm should define the mood of India's youth, and why nobody has built an AI-native social product yet. We decided to try. 00:00 Introduction 06:35 Voice as the next tech interface 13:24 Competing with OpenAI and big labs 20:12 Preserving emotion in dubbed content 27:39 Building profitable voice businesses today 35:09 AI valuations and global opportunity 42:29 Geopolitics reshaping trust in tech platforms 49:51 Designing a new social media platform 56:40 Incentivising authenticity over negativity #nikhilkamath Co-founder of Zerodha and Gruhas Host of 'WTF is' & 'People By WTF' Podcast Twitter: https://x.com/nikhilkamathcio/ Instagram: https://www.instagram.com/nikhilkamathcio/ LinkedIn: https://www.linkedin.com/in/nikhilkamathcio?utm_source=share&utm_campaign=share_via&utm_content=profile&utm_medium=ios_app Facebook: https://www.facebook.com/nikhilkamathcio/ #matistaniszewski LinkedIN- https://uk.linkedin.com/in/matiii Twitter - https://x.com/matiii Watch 'WTF is' Podcast on Spotify https://tinyurl.com/4nsm4ezn Watch 'People by WTF' Podcast on Spotify https://tinyurl.com/yme92c59 Watch 'WTF Online' on Spotify https://tinyurl.com/4tjua4th #WTFiswithnikhilkamath #PeopleByWTF #WTFOnline

Nikhil KamathhostMati Staniszewskiguest
Mar 11, 202659mWatch on YouTube ↗

CHAPTERS

  1. 0:07 – 4:19

    AI-native hardware, Nothing’s bet, and real-time translation earbuds

    The conversation opens with Mati and Nikhil discussing hardware as a difficult but important frontier for AI—using Nothing (Carl Pei) as a case study. They explore how voice-first devices like earbuds could become the always-on interface, especially if real-time translation becomes seamless.

    • Nothing’s design-led approach and why hardware innovation is uniquely hard
    • Earbuds/headphones as a plausible ‘AI-native’ device form factor
    • Real-time translation as a killer use case for voice interfaces
    • Why device adoption (not model capability) may be the main bottleneck
  2. 4:19 – 5:56

    Voice agents for events: the “pendant” idea for frictionless capture and follow-ups

    Mati describes how his team uses a voice agent to capture feedback and action items after events, aiming to shorten administrative loops. They imagine an opt-in wearable (pendant) that transparently records, transcribes, and summarizes conversations to produce end-of-day follow-ups.

    • Event and meeting workflows as immediate, practical voice-agent use cases
    • Opt-in recording as a requirement for safety and trust
    • Hardware constraints: size, battery, usability, and social acceptance
    • Partnerships with stealth startups exploring wearable capture
  3. 5:56 – 11:21

    Why voice becomes the next interface: the 3 prerequisites (quality, knowledge, form factor)

    Nikhil asks why voice hasn’t fully “clicked” yet despite broad belief it’s next. Mati lays out three conditions needed for voice to replace current interfaces: human-level conversational quality, robust knowledge/memory/integrations, and the right deployment form factor.

    • Voice must feel human: interruptible, emotional, fast, and intelligent
    • Agents need knowledge access: integrations, memory, CRM and system context
    • Form factor remains unsolved: phone vs glasses vs headphones vs future neural interfaces
    • Mati’s personal bet: behind-the-ear audio devices and silent speech detection
  4. 11:21 – 13:23

    What Sam Altman & Jony Ive might build—and why hardware becomes strategic for AI

    They speculate about what a new AI device could look like and why adoption may start with familiar hand-held behavior. The discussion returns to the strategic importance of bundling AI-first software with hardware distribution over time.

    • Likely early scaling path: a phone-like device with a new UI plus a ‘buddy’ device
    • Why AI companies may need hardware to preinstall the assistant experience
    • OpenAI’s personal assistant trajectory and increasing voice focus
    • How hardware positioning could shape platform power (and acquisition interest)
  5. 13:23 – 16:09

    Competing with OpenAI and big labs: research + product as the core defense

    Nikhil raises the classic platform risk: big model providers moving up the stack and copying applications. Mati explains ElevenLabs’ approach—owning foundational audio research while also shipping products—plus additional moats via customer value delivery.

    • Why ElevenLabs doesn’t depend on OpenAI for voice models
    • Staying ahead via multilingual model quality (India + Europe languages)
    • Defense beyond research: product value, workflows, and customer outcomes
    • Two core offerings: Creative platform and Agents platform
  6. 16:09 – 18:42

    ElevenLabs explained: creator workflows, enterprise agents, and real examples

    Mati breaks down ElevenLabs in plain terms: foundational speech generation/understanding plus products for creators and enterprises. He gives concrete creator workflows (patching audio, dubbing) and references high-profile localization work.

    • Simple definition: audio AI models + products for creators and agents
    • Creator use case: post-processing and recreating a speaker’s voice for edits
    • Localization at scale: dubbing podcasts into multiple languages
    • Example: dubbing Lex Fridman–PM Modi episode with human-in-the-loop review
  7. 18:42 – 22:16

    Preserving emotion in dubbing—and the next layer: lip reanimation

    Nikhil challenges the biggest weakness of dubbing: emotion often sounds robotic and doesn’t match facial cues. Mati explains how their system preserves context and intonation, the limitations caused by different sentence structures, and the likely future of automated lip animation.

    • Emotion preservation: capturing intonation and context across turns
    • Hard cases: laughter/screams and reordering due to language grammar (e.g., German)
    • Present vs future: audio-first fixes today, lip reanimation soon
    • Opportunity for startups focused purely on lip-sync and facial alignment
  8. 22:16 – 27:11

    Voice agents as products: MasterClass-style interactive experts and AI SDR augmentation

    The discussion shifts to agent experiences where users converse with an AI persona to learn or get guided help. Mati shares examples like MasterClass’ AI instructors and explains why agents often work best as augmentation (e.g., better lead qualification) rather than full replacement.

    • Interactive learning agents: AI Gordon Ramsay, AI Chris Voss negotiation practice
    • Concept: an AI ‘Nikhil’ that teaches from existing content and knowledge
    • Deployment patterns: websites and increasingly phone-number call-in agents
    • Business agents: AI SDR as a funnel accelerant, improving lead detail and readiness
  9. 27:11 – 35:02

    Building profitable voice businesses “tomorrow”: go deep in boring domains

    Nikhil asks for advice aimed at young entrepreneurs looking for near-term opportunities. Mati recommends pairing voice tech with domain expertise—especially in traditional, slow-to-innovate sectors—then building repeatable deployments across similar customers.

    • Best entry point: domain expertise + voice agent execution
    • Targets: automotive, healthcare, financial services, e-commerce
    • Start with easy-to-deploy ROI: customer support and customer experience
    • Example: Meesho voice support at scale and the future ‘AI concierge’ in shopping
  10. 35:02 – 38:09

    AI valuations and global opportunity beyond Silicon Valley

    Nikhil asks whether AI valuations look bubble-like given revenue vs valuation mismatches. Mati argues many companies show real value and revenue, flags skepticism in certain infrastructure ‘middlemen’ plays, and emphasizes global innovation from Europe and Asia.

    • Why this cycle differs: clearer value capture and real revenue in many AI companies
    • Skeptical pocket: some GPU/inference providers built on top of Nvidia
    • High risk–high reward logic for research-first companies
    • Hope for global parity: great companies outside the Valley should command similar valuations
  11. 38:09 – 46:39

    Geopolitics, data residency, and trust: will platforms fragment by country?

    Nikhil predicts a more multipolar world where countries build local versions of major platforms and keep data within borders. Mati agrees on data residency and local fine-tuning needs, but argues foundational models and network effects still favor global infrastructure—if trust can be maintained.

    • Thesis: platform fragmentation as countries seek sovereignty and control
    • Counterpoint: foundational models likely concentrated due to resource requirements
    • Local adaptation via fine-tuning, residency, and product variants
    • Trust as the key variable: geopolitical tension can break perceived neutrality
  12. 46:39 – 59:39

    Designing a new social media platform—and using voice to incentivize authenticity

    Nikhil outlines why social media feels broken: brand-heavy feeds, engagement incentives toward outrage, and foreign-controlled algorithms shaping youth culture. Mati proposes voice-native interaction patterns (summarized feeds via assistants, voice comments, instant translation) and they debate verification and incentive design to reward authenticity/curiosity over knee-jerk reactions.

    • Problems: declining organic sharing, algorithmic outrage loops, and perceived external control
    • Concept: assistant-driven social feed that summarizes and enables voice-first engagement
    • Voice features: audio posts/comments, listening to written posts in creator voice, auto-translation
    • Incentive design: verification/real-human identity, discourage reaction-maximizing ranking, promote curiosity/neutrality

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.