Skip to content
YC Root AccessYC Root Access

This Startup Built the Infrastructure Powering Voice AI

In this episode of Founder Firesides, YC Managing Partner Jared Friedman talks to Dylan Fox, the Founder of Assembly AI (S17), which has raised $160M to date. AssemblyAI is the voice AI infrastructure platform powering 10,000 companies, including Granola, Zoom and Delta Airlines. https://www.assemblyai.com/ Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs Chapters: 02:08 - What AssemblyAI actually does 05:23 - Dylan learns to code and discovers ML 07:11 - The Amazon Echo moment 09:32 - Why Dylan built voice AI infrastructure 13:02 - Building AI before anyone cared 16:50 - The 2021 inflection point 24:13 - Real-time voice agents are here 28:26 - Inside AssemblyAI’s new voice models 45:33 - Lessons from hypergrowth 52:00 - The future of voice AI

Jared FriedmanhostDylan Foxguest
Mar 5, 202653mWatch on YouTube ↗

CHAPTERS

  1. 0:05 – 1:37

    AssemblyAI today: voice AI infrastructure at massive scale

    Dylan explains AssemblyAI’s role as the underlying platform that lets other companies build voice AI features like note-takers, contact center analytics, and real-time voice agents. He shares scale metrics (developers, customers, and voice hours processed) to frame how foundational the company has become.

    • AssemblyAI provides voice AI primitives/infrastructure rather than end-user apps
    • Scale: ~1M developers, ~10K customers, hundreds of millions of voice hours processed
    • Common applications: note-taking, contact centers, real-time agents, healthcare ambient scribes
    • Focus on enabling fast developer innovation with a strong API/UX
  2. 1:37 – 3:41

    Customer examples: note-takers, contact centers, and enterprise deployments

    Jared asks for recognizable examples, and Dylan names products and sectors where AssemblyAI is embedded. The discussion highlights how AssemblyAI often powers experiences indirectly through partners and large platforms.

    • Note-taking customers include Granola and Fireflies; also hiring workflows like Ashby/MetaView note-taking
    • Contact center deployments via vendors (e.g., Delta-related call flows through partner stack)
    • Enterprises and Fortune 500s building internal voice automation for ops and trust/safety
    • Zoom referenced as a major user of AssemblyAI infrastructure
  3. 3:41 – 6:38

    From college startups to machine learning: Dylan’s path into AI

    Dylan recounts teaching himself to code while building small college-era startups, then progressively deepening into ML. He describes moving to San Francisco, joining Cisco, and encountering early neural network momentum (TensorFlow-era).

    • Self-taught programming via books; built early SaaS side projects
    • Fell in love with the rapid creative loop of coding and shipping
    • Early ML interest pre-deep-learning boom (SVM-era), then shifted to neural nets
    • Joined Cisco ML team around 2015; early exposure to TensorFlow community
  4. 6:38 – 7:38

    The Amazon Echo “it finally works” moment for voice

    Buying an Amazon Echo in 2015 changed Dylan’s perception of what voice could be because far-field recognition worked reliably in noisy environments. That reliability created new user habits and motivated him to explore building on voice tech himself.

    • Prior voice UX (Siri-era) was unreliable and not habit-forming
    • Echo’s far-field accuracy and noise robustness felt like a step-function improvement
    • Reliability drives behavior change (timers, weather, music)
    • Sparked Dylan’s experimentation mindset: “What can I build with this?”
  5. 7:38 – 10:13

    Why AssemblyAI: Twilio/Stripe for voice, not CD-ROM SDKs

    Dylan contrasts the developer experience he wanted with the reality of incumbent voice tooling (Nuance’s expensive, clunky SDK distribution). He outlines the founding thesis: use deep learning to improve voice capabilities and deliver them through an easy, modern developer platform.

    • Market gap: voice tooling was either poor quality or inaccessible to developers
    • Incumbent experience (Nuance): high upfront cost, mailed SDKs, poor DX
    • Inspiration from Twilio/Stripe: self-serve APIs unlock creativity
    • AssemblyAI positioned as infrastructure (not an app company) for all developer types
  6. 10:13 – 12:24

    Building before AI was cool: obsession, conviction, and early believers

    The conversation turns to how early the company was—when “AI” was considered hype or a scam label. Dylan argues founders need obsession with the problem to survive long timelines, and he credits early support (including YC partner Daniel Gross) for sustaining the company.

    • 2017 sentiment: “AI” label was stigmatized; “deep learning” was safer language
    • Founder advice: choose a problem you’re personally obsessed with and would use
    • Dogfooding and demos kept motivation high despite slow progress
    • Daniel Gross as an early believer and supporter post-YC
  7. 12:24 – 17:13

    The long slog (2017–2020): slow iteration cycles and missing ecosystem pieces

    Dylan describes how hard it was to iterate on ML models inside YC-style timelines and why voice apps didn’t take off yet. He explains the broader ecosystem dependencies needed for voice AI applications to flourish beyond raw transcription.

    • ML iteration was slow: user feedback could take a month to address via new models
    • In 2017 the market was tiny and the enabling tech stack was incomplete
    • Beyond STT, needed: LMs, vector DBs, WebRTC, 5G, and better real-time infra
    • Early “interesting apps” (summaries/sentiment) required training multiple bespoke models back then
  8. 17:13 – 18:19

    2021 inflection point: COVID data surge + better transformers + expanding TAM

    Dylan pinpoints why adoption accelerated around 2021: remote work created more voice data, models improved, and NLP advances made downstream tasks easier. This culminated in the first major customer, Series A in early 2022, and rapid capital formation.

    • COVID increased internet-captured voice data (meetings, podcasts, remote workflows)
    • Model improvements: more data, transformers, better/cheaper transcription
    • NLP building blocks (e.g., BERT-era) enabled summarization and sentiment more easily
    • First “real” customer in 2021 (contact center); Accel-led Series A Jan 2022; rapid fundraising thereafter
  9. 18:19 – 23:09

    From batch to real time: why real-time voice agents are exploding now

    The interview distinguishes earlier non-real-time batch processing from today’s real-time surge. Dylan explains that only recently have real-time models crossed key thresholds in accuracy, latency, and cost—unlocking new product categories.

    • Early usage was almost entirely pre-recorded audio analytics and processing
    • Non-real-time APIs still growing rapidly, but real-time is accelerating faster
    • Real-time required crossing thresholds in latency + accuracy + cost in last ~18 months
    • Remaining gaps: speaker ID and other real-world robustness challenges
  10. 23:09 – 27:28

    Where voice AI is going: agents, robotics/hardware, and ambient intelligence

    Dylan shares the biggest product directions he’s seeing: voice agents in customer support, voice as a UI for devices and robots, and ambient capture in healthcare and sales. He gives concrete examples of how improved models enable high ROI workflows.

    • Real-time voice agents now deliver acceptable UX and strong ROI in frontline tasks
    • Growing demand from robotics and consumer hardware for voice interaction
    • Ambient healthcare scribes: noisy, far-field conversations transcribed with high accuracy to automate documentation and insurance workflows
    • Ambient sales coaching (e.g., field sales) providing real-time advice with measurable earnings impact
  11. 27:28 – 31:37

    Inside AssemblyAI’s new models: “smarter” voice understanding via instructions

    Dylan introduces Universal 3 Pro and the goal of making voice models more intelligent and controllable—not just transcription. He frames it as a middle ground between traditional STT and multimodal LLMs, emphasizing reliability and real-time deployment.

    • Roadmap focus: models that understand context, roles, noise, stress, multilingual speakers, and command-source separation
    • Universal 3 Pro: instruction-following audio model designed to stay “on the rails” for transcription-like tasks
    • Positioned between STT and multimodal LLMs: more controllable and reliable than general models
    • Supports real-time usage and self-hosting to reduce latency
  12. 31:37 – 34:52

    Live demo: verbatim accuracy, alphanumerics, whispering, and translation prompts

    Dylan demonstrates Universal 3 Pro’s real-time transcription strengths: capturing stutters, complex emails/ticket-like strings, and whisper audio. He then shows prompt-based control, including translating transcription output into Spanish and expanding multilingual support.

    • Low-latency verbatim capture (including disfluencies)
    • Strong performance on emails, long digit strings, and alphanumerics (common call-center requirements)
    • Robustness to challenging conditions (e.g., whispering)
    • Prompted behavior: translate-to-Spanish transcription; multilingual roadmap expansion beyond initial languages
  13. 34:52 – 38:47

    Controllability for real-world audio: cross-talk handling and app-specific requirements

    The discussion moves from demo wow-factor to developer value: different applications want different behaviors (ignore background speakers, mark cross-talk, or transcribe everything). Dylan shows how instruction prompting can control these tradeoffs without resorting to a full general-purpose LLM.

    • Apps vary: some need primary-speaker-only; others need background context
    • Prompt can mark cross-talk segments without transcribing them
    • Prompt can be adjusted to transcribe background speech when needed
    • Emphasis: model remains STT-like and predictable while adding instruction-based control; rapid iteration planned
  14. 38:47 – 44:27

    Competing with giants: winning through subject-matter expertise and feedback loops

    Jared asks about competing against Google and Nuance; Dylan argues incumbents’ resources don’t guarantee superior products. He explains AssemblyAI’s advantage as deep, hands-on expertise driven by scale, direct customer proximity, and a culture that surfaces unfiltered feedback.

    • Investors were skeptical due to entrenched incumbents and perceived commoditization
    • Differentiator: being an intense user of the tech and understanding where incumbents fall short
    • Wright brothers vs. Langley analogy: iteration and direct learning beat funding alone
    • Customer feedback at scale (Slack channels, unfiltered critique) informs rapid product fixes and priorities
  15. 44:27 – 50:47

    Hypergrowth lessons and operating lean: hiring, focus, and minimal overhead

    Dylan reflects on scaling from a tiny team to rapid growth after fundraising, and the mistakes that come with hiring to ‘explore’ without conviction. He outlines AssemblyAI’s current operating philosophy: a dense team, speed over heavy process, and strict hiring non-negotiables tied to mission fit.

    • Key mistake: investing/hiring before clarity on what’s core vs exploratory
    • New approach: explore with existing team first, then invest when signal is strong
    • Hiring rigor: role-specific non-negotiables and passion for voice AI specifically (not just ‘AI’)
    • Lean ops philosophy: avoid heavy OKR cascades; transparency + fast adaptation; ~80-person team with open roles across functions
  16. 50:47 – 53:00

    Living in the voice-first future: company-wide transcripts as a shared brain

    The episode ends with how AssemblyAI uses its own tech internally: pervasive note-taking, centralized transcripts, and AI-driven synthesis of customer feedback. Dylan argues that making “truth” broadly accessible reduces subjective filtering and gives teams direct line-of-sight to customers.

    • AI note-takers in most meetings; transcripts feed a company knowledge base
    • Ability to query: “What should our roadmap be based on customer feedback?”
    • Democratizes customer truth for engineers and teams—fewer layers and less bias
    • Belief: AI-augmented organizations will outcompete those without augmentation

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.