Skip to content
YC Root AccessYC Root Access

David AI: Powering the Voice Era of AI

Tomer Cohen and Ben Wiley launched David AI just days before the Y Combinator deadline—submitting their application at midnight and hoping it counted. A year later, their company is now one of the market leaders for voice training data in AI, having just closed a $25 million Series A. They met while working at Scale AI, where they bonded over the belief that the next big leap for AI would be moving beyond screens, into real-world interactions powered by voice. That idea became David AI, a company that collects, produces, and refines massive volumes of audio data for training voice models. So far, they've built a library of 100,000 hours of audio in over 15 languages, complete with rich metadata like accents and dialects. YC Partner Diana Hu recently sat down with the David AI founders to talk about how they got here, their founding story, and the kind of company they are building. Learn more about David AI at https://www.withdavid.ai. Apply to Y Combinator: https://ycombinator.com/apply Chapters 00:00 - Introduction 00:12 - What is David AI? 00:31 - Challenges in Audio Data 01:11 - Origin Story of David AI 01:46 - Building the First Product 04:12 - Early Success and Growth 05:24 - Business Model and Approach 07:40 - Future Plans and Hiring

Diana HuhostTomer CohenguestBen Wileyguest
May 28, 20259mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:31

    David AI’s milestone: YC S24 to $25M Series A, and what they do

    Diana opens by introducing David AI’s recent progress—YC Summer ’24 and a $25M Series A. Ben and Tomer define the company as an audio data research business focused specifically on conversational speech across languages, accents, and contexts.

    • Company context: YC S24 + newly announced Series A
    • Core focus: audio → speech → conversational data
    • Emphasis on diversity: languages, dialects, accents, conversational settings
  2. 0:31 – 0:36

    Why high-quality conversational audio data is uniquely hard to source

    The team explains why audio lacks a “common crawl” equivalent and why internet audio is often unusable for training frontier speech models. They highlight how mono-channel recordings and channel bleed break the requirements of modern end-to-end speech architectures.

    • No “common crawl” for large-scale usable audio data
    • Most online audio is mono/single-track, not multi-channel
    • End-to-end speech models require extremely clean channel separation
    • Off-the-shelf source separation isn’t sufficient for low-bleed needs
  3. 0:36 – 1:08

    The key technical insight: collect audio separated at the source

    Ben describes reaching the conclusion that the only reliable path to model-grade conversational data is collecting it correctly from the start. Instead of trying to “fix” flawed recordings after the fact, David AI designs collection so separation is inherent.

    • Model tolerance for cross-talk/bleed is very low
    • Post-hoc separation tools didn’t meet quality thresholds
    • Highest-quality datasets require source-separated collection
    • This insight drives the company’s product and operations strategy
  4. 1:08 – 1:44

    Origin story: from Scale friendships to betting on the voice era

    Tomer shares how he and Ben met at Scale, bonded, and became excited about multimodal AI—especially voice as a pathway to real-world interactions. They applied to YC with an exploratory idea, got in, left their jobs, and moved to SF to build.

    • Founders met at Scale and wanted to build together
    • Thesis: voice AI as the next evolution for real-world interfaces
    • Applied to YC pre-product, then committed full-time post-acceptance
    • Early stage focused on exploration and customer discovery
  5. 1:44 – 2:22

    Customer discovery “aha”: robotics needed audio more than anything else

    Early outreach to YC companies revealed a surprising signal: a humanoid robotics company’s biggest pain point was audio data for voice interaction. That insight clarified the opportunity—audio data was an under-served bottleneck even for frontier teams.

    • Systematic outreach to multimodal/AI model builders
    • Robotics customer highlighted audio as the critical missing piece
    • Reframed audio data as a foundational dependency for embodied AI
    • Helped the team narrow focus and commit to voice data depth
  6. 2:22 – 3:42

    Contrarian focus: why going deep on audio beats going broad on “all data”

    Diana presses on confidence given incumbents like Scale. Tomer argues audio is central to many non-keyboard interfaces (robots, wearables, games, avatars) and that a vertical, deep approach can outperform horizontal generalists by solving the hardest modality-specific problems.

    • Audio is broader than call centers: robots, wearables, games, avatars
    • Real-world AI interfaces often require voice/audio
    • Counterpoint to “horizontal data companies win” narrative
    • Strategy: pick one modality and build deep, repeatable capabilities
  7. 3:42 – 4:13

    First product in a weekend: a phone-calling app to generate initial datasets

    Ben describes building a simple calling application over a weekend to validate data collection hypotheses using friends and family. That scrappy prototype evolved into a global platform enabling both scripted and unscripted conversational collection.

    • Rapid prototyping to test collection hypotheses
    • Used personal networks to create an initial dataset quickly
    • Progression from prototype to worldwide data collection platform
    • Supports multiple conversation modes: scripted and unscripted
  8. 4:13 – 4:46

    Revenue traction: from $1K pilot to six-figure deal by the end of the batch

    Tomer outlines how early small contracts helped them build a repeatable perspective on audio data value. Starting with a $1K robotics deal, they iterated customer-to-customer until landing a six-figure contract with a major AI lab by batch end.

    • First customer: robotics, $1K contract
    • Each engagement sharpened their understanding of model-ready audio data
    • Compounding learning enabled larger follow-on deals
    • By end of YC batch: first six-figure contract with a big AI lab
  9. 4:46 – 5:24

    Scaling beyond YC: seven-figure deals, big tech customers, and “data sells itself”

    Post-batch growth accelerated quickly into seven-figure contracts and relationships with major tech companies. Tomer frames the sales motion as evidence-driven—labs evaluate whether the dataset works, making the transaction feel less like selling and more like adoption.

    • Rapid post-batch step-up to seven-figure contracts
    • Now serves major tech companies with massive audio needs
    • Sales motion strengthens as data volume and quality compound
    • Positioning: labs choose data based on utility; minimal hard selling
  10. 5:24 – 6:58

    Operating model: an audio data research lab, not bespoke professional services

    David AI differentiates itself from traditional labeling/pro-services models. They research where models are going, validate internally via R&D, then scale winning datasets and offer them broadly—rather than building one-off datasets owned exclusively by a single customer.

    • Self-definition: “data research lab” with its own model/data viewpoint
    • Internal R&D validates which data shapes improve models
    • Once validated, they scale datasets and release to the market
    • Contrast: bespoke collection + customer-owned data + take-rate model
  11. 6:58 – 7:40

    The “picks and shovels” layer behind voice agents and voice AI apps

    Diana connects David AI’s work to the surge in voice agents across verticals. Tomer explains the dependency stack: great apps require great models, and great models require great data—especially in audio, where the data layer receives less attention.

    • Voice agent category growth increases demand for high-quality speech data
    • Dependency chain: apps → models → data
    • Audio data is under-appreciated relative to app-layer momentum
    • David AI positions itself as foundational infrastructure for the ecosystem
  12. 7:40 – 9:00

    Next five years: building research depth, scaling collection 10x, and hiring

    Tomer lays out the roadmap: strengthen the audio research function to anticipate model needs, and build product/operations that can scale data collection by orders of magnitude. They close with near-term hiring priorities across research, engineering, and operations.

    • Priority: build a strong audio research function to predict model direction
    • Parallel push: scale collection capabilities 10x, then 10x again
    • Team growth as the main execution lever
    • Hiring focus: researchers, engineers, and operators

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.