CHAPTERS
- 0:00 – 0:31
David AI’s milestone: YC S24 to $25M Series A, and what they do
Diana opens by introducing David AI’s recent progress—YC Summer ’24 and a $25M Series A. Ben and Tomer define the company as an audio data research business focused specifically on conversational speech across languages, accents, and contexts.
- •Company context: YC S24 + newly announced Series A
- •Core focus: audio → speech → conversational data
- •Emphasis on diversity: languages, dialects, accents, conversational settings
- 0:31 – 0:36
Why high-quality conversational audio data is uniquely hard to source
The team explains why audio lacks a “common crawl” equivalent and why internet audio is often unusable for training frontier speech models. They highlight how mono-channel recordings and channel bleed break the requirements of modern end-to-end speech architectures.
- •No “common crawl” for large-scale usable audio data
- •Most online audio is mono/single-track, not multi-channel
- •End-to-end speech models require extremely clean channel separation
- •Off-the-shelf source separation isn’t sufficient for low-bleed needs
- 0:36 – 1:08
The key technical insight: collect audio separated at the source
Ben describes reaching the conclusion that the only reliable path to model-grade conversational data is collecting it correctly from the start. Instead of trying to “fix” flawed recordings after the fact, David AI designs collection so separation is inherent.
- •Model tolerance for cross-talk/bleed is very low
- •Post-hoc separation tools didn’t meet quality thresholds
- •Highest-quality datasets require source-separated collection
- •This insight drives the company’s product and operations strategy
- 1:08 – 1:44
Origin story: from Scale friendships to betting on the voice era
Tomer shares how he and Ben met at Scale, bonded, and became excited about multimodal AI—especially voice as a pathway to real-world interactions. They applied to YC with an exploratory idea, got in, left their jobs, and moved to SF to build.
- •Founders met at Scale and wanted to build together
- •Thesis: voice AI as the next evolution for real-world interfaces
- •Applied to YC pre-product, then committed full-time post-acceptance
- •Early stage focused on exploration and customer discovery
- 1:44 – 2:22
Customer discovery “aha”: robotics needed audio more than anything else
Early outreach to YC companies revealed a surprising signal: a humanoid robotics company’s biggest pain point was audio data for voice interaction. That insight clarified the opportunity—audio data was an under-served bottleneck even for frontier teams.
- •Systematic outreach to multimodal/AI model builders
- •Robotics customer highlighted audio as the critical missing piece
- •Reframed audio data as a foundational dependency for embodied AI
- •Helped the team narrow focus and commit to voice data depth
- 2:22 – 3:42
Contrarian focus: why going deep on audio beats going broad on “all data”
Diana presses on confidence given incumbents like Scale. Tomer argues audio is central to many non-keyboard interfaces (robots, wearables, games, avatars) and that a vertical, deep approach can outperform horizontal generalists by solving the hardest modality-specific problems.
- •Audio is broader than call centers: robots, wearables, games, avatars
- •Real-world AI interfaces often require voice/audio
- •Counterpoint to “horizontal data companies win” narrative
- •Strategy: pick one modality and build deep, repeatable capabilities
- 3:42 – 4:13
First product in a weekend: a phone-calling app to generate initial datasets
Ben describes building a simple calling application over a weekend to validate data collection hypotheses using friends and family. That scrappy prototype evolved into a global platform enabling both scripted and unscripted conversational collection.
- •Rapid prototyping to test collection hypotheses
- •Used personal networks to create an initial dataset quickly
- •Progression from prototype to worldwide data collection platform
- •Supports multiple conversation modes: scripted and unscripted
- 4:13 – 4:46
Revenue traction: from $1K pilot to six-figure deal by the end of the batch
Tomer outlines how early small contracts helped them build a repeatable perspective on audio data value. Starting with a $1K robotics deal, they iterated customer-to-customer until landing a six-figure contract with a major AI lab by batch end.
- •First customer: robotics, $1K contract
- •Each engagement sharpened their understanding of model-ready audio data
- •Compounding learning enabled larger follow-on deals
- •By end of YC batch: first six-figure contract with a big AI lab
- 4:46 – 5:24
Scaling beyond YC: seven-figure deals, big tech customers, and “data sells itself”
Post-batch growth accelerated quickly into seven-figure contracts and relationships with major tech companies. Tomer frames the sales motion as evidence-driven—labs evaluate whether the dataset works, making the transaction feel less like selling and more like adoption.
- •Rapid post-batch step-up to seven-figure contracts
- •Now serves major tech companies with massive audio needs
- •Sales motion strengthens as data volume and quality compound
- •Positioning: labs choose data based on utility; minimal hard selling
- 5:24 – 6:58
Operating model: an audio data research lab, not bespoke professional services
David AI differentiates itself from traditional labeling/pro-services models. They research where models are going, validate internally via R&D, then scale winning datasets and offer them broadly—rather than building one-off datasets owned exclusively by a single customer.
- •Self-definition: “data research lab” with its own model/data viewpoint
- •Internal R&D validates which data shapes improve models
- •Once validated, they scale datasets and release to the market
- •Contrast: bespoke collection + customer-owned data + take-rate model
- 6:58 – 7:40
The “picks and shovels” layer behind voice agents and voice AI apps
Diana connects David AI’s work to the surge in voice agents across verticals. Tomer explains the dependency stack: great apps require great models, and great models require great data—especially in audio, where the data layer receives less attention.
- •Voice agent category growth increases demand for high-quality speech data
- •Dependency chain: apps → models → data
- •Audio data is under-appreciated relative to app-layer momentum
- •David AI positions itself as foundational infrastructure for the ecosystem
- 7:40 – 9:00
Next five years: building research depth, scaling collection 10x, and hiring
Tomer lays out the roadmap: strengthen the audio research function to anticipate model needs, and build product/operations that can scale data collection by orders of magnitude. They close with near-term hiring priorities across research, engineering, and operations.
- •Priority: build a strong audio research function to predict model direction
- •Parallel push: scale collection capabilities 10x, then 10x again
- •Team growth as the main execution lever
- •Hiring focus: researchers, engineers, and operators
