YC Root AccessWhat It Actually Takes to Deploy a Voice Agent to a Fortune 500
CHAPTERS
- 0:05 – 1:16
Coval’s mission: simulation + observability for voice agents at massive scale
Harj introduces Brooke Hopkins and Coval, framing the company as infrastructure for testing and monitoring voice agents before and after deployment. Brooke explains Coval’s core promise: scale voice agents to millions of conversations without relying on risky production-only learning.
- •Coval simulates and evaluates voice agents before they hit real customers
- •Observability in production: understand what happens “in the wild”
- •Enterprise-grade scale: monitoring tens of millions of calls/month
- •Roots in Waymo-style simulation and evaluation infrastructure
- 1:16 – 2:37
Why voice agents are the breakout ‘autonomous agent’ interface
Brooke argues voice is taking off because it’s the most natural interface and one of the first truly productionized autonomous-agent use cases. Voice also lowers adoption friction by meeting users in real-world contexts where traditional software isn’t pervasive.
- •Voice as a natural UI for headless AI/agents
- •First widely deployed autonomous-agent category
- •Works in logistics, healthcare, and environments without rich software
- •Enables automation with minimal behavior change from users
- 2:37 – 4:30
Enterprise adoption: leveraging existing call infrastructure and expanding beyond support
Enterprises are adopting voice agents quickly because they already have call flows, SOPs, and IVR infrastructure—making the leap to autonomy smaller. Teams often start with support and then discover many other internal and customer-facing opportunities.
- •Existing IVR/call flows accelerate deployment readiness
- •Start in customer support, then expand across the org
- •New use cases: concierge/product discovery, logistics, back office
- •Voice evolution mirrors web→mobile: from basic replication to native experiences
- 4:30 – 5:29
The ‘positive vision’ for voice: better experiences, not just labor replacement
Harj and Brooke discuss how voice agents can create new value, such as improving customer experience and unlocking complex actions via simple requests. Brooke uses airlines as an example of moving from long hold times to proactive, context-aware assistance.
- •Voice can drive sales/adoption, not only cost reduction
- •Airline example: rebooking and decision-making via a simple call
- •Voice distills complex options into goal-based interaction
- •Rising expectations: long hold times become unacceptable
- 5:29 – 6:47
What Coval provides: missing infrastructure for scalable voice apps
Brooke compares today’s voice stack to early web infrastructure—fragile and hard to scale without specialized tooling. Coval focuses on enabling reliability, compliance, and insight across millions of conversations, turning voice interactions into actionable product data.
- •Voice lacks mature ‘serverless-like’ infrastructure for scale
- •Detect failures, compliance risks, and security issues
- •Analyze unstructured conversation data for product/upsell insights
- •Goal: make enterprise voice apps scalable and understandable
- 6:47 – 8:54
How voice agents fail: from hallucinated audio to workflow mistakes
Voice agents excel at instantly reflecting updated policies, but fail in surprising and sometimes dramatic ways. Brooke outlines key evaluation dimensions: task success, correct workflow/tool use, and audio-quality factors that shape user trust.
- •Strength: rapid propagation of product/policy updates
- •Brittleness: egregious errors (wrong info) and ‘vocal hallucinations’
- •Audio anomalies: screaming/whispering/voice shifts
- •Core evaluation axes: outcome, workflow/tool calls, audio/latency/noise
- 8:54 – 10:53
Building an enterprise evaluation strategy (and what people mis-measure)
Coval works closely with enterprises to create scalable eval processes, inspired by self-driving’s safety-and-improvement flywheels. Brooke notes a common misconception: overvaluing word error rate instead of measuring intent understanding and task completion.
- •Coval helps design a scalable evaluation system, not just metrics
- •Self-driving-style continuous improvement flywheel applied to voice
- •Misconception: word error rate often matters less than intent/outcome
- •Agents can struggle with conversations that restart or front-load info
- 10:53 – 12:32
Next unlocks: controllability in real-time voice models and better model integration
Brooke predicts major progress will come from making real-time models more controllable and from better integration across components. She maps voice architecture to autonomy loops in robotics: perception, planning, and control—suggesting the future is neither one monolith nor fully separated components.
- •Today’s common stack: STT → LLM → TTS (cascaded)
- •Analogy to autonomy: perception → planning/reasoning → control
- •Future: share embeddings/context across steps while maintaining specialization
- •Trend mirrors self-driving: model condense + specialize over time
- 12:32 – 14:50
Waymo to Coval: datasets, developer tools, and why edge cases matter
Brooke describes her work at Waymo: building dataset infrastructure focused on rare but critical scenarios and later leading developer tooling for large-scale simulation runs. These experiences shaped Coval’s approach to realism, determinism, and targeted simulation in voice.
- •Waymo datasets prioritized critical edge cases over average performance
- •Built tooling to combine datasets + configs and run on distributed compute
- •Realization: many AI deployment problems resembled Waymo-era challenges
- •Simulation concerns: realism, determinism, partial substitution of audio/context
- 14:50 – 16:38
Finding the wedge: from generic evals to voice via intense customer pull
Brooke recounts how the company pivoted from a broader evals idea to voice after direct customer demand—so strong that a customer offered to pay even before software existed. Fonely became the first key design partner, validating the wedge and shaping the product.
- •Initial direction: general evals felt crowded/less urgent
- •Voice customer pain was acute; willingness to pay signaled strong pull
- •YC principle: don’t just listen—interpret needs as a window into their world
- •Design-partner approach with Fonely helped refine the product early
- 16:38 – 18:20
Recognizing real PMF: procurement momentum and the cost of manual testing
Brooke contrasts ‘tire kickers’ with true product-market fit, described as customers pushing the deal through procurement. The pain is amplified by the time and cost of repeated call testing, and by the high stakes of deploying to millions of conversations.
- •Earlier: interest without revenue vs later: customers “chasing” and unblocking procurement
- •Manual testing is expensive and low-signal (10 calls ≠ confidence for millions)
- •Reliability is critical; failures are high-impact in voice
- •Scale forces the need for systematic evaluation and monitoring
- 18:20 – 22:29
Why focus on enterprise early: scale problems, roadmap clarity, and founder-market fit
Brooke explains the deliberate choice to go enterprise-first: the hardest scaling challenges appear when hundreds of engineers collaborate on one agent system. She also notes the balance—don’t ignore fast-moving startups—but enterprise depth can be a strong signal when the product is inherently about scale.
- •Enterprise offers consistency and longer-term roadmaps for collaboration
- •Voice quality + compliance challenges intensify at large scale
- •Founder-market fit: experience building tools for hundreds of engineers at Waymo
- •Advice: don’t go too fast upmarket, but enterprise pull can validate value
- 22:29 – 25:35
A new category: agent testing + observability, and why validation time is growing
Brooke positions Coval as akin to Datadog/Applied Intuition—but for AI agents—combining testing and observability. She argues developer time is shifting away from ‘building’ toward planning, validation, rollout, and continuous correctness over time.
- •Category framing: testing + observability for AI agents
- •Engineering time shifts: planning/validation dominate as building gets cheaper
- •“Software engineering is the integral of programming over time”
- •Need to ensure systems keep working continuously, not just once
- 25:35 – 30:46
Founder journey: solo-founder rationale, YC’s bar, and Coval’s next roadmap
Brooke shares her motivation to found a company and why solo founding worked for Coval’s technical, category-creation challenge. Harj explains YC’s exception criteria for solo founders, then Brooke closes with roadmap focus on agentic evals and self-improving systems—plus founder advice to be obsessive and move fast.
- •Entrepreneurial roots; desire for creativity and building despite uncertainty
- •YC solo-founder bar: technical + build-and-sell + exceptional founder-market fit
- •Roadmap: agentic evals, high-signal diagnosis, self-improving optimization loops
- •Advice: obsession, curiosity, and speed can compensate for shortcomings