Skip to content
YC Root AccessYC Root Access

What It Actually Takes to Deploy a Voice Agent to a Fortune 500

Brooke Hopkins is the founder and CEO of Coval (S24), a simulation and observability platform for voice agents that helps enterprises test, monitor, and evaluate AI-powered phone systems at scale — working with customers like Perplexity and Deepgram to process tens of millions of calls per month — and has just raised a $28.2M Series A. In this fireside, Brooke sat down with Harj Taggar, Managing Partner at YC to talk about how her years building evaluation infrastructure and developer tools at Waymo turned out to be surprisingly transferable to the world of voice agents, why voice is emerging as the first truly productionized use case for autonomous agents, and what it took to go from a broader evals idea to a deeply focused enterprise platform — including the moment a customer offered to pay her before she'd written a single line of code. https://www.coval.dev Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Harj TaggarhostBrooke Hopkinsguest
Jun 24, 202630mWatch on YouTube ↗

CHAPTERS

  1. 0:05 – 1:16

    Coval’s mission: simulation + observability for voice agents at massive scale

    Harj introduces Brooke Hopkins and Coval, framing the company as infrastructure for testing and monitoring voice agents before and after deployment. Brooke explains Coval’s core promise: scale voice agents to millions of conversations without relying on risky production-only learning.

    • Coval simulates and evaluates voice agents before they hit real customers
    • Observability in production: understand what happens “in the wild”
    • Enterprise-grade scale: monitoring tens of millions of calls/month
    • Roots in Waymo-style simulation and evaluation infrastructure
  2. 1:16 – 2:37

    Why voice agents are the breakout ‘autonomous agent’ interface

    Brooke argues voice is taking off because it’s the most natural interface and one of the first truly productionized autonomous-agent use cases. Voice also lowers adoption friction by meeting users in real-world contexts where traditional software isn’t pervasive.

    • Voice as a natural UI for headless AI/agents
    • First widely deployed autonomous-agent category
    • Works in logistics, healthcare, and environments without rich software
    • Enables automation with minimal behavior change from users
  3. 2:37 – 4:30

    Enterprise adoption: leveraging existing call infrastructure and expanding beyond support

    Enterprises are adopting voice agents quickly because they already have call flows, SOPs, and IVR infrastructure—making the leap to autonomy smaller. Teams often start with support and then discover many other internal and customer-facing opportunities.

    • Existing IVR/call flows accelerate deployment readiness
    • Start in customer support, then expand across the org
    • New use cases: concierge/product discovery, logistics, back office
    • Voice evolution mirrors web→mobile: from basic replication to native experiences
  4. 4:30 – 5:29

    The ‘positive vision’ for voice: better experiences, not just labor replacement

    Harj and Brooke discuss how voice agents can create new value, such as improving customer experience and unlocking complex actions via simple requests. Brooke uses airlines as an example of moving from long hold times to proactive, context-aware assistance.

    • Voice can drive sales/adoption, not only cost reduction
    • Airline example: rebooking and decision-making via a simple call
    • Voice distills complex options into goal-based interaction
    • Rising expectations: long hold times become unacceptable
  5. 5:29 – 6:47

    What Coval provides: missing infrastructure for scalable voice apps

    Brooke compares today’s voice stack to early web infrastructure—fragile and hard to scale without specialized tooling. Coval focuses on enabling reliability, compliance, and insight across millions of conversations, turning voice interactions into actionable product data.

    • Voice lacks mature ‘serverless-like’ infrastructure for scale
    • Detect failures, compliance risks, and security issues
    • Analyze unstructured conversation data for product/upsell insights
    • Goal: make enterprise voice apps scalable and understandable
  6. 6:47 – 8:54

    How voice agents fail: from hallucinated audio to workflow mistakes

    Voice agents excel at instantly reflecting updated policies, but fail in surprising and sometimes dramatic ways. Brooke outlines key evaluation dimensions: task success, correct workflow/tool use, and audio-quality factors that shape user trust.

    • Strength: rapid propagation of product/policy updates
    • Brittleness: egregious errors (wrong info) and ‘vocal hallucinations’
    • Audio anomalies: screaming/whispering/voice shifts
    • Core evaluation axes: outcome, workflow/tool calls, audio/latency/noise
  7. 8:54 – 10:53

    Building an enterprise evaluation strategy (and what people mis-measure)

    Coval works closely with enterprises to create scalable eval processes, inspired by self-driving’s safety-and-improvement flywheels. Brooke notes a common misconception: overvaluing word error rate instead of measuring intent understanding and task completion.

    • Coval helps design a scalable evaluation system, not just metrics
    • Self-driving-style continuous improvement flywheel applied to voice
    • Misconception: word error rate often matters less than intent/outcome
    • Agents can struggle with conversations that restart or front-load info
  8. 10:53 – 12:32

    Next unlocks: controllability in real-time voice models and better model integration

    Brooke predicts major progress will come from making real-time models more controllable and from better integration across components. She maps voice architecture to autonomy loops in robotics: perception, planning, and control—suggesting the future is neither one monolith nor fully separated components.

    • Today’s common stack: STT → LLM → TTS (cascaded)
    • Analogy to autonomy: perception → planning/reasoning → control
    • Future: share embeddings/context across steps while maintaining specialization
    • Trend mirrors self-driving: model condense + specialize over time
  9. 12:32 – 14:50

    Waymo to Coval: datasets, developer tools, and why edge cases matter

    Brooke describes her work at Waymo: building dataset infrastructure focused on rare but critical scenarios and later leading developer tooling for large-scale simulation runs. These experiences shaped Coval’s approach to realism, determinism, and targeted simulation in voice.

    • Waymo datasets prioritized critical edge cases over average performance
    • Built tooling to combine datasets + configs and run on distributed compute
    • Realization: many AI deployment problems resembled Waymo-era challenges
    • Simulation concerns: realism, determinism, partial substitution of audio/context
  10. 14:50 – 16:38

    Finding the wedge: from generic evals to voice via intense customer pull

    Brooke recounts how the company pivoted from a broader evals idea to voice after direct customer demand—so strong that a customer offered to pay even before software existed. Fonely became the first key design partner, validating the wedge and shaping the product.

    • Initial direction: general evals felt crowded/less urgent
    • Voice customer pain was acute; willingness to pay signaled strong pull
    • YC principle: don’t just listen—interpret needs as a window into their world
    • Design-partner approach with Fonely helped refine the product early
  11. 16:38 – 18:20

    Recognizing real PMF: procurement momentum and the cost of manual testing

    Brooke contrasts ‘tire kickers’ with true product-market fit, described as customers pushing the deal through procurement. The pain is amplified by the time and cost of repeated call testing, and by the high stakes of deploying to millions of conversations.

    • Earlier: interest without revenue vs later: customers “chasing” and unblocking procurement
    • Manual testing is expensive and low-signal (10 calls ≠ confidence for millions)
    • Reliability is critical; failures are high-impact in voice
    • Scale forces the need for systematic evaluation and monitoring
  12. 18:20 – 22:29

    Why focus on enterprise early: scale problems, roadmap clarity, and founder-market fit

    Brooke explains the deliberate choice to go enterprise-first: the hardest scaling challenges appear when hundreds of engineers collaborate on one agent system. She also notes the balance—don’t ignore fast-moving startups—but enterprise depth can be a strong signal when the product is inherently about scale.

    • Enterprise offers consistency and longer-term roadmaps for collaboration
    • Voice quality + compliance challenges intensify at large scale
    • Founder-market fit: experience building tools for hundreds of engineers at Waymo
    • Advice: don’t go too fast upmarket, but enterprise pull can validate value
  13. 22:29 – 25:35

    A new category: agent testing + observability, and why validation time is growing

    Brooke positions Coval as akin to Datadog/Applied Intuition—but for AI agents—combining testing and observability. She argues developer time is shifting away from ‘building’ toward planning, validation, rollout, and continuous correctness over time.

    • Category framing: testing + observability for AI agents
    • Engineering time shifts: planning/validation dominate as building gets cheaper
    • “Software engineering is the integral of programming over time”
    • Need to ensure systems keep working continuously, not just once
  14. 25:35 – 30:46

    Founder journey: solo-founder rationale, YC’s bar, and Coval’s next roadmap

    Brooke shares her motivation to found a company and why solo founding worked for Coval’s technical, category-creation challenge. Harj explains YC’s exception criteria for solo founders, then Brooke closes with roadmap focus on agentic evals and self-improving systems—plus founder advice to be obsessive and move fast.

    • Entrepreneurial roots; desire for creativity and building despite uncertainty
    • YC solo-founder bar: technical + build-and-sell + exceptional founder-market fit
    • Roadmap: agentic evals, high-signal diagnosis, self-improving optimization loops
    • Advice: obsession, curiosity, and speed can compensate for shortcomings

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.