Skip to content
YC Root AccessYC Root Access

Building the Safety Layer for AI Agents

Raindrop is building the safety layer for AI agents. As agents get more capable and take on more complex work, Raindrop detects when they go wrong in production, from failed tool calls and hallucinations to problems companies didn’t even know to look for. They’re already used by companies including Vercel, Clay, Framer, and Speak, and just raised a Series A, bringing their total funding to $50 million. In this episode of Founder Firesides, the founders sat down with YC's Diana Hu to share how a tool they originally built to debug their own coding agent became the company, why better AI agents actually create more ways for things to go wrong, and how their new simulations product can catch failures before they ever reach production. https://www.raindrop.ai 00:00 — Intro 00:05 — What Raindrop Does 01:32 — Why Better AI Agents Create More Problems 03:25 — Finding the Failures Evals Miss 04:38 — How Raindrop Got Started 06:32 — Why They Kept Going 07:58 — Catching Agent Failures Before Production Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Diana HuhostRaindrop foundersguest
Sep 17, 202610mWatch on YouTube ↗

CHAPTERS

  1. 0:05 – 0:33

    Raindrop’s mission: detect and prevent AI agent issues

    Diana Hu introduces Raindrop following their Series A and asks for the simple explanation of what the company does. The founders frame Raindrop as both production monitoring and prevention—stopping agent issues before they ship.

    • Series A context and rapid growth since YC batch
    • Core value prop: detect issues in production
    • Extension of the mission: prevent issues from entering production
    • High-level positioning as a safety/quality layer for agents
  2. 0:33 – 0:48

    Who uses Raindrop: agent-native startups and Fortune 500 deployments

    The conversation quickly moves to customer traction and the kinds of organizations adopting agent monitoring. Raindrop cites both well-known AI-forward software companies and large enterprises across industries.

    • Named customers (e.g., Vercel, Clay, Framer, Speak)
    • Adoption by Fortune 500 companies
    • Cross-industry relevance as agents spread into more workflows
    • Signals that the problem is not niche—applies broadly
  3. 0:48 – 1:10

    Concrete failure example: tool-call errors and “agent response mismatch”

    Raindrop explains a specific, high-impact production failure pattern: the agent claims it executed an action even when the underlying tool call failed. This anchors the product in real operational risk rather than abstract hallucination concerns.

    • “Agent response mismatch” in customer support/refunds
    • Tool failures that lead to incorrect user-facing claims
    • Why this creates trust, compliance, and financial risk
    • Monitoring needs to connect tool execution to agent assertions
  4. 1:10 – 1:32

    A taxonomy of agent failure modes and metrics companies rely on

    The founders broaden from one example to a wider set of failure categories and explain how teams use Raindrop’s metrics to assess agent quality. The emphasis is on many distinct ways agents can fail beyond simple hallucinations.

    • Action-claim vs action-execution mismatches
    • Capability gaps (user asks for what agent can’t do)
    • Task failures and incomplete trajectories
    • Hundreds of tracked metrics used across the company to judge agent performance
  5. 1:32 – 2:21

    Why better models can create bigger problems: complexity and higher stakes

    Diana highlights the counterintuitive trend: demand for safety tooling rises with more capable models. The founders explain that as agents gain access to more tools and critical data, the number and cost of failure modes increases sharply.

    • More capable agents call many tools and touch critical systems
    • Deployment in high-stakes domains (healthcare, military, finance)
    • Complexity growth introduces unanticipated failure modes
    • Rising blast radius and cost of mistakes drives demand
  6. 2:21 – 3:25

    Why “LLM-as-judge” isn’t enough: cost and the hard nature of binary decisions

    Raindrop argues that using LLMs to judge every trace is prohibitively expensive and still unreliable for crisp pass/fail classification. They connect this to the inherent difficulty of defining correctness and the need for company-specific alignment.

    • Running an LLM judge on every trace can at least double costs
    • Binary classification (right/wrong, true/false) is intrinsically hard
    • Correctness depends on a company’s definitions and policies
    • Raindrop as an “alignment” layer between humans and agent behavior
  7. 3:25 – 4:38

    Full-trace visibility and proactive detection beyond sampling and known checks

    The discussion shifts to why sampling-based approaches miss critical issues and why proactive discovery matters. Raindrop positions itself as finding problems teams didn’t know to look for—reducing risk, liability, and token waste.

    • Customers can inspect full traces rather than small samples
    • Sampling + predefined prompts only catch known failure patterns
    • Proactive detection surfaces unknown/novel failure modes
    • Benefits include reduced risk/liability and lower excess token spend
  8. 4:38 – 5:36

    Origin story: from a terminal coding agent to an internal debugging tool

    Diana revisits Raindrop’s YC-era direction—building a terminal coding agent before “agents” were mainstream. The founders explain that Raindrop emerged as an internal tool to understand why their own agent was failing.

    • Initial product: terminal-based coding agent
    • Agents weren’t yet a mainstream category at the time
    • Founders experienced debugging/observability gaps firsthand
    • Raindrop began as an internal tool that became the company
  9. 5:36 – 6:32

    The inflection: agent adoption accelerates and guardrails come off

    They describe how agent usage ramped materially later, coinciding with stronger agentic coding. As teams loosen permissions and connect agents to production resources, stakes rise and failure monitoring becomes mandatory.

    • Agents saw slower adoption until capabilities improved
    • Stronger models drove broader deployment and revenue growth
    • Permissions/guardrails removed (“hook it up to production”)
    • Higher-stakes access (prod logs, databases) increases risk of failures
  10. 6:32 – 7:59

    Why they persisted: friendship, conviction, and a long-term view of agent risk

    The founders share motivational factors during slower early progress, emphasizing team dynamics and belief in inevitable agent proliferation. They outline core company convictions: capabilities will rise, and the cost of mistakes will rise with them.

    • Founding team’s long friendship and enjoyment of working together
    • Prior startup experience (crypto company acquired by Coinbase)
    • Early conviction that agents will run critical parts of the economy
    • Two beliefs: agents get more capable; mistakes get more expensive
  11. 7:59 – 8:55

    New launch: Raindrop Simulations to catch failures in CI before production

    Raindrop announces a new product direction: simulations that evaluate every PR/change to predict how an agent might fail before deployment. They contrast this with traditional evals that only test known scenarios and position the approach as bringing “frontier lab” techniques to other companies.

    • Launch: Raindrop Simulations
    • Run issue detection pre-merge / in CI for every change
    • Find unimagined failure cases, not just known eval targets
    • Goal: democratize techniques typically available to top labs
  12. 8:55 – 10:17

    From “agent civilization” failures to hiring and scaling distribution

    They connect simulations to real incidents where issues lived in logs for weeks, underscoring the need for proactive discovery. The conversation closes with hiring plans across engineering and go-to-market to bring the tooling to more teams.

    • Example: failures can hide in logs for weeks before detection
    • Simulations help catch rare, emergent behaviors earlier
    • Hiring broadly: engineering plus marketing, sales, devrel
    • Scale requires both deep tech and strong distribution

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.