Skip to content
YC Root AccessYC Root Access

Building the Safety Layer for AI Agents

Raindrop is building the safety layer for AI agents. As agents get more capable and take on more complex work, Raindrop detects when they go wrong in production, from failed tool calls and hallucinations to problems companies didn’t even know to look for. They’re already used by companies including Vercel, Clay, Framer, and Speak, and just raised a Series A, bringing their total funding to $50 million. In this episode of Founder Firesides, the founders sat down with YC's Diana Hu to share how a tool they originally built to debug their own coding agent became the company, why better AI agents actually create more ways for things to go wrong, and how their new simulations product can catch failures before they ever reach production. https://www.raindrop.ai 00:00 — Intro 00:05 — What Raindrop Does 01:32 — Why Better AI Agents Create More Problems 03:25 — Finding the Failures Evals Miss 04:38 — How Raindrop Got Started 06:32 — Why They Kept Going 07:58 — Catching Agent Failures Before Production Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Diana HuhostRaindrop foundersguest
Sep 17, 202610mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Raindrop builds a safety layer to detect and prevent agent failures

  1. Raindrop provides a safety and reliability layer for AI agents by detecting and preventing failures using comprehensive production trace analysis.
  2. The founders argue that as agents become more capable, they create more complex and higher-stakes failure modes, increasing demand for monitoring and prevention.
  3. They contend that LLM-as-judge approaches are constrained by cost and by the difficulty of binary correctness decisions that vary by company context.
  4. Raindrop differentiates by proactively discovering unknown issues in logs and by launching “Raindrop Simulations” to catch failures in CI before code ships.
  5. The company evolved from building a coding agent to building the internal debugging/issue-detection tool that became its core product, now used by startups and Fortune 500s.

IDEAS WORTH REMEMBERING

5 ideas

Production observability for agents requires trace-level monitoring, not just sampled evaluation.

Raindrop monitors agent behavior using full production traces and a large set of failure-mode metrics (e.g., tool-call failures, loops, hallucinations, capability gaps) so teams can see exactly where and how agents break in real workflows.

Better models increase safety demand because capability growth amplifies risk and unexpected failure modes.

As agents become more capable, they call more tools, touch more critical data, and get deployed in higher-stakes domains (healthcare, finance, military), which increases both complexity and the cost of failures.

Judging agent correctness is a hard, domain-defined binary decision that doesn’t outsource cleanly to a generic judge model.

They argue “LLM-as-judge” approaches are limited by inference cost at scale and by the inherent difficulty of binary classification (deciding if behavior is right/wrong), which is company- and context-specific.

The biggest failures are often the ones you didn’t know to write an eval for.

Unlike evals that check known criteria, Raindrop aims to surface novel, unanticipated issues that can sit unnoticed in logs for weeks—creating liability and wasted spend (e.g., excess token costs).

Catching agent failures pre-production requires simulation-driven CI, not only post-hoc monitoring.

Raindrop Simulations shifts detection left: for every PR/change, it predicts how the agent might fail before deployment by bringing production-grade issue detection into CI.

WORDS WORTH SAVING

5 quotes

Really simply, we detect issues with agents in production, and now prevent those issues from getting into production in the first place.

Raindrop founders

A customer support agent is trying to issue a refund, but the tool called to issue the refund fails, and the agent says, "I issued you a refund." So that's a problem.

Raindrop founders

As capabilities go up, complexity goes up, and the cost of issues goes up dramatically, and you have all these failure modes that you never anticipated.

Raindrop founders

Binary classification is extremely hard.

Raindrop founders

When your life is you wake up in the morning and you walk to an office and work on the most interesting problems in the world with your best friends, you're like, "Li- life is pretty good."

Raindrop founders

Agent observability and tracingProduction issue detection vs evalsLLM-as-judge cost and limitsBinary classification and alignment to company policiesTool-call failures and action/claim mismatchProactive discovery of unknown failure modesCI simulations to prevent pre-production regressions

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.