Skip to content
YC Root AccessYC Root Access

Building the Safety Layer for AI Agents

Raindrop is building the safety layer for AI agents. As agents get more capable and take on more complex work, Raindrop detects when they go wrong in production, from failed tool calls and hallucinations to problems companies didn’t even know to look for. They’re already used by companies including Vercel, Clay, Framer, and Speak, and just raised a Series A, bringing their total funding to $50 million. In this episode of Founder Firesides, the founders sat down with YC's Diana Hu to share how a tool they originally built to debug their own coding agent became the company, why better AI agents actually create more ways for things to go wrong, and how their new simulations product can catch failures before they ever reach production. https://www.raindrop.ai 00:00 — Intro 00:05 — What Raindrop Does 01:32 — Why Better AI Agents Create More Problems 03:25 — Finding the Failures Evals Miss 04:38 — How Raindrop Got Started 06:32 — Why They Kept Going 07:58 — Catching Agent Failures Before Production Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Diana HuhostRaindrop foundersguest
Sep 17, 202610mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:000:05

    Intro

    1. DH

      [upbeat music]

  2. 0:051:32

    What Raindrop Does

    1. DH

      I'm excited today to welcome the founders of Raindrop on closing a big Series A led by CRV, bringing your total funding to $50 million, where Lightspeed did your seed round. So this was just two years after the batch. That's pretty cool. Tell us, what do you guys do?

    2. RF

      Yeah. Really simply, we detect issues with agents in production, and now prevent those issues from getting into production in the first place.

    3. DH

      So who are some of your big customers right now?

    4. RF

      So we have a number of really amazing AI agents that we work with, companies like Vercel, Clay, Framer, Speak, et cetera, and then als- also Fortune 500s of multiple industries, yeah.

    5. DH

      Do you have concrete examples of how these customers use you?

    6. RF

      Yeah, there's all sorts. So one common one with customer support agents is basically a tool called agent response mismatch. So let's look at something really specific. A customer support agent is trying to issue a refund, but the tool called to issue the refund fails, and the agent says, "I issued you a refund." So that's a problem.

    7. DH

      Hmm.

    8. RF

      That's definitely one, right? Around, like, an agent saying it took an action and it didn't. There's, you know, capability gaps, like the user asking for something the agent can't do. Task failure, like the agent isn't able to complete something for some reason. And we really just h- have hundreds of these sort of metrics, and what we see is that companies, like the entire company, use these metrics to determine how good or bad their agent is.

  3. 1:323:25

    Why Better AI Agents Create More Problems

    1. DH

      So you guys basically are watching whenever agents go rogue, whenever you have problems like hallucination, errors with tool calls, broken loops, or any of these kinds of issues. What is very interesting to me, and counterintuitive, is that as model capabilities increase, it seems like the big labs should solve this, but it's the opposite problem. You guys are seeing more demand than ever with the newest models. Why is that?

    2. RF

      Yeah. So you can see with Hugging Face, for example, right, that as agents get better, they're calling hundreds of tools. They're reading and writing from critical data sources. They're being deployed in healthcare, the military, and finance. And so as capabilities go up, complexity goes up, and the cost of issues goes up dramatically, and you have all these failure modes that you never anticipated.

    3. DH

      And a regular LLM as a judge doesn't cut it. Why is that?

    4. RF

      I think there's really two reasons. The first is easy, and it's cost. Uh, I think that if you were gonna have an LLM look at every single trace, it would double your costs at least. I think the second is actually a lot trickier, and something I think our company sort of exists to solve, which is that binary classification is very, very hard. And so we were talking earlier, there's this paper from OpenAI, uh, maybe two years ago or something now, where they said, you know, "Hallucination is a binary classification problem." Like, something is either true or it's not true. And there were all these people on Twitter that were kinda like, "Wow, OpenAI solved hallucination," because it's just binary. And it's like, no, actually, binary classification is extremely hard. Like, every choice in your life you could say is either you should do it or you should not do it. And I think that similarly, something in your agent, some behavior it's doing, is either wrong or it's right, and it's really up to the company to define what's wrong and what's right. And so we sort of think of ourselves as an alignment company, aligning humans inside of a company with their agent's behavior.

  4. 3:254:38

    Finding the Failures Evals Miss

    1. DH

      So I think the cool thing that you do then is that you enable your customers to basically watch the full traces of all of their agents rather than just doing sampling-

    2. RF

      Yes

    3. DH

      ... and then getting a judge to run, which is how people do it today, but it doesn't catch all the errors.

    4. RF

      Yeah.

    5. DH

      And you guys basically catch all of them.

    6. RF

      Yes.

    7. RF

      Yeah, and I think it's more than just catching errors. It's proactive issue detection. So [clears throat] when you have LLMs as a judge, again, you can only run it on 1% of data, 10%, but also you can only look for problems you know about ahead of time. You tell your LLM, "Look at this. You know, rate the answer on, like, XYZ parameter I care about." But what Raindrop does is proactively find issues you didn't know about, right? So if you take the OpenAI, like Hugging Face, uh, problem, these agents are talking. They have all these, like, tons and tons of logs, and it's going unnoticed for weeks because you didn't know to look for it. And Raindrop, by having powerful signals but also having proactive issue detection, you're finding in your... You know, our customers are finding issues that are creating risk, creating liability, and also creating excess token costs, like, across things they didn't know about, and that's what's interesting.

  5. 4:386:32

    How Raindrop Got Started

    1. DH

      This journey sounds pretty straightforward, but the reality is when you guys went through the batch about two years ago, you were a very different company. When I interviewed you guys, you were building basically a coding agent for the terminal before agents were even a thing. That was not even, not even the spoken word in terms of how people thought about agents, but you were bu- bu- basically building a coding agent on the terminal.

    2. RF

      Yeah.

    3. DH

      So tell us what happened.

    4. RF

      Totally. Yeah, we were... I mean, we've all been best friends for many years, and we were all just hacking on things that we wanted for ourselves, and very early, we found ourselves wanting what are now called coding agents. Agents are, like, basically at the time we'd say chatbots that can do things, and this was, like, the earliest way we thought about agents. And as we were building those coding agents, we ran into this issue ourselves where we had no idea why these issues were popping up or what was going wrong. And so what's now Raindrop started as an internal tool, um, that we used for our own coding agents.

    5. DH

      Agents really had a very slow trickling since 2024.

    6. RF

      Yeah.

    7. RF

      Yeah.

    8. DH

      It was not even a thing until probably 2025, and it really started to work around Clockcode Opus 4.5 with agentic coding starting to work.

    9. RF

      Yeah.

    10. RF

      Yeah.

    11. DH

      And that's when things took off for you guys. Basically, your revenue went like this.

    12. RF

      Yeah.

    13. DH

      So tell us what, what, what happened. Is, like, more coding agen- agents created more- More issues for people to follow up?

    14. RF

      It, it's exactly what Zubin was saying before, is that, like, the more capabilities, the more complexity, and the more ways that they go wrong. I think what you also saw with Opus 4.5 is that the, the sort of guardrails start coming off, right? The permissions start coming off. Like, everyone's like, you know, dangerously skip all permissions.

    15. DH

      [laughs]

    16. RF

      Everyone's like, you know, "Oh, maybe we should hook it up into production. Maybe-

    17. RF

      Yeah

    18. RF

      ... it should have access to production logs." And so it's exactly the second part of what-

    19. RF

      Yeah

    20. RF

      ... Zubin was saying, is that the stakes get higher.

  6. 6:327:58

    Why They Kept Going

    1. DH

      What was actually really cool about all of you is that, as you mentioned, you are very close friends.

    2. RF

      Yeah.

    3. DH

      And this was basically your second company. Your previous one was a crypto company that you sold to Coinbase, and you were basically had this company kind of not making a lot of progress for a bit until things took off last year. What, what made you keep going?

    4. RF

      I think-

    5. RF

      [laughs]

    6. RF

      ... one of the biggest things is, like, getting to work with each other, honestly.

    7. RF

      [laughs]

    8. RF

      Yeah.

    9. RF

      Like, when your life is you wake up in the morning and you walk to an office and work on the most interesting problems in the world with your best friends, you're like, "Li- life is pretty good."

    10. RF

      Yeah.

    11. RF

      Uh, like, you're learning a lot. You're creating things that you genuinely wanna use, and I think we just had this also conviction in the future-

    12. RF

      Yeah

    13. RF

      ... where like, these AIs are going to do things, and eventually, you know, people are gonna hook them up to production databases. Eventually, they're gonna be issuing prescriptions. Like, I think it was obvious to us very early that these agents would basically run a lot of the economy, and now that they are, people are realizing that it's kind of a hair-on-fire problem.

    14. RF

      Yes. I think that we definitely frame things a lot as just, like, fundamental truths. Like, I think the, our fundamental beliefs, and everything e- else around your fundamental beliefs as a company can change. Like, the way... I w- I, like, e- in a year, our product could look entirely different. I have no idea, but, like, the things we know as a company are that, like, one, agents will get more capable, and two, that the cost of their mis- mistakes will increase. And so if both of those things happen, then people will pay us money to make sure those things don't

  7. 7:5810:17

    Catching Agent Failures Before Production

    1. RF

      happen.

    2. DH

      So you guys have a pretty cool announcement that you're solving a very hard problem that was described by Dwarkesh's post on the rise and fall of agent civilization that basically describes how Hugging Face was pwned by [laughs] the OpenAI agents, and you guys have this new solution that's gonna solve this. Tell us more about that.

    3. RF

      We're gonna be launching Raindrop Simulations. So what this does is every single change you make, any PR you make, we show you how it might go wrong before it ever hits production. So we take all the issue detection we do in production and bring that to CI. We try to find cases where things might go wrong that you haven't even imagined. And so this is a very, very different strategy than evals, which mostly are testing things you already know about. And this really, today, OpenAI and Anthropic are the only ones that have access to this sort of technology, um, and we're bringing that to every company in the world.

    4. DH

      Pretty cool. So you guys can detect the rise and fall of agent civilizations.

    5. RF

      Yes. [laughs] Quite literally. I think that that is a g- extremely good example of something, where it's like, sure, now that you know about it, you could add an eval, right? But the reality, and this is true for most companies, if it's true for OpenAI, is that that sat in logs for weeks before anyone found it. It took, you know, uh, crashing their internal package manager, um, before it was detected.

    6. DH

      So tell us about the roles that you're hiring for.

    7. RF

      Yeah. Hiring across the board, right? This is such an important problem, hiring across engineering. I think these are some of the most interesting engineering problems, period, right? As Ben said earlier, we're taking some of the tools that only OpenAI had and Anthropic have in-house and just making them available to the entire world as agents proliferate. And, um, also in, when we think about agents pro- proliferating and getting these tools to everyone in the world, every developer in the world, that requires a huge distribution and sort of, like, GTM. And so we're hiring a lot for everything from marketing to sales, dev, dev rel, et cetera, et cetera. Yeah.

    8. DH

      All right, guys. Congratulations.

    9. RF

      Thank you.

    10. RF

      Thank you.

    11. RF

      Thank you. [outro music]

Episode duration: 10:17

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode 5XO7ZEOGpJc

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.