Skip to content
ClaudeClaude

Patrick Collison on Claude Code at Stripe

Stripe CEO Patrick Collison joined Boris to talk about how Stripe builds with Claude Code. Stripe runs its core APIs at five and a half nines of reliability, and about 55% of its pull requests now start as a prompt to Minions, Stripe's internal tool for orchestrating Claude Code on throwaway VMs. Patrick explains why isolated dev boxes and guardrails built as infrastructure made that possible, how a two-to-three person team shipped Stripe Projects in about two months, and why he thinks code quality at Stripe will be higher in three years, not lower. Have a question? Let us know in the comments. Learn more about Claude Code: https://claude.com/claude-code Chapters 0:00 Claude-powered weather model and throwaway devboxes 2:11 600 AI-written pull requests, Minions, one revert 4:23 AI guardrails as infrastructure 6:34 Stripe Projects with Claude Code 9:10 AI code quality: every codebase is a prompt 10:44 Claude Code plan mode and 10 parallel devboxes 13:06 Stripe's data: new startups up 2x with AI 16:33 How AI agents will pay each other on Stripe

Patrick CollisonguestBorishost
Sep 24, 202617mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:00 – 2:11

    Claude-powered weather model and throwaway devboxes

    1. PC

      At our house, we install the weather station. With Claude, I could design from scratch a multimodal model, and it predicts afternoon weather better than the National Weather Service.

    2. BO

      [upbeat music] I think a lot of people in the industry, like they tried to move to cloud VMs. They tried to move to these remote environments. What, what gave you the conviction to, to do this?

    3. PC

      Well, reliability and security are really important for us. Stripe operates with five and a half nines of reliability, um, so extreme, uh, reliability for our core APIs, but at the same time, we want to be developing and launching new products and adding new features extremely quickly. We want to have continuous deployment of our API, and across financial services, the way things typically work is maybe things get deployed once a month, once a quarter, certain places once a half. And I mean, that provides a kind of local stability, um, but of course at too enormous cost when you don't get sort of quick feedback from customers. You can ship things on a regular basis, but then also it does mean that these moments of migration are incredibly fraught and scary because you've this accumulated, uh, detritus of, you know, months of, uh, of progress. Um, and so we thought that was totally untenable. In fact, like when we're developing things, we want feedback from our customers multiple times a day, uh, and so we really weren't willing to, uh, to compromise on this. And so in order to achieve five and a half nines of reliability, but with this continuous process of feature evolution development, we thought we really needed to invest in the end-to-end, uh, development and quality assurance process. And so the dev boxes are just the first part of that, and we can have some instrumentation there and observability and operational characteristics and so forth, and then a whole process of incremental deployment where you first roll it out to, you know, some a handful of machines and then, uh, uh, one percent of machines and, and, uh, some progressive, uh, rollout strategy from there. But I guess the general reason for our, you know, for our conviction, we needed extreme reliability and security, um, extreme velocity, and I think it's the only way to achieve both.

  2. 2:11 – 4:23

    600 AI-written pull requests, Minions, one revert

    1. BO

      It's incredibly impressive, the, the bar that you've been able to hit, and I, I feel like everyone in the industry looks to Stripe as the company that sets the bar for reliability.

    2. PC

      That's, um, it's very nice to hear. We obviously, we're working extremely hard to, um, to sustain that. It-it's been interesting for us in the era of agentic development, where obviously there's a question of, well, will AI be a tailwind or a headwind here? And there are obviously a lot of concerns of, well, if people are writing code so much faster, and maybe they're not reviewing each line of the code with quite as much scrutiny, uh, lots of speculation as to sort of which way this will cut. I was speaking with one engineer at Stripe this morning, and over the course of H one, um, he had more than six hundred pull requests merged. Every single one was written with AI, and exactly one of those pull requests had to be reverted, which at least suggests that it is possible to have this, you know, enormously accelerated development rate with still empirically quite, quite high reliability. Minions are this thing on top of it, uh, which are a way to orchestrate VMs, uh, with prompts from Slack or from a web interface or, I mean, in principle, from any tool. You can just ask for some feature to be implemented, uh, or some task to be completed. It will create a, a fully new VM for that, and then, you know, go and perform all the tasks, uh, along the way, a-and then it'll package it up and submit it and build it and run it through our test suite and the whole end-to-end process. We have seen that quality per pull request over that-- over the last eighteen months, um, has gone up. Now, there's still some challenge because the incidents per unit time has gone up slightly, um, and so now they're, they're mostly very minor incidents and all of our kind of secondary, uh, mechanisms for catching things before they become problematic have meant that our total reliability is, is essentially unchanged, which we, we take as quite heartening. And then I think the, the cool thing is, obviously with LLMs and AI and everything else, there's now, of course, the possibility to build new observability and new instrumentation and new harnesses and new kinds of automated scrutiny, uh, that we couldn't build before. And so I feel pretty confident that over the next year or two, that AI on net will make Stripe quite a bit more reliable.

  3. 4:23 – 6:34

    AI guardrails as infrastructure

    1. BO

      So tell, tell me a little bit more, like nuts and bolts, how do you ship faster using AI without trading off against reliability and quality? Like what, what are the specific guardrails that you have? We wanna ship a model that's really aligned, and there's a bunch of protections. There's like pump injection protection, and then also like as the number of pull requests accumulates, how, how do you make sure that ev-everything remains reliable and that code quality goes up?

    2. PC

      It does mostly come back to this idea of, um, relying on invariance and, uh, hard barriers, uh, rather than, um, things that are somewhat more subjective or discretionary or, um, or probabilistic. And so obviously it's nice if the model is aligned and is more likely to write correct code or secure code or whatever. I mean, nothing's ever perfect. You really do want to, uh, rely on guarantees.

    3. BO

      That's really interesting. So what, what I'm hearing is like before AI, there was a lot of benefits to kinda encoding guardrails and constraints as infrastructure-

    4. PC

      Yes

    5. BO

      ... because then engineers can make less mistakes, and there's also-

    6. PC

      Yes

    7. BO

      ... benefits to data segregation and making sure that-

    8. PC

      Yes

    9. BO

      ... people just literally can't access data they shouldn't access.

    10. PC

      Yes.

    11. BO

      And, and now this is kinda paying dividends.

    12. PC

      Yeah. I mean, this is full agreement with all of that. And actually, um, we started seriously investing a lot of this, and, um, the apparatus and the tagging and the sort of semantically aware guardrails around our data back in twenty seventeen. And, and it's been a multi-year journey because there's a lot of annotating to do, and there's a lot of granular permissioning to do and so forth. And we did all of this because we thought the security guarantees were, were so important. But you're right that it kind of accidentally put us in a better position when, when agentic development came along.

    13. BO

      Yeah. And then, and then luckily Claude is great at writing guardrails, so.

    14. PC

      Indeed, indeed. Yes. Well, this comes back to the point where I, I really-- I mean, there's kind of a general question in the world about- Whether AI and LLMs will be, um, on net a, um, you know, a-a offense advantaging or defense advantaging. But my guess is that in equilibrium in a couple of years, this is all going to be substantially defense advantaging.

  4. 6:34 – 9:10

    Stripe Projects with Claude Code

    1. BO

      I wanna come back a bit more to how Stripe uses Claude Code. Tell me about how, how do you use Claude? What are some of the products that you've shipped?

    2. PC

      Every dev box has Claude Code pre-installed. It's the first place people go when just, um, the, the median person is going to complete some task. It really has delivered meaningful acceleration. One case that I thought was quite informative here was we had the idea at the beginning of the year of building a thing called Stripe Projects, and this was kind of inspired by Claude Code, uh, because, uh, when you're building some project, uh, with, um, with Claude Code, for anything of any materiality, uh, or substance, you invariably need to couple that to some other set of services, right? I want PostHog for logging, or I want, uh, I wanna host it on Vercel, or I want a database, or, you know, whatever. Because almost all these companies are Stripe customers already, we thought we could, uh, work with them to expose their capabilities in a new way and make it incredibly easy for an agent to, uh, to instantiate an account with them. We had this idea at the beginning of the year. We decided to go and build it. It was, um, between two and three engineers, uh, for about two months. Um, but from first idea to public launch, and that's not only building the internal APIs and services and harnesses and all the things, um, but also integrating with now about fifty different services. Uh, and those services all have their own idiosyncrasies and, you know, occasional bugs, and I don't think you could have done that with two to three people in eight weeks, uh, in the before times.

    3. BO

      What would that have taken before?

    4. PC

      So I asked one of the engineers involved, um, and his estimate was a bigger team. He didn't say exactly much bigger, but a bigger team and six months. Let's just say twice as big a team. I d- I don't know. And, uh, I guess that would've been three X longer. So, you know, that's, uh, a six X relative change. Uh, again, maybe that, maybe that's an underestimate, who knows? Uh, in that, uh, software engineers are famously, uh, uh, you know, optimistic [chuckles] uh, when they, uh, estimate software projects. Uh, so maybe we can say at least a, uh, a six X speed up, which like, I mean, Stripe has thousands of software engineers, and who knows if every project is being accelerated by that magnitude. Uh, I- I suspect there's a distribution. Some things are massively faster, some things aren't. Um, but again, even we're really pessimistic if it's only-- let's just say it's only a two X improvement across the, uh, the sort of entirety of what we do, I mean, that, that's obviously still a, a preposterously large deal.

    5. BO

      Yeah, yeah. That's huge. So it takes less engineers, they do it faster. What do engineers do? Do they-

    6. PC

      Again, with higher quality, at least per pull request.

    7. BO

      With higher quality, maintain the reliability bar.

    8. PC

      Uh-huh.

    9. BO

      Um, so, so what do engineers

  5. 9:10 – 10:44

    AI code quality: every codebase is a prompt

    1. BO

      do? Like, you, you just do more projects? Do you try more experiments?

    2. PC

      I think we, um, we definitely create more products and try more experiments. Um, and we can, we can see this in the numbers. Like we, we just had our annual conference sessions and the number of new products and the number of new features, uh, yeah, and again, the sort of the quantum of a feature is somewhat subjectively and qualitatively defined, but, you know, we, we try to be reasonably consistent year to year. The number of new products and features at sessions this year was wildly ahead of any prior year. Um, uh, the engineering organization was somewhat larger, but clearly there was some kind of productivity effect, uh, happening there. So, so we're certainly building more for customers. But the other thing that I find kind of interesting and that cuts against some of the slop concerns is we're undertaking more projects now to improve our architecture and improve our code base. And the line, um, from a friend we often mention at Stripe is that every code base is now the prompt for another code base. Um, and you know, of course, Jared Sumner, um, a former Stripe, um, has demonstrated this, uh, with the, the bun rewrite. Um, my estimate of the-- there could be a question of, you know, with AI and with LLMs, should we expect the quality of code at Stripe to be higher or lower in three years? And I think the answer has to be higher.

    3. BO

      It, it's interesting. There's this sort of like fluidity between tokens and infrastructure. You could use the tokens to write product code, you can use the tokens to write infrastructure, you can use it to write guardrails.

    4. PC

      Mm-hmm.

    5. BO

      Um, and I think people kinda over-focus on using it to directly build product, but-

    6. PC

      That's right

    7. BO

      ... it's all these intermediary use cases that are some of the most powerful.

  6. 10:44 – 13:06

    Claude Code plan mode and 10 parallel devboxes

    1. BO

      What were some of the barriers that you hit, um, getting AI to this kind of scale internally?

    2. PC

      In general, it's been adopted, um, very enthusiastically and very completely and very quickly and so forth. So i-i-in some sense, it feels funny to ask what the barrier is given the, the adoption fervor. The phase changes were definitely, um, Claude Code itself as a modality. Um, so thank you, thank you for that. And the model improvements of late last year, uh, another, another inflection point. I think over the course of this year, I think the main, um, thing that people had to kinda wrap their minds around is getting a sense for what's possible and how one can now work. One person at Stripe, um, the way he likes to work is he will work with some, you know, the, the smartest model available to construct a very complicated, extensive plan, but like really invest a lot in the planning process, and then maybe dispatch ten different dev boxes, um, uh, with different agents executing different parts of the plan, and they might work all day or in, in the extreme case, even for multiple days, um, on, um, on implementing this. That's just a very different way of working, right? It-it's not only very different to how things worked pre-AI, but it's even-- it's kinda non-obvious, even given Claude Code that, oh, I can-- like the, the, the quantum of labor can actually be, uh, so, um, so extensive and, uh, again, if properly orchestrated, um, uh, and guided that the model is capable of, uh, such extensive self-verification. I think until recently, I myself undervalued The plan modality, and I guess goals, uh, have also made a lot of progress over the course of this year. And, you know, I've changed how I use the model. So anyway, I think there's just kind of getting intuition for the ever-changing and moving target of the modality.

    3. BO

      Yeah, I, I have this like giant graveyard of these projects that I started but never finished.

    4. PC

      Right. Yeah.

    5. BO

      And that's just like not really a thing anymore.

    6. PC

      Yeah.

    7. BO

      You kind of finish every project, and then you decide-

    8. PC

      Yes.

    9. BO

      -do I want this or not?

    10. PC

      Yes, exactly. And it's definitely the case that, um, sort of shower thoughts now get translated to actual existence, you know, at a far higher rate. I, I think the returns to being curious, uh, have gotten much higher.

    11. BO

      I

  7. 13:06 – 16:33

    Stripe's data: new startups up 2x with AI

    1. BO

      wanna talk a little bit about what you're saying on the Stripe side about how AI is changing business. One of the things I've been thinking about is with AI, it is much easier for small teams to compete against much larger teams in, in a way that you just couldn't do before.

    2. PC

      Yep.

    3. BO

      And you know, like I do these talks at Y Combinator, you know, every three months or whatever, and, and I used to ask, you know, "Who uses QuadCode? Raise your hand." And when we first released QuadCode, a few hands went up, then at some point every hand goes up. And now I ask a different question 'cause that's not a useful question anymore. I ask, "Who writes a hundred percent of their code using AI?" And, um, now, now it's roughly I think like seventy percent of people write a hundred percent, um, at these very small startups.

    4. PC

      Mm-hmm.

    5. BO

      Um, I think there's more startups than before. They're moving quick- more, more quickly than they were before. They're, they're tackling more ambitious problems.

    6. PC

      Yeah.

    7. BO

      I'm, I'm curious if you're seeing this in the data.

    8. PC

      Stripe launched fifteen years ago, so we have a decade-and-a-half longitudinal, uh, time series of firm creation. But Stripe has been fairly popular among developers for, you know, a reasonable number of years, and so I think what we see in our time series is some kind of proxy for what's happening in the ecosystem, uh, overall. And you know, for example, during March of twenty-twenty, uh, we saw a significant, uh, acceleration in, uh, in new business creation on Stripe, uh, for, for kind of obvious reasons. The reason I bring this up is because the acceleration that we've seen over the last year is, on a relative basis, much larger than any prior acceleration that we've ever seen. The number of new businesses, uh, launching on Stripe per unit time is up by roughly a factor of two. And again, as kind of a one proxy metric for coverage, uh, about a quarter of all Delaware corporations are now incorporated with Stripe. Um, and of the ones that aren't, lots of them are like random subsidiaries for a, you know, multinational or something. So I think of the true startups, we, we-- it's a, it's a, you know, an even larger fraction than twenty-five percent. So I think what we're seeing is pretty representative, and what we see is this huge acceleration. And interestingly, the acceleration is... it's quite broad-based in the sense that we see the same thing in, in most countries. And in fact, in many cases, government data hasn't even woken up to this yet. Um, and we recently published a piece in the Stripe Economics, uh, Substack about how the UK's official company statistics show a year-over-year decline in new company creation, whereas in fact, and we kind of walk through the methodology in the post, we think the UK is seeing a huge surge in, in new entrepreneurship. And what we see is that the average revenue per new business is going up, not down. So it's way more companies, and the average outcome is improving. And if you, um, partition it, uh, to look at businesses reaching a hundred K or a million or again, any kind of arbitrary revenue threshold, those are also, uh, uh, all up very substantially. Uh, and so I think it's just a, um, a, a straightforwardly better time for entrepreneurship, uh, than it was in the relatively recent past. Uh, and that's, that's pretty exciting. And I think it's also one of the concerns with AI is that, uh, it will bring about a more centralized economy. But in the microdata that we observe, we certainly see evidence for the opposite, where there is m- more new firm creation happening, those firms are more likely to succeed, and growth rates overall are, are improving. So we're, we're pretty excited.

    9. BO

      Ha- has

  8. 16:33 – 17:49

    How AI agents will pay each other on Stripe

    1. BO

      this changed the kinds of products that you're building? Has it changed how you think about your own product?

    2. PC

      We're thinking about how are all the agents and the Claude code instances, how are they gonna use Stripe directly themselves, right? Um, both the Stripe product and h- just how will they transact more broadly. And so we're just thinking a lot about a world where most transactions, agents are the counterparties on both sides. Uh, and so how does an agent sign up for Stripe? How does an agent use Stripe? Um, is MCP enough? Um, what should the Stripe CLI do? How to make sure that everything in Stripe can be orchestrated and conducted from, um, from the CLI or in some other way that is accessible to agents. How will agents pay each other? What currency will they use?

    3. BO

      It's not just agent economy, but it's like agent-to-agent economy.

    4. PC

      That's right. Yes. Yes. Again, the, the, the Stripe house view is that most transactions, uh, will be between agents within, call it three years. Um, it may not be the case that most, um, most of the dollar volume is directly between agents, but I'll be sort of surprised if there isn't just a whirling vortex of reasonably small agent-to-agent payments.

    5. BO

      Patrick, thank you so much for coming by.

    6. PC

      Thank you for having me. That was super fun.

    7. BO

      Yeah, that was great. I think I, I got to like half the questions. [chuckles] [outro jingle]

Episode duration: 17:49

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode S_lzYIvtEaQ

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.