Skip to content
ClaudeClaude

Inside Shopify's AI strategy with CTO Mikhail Parakhin

Shopify's CTO on how Shopify uses Claude: unlimited tokens for everyone, always the largest model, and a digital twin of the company. In this conversation with Boris, Shopify's CTO explains why the real value of working with Claude is in building things that weren't possible before. He describes solving a problem he'd worked on for seven years, how Shopify built a digital twin of the company, and the unique barbell strategy he recommends for getting the most out of Claude. Have a question? Let us know in the comments. Learn more about Claude Code: https://claude.com/claude-code Chapters 0:00 How Shopify's CTO raises the ceiling with Claude 2:07 Measuring AI ROI and engineering productivity 3:32 Why Shopify gives every engineer unlimited tokens 4:42 Building a digital twin to predict merchant growth 7:28 AI for managing engineering teams 9:41 Advice for CTOs: the barbell token strategy

Mikhail ParakhinguestBorishost
Oct 8, 202612mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:00 – 2:07

    How Shopify's CTO raises the ceiling with Claude

    1. MP

      You know how text is just sequence of letters and LLMs learn really well how to predict the next, and then they, they can complete the text? Human actions are very similar, but in our case, more, more importantly, companies' actions are similar. You can represent a company as a sequence of actions, and then what happens is that you can create a digital twin of the company.

    2. BO

      [on-hold music] What, what are, like, the hardest problems that you're throwing at Feeble, that you're throwing at Claude, where you've been surprised by success?

    3. MP

      I came up with a solution to a problem that for seven years I've been trying to solve, right? And I was unable to solve it, and LLMs are not able to solve it. Uh, something called wristband loss function, uh, making everything Gaussian. It's not that LLMs take away the drudge work or the work that you didn't wanna do and now they do it, and so now you can do, like, interesting things more, uh, sophisticated. It's that you can do things that you previously couldn't, like, in no amount of time, no amount of helpers. Like, I could have had thousand, you know, best mathematicians at my disposal, I still wouldn't be able to do it. But with the model, I, I can. And the, the... um, for the best problems right now I see is that I cannot solve it without the model, but the model cannot solve it without me either.

    4. BO

      Mm.

    5. MP

      It requires, you know, this back and forth and, uh, you know, creating that centaur, you know, using chess analogy, uh, that, that, that together, like, is better than the model and better than a human being.

    6. BO

      So you're raising the floor, but actually don't forget, you can also raise the ceiling.

    7. MP

      Exactly. That's, that is my point, yes. Uh, this, uh, re- and raising, raising ceiling is, is much more impactful-

    8. BO

      Mm

    9. MP

      ... 'cause, uh, raising the floor, yeah, you, you could have just been more diligent or throw more people at it. Ceiling is there's no easy way to extend, right? Like, that's, that's a real unlock.

    10. BO

      So,

  2. 2:07 – 3:32

    Measuring AI ROI and engineering productivity

    1. BO

      so when you do this, I, I, I guess for, you know, like research or some problem that hasn't been solved before, you can think about it that way. Like, this is a novel problem. The, the solution was out of reach, now we've solved it. In a, in a company context, in a business product context, how do you think about the ROI for something like this? 'Cause so- something I hear a lot is when you raise the floor, it's really easy to measure because you're like, "Okay, this amount of work would've been done by engineers. Now, you know, Claude can do it. Engineers are freed up." That's fairly easy to quantify. How do you think about the other kind of problem where you've raised the capability ceiling and now engineers are doing something they couldn't have before?

    2. MP

      It is a very good question, and I feel like at Shopify we're... I'm just gonna go ahead and say it, like we're probably best in the world right now at measuring how, how high is the floor. Uh, we, we have not only, of course, just like everybody else, like estimate nu- uh, number of PRs and complexity of PRs, but we, we actually do things like estimating complexity of the project and in our, uh, system, you know, GSD system, Get Stuff Done, we could, we could see that, hey, the teams of people are being, uh, able to execute on more projects faster and, uh, normalized for complexity. You can be very precise at measuring percentages of, of productivity you're gaining.

  3. 3:32 – 4:42

    Why Shopify gives every engineer unlimited tokens

    1. BO

      So how do, how do you have this conversation with, you know, with Toby? And, like, what's your advice to, you know, another CTO at a different company that has to have the conversation with their CEO about this?

    2. MP

      Always use the largest model. Like, you... because the largest model raises the ceiling the most. First of all, you don't know unless you use the large model and then, and compare the results, you don't know what you're missing, right? It's a unknown unknown, uh, and that, and that ta-takes time. Second, always throw maximum, uh, thing at a problem, and then until I see the evidence otherwise, that's what we're gonna be doing. And that's why, you know, at Shopify we have unlimited tokens for everybody. That's why we run the heaviest models we can find. Uh, that's why we, we throw a lot of, a lot of resources in our both, uh, GPU in-inference and training and ability to optimize.

    3. BO

      So, so in a way it's like, it's like you just have to y- you have to make sure that everyone that's making the decisions is actually using the model and using the product so, so they understand what this is. And then you have to just keep pushing the ceiling and seeing those returns, and then it's just obvious if it works or it doesn't work. So what's an example of one of these projects you shipped

  4. 4:42 – 7:28

    Building a digital twin to predict merchant growth

    1. BO

      that, that raised the ceiling in a obvious way?

    2. MP

      Well, my-- the one that I'm probably I'm most proud of is, is our HSTU prediction system. So the idea is basically, you know how text is just sequence of letters and LLMs learn really well how to predict the next, and then they, they can complete the text? Human actions are very similar, but in our case, more, more importantly, companies' actions are similar. You can represent a company as a sequence of actions. The company's, I don't know, open credit line, ship a product, like add, make changes, like start accepting credit cards, take a loan, and I mean, these are large changes, but all the immediate things that, that company, all of its employees, you know, in every second are doing, it ca- can all be turned into a sequence of actions. And then what happens is that you can create a digital twin of the company, and then you can start doing things to it instead of to real company, right? [chuckles] And then see what happens. So you can have, uh, counterfactual interventions. You could say, "Oh, what if we edit advertising campaign? Or what if they started shipping, you know, one day faster? Or what if, you know, it's all about our merchants. What if we give them a loan?" And then we find sequence of interactions that would maximize their chances of growth. Uh, and then we, we actually do that in production, right? So, so we, we go offer them a loan or we, you know, uh, uh, pu-push them to send them communication like for, for folks like your most important thing is to, uh, to start shipping faster or we, uh, start advertising campaign on, I don't know, Twitter. So, uh, having ability to play with a company without endangering anything, you know, real people is, is huge unlock because now, now you can, you can, uh, help more people to, to be successful. So this is very clear unlock that wouldn't have happened and that has like very direct impact on our bottom line and then the financial results.

    3. BO

      And, and, and just building this, like let's say you could have built this, how long would this have taken in the past?

    4. MP

      But this is the thing. I don't think it was possible in the past. Like, just think about it, like you would need to get the concentration of talent. You would have to bring that talent in, and then you have to have, have this premonition that you need to do that, and then you, you need to start getting, collecting data and all, all this. I don't think, I don't think that was humanly feasible previously. And so the, the, the trick here is the answer is like previously it would take infinity. We just wouldn't have done it. And then that's the def-- my definition of raising the ceiling, right? Like you start doing something that pre-previously was not even in consideration set.

  5. 7:28 – 9:41

    AI for managing engineering teams

    1. BO

      What's something that you changed your mind about because of, because of Claude, because of the model?

    2. MP

      I thought that LLMs cannot really help me manage the team better. I could not have been more wrong and after a while I realized, uh, I, I flipped so much I, I created a little system that actually analyzes what everything is happening and proactively tells me it looks like this project is gonna start slipping their deadline. It allows me to detect issues before they actually happen and go resolve it and so the whole system starts functioning much better.

    3. BO

      Mm. Has, has the role of a EM at, at Shopify changed?

    4. MP

      Now, technical side is even more important because the simple things LLM will go caught up. But you know, LLM kind of is very good at picking up signals and it tries to please. And so whatever you ask, the answer is very often in the question and it will just elucidate that and like give you exactly what you wanted, right? So it's really important to want what's actually needed, not put in words first thing that comes to mind because then, then your consequences could be actually worse than before because before at least it wouldn't be immediately built. There would be some discussion and now it's so much easier to just go and build it. So technical acumen becomes more important and the human component of supervising now engineers who are supervising LLMs rather than supervising engineers who are typing, you know, pressing buttons. That's a very different thing. That's-- it's almost like you get second-- become second level manager because now you, you want to share, hey, what are the best, best ways to dispatch the work to different LLMs and how you're gonna not forget what they're doing and then how you're gonna avoid conflicts between them. I think, uh, that's why I wouldn't say it's dramatically different in, in essence, but I, I say the accent has became much more on those two things, like keeping the team more efficient and, um, making them more, uh, making technical decisions, uh, and less on checking day-to-day, "Hey, has, has that thing been done?"

  6. 9:41 – 12:18

    Advice for CTOs: the barbell token strategy

    1. BO

      Um, okay, I wanna end on this. What is your advice to other CTOs, other technical leaders that are also trying to figure out how do I adopt Claude? How do I use it? How do I scale it out to my org? How do I get my org proficient?

    2. MP

      My main advice would be try to not limit tokens. You always need like circuit breakers because there are always runaway processes that suddenly start burning tokens and, and, uh, doing things bad. But it might not look like there's big difference between big models and small models. There's a hu- world of difference. And you want it-- you don't know the counterfactual, so you, you won't know which one, which one's good. So always use this dichotomy barbell strategy for anything coding, for anything, uh, development, testing, auto research. Use the largest model, budget for it, eat that cost. For everything production, use fine-tuning, use the right model for that production, auto research the hell out of it so that you balance both, uh, cost throughput and, and quality. And save money on production by using cheaper models, uh, uh, often fine-tuned and invest that money in development side on the, on the place where there's huge potential in upgrading your model to, you know, in Fable 5 I was on the record to say that it's, it was such a huge step, like much, much bigger than like all the previous steps. You know? I've seen repeatedly people doing the opposite and trying to save money on coding tokens on, on Claude on development and then at the same time wastefully running, uh, models larger than they need in, in the actual production services. I feel like I will out-compete anybody who is not on that strategy. I was sharing with, uh, my friends here that, uh, Boris destroyed my, uh, my will to live really and my... I kept loudly proclaiming that I will end up, uh, deciphering, uh, Linear A which is a like famous writing was used on Crete by Minoans, uh, uh, in Bronze Age, kind of personal area of interest of mine and then, and then Boris tweeted that somebody did it using Claude and I'm like, "Damn it. That was my whole spiel like I had. That was my retirement project I was going to work on." [laughs]

    3. BO

      You just didn't finish in time. [laughs]

    4. MP

      [laughs] Exactly. Exactly.

    5. BO

      Thanks so much for stopping by.

    6. MP

      Thank you. [outro jingle]

Episode duration: 12:18

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode 9GtcbnmldcY

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.