Skip to content
Aakash GuptaAakash Gupta

The ONE AI Skill Every Product Manager NEEDS in 2026

Today, we’ve got some of our most requested guests yet: Hamel Husain and Shreya Shankar, creators of the world’s best AI Evals cohort. You’ll learn everything you need to know about AI Evals, how to build them, common mistakes to avoid, and much more! If I were you, I’d stop everything right away and binge watch it right now & make my action plan to execute tomorrow. Also, we’ve also done a Newsletter deep dive with them, check it out - AI Evals: Everything You Need to Know to Start: https://www.news.aakashg.com/p/ai-evals 🎥 Timestamps: Preview - 00:00 Three reasons PMs NEED evals - 02:06 Why PMs shouldn't view evals as monotonous - 04:40 Are evals the hardest part of AI products solved? - 06:23 Why can't you just rely on human "vibe checks"? - 07:37 Ads - 12:11 Are LLMs good at 1-5 ratings? - 14:06 The "Whack-a-mole" analogy without evals - 15:45 Hallucination problem in emails (Apollo story) - 16:26 How Airbnb used machine learning models? - 21:22 Evaluating RAG Systems - 23:56 Ads - 29:52 Hill Climbing - 31:42 Red flag: Suspiciously high eval metrics - 35:51 Design principles for effective evals - 39:02 How OpenAI approaches evals - 42:42 Foundation models are trained on "average taste" - 44:39 Cons of fine-tuning - 49:36 Prompt engineering vs. RAG vs. Fine-tuning - 51:27 Introduction of "The Three Gulfs" framework - 53:00 Roadmap for learning AI evals - 56:04 Why error analysis is critical for LLMs - 01:01:41 Using LLM as a judge - 01:08:29 Frameworks for systematic problem-solving in labels - 01:10:15 Importance of niche and qualifying clients (Pro tips) - 01:17:42 $800K for first course cohort! - 01:18:43 Why end a successful cohort? - 01:20:15 GOLD advice for creating a successful course - 01:25:49 Outro - 01:33:39 ---- Podcast transcript: https://www.news.aakashg.com/p/hamel-shreya-podcast 💼 Check out our sponsors: 1. The AI Evals Course for PMs & Engineers :Get $800 off with this link - https://maven.com/parlance-labs/evals?promoCode=ag-product-growth 2. Jira Product Discovery: Plan with purpose, ship with confidence - https://www.atlassian.com/software/jira/product-discovery 3. Vanta: Automate compliance, security, and trust with AI (Get $1,000 with our link) - https://www.vanta.com/lp/demo-1k?utm_campaign=1k_offer&utm_source=product-growth&utm_medium=podcast 4. Product Faculty: Get $500 off the AI PM certification with code AAKASH25 - https://maven.com/product-faculty/ai-product-management-certification?promoCode=AAKASH25 👀 Where to Find Hamel & Shreya Hamel’s LinkedIn: https://www.linkedin.com/in/hamelhusain/ Shreya’s LinkedIn: https://www.linkedin.com/in/shrshnk/ 👨‍💻 Where to find Aakash: Twitter: https://www.twitter.com/aakashg0 LinkedIn: https://www.linkedin.com/in/aagupta/ Instagram: https://www.instagram.com/aakashg0/ 🔑 Key Takeaways: 1. Stop Guessing. Eval Your AI. Your AI isn’t an MVP without robust evaluations. Build in judgment — or you’re just shipping hope. Without evaluation, AI performance is a happy accident. 2. Error Analysis = Your Superpower. General metrics won’t save you. You need to understand why your AI messed up. Only then can you fix it — not just wish it worked better. 3. 99% Accuracy is a LIE. Suspiciously high metrics usually mean your evaluation setup is broken. Real-world AI is never perfect. If your evals say otherwise, they’re flawed. 4. Fine-Tuning is a Trap (Mostly). Fine-tuning is expensive, brittle, and often unnecessary. Start with smarter prompts and RAG. Only fine-tune if you must. 5. Your Data’s Wild. Understand It. You can’t eyeball everything. Without structured evaluation, you’ll drown in noise and never find patterns or fixes that matter. 6. Models Fail to Generalize. Always. Your AI will break on new data. Don’t blame it. Adapt it. Use RAG, upgrade inputs, and stop expecting out-of-the-box magic. 7. Your Prompts Are S**T. If your AI is bad, it’s probably your fault. The cheapest, most powerful fix? Sharpen your prompts. Clearer instructions = smarter AI. 8. Let AI Teach You. Seriously. LLM judges aren’t just scoring you — they can teach you. Reviewing how your AI fails is the best way to learn what great outputs should look like. #ai #aievals #aiproducts #aiprompt 🧠 About Product Growth: The world's largest podcast focused solely on product + growth, with over 175K listeners. Hosted by Aakash Gupta, who spent 16 years in PM, rising to VP of product, this 2x/ week show covers product and growth topics in depth. 🔔 Subscribe and like the video to support our content! And turn on the bell for notifications.

Aakash GuptahostHamel HusainguestShreya Shankarguest
Jul 11, 20251h 34mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:005:25

    Why AI evals are a must-have PM skill: taste, iteration, and scale

    1. AG

      Why do PMs need to be good at AI Evals?

    2. HH

      Okay, so there's three things that are really important. One, evals give you a way as a PM to inject your taste and your judgment directly into the critical path of the AI product being developed. The second thing is, like, evals are really important in helping you iterate. The most effective way to do that is using evals, specifically looking at data in a very structured way. And then the third thing is scale. By mastering evals, what you can do is you can make sure that you can scale your taste, judgment, uh, so on and so forth, user requirements across all the AI workloads that are running.

    3. AG

      When it comes to AI Evals, Hamel Husain and Shreya Shankar are known as the worldwide leading experts. Companies like OpenAI and Arize go to them, and today we're gonna learn everything you need to know about evals from them. What is the most critical skill for PMs who want to build AI features to develop?

    4. SS

      Hands down, error analysis, the ability to look at your outputs and systematically figure out what makes for a bad output, quantify how many of these failure modes you see in a big batch of traces for your system, and then figure out how to turn that in measurement into a continuous flywheel of improving your product.

    5. AG

      If you guys had to build a roadmap for people who wanted to get really deep on AI Evals, what topics should they learn? Really quickly, I think a crazy stat is that more than 50% of you listening are not subscribed. If you can subscribe on YouTube, follow on Apple or Spotify podcasts, my commitment to you is that we'll continue to make this content better and better. And now on to today's episode. Hamel Husain and Shreya Shankar are the people who the experts go to for evals, OpenAI, Arize AI. Those people are going to them for evals, and we have them on the podcast today. Welcome, Shreya and Hamel.

    6. SS

      Hey.

    7. HH

      Thank you. Nice to be here.

    8. AG

      Why do PMs need to be good at AI Evals?

    9. HH

      Okay, so there's three things, um, that are really important. One, uh, evals give you a way as a PM to, you know, inject your taste and your judgment directly into the critical path of the AI product being de- uh, developed. So, like, you know, as we all know, like, PMs, they spend a lot of time gaining context from customers, user feedback, so on and so forth. They're writing PRDs. They're, you know, trying to give context to engineers and, you know, they're hoping, like, kind of engineers are faithfully carrying out their vision. Now, what evals give you is y- you know, you can directly make sure that your taste and all of that context, if done correctly, is now on the critical path when your engineering team is developing those AI products. The second thing is, like, evals are really important in helping you iterate. So, you know, nothing is, like, set in stone. You have to constantly, like, change your requirements. You're learning more about your customers, so on and so forth. The most effective way to do that is, uh, you know, using evals, specifically looking at data in a very structured way, and which this is one of the things that Shreya and I teach. Um, you know, that allows you to, to refine and have really fast feedback loops and really fast cycles of feedback. And then the third thing is scale. So, you know, by mastering evals, what you can do is you can make sure that you can scale your taste, judgment, uh, so on and so forth, user requirements across, you know, all the AI workloads that are running in a way that you just couldn't before, because ultimately there's a lot of... You know, you can bake a lot of these evals. They're using AI themselves. You just have to make sure that you do it correctly. So you have to make sure that you, uh, align the AI with yourself as a PM in a very kind of, uh, process that we teach. Um, and as long as you do that correctly, and you do it in such a way that you develop trust in the AI that is doing the eval, and there's a way to do that, that alignment, then you can really scale yourself. So a lot of times, PMs, or not just PMs, but people at large, they kind of view evals as a very monotonous task that, you know, you just want someone else to do it. It's like, "Oh, like, I have to look at data. I have to annotate data. Um, you know, who's gonna do this?" You don't wanna give up that leverage because when you, when you build that foundation of evals, um, you develop... You have immense leverage, and you can... You know, it's a really quick way to kind of exert lots of influence over the process in, in a good way. And so this is why I would encourage PMs to really pay attention to this.

  2. 5:257:37

    Defining “evals” and why they’re the hardest part of AI products

    1. AG

      Can you guys precisely define evals?

    2. SS

      Yeah, I can take this one. An eval is some systematic measurement of some aspect of quality. So what varies in an eval is what that criterion is. For example, maybe it's conciseness of a response, and then how you want to measure it. So maybe that is, I'm gonna define it by, you know, word length. I'm gonna define it by sentence length. Um, maybe it is some, you know, very, very complex bespoke human judgment or something that's more subjective. Um, but those two things make up an ema- eval.And oftentimes, products actually have a suite of evals. I've never seen just one eval doing the job. I see three to five, sometimes even up to 10 evals that are really important for a product.

    3. AG

      People say that if you get evals right, you've gotten the hardest part of the AI product solved. Is that accurate?

    4. SS

      I think it's accurate now. Hamel, what do you think?

    5. HH

      I think it's totally accurate.

    6. SS

      Yeah.

    7. HH

      Just like anything else, it's the process of creating the evals that provides all the value. It's not necessarily the eval itself. It's the journey that creates all the value. And so once you've done all of that work, you've looked at all of your data, you've iterated on your system, you've thought very carefully and s- and, you know, oftentimes scientifically about how to improve your system, you've already got 99% of the way there.

    8. SS

      The way that I like to think about it is if you ever want your product to make it past one iteration, you need evals. I've never seen somebody make it through multiple iterations of their product without any evals. But once you have good evals in place, then evals are not necessarily the bottleneck for you. But that's a good thing. That's how it should be, right? You should be able to focus on building out other aspects of the product, making things faster, making things feel better, more intuitive, um, you know, everything beyond that.

  3. 7:378:53

    Why “vibe checks” don’t scale—and turning judgment into rubrics

    1. AG

      Why can't you just rely on, like, human evals? Like, the PM looks at the feature, the engineers look at the feature, and they feel like those outputs are good enough.

    2. SS

      Oh, I love this question on the vibe checks and why... So, so Hamel and I teach our course and pitch it in a way that we, we are helping you codify, operationalize, and scale up your vibe checks. Your vibe checks are very important, but they don't scale, right, 'cause they involve you, the human. It's very hard to onboard other people to do the vibe checks in the same way as you are. So, like, I would have to observe you do this thousands of times, look at outputs, try to build my own rubric or mental model of what you're doing, and then I have no good way of teaching other people of how to do this. So being able to do evals just means taking your vibe checks and translating them to something concrete. In our course, we define that as a rubric of binary criteria. Every criteria could be complex, that's fine, could be subjective, that's fine, but you better have a very precise definition for pass, fail, have some examples of pass, have some examples of fail, and we also teach people ways to measure alignment on those results. Um, that's really what this whole process is about.

  4. 8:5314:11

    Binary evals over 1–5 ratings: clarity, calibration, and LLM-judge reality

    1. AG

      I think the critical phrase there is binary criteria. Why binary?

    2. HH

      Yeah, so binary really is a kind of a heuristic in a w- in a sense, like, that is, like, a simplification for, that works for most people. And the, the thing is, like, you know, a lot of people try to, like, assign scores, let's say, on a rating scale of one to five. That's usually a really bad idea because no one knows what that means. If you have a average score of 3.2 versus an average score of 3.7, what does that really mean? And, you know, that can be very hard to calibrate, and you have to work incredibly hard to, you know, make sense of that. So binary judgments force you to kind of make a pass-fail decision, and that tends to also correlate with the fact that you're going to have to ship this product. Do you wanna ship it or not? And it really distills that decision-making down into the annotation. And for the vast majority of people, that's the right choice.

    3. SS

      Yeah, and to provide a little bit extra context on the background of LLM as judge and why, you know, people have a lot of variance in whether they want it to be, you know, binary or rating-based scale, um, LLM as a judge has been around, you know, before these foundation models even, just regular language models, fine-tuning models to serve as judges. Um, and in those cases, people, A, had a lot of preference data of what is good and bad, and maybe even a fine grain scale, and B, could fine-tune models to be aligned with that preference data. Today's world of LLM judge is very different. We don't see people fine-tuning judge models as much. We see people trying to use off-the-shelf models, still want to align with their complex subjective criteria, and now the alignment problem is much harder, right? You can't, you know, steer the LLM in a way that you could before. Um, and for that reason, we say limit yourself to binary because that is what the LLM can do very well. All you have to do is provide examples of pass, p- provide examples of fail, and have, you know, very simple or, like, good rubrics. Um, and people find that much easier to do than, say, you know, rating on a scale of one to five. Okay, now you need to provide examples for one, for two, for three, for four, for five. You need to, you know, have descriptions of what makes a one different from a two. All of these things, you know, the pairwise interactions between all these ratings just explode in complexity, and we never see people successfully able to operationalize that at the rate at which they can do binary evals.

    4. HH

      And a lot of times, um, non-binary evals, like ratings of one to five, that is a s- a smell of intellectual laziness.

    5. SS

      [laughs]

    6. HH

      Like, the work hasn't been done-

    7. SS

      Contic

    8. HH

      ... to actually, you know, to make a call of, like, what is good enough and what's not good enough. And it's kind of like, "Oh, we don't really know. Let's just capture these, like, rough things, um, you know, in, in this, like, s- uh, score 'cause we're gonna lose something." And it doesn't... You know, like, the binary scale a lot, like, really forces you to be very clear about what you want.

    9. AG

      AI Evals are one of the most important skills for PMs, and I know you know they matter. The question is, are you doing them right? Most teams are winging it with basic metrics and hoping for the best. Meanwhile, the teams that actually ship reliable AI, they've cracked the code on systematic evaluation. Today's episode is brought to you by the AI Evals for Engineers and PMs course by Hamel Husain and Shreya Shankar. This live Maven course will teach you the battle-tested frameworks from Hamel and Shreya, who are the engineers behind GitHub Copilot's evaluation system and 25-plus production AI implementations. Four weeks, live instruction, next cohort starts July 21st. Start shipping AI that actually works. Enroll at maven.com with my code AG-PRODUCT-GROWTH for over $800 off. That's AG-PRODUCT-GROWTH. Today's episode is brought to you by Jira Product Discovery. If you're like most product managers, you're probably in Jira, tracking tickets and managing the backlog. But what about everything that happens before delivery? Jira Product Discovery helps you move your discovery, prioritization, and even roadmapping work out of spreadsheets and into a purpose-built tool designed for product teams. Capture insights, prioritize what matters, and create roadmaps you can easily tailor for any audience. And because it's built to work with Jira, everything stays connected from idea to delivery. Used by product teams at Canva, Deliveroo, and even The Economist, check out why and try it for free today at atlassian.com/product-discovery. That's A-T-L-A-S-S-I-A-N.com/product-discovery. Jira Product Discovery, build the right thing. I've heard that LLMs are also not very good at one to five ratings.

  5. 14:1116:26

    Skepticism + scientific method: avoiding whack-a-mole iteration

    1. AG

      Is that true?

    2. SS

      They're good at what they're trained on. [laughs] Somewhere out there in the world, I am sure there is a task with very clear or simple one to five ratings, and the LLM is good for that. But to make such a blanket statement for all products and all use cases is very hard to do. And that's the thing, that's the message we want to hammer home to every single product manager who takes the course. Like, look, you think the LLM might be able to do something. You saw an instance of an LLM being able to do the task for some other domain. That doesn't mean it's gonna translate to your domain or your use case. You still have to put in this work, um, and just don't trust any... Hamel has a great way of saying this. Maybe he should talk about it, but he always tells people, "Never trust it. Always put on your detective hat." Hamel, you wanna talk about that?

    3. HH

      Yeah. What underlines the entire process of evals is the scientific method, something that we've all learned in high school education, um, but it's really applied, you know, in this context. And what you have to do is be very skeptical of everything and do lots of experiments, and prove to yourself that the, the thing that you're trying to achieve or some new complexity you wanna add, whatever it is, that it's actually working, and try to do it in the simplest way. Um, and build intuition doing, by doing lots of experiments. Um, but the point is to, like, measure those and, like, you know, record those and go through it in a structured way, uh, rather than those vibe checks. You, you, you asked about vibe checks earlier. The analogy that I like to use, and I can give this to you, uh, if you ask me later, there's a, uh, there's a little video of my friend Greg Secarelli playing Whack-a-mole, and it's my favorite meme to use when telling people about the need for evals. It's, you know, it's like you're playing Whack-a-mole, uh, without evals, you know? So you see a problem, okay, hammer it over with some tool or a prompt change. Then another problem comes up, you hammer that, and you keep going. You don't really make any progress. It's really with evals that you can systematically try to solve the problem without, like, going in circles.

  6. 16:2619:34

    Case study: hallucinated sales emails—and why domain-specific evals work

    1. AG

      I want to talk about some stories. So one of the features I implemented in my last job, I was VP of product at Apollo.io. It's a unicorn startup that does sales technology. So what do salespeople need to do, right? They need to write emails. [laughs] So as soon as, uh, I think it was ChatGPT 3.5 came out, we're like, "Okay, we're gonna use GPT 3.5 to write people's emails." But the very first thing we found was that it was hallucinating all sorts of crazy details, and ultimately, we had to set up a bunch of evals instead of vibe checks. Can you guys give a little bit more context into why evals helped solve that Hallucination problem for us?

    2. HH

      Yeah. So with any kind of problem that you see, um, you know, hallucination or whatever it is, so the first thing to know about evals, where people go off the rails, is do not reach for generic metrics. So the industry is full of tools and vendors wanting to sell you, like, a magic pill to solve your evals problem. Like, "Hey, don't worry. Just buy our tool, plug it in. We'll show you a dashboard of all of these metrics and things like that." The problem is it doesn't work because those off-the-shelf evals' hallucination score is not going to work. What you need to do is take a look at, in your case, those emails and understand, like, what exactly is the failure mode, and what... If you do, uh, observe a hallucination, what is the hallucination? And kind of give more life to the domain specificity of that hallucination so that you can then start crafting an LLM as a judge that is prompted in such a way that is very specific to the types of hallucination that you are seeing.Um, and then you go to, through an iterative process. So you kind of hand label when hallucinations happen, you know, how they're happening. You, you're building an LLM as a judge, and you're measuring that judge. You're being skeptical again through the scientific process, and you're saying, "Okay, I have this judge. Can I get it to agree with me? Can I have it be..." Let's say if, if it's you, Aakash, making this judge, this email hallucination thing, um, you know, how do I make this judge a proxy of me, and how do I trust it almost like an employee? And so the only way to do that is to check it, and you can do that iteratively through this process. Um, and when you do that, then you not only do you have something that scales, it's an automated way of checking that problem, but it's also something that you trust. And that second part, that something that you trust, is key because the last thing you want is the whole bunch of evals that you put up on a dashboard, and then people stop looking at them because they're like, "Uh, we have these evals, but you know what? The product doesn't really work." No. That, that's the death of your AI product because then-

    3. AG

      Mm-hmm

    4. HH

      ... no one's gonna ever look at evals again, and then you don't have any leverage.

  7. 19:3422:19

    From classic ML to LLMs: what Airbnb taught about evaluating stochastic systems

    1. AG

      That's exactly [laughs] what we experienced, and you weren't even there. I wanna talk a little bit about your experiences. I know you worked at Airbnb on these products. Can you tell us a little bit more about that, and what was the most difficult part of building evals there?

    2. HH

      So I didn't work on LLMs at Airbnb because it was prior to, you know, way before ChatGPT. Um, but what I did work on my entire career is machine learning and, um, you know, building predictive models. And a lot of the same machinery of evals comes directly from machine learning. Um, the reason that is is because machine learning systems are systems that produce stochastic outputs. You know, they'll give you predictions of various kinds, um, or, like, classify things, and they're, they're non-deterministic, and you have to evaluate them. You know, like, you know, they're g- giving you different outputs every time, and they have noise. And, um, so how do you do that? That's something that's very well established in machine learning that, you know, a lot of people haven't been exposed to. So when it comes to, like, AI more generally, now you have really a very similar thing. It's like, hey, you have, like, a stochastic system. You know, it's non-deterministic. It can output anything, like these emails. It's like, how do you go about measuring that? And, um, yeah, and so basically you can kind of bring that over.

    3. AG

      Yeah.

    4. HH

      And so what we teach is instead of going through the entire data science machine learning curriculum, how do you have a very focused way of learning that, that is contextualized to LLMs?

    5. AG

      That's actually fascinating. How was Airbnb using machine learning models? I'm sure people wanna know under the hood.

    6. HH

      Yeah. Airbnb was using it for a lot of things such as detecting fraud in payments. Um, also the biggest use case at Airbnb was search ranking. So if you're searching for a listing, let me show you other listings that you might be interested in based upon what you've been looking at and what you're searching for. So basic recommendation systems, search ranking, things like that. Um, a lot of, like, growth marketing initiatives, like trying to figure out the lifetime value of a specific guest so you can allocate marketing to them appropriately. That's what I worked on. Um, you know, and there's many other use cases, but those are the ones that sort of one of, like, they're bigger ones. And now they have, you know, now Airbnb's doing generative AI stuff, uh, as well, just like any other company.

  8. 22:1925:22

    Evaluating search and RAG: changing metrics when the consumer is an LLM

    1. AG

      Search ranking is something that probably we've been dealing with, right, for, like, 30-plus years, I guess, ever since, you know, these search engines ever came out. So what can we learn from how people are evaluating search and apply into how we're evaluating our LLMs?

    2. SS

      The number one difference now is we definitely search systems to try to get context to improve our LLMs, sure, but the consumer of the search results is now the LLM. In the past, the consumer of search has been the human, and humans and LLMs are good at different things. LLMs are good at finding needle in a haystack in very, very long, complex windows. Humans have very short attention spans. They'll read through things, but after a few paragraphs, they're done. So some of the metrics, like how high up a relevant result was ranked, are much more important for humans than they are for LLMs. You know, you could try to retrieve, like, 500 results, and as long as it's, you know, even if it's, like, 150, 200, the LLM will pick it out and figure out how to give that result to the hu- the user. So I think the bottom line differences are, you know, we still wanna use the same metrics, but our tolerance has changed slightly. We still wanna prioritize recalling the right information, but now that we have long context windows, it's okay for it to kind of be at the end, um, as long as we carefully, you know, leverage LLMs to go and, uh, iteratively refine those search results, pick out the bottom results, and then go and show that back to the human.

    3. HH

      And one thing I'll add onto this is people often wonder, "Okay, how do I evaluate RAG systems?" So RAG is a big thing. And so what i- what is RAG? You know, Retrieval-Augmented Generation. There's the retrieval part, and then there's the generation part. And so the retrieval part, you evaluate that pretty much the exact same way that you would evaluate any search system. So all those classic search systemsUm, and informational retrieval, like that entire scientific discipline can be applied onto the retrieval part. So like optimizing that, making sure you're getting the right documents, the right context, so on and so forth. Um, and a lot, and a lot of it applies, and Shreya's right, there's a lot, there's some nuance in terms of different tolerances and things like that.

    4. SS

      Yeah. The, the number o- one thing I see that's different is instead of measuring like recall at 10, like we can measure recall at 500 and have a long context model like Gemini be able to consume those results. Uh, at the end of the day, maybe you wanna measure precision for the result of the LLM call that feeds in the result to the human. Sure, whatever humans consume, they should all be measured the same way. But if LLMs are there refining search results, then we can use that to our advantage to be able to re-rank outputs from a search engine.

  9. 25:2231:25

    Code generation evals: verifiable domains, test harnesses, and why dev tools lead

    1. AG

      So that's search. One of the other sort of really big things in the AI space that I think, Hamel, you worked on a little bit, was the precursor to GitHub Copilot, right? LLMs for code generation. What was the biggest problem you guys faced with evals for that?

    2. HH

      So with... Okay, things like code generation are actually really interesting. So in evals, there's gen- more, more generally, in AI products more generally, is you want your domain expert in, inside the inner loop. So what you don't want to do is give your developer the, the task of annotating your data and, you know, writing evals necessarily, because they don't know enough. They don't have enough context. And that really bottlenecks a lot of teams because they don't get that right. But there's an exception. The exception is developer tools. That is one case where the domain expert is the developer. And so that's why we saw developer tools like as the first, you know, AI products, because that was the sweet spot. That was like where it was easiest to develop. That's one property of, you know, developer tools. The second property is it's a highly verifiable domain, so there's some structure to code, and it turns out, like on GitHub, um, you know, you have all this code, and you also have lots of tests defined against that code. So a lot of time was spent in that verifiable domain, and it's excellent when you have a verifiable domain, um, is to develop kind of a test harness, and the test harm- harness was very impressive. Basically, what it did is it took all of the code at scale, not all of it, but a select, like filtered quantity of it, and basically recreated the environment that code is gonna run in and ran the test at scale. And basically, it, you know, things like asking the LLM to, you know, fix certain things or complete certain code, and it would run all the tests. It's kind of a very impressive engineering effort because you're talking about running all this random software from everywhere with all kinds of dependencies at scale all the time. Um, so the, the, the details don't matter. What I wanna say is like the evals really mattered because a lot of upfront work went into constructing the eval system, and it's really after that that the team was able to iterate really fast. Like when, you know, G- GitHub Copilot was first released internally, it, it didn't really work. You know, it was something like, you know, 20% or so of, you know, uh, like, uh, acceptance rates of like suggestions, things like that. And then after the evals, the team was able to like iterate really fast and like climb that, uh, you know, to where it did work. Um, and so, you know, it was really like... Or let me step back. It w- it was, yeah, it was, it was really like the key that unlocked, uh, you know, progress on that. Um, you know, I didn't work directly on GitHub Copilot, like, you know, in that phase. I was kind of working on some research before that. Um, I was actually working on some research that led to some of the benchmarks. So, um, basically one thing I did at GitHub is, uh, you know, I worked on this project called CodeSearchNet, which is a semantic search of code, and basically, like you type in like what you're, what code you're looking for and like semantically try to find it, and it leveraged a lot of the, um... There's, so a lot of code has comments in it, in the documentation embedded in that code, so it was just a really large dataset. We opened a large dataset with benchmarks of, uh, you know, uh, code retrieval. And so that was like used by OpenAI in their very first iteration of Codex, which is like a very, uh, old model. It's not, not the current Codex. This is like a different, the same name, but different, different thing. Um, but it was, it was an eval. It was an eval that was used back in the day. So you can say I've like been working on LLM evals for a really long time.

    3. AG

      Trust isn't just earned, it's demanded. Whether you're a startup founder navigating your first audit or a seasoned professional scaling your GRC program, proving your commitment to security has never been more critical or more complex. That's where Vanta comes in. Businesses use Vanta to establish trust by automating compliance needs across over 35 frameworks like SOC 2 and ISO 27001.Centralize security workflows, complete questionnaires up to five times faster, and proactively manage vendor risk. Vanta can help you start or scale your security program by connecting you with auditors and experts to conduct your audit and set up your security program quickly. Plus, with automation and AI throughout the platform, Vanta gives you time back so you can focus on building your company. Join over 9,000 global companies like Atlassian, Quora, and Factory who use Vanta to manage risk and improve security in real time. For a limited time, my listeners get $1,000 off Vanta at vanta.com/aakash. That's V-A-N-T-A.com/A-A-K-A-S-H for $1,000 off. Today's episode is brought to you by the AI PM certification on Maven. Run by Miqdad Jaffer, who is a product leader at OpenAI, this is not your typical course. It's eight weeks of live cohort-based learning with a leader at one of the top companies in tech. OpenAI just doesn't stop shipping, and this is your chance to learn how. Run along with product faculty and Mo Ali, the course has a 4.9 rating with 133 reviews. Former students come from companies like OpenAI, Shopify, Stripe, Google, and Meta. The

  10. 31:2535:12

    Hill climbing needs PM-defined metrics—and the risk of overfitting

    1. AG

      best part? Your company can probably cover the cost. So if you want to get $500 off, use my code AAKASH25 and head to maven.com/product-faculty. That's M-A-V-E-N.com/P-R-O-D-U-C-T-F-A-C-U-L-T-Y. You mentioned hill climbing, and I think that's a really interesting concept because I was reading Daniel McKinnon's piece on evals. He is a PM on Meta's Llama team, and he talked about how it's so important for the PM to define the evals so that then the AI engineers and research teams that he works with can hill climb. Can you guys explain that?

    2. SS

      Engineers are really good at hill climbing, and especially ML, AI, data science, stats people, people who are trained to go and look at, you know, ML metrics, figure out why they're bad, and how to improve them. They're good at that in structured data and traditional machine learning because those metrics are very well defined, like accuracy is well defined, loss is well defined. If you're, like, doing a binary classifier, and you already have labeled data, you know, it's all there for you. You're just doing your job and trying to improve on those metrics. Now we're in a role in which none of the metrics are defined, and so there's no way to go and exercise all the skills that the AI engineers have to improve those metrics. I think it is so spot on that this article talks about, you know, PMs need to set these metrics because they have the best context to set them. And once you set them, once you provide the definition or, you know, examples of good and bad, then engineers will figure out a way to encode this into the product and improve, and even get to self-impruding, improving products. Like, that is exactly what I see also in my experience.

    3. HH

      One thing you have to be careful with, with hill climbing, and when that term is said, it does tickle my brain, is make sure that people don't overfit. 'Cause, you know, engineers, sometimes they use that term, and they are overfitting in some sense. Um, us- if they're not, if, you know, machine learning people know, data science people know, but it, it is a, it is a thing.

    4. AG

      And what is, what does overfitting look like in practice?

    5. HH

      Yeah, overfitting means that you, um, you have hill climbed against some data or something that doesn't generalize. So the key phrase is "does not generalize." And, um, you know, maybe you have some evals. So, like, a one very trivial way to mess this up, and this is from a real client, is you have some, uh, you have a test data set of, you know, cases that you want to do well on in your evals, and you use the same data as few short examples in your prompt. So that is a hilariously kind of straightforward way of overfitting because you have given the LLM the answer directly-

    6. AG

      Yeah

    7. HH

      ... in the prompt. Um, and there's, you know, we can relax that a little bit. There's, there's more subtle ways that you might do that. Um, but essentially, you know, you might over-engineer your prompt with those specific cases in mind with the details of those cases that are a little bit too specific, you know? Even though you don't have the exact, like, information there. You know, that might lead to overfitting. Um, and there's many ways you can overfit, and what we, Shreya and I teach is, well, how do you know that you overfit? How can you guard against it? It turns out you can use a lot of the same techniques from machine learning to kind of give yourself a early warning system to say, "Hey, like, you overfit."

  11. 35:1238:08

    Anti-overfitting toolkit: holdout sets, suspiciously high scores, and hard tests

    1. AG

      What are some of those techniques?

    2. SS

      Uh, number one thing is collect label data, reserve some of it for testing, as Hamel kind of alluded to. You know, you wanna have some test cases that you wanna try out your product to see if it performs well on. You know, reserve that set and never, never look at it. Never look at it when you're developing your product, when you're developing your prompts. Even when you're kind of trying to test out your prompts to figure out how to improve it, just make sure that the test data is, like, in the sandbox that you never, never see. That's the number one strategy that I tell people. Um, another thing that I tell people, and PMs should also know this, is anytime you see a metric that is suspiciously too high, like I achieved 99% alignment or 95 or 100 or whatever, immediate smell warning flags that something is off. We're s- leaking some data into our prompts from the test set maybe. Maybe our test set doesn't have enough examples. Um, so go back and question the process if you see any numbers that are too high.

    3. HH

      That's absolutely true. Well, one of my favorite examples is there's, um, there is a interview that's been going around for a while now. It's from a company called Casetext, who's a, you know... It was bought by, I believe, uh, Thomson Reuters or maybe LexisNexis, one of the two big ones, uh, legal companies. And the CEO went on the, you know... He said, "Hey, like, you need to keep iterating on your evals until you get to 100% accuracy." And so, um, you know, if you're iterating or if you're getting 100% accuracy in your evals, it's likely your evals are worthless. They, because they are providing you with no signal. Just think about anything else. If every, every student in your class is getting a perfect score, is it a good test? You know, a- any kind of test, if everybody is passing it, if everything is passing, should you be celebrating? No. It means the test doesn't have, is not differentiating good and bad, and it probably is not really worth it. And so, um, it's really good... You know, you have to keep that in mind. You want your tests to be hard.

    4. SS

      I think there's a bit of-

    5. HH

      Yeah

    6. SS

      ... a bit of nuance to it. You want some tests that are, you know, basic regression tests or functionality tests that absolutely you wanna pass them because otherwise you're shipping a broken product. But I think the broader point Hamel and I want to make is you also should have some aspirational evals. Um, and those, you know, it doesn't need to be 0% or, like, designed adversarially, but some benchmark that helps you know whether, you know, the general AI or intelligence in your system is getting better.

  12. 38:0842:42

    What “good evals” look like in the wild: interfaces, scoped judges, and product metrics

    1. AG

      In talking to so many AI companies about evals, what's one AI company that sticks out to you as doing it particularly well, and why?

    2. SS

      Everyone's struggling. [laughs] Okay, I'll, I'll talk about some design principles that I see people converging on that are solving parts of their problems. So, so I don't think anybody is doing it perfectly. I mean, otherwise you would have, like, a rocket ship AI startup that's, like, already perfect and, like, gone public and everything works. So we just don't have that, right? It's, it's not there, so, like, let's be honest. Um, some of the things that I've been seeing teams that really seem to give them leverage are, one, building custom data labeling and annotation interfaces, um, or providing ways for domain experts or other people that they trust to be able to give feedback on traces. Um, a lot of the observability tools for LLMs have started to build in these features. They didn't have them even in the last, like, two months, three months, um, so it's gonna take some time until we can really see the effect of that. Um, another thing that I've been seeing is these really well-scoped LLM judges that are being deployed as part of products. Um, Claude, Anthropic just had a new post on their research agent about how, you know, they really break down the complex task into a multi-agent system, and then critically they have a bunch of LLM judges that are very, very well scoped to each task, much like we teach in our course. Um, and I'm sure they're making a bunch of revenue off of this product. Um, and then of course, I think there's a lot around, you know, Cursor famously, you know, very dutifully measures your, you know, next token prediction. They have their own models for this. Um, I think next token prediction for coding is interesting because that metric is very well defined. People know that it's useful. It directly correlates with product success. Not all products have such a metric like that. Um, PMs need to really go figure that out. But these kinds of things I think are broadly helpful across the board that I've been seeing. But I, I think we're pretty far out from, like, you know, the product that's perfect right now.

    3. AG

      What is next token prediction?

    4. SS

      Uh, like, in, when you're writing code, like the next snippet, um, that you're gonna write in your code, if the AI model predicted that correctly, then the metric goes up. If not, metric goes down.

    5. AG

      Kind of like an advanced autocomplete.

    6. SS

      Yeah.

    7. AG

      Why haven't we seen better autocomplete in, like, regular products like email and text message and things like that?

    8. HH

      It's incredibly frustrating. Um, if, especially if you're writing code, if you're a programmer using Cursor, Shreya's using Cursor I know, like, six hours a day. Um-

    9. SS

      [laughs]

    10. HH

      ... and you know, the developer tools are very far ahead of the other tools, and it's incredibly frustrating if you try to use AI in Gmail or in Google Docs or in PowerPoint or anything. And it's, like, even more frustrating 'cause then you open LinkedIn or whatever and it's like, "We have AI in everything." And we're like, "No, you don't. Like, you barely do." Um, and, um, yeah, I'm not really sure why. Um-

    11. SS

      I have some hypotheses.

    12. HH

      Yeah.

    13. SS

      Yeah. So it goes back to Hamel's comment on verifiable domains. Code is a verifiable domain to some extent, right? You can just make sure the code runs, like, that's already great. The other thing is there's a bunch of data of how people write code that's already there and available on the internet in ways that, you know, we don't have that for emails. Um, emails are not very verifiable, right? Like, nobody is telling you, like, this is ground true, like, this is correct or, like, this is not correct. Um-

    14. HH

      Yeah

    15. SS

      ... PMs have to do this job of figuring out how to design a system around good and bad emails. Um, and that's where it becomes really, really bespoke. And, and I think the, that it goes also back to Hamel's comment around code where it's like developers knew what the evals were, so they were able to implement that.Um, but when... Like, this is why we need AI PMs. We need people to be able to help developers craft these verifiable or even, like, loosely verifiable signals.

  13. 42:4246:11

    Foundation benchmarks vs business evals: why OpenAI can’t define your quality

    1. AG

      I know you guys have spent some time with, uh, OpenAI. How is OpenAI approaching evals and what can we learn from them?

    2. HH

      So one thing I'll say is, um, you know, there's a really big difference between foundation model benchmarks, so MMLU score, HumanEval score. These are, like, general purpose benchmarks which try to assess the general capability of models. And then there's another kind of eval which is your domain-specific eval, evals for your business, and they're very, very different. Um, and it's important that people know that. Um, now rightfully so, uh, the foundation model labs like OpenAI, they're very much focused on the former, the, the general purpose benchmarks, because that is, you know, relevant to their product. Um, but I think, you know, they're kind of at the same level of playing field as everybody else when it comes to domain-specific evals and, like, how the companies should be defining their evals.

    3. SS

      The other thing is they're not gonna do it, right? Because, like, your definition of a good email is different from my definition of a good email, and we're both building email assistant companies. They're gonna be different products, and they should be because they should reflect our taste and our company's vision. Um, OpenAI just, like, cannot solve both of them with one model, right? Does- that doesn't make sense, and they're not in that business, right? Like, they're trying to make the model that you wanna hire, for lack of a better term, to do the job. Um, but you need to train, not just... not really train in terms of model parameters, but you need to figure out how to make an environment, how to elicit the right signals, how to reject emails that are bad according to your definitions. Like, that's all the stuff that you've gotta build that's domain-specific that no foundation model company is gonna do for you.

    4. HH

      Definitely. One example that is top of mind right now is Shreya just wrote this excellent blog post, Writing in the Age of LLMs. And, you know, if I was to... So the problem is, is, like, foundation models, you know, they're trained. They have, like, sets of labelers, and basically it's the, it's by definition trained on the average taste of some people. And, you know, if you're, if for... So if you're writing a lot, you know, like I know you are, Aakash, you know, as well, is, you know, what comes out of the LLM, like, almost everybody I know that writes a lot kinda hates it.

    5. AG

      [laughs]

    6. HH

      They're like, "Why can't it be better?" Like, "Why is it giving me, like, the same... Like, why is it, like, using so many words, and why is it, like, doing this? Why does it keep using em dash everywhere?"

    7. AG

      [laughs]

    8. HH

      "Why is it using em- so many emojis? Like, stop it." It's because, like, y- you know, whoever's labeling the data, you know, they kind of bias towards, like, longer explanations and, like, overusing bullet points and all this stuff. And so, like, you know, if you want to have, like, a, a thing that writers love, you know, then you have to get someone... You know, you have to get, like, perspective like Shreya has and, like, iterate on your product, like, over time and make it do those things that incorporate the taste that you have, you know? Um, and, s- and so that's kind of where i- it diverges from the foundation models.

  14. 46:1156:04

    Evals as the moat—and how fine-tuning fits the lifecycle (after evals)

    1. AG

      So do those wrappers where you're putting in your own taste into the evals, do they actually have some sort of moat where people should consider building those startups?

    2. SS

      Oh, absolutely. Um, I, I think maybe I'm crazy. I think evals are the moat for AI products, and, like, truly nothing else. Like, tomorrow you can use a different model, you can have a different stack for serving, whatever it is. Like, they're... The secret sauce is your evals and how you're able to operationalize or scale that out as quickly as possible. Um, so that means having good LLM judges, for example, that are, like, very well aligned with your preferences that you can just automatically run on everything, um, and build that flywheel for yourself to, like, go then look at what failed and improve your product and so forth.

    3. HH

      And it's not just, like, it's not just the eval itself. It's the whole system, like Shreya is saying, around the eval, the entire eval pipeline. It basically should be portable. You should be able to switch models or, you know, switch components and see what the effect of that is. And the key thing is, like, evals open up a whole lot of doors. So it opens up the door for easy fine-tuning. You know, once you've done an eval, you've already done 99% of the work of fine-tuning. Um, fine-tuning is just a kind of a formality that you can go through after that almost, to, to an extent, 'cause you've already done... You already have all the tools for data curation. You know how to measure things. You know how to, like, see, measure the effects of the fine-tuning, so on and so forth. Um, you know, and I suspect that that will become... That just accelerates your advantage, like, that much more.

    4. AG

      It's kinda funny 'cause I see people tend to focus on fine-tuning [chuckles] a lot more, and there's a lot more attention given to that. I think it's because, like, in the OpenAI and Anthropic developer docs for when they u- give you a model, they're like, "Yeah, this is how you can fine-tune it," and things like that. How does fine-tuning really fit in with the overall life cycle? When should you be putting attention into fine-tuning?

    5. HH

      We, we should put it in last. So you should do evals first because, like, what are you... You have to know-

    6. SS

      How do you know?

    7. HH

      ... if you're fine-

    8. SS

      Yeah.Oh, sorry.

    9. HH

      No, no. You're good

    10. SS

      My, my connection is bad. Uh, there's, like, lag here. Um, yeah, you need to know... This, this goes back to you wanna iterate on your product. You wanna know, "Okay, my product is not good, so I'm gonna improve it." You don't have any way of knowing if your product is not good unless you have evals. So evals are step zero. Quantify your performance. Then step one is, okay, how am I gonna improve its weak points? Some low-hanging fruit ways of improving are just use a more powerful model. Switch from f- GPT-4o mini to GPT-4o. See if that works. Because you have good evals, right, you can just make that switch, run it, see if your numbers go up. There are other more complicated strategies you can do, like take a complex LLM call, a complex task described in them, and break them down into multiple LLM calls. So instead of, you know, having an LLM extracting 10 things from, uh, this document, I'm gonna have one LLM call each extracting one of those 10 things. Um, so like that you can kind of do task... That's what we call task decomposition. When you exhaust these strategies, then it makes sense to move into fine-tuning, where it's like, "I have no other way of solving my problem. The model's just not there yet. I'm gonna collect a bunch of data or use the data I've already labeled and fine-tune a model." Fine-tuning has cons because now you have to make sure that you're continually fine-tuning that model as you get new data, you learn something new about your preferences of your customers. Um, they don't just automatically update if the base model updates, right? GPT-4o or 5o or whatever comes out, now you have to redo your all fine-tuning all over again. And then if you don't use a f- model provider's fine-tuning service, like you're fine-tuning an open source model, you need to deal with the MLOps complexity of serving that, making sure it has good uptime, latency is low, load. This is all actually pretty complicated, right? Costs a lot to maintain this infrastructure, both in human personnel as well as in money. Um, and, and using a off-the-shelf LLM is, like, pretty much the right way to go in most people's use cases. Um, so for that reason, for those reasons, we really try to not encourage people to do fine-tuning unless you really, really have a good reason for doing your own fine-tuning.

    11. HH

      One example where I might e- go to fine-tuning, hilariously, is this writing thing, 'cause, like, Shreya, we, uh, we talk about it a lot, like writing, writing. We bash our head against the wall. Even right before this call it's like, oh, like, did you try 4.1? And Shreya made, uh, good comments like, "No, like, you know, you can't prompt it with all of these rules. It's not gonna follow these rules." So I'm like, "Okay, there's no... This really feels like there's only one avenue left here." Um, you know, we had put a lot of work into it, but it's like if you wanted Shreya GPT, that feels like the only way to get it.

    12. AG

      Fine-tuning is always talked in the same conversation as prompt engineering and RAG. So when do you really think about using each of those three techniques?

    13. HH

      Um, okay, so, like, yeah. Prompt engineering is basically, you know, everybody sh- That's just basic, you know, you're using the LLM, you have to communicate what you want to it. You have to specify what you want in whatever way, so you should always be writing prompts. You don't even need the word, put the word engineering at the end. We are kind of, you know, adding a lot of ceremony. Um, you know, we're just-

    14. SS

      [laughs]

    15. HH

      ... you're writing. You know, write English or, you know, in some cases you might, you can write other languages too, but it is mostly English. Um, you know, and refine your thinking and refine your instructions. RAG is, you know, any time you're, you need external context, um, which there's a very narrow s- set of use cases where you don't need external context. Most of the time you do need, in many applications, you need some external context. Um, you know, that's not some general knowledge of the world. Like... And so, um, that's when you need RAG, you know, and you're not trying to bake all that external context into your prompt. Um, and then, yeah, fine-tuning is sort of, hey, like, if you can't prompt this behavior, if, if the, if the model's not doing what you want... Shreya has a, uh, has, uh, created this, like, very powerful, um, model called The Three Gulfs. Maybe we should bring it up on the screen, actually.

    16. SS

      Yeah, this is The Three Gulfs question.

    17. HH

      Yeah. Let's, let... We should dive into The Three Gulfs.

    18. SS

      I think they're in lesson one.

    19. AG

      Okay.

    20. SS

      Or, great, you have them.

    21. HH

      [laughs]

    22. SS

      Oh, you made your own gulfs. Nice. [laughs]

    23. HH

      [laughs]

    24. SS

      Yeah, I, I can take this out. So your question was, you know, you have all these tools available to you for improving AI products. You have prompt engineering, you have RAG, you have fine-tuning. Are they the same? Are they the different? When do I use them? Blah, blah, blah. This is a good question that trips people up a lot because people are just thinking of them all equally as ways to improve their product. They're all complementary strategies. You can do multiple of them. You can do none of them. You can do one of them, whatever it is. Prompting is very good when you have specific requirements that your task follows, and you need to communicate that to the LLM, right? So say you were to hire some human to do some job or contract some job for your company, you might give them a s- task specification. Say, "Do this, then t- do step A, then step B, then step C. Make sure your output follows these requirements." Right? That's, that's prompting. You know, there's no... It almost feels like there's no ceiling sometimes to how much value you can get from good prompting, and this is the first thing that you should be doing. And this solves this gulf of specification problem that Hamel and I see a lotUm, in how people build LLM products, which is, you know, they have some latent, this hidden criteria, this hidden task that they want the LLM to do, but they just don't know how to specify it well in a clear way and completely so that LLM fully understands all of those hidden preferences. Now, fine-tuning RAG strategies are for this gulf of generalization, which targets this problem of I have a very good specification, but the model is simply not good enough for some reason. Maybe the model is not powerful enough, so I should do fine-tuning. Maybe the model just doesn't have all of the context that it needs to make the decision. That's where RAG comes in. I need to pull in external data sources to improve my product. But yeah.

    25. AG

      Very helpful. And-

    26. HH

      And I would submit to you, I know that you ask this question a lot in your podcast, 'cause I think it's confusing, and rightfully so you ask this question. I would say this might be a good framework for answering that question. Um, you know, since it comes up a lot.

    27. AG

      Yeah, I try to play the role of the, the watcher, right? [laughs] And they keep asking me that question, so I gotta ask the experts. [laughs]

    28. HH

      Makes sense.

  15. 56:041:15:27

    Learning roadmap: error analysis, grounded theory coding, judge calibration, and productionization

    1. AG

      If you guys had to build a roadmap for people who wanted to get really deep on AI evals, what topics should they learn?

    2. HH

      Great question. So Shreya and I have developed this course reader. It's actually a very extensive set of notes. One might even call it a book. It probably will become a book. It's 150 pages. This is the detail that we go through in our course. We arm our students with a lot of information to make sure that they get materials in many different ways, including, uh, you know, live instruction, office hours, but also this very detailed sort of treatise on, like, how you go through eval step-by-step process. So we start off with, um, you know, like what is evaluation. We go through The Three Gulfs framework that we just described. Um, you know, we talk about why you need evals. We motivate that. Uh, then we kind of do a little bit of overview of, like, okay, the strengths and weaknesses of LLMs, you know, what kinds of things you need to intuitively understand when you're doing evals before you even get to evals. The third chapter is very important. It's probably where we spend a lot of our time in practice, and a lot of people we don't know about this step. It's called error analysis. So what is error analysis? So you might hear Shreya and I talk a lot about looking at your data. We keep beating people over the head with this phrase, "Look at your data. Look at your data." What does that mean, look at my data? What data do you look at? Do you look at all your data? Do you just quit your job and look at your data and do nothing else? Do you not have time? How do you look at your data? And what do you do when you look at the data? Like, how do you make sense of it? How do you make it... How does it make it tractable and learn something from your data? What if you don't have any data? Like, what data am I talking... What if you haven't built anything yet? Um, the key thing is, like, you need to look at some sort of data, go through a structured process. It's not as painful as it sounds, looking at data. It's actually really, um, beneficial. A lot of my clients, you know, we have a plan to go through this whole process, and they get so much value just out of error analysis that they just, they're like, "This is amazing. I'm done." I'm like, "Wait, what do you mean you're done? Like, I can do all this other stuff for you." Like, "No, no, this is great. Like, I, this is like I'm busy for a while. Error analysis has taught me so much." And, you know, error analysis is, you know, to go to error analysis, um, we can scroll through, um... Let me see if I can, there's a diagram here. But, you know, um, this is, this is, like, one way to generate synthetic queries, for example. Um, it's hard to dive into the details just without the context, but we go through the process of, like, how to generate, you know, synthetic data. Um, you know, and then also we go through, like, how to look at your data. Um, uh, and, you know, so we describe it in a lot of detail. Um, a lot of it is supplemented with videos. Let me see if I can scroll up. Maybe Shreya, you can, uh...

    3. SS

      Yeah, that diagram is pretty good-

    4. HH

      Yeah, this one

    5. SS

      ... for what people do.

    6. HH

      So there's a concept called axial coding, or open coding and axial coding. Um, Shreya, you wanna tell, tell us everyone where that comes from-

    7. SS

      Yeah

    8. HH

      ... and the history behind it?

    9. SS

      Yeah, definitely. So, um, in creating this curriculum, we took a lot of inspiration from social science research actually, because w- we were thinking, okay, in what field and domain do people need to look at vast amount of unstructured free-form text and labels and come up with meaningful, actionable insights out of it? Well, turns out social science researchers have done this for a very long time, and there's this process called grounded theory, um, that gives this systematic structure to the process. And first what people do is go through their open-ended data co- outputs or whatever it is. They read them, and they write free-form open codes is what they call it. Um, free-form notes on what's good, what's bad, any themes that emerge, whatnot, on a trace-by-trace basis. Then after going through about 100 of those, they will try to merge similar notes together into clusters, and this determines your failure modes. If you find that a bunch of-Notes around, you know, a specific type of hallucination get merged together. Well, looks like that's a huge problem that you need to solve in your product. Um, and this kind of loop keeps going until what we call theoretical saturation in qualitative research, which means I've not learned any new failure modes. That's all. I keep looking at data, and I just keep adding to my existing failure modes. Um, at that point, you can kind of stop, and then you can move on to, okay, how am I going to turn those into automated evaluators so I don't have to do it all the time? Maybe I'll build LLM as judges based on my labeled failure modes. Maybe I'll write some code-based evaluators for things that code can check. Um, and then we find that, you know, we teach people this, this diagram, error analysis, and then suddenly they're, they're like, "Oh, this is great. Okay, bye." [laughs] And it's like, I mean, I guess, right? Like, it, it makes sense. This is where most people are actually bottlenecked, right? Because M- ML, traditional machine learning people didn't have to do this. They didn't... When they looked at their data, they would just look at all these, you know, tables of features, which are all numbers, and the outputs are also numbers, and so there are ways to kind of debug those. LLMs are different because now the inputs are free-form text, and the outputs are free-form text, and people are like, "I don't know how to..." They're... It's new technol- It's new, new skills that you need to be able to make sense of that. Um, but fortunately, once you do this, then you can go and implement automated evaluators as you might with traditional machine learning, um, and then you can do improvement strategies, you can do fine-tuning, all of these things like a traditional machine learning person would.

    10. HH

      Um, one thing I wanna share that might be useful here, or interesting perhaps, for a product manager, um, we have Teresa taking our class and, you know, just today she was discussing, uh, you know, you know, uh, okay, I love the class too. Um, you know, when open and axial coding, are we analyzing opportunities towards an outcome? And she responds, "I've been thinking about this a lot." Um, you know, and then she has some comments about generating synthetic data. You know, but she also recognizes, and this is really fun, this is why we have product managers in our class because, you know, um, these kind of like tying together of things like, okay, identifying opportunities from interviews is based on grounded theory also. Um, you know, and then she's experimenting with methods for capturing annotations directly from the customer. And, you know, she also kind of points out that, okay, these are the scientific method applied, look at the data. So, um, I think project managers have a lot to offer here, you know, bringing their customer and user research into the whole process. That's why w- we think it's really important that they are in the driver's seat of these evals.

    11. SS

      I will go even further to say that I think that PMs are, like, vital in building AI products. Like, we're not gonna have successful AI products across different domains unless we have good AI PMs. Like, I cannot emphasize that enough. Um, without good AI PMs, like, we're just gonna have failed AI products, so I hope this is a call to action for PMs. Like, this is a skill that you absolutely need to develop. Um, it's going to set you apart, obviously, of course, from a career perspective, but also we need this, right? In order to really realize this vision of AI products changing people's lives.

    12. HH

      Um, okay, so to get back on what, you know, the things to know. So once you do error analysis, now you have a grounded way of knowing, like, what to focus on. Because, like, the question is, like, what are you even evaluating? You can come up with, like, millions of metrics, like hallucination score, toxicity score, conciseness score. You can name these scores all day long, and you can get really confused and overwhelmed and say, and say, "Oh my goodness," like, "I can't do it." So e- error analysis is very important. Um, you know, you would be surprised, in our Discord channel, almost every other question is answered with the phrase error analysis. Because, um, you know, one of the main trouble people have with evals is what do you eval? You can, you know, you can eval anything. There's an infinite number of things, ideas of things that can go wrong. Um, you know, instead of just being paranoid and sort of, you know, working yourself up about, like, what can go wrong, you should be grounding it in things that are going wrong or the things that will probably go wrong. Um, and that's what error analysis helps you with. It helps you focus on high-value things because evals are not free. Um, it takes a little bit of effort, and so there's a... This entire time that you're doing evals, there has to be a cost-benefit analysis a- around, like, what do you even eval? Um, should you be evaling? What do you eval? Whatever. And so a error analysis is the answer to all those questions. Um, and error analysis is so valuable that, uh, you know, a sizable portion of my clients, they sign up with me to go through this whole evaluation process, and they do the error analysis part, and they're like, "Great, Hamel. We love it. We're done." Like, "This is..." I'm like, "No, wait. I have all this other stuff, like all this other stuff in this table of contents I can take you through." They're like, "No, this is so much value." Like, "We've... We're so happy." And it's true because, like, you know, error analysis is like this thing about looking at your data. Um, it actually drives you to develop a very deep intuition about your system and what's going wrong with it, and it makes you develop a nose for everything and, um, kind of gives you a sixth sense of, like, what is even gonna break if you're doing it enough. And so it's extremely powerful, probably the most powerful technique in evalsSo, um, you know, I can't highlight it more than that. Um, but okay, once you move past error analysis, now it's... Now you know what is wrong and where to focus and what to prioritize, now how do you actually go about the process of writing the evals? And so that's where these other chapters come in around, okay, um, how-- who does the-- who writes the evals? Who does the annotations? So we have this, uh, we have some strong guidance in here. We're not, we're not afraid to, you know, cut right to the chase and have opinions. And so, like, one opinion that we have, for example, of the many opinions, but I'll just highlight it here, is benevolent dictators. So we say, okay, the people always ask, like, "Well, who is going to make the final call of whether or not this is a pass or fail? Or who is it gonna be the one writing the eval? Who's in charge?" You know, I have a whole team. Should I just involve the whole team? The answer is no. You need to have a benevolent dictator in most situations. Um, and there's some, you know, there's some reasoning behind that. It just, you know, it's like the binary classification thing. It forces you to kind of make a decision, and it makes the whole process tractable, 'cause it can quickly become intractable if you add various complexity. So going back to the table of context, let me... context. Let me just scroll back up. Um, feel free to interrupt me, Shreya, any time. Um, so we talked through implementing, you know, making this a process automated with LLM as a judge. Now, with LLM as a judge, you're writing prompts, and you're asking an LLM to grade something. Now, almost everybody, I would say 80% of people who are using LLM as a judge, are just writing a prompt and just praying that it's doing the right thing. But that's not what you should do. Y- When you... The right way to use an LLM as a judge is to label, to get your label data, which you've already done in a previous step, that's what we were teaching, and measure the LLM as a judge against your label data to understand, is the judge any good, number one. But number two, more importantly, you have to iterate on your LLM as a judge. You have to see, like, where is it wrong? And you're like, "Oh," it's al- always the case that not only do you adjust the LLM as a judge, but Shreya has a lot of research that shows that the annotator also adjusts their requirements seeing the LLM output. Uh, Shreya, you wanna talk a little bit about that component?

    13. SS

      Yeah. I have a paper called "Who Validates the Validators," um, with 2024, if people are interested in the venue. But yeah, one of the things we did was we built an interface for people to go and annotate outputs to train LLM judge evaluators. And then we found that people will keep changing the rubric as they encounter new failure modes when looking at the data. And this underscores, right, how tricky it is to figure out what the rubric is, how you need to go through this iterative process, how sometimes when the criteria is a little bit too subjective, maybe you wanna have multiple people weigh in on what correctness means. And there are ways to do this, like the interannotative, interannotator agreement metric that we talk about, um, Cohen's Kappa in this course reader, a lot of stuff. Um, point is that we just provide a lot of frameworks to really systematically solve the problems when you encounter challenges in coming up with good labels and ground truth.

    14. HH

      Um, and so, you know, we teach you a lot of things, like also how to, you know, correct the LLM as a judge error rate based upon your real error rate, and, you know, all kinds of advanced things. We also, uh, teach you how to think about multi-turn evaluations. Like, okay, if you have a long conversation, um, that's like multiple turns between an AI and a human, how do you actually evaluate that? Um, do you evaluate the entire conversation? Um, you know, one piece of advice we have is, like, you should st- you know, evaluate... You should... When you're doing open coding, for example, you should stop on the first error that you see, because these things have causal relationships, and it's usually the case that the first error is the blocking, uh, is the blocking problem. And so it, you know, a simplifying heuristic is to, is to, is to y- anchor on the first error. And there's all kinds of heuristics like this that make the entire process tractable. Um, but also when it comes to evaluating multi-turn conversations, like how do you actually have test data sets for that? How do you, how do you simulate a multi-turn conversation, or do you need to simulate a multi-turn conversation? There's a lot of, uh, com- you know, there's a lot of things to dive in there, so we cover that. We also cover how to evaluate retrieval augmented generation. And so we talk about when to treat certain aspects of the problem like a search problem, when, what components of the problem, like the generation problem, how to think about that separately, and, you know, all the nuances there. Um, and then we talk about other specific architectures and components, so like how, how to think about tool calling, what you need to know about agents. So agents are systems that can be very dynamic. You know, they might have trajectories that are unpredictable and do various things that you can't anticipate. How do you evaluate that? How do you tame that m- monster, that spaghetti? It turns out you can use various analytical tools that simplify the problem and allow you to attack it and make sense of, you know, agentic systems. And we show you, uh, things like how to use transition matrices and other analytical tools to help wrangle that problem. Um-And then we talk about, you know, how to evaluate specific input data modalities, like different types of modalities and, like, not just text. Um, then we talk about production, uh, productionizing these things, so like CI/CD. Um, you know, kind of automating these and running these at scale. And then we talk about also kind of probably my second favorite subject in chapter 10 is interfaces, and I'll, I'll give it to Shreya to talk about the interfaces, actually.

    15. SS

      Yeah. Chapter 10 and 11, they're the last week of the course, and they're kind of like special advanced topics that you will, uh, oh, I can pretty... I think I can guarantee that you won't find them anywhere out there on the internet. So we wanted to do something special, um, for our students. Chapter 10 is about how to build perhaps even vibe code your way to an effective interface that has people really, really labeling things very quickly, maximizing the throughput of how many labels you can get from human reviewers. We talk about principles there to case studies of good, bad interfaces. Um, and then chapter 11, I think, is about improvement. You know, obvious strategies that you might know about, like decomposing tasks, fine-tuning models, so forth, but also cost optimization and cost improvement. A lot of people actually say like, "Oh, my pipeline is working, but it's too expensive," especially if it's working on really long documents. Um, how do we keep the same quality but reduce the cost by an order of magnitude or even more? Well, we talk about some cost optimization techniques there. So that's a very long answer, I think, to your question on, you know, what's the roadmap like? Here is a roadmap. We think it's pretty good. A lot of places you can dive in. You can dive into each chapter that you're interested in. Um, and chapters 10 and 11 here, I think, are pretty open territory, um, in b- in research to explore more as well.

    16. HH

      So upon seeing this book, a natural question is, where can I get the book? Or are you going to release this book? And so the answer to that question is eventually yes, maybe early next year. However, um, you know, if you want the kind of hands-on approach of, you know, doing this in, in a very guided way, then, of course, take our course.

  16. 1:15:271:34:22

    Business of teaching evals: consulting to course, pricing, and why cohorts end

    1. AG

      That's actually what I wanted to ask next, Hamel, is you were working at Airbnb and GitHub. Why did you transition into consulting and courses?

    2. HH

      Good question. Um, it kind of happened somewhat organically. Um, you know, I took a sabbatical after, uh, GitHub for a bit. I worked at some startups, and then I... Sorry. I, I, uh, after GitHub, I worked at a s- a, a few startups, then decided to take a sabbatical. Um, and a company decided to contract me and convince me kind of to, uh... The name of the company was Weights & Biases. They convinced me to, uh, do some consulting work for them. And so I, yeah, I always thought that, like, I would hate consulting, 'cause I did consulting with a large consulting company very early in my career before tech. And it j- and I realized, like, hey, I really like it. When you're doing it on your own, becoming, like, an independent indie kind of entrepreneur, it's very enjoyable. Um, it's a lot of freedom. And I found that a lot of people found it very valuable, like, in terms of, like, helping them sort of navigate LLMs and AI. Um, and then I just, yeah, I just started really enjoying it, and one thing led to another, and here I am. Like, uh, there's no, let's say, grand plan as such. It's just a kind of, you know... Yeah, it just, like, happened to be in this place. And it happened that everybody building with AI, they're always stuck on evals every single time. And it was the exact same problem almost every single time I needed to start with error analysis, and this is, like, you know, working with 25, 30 companies. Um, you know, and then I just started writing about it a lot, and it was really apparent to me that there's not that much education on this topic, and it's where everyone's struggling. Um, you know, and consulting is expensive. And so that's why we created this course, to make this way more accessible to a larger audience.

    3. AG

      And do you tell people who wanna consult you as a consultant to not reach out if they can't spend $38,000?

    4. HH

      Yeah. So I think one thing that you have to do as an entrepreneur is to... A lot of different things, but one is, like, it's important, can be important to have a niche so that you customers know, uh, what you can help them with. So in this case, mine is evals. And, you know, just from a practical perspective, there's a lot of overhead that comes with being an entrepreneur. Like, you have, you're in charge of your own sales, your own marketing, your own, you know, all the administrative overhead of it. And, um, you know, you don't wanna be on sales calls all day. You have to qualify people, and, you know, the problem has to be painful enough for them to wanna solve it. And so that's just a kind of a basic kind of blocking, tackling approach to say, "Okay," like, "how do I not drown in Zoom meetings?"

    5. AG

      And then the course, it costs roughly $2,000, and you have over 600 people in this cohort that me, Theresa, Pavel, other people are participating in. So does that mean you earned over a million dollars on this cohort of the course?

    6. HH

      Um, [clears throat] not n-No. So we didn't quite cross the million dollar threshold because we gave a lot of stuff away for free. A lot of friends, uh, we, you know, I let all my friends in, basically, uh, Shreya-

    7. SS

      We're also very generous with any discounts for people who have any need. Basically, if your company can't reimburse it and you don't want to pay fully out of pocket, you know, just email us with how much your company can reimburse, and we are super flexible. Like, we want people to join the course. We have a steep price point because we find that we want, we want to attract people who are serious about building AI products and evals, right? It benefits nobody if our course has 10,000 people and everyone's, like, only mildly interested and serious, right? We want to focus on the people who are going to go and put these techniques into production, who are going to go build products. Um, and for that, turns out you need to have a steep price point.

    8. HH

      So almost, yeah, almost a million. I don't want to dodge the question because I know, like, whatever.

    9. SS

      Oh, yeah. It's, uh-

    10. HH

      It's, like, almost a million. It was like 800K is what it was, the first cohort.

    11. AG

      Which is insane. That's mind-blowing, right? Especially for the first cohort of a course. And what's even more insane, I think, is that your next cohort is gonna be your last. Why?

    12. HH

      Yes. So I don't want it to be... You know how you, like, watch a movie and, like, the first one is good, and usually, like, it's really hard to make a good sequel? I mean, this is, like, really impossible, like, like, to make it, like, continue to be good. And honestly, like, um, you know, I don't want to... You know, I have a lot of different interests, but also, um, there's a lot of things that we can do with the course. So we could, for example, you know, you see this, like, 150 page. This is just a draft, honestly. I mean, you know, we, we keep iterating on this, on this book that we have open right now. Um, so one, you know, one sort of motion is to, like, okay, focus time on writing a book. Then can... You know, it's, currently it's a $2,300 course. We're gonna actually be increasing the price to $2,500 on Friday. Um, you know, can there be a recorded version that doesn't involve us, like, live, that doesn't involve any office hours or anything else, um, you know, that's, like, a lower cost? Okay, that's fine. So we think, like, there's different products out there that we can offer or different ways that we can offer this. Um, but also, like, you know, this course takes an immense amount of time to deliver. And so one of the things that we like to do is actually build these AI projects. Like, I like to work with customers and build these things. I know Shreya does, too. She likes to, you know, do research on this topic. Her research is very applied. She works with companies on this. She builds tools in the space. And it's really, yeah, if we were to do the course over and over again like this, we wouldn't be able to do that. It would actually dilute the course. So we want to keep it special. We want to, we want people to feel like, okay, they were here for this course. It's something they could remember. It, you know, um, last year I had a course like this. Um, it was not on this. It was on fine-tuning. It was, it was provisionally about fine-tuning and then became, like, basically a conference. But it was basically the same scale as this course. It was, like, also a $900,000, uh, revenue course in one cohort. Um, and yeah, we decided, like, we don't want to do that again because how can we possibly recreate that, that subject and that feeling of that time? You know, we don't want to do a second one and just have it be underwhelming because we try to repeat something. So like, you know, this is different. This is, like, that was, you know, kind of a conference in a way. It turned into a conference. It was just, like, everybody about anything about LLMs. It was really fun. Um, you know, this is, like, really focused on evals. We think that, okay, there's a lot more repeatability to this. But, you know, our goal is not necessarily, you know, trying to milk all the money out of it, per se. We're actually, like, you know, all this, this money that made, we're actually, like, gonna invest it in all these things. So, like, writing a book, um, you know, doing this additional, like, different kinds of courses, like, that are, can be delivered in different ways. You know, it's all gonna be reinvested into that to make this material, like, accessible, um, because we really believe, we have a lot of conviction in this message and this topic and the impact of, of learning this.

    13. AG

      I think there's a lot of lessons there for aspiring and current course creators, but I think one of the most interesting ones is how of the time you've made this course. Like, you're addressing the problems that you're seeing with your consulting clients and with AI engineers and AI PMs that they're facing with evals today. So you're able to write about it that way. You're not trying to just create timeless content, and I think that makes it really effective.

    14. HH

      Yeah, I hope so. I mean, uh, you know, people message me and they're like, "Oh, this is a lot of money for, like, a month," or whatever. I mean, I remind them, like, been writing about this for two years. I think Shreya has reviewed all of my blog posts throughout the years. Um, you know, I send it to her, and she, like, reviews it. Uh, you know, I'm really grateful for that. Um, you know, and she's been writing about this for many years as well. Um, and by some stroke of luck, I was able to convince her to nerd sniper also and say, like, "Okay, let's do this course together." Um, and I wouldn't have done it if she said no. It was only, like, if she does it, then I will do it kind of situation. Um, you know, because Shreya is, you know, if you look at her writing, it's actually really impressive. Like, she, yeah, she, like, um, is a very good complement because I'm, like, really focused on consulting and stuff like that. I don't have time to survey the field, bring in, you know, like-This, like, generalized perspective in theory, but there's a really good complement there where Shreya brings to, you know, structuring this material, bringing in like, you know, a lot of the kind of... I mean, I didn't even know what axial coding and... I was just doing axial coding and open coding. I didn't know it had a history in social sciences. So it was really good to know that there's a, there's actually some, like, history of this working. So things like that. I mean, um, I don't even know if I'm answering the question anymore. [laughs]

    15. AG

      There was no question, so that was great. If you were to give advice to somebody who wanted to create an $800,000, $900,000 course, what would be your advice to them?

    16. HH

      Yeah. So one is... Okay, so a lot of people ask me this question. Uh, let me try my best. One is, is to really find a niche that you are passionate about. Like, I didn't think I was gonna create a course around this, honestly. Um, you know, I've been writing about this forever. In fact, evals is, like, one of the most unpopular things you can possibly write about or talk about.

    17. AG

      [laughs]

    18. HH

      It's, like, deeply, deeply unpopular. Um-

    19. SS

      Yeah.

    20. HH

      Yeah.

    21. SS

      I really wanna drive that point. Had we talked about something very cool and sexy like, you know, I don't even... There's so many buzzwords out there, like MCP and, like, agents and stuff, right? Like, that is a easy topic to write a course on because everyone wants to learn how to do it. Evals is something that everybody knows is a problem but doesn't wanna go and put in the effort to doing. So I think, yeah, my number one advice would be to find a... If you wanna do it the easy way, find a problem that people are excited about going and actually doing the work for, 'cause you could probably 10X what we did [laughs] for evals if you pick a ex- exciting topic. Um, but Hamel, you should continue.

    22. HH

      Yeah, I mean, I actually went through several nights or weeks or months or whatever of thinking to myself, "Man, am I swimming upstream?" Like, you know, I talk about evals, I built... Sorry. I built my, uh, consulting around evals, but like, you know, you know, it's like telling people to eat their vegetables. Um, you know, it's not really that popular. And, um, you know, it's much easier to talk about agents. You know? If, if you did some, like, background research, if you tried to come out of the perspective like, "Hey, I should build a course. I wanna, like, build a course," you would never pick evals. You would pick agents.

    23. AG

      Yeah.

    24. HH

      You'd be like, "Okay, I'm gonna sell an agent course. Everyone wants to learn about that." People are actively googling. If you do, like, SEO research, like what are, like, the keywords people... Like, almost nobody is, like, searching for evals. There's very little competition for that keyword, actually. Um, like, if you search the evals, like Shreya, myself, Eugene, basically our friends. Like Shreya, me, and our friends. And I'm talking about our friends meaning the friends that Shreya and I have in common. We all are come, come up at on, like, you know, at the top of the list. It's pretty insane. It's like that's how niche it, or quote, "niche" it is. But you know, at some point I just didn't care. I was like, "We need to create the category." So you know, it's not popular. We need to make it popular. So I don't know if that's good advice. That was my mindset, but I can tell you the only reason that allowed me to get into the mindset is because Shreya partnered with me. Shr- And, and so one thing I would say is like, find a good partner to partner with you. So like, to get into the mechanics, like, you know, a lot of course sales has to, you know, marketing is very important. You need to let people know that a course exists. Otherwise they not can't buy the course if you don't know. And, uh, they not only need to know it exists, they need to know, like, what it's about, um, they need to get excited about it. They will need to know, like, it's good. They need to know, like, you know, need to have some social proof, or need to have, you know, some idea, like, people like it, they, and their peers get value out of it, all of these things. And they need to be reminded about it constantly because, you know, I need to be reminded about things constantly for anything. So, um, you know, all of this is very, very important. And so, uh, that's a full-time job. You know? And so s- s- if you want a course, you wanna sell a course like that, like, you have to spend a lot of time doing that. And can you outsource that marketing? You can definitely partner with people on marketing, but you kinda have to drive it 'cause you are the subject matter expert. You're the domain expert. You know, you know, you have to build that authority and that connection with your audience because, you know, whatever. Um, you know, just relentless focus on that. Um, but-

    25. SS

      Hamel is not also-

    26. HH

      Yeah

    27. SS

      ... talking about all of the work that he spends putting together, you know, like, a lot of really great guest speakers, really thoughtfully figuring out, okay, you know, we're talking about X, Y, and Z, so and so is like the world's leading expert on X. Like Y has been applied in these three companies, so let's get somebody from there. Um, this, the guest speakers are really diverse, and I think really distinguish also our course from a lot of other courses out there. It's not just us monotonously lecturing. It's we've got other people coming in, showing you how these things are implemented in practice at big companies, at small companies, at specific domains or whatnot.

    28. HH

      Another thing is focus. Yeah, this is like, I get up in the morning, all I think about is the course and how can I let more people know about the course or like, you know, make it better for students. Um, how can I bring more perspective in, in to there, so on and so forth.Yes, all that stuff is really important. Um, and just constant experimentation. So kind of going back, it's kind of like a, you know, I'm a data scientist kind of thinking person. I'm just constantly doing experiments and looking at the data. Um, and you know, I send Shreya, like, probably average of 200 text messages a day or something like that-

    29. AG

      Wow

    30. HH

      ... with, like, different experiments. Like, I'm like, "Oh, I'm trying this. I'm doing this thing. Look at these numbers. Okay, like, what do you... Like, this marketing copy, this, here. I'll put it... Like, if you put an ad here, this is what hap- this is what happens. Do you have any ideas for, like, how we can go about this differently?" Whatever. Um, you know, hopefully, she has do not disturb on, I hope.

Episode duration: 1:34:32

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode g_3LJ2QBOQE

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.