Skip to content
Aakash GuptaAakash Gupta

How to Build AI Evals Step-by-Step | Daniel McKinnon | Product Growth

Every PM is about to start building AI features, and the ones who understand evals will separate themselves from everyone else. Daniel McKinnon, former PM on the Llama models at Meta and Google, builds a real agentic eval from scratch on screen. He shows why subject matter expertise, not tooling, is what actually drives a good eval. Full Writeup: https://www.news.aakashg.com/p/how-to-build-your-first-eval --- Timestamps: 00:00 - Intro 02:41 - What an eval actually is 03:25 - Offline evals vs shipping to prod 05:22 - How evals got more complex 07:03 - The mechanical process of writing an eval 08:27 - Why easy and hard evals both fail 10:05 - Ads 12:35 - Why old benchmarks are saturated 13:49 - From QA thinking to task thinking 15:32 - Building an agentic eval in real time 17:59 - You must deeply understand the problem 18:49 - The cystic fibrosis eval 23:11 - Running the variance file eval 27:32 - Testing the agent on the disease 29:44 - Ads 33:07 - The congenital heart disease eval 36:07 - Moving to a harder problem 40:39 - Why you sample multiple times 44:10 - Finding the model ceiling 47:04 - It is all subject matter expertise 50:01 - Product management at Meta vs Google 52:59 - Building Gamoff Labs 55:33 - Closing thoughts --- 🏆 Thanks to our sponsors: 1. SerpApi: Clean, structured search results from Google, YouTube, Bing, Google News and more. Get started with 250 free credits - https://serpapi.com/?utm_source=youtube&utm_campaign=aakashgupta_july_2026 2. Product Faculty: Get $550 off their #1 AI PM Certification with code AAKASH550C7 - https://maven.com/product-faculty/ai-product-management-certification?promoCode=AAKASH550C7 3. Ariso: Ship AI agents and features faster, with fewer regressions - https://ariso.ai/aakash 4. Land PM Job: A 12-week experience to master getting a PM job - https://www.landpmjob.com/ 5. Pendo: The #1 software experience management platform - http://www.pendo.io/aakash --- Key Takeaways: 1. An eval is a trivia question for the model - At its core, an eval is a prompt with a correct or plausibly correct answer plus a way to score whether the output is good. It is the clearest way to communicate what your product should do in the AI era. 2. Offline evals catch problems before you ship - Test the model offline against a fixed prompt set before pushing to production. If it fails, you change the model, the prompt, or the approach before real users ever see it. 3. The best eval sits between too easy and too hard - An eval that scores 100% gives your engineering team nothing to optimize. An eval that scores 0% is equally useless. Aim for a 25% to 50% success rate so there is room to run. 4. Old benchmarks are already saturated - MMLU, HellaSwag, ARC and the rest were built for a simpler question-and-answer world. Frontier models now score effectively 100% on them, which is why you have to keep building new evals and throwing away old ones. 5. Writing an eval is mechanical once you understand the problem - Come up with roughly 100 prompts that match the real distribution of tasks. The hard part is not the writing. It is deeply understanding the domain first. 6. Subject matter expertise drives everything - The cystic fibrosis and congenital heart disease evals worked because Daniel understood the genetics, not because of any template or tool. There is no eval template the way there is a PRD template. 7. Modern evals are agentic, not just Q&A - The genetics eval hands the agent a file with billions of variants and asks it to find the cause of a disease. This is a task, not a lookup, and it mirrors how real AI products now work. 8. Find the model ceiling on purpose - The easy cystic fibrosis case gets solved by most models. The harder digenic congenital heart disease case exposes where even strong models fail. Knowing the ceiling is the point of the exercise. 9. Sample multiple times before you trust a result - Models are non-deterministic. Run the same task several times so you understand the real distribution of outcomes rather than a single lucky or unlucky pass. 10. Meta and Google build products very differently - Google is seen as more engineering-led, Meta as more product-led and far more aggressive culturally. Daniel worked on both Gemini and Llama and saw everything from Llama 3 highs to Llama 4 lows. --- 👨‍💻 Where to find Daniel McKinnon: LinkedIn: https://www.linkedin.com/in/daniel-mckinnon-8414649/ Twitter: https://x.com/danielmckinn0n 👨‍💻 Where to find Aakash: Twitter: https://www.x.com/aakashg0 LinkedIn: https://www.linkedin.com/in/aakashgupta/ Newsletter: https://www.news.aakashg.com #AIProductManagement #Evals --- 🧠 About Product Growth: The world's largest podcast focused solely on product + growth, with over 200K+ listeners. 🔔 Subscribe and turn on notifications.

Daniel McKinnonguestAakash Guptahost
Jul 27, 202656mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:002:41

    Intro

    1. DM

      The Metas and the Googles and all the other large companies have to reinvent themselves right now in the age of AI. Every single PM is gonna start building AI features. You cannot exist as a PM without understanding this.

    2. AG

      Meet Daniel McKinnon, former PM on the Llama models at Meta, now a startup founder and a master of evals.

    3. DM

      PMs right now from a career perspective are in a really tough situation. The average PM is an orchestrator, a motivator, and an analyst, but a lot of this is easy to do with AI.

    4. AG

      What really is the difference between product management at Meta versus Google?

    5. DM

      Meta is, like, a much, much more aggressive culture, uh, in many ways. Google is considered to be more of an engineering-led company, whereas Meta is more of a product-led company.

    6. AG

      If you're a PM who's never worked on an AI feature before, when would you be going through this evals process?

    7. DM

      You work at Pinterest, and you wanna have, uh, better image generation that doesn't look like slop, that actually pleases users. You need to think about from day zero, what does success look like?

    8. AG

      If it's the best way to measure it, we've gotta learn it, where should we start?

    9. DM

      Let me walk through-

    10. AG

      Before we get into today's show, please take a second to check that you're subscribed on YouTube and following on Apple and Spotify podcasts. If you want access to all of my favorite AI tools, I've gotten them to give you an entire year of their paid plans. Check out bundle.aakashg.com for an entire year of Bolt.new, Airtable, Speechify, Descript, Magic Patterns, Linear, Dovetail, Arise, and Mobbin. And now, into today's show. Daniel, welcome to the podcast.

    11. DM

      Yeah, thanks so much for having me, and, uh, it'll be fun to talk about this stuff.

    12. AG

      So I wanna start with your article. You had this provocative claim and this funny meme here. Do evals replace the PRD? What is the role of evals?

    13. DM

      Yeah, so I think it was, like, a little bit spicy in saying it replaces the PRD, because without a product strategy or a particular customer, like, your product is nothing. But again, that's a paragraph or even a sentence, depending on what the product is. The majority of most PRDs that I've seen in my career have spent most of the document talking about specifics, how it will work, how it will behave in certain situations, how the user can expect to get value from that. But that gets really turned on its head in this kind of gen AI world, where these products really need to do, like, everything, or at least a lot more things than previous products. And it's very hard to describe that, say, "Oh, this thing just does everything." And the best way to actually communicate what the product should do is through examples,

  2. 2:413:25

    What an eval actually is

    1. DM

      and that's what an eval is. It's really just, like, a trivia question for the model, and it's saying, "This is, like, the shape of the things the model needs to do well." And if the model does it well, it means it's getting these answers. And if it gets these answers, the users will probably like it. And if the users probably like it, let's ship it into prod and see if they actually like it with an online eval. But the key way to communicate how a product should work in this kind of gen AI era is, is it performing well on an offline eval? And if not, you either need to change the model, change the harness, or change the product. It's possible that what you wanna do is not possible with the models today, but it's better to find that out early with an offline eval versus just shipping it to prod and getting frustrated users.

    2. AG

      And for people who don't quite understand that nuance, what's

  3. 3:255:22

    Offline evals vs shipping to prod

    1. AG

      an offline eval versus just shipping to prod and looking at them?

    2. DM

      Yeah. This is a really, really important nuance, and I touched on it in this blog post. But when we usually talk about evals in this AI world, it's something that's run offline. There's, like, a little bit of gray areas in terms of RL environments and stuff, but think about it as, like, a pre-baked set of trivia questions that you ask the model. So for example, let's say you have a recipes website, and you want to tell users how to make their favorite kinds of ice cream. An offline eval would be a prompt set, say, 100 prompts of different ice creams that users might like, and the answer key would be, uh, a correct answer or a plausibly correct answer with a way to score whether it's good. So that's run offline during the development of your product, and you use that as a proxy for real user traffic. Once you do well on your offline evals, you can ship that product online, and you have a website, and it lets users generate ice cream recipes. And you think because you're a good PM and you really thought deeply about what the customer wanted, that performance on that offline eval set will reflect the satisfaction of the user in the online eval. Um, this obviously doesn't always happen, but that's the idea, and it's the best way to measure whether a gen AI product is likely to satisfy users.

    3. AG

      If it's the best way to measure it, we've gotta learn it. Where should we start?

    4. DM

      Let me walk through what's changed since I wrote that article. So the key thesis has remained true, and this has been now two years, in that an eval is the best way to communicate what your product should be doing and to explain to the engineering team working on making the product work what success looks like. This can be very, very challenging in an AI world when they, these products do so many different things that it's hard to necessarily understand what good is. But things

  4. 5:227:03

    How evals got more complex

    1. DM

      have gotten a lot more complex. So back in the day, evals were really simple. My blog post basically just covered these simple evals. They were question and answer. And this seems crazy to think back two years about what these models looked like and what the use cases were like, but this is before Claude Code and agentic coding and all of these crazy business applications that are getting built right now, uh, Claude Cowork and the like. Really, the core thesis two years ago was that gen AI was essentially a search replacement. I don't know if everyone remembers when Google's stock tanked because, uh, this was going to be the replacement for search, and fundamentally, these were, like, question and answer products. So what I have right now is the benchmarks that OpenAI reported on GPT-4, and you might remember a lot of these. MMLU, HellaSwag, ARC, Winograd, HumanEval. HumanEval is ironically an automated eval of Python. Uh, DROP, and- These are all just question and answers. So I dropped one example here from MMLU. If you know the actual brightness of an object and its apparent brightness from your location, then with no information, you can estimate, A, speed relative to you; B, composition; C, size; D, distance from you. And if you want to go ahead and look at some of these just for, like, historical fun, uh, they're, they're all on Hugging Face. So this was actually a very, very straightforward thing to do, is you had to think about the types of questions users would ask and the types of answers they would expect, and how to score that. So the key thing was just matching the questions to the domain of interest and scoring the answe- answers. So if we

  5. 7:038:27

    The mechanical process of writing an eval

    1. DM

      go back into my post here, we can see a little bit how I said to do this. And, uh, first thing is just figure out your problem, and it doesn't need to be perfect. What is a set of problems that your users might ask about? So the example I used was, uh, for example, generating recipes for videos. I guess I have food on my mind, 'cause I randomly came up with the ice cream, uh, example earlier. And new problem is I have a video on my social media site, and I wanna be able to generate a recipe that somebody can use to make that thing as they're walking the video. And you say, "Okay, this is really well-defined." Um, and now you have to measure h- how do you know if it's good? You know, it might be formatted right. It might have all the ingredients listed. It might be written in the right style. And then you select all of these components and figure out how to judge if it's correct or hypothesize how to judge it's correct. This can be a auto scorer. This can be another LLM. This can be a human. You can judge correctness many different ways. Once you have that, it's really a mechanical process to actually write the eval. Just come up with probably 100 prompts that are in this distribution. Can be less, can be more, but this is a typical size of an eval in a genI- gen AI world. And just send it through the model and figure out what is a hard prompt, what is an easy prompt, and have something that scores maybe, like, 50%, 'cause you have to have room to

  6. 8:2710:05

    Why easy and hard evals both fail

    1. DM

      run. If you create a very easy eval that scores 100%, there's no way for your engineering team to optimize on that. And if you create a very hard eval that scores 0%, you also don't even know if this is kind of possible with today's technologies. Um, and that's pretty much it. And once you have this, uh, you can give it to the team. They can improve on it. You can ship it out to users. If you achieve a score high enough, you can see if you're actually online performance matches what you expect it to offline. And then, uh, this, uh, blog has a bunch of examples of, like, you know, some kind of other nice tips. But why is this not that relevant today, or why does it need to change today? The real answer is that models have generally saturated QA. This isn't 100% true, but when we think of a good model right now, we don't think of one that can answer a relatively challenging high school physics question, like the, uh, example I gave above. We think of models that are, like, winning gold med- medals at International Math Olympiads. Like, there's almost no question that any human, beyond some super, super specialist, can ask the model to do and it not have a good answer back. And also, QA is not the most useful application now. I mean, I think we all remember this narrative that ChatGPT was gonna be the next great consumer app, and they were gonna get all these users, and it's QA, and you're answering all these problems, and, you know, Google did a generative search experience. But if you look at all the headlines in gen AI right now, it's not QA, it's agents. And this is actually reflected with how the labs communicate progress.

  7. 10:0512:35

    Ads

    1. AG

      Here's a quick word from our sponsors. If you're building anything that uses live data from the web, eventually you hit the same wall. An agent, a research tool, a trends dashboard; they all need fresh data. Scraping that data is the worst part. CAPTCHAs, proxies, layouts that change every week. It's a whole side project you didn't sign up for. That's where SerpApi comes in. SerpApi gives you clean, structured results from Google, YouTube, Bing, Google News, Google Scholar, and more. One API call, one clean JSON response. They handle the CAPTCHAs, proxies, and layout changes for you. Take the Google Scholar API as one example. Say you're building a research assistant or pulling sources for a literature review. You hit one endpoint. You get peer-reviewed articles back with full text or metadata, titles, links, publications, citation info, all of it, across publishers and formats. No scraping a dozen publisher sites and gluing the data together yourself. The same idea extends across the rest of their APIs. Real-time Google search for an agent, pre-classified images for training data, Google News for monitoring, 99.9% uptime, 1.2-second response time. Get started with 250 free credits. Link is in the description, or scan the QR code on screen. Thanks to SerpApi for sponsoring. Are you looking to up your AI product management chops? I highly recommend the AI Product Management Certification by Product Faculty. It has a 4.7 with 1,249 reviews on Maven for a reason. I myself took the course back in 2024, and it was awesome. Since then, they have upgraded it, so now you get to learn from product leaders at OpenAI and Anthropic. On top of that, Paul Hearn, author of The Product Compass, leads the build labs. So you will go from theoretical knowledge about AI PM to a very practical course. It's gonna help you identify AI leverage opportunities. It's gonna help you design trustworthy AI experiences. It's gonna help you systematically optimize outputs for accuracy and relevance, build rigorous evaluation suites, architect AI agentic systems that work, and select the perfect LLM for your use case. It's normally $2,500, but you get a discount when you use my link. The next cohort starts June 22nd and goes to August 9th, so do check it out with my link in the description. They have been one of my longest sponsors for a reason. I trust this product, and I think you should consider the cohort

    2. DM

      So back here, two years

  8. 12:3513:49

    Why old benchmarks are saturated

    1. DM

      ago, if you look at w- how OpenAI communicated progress, it was these five evals, I guess six evals. If we fast-forward to look at how Anthropic communicated progress for Opus 4.8, you can see they're using entirely different benchmarks. You don't see any continuity. It's partially because those evals are saturated, and Opus 4.8 would score effectively 100% on all of them. But it's partially because the task is really different. You'll notice we have agentic coding, agentic terminal coding, multidisciplinary reasoning. This is actually some agentic reasoning, agentic computer u-use, knowledge work. This is actually agentic knowledge work and agentic fin-fi-financial analysis. All this means is what the core model task is, is no longer to get a prompt from a user and come back with an answer, but it is actually to get a task from a user that requires many, many steps. Some of these steps might involve just thinking, which is called reasoning in this world. Some of it might involve tool calling, something like search. Some of it might involve more advanced tool calling, like something we'll go over today. And this is a totally new paradigm

  9. 13:4915:32

    From QA thinking to task thinking

    1. DM

      of writing evals because you're no longer thinking about QA, you're thinking about tasks. Fortunately for us, the framework is largely the same. We still have to define the problem. We still have to be good PMs and know what we're solving. We still have to collect representative prompts. I call this Goldilocks style. Again, they can't be too hard, and they can't be too easy. There has to be some room to run. A typical good eval will have something like 25%, 50% success rate, and then over, you know, months, that will go to 100%, and then you'll have to throw it away and create a new one that is harder. And then you also have to figure out how to score. And one thing that has changed is QA is relatively easy to score with humans worst case scenario. There's exceptions to this, of course. The reason these models are so bad at things like creative writing is because it's hard to score, and there's different preferences, and, um, different users like different things, and why they're so good at math and coding is there is, like, some right answer, and this is much easier to hill climb. But for, to some extent, QA style questions, you can ask human raters to review worst case scenario if you can't find a better way to score it. With agentic work, it's much more challenging because the time horizon tends to be very long, and the final output is a collection of many, many, many steps that it took. Some steps could be correct and lead to the wrong outcome, some steps could not be correct, and it's just from a labor perspective and a, um, like, defining success perspective, it's much, much more important to get something that can be automatically scored, uh, to make more of these rollouts and understand how you can do more experiments. But again, it's largely the same, and the tasks are much longer time horizon. So kind of the goal of this

  10. 15:3217:59

    Building an agentic eval in real time

    1. DM

      podcast is to walk the audience through creating an actual agentic eval in real time. I want to caveat this. This is a little bit pre-baked. It is unrealistic in 45 minutes to come up with a brand-new eval. This is kind of like a weeks or months problem of deep thinking, but we're gonna kind of pretend, and we'll, we'll go through some of the steps together. So first problem, I want to measure and improve the model's ability to help with clinical genomics. This is a problem that I care deeply about. It's one that I think, uh, can improve the world, and it's something that I've launched a new startup to solve. And the problem is, is that, uh, whole genome sequencing has become the absolute gold standard in diagnostics in NICU settings, so for sick babies. Unfortunately, interpreting the results of a whole genome sequence is very labor-intensive, and it limits access to this life-saving technology. So I wanted to see if I could distill some of this human expertise into a model to help broaden the accessibility of this technology. So first thing is, like before, we have to deeply understand the problem, and I'm not going over these flow sheets, but this is just kind of how complex this is. Generally, you start with the raw reads off the sequencer. You do a lot of processing work to identify how the particular patient differs from the reference human genome, and then you do another set of work to determine whether those changes to the genome matter. For example, if my genome were sequenced, I would get about a billion reads that are 150 base pairs long. They would come out in a giant text file, and I would need to transform that into a diagnosis that says that this gene may or may not be responsible for this patient's condition, and there's a very structured way of doing this. So I'm not gonna spend too much time on this here because, uh, we're gonna go over some real examples, but you really need to deeply understand the problem. This is why people like Anthropic and OpenAI are hiring investment bankers, accountants, lawyers. You see job ads for all these vertical specific teams.

  11. 17:5918:49

    You must deeply understand the problem

    1. DM

      You must deeply understand the problem. You will unlikely be successful in creating an eval for some topic if you don't have some background in it or haven't really educated yourself on it. So then let's say we've understood the problem. I think I understand this problem pretty well, and by the end of this podcast, you will too. Let's go to the prompts. So again, we want to find this, like, Goldilocks set of prompts. So the first thing I like to do is just start with something easy. So you want to make sure the model can actually do this. And when I say the model for agentic stuff, I'm usually talking about the model plus the harness. So I'll use those words interchangeably, but this is a ... frontier model harnessed in a way that it can use these tools and it can do this reasoning and it can come back with a solution. So for this problem of genome interpretation,

  12. 18:4923:11

    The cystic fibrosis eval

    1. DM

      I picked like one of the easiest genetic diseases possible, and this is cystic fibrosis. This was something that we have known the genetic cause for quite some time, and there are like canonical genes that cause cystic fibrosis. So to save you the effort of me googling for this, I just had the link right here, and let's just go ahead and look at this. So what we see here is the canonical cystic fibrosis mutation in ClinVar, which is an NIH database for, uh, a lot of genetic disease. And what we see here is it's got four stars and three stars. This really should be four and four. This is like the canonical, uh, genetic defect for cystic fibrosis. So I'm gonna make sure that my agent can actually get this before I go forward. And so this is like the easy thing to start. So we're gonna have our agentic genetics eval, and we're gonna say gene, and we're gonna say CFTR2. This is the, again, canonical gene. And then we'll say variant. And what a variant is, is how a particular gene is mutated. So right here, this is, again, this is a somewhat niche eval, but what this is saying is that on this particular transcript of CFTR, at this position, there is one base deleted, and what that results in is the five hundred and eighth phenylalanine deleted. So this is the actual variant that's gonna exist in our eval. Sounds good. Okay. So what we see is that this is the exact variant that we care about. And what you'll notice is this is like very complex and nuanced, and this kind of comes back to the absolute first point I was making, is you really need to know the space to do these evals. A lot of the evals are kind of like picked up. Like, if you're trying to come up with an eval for Python coding, like many of these are pretty good now. And for any of you who have used these models, saying Python coding is a solved problem is a little bit of a strong statement, but it is a very, very well understood and well-characterized problem. So we will go to this, and we're gonna say, okay, so this is the thing we wanna know, and I should add a column here, and this is the phenotype, is cystic fibrosis. [lip smack] So what you see here is I'm starting to build a table of question and answer. So the question is, I have this phenotype of cystic fibrosis, which, uh, is, is a lung disease and, you know, we'll describe exactly what happens there. And then we have the genome of this patient, and then we have the answer, which is this variant. So let's walk through like how we would do that. And when you're constructing these evals, you're gonna really, really, really use gen AI a lot to construct them. So what I'm gonna do is now I'm gonna go over to a terminal window I have open here, and I'm using Codex. Any of the tools will work. I have it on 53 spark low because I want it to be fast for this demonstration. But, uh, you know, you can use any, any model, any- anything you like, and for this task, this will be fine. So what I have here, and I pre-baked some of these just to kind of make it go faster, but I wanna walk through any step anyway, is I have actually two files here that are representing my genome. And what I want to do is create a synthetic version of this genome that has these variants that have the question and answer through this agentic flow that I want to get at. So what I'm gonna do is I'm gonna say, "Please add," and then this is gonna be this variant, "two," and this is a small variant, so it's gonna come through here and create a new file in a new folder, and we're gonna call this, uh, Dan CF live.vcf.gz, and we're gonna have this be Dan CF live. And we're actually asking AI to help us create

  13. 23:1127:32

    Running the variance file eval

    1. DM

      the eval for AI. So we're gonna do this, and what the model is going to do is a variance file is literally just a text file. Actually, we can see what it looks like here, uh, just for fun. So while, while this is running, let's just take a look just so we know what these variant files look like. And, uh, this is, this is fine. We don't need to show all of it. But what we basically see here is a chromosome. So, uh, you might remember from things like 23andMe, we have twenty-three chromosomes, so this is one. This is the biggest one, and this is a position. So your chromosomes have different number of bases. You might remember we have like three billion bases in our, in our genome. And then these are swaps. So what we see is in this position, a reference human has C, and I have a CA here. So that means that, um, you know, during something I inherited from my parents or maybe something that emerged during my development, um, I got a, a base swap here. And then there's a bunch of metrics around quality, how real it is. You, you might remember you have two copies of each gene, so this is actually heterozygous, meaning only one copy is impacted. And, you know, really, it's just a text file of all of the letters in your alphabet. And what I'm saying is I want to add, uh, okay, this is still running, and if this is still running in a while, I'll just use the pre-baked one, is I just want to add this particular variant into my genome to see if our system can catch it. And This is what Codex is doing right now, and I actually don't know why this is taking so long because this is like a one-liner. Um, but-

    2. AG

      [laughs] Find the line, I think or something.

    3. DM

      Yeah, yeah. Right, right, right. I guess I probably-- I didn't wanna make this demo too pre-baked. I thought about it, should I just have a, uh, a, a, a Python program that just does all this for you? But then I'm like, that would not help the users at all because when they're constructing their own, they would know how to do it. So just believe me that this will work, and for the sake of time, we'll go to the pre-baked ones.

    4. AG

      Mm-hmm.

    5. DM

      So, uh, so basically what's gonna happen is that Codex is gonna add... This is in chromosome-- Where is it? I'm actually not sure. But in whatever chromosome the cystic fibrosis gene is in, it's just gonna add one row, and it's gonna say, "We're gonna have a deletion." So instead of having like a CA here, it'll just have a C. So like here's a deletion. You'll notice that we've, we've lost an A. We went from TA to T. And then it's gonna have a, a fake cystic fibrosis patient. So now let's just check to see how this works, and we'll say, uh, use our agent. And again, these are all agentic evals, and I'm using Codex as our agent, but you could use anything, Cline, OpenCode, Claude Code, your custom harness, any way you could to get these agents to actually operate. And in fact, even ChatGPT and Claude, actually in the web UI, they use agents right now. They're not just model in, model out. They're model in, reasoning tools, everything. Okay, cool. All right. So the agent finished here. So we see we added this in a record. Okay, and we see here it is. It's in chromosome seven, and you'll notice this is deletion. It says TCTT, and instead it's a T. And so this means that this is like the canonical CF. So we're going to see if our agent is able to do it, and we're actually gonna try a few different agents because one thing I mentioned is you wanna score like, you know, 25 to 50% on these evals. You have to think about what tool are you using, the, you know, Mythos-5 for some biodefense thing, then it's gotta be really, really hard. Or maybe this is something that for infrastructure reasons or cost reasons, you need to use a very small model. You need to use Haiku or something like that. So we're actually gonna try these simultaneously on a few agents and see what happens. So we're gonna say inside-- What was this file we had? Uh, Dan CF Live. Again, this is the one that we just made.

    6. AG

      Hmm.

    7. DM

      We have the genome of a patient suffering from... And let's just very quickly copy and paste some cystic fibrosis

  14. 27:3229:44

    Testing the agent on the disease

    1. DM

      symptoms. Uh, this one. We suspect cystic fibrosis. Please find a genetic cause. And then we are going to-- we're actually gonna copy this prompt so we can use it across multiple agents. So now we're going and we're, we're, we're checking GPT 3.5 Codex Spark Low. Uh, we have a few other tabs open. So I mentioned, you know, maybe we wanna actually see if Haiku can do this. So we'll just upload this. And again, we're going to the CF Live. And let's also try ChatGPT 5.5 Extra High. And I'm not gonna use Pro 'cause it will take too long, and this is also a relatively easy task. I suspect all the agents will get them. Okay, great. So we are cooking with Haiku, and we are cooking with... You can see what's al- al- already happened with our first agent, is after one minute of thinking, you've actually find the CFTR mutation pattern consistent with this deletion, right? This is the canonical cystic fibrosis gene. So what this is telling us, going back to our steps, is this task is not too hard for these agents, at least in the easy case.

    2. AG

      Hmm.

    3. DM

      Especially not with a powerful system like Codex. We'll see if Haiku gets it. I suspect Haiku will also get it. But you can see Haiku even itself knows the CFTR region is, uh, important. So while that's cooking, let's go back to our next step. So again, we start with something easy just to make sure it's possible. I knew that this was possible, but if you just told a layman and say, "Hey, could, uh, you know, an AI agent find the canonical cause of cystic fibrosis inside a file with billions of variants?" They might say yes. They might say no, right? You just need to know. You need to kind of try it to get a sense of if it's possible.

    4. AG

      Quick thought experiment for you. Is there anything in this video you should be trying on your own? If there is, try it, take a screenshot, post it on LinkedIn

  15. 29:4433:07

    Ads

    1. AG

      or X, and tag me. I'd love to see what you're learning. Now, a quick word from our sponsors before we get into the back half of the pod. If you've worked at any company bigger than 30 people, you know this one. The CEO sets strategy. By the time it reaches the people actually doing the work, it goes through three or four layers of translation. Half of it gets lost, and nobody finds out until the quarter is over. That's the problem Ariso is built for. It's an AI operating partner for every manager and team. It connects to where work actually happens, the meetings, the messages, the docs, and it turns all that fragmented activity into a clear picture of execution. Managers get real coaching grounded in their team's actual work, not generic advice. Teams stay aligned with strategy as it changes, not as it was last quarter. And leaders see where execution is drifting in weeks, not in the postmortem. One shared memory for the whole org. Everyone finally working from the same picture. If you lead a team, check out ariso.ai/aakash. That's A-R-I-S-O.A-I/A-A-K-A-S-H. I want to take a second to talk to you about the fourth cohort of Land PM Job. I trained thirty students in cohort one, fifty students in cohort two, and seventy-five students in cohort three, and I am bringing back the program for cohort four. It starts in August, and it lasts three months, where you're going to have intense sessions. A Monday morning session where I go over your resume, behavioral interviews, LinkedIn. On top of that, Bart Jaworski is going to be teaching you the PM fundamentals in 2026, how to write AI PRDs, how to AI prototype with Claude Code, all of the key skills you need to freshen up your knowledge for this market. And Ankit Virmani is going to be teaching you AI product management. He is an AI product manager at Uber, and he is going to teach you how to build AI features that actually work successfully. On top of that, Prasad Reddy is going to be doing one-on-ones with you for mock interviews, LinkedIn review, candidate market fit review. So it is a full package. It is three courses in one for one low fee. So join at landpmjob.com. Today's podcast is brought to you by Pendo, the leading software experience management platform. McKinsey found that 78% of companies are using gen AI, but just as many have reported no bottom line improvements. So how do you know if your AI agents are actually working? Are they giving users the wrong answers, creating more work instead of less, improving retention or hurting it? When your software data and AI data are disconnected, you can't answer these questions. But when you bring all your usage data together in one place, you can see what users do before, during, and after they use AI, showing you when agents work, how they help you grow, and when to prioritize on your roadmap. Pendo Agent Analytics is the only solution built to do this for product teams. Start measuring your AI's performance with Agent Analytics at pendo.io/aakash. That's P-E-N-D-O.I-O/A-A-K-A-S-H.

    2. DM

      But then if you know that the easy thing works and, you know, we have al- already have early evidence the easy things works, you have to, like, establish the ceiling. It's like, what is, is, is the hard thing working? 'Cause, like, if it's just totally saturated, then, like, w- what's the point of even having eval? This task is al- already solved. Um, so I'm, I'm gonna show off something that's pretty hard to do today. And what we have here is a recent paper, so this is from, uh, last year and, and I will make this bigger. The authors here are deciphering the digenic architecture

  16. 33:0736:07

    The congenital heart disease eval

    1. DM

      of congenital heart disease. So what does this mean? This means congenital heart disease is if you're a baby and you're born with problems with your heart, and digenic means it involves two genes. So single gene, single variant genetic diseases are actually sometimes a solved problem, like with cystic fibrosis. Not all cases, but many cases like this one are totally understand. Digenic genetic diseases are, like, a very, very new thing that people are studying. So this is, like, a very hard task to do. And even though this is published, and in theory, an agent should be able to search the internet and find every publication and, uh, you know, deeply understand all of this, it's actually not that simple. They're not perfect, and they actually need a lot of guidance, which is why there are a lot of these companies, including my own, that are called, like, harness engineering companies or vertical AI companies or agentic AI companies, because you need some specialized capability to be able to, um, have the LLMs do stuff like this. So let's briefly return to our, our agents, and let's just make sure they got it. Okay, so, uh, Codex with GPT 5.3 says, "Okay, this," if you recall, this is the deletion that we added. Boom. I gave it the phenotype. I gave it the genome. This is correct. So how would we mark this correct? We would actually probably have another LLM. I'm not gonna do this right now for the sake of time. Just compare my scorecard, this is the correct answer, with the response the model is giving right here. And then we can check Haiku, and even Haiku... Oh, wait, did Haiku not get this? Okay, so, so this is interesting. This is actually harder than I would have thought. I would have expected Haiku to get this because this problem is so easy.

    2. AG

      Mm-hmm.

    3. DM

      But you'll notice what Haiku says is there's forty-eight variants spanning the gene. So it's looking at the gene, but it fails to actually find the particular... Oh, this is so interesting. It also hallucinates a hemizygous large deletion. So coming back to this, coming back to our point, is start with something easy. I thought I started with something easy here. It's a good thing I did this because if I were benchmarking Haiku, this is too hard, and I'd have to make it even easier. And the things I could do to make it even easier would be potentially, uh, you know, limit the region of the genome of interest, give it more hints, maybe provide access to more, uh, external information more easily. But we can see that Haiku even fails this easy task, and I would be absolutely shocked if 5.5 did. Okay. It, it, it's not finished yet, but you can already see that it, it found the correct answer. So if we were to score this, we could easily have an LLM say, "Okay, Haiku, this is not correct. This does not match what I have in this table. This is correct, and this is correct." So that's kind of... And then in our, in our spreadsheet, we would just say, you know,

  17. 36:0740:39

    Moving to a harder problem

    1. DM

      Haiku bad, others good.

    2. AG

      Mm-hmm.

    3. DM

      Um, and but now let's move on to something where we want to understand the hard cases. And again, I unexpectedly actually picked out a hard case for Haiku, but this paper is quite challenging, and I believe it is unlikely that any of the models will solve this. So in the supplementary information of this paper is a table, and it is a list of patients, and a, a proband is a medical term for the patient you're evaluating, and it has digenic causes for congenital heart disease. So this particular patient has ACACB, I have no idea what this is. Some gene. It's het, meaning it only has one copy of this variant, and MYOCD, which is also het, which is one copy. And these researchers discovered that the combination of these two diseases leads to congenital heart disease. So let's see if the models can figure this out. So what we'll do again is we'll go to our table and our phenotype-

    4. AG

      Is it two diseases or is it two, uh, abnormalities in their DNA?

    5. DM

      Yeah, that's a great question. It's one disease. It's a congenital heart disease, and I don't know exactly which one it is from this paper. Um, and, you know, we could read the paper and figure out exactly what the phenotype is, but it's some defect with the heart. And what's unusual about this and why this is hard is it's two heterozygous variants on two different genes that is causing this single disease. So it's complicated. So our phenotype here is congenital heart disease, and our gene here, we have two of them. One is this guy and, oops, and the second one is this guy. And then the variants are these guys. And this is... You'll notice the notation is a little bit different, but this is something that you'll just have to deal with in these evals is like, you know, no matter what you're doing, because these tend to be in very technical, specialized domains at this point, you know, no one wants evals for, for boring stuff like ice cream flavors. Y- you'll just have to get comfortable with all these different mutations. And now let's go and let, let's try this again. So what I would do is I would say something like, "Please add these to dan.dvariant.vcf." But I'm actually not gonna do this because you already saw how this worked and basically how the VCF file was structured. It would add these two rows, and for the sake of time, I've already done it. But then let's go and let's check and see, oh, how do we do on this use case? So in this case, I've already pre-baked it, and I'm going to say, "This is Dan CHD," and I'll say, "This contains the genome of a patient with congenital heart disease. Please identify the genetic cause." Okay. And while we're going to have this one running, we're gonna try our other two agents just to see how they do. Um, we can almost guarantee that Haiku will not get this because it didn't get the much, much easier task. But for the sake of completion, uh, completeness, we'll, we'll do this as well. And, and then we'll do this with, um, GPT 5.5 extra high as well. And again, we would do probably-

    6. AG

      There is probably like some non-deterministic nature, right? Like do you need to like test the same model a couple times to just see if like maybe two out of three times it gets it right, or is that not important?

    7. DM

      Yeah, that's a really good point. Um, so this is a question about sampling. Um, so sampling is actually really important. And you might remember, um, like all of this old research where you would basically sample for good traces, and this is kind of what like RL environments do, is you do a rollout, you do a rollout, you do a rollout, and then you get the correct answer, and boom, you give it a good reward for that. And the key thesis here is inside the weights of the model, the right answer might live there. It just might not get the right answer each time. So when you're

  18. 40:3944:10

    Why you sample multiple times

    1. DM

      doing these evals, you do want to try multiple times. In this particular case, I actually know from having done it that sampling has very little effect, and it's essentially deterministic based on model capabilities. I've seen slightly the same model get to the same conclusion with slightly different approaches. But in general, sampling in my experience is less important than it used to be, where sampling used to be a big deal. Um, like if you look at, uh, let's just look at, this is a funny story. Um, let's look at Gemini Ultra, uh, scorecard. So if you'll remember Gemini Ultra... Oh wow, did Google actually bury it?

    2. AG

      [laughs]

    3. DM

      Okay, here it is. This is-- So you'll remember way back in 2023 when people thought Google was kind of out of the AI race. I actually worked on this model, so I know this story very well. Google released Gemini Ultra, which was, uh, I believe it was a 660B dense model, which was crazy back then. That was like one of the largest dense models ever trained. And they released this scorecard, and what you see is this. This was very controversial, and this comes back to your question about sampling. If you remember MMLU, this used to be like the canonical benchmark for LLMs, and let's just go back to look at what a question is to remind you is just simple question answer. If you know the actual brightness of an object and its apparent brightness from a location with no underest-- information, you can estimate this, okay? Models used to be bad at this, which is hilarious because this seems so distant right now. And Google wanted to be the best, and GPT-4 was the best at this point and got 86.4% on MMLU, and it was on five shots, meaning it had five samples and it picked the best one. And that's why you see five shot, three shot, three shot, 10 shot. It really is like kinda like a way of cheating. It's like how many times can you sample, um, from this model? And what you see is that Gemini Ultra actually had 32 shots, so they got more shots on goal. And actually, now that I'm remembering this, this actually might be pre examples, but i- if this is actually not the number of examples in the context window and just the shots or, or, or the number of times sampled, it, it, it doesn't really matter for the sake of this argument. But The answers to MMLU might be inside the model weights, but it might just be not enriched enough in terms of the probabilities. So by sampling more times, you actually get a higher chance of getting the correct answer. So this used to be a really big thing back in the day. Today, I don't think it's as a big thing. The labs don't really publish anymore. I don't think there's a lot known about this. In my personal experience, I've not found that running the same prompt through the model multiple times generates different answers. In fact, I don't know if I've ever seen that for this task, but it's a really, really good thing to do, and you should test for your use case. Little, little side, side, uh, conversation while we, um, look at the answers. And we've got all these cooking. These are all cooking, and we can see that 5.3 Spark finished first. This is actually why we did it, and what you'll notice is it's totally wrong. It found multiple variants, TBX1, MYH6, JAG1. You'll notice these aren't even genes of interest for us. These variants aren't even relevant. No high confidence variants. Um, so what we've done here is we've done the second step in our process, is we've established the floor.

  19. 44:1047:04

    Finding the model ceiling

    1. DM

      We found something easy, that cystic fibrosis gene. Now we found the ceiling is this model didn't get it, and plot twist, no model on the planet gets this without, like, a very, very strong harness. Again, I'm working on that very strong harness, so, you know, we can make systems get this, but this is hard. And, uh, then you just kind of go through, and it's almost like a binary search process where you say easy, medium, hard, and then just assemble a list of prompts. You know, you might have 100 of these, which again would be phenotype, gene, phenotype, gene, phenotype, gene, and then you understand how the models do on it. And then you're done, and then you have your eval, and then you understand what is good enough to actually ship product. So if you're scoring 50%, is that good enough to ship your product? Probably not. So then you look at which phenotypes am I better at? Maybe you put guardrails on the product to make sure that it only will answer the types of questions that it can get 80% on or something like that. That's a product decision. That's a product manager's decision to do that. And then you also hand all the hard ones to the research team, and you say, "Hey, guys, you didn't get this. Fix the model or fix the harness and make sure that it can get these in the future, so I can ship a product with these capabilities." And, uh, not shockingly, Haiku is totally off base. Uh, clearly Haiku is not good at all at this, um, task, um, so that's surprising it missed the first one, not surprising it missed this one. Um, and we again, we have... Oh my God. Okay, this is so interesting. Okay, so this is actually a good example of something that came out and sampled a second time and worked, 'cause I actually tried this. We were just talking about sampling. But we can see right now that GPT 5.5 extra high this time actually did identify this digenic pairs, and it did almost certainly find the paper. Yeah, so it did find the paper with all these digenic pairs. So this is actually a very interesting reasoning trace, where it was able to turn this congenital heart defect phenotype into a search for a very specific paper and then pull out the results from this paper. So actually this is quite impressive from, uh, GPT 5.5. But, uh, this is, this is correct. So in this case, I would have to find an even harder one if I were benchmarking this model in particular. But this is basically the, the key set of steps, and I th- don't think we have time to do a bunch more, but it's basically running through all the different types of scenarios and then coming up with prompts that will challenge the model but not totally stump the model. So yeah, with that, that's, that's how you write an agentic eval, and here's two lines in our new one.

    2. AG

      Wow, so it's just a spreadsheet, and the key thing

  20. 47:0450:01

    It is all subject matter expertise

    1. AG

      here is the domain subject matter expertise. It's not like how it's written or anything like that. You're not giving us an eval template like we might have given people a PRD template before. It's really the subject matter expertise that's driving all of this.

    2. DM

      Yeah, exactly. There are a lot of companies over the last, you know, N years who have tried to build better tools for evals, and I'm not saying that tools for evals don't need to exist. There's plenty of ways to improve. But when you're creating evals like this, it is literally just prompts, responses, and ways of scoring whether response is correct.

    3. AG

      Fascinating. So just to bring it all back full circle, if you're a PM who's never worked on an AI feature before, when would you be going through this evals process, and when wouldn't you, and how would you be using it?

    4. DM

      Yeah, that's a, a good question. So if you've never worked on an AI feature before, I would actually try to find somebody who has, who can help you through this. This is deceptively simple. I made this really simple because we had 45 minutes today. But this, and I don't know why it's so complex, honestly. I had many, many conversations with people about how to build evals, but there's just something kind of like taste-based or, or nuanced about how to build them, and, uh, you know, it is what it is. But fi- find somebody who can help you. But I would say you start from the beginning, is like you are building an AI feature, like, I don't know, I'm just making it up. You work at Pinterest, and you want to have, uh, better image generation that doesn't look like slop that actually pleases users. So it's like an image generation feature. You need to think about from- Day zero, what does success look like? What is unique about those Pinterest users? What do they want to see? And you need to translate that. You can't just write down a PRD and say they want beautiful kitchens. You have to explicitly define what a beautiful kitchen is, and not in words, in examples, and a way to score those examples. And I am not in the image generation space, so I don't exactly know what that looks like, but there are many, many, many examples like this where you need to understand what the user wants and transl- translate that into prompts and responses and a way to score those responses, and that is the first thing you should do when you're starting to build a new AI feature.

    5. AG

      Okay. So just like you have ramped up your expertise in the genomic space, if you were tackling that problem, you'd go learn, you'd go talk to people who have built Imogen evals. You'd understand, "Oh, this is how I build an LLM judge that generates images of this type," and then that would really be the basis for your eval.

    6. DM

      Yep, that's correct.

    7. AG

      Okay. Wow, there's so much, so many layers. I've done, like, five or six episodes on evals, but I think this was one of the most tactical that really helped me understand how things change, and I think that's a function of your experience, which I wanted to talk about for a little bit. So your evals piece crossed my radar. I think another really interesting

  21. 50:0152:59

    Product management at Meta vs Google

    1. AG

      piece you wrote about was product management at Meta versus Google. You've worked on Gemini, you've worked on Llama, you've seen both of these cultures. What really is the difference between product management at Meta versus Google?

    2. DM

      Yeah, so I would caveat that and say I wrote this, like, two and a half years ago, and I was at Google three and a half years ago, I think. So a lot has changed. When I was at Google, Google was a dead company. I think the stock fell to, like, $80, and I think it's, you know, 300 or 400 right now, and, uh, they've really changed how they think about things. I am a boomerang at Meta, so I think I spent seven years there in, in, in total. And, you know, I saw everything from Cambridge Analytica lows to highs of, like, Llama 3 really wowing people to lows of Llama 4 disappointing. So I, I saw a, a very large spectrum. And what I would say, like, my key takeaways for what, uh, Google versus Meta was like is Meta is, like, a much, much more aggressive culture, uh, in many ways. I think that it comes from, like, the founder leadership of Mark Zuckerberg is he is the last man standing, well, I guess besides Elon. But he is the last man standing who's got, like, you know, the founder leading a FAANG company who has utter and absolute control, who's gonna do what he wants. And sometimes it's really empowering because he says, "This is super important to me. You have all the resources in the world, and you should go do it." And sometimes it's, like, not what you want. For example, I worked on Llama 4. Llama 4 had problems. I think the main problem was actually how it was evaluated. Huh, funny. Those evals are important. If you want to read about that story, Google it. I had nothing to do with that. And I, I loved working on Llama 4, and Mark just said, "You guys all suck. You need to go find new jobs." So basically, the whole Llama 4 team is gone because of, you know, Mark's, his decisions. So I think that, you know, that, that cuts both ways. Uh, I'd say my overall preference is for, like, a very, like, high-conviction, founder-led company. Um, I actually have, like, incredible respect for Mark. I've only, you know, met him a couple times, and, uh, every time has been just like, "Wow, this is, like, a really smart guy." But it creates a lot of problems too, right? Because, you know, Google is much more consensus-driven. I think Google has much weaker product management function, at least it did when I was there. So, uh, you know, Google is considered to be more of an engineering-led company, whereas Meta is more of a product-led company, at least that historically has been the case. And, um, it was a really great experience working at both places. I think I learned a lot from both places. I left both places with a lot of friends. And, uh, yeah, if you want to read, like, kind of, this blog post actually went pr- pretty viral. If you want to read, like, kind of a interesting snapshot of what it was like in, say, 2023 between both of the places, uh, you know, give it a read.

    3. AG

      Highly recommend it to everybody. As you guys can see, I'm itching to ask many more questions, so Daniel, we're gonna need to have you back. Before you go, tell us a little bit about your startup.

    4. DM

      Oh, cool. Yeah. So I've

  22. 52:5955:33

    Building Gamoff Labs

    1. DM

      started a company called Gamoff Labs, and as I hinted at during these evaluations, the core problem I want to solve is to make it much, much easier to get whole genome sequencing into every single NICU in the entire world. This is the absolute gold standard of helping sick babies. Um, there's overwhelming clinical and economic evidence that it's effective, but the problem is it's just too damn hard and expensive. So where you see this used is in places like Stanford, Boston Children's, uh, CHOP, and these are the absolute top facilities in the world. And I want to see them in, you know, rural Arkansas, rural India, you know, rural China, all the places where, um, this, this kind of life-saving technology is not being harnessed. And my key thesis is that a lot of the human work involved in interpreting these genomes could be augmented by AI. And we've already shown that using the system that we've built, we can identify variants that have never been discovered before. We've actually allowed one family to have a child, and they, they couldn't before because they didn't know. Again, this is small scale. We started this five weeks ago. So, you know, but one person is, is really crazy to have that impact on their life. And, um, yeah, like it's a deeply, deeply mission-driven thing. I think it's very, very interesting technically because it's all about building the best agentic harnesses. It's all about understanding how AI can help with biology. And if you're interested in joining me on this journey, we are hiring right now. Um, we're a very small team. Uh, ra- raised our pre-seed round and are basically planning on building, like, the operating system for rare disease and genomic medicine. And I couldn't be more excited to wake up to work on this every morning, and I would love, uh, I would love it if you would reach out if you're interested.

    2. AG

      Wow. So a lot of you guys I know, [chuckles] at least in my audience, you want to become that AI PM at Meta or Google. This is often the next step after that. So if You know, we always say the grass is greener at some point. This is where I've seen those AI PMs at Meta and Google go, just like Daniel, into starting their own companies, and that's actually the cool thing, is it helps prepare you for that. You can see how his own evals and deep AI knowledge has now applied to his startup. Daniel, thank you so, so much for lending your expertise today.

    3. DM

      Yeah, and I want to leave with just one parting thought, is the Metas and the Googles and all of the other large companies have to reinvent themselves right now in the age of AI. If you're inside these companies, it is

  23. 55:3356:46

    Closing thoughts

    1. DM

      very, very interesting to see how this classic consumer software building factory has changed. But if you come and you do a startup or you start your own thing, you get to build the future from scratch, and sometimes that's actually easier.

    2. AG

      We'll leave it there. See you all in the next episode. I hope you learned as much from today's episode as I did. If you can do one thing that's totally free that would help the show, it would be to check that you're following on Apple and Spotify podcasts, check that you've left ratings and reviews on those platforms, check that you're subscribed on YouTube, leave a like and a comment on this video, and then share it with your friends. We're trying to make better and better podcasts. After two years, we think we've gotten something pretty good going. So let us know what we can do to make it even better, who else we should interview, and we will put on the best shows we possibly can. Finally, don't forget my offer for the bundle. You get an entire year of my paid newsletter, plus my favorite AI tools, Bolt.new, Airtable, Speechify, Descript, Magic Patterns, Linear, Dovetail, Arise, and Mobbin'. That's $27,000 worth of value for just $150. So check that out at bundle.aakashgee.com if it interests you, and I can't wait to share our next episode soon.

Episode duration: 56:56

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode ztN6bE_FuQQ

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.