Aakash GuptaHow to Build AI Evals Step-by-Step | Daniel McKinnon | Product Growth
Episode Details
EPISODE INFO
- Released
- July 27, 2026
- Duration
- 56m
- Channel
- Aakash Gupta
- Watch on YouTube
- βΆ Open β
EPISODE DESCRIPTION
Every PM is about to start building AI features, and the ones who understand evals will separate themselves from everyone else. Daniel McKinnon, former PM on the Llama models at Meta and Google, builds a real agentic eval from scratch on screen. He shows why subject matter expertise, not tooling, is what actually drives a good eval. Full Writeup: https://www.news.aakashg.com/p/how-to-build-your-first-eval --- Timestamps: 00:00 - Intro 02:41 - What an eval actually is 03:25 - Offline evals vs shipping to prod 05:22 - How evals got more complex 07:03 - The mechanical process of writing an eval 08:27 - Why easy and hard evals both fail 10:05 - Ads 12:35 - Why old benchmarks are saturated 13:49 - From QA thinking to task thinking 15:32 - Building an agentic eval in real time 17:59 - You must deeply understand the problem 18:49 - The cystic fibrosis eval 23:11 - Running the variance file eval 27:32 - Testing the agent on the disease 29:44 - Ads 33:07 - The congenital heart disease eval 36:07 - Moving to a harder problem 40:39 - Why you sample multiple times 44:10 - Finding the model ceiling 47:04 - It is all subject matter expertise 50:01 - Product management at Meta vs Google 52:59 - Building Gamoff Labs 55:33 - Closing thoughts --- π Thanks to our sponsors:
1. SerpApi: Clean, structured search results from Google, YouTube, Bing, Google News and more. Get started with 250 free credits - https://serpapi.com/?utm_source=youtube&utm_campaign=aakashgupta_july_2026
1. Product Faculty: Get $550 off their #1 AI PM Certification with code AAKASH550C7 - https://maven.com/product-faculty/ai-product-management-certification?promoCode=AAKASH550C7
1. Ariso: Ship AI agents and features faster, with fewer regressions - https://ariso.ai/aakash
1. Land PM Job: A 12-week experience to master getting a PM job - https://www.landpmjob.com/
1. Pendo: The #1 software experience management platform - http://www.pendo.io/aakash --- Key Takeaways:
1. An eval is a trivia question for the model - At its core, an eval is a prompt with a correct or plausibly correct answer plus a way to score whether the output is good. It is the clearest way to communicate what your product should do in the AI era.
1. Offline evals catch problems before you ship - Test the model offline against a fixed prompt set before pushing to production. If it fails, you change the model, the prompt, or the approach before real users ever see it.
1. The best eval sits between too easy and too hard - An eval that scores 100% gives your engineering team nothing to optimize. An eval that scores 0% is equally useless. Aim for a 25% to 50% success rate so there is room to run.
1. Old benchmarks are already saturated - MMLU, HellaSwag, ARC and the rest were built for a simpler question-and-answer world. Frontier models now score effectively 100% on them, which is why you have to keep building new evals and throwing away old ones.
1. Writing an eval is mechanical once you understand the problem - Come up with roughly 100 prompts that match the real distribution of tasks. The hard part is not the writing. It is deeply understanding the domain first.
1. Subject matter expertise drives everything - The cystic fibrosis and congenital heart disease evals worked because Daniel understood the genetics, not because of any template or tool. There is no eval template the way there is a PRD template.
1. Modern evals are agentic, not just Q&A - The genetics eval hands the agent a file with billions of variants and asks it to find the cause of a disease. This is a task, not a lookup, and it mirrors how real AI products now work.
1. Find the model ceiling on purpose - The easy cystic fibrosis case gets solved by most models. The harder digenic congenital heart disease case exposes where even strong models fail. Knowing the ceiling is the point of the exercise.
1. Sample multiple times before you trust a result - Models are non-deterministic. Run the same task several times so you understand the real distribution of outcomes rather than a single lucky or unlucky pass.
1. Meta and Google build products very differently - Google is seen as more engineering-led, Meta as more product-led and far more aggressive culturally. Daniel worked on both Gemini and Llama and saw everything from Llama 3 highs to Llama 4 lows. --- π¨βπ» Where to find Daniel McKinnon: LinkedIn: https://www.linkedin.com/in/daniel-mckinnon-8414649/ Twitter: https://x.com/danielmckinn0n π¨βπ» Where to find Aakash: Twitter: https://www.x.com/aakashg0 LinkedIn: https://www.linkedin.com/in/aakashgupta/ Newsletter: https://www.news.aakashg.com #AIProductManagement #Evals --- π§ About Product Growth: The world's largest podcast focused solely on product + growth, with over 200K+ listeners. π Subscribe and turn on notifications.
SPEAKERS
Daniel McKinnon
guestFormer PM on Llama at Meta and Gemini Ultra at Google; founder of Gamoff Labs focused on genomic medicine and AI evaluation systems.
Aakash Gupta
hostHost of Product Growth with Aakash Gupta; product/growth creator who interviews experts and shares AI/product resources.
EPISODE SUMMARY
In this episode of Aakash Gupta, featuring Daniel McKinnon and Aakash Gupta, How to Build AI Evals Step-by-Step | Daniel McKinnon | Product Growth explores building agentic AI evals with domain expertise and scoring rigor Evals are a practical way to specify and communicate expected AI product behavior through concrete examples, often covering much of what traditional PRDs tried to capture in prose.
RELATED EPISODES