Skip to content
Aakash GuptaAakash Gupta

AI Product Metrics Mock Interview (Meta/Google/OpenAI Case)

We break down the complete AI product metrics framework. The 40-minute case walkthrough, visual framework approach, and why output metrics matter more than you think. Full Writeup: https://www.news.aakashg.com/p/ai-success-metrics-interview ---- Timestamps: 0:00 - Why AI Product Execution Interviews Matter 1:18 - The Underlord Case: Live Mock Interview Begins 4:29 - Why I Pulled Up The Product Live (Don't Skip This) 6:42 - The User Segmentation Push-Back 9:16 - Building The Visual Framework in Real-Time 12:52 - Value Enumeration: The 4 Core Values 17:24 - The Positive Metrics Bank 23:00 - North Star Selection: Why Exports Won 25:47 - Breaking Down The North Star (3 Vectors) 31:00 - Trade-offs & Guardrails Deep Dive 34:12 - The "Genie Metric" Curveball 35:01 - The Output Metrics Miss (My 8.5/10 Moment) 38:08 - The Power Move: Post-Interview Follow-Up Strategy ---- Key Takeaways: 1. Visual frameworks are non-negotiable - Draw your structure live. Interviewer needs to follow you for 40 minutes. Without visual anchor, even great ideas get lost. This separates you from every other candidate. 2. Always push back on user assumptions - Bart said "beginners." I challenged it. Underlord is on homepage, so metrics need to work for ALL users. This type of thinking turns 8/10 into 10/10 answers. 3. Build metrics bank BEFORE choosing North Star - Generate 15+ positive metrics first. Time, volume, adoption, discovery, engagement, retention, output. Then evaluate. Don't pick your favorite and work backwards. 4. Control for complexity in time metrics - Underlord might INCREASE time to edit because users do more. Create table: 1 tool = 3min, 2 tools = 4min. If Underlord takes longer for SAME complexity, that's a problem. 5. The output metrics mistake cost me 2 points - I forgot upgrades, renewals, referrals. Bart had to prompt with "genie metric" question. Always include input metrics AND output metrics. Non-negotiable. 6. North Star selection needs explicit reasoning - Don't just pick. Evaluate out loud. "Time to export doesn't work because... Number of tools is hard to operationalize because... Number of exports works because..." Show your math. 7. AI guardrails are different from traditional metrics - Hallucination rate. Support requests: no increase. Build evals that verify AI actually did what it claimed. This is table-stakes. 8. Break down your North Star by segments - New users vs power editors. Free vs paid. Short-form vs long-form. By equation: sessions × completion rate × exports per session. Makes it operationalizable. 9. The eval-driven approach for discovering failures - Ship to 10%. Collect traces. Review failure modes weekly. Create synthetic evals. Add new guardrails. This is the Hamel Husain / Shreya Shankar methodology. 10. The 30-minute follow-up wins jobs - Take your framework. Find gaps. Fix them. Mock up dashboard. Email interviewer. At Descript, you'll be the ONLY person who does this. Immediate differentiation. ---- 👨‍💻 Where to find Dr. Bart Jaworski: LinkedIn: https://www.linkedin.com/in/drbartpm/ Land PM Job: https://www.landpmjob.com 👨‍💻 Where to find Aakash: Twitter: https://x.com/aakashgupta LinkedIn: https://www.linkedin.com/in/aagupta/ Newsletter: https://www.news.aakashg.com #productmetrics #pminterview ---- 🧠 About Product Growth: Aakash Gupta's newsletter with over 200K+ subscribers. 🔔 Subscribe and turn on notifications to get more videos like this.

Aakash GuptahostDr. Bart Jaworskiguest
Jan 30, 202640mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:00

    Why AI product execution & metrics interviews are becoming core for top AI PM roles

    Aakash frames PM case interviews into product sense/design vs execution/success metrics, and explains why AI adds a distinct flavor. He sets the stakes: top AI labs increasingly test how you define and defend metrics for AI launches and agentic products.

    • Two main PM interview buckets: product sense/design vs execution/success metrics
    • AI companies increasingly ask AI-specific success-metrics questions (dashboards, launches)
    • This video is positioned as an end-to-end AI execution mock interview example
    • Context on why mastering metrics can be career-defining for AI PM roles
  2. 1:00 – 2:29

    Program pitch and transition into the live mock interview

    Aakash briefly describes the cohort program outcomes and directs viewers to the enrollment page. Then the video shifts into the live mock interview format with Dr. Bart as interviewer.

    • Cohort results and how coaching maps to real interview performance
    • Where to learn more (landpmjob.com) and timing/availability notes
    • Set-up: mock interview will demonstrate a 10/10 metrics answer structure
    • Hand-off into the interview environment and roles
  3. 2:29 – 2:53

    Case prompt: Descript’s ‘Underlord’—define success for a natural-language video editing agent

    Bart presents the case: Descript is launching Underlord, and the candidate must measure its success. Aakash immediately seeks alignment on what the feature is by inspecting the product live.

    • Prompt: measure success of a new feature called Underlord
    • Aakash asks to align on product definition before jumping into metrics
    • Live product review is used to reduce ambiguity and anchor metrics to reality
    • Underlord positioned as a natural-language alternative to manual editing
  4. 2:53 – 4:56

    Live product walkthrough: clarifying capabilities, tool access, and core job-to-be-done

    Aakash explores Underlord’s UI, suggestion flow, and the extent of its agentic tool access. Bart confirms it can access all Descript tools, reframing it as a language-first editing interface.

    • Underlord provides suggestions and can execute editing actions
    • Clarifies whether the agent can call ‘all tools’ vs a limited subset
    • Positions Underlord as an editing partner and a natural-language interface
    • Locks key assumption: broad tool access materially shapes success metrics
  5. 4:56 – 5:49

    Interview framework setup: clarifications → users → value → metric bank → North Star → guardrails

    Aakash outlines a structured approach for the entire case, signaling the roadmap to the interviewer. This creates a shared mental model and a clear sequence for metric selection and evaluation.

    • Step-by-step structure for AI metrics cases
    • Creates a ‘bank’ of positive metrics before selecting a North Star
    • Explicitly includes trade-offs/guardrails and a closing summary/dashboard
    • Uses the structure to keep pace and avoid missing key metric categories
  6. 5:49 – 9:23

    User segmentation debate: designing metrics for both novices and power editors

    Bart initially emphasizes beginners; Aakash pushes back that the feature is homepage-level and must serve all users, including high-fluency editors. They align that success metrics should not overfit to a single segment.

    • Initial target: new editors lacking skills/knowledge
    • Pushback: homepage/persistent agent implies broad user coverage
    • Recognizes power users may gain scale/time benefits from agentic editing
    • Decision: do not prioritize a single user group; segment in analysis instead
  7. 9:23 – 15:07

    Value enumeration: four core benefits that drive the metric landscape

    Aakash maps Underlord’s value to measurable outcomes, using the product to inspire concrete value statements. He distinguishes overall Underlord success from tool-specific quality metrics owned by other PMs.

    • Core values: reduce time-to-edit; enable more edits/outputs; help first edit completion; generate publish-ready info (e.g., chapters)
    • Acknowledges tool-level quality metrics (captions, chapters) but deprioritizes them for this case
    • Connects values directly to candidate success metrics
    • Begins surfacing risks like hallucinated edits and quality bars as future guardrails
  8. 15:07 – 19:34

    Building the positive metrics bank: efficiency, volume, activation, and feature unlocking

    Aakash translates each value into measurable signals, adding engagement/retention interpretations that fit an efficiency product. Bart adds the important angle of improving final output quality and unlocking features users wouldn’t otherwise use.

    • Efficiency: reduce time to export/publish; avoid users ‘editing the AI’ excessively
    • Volume: more exports/publishes; more short clips per long-form project
    • Activation: first edit completion rate for new users
    • Breadth: ‘unlocking’ unused features (increased AI tool usage)
    • Health signals: retention (D7/D30) and editing more videos (without increasing time per edit)
  9. 19:34 – 22:55

    North Star selection: why ‘exports/publishes in 7–30 days’ wins (with guardrails)

    Aakash evaluates candidate North Stars and chooses a behaviorally grounded metric that captures value across segments. He adds guardrails like support ticket rate to ensure the metric isn’t improved by creating user pain.

    • Rejects time-to-export as sole North Star (may miss quality; could rise with more tool use)
    • Considers ‘# AI tools used’ but notes complexity and risk of superficial inflation
    • Chooses ‘# exports/publishes in 7–30 days’ as the all-encompassing success signal
    • Adds guardrail: support requests should not increase
    • Deprioritizes ‘write-up info/copy-paste’ as supportive but not North Star material
  10. 22:55 – 24:53

    Breaking down the North Star: three vectors to diagnose performance

    Aakash decomposes the North Star to ensure it’s actionable and interpretable. He proposes segmenting by user type and export type, and ties the metric to the actual user journey of exporting/publishing.

    • Vector 1: user type (new users vs power editors)
    • Vector 2: export type (short-form vs long-form)
    • Vector 3: equation/journey clarity (explicitly counts publish/export actions)
    • Highlights need to monitor ‘accept with minimal edits’ as a key companion signal
  11. 24:53 – 28:41

    Trade-offs & guardrails: evals for hallucinations, time regressions, support load, and ‘rage’ signals

    The discussion turns to what could go wrong and how to operationalize protections. Aakash proposes AI-evals for hallucinated edits and careful time-to-edit comparisons normalized by tool usage, plus behavioral signals of user frustration.

    • Guardrail: hallucination rate (agent claims it edited but didn’t) with strict thresholds
    • Guardrail: time-to-edit should not increase when normalized by number of tools used (A/B table)
    • Guardrail: support requests per user should not rise
    • Guardrail proxy: ‘rage interactions’ in chat as a frustration signal (acknowledged as needing calibration)
    • Notes that detailed thresholds are illustrative and should be data-informed
  12. 28:41 – 29:31

    Curveball: privacy concerns about monitoring chats—and a product/measurement mitigation

    Bart challenges the ethics/privacy implications of analyzing conversation content. Aakash responds with an opt-in/opt-out setting and a user-driven reporting mechanism to balance safety, quality control, and privacy.

    • Risk: users may feel ‘listened in on’ if chat is analyzed for rage/quality signals
    • Mitigation: explicit setting to control whether conversations are used for training/analysis
    • Alternative: allow users to report specific conversations even if opted out
    • Reframes monitoring as necessary for policing agent behavior and safety
  13. 29:31 – 34:13

    Wrap-up dashboard + ‘genie metric’ prompt reveals a missing bucket: output/business metrics

    Aakash summarizes the metric dashboard (positive + negative) and ties guardrails to eval-driven monitoring. Bart then asks for a dream ‘genie metric,’ prompting Aakash to add business outcome metrics he initially omitted.

    • Dashboard summary: North Star exports + supporting engagement/retention/activation metrics
    • Negative set: time regressions, hallucinations, support tickets, minimal-edits acceptance proxies
    • Mentions eval-driven approach combining production + synthetic traces for failure modes
    • Genie metric concept: users saying ‘that was awesome’ + strong social chatter + business lift
    • Adds missing output metrics: upgrades, renewals, referrals (revenue/K-factor)
  14. 34:13 – 40:44

    Post-interview debrief: score, how to recover from hints, and the follow-up ‘power move’

    Aakash self-critiques (8.5/10) for missing output metrics and recommends building input/output and leading/lagging into the default framework. Bart emphasizes visual structure and ownership of the answer, then Aakash outlines a standout post-interview follow-up strategy.

    • Lesson: always include input vs output metrics; leading vs lagging indicators
    • When interviewer hints, revisit and amend earlier parts—don’t defend the miss
    • Visual anchors (Miro/board) help interviewer track a long narrative and enable backtracking
    • Follow-up strategy: send a refined metric/dashboard note 30 minutes after the interview
    • Context: this follow-up helps most at smaller companies; less impact at Meta/Google due to process rules

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.