Skip to content
Aakash GuptaAakash Gupta

Stop Applying to AI PM Jobs Until You Watch This Safety & Ethics Mock

Apply to Land a PM Job Cohort 3 (starts May 4): https://www.landpmjob.com/ Most AI PM candidates underrate the safety and ethics round. In this episode, Ankit Virmani (AI PM at Uber, formerly GPM at Meta), Prasad Reddy (former CPO at L-Nutra, ex-VP at Danaher), and Dr. Bart Jaworski (coach to 12,000+ at Amazon, Microsoft, Zalando) join Aakash for four live mock rounds with real-time scoring, plus a framework you can use the next time medical chatbots, hiring bias, or autonomous agents come up in your loop. Full Writeup: https://www.news.aakashg.com/p/safety-ethics-interview --- Timestamps: 00:00 The safety round most AI PM candidates underrate 01:36 Why senior candidates freeze on safety questions 03:39 The SHIR framework: severity, harm scope, immediacy, reversibility 06:34 Mock 1: Medical chatbot contradicting clinical guidelines 11:06 Mock 2: Hiring tool with a 15% demographic gap 16:50 Mock 3: AI agent booking flights and sending emails 21:17 Mock 4: Right for users, wrong for short-term metrics 27:09 Bart's full scoring reveal 32:50 The 40 minute rule for proactive safety mentions 33:38 Anthropic vs OpenAI vs Google: hardest safety round 34:48 The one question every AI PM candidate should be ready for --- 🏆 Two things to consider: 1. AI Tools Bundle: A full year of Mobbin, Arize, Relay, Dovetail, Linear, Magic Patterns, Deepseek, Reforge, Build, Descript, and Speechify with an annual paid newsletter sub - https://bundle.aakashg.com 2. Land a PM Job: 12-week cohort with Aakash, Ankit, Prasad, and Bart - https://www.landpmjob.com/ --- Key Takeaways: 1. SHIR is the framework that buys you 30 seconds of structured thinking - Severity, Harm scope, Immediacy, Reversibility. Run any safety question through these four words before you say a word about your solution. Most candidates jump to "pull the feature" or "ship it anyway." SHIR gets you to a guardrail-plus-audit answer that actually matches how senior PMs think. 2. At the CPO and VP level, sizing business impact is table stakes - The pull costs $50M. The guardrails cost $200K and two weeks. The full retrain costs $2M and three months. If you cannot put numbers next to each path, you are not interviewing at the right altitude. 3. Safety is evaluated across the entire loop, not in one round - Meta embedded safety thinking inside the product sense rubric itself. If you make it 40 minutes into a 60-minute interview without mentioning safety, you have probably already lost points you cannot recover. 4. Reframe revenue arguments as headline arguments - When the VP says "we cannot pull this before earnings," your move is to ask whether the company can afford the headline that you knew the AI was giving dangerous medical advice and let it ship anyway. That converts a $50M quarter risk into a $5B brand risk in one sentence. 5. Agent safety has three pillars - Scope (spending caps and category limits), confirmation (forked by stakes, with push notifications and undo windows for medium actions), and reversibility (pending states, send delays, anomaly detection on top). Memorize this stack for any agent question. 6. Liability for AI agents almost always lands on the platform - Because you designed the guardrails. Frame your answer around how you reduce risk through scope limits and confirmation flows, then acknowledge the legal gray area and the jurisdiction-by-jurisdiction nuance. 7. No questions asked refunds create moral hazard - Prasad's pushback on Aakash here is the lesson. Refunds are the safety net. Scope limits are the railing. Build both. If you only build the refund, users will test the limit. 8. Anthropic has the hardest safety round in the industry - Expect 45 to 60 minutes on safety alone. Read up on constitutional AI and the founding story before you walk in. Practice both situational and historical behavioral answers out loud, and watch the recording back. --- 👨‍💻 Where to find Ankit Virmani: LinkedIn: https://www.linkedin.com/in/ankitvirmani/ 👨‍💻 Where to find Prasad Reddy: LinkedIn: https://www.linkedin.com/in/prasad-09/ 👨‍💻 Where to find Dr. Bart Jaworski: LinkedIn: https://www.linkedin.com/in/bart-jaworski/ 👨‍💻 Where to find Aakash: Twitter: https://www.x.com/aakashg0 LinkedIn: https://www.linkedin.com/in/aakashgupta/ Newsletter: https://www.news.aakashg.com #aipm #pminterview --- 🧠 About Product Growth: The world's largest podcast focused solely on product + growth, with over 200K+ listeners. 🔔 Subscribe and turn on notifications to get more videos like this.

Aakash GuptahostAnkit VirmaniguestPrasad ReddyguestDr. Bart Jaworskiguest
May 3, 202638mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 2:40

    Why AI PM safety & ethics interviews are a hidden filter across the whole loop

    Aakash, Ankit, and Prasad argue that candidates consistently underrate safety/ethics—and that companies often evaluate safety thinking throughout product sense and execution, not only in a dedicated round. They set the stakes: in many AI products, harm is real, physical, and legally consequential.

    • Safety is embedded into product thinking at companies like Meta, not a separate checkbox
    • Candidates often fail by not proactively identifying harms and mitigations
    • Higher-stakes domains (rides, healthcare) raise the bar significantly
    • Senior candidates can fail fast if they can’t discuss liability and board-level implications
  2. 2:40 – 3:47

    Why even senior PMs freeze: translating intuition into formal safety reasoning

    Prasad explains that experienced candidates often have instincts about risk but haven’t practiced articulating a structured safety rationale. The group frames safety as table stakes at senior levels, where the ability to reason under pressure matters as much as the decision itself.

    • Freezing happens when safety reasoning hasn’t been formalized into a repeatable approach
    • CPO/VP interviews expect board-level judgment and liability awareness
    • In regulated areas, inability to handle safety scenarios can end the interview
    • Safety evaluation is often implicit across multiple rounds
  3. 3:47 – 5:11

    The SHIR framework (Severity, Harm scope, Immediacy, Reversibility) + asking for thinking time

    Aakash introduces SHIR as a compact mental model to triage risk and structure answers quickly. He also recommends explicitly taking 30 seconds to gather thoughts, then using SHIR to guide what data to request and what actions to prioritize.

    • Severity: worst-case impact (e.g., medical misinformation vs rudeness)
    • Harm scope: number of users affected and exposure rate
    • Immediacy: active harm now vs latent risk later
    • Reversibility: whether damage can be undone (leaks vs bad recs)
    • Use SHIR as a quick triage or a deep case structure depending on interview format
  4. 5:11 – 6:34

    Executive-level add-on: quantify business impact and compare options side by side

    Prasad and Ankit stress that strong candidates don’t just assess risk—they size business impact of mitigations and tradeoffs. Putting costs, timelines, and risk reduction next to each other makes decisions clearer and demonstrates executive-level judgment.

    • Top candidates quantify cost/time of each mitigation path
    • Compare ‘pull’, ‘guardrails’, and ‘retrain’ against SHIR risk assessment
    • Sizing before solving is a repeated signal of seniority
    • Business framing strengthens safety recommendations
  5. 6:34 – 8:57

    Mock 1: Medical chatbot contradicts clinical guidelines—contain risk without nuking revenue

    Ankit runs a scenario where a consumer chatbot occasionally gives medical advice contradicting clinical guidelines, with major revenue/user impact if pulled. Aakash proposes a phased approach: immediate guardrails and disclaimers, fast auditing to quantify harm, escalation thresholds, and same-day legal involvement.

    • Start by sizing exposure: % of medical queries and failure rate
    • Immediate mitigation: classify medical content and add disclaimers + verified links
    • Audit recent queries to quantify harm and set escalation thresholds
    • Temporarily filter medical topics if risk exceeds thresholds; retrain in parallel
    • Involve legal quickly due to liability and physical harm risk
  6. 8:57 – 10:54

    Mock 1 follow-up: VP blocks changes pre-earnings—reframe to headline/brand catastrophe risk

    When the VP refuses to pull the feature due to earnings pressure, Aakash reframes the decision from quarterly revenue to reputational/brand risk and potential massive downside. He also discusses documenting risks and escalating to safety channels while staying collaborative.

    • Shift frame from ‘can we afford to act’ to ‘can we afford the headline’
    • Re-size the true business impact (guardrail may reduce impact vs full pull)
    • Escalate appropriately and document risk/decision rationale
    • Be a team player while ensuring risk is clearly communicated and recorded
  7. 10:54 – 13:24

    Mock 2: Hiring tool shows a 15% demographic gap—pause auto-reject, audit, and brief the board

    Aakash presents Prasad with a biased hiring recommendation gap. Prasad treats the root cause debate (data vs model) as secondary to immediate risk mitigation, halting auto-reject for affected groups, initiating an audit, and preparing transparent board communication.

    • Outcome-focused framing: regardless of source, discriminatory impact is the issue
    • Immediate mitigation: pause automated rejection; keep humans in the loop
    • Legal/regulatory context: EEOC risk and class action exposure
    • Board strategy: disclose early with a clear plan and timeline
    • Uses prior real-world precedent to justify decisive action
  8. 13:24 – 16:51

    Mock 2 follow-up: ‘Competitors don’t test this much’—position safety as long-term advantage

    Pressed by a CEO worried about speed, Prasad argues that safety diligence is a strategic advantage, not a drag. He cites rising enforcement and enterprise procurement requirements, contrasting short audit delays with the massive cost and time of litigation and reputational damage.

    • Short-term speed vs long-term survivability and trust
    • Enforcement trends and market expectations are increasing
    • Enterprise buyers increasingly demand audits and governance
    • Cost framing: days of audit vs years of legal exposure
    • Safety can be a differentiator (‘moral high ground’)
  9. 16:51 – 19:40

    Mock 3: Agent safety for bookings/emails—design guardrails for autonomous real-world actions

    Prasad asks Aakash how to keep an AI agent safe when it can spend money and act on a user’s behalf. Aakash proposes a product safety framework based on scope limits, graduated confirmations, reversibility buffers (undo/pending), plus anomaly detection.

    • Set spending caps during onboarding (per transaction and per trip)
    • Use tiered confirmation based on stakes (none / soft / hard)
    • Build reversibility: pending windows and send delays to enable undo
    • Add anomaly detection for unusual behavior compared to user baseline
    • Design assumption: bots lack legal agency, so guardrails are essential
  10. 19:40 – 21:25

    Agent liability debate: refunds, moral hazard, and designing to prevent failures

    Aakash answers who is liable when the agent books expensive flights by separating legal uncertainty from product/business strategy and favoring user-friendly remediation. Prasad pushes back: always-refund policies can create moral hazard, so prevention guardrails must be primary, with refunds as a backstop.

    • Distinguish legal answer vs product strategy answer
    • Consider guardrail compliance when evaluating responsibility
    • User trust may require generous remediation, but it has tradeoffs
    • Moral hazard risk: users may exploit ‘no questions asked’ refunds
    • Prioritize prevention (rails) with remediation (safety net)
  11. 21:25 – 24:52

    Mock 4: ‘Right for users, wrong for short-term metrics’—rebuilding a value model for quality

    Aakash interviews Ankit on a behavioral tradeoff scenario. Ankit describes changing Reels ranking from click-optimized incentives to engagement-quality and retention, sequencing evidence to overcome revenue fears and demonstrating measurable gains in users while revenue stabilized over time.

    • Problem: click optimization created clickbait and reduced diversity/quality
    • Reframe: link quality engagement to long-term retention vs raw clicks
    • Model change: additive to multiplicative value model across funnel stages
    • Execution: sequence low-risk signal changes to build credibility
    • Outcome: DAU and marginal user lift; revenue stabilized via better sessions/inventory
  12. 24:52 – 27:09

    The ‘one question’ scenario: discovering a known safety issue leadership ignored—how to escalate

    Ankit outlines how to respond when leadership knowingly ships with a safety issue: gather context, validate the risk, document concerns formally, escalate through official channels, and consider leaving if active harm persists and internal pathways fail.

    • Start with context: understand why leadership decided as they did
    • Validate whether harm is real/active and assess severity
    • Document in durable formats (memos) and include specific risks + fixes
    • Escalate via ethics channels if normal routes fail
    • Personal boundary: if active harm continues, reassess whether to stay
  13. 27:09 – 32:50

    Scoring reveal + coaching: what distinguished answers (and common pitfalls like sounding too polished)

    Bart scores the mock answers and explains why Prasad edged out by a small margin, while noting all would likely pass. The group reflects on what makes safety answers strong—clear structure, real examples, and natural delivery—while warning that overly polished responses can backfire.

    • Bart’s rubric focus: framework application, tradeoff reasoning, stakeholders, clarity
    • Prasad wins narrowly; examples and executive framing stand out
    • Aakash self-critiques: note-checking and over-structured delivery risks
    • Over-polished answers can trigger suspicion (especially with AI tools)
    • Strong narratives feel linear and human, even when well-structured
  14. 32:50 – 38:16

    Rapid-fire: the 40-minute proactive safety rule, hardest safety rounds, and must-prepare questions

    In a Q&A sprint, Aakash shares tactical prep advice and what interviewers look for most. He emphasizes proactively bringing up safety, notes Anthropic’s especially intense safety focus, recommends practice on video, and highlights key prompts like unintended harm and agent liability.

    • Biggest red flag: failing to mention safety proactively across interviews
    • 40-minute rule: if you haven’t mentioned safety by minute 40, bring it in
    • Hardest safety round: Anthropic (often 45–60 minutes; constitutional AI context matters)
    • Two-hour prep plan: SHIR + practice out loud on video (not just AI dictation)
    • Must-prepare question: ‘Tell me about a time your product caused unintended harm’
    • Agent liability: usually platform responsibility, with legal nuance by jurisdiction

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.