Skip to content
Lenny's PodcastLenny's Podcast

Why the people building AI can’t tell you what’s next | Dianne Penn (Anthropic)

Dianne Penn is Head of Product for Anthropic’s AI Research and Labs teams. She joined in 2023 as Anthropic’s first technical product manager, when the entire product team was five engineers, and has since helped ship every model from Claude 2 through Fable, and helped incubate Claude Code, MCP, Skills, computer use, tool use, and reasoning. Before Anthropic, she helped build Alexa’s AI at Amazon and, before that, traded high-yield bonds at JP Morgan Chase. *In our in-depth conversation, we discuss:* 1. What Anthropic’s early days were like 2. The inflection points that turned Anthropic from an underdog into the fastest-growing company in history 3. How exactly Claude got so good at coding 4. The eval-driven development loop her team is pioneering 5. How to find joy in AI when everything is moving this fast 6. Why Claude’s willingness to push back is key to its success 7. Where human judgment remains irreplaceable *Brought to you by:* WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more: https://workos.com/lenny Mercury—Radically different banking, now with Command: https://mercury.com/ *Episode transcript:* https://www.lennysnewsletter.com/p/anthropics-first-technical-pm-on *Archive of all Lenny's Podcast transcripts:* https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&st=ahz0fj11&dl=0 *Where to find Dianne Penn:* • LinkedIn: https://www.linkedin.com/in/dianne-na-penn Where to find Lenny:* • Newsletter: https://www.lennysnewsletter.com • X: https://twitter.com/lennysan • LinkedIn: https://www.linkedin.com/in/lennyrachitsky/ *In this episode, we cover:* (00:00) Introduction (02:31) Early Anthropic days (08:55) Big milestones (13:50) Inside the exponential (20:02) Token maxing (23:30) Anthropic Labs and the incubation model (27:30) How the research role works (31:35) How to become a top researcher (35:18) Frontier model safeguards (39:38) Hiring in the AI era (44:16) Building an eval set (47:48) Evals vs PRDs (49:55) The importance of hands-on leadership (52:46) Finding joy in AI (58:10) How Dianne uses Claude (01:01:05) Avoiding overreliance on AI (01:03:50) The constitution that makes Claude better (01:07:11) AI writing and verification (01:11:40) Where human brains will continue to be valuable (01:14:10) Navigating AI with kids (01:16:26) Alignment, the future of the PM role, and burnout (01:21:54) Lightning round and final thoughts *Referenced:* • Anthropic: https://www.anthropic.com • Golden Gate Claude: https://www.anthropic.com/news/golden-gate-claude • Dario Amodei’s website: https://darioamodei.com • Scaling Laws and Interpretability of Learning from Repeated Data: https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data • Tokenmaxxing: How Top Builders Use AI To Do The Work Of 400 Engineers: https://www.ycombinator.com/library/Pa-tokenmaxxing-how-top-builders-use-ai-to-do-the-work-of-400-engineers • Garry Tan on X: https://x.com/garrytan • Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night | Ben Mann: https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann • Anthropic’s CPO on what comes next | Mike Krieger (co-founder of Instagram): https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next • Introducing Labs: https://www.anthropic.com/news/introducing-anthropic-labs • Louis CK | about airplane Wi Fi: https://www.youtube.com/watch?v=me4BZBsHwZs • What happens after coding is solved? | Fiona Fung (Manager of the Claude Code and Cowork Teams): https://www.lennysnewsletter.com/p/building-the-most-ai-pilled-engineering • The Anthropic Hive Mind: https://steve-yegge.medium.com/the-anthropic-hive-mind-d01f768f3d7b • How to build a company that withstands any era | Eric Ries, Lean Startup author: https://www.lennysnewsletter.com/p/how-to-build-a-company-that-withstands • Fallout on Prime Video: https://www.amazon.com/dp/B0CN4GGGQ2 • Fallout (video game): https://fallout.bethesda.net • Claude Tag: https://www.anthropic.com/news/introducing-claude-tag *Recommended books:* • Crucial Conversations: Tools for Talking When Stakes Are High: https://www.amazon.com/dp/0071771328 • How to Raise an Adult: Break Free of the Overparenting Trap and Prepare Your Kid for Success: https://www.amazon.com/How-Raise-Adult-Overparenting-Prepare/dp/1627791779 • Incorruptible: Why Good Companies Go Bad... and How Great Companies Stay Great: https://www.amazon.com/dp/B0FWZZBPZB _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com._ Lenny may be an investor in the companies discussed.

Dianne PennguestLenny Rachitskyhost
Jul 26, 20261h 33mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 8:47

    Anthropic’s early days: startup pace, strong culture, and “Golden Gate Claude”

    Dianne describes joining Anthropic in 2023 when the product org was tiny and the company was still searching for its identity. She shares the “Golden Gate Claude” demo as a revealing example of cross-functional speed, bottoms-up initiative, and translating research into a playful public experience.

    • Joined in 2023: ~5 product engineers, 1 engineer covering the API business
    • Early focus: not just building models, but figuring out how they create real user/societal value
    • “Golden Gate Claude” interpretability demo: dialing up a feature made Claude obsess over the Golden Gate Bridge
    • Built and shipped the experience in ~24 hours across product/design/engineering/research
    • Culture/values and bottoms-up experimentation were core from the beginning
  2. 8:47 – 12:10

    Major milestones: Opus 3, building frontier confidence, and betting on coding

    The conversation shifts to the inflection points that helped Anthropic establish itself against stronger incumbents. Dianne highlights Opus 3 as a trust-building, company-wide rally and explains how a relatively small training emphasis on long-form coding became a key differentiator.

    • Opus 3 training/testing as a company-wide milestone and “frontier model” moment
    • Intense cross-team collaboration created long-term trust between product and research
    • Early core question: “Why choose Claude?” drove focus and prioritization
    • Coding wasn’t initially associated with Anthropic; shifted after observing long-form coding usage
    • Small training changes toward coding produced outsized competitive differentiation
  3. 12:10 – 13:50

    Opus 4.5 + Claude Code: models need frontier products (and vice versa)

    Dianne explains why Opus 4.5 felt different: it paired a major capability jump with a product vehicle that let users experience it immediately. She argues that frontier models and frontier products reinforce each other, accelerating adoption and unlocking agentic workflows.

    • Opus 4.5 as a major moment because it arrived with a strong product experience (Claude Code)
    • Thesis: you need frontier products for people to feel frontier model “magic”
    • Agentic, end-to-end workflows became broadly usable at a new intelligence level
    • Claude Code adoption accelerated because the model improved; the model’s impact grew because the product existed
    • Product-model co-evolution as a repeatable growth pattern
  4. 13:50 – 20:03

    Living “inside the exponential”: emerging capabilities, adaptability, and eval-driven discovery

    Lenny and Dianne discuss the sensation of rapid capability jumps and why prediction is hard even for insiders. Dianne emphasizes adaptability, first-principles decision-making, and the role of evals in detecting discontinuous “emergent” capabilities that appear suddenly as scaling progresses.

    • Analogy to the internet becoming mainstream: adaptability becomes a core organizational skill
    • Hard to predict which model release unlocks which capability; teams must respond to new information fast
    • Scaling law papers: smooth loss curves vs discontinuous emergent capability jumps
    • Evals are essential to detect capability jumps and to manage safety challenges
    • “Product overhang” and “user overhang”: current models can do more than products/users realize
  5. 20:03 – 23:33

    Token maxing vs experimentation: why discovery is communal, not solo

    They unpack the idea that heavy token spend can let you ‘live in 2028’ today, but Dianne reframes it as an experimentation outcome rather than a spend target. She describes Anthropic’s internal “working in public” behaviors where shared discovery in Slack turned individual prompts into repeatable use cases.

    • Token spend is an input; the real goal is experimentation and learning velocity
    • Best prototypers spend lots of time with new model versions—no substitute for hands-on use
    • Early Anthropic: broad internal testing in public channels accelerated discovery
    • Communal iteration: one person’s use case becomes many people’s variations, quickly revealing patterns
    • Broad capabilities are known (e.g., writing), but user-level pain points require exploration
  6. 23:33 – 27:30

    Inside Anthropic Labs: discontinuous bets, incubation pods, and revisiting ideas by model generation

    Dianne outlines Labs’ charter: pursue large, discontinuous bets outside the core roadmap and pressure-test whether there’s a 10x–1000x opportunity. She explains why small pods, strong opinions about themes, and a fast experimentation culture let Labs repeatedly produce breakout products.

    • Labs thesis: pull on discontinuous threads and validate 10x/100x/1000x potential
    • Examples: Claude Code, Skills, Claude Design, MCP
    • Strongly held opinion on theme; loosely held about prototype details
    • Small pods (sometimes 1 engineer) avoid coordination drag on ambiguous ideas
    • “Not shipping” can still be valuable learning; revisit ideas in 1–2 model generations
  7. 27:30 – 31:36

    How AI research teams work—and how PMs translate user pain into actionable model work

    Dianne demystifies researcher workflows: a mix of bold, long-horizon vision and short-horizon iteration to improve real-world performance. She then details the PM’s critical bridge function—turning vague user feedback (e.g., ‘hallucinated’) into specific failure modes and measurable, researcher-friendly tasks.

    • Researchers operate with founder-like ambition (e.g., early visions like “AI using a computer”)
    • Dual focus: medium/long-term visions plus immediate product-impact improvements
    • PM role: diagnose vague feedback into concrete categories (tool use vs search vs alignment, etc.)
    • Actionability requires deep trace-level analysis of what the user asked and what the model did
    • Evals and structured feedback loops convert user experience into train/test objectives
  8. 31:36 – 35:19

    Becoming a top researcher: first-principles thinking, ambition, and staying close to the details

    Lenny asks what makes researchers (and research-adjacent PMs) exceptional in today’s talent market. Dianne highlights reasoning from first principles, bold visions, and a strong “taste” developed by being deeply engaged with training runs, data, and evals.

    • Top researchers: strong first-principles reasoning and problem decomposition
    • Bold ambition (shoot for transformative outcomes) is a recurring differentiator
    • Staying close to details: leaders still inspect training runs, evals, and data
    • Developing “taste” through hands-on iteration
    • Being stubborn about the problem area but flexible about approach as tech changes
  9. 35:19 – 39:23

    Frontier model safeguards: rising scrutiny, red teaming, and “fallback” user experiences

    They discuss how increasingly capable frontier models trigger more scrutiny and access constraints. Dianne focuses on product implications: evolving pre-release safety processes and building fallback UX so users still receive high-quality outcomes even when advanced models are gated or restricted.

    • Frontier capability increases require stronger safeguards and pre-release testing/red teaming
    • Product challenge: maintain great UX while expanding safety systems
    • Example: building fallback systems so users still get strong answers from a safer/available model tier
    • Anthropic goal: make general-purpose AI inclusive and accessible while managing risk
    • Expectation: ongoing innovation in “model safeguards package”
  10. 39:23 – 44:16

    What’s changing for PMs: first principles, “Evals are the new PRDs,” and sweating tokens like pixels

    Dianne explains why Anthropic hasn’t changed its PM hiring loop: the core traits still matter, but execution artifacts shift. In AI product work, the PM’s job increasingly starts with defining and measuring behavior through evals—often faster and more actionable than traditional documents.

    • Hiring loop unchanged; core trait emphasized: first-principles thinking over pattern-matching past playbooks
    • “Evals are the new PRDs” as a redefinition of how user value is specified and measured
    • New PM craft: reading transcripts, diagnosing failure trajectories, and turning them into evals
    • “Sweat the tokens as much as you sweat the pixels” as a new product rigor
    • Goal: shorten distance from user pain to researcher actionability
  11. 44:16 – 50:19

    What evals look like in practice—and when PRDs still matter

    Dianne gives a concrete eval example: early Claude struggled with structured outputs like JSON, so the team built a targeted eval set from real failures and tracked improvements over time. She also clarifies that PRDs remain valuable for alignment across large stakeholder groups and for ambiguous, visionary work (e.g., computer use).

    • Example eval: instruction-following failures often meant “wrong JSON output”
    • Create 30–40 representative examples to form an eval set; run it every model iteration
    • Evals enable measurable progress and help handle non-determinism and jagged performance
    • PRDs still used for launches and cross-functional alignment (engineering, legal, safety, etc.)
    • PRDs especially useful for ambiguous opportunities where user pain points aren’t yet well-defined
  12. 50:19 – 58:11

    Hands-on leadership and finding joy: why managers must ship, and why experimentation is social

    Dianne argues that managers can’t lead AI product work without being hands-on and shipping with the tech themselves. To sustain motivation, she recommends pairing with excited peers, going deep on a few tools, and building virtuous cycles of shared discovery rather than treating AI as a checkbox.

    • Leaders must be hands-on: onboarding is the same regardless of tenure; everyone must learn via real usage
    • Carve time to own workstreams to maintain “theory of mind” for model/product pace
    • Joy comes from shared discovery—pair with someone excited and explore together
    • Go deep on 1–2 tools/use cases instead of shallowly trying everything
    • Internal “experimenting in public” as a culture mechanism for sustained learning
  13. 58:11 – 1:03:53

    How Dianne uses Claude: agentic workflows, management coaching, and avoiding overreliance

    Dianne shares practical personal use: using Claude/Skills to prepare for difficult conversations and become a better coach, not just to boost “IQ” tasks. She then addresses overreliance by describing when she forms her own POV first versus delegating standardized writing, emphasizing that verification and accountability matter as much as authorship.

    • Using Claude for coaching/management prep (e.g., “Crucial Conversations” guidance)
    • AI as EQ augmentation: brainstorming phrasing, trust-building, and directness
    • Avoiding overreliance: form your POV first, then use Claude as sparring partner
    • Delegate low-leverage writing (e.g., business reviews) while focusing human effort on thinking/judgment
    • Key distinction: who verifies/signs off is often more important than who drafts
  14. 1:03:53 – 1:14:11

    Claude’s constitution and better thinking partners: why pushback beats compliance, plus what humans retain

    They explore why alignment and Claude’s “constitution” can make the model more useful—because it knows when to push back and help users reach better outcomes. The discussion broadens to what remains distinctly human for now: judgment, persistence, proactivity, and domain expertise in fields not yet on the steep part of the curve.

    • A good thinking partner shouldn’t just agree; pushback improves outcomes
    • Alignment traits support proactivity and principled disagreement, not just refusal behavior
    • Using Claude even for meta decisions (e.g., pricing strategy brainstorming) to refine reasoning
    • Human value: judgment built from nuanced lived experience; choosing what to build among infinite options
    • Emerging domains beyond software: biology/life sciences; Anthropic investing in specialized efforts (e.g., Claude Science)
  15. 1:14:11 – 1:33:50

    Kids, burnout, culture—and closing: sustainability via team trust, low ego, and feedback loops

    Dianne emphasizes teaching kids curiosity, persistence, and an inner voice in an AI-saturated world. She explains burnout avoidance through high-trust teamwork (“not an individual sport”), then closes with a defense of PM value: user-centric detail work and actionable translation are more needed, not less, as building becomes easier.

    • Raising kids for an AI future: curiosity, persistence, and developing an independent inner voice
    • Burnout prevention: high-performance pace requires team collaboration, not heroics
    • “Hive mind” and low-ego culture let people take real PTO without the system breaking
    • PMs still matter: deciding what to build and grounding tech in user needs is the bottleneck
    • Call to action: give product feedback (thumbs up/down, customer channels) and hiring for curious, first-principles tinkerers

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.