Skip to content
Lenny's PodcastLenny's Podcast

Why the people building AI can’t tell you what’s next | Dianne Penn (Anthropic)

Dianne Penn is Head of Product for Anthropic’s AI Research and Labs teams. She joined in 2023 as Anthropic’s first technical product manager, when the entire product team was five engineers, and has since helped ship every model from Claude 2 through Fable, and helped incubate Claude Code, MCP, Skills, computer use, tool use, and reasoning. Before Anthropic, she helped build Alexa’s AI at Amazon and, before that, traded high-yield bonds at JP Morgan Chase. *In our in-depth conversation, we discuss:* 1. What Anthropic’s early days were like 2. The inflection points that turned Anthropic from an underdog into the fastest-growing company in history 3. How exactly Claude got so good at coding 4. The eval-driven development loop her team is pioneering 5. How to find joy in AI when everything is moving this fast 6. Why Claude’s willingness to push back is key to its success 7. Where human judgment remains irreplaceable *Brought to you by:* WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more: https://workos.com/lenny Mercury—Radically different banking, now with Command: https://mercury.com/ *Episode transcript:* https://www.lennysnewsletter.com/p/anthropics-first-technical-pm-on *Archive of all Lenny's Podcast transcripts:* https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&st=ahz0fj11&dl=0 *Where to find Dianne Penn:* • LinkedIn: https://www.linkedin.com/in/dianne-na-penn Where to find Lenny:* • Newsletter: https://www.lennysnewsletter.com • X: https://twitter.com/lennysan • LinkedIn: https://www.linkedin.com/in/lennyrachitsky/ *In this episode, we cover:* (00:00) Introduction (02:31) Early Anthropic days (08:55) Big milestones (13:50) Inside the exponential (20:02) Token maxing (23:30) Anthropic Labs and the incubation model (27:30) How the research role works (31:35) How to become a top researcher (35:18) Frontier model safeguards (39:38) Hiring in the AI era (44:16) Building an eval set (47:48) Evals vs PRDs (49:55) The importance of hands-on leadership (52:46) Finding joy in AI (58:10) How Dianne uses Claude (01:01:05) Avoiding overreliance on AI (01:03:50) The constitution that makes Claude better (01:07:11) AI writing and verification (01:11:40) Where human brains will continue to be valuable (01:14:10) Navigating AI with kids (01:16:26) Alignment, the future of the PM role, and burnout (01:21:54) Lightning round and final thoughts *Referenced:* • Anthropic: https://www.anthropic.com • Golden Gate Claude: https://www.anthropic.com/news/golden-gate-claude • Dario Amodei’s website: https://darioamodei.com • Scaling Laws and Interpretability of Learning from Repeated Data: https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data • Tokenmaxxing: How Top Builders Use AI To Do The Work Of 400 Engineers: https://www.ycombinator.com/library/Pa-tokenmaxxing-how-top-builders-use-ai-to-do-the-work-of-400-engineers • Garry Tan on X: https://x.com/garrytan • Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night | Ben Mann: https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann • Anthropic’s CPO on what comes next | Mike Krieger (co-founder of Instagram): https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next • Introducing Labs: https://www.anthropic.com/news/introducing-anthropic-labs • Louis CK | about airplane Wi Fi: https://www.youtube.com/watch?v=me4BZBsHwZs • What happens after coding is solved? | Fiona Fung (Manager of the Claude Code and Cowork Teams): https://www.lennysnewsletter.com/p/building-the-most-ai-pilled-engineering • The Anthropic Hive Mind: https://steve-yegge.medium.com/the-anthropic-hive-mind-d01f768f3d7b • How to build a company that withstands any era | Eric Ries, Lean Startup author: https://www.lennysnewsletter.com/p/how-to-build-a-company-that-withstands • Fallout on Prime Video: https://www.amazon.com/dp/B0CN4GGGQ2 • Fallout (video game): https://fallout.bethesda.net • Claude Tag: https://www.anthropic.com/news/introducing-claude-tag *Recommended books:* • Crucial Conversations: Tools for Talking When Stakes Are High: https://www.amazon.com/dp/0071771328 • How to Raise an Adult: Break Free of the Overparenting Trap and Prepare Your Kid for Success: https://www.amazon.com/How-Raise-Adult-Overparenting-Prepare/dp/1627791779 • Incorruptible: Why Good Companies Go Bad... and How Great Companies Stay Great: https://www.amazon.com/dp/B0FWZZBPZB _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com._ Lenny may be an investor in the companies discussed.

Dianne PennguestLenny Rachitskyhost
Jul 26, 20261h 33mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Anthropic’s Dianne Penn on AI product, evals, and pace ahead

  1. Anthropic’s early momentum came from a strong mission-led culture and fast, cross-functional execution that turned research moments (e.g., “Golden Gate Claude”) into public product experiences quickly.
  2. Key inflection points included successfully shipping a frontier-class model with Claude/Opus 3, then pairing later frontier capability (Opus 4.5) with a “vehicle” product (Claude Code) to unlock adoption.
  3. Penn argues AI progress is characterized by discontinuous “emergent capabilities,” making precise prediction hard and making adaptability, first-principles reasoning, and strong eval infrastructure essential.
  4. In Anthropic’s research product workflow, “evals are the new PRDs”: translating vague user feedback into measurable, actionable eval sets that researchers can directly optimize against.
  5. She outlines how PM and leadership expectations are shifting toward deeply hands-on building (“sweat the tokens as much as pixels”), collaborative experimentation, and safeguards that preserve user experience as safety requirements tighten.

IDEAS WORTH REMEMBERING

5 ideas

Frontier models need frontier products to reveal their value.

Penn describes a product–model flywheel: Opus 4.5’s breakthrough mattered because Claude Code made the capability usable, and Claude Code’s adoption accelerated because the model crossed an intelligence threshold for more end-to-end agentic work.

Expect capability “jumps,” not smooth progress—so build for adaptability.

Scaling may reduce loss smoothly, but user-relevant abilities can appear discontinuously; teams should be ready to pull plans forward or change direction quickly when a new model makes something suddenly feasible.

Treat experimentation as the goal; token spend is just one input.

Rather than “token maxing” for its own sake, Penn frames success as systematic experimentation—often communal—where people share prompts, failures, and discoveries to uncover new use cases faster than solo tinkering.

In AI product, measurable evals increasingly replace PRDs as the core artifact.

“Evals are the new PRDs” because they encode real user pain into tests researchers can optimize; PRDs still matter for cross-functional alignment and ambiguous zero-to-one efforts, but evals shorten the path from feedback to action.

A practical research-PM superpower is translating vague feedback into specific failure modes.

“Claude hallucinated” isn’t actionable; Penn’s team traces the trajectory to determine whether it’s tool-use, retrieval, reasoning, or alignment, then builds an eval set that captures both failing and non-failing cases to prevent regressions.

WORDS WORTH SAVING

5 quotes

We actually have a saying on the team of, "Evals are the new PRDs."

Dianne Penn

You have to sweat the tokens as much as you sweat the pixels.

Dianne Penn

It's very hard to predict the exact moment or the exact model, and so the adaptability of when you're faced with new information, how do you then make better decisions versus keeping the same plan?

Dianne Penn

You should have fun working with this technology.

Dianne Penn

If you have a AI that can be a thinking partner, a thinking partner doesn't just agree with you. It should add to you, and you should come away at the end of the day having better ideas because you worked with Claude.

Dianne Penn

Early Anthropic culture and fast prototypingOpus 3 and coding differentiationOpus 4.5 + Claude Code product-model flywheelEmergent capabilities and unpredictability in scalingToken maxing vs experimentation framingAnthropic Labs incubation modelResearch PM work: user feedback → actionable evalsSafeguards, red-teaming, and fallback UXHands-on leadership and AI-era PM hiring signalsOverreliance, judgment, and human valueClaude’s “constitution,” pushback, and alignment benefitsKids, curiosity, inner voice, and resilience

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.