Lenny's PodcastWhy AI is going vertical (again) | Dianne Penn (Anthropic)
At a glance
WHAT IT’S REALLY ABOUT
Anthropic’s Dianne Penn on AI product, evals, and pace ahead
- Anthropic’s early momentum came from a strong mission-led culture and fast, cross-functional execution that turned research moments (e.g., “Golden Gate Claude”) into public product experiences quickly.
- Key inflection points included successfully shipping a frontier-class model with Claude/Opus 3, then pairing later frontier capability (Opus 4.5) with a “vehicle” product (Claude Code) to unlock adoption.
- Penn argues AI progress is characterized by discontinuous “emergent capabilities,” making precise prediction hard and making adaptability, first-principles reasoning, and strong eval infrastructure essential.
- In Anthropic’s research product workflow, “evals are the new PRDs”: translating vague user feedback into measurable, actionable eval sets that researchers can directly optimize against.
- She outlines how PM and leadership expectations are shifting toward deeply hands-on building (“sweat the tokens as much as pixels”), collaborative experimentation, and safeguards that preserve user experience as safety requirements tighten.
IDEAS WORTH REMEMBERING
5 ideasFrontier models need frontier products to reveal their value.
Penn describes a product–model flywheel: Opus 4.5’s breakthrough mattered because Claude Code made the capability usable, and Claude Code’s adoption accelerated because the model crossed an intelligence threshold for more end-to-end agentic work.
Expect capability “jumps,” not smooth progress—so build for adaptability.
Scaling may reduce loss smoothly, but user-relevant abilities can appear discontinuously; teams should be ready to pull plans forward or change direction quickly when a new model makes something suddenly feasible.
Treat experimentation as the goal; token spend is just one input.
Rather than “token maxing” for its own sake, Penn frames success as systematic experimentation—often communal—where people share prompts, failures, and discoveries to uncover new use cases faster than solo tinkering.
In AI product, measurable evals increasingly replace PRDs as the core artifact.
“Evals are the new PRDs” because they encode real user pain into tests researchers can optimize; PRDs still matter for cross-functional alignment and ambiguous zero-to-one efforts, but evals shorten the path from feedback to action.
A practical research-PM superpower is translating vague feedback into specific failure modes.
“Claude hallucinated” isn’t actionable; Penn’s team traces the trajectory to determine whether it’s tool-use, retrieval, reasoning, or alignment, then builds an eval set that captures both failing and non-failing cases to prevent regressions.
WORDS WORTH SAVING
5 quotesWe actually have a saying on the team of, "Evals are the new PRDs."
— Dianne Penn
You have to sweat the tokens as much as you sweat the pixels.
— Dianne Penn
It's very hard to predict the exact moment or the exact model, and so the adaptability of when you're faced with new information, how do you then make better decisions versus keeping the same plan?
— Dianne Penn
You should have fun working with this technology.
— Dianne Penn
If you have a AI that can be a thinking partner, a thinking partner doesn't just agree with you. It should add to you, and you should come away at the end of the day having better ideas because you worked with Claude.
— Dianne Penn
High quality AI-generated summary created from speaker-labeled transcript.