Lenny's PodcastWhy AI is going vertical (again) | Dianne Penn (Anthropic)
CHAPTERS
- 0:00 – 8:47
Anthropic’s early days: startup pace, strong culture, and “Golden Gate Claude”
Dianne describes joining Anthropic in 2023 when the product org was tiny and the company was still searching for its identity. She shares the “Golden Gate Claude” demo as a revealing example of cross-functional speed, bottoms-up initiative, and translating research into a playful public experience.
- •Joined in 2023: ~5 product engineers, 1 engineer covering the API business
- •Early focus: not just building models, but figuring out how they create real user/societal value
- •“Golden Gate Claude” interpretability demo: dialing up a feature made Claude obsess over the Golden Gate Bridge
- •Built and shipped the experience in ~24 hours across product/design/engineering/research
- •Culture/values and bottoms-up experimentation were core from the beginning
- 8:47 – 12:10
Major milestones: Opus 3, building frontier confidence, and betting on coding
The conversation shifts to the inflection points that helped Anthropic establish itself against stronger incumbents. Dianne highlights Opus 3 as a trust-building, company-wide rally and explains how a relatively small training emphasis on long-form coding became a key differentiator.
- •Opus 3 training/testing as a company-wide milestone and “frontier model” moment
- •Intense cross-team collaboration created long-term trust between product and research
- •Early core question: “Why choose Claude?” drove focus and prioritization
- •Coding wasn’t initially associated with Anthropic; shifted after observing long-form coding usage
- •Small training changes toward coding produced outsized competitive differentiation
- 12:10 – 13:50
Opus 4.5 + Claude Code: models need frontier products (and vice versa)
Dianne explains why Opus 4.5 felt different: it paired a major capability jump with a product vehicle that let users experience it immediately. She argues that frontier models and frontier products reinforce each other, accelerating adoption and unlocking agentic workflows.
- •Opus 4.5 as a major moment because it arrived with a strong product experience (Claude Code)
- •Thesis: you need frontier products for people to feel frontier model “magic”
- •Agentic, end-to-end workflows became broadly usable at a new intelligence level
- •Claude Code adoption accelerated because the model improved; the model’s impact grew because the product existed
- •Product-model co-evolution as a repeatable growth pattern
- 13:50 – 20:03
Living “inside the exponential”: emerging capabilities, adaptability, and eval-driven discovery
Lenny and Dianne discuss the sensation of rapid capability jumps and why prediction is hard even for insiders. Dianne emphasizes adaptability, first-principles decision-making, and the role of evals in detecting discontinuous “emergent” capabilities that appear suddenly as scaling progresses.
- •Analogy to the internet becoming mainstream: adaptability becomes a core organizational skill
- •Hard to predict which model release unlocks which capability; teams must respond to new information fast
- •Scaling law papers: smooth loss curves vs discontinuous emergent capability jumps
- •Evals are essential to detect capability jumps and to manage safety challenges
- •“Product overhang” and “user overhang”: current models can do more than products/users realize
- 20:03 – 23:33
Token maxing vs experimentation: why discovery is communal, not solo
They unpack the idea that heavy token spend can let you ‘live in 2028’ today, but Dianne reframes it as an experimentation outcome rather than a spend target. She describes Anthropic’s internal “working in public” behaviors where shared discovery in Slack turned individual prompts into repeatable use cases.
- •Token spend is an input; the real goal is experimentation and learning velocity
- •Best prototypers spend lots of time with new model versions—no substitute for hands-on use
- •Early Anthropic: broad internal testing in public channels accelerated discovery
- •Communal iteration: one person’s use case becomes many people’s variations, quickly revealing patterns
- •Broad capabilities are known (e.g., writing), but user-level pain points require exploration
- 23:33 – 27:30
Inside Anthropic Labs: discontinuous bets, incubation pods, and revisiting ideas by model generation
Dianne outlines Labs’ charter: pursue large, discontinuous bets outside the core roadmap and pressure-test whether there’s a 10x–1000x opportunity. She explains why small pods, strong opinions about themes, and a fast experimentation culture let Labs repeatedly produce breakout products.
- •Labs thesis: pull on discontinuous threads and validate 10x/100x/1000x potential
- •Examples: Claude Code, Skills, Claude Design, MCP
- •Strongly held opinion on theme; loosely held about prototype details
- •Small pods (sometimes 1 engineer) avoid coordination drag on ambiguous ideas
- •“Not shipping” can still be valuable learning; revisit ideas in 1–2 model generations
- 27:30 – 31:36
How AI research teams work—and how PMs translate user pain into actionable model work
Dianne demystifies researcher workflows: a mix of bold, long-horizon vision and short-horizon iteration to improve real-world performance. She then details the PM’s critical bridge function—turning vague user feedback (e.g., ‘hallucinated’) into specific failure modes and measurable, researcher-friendly tasks.
- •Researchers operate with founder-like ambition (e.g., early visions like “AI using a computer”)
- •Dual focus: medium/long-term visions plus immediate product-impact improvements
- •PM role: diagnose vague feedback into concrete categories (tool use vs search vs alignment, etc.)
- •Actionability requires deep trace-level analysis of what the user asked and what the model did
- •Evals and structured feedback loops convert user experience into train/test objectives
- 31:36 – 35:19
Becoming a top researcher: first-principles thinking, ambition, and staying close to the details
Lenny asks what makes researchers (and research-adjacent PMs) exceptional in today’s talent market. Dianne highlights reasoning from first principles, bold visions, and a strong “taste” developed by being deeply engaged with training runs, data, and evals.
- •Top researchers: strong first-principles reasoning and problem decomposition
- •Bold ambition (shoot for transformative outcomes) is a recurring differentiator
- •Staying close to details: leaders still inspect training runs, evals, and data
- •Developing “taste” through hands-on iteration
- •Being stubborn about the problem area but flexible about approach as tech changes
- 35:19 – 39:23
Frontier model safeguards: rising scrutiny, red teaming, and “fallback” user experiences
They discuss how increasingly capable frontier models trigger more scrutiny and access constraints. Dianne focuses on product implications: evolving pre-release safety processes and building fallback UX so users still receive high-quality outcomes even when advanced models are gated or restricted.
- •Frontier capability increases require stronger safeguards and pre-release testing/red teaming
- •Product challenge: maintain great UX while expanding safety systems
- •Example: building fallback systems so users still get strong answers from a safer/available model tier
- •Anthropic goal: make general-purpose AI inclusive and accessible while managing risk
- •Expectation: ongoing innovation in “model safeguards package”
- 39:23 – 44:16
What’s changing for PMs: first principles, “Evals are the new PRDs,” and sweating tokens like pixels
Dianne explains why Anthropic hasn’t changed its PM hiring loop: the core traits still matter, but execution artifacts shift. In AI product work, the PM’s job increasingly starts with defining and measuring behavior through evals—often faster and more actionable than traditional documents.
- •Hiring loop unchanged; core trait emphasized: first-principles thinking over pattern-matching past playbooks
- •“Evals are the new PRDs” as a redefinition of how user value is specified and measured
- •New PM craft: reading transcripts, diagnosing failure trajectories, and turning them into evals
- •“Sweat the tokens as much as you sweat the pixels” as a new product rigor
- •Goal: shorten distance from user pain to researcher actionability
- 44:16 – 50:19
What evals look like in practice—and when PRDs still matter
Dianne gives a concrete eval example: early Claude struggled with structured outputs like JSON, so the team built a targeted eval set from real failures and tracked improvements over time. She also clarifies that PRDs remain valuable for alignment across large stakeholder groups and for ambiguous, visionary work (e.g., computer use).
- •Example eval: instruction-following failures often meant “wrong JSON output”
- •Create 30–40 representative examples to form an eval set; run it every model iteration
- •Evals enable measurable progress and help handle non-determinism and jagged performance
- •PRDs still used for launches and cross-functional alignment (engineering, legal, safety, etc.)
- •PRDs especially useful for ambiguous opportunities where user pain points aren’t yet well-defined
- 50:19 – 58:11
Hands-on leadership and finding joy: why managers must ship, and why experimentation is social
Dianne argues that managers can’t lead AI product work without being hands-on and shipping with the tech themselves. To sustain motivation, she recommends pairing with excited peers, going deep on a few tools, and building virtuous cycles of shared discovery rather than treating AI as a checkbox.
- •Leaders must be hands-on: onboarding is the same regardless of tenure; everyone must learn via real usage
- •Carve time to own workstreams to maintain “theory of mind” for model/product pace
- •Joy comes from shared discovery—pair with someone excited and explore together
- •Go deep on 1–2 tools/use cases instead of shallowly trying everything
- •Internal “experimenting in public” as a culture mechanism for sustained learning
- 58:11 – 1:03:53
How Dianne uses Claude: agentic workflows, management coaching, and avoiding overreliance
Dianne shares practical personal use: using Claude/Skills to prepare for difficult conversations and become a better coach, not just to boost “IQ” tasks. She then addresses overreliance by describing when she forms her own POV first versus delegating standardized writing, emphasizing that verification and accountability matter as much as authorship.
- •Using Claude for coaching/management prep (e.g., “Crucial Conversations” guidance)
- •AI as EQ augmentation: brainstorming phrasing, trust-building, and directness
- •Avoiding overreliance: form your POV first, then use Claude as sparring partner
- •Delegate low-leverage writing (e.g., business reviews) while focusing human effort on thinking/judgment
- •Key distinction: who verifies/signs off is often more important than who drafts
- 1:03:53 – 1:14:11
Claude’s constitution and better thinking partners: why pushback beats compliance, plus what humans retain
They explore why alignment and Claude’s “constitution” can make the model more useful—because it knows when to push back and help users reach better outcomes. The discussion broadens to what remains distinctly human for now: judgment, persistence, proactivity, and domain expertise in fields not yet on the steep part of the curve.
- •A good thinking partner shouldn’t just agree; pushback improves outcomes
- •Alignment traits support proactivity and principled disagreement, not just refusal behavior
- •Using Claude even for meta decisions (e.g., pricing strategy brainstorming) to refine reasoning
- •Human value: judgment built from nuanced lived experience; choosing what to build among infinite options
- •Emerging domains beyond software: biology/life sciences; Anthropic investing in specialized efforts (e.g., Claude Science)
- 1:14:11 – 1:33:50
Kids, burnout, culture—and closing: sustainability via team trust, low ego, and feedback loops
Dianne emphasizes teaching kids curiosity, persistence, and an inner voice in an AI-saturated world. She explains burnout avoidance through high-trust teamwork (“not an individual sport”), then closes with a defense of PM value: user-centric detail work and actionable translation are more needed, not less, as building becomes easier.
- •Raising kids for an AI future: curiosity, persistence, and developing an independent inner voice
- •Burnout prevention: high-performance pace requires team collaboration, not heroics
- •“Hive mind” and low-ego culture let people take real PTO without the system breaking
- •PMs still matter: deciding what to build and grounding tech in user needs is the bottleneck
- •Call to action: give product feedback (thumbs up/down, customer channels) and hiring for curious, first-principles tinkerers