Skip to content
Aakash GuptaAakash Gupta

How AI PMs Ship Features Users Love (Descript CEO Explains)

Laura Burkhauser went from IC PM to CEO of DeScript ($550M AI video editing platform) in just 3 years. Here's every AI feature she shipped along the way - and the exact career playbook to rise through product. Summary: https://www.news.aakashg.com/p/descript-ceo-laura-burkhauser Transcript: https://www.aakashg.com/laura-burkhauser-descript-ceo/ ---- Timestamps 0:00 Intro 1:35 Laura Welcome 1:49 Features That Led to CEO Promotion 4:19 The Great AI Boom 6:27 Feature Timeline 10:05 Rolling Out AI Tools 12:34 Ads 14:13 Measuring Success (topic after ads) 17:19 PM's Role in AI Features 24:46 Building Underlord 32:02 Ads 34:32 Quantifying Success (topic after ads) 41:31 Career Before Descript 48:25 IC to CEO Progression 53:32 Outro ---- Thanks to our sponsors: 1. Maven: Improve your PM skills with awesome courses. Discount with my link - https://maven.com/x/aakash 2. Pendo: #1 Software Experience Management Platform - http://www.pendo.com/aakash 3. Vanta: Automate compliance across 35+ frameworks like SOC 2 and ISO 27001 - http://vanta.com/aakash 4. NayaOne: Airgapped cloud-agnostic sandbox - https://nayaone.com/aakash/ 5. Kameleoon: Leading AI experimentation platform - http://www.kameleoon.com/ ---- Key Takeaways 1. Map your user journey BEFORE picking AI features. Descript identified pain points (retakes, eye contact, rambling), then asked "what just became possible with LLMs?" Build that intersection. 2. Build prepackaged buttons, not blank chat boxes. Each Descript AI tool is a carefully crafted prompt behind a single button that delivers reliable results every time. 3. Use human evals on production data before shipping. Test on real customer data, ask "would I use this as a customer?" If yes, ship. If no, don't. 4. The ultimate metric is export rate. If users apply your AI feature then remove it before exporting, it didn't meet their quality bar. 5. Switch from buttons to chat when you hit 30+ parameters. When users wanted topic selection, speaker choice, and platform optimization, chat became better than buttons. 6. Match your eval data to actual use case. Descript failed with Studio Sound because they tested on terrible audio (vacuuming, jackhammers) when real users had laptop microphones. Different models handle different quality levels. 7. Test agents with real customer language early. Don't use toy data or employee terminology. Mix sophistication levels—some advanced at video and AI, some complete beginners—to understand how real people prompt. 8. Launch AI agents to new users first. Video editing is hard and many people quit. Descript tested Underlord on activation and it won, so new users got it first before existing users. 9. Choose breadth over depth for product-wide agents. Descript chose breadth—Underlord works across all features because "we're not a point solution." Requires more context, tool coverage, and evals but serves the product vision. 10. Earn founder trust by getting command, not by being strategic. Use the product extensively. Talk to customers constantly. When you speak, people think "Smart" and invite you to more rooms. Ship features before focusing on strategy. ---- 👩‍💼 Where to find Laura Burkhauser: LinkedIn: https://www.linkedin.com/in/burkhauser/ Company: https://web.descript.com/ 👨‍💻 Where to find Aakash: Twitter: https://www.x.com/aakashg0 LinkedIn: https://www.linkedin.com/in/aagupta/ Newsletter: https://www.news.aakashg.com #Descript #AIProductManagement #CareerGrowth --- About Product Growth: The world's largest podcast focused solely on product + growth, with over 187K listeners. Hosted by Aakash Gupta, who spent 16 years in PM, rising to VP of product, this 2x/week show covers product and growth topics in depth. Subscribe and turn on notifications to get more videos like this.

Laura BurkhauserguestAakash Guptahost
Dec 15, 202554mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 4:10

    Why great products transform identity (and why Descript clicked)

    Laura opens with a product philosophy: the best tools don’t just complete tasks—they change how users feel about themselves. She explains how Descript’s “edit video like a doc” experience created an immediate product-love moment that later shaped her approach to shipping AI features.

    • Products can be identity-changing, not just utility-providing
    • Descript’s core UX: transcript-first editing instead of timeline-first
    • Immediate product-love/PMF moments often come from removing dread (timeline scrubbing)
    • User joy becomes a compass for what to build next
  2. 4:10 – 5:11

    From AI-native foundation to the ‘Great AI Boom’—choosing real user jobs

    Laura frames Descript as AI-native before AI hype and describes the shift once LLMs became mainstream. She explains the guiding principle for new AI features: package prompts behind reliable, job-focused buttons grounded in known editing pain points.

    • Descript used AI early, but users didn’t care until the AI boom
    • LLMs excel at language—perfect fit for script-based editing workflows
    • AI features should map to explicit user jobs, not vague ‘AI’ capabilities
    • “Parameterised, job-based buttons” prioritize repeatable reliability
  3. 5:11 – 6:58

    First wave AI tools demo: remove retakes (and why it works)

    Laura walks through Descript’s AI tools panel and demonstrates ‘Remove Retakes’ as a concrete LLM-enabled editing workflow. She highlights how straightforward prompting, aligned to a high-frequency customer need, can produce high-leverage editing automation.

    • AI tools menu includes: edit for clarity, remove filler words, remove retakes, add chapters
    • Remove Retakes stitches best takes together automatically
    • Simple prompting can be powerful when paired with clear user intent
    • Reliability matters more than novelty for core editing tasks
  4. 6:58 – 7:58

    Feature timeline constraints: context windows, chunking, and picking feasible use cases

    Aakash probes how Descript selected early LLM features amid technical limits. Laura explains how chunking enabled some tasks (retakes) while broader rewriting needed longer context, forcing careful scoping and sequencing.

    • Early LLM limitations: context window made full-transcript tasks hard
    • Chunking works for localized tasks like retake detection
    • Some jobs (rewrite) require global context and were gated by tech maturity
    • Sequencing features depended on feasibility + user value
  5. 7:58 – 10:36

    Customer journey mapping: scripted vs improvised creators (Eye Contact, retakes, clarity)

    Laura describes how Descript’s understanding of creator workflows drove which AI tools to build. She segments customers by recording style and maps each segment to a best-fit capability: Eye Contact for script readers, Retakes for line-by-line recorders, and Edit for Clarity for improvised speakers.

    • Deep customer observation preceded LLM adoption
    • Two scripted creator problems: eye line vs frequent retakes
    • Eye Contact feature integrates a model to correct gaze
    • Edit for Clarity targets rambling/improvised recordings
  6. 10:36 – 14:43

    Shipping AI buttons: ‘six killer apps,’ fast-moving backlog, and public beta confidence

    Laura explains the initial ‘Underlord’ concept as an AI toolbar composed of multiple high-value actions. The team balanced what was possible now vs soon, kept a rolling backlog for tech-unblocked ideas, and relied heavily on hands-on testing to decide readiness.

    • Toolbar needed a portfolio: roughly six strong AI actions at launch
    • Mixture of editing actions + ‘publish’ helpers (titles, show notes, descriptions)
    • Roadmap included “blocked by current model limits” ideas for rapid later shipping
    • Readiness test: run on production data and ask “would I use this?”
  7. 14:43 – 16:14

    Rollout strategy: public beta, iterative tweaks, and when production data matters most

    Aakash challenges how Descript avoided lab-only validation. Laura outlines a pragmatic release approach: public beta for straightforward tools, A/B tests for refinements, and a higher bar for open-ended agent experiences where real-world usage produces surprising edge cases.

    • Beta approach: public beta (not private) for early, bounded tools
    • Team relied on internal heavy usage + production data access
    • A/B tests used for safe iteration on prompts/behavior
    • Open-ended agents amplify weird edge cases—production data becomes critical
  8. 16:14 – 17:27

    Measuring AI tools: adoption, retention, exports, and lightweight user feedback

    Laura details how Descript defined success for AI tools using behavior-based metrics rather than abstract model scores. They benchmarked new tools against the proven ‘Remove Filler Words’ feature and validated quality by whether users exported final content with the AI edits intact.

    • Primary metrics: adoption and retention
    • Baseline comparison: Remove Filler Words as a proven reference point
    • Outcome quality proxy: did users export with the AI edits applied?
    • Thumbs up/down feedback supports training and iteration
  9. 17:27 – 20:41

    The PM’s unique role in AI: writing eval criteria (and adding domain taste)

    Laura argues PMs still own core product thinking, but AI makes eval design central. PMs are best positioned to define what ‘great’ vs ‘acceptable’ vs ‘harmful’ output means, while also bringing in domain experts (e.g., audiophiles) when the work requires specialized taste.

    • PM owns job-to-be-done framing and what success means
    • PM is uniquely qualified to codify eval criteria for outputs
    • Example nuance: limit jump cuts per time window for ‘Edit for Clarity’
    • For creative quality, bring expert judges (e.g., pro musician for audio)
  10. 20:41 – 22:48

    How Descript operationalizes evals: human criteria → human judging → LLM judge

    Aakash asks what the PM’s output looks like in practice. Laura describes a pipeline where engineers set up eval tooling, while PMs define pass/high-pass/fail criteria on representative production queries, gradually training and trusting an LLM judge with humans as the tiebreaker.

    • Engineers/research set up eval infrastructure; PM defines decision criteria
    • Representative production queries are sampled and withheld from engineers
    • Process: human judge → human+LLM judge → reconcile disagreements
    • Goal: eventually delegate routine judging to LLM with confidence
  11. 22:48 – 26:12

    Failure case study: Studio Sound evals broke when the use case wasn’t defined

    Laura shares a concrete mistake: after replacing an expert ‘taste’ evaluator with a checklist, the model got worse because the test data didn’t represent the primary customer scenario. Optimizing for extreme ‘terrible audio’ cases degraded performance for the far more common ‘okay laptop mic’ use case.

    • Codifying taste into a checklist can lose crucial nuance
    • Different models excel at different audio conditions—no one model rules all
    • Using unrepresentative eval samples leads to wrong optimization target
    • Define the primary use case explicitly before tuning models
  12. 26:12 – 30:12

    Why agents: Create Clips ‘knob explosion’ and the case for objective-based chat

    Laura explains the product pressure that led from buttons to an agent: feature requests kept adding parameters to ‘Create Clips’ until the UI became unwieldy. Underlord addresses this by letting users express objectives or highly customized workflows in natural language instead of endless controls.

    • Create Clips workflow surfaced escalating customization demands
    • Too many parameters turns buttons into an unusable control panel
    • Underlord is for objectives (e.g., ‘get this to 90 seconds’) and custom workflows
    • Chat isn’t always better than buttons—but becomes better past a complexity threshold
  13. 30:12 – 35:15

    Building Underlord: breadth vs depth, tool coverage, context, and ‘woolly mouse’ reality

    Laura describes Underlord as a harder, open-world co-editor and explains the architectural commitments behind it. Descript chose a breadth agent spanning many editor tools, requiring deeper context, broad tool access, and a report-card system—while acknowledging the current version is still a ‘woolly mouse’ on the path to a ‘mammoth.’

    • Underlord positioning: user stays creatively in control; agent executes
    • Hard choice: breadth agent across Descript vs narrow depth agent
    • Needs: context injection, comprehensive tool coverage, eval/report card system
    • Current system is useful but brittle; team is iterating toward a stronger harness
  14. 35:15 – 39:43

    Agent rollout and quantifying impact: regression tests → private alpha → activation lift

    Laura outlines a staged approach to shipping an open-ended agent: define regression suites, gather real customer prompts in private alpha, convert them into broader regression coverage, then test whether Underlord improves new-user activation. The team gradually expanded access from new users and opt-in cohorts toward a more default experience.

    • Didn’t quantify until broad exposure; early focus was capability + correctness
    • Private alpha captured real-world language from diverse skill levels (video + AI)
    • Regression prompts used in bug bashes; acknowledged overfitting risk
    • Key business test: Underlord improved activation by helping users ‘get over the hump’
  15. 39:43 – 41:21

    Underlord-native vs hybrid users—and what that means for product strategy

    Laura describes a split in the customer base: some prefer the traditional editor, others use a hybrid approach for bulk edits, and a growing segment is fully agent-native. She compares this to how non-engineers use coding agents—willing to re-prompt repeatedly to avoid learning the underlying skill.

    • Some customers don’t use Underlord; core editor remains valuable
    • Hybrid users delegate tedious bulk operations to the agent
    • Underlord-native users prefer prompting over learning video editing mechanics
    • Agent products shift UX expectations toward iterative dialogue loops
  16. 41:21 – 48:56

    Career path to CEO: consulting strategy, org leadership lessons, and cold-emailing Descript

    The conversation shifts to Laura’s background and how it prepared her for leadership at Descript. She credits consulting for strategic thinking, startups for operating instincts, Twitter for learning org leadership, and a product-love cold outreach that landed her at Descript without optimizing for title or pay.

    • Non-traditional background (German literature) → consulting → product
    • Consulting taught strategic decision-making; building a learning platform sparked product interest
    • Startup vs big-company fit is visceral; Twitter helped develop org-leader skills
    • Cold outreach driven by genuine product love led to joining Descript and quickly becoming VP Product
  17. 48:56 – 54:50

    From IC PM to ‘in the room’: earning founder trust through mastery and shipping

    Laura explains how she progressed at a founder-led startup by aligning with the founder’s vision and consistently adding value. Her advice: earn credibility through deep product/customer/business command, maintain excellence in core execution, and ship reliably before trying to ‘do strategy.’

    • Founder-led dynamic: prove why the founder should trust your instincts
    • Earn trust by using the product, talking to customers, and knowing the business
    • Avoid skipping fundamentals in pursuit of ‘strategy’ visibility
    • Credibility compounds through shipping; invitations to decision rooms follow

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.