Skip to content
Lenny's PodcastLenny's Podcast

Jason Droege: Why AI still needs someone to dig up the road

Through Meta's $14B stake and expert networks of doctors, engineers, and PhDs: production AI now takes 6 to 12 months, and Uber Eats lessons still apply.

Lenny RachitskyhostJason Droegeguest
Oct 9, 20251h 24mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:31

    Cold open: Why enterprise AI progress feels slower than headlines

    Lenny and Jason set expectations: real enterprise automation takes time, operational work, and iteration. Jason frames today’s shift from “models knowing things” to “models doing things,” and why that changes what training data must look like.

    • Enterprise AI often needs 6–12 months to become robust enough for mission-critical workflows
    • Tech revolutions require unglamorous operational build-out (the “digging up roads for broadband” analogy)
    • Key near-term trend: models moving from knowledge to action/decision-making (agents)
  2. 1:31 – 6:10

    Who Jason Droege is—and what this episode will cover

    Lenny introduces Jason as the new CEO of Scale AI and the first to speak publicly after the Meta investment and Alex Wang’s move. The episode roadmap: how expert data, labeling, and evals are shaping frontier model progress, plus Jason’s product and leadership lessons from Uber Eats and beyond.

    • Jason’s background: co-founding with Travis Kalanick, building Uber Eats, now leading Scale AI
    • Scale’s role in training data, expert labeling, and evals for frontier labs
    • Why experts (doctors, engineers, PhDs) increasingly drive model improvements
    • Promise: tactical product/leadership insights in addition to AI-market depth
  3. 6:10 – 9:12

    Early career at Scour with Travis Kalanick: “Everything’s negotiable”

    Jason recounts building Scour from a UCLA dorm room and learning how deals and power dynamics really work. The takeaway: there’s no fixed playbook—outcomes are shaped by negotiation, incentives, and persistence.

    • Building and running infrastructure scrappily from a dorm room
    • Fundraising terms can change dramatically; reality differs from “how it’s supposed to work”
    • Lesson: business norms are flexible if incentives align
    • How early experiences influenced later operating style at Uber
  4. 9:12 – 10:28

    Getting sued for a quarter-trillion dollars: the real-world shock of regulation and incumbents

    Scour’s peer-to-peer use cases attracted massive legal pressure from entertainment industry groups. Jason describes the mismatch between headline claims and final outcomes—and how incumbents use legal leverage to force startups out.

    • Scour was used to find free content in a legally ambiguous era
    • RIAA/NBAA sued for $250T and settled for $1M—signaling intent to bankrupt, not “collect damages”
    • A lesson in how established players apply pressure without a clear “rulebook”
    • How early existential threats shape later risk calculus
  5. 10:28 – 12:30

    Meta’s $14B investment: what actually changed at Scale AI (and what didn’t)

    Jason clarifies the structure and implications of Meta’s investment: Scale remains independent, with unchanged governance and security boundaries. He also addresses misconceptions about preferential access and the company’s ongoing growth.

    • Meta invested ~$14B for 49% non-voting stock; no new board seat taken
    • Scale remains independent; governance and privacy/security practices remain intact
    • Only ~15 people moved over as part of the transaction; Scale ~1,100 employees
    • Scale has two major businesses with hundreds of millions in revenue each
  6. 12:30 – 17:02

    From autonomous vehicles to GenAI: how Scale’s data work evolved with model needs

    Jason explains Scale’s throughline since 2016: models improve when data improves, and data needs change as capability rises. He argues competitors’ “Scale can’t do expert work” positioning is misleading—expert workflows are now core.

    • Scale began with AV labeling, expanded into computer vision and government use cases
    • GenAI accelerated demand for more sophisticated, high-skill tasks
    • Labeling shifted from simple preferences (e.g., story comparisons) to hours-long expert tasks
    • Scale’s expert network: ~80% bachelor’s+; ~15% PhDs
  7. 17:02 – 18:48

    Building and sustaining an expert network: referrals, campuses, and experience design

    Lenny probes the hardest part: finding and retaining high-end experts. Jason shares the multi-channel playbook and emphasizes that the best acquisition channel is a great contributor experience that drives referrals.

    • Experts are hard to recruit; no single channel works
    • Strongest funnel: expert-to-expert referrals driven by good experience and meaningful work
    • Campus programs: engaging professors and students directly
    • Traditional sourcing (e.g., LinkedIn) complements grassroots tactics
  8. 18:48 – 22:03

    Reinforcement learning environments: training agents to operate real systems

    Jason describes the industry’s move toward RL environments—sandboxes where agents learn to complete goals reliably. The challenge shifts to designing environments and tasks that generalize across countless real-world permutations.

    • RL environments simulate tools like Salesforce, with real configurations and data types
    • Training must include escalation to humans when confidence is low
    • Key research question: what data/tasks generalize vs. require bespoke collection
    • Goal: maximize usefulness and reliability of agents in real workflows
  9. 22:03 – 23:33

    Concrete examples: what “expert data” actually looks like in practice

    They get specific about artifacts being produced—code, annotations, rationale, and debugging traces—not just outputs. The emphasis is on capturing decision-making, not merely final answers.

    • Web-dev tasks may include both the final site and the reasoning behind decisions
    • Data can include annotations like “why this choice, why not the alternative”
    • Debugging datasets can include broken examples plus explanations
    • Data format depends on what model builders are training for (generation vs. diagnosis vs. repair)
  10. 23:33 – 28:18

    Enterprise AI: digitizing judgment inside a healthcare system

    Jason shares a real enterprise implementation: AI that summarizes massive patient documentation and surfaces key diagnostic/treatment considerations. The deeper lesson: enterprises increasingly must label their own internal judgments because “off-the-shelf + RAG” hits a ceiling.

    • Doctors face 200–300 pages of mixed-format documentation per case
    • AI can surface top considerations (e.g., allergies/medication conflicts) and reduce revisits
    • Enterprises are becoming labeling sites: internal experts encode local context and standards
    • Bottleneck: converting deep, organization-specific judgment into usable training signals
  11. 28:18 – 31:14

    Humans in the loop (for a long time): why new knowledge keeps arriving

    Lenny asks the existential question for data-labeling businesses: when do humans become unnecessary? Jason argues the need persists because human knowledge and skills keep evolving, and mission-critical systems require ongoing oversight and updates.

    • Scale’s history is “new beginnings”: as one labeling need fades, new ones emerge
    • If models need no new human data, society has reached an almost unfathomable state
    • Operations must continually discover new model gaps and recruit relevant expertise
    • Human adaptability matters more than doom-or-boom narratives suggest
  12. 31:14 – 41:34

    Evals, agentic futures, and why enterprise AI adoption is ‘easy to learn, hard to master’

    Jason explains evals as establishing “what good looks like,” especially in enterprise and government. He then forecasts the next 2–3 years as an agentic shift—and explains why real deployments lag: the last mile of reliability, approvals, and change management is the hard part.

    • Evals define benchmarks for acceptable behavior; in enterprise/government they dominate the work
    • These systems are probabilistic, so “good” can matter more than a single “correct” answer
    • Near-term trajectory: models moving from knowing to doing; change management becomes the constraint
    • POCs often stall at ~60–70%; production takes months due to reliability, legal, regulatory, and adoption hurdles
  13. 41:34 – 53:04

    Uber Eats origin story: customer incentives, unit economics, and finding urgency

    The conversation shifts to product building: Jason’s method is deep incentive analysis and independent triangulation rather than literal customer requests. He recounts reverse-engineering restaurant economics to set pricing and ensure marketplace viability.

    • Customer closeness means understanding incentives (financial, ego, career risk), not just listening
    • Uber Eats early research: reverse-engineering ingredient and labor costs to model margins
    • Marketplace pricing: finding the clearing rate that gets all sides to participate
    • Key product insight: urgency matters—solve what buyers think about daily, not annually
  14. 53:04 – 57:07

    Picking the right bets: wide aperture experimentation → scaling to $80B

    Jason describes exploring many delivery-adjacent ideas before committing to food delivery. He shares failed experiments (like convenience-store trucks), why grocery scared him, and what signals made Eats the clear winner—then the speed of scaling once it clicked.

    • Strategy: keep a wide aperture until evidence coalesces; test “bad” ideas long enough to be sure
    • Failed experiment: convenience-store truck (wrong SKU mix; misunderstood demand drivers)
    • Food delivery showed strong early signals + workable unit economics; also supported local restaurants
    • Scaling milestones: launch in 2015, ~$20B in ~4.5 years; accelerated dramatically during COVID and later reached ~ $80B
  15. 57:07 – 1:00:13

    The McDonald’s deal: strategy, stubbornness, and operational mayhem at global scale

    Jason recounts initially rejecting McDonald’s because it didn’t fit the “help the little guy” narrative—then learning the value of the partnership. The rollout demonstrates how startups and large incumbents collide in speed, expectations, and process maturity.

    • McDonald’s approached Uber Eats; Jason initially said no for brand/mission reasons
    • Delaying engagement helped secure a strong (and effectively exclusive) partnership
    • Chains were hesitant due to basket-size sensitivity; Uber culture pushed to “make it work”
    • Global onboarding in ~6 months created massive operational complexity for a young org
  16. 1:00:13 – 1:24:01

    Jason’s operating principles: gross margin realism, risk management, hiring, and personal AI workflows

    Jason shares a practical toolkit for evaluating business ideas (gross margin as a fast filter), avoiding existential risk (“not losing” as prerequisite to winning), and building resilient teams. He closes with how he uses AI daily (tutoring and summarizing), then a lightning round of books, media, and personal mottos.

    • Gross margin is a coarse but powerful filter for differentiation and future margin compression risk
    • “Not losing is a precursor to winning”: survive long enough for timing/insight to turn
    • Hiring: optimize for curiosity/problem-solving, cross-functional humility, and leadership; reserve “must-have experience” for a small set of roles
    • AI usage: voice-mode tutoring for fast learning; summarizing internal docs to find what matters
    • Lightning round: key books (The Selfish Gene, Good to Great, Thinking Fast and Slow), motto (“the end is never the end”), and more

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.