Skip to content
Lenny's PodcastLenny's Podcast

Brendan Foody: How Mercor hit $500M run rate in 16 months

Lawyers, bankers, and engineers design rubrics and RL environments for top AI labs; the work pays up to $500 an hour as an entirely new job category.

Brendan FoodyguestLenny Rachitskyhost
Sep 18, 20251h 7mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 5:51

    Why AI labs will spend “whatever it takes” to improve models

    A cold open sets the stakes: the richest companies are pouring money into making models better, and evals are becoming the key lever. Lenny frames the episode around a new paradigm—measurement as the bottleneck—and previews Mercor’s explosive growth.

    • AI labs treat capability gains as worth near-unlimited spend
    • “Era of evals” as the central theme for model progress
    • Human expertise as the limiting factor for what models can’t yet do
    • Mercor’s breakout growth teased as evidence of a massive market shift
  2. 5:51 – 9:26

    Evals as PRDs and sales collateral: what “the era of evals” really means

    Brendan defines evals as the product requirements document for AI models—and increasingly the proof of capability. They discuss how reinforcement learning can rapidly “hill climb” a good eval set, making eval design the gating step to real-world automation.

    • If the model is the product, the eval is the PRD
    • Researchers iterate on eval sets through many experiments
    • RL makes progress fast once a strong eval exists (e.g., SWE-Bench)
    • Enterprises should eval their value chain to deploy AI effectively
    • “Evals are your new marketing” as the new capability demo layer
  3. 9:26 – 13:12

    From crowdsourcing to experts: the AI training market’s new center of gravity

    Lenny zooms out on the fastest-growing company categories and positions data/training companies alongside foundation models and “vibe coding” apps. Brendan explains the market transition from low-skill labeling to sourcing and vetting high-caliber professionals who can define and measure model capabilities.

    • Three hypergrowth buckets: model labs, coding apps, and training/data companies
    • Shift away from low/medium-skill crowdsourcing tasks
    • New bottleneck: sourcing vetted experts (engineers, bankers, doctors, lawyers)
    • Mercor’s origin: automating hiring workflows before pivoting into AI labs
  4. 13:12 – 15:23

    What experts actually do day-to-day: rubrics, verifiers, and eval design

    Brendan makes the work concrete with examples like legal redlining: experts define rubrics that score model outputs. These rubrics become both benchmarks and reinforcement signals—blurring the line between “evals” and “RL environments.”

    • Experts create rubrics/verifiers (e.g., what a great contract redline includes)
    • Evals and RL environments are often the same data viewed differently
    • Benchmarks guide research; verifiers guide reinforcement learning
    • Model improvement is bounded by what humans can still do better than models
  5. 15:23 – 17:10

    Understanding the training stack: SFT, RLHF, and the shift to RLAIF

    They outline the historical data types (supervised fine-tuning and RLHF) and why the ecosystem is moving toward reinforcement learning from AI feedback (RLAIF). Humans increasingly define success criteria while AI systems scale the feedback and optimization loop.

    • SFT as input/output training examples; RLHF as preference comparisons
    • Trend toward RLAIF: humans define success criteria; AI generates/grades at scale
    • Unit tests in code and rubrics in other domains as scalable evaluators
    • Data efficiency improves when rewards are well-specified
    • Post-training becomes the locus of practical capability gains
  6. 17:10 – 22:17

    Future of work: the economy as an RL environment—and new jobs created

    Brendan argues the job-displacement narrative misses the parallel creation of new work: building evals and environments that teach models. They discuss why human contribution persists as long as there are tasks humans can do that models can’t, and how people should upskill.

    • Human roles persist while any human advantage remains
    • The economy may become an “RL environment machine” of worlds and contexts
    • AI creates new categories of jobs around evaluation and feedback design
    • Advice for students: learn to use AI tools rather than fight them
    • Assessments should allow AI use and measure outcomes (e.g., build a product in 1 hour)
  7. 22:17 – 25:54

    Elastic demand skills: where AI increases opportunity instead of shrinking it

    Brendan introduces “elastic demand” as the key lens for career resilience: fields where making people more productive leads to more total output and thus more demand. They contrast accounting (bounded demand) with software (seemingly unbounded) and extend the idea to roles across company-building.

    • Elastic demand industries expand when productivity rises
    • Software as the most elastic: 10–100x more features/products become feasible
    • Product and engineering remain strong bets; leverage AI as a multiplier
    • Potential expansion in consulting/ops-style work when costs drop
    • Winning mindset: lean into abundance rather than resist displacement
  8. 25:54 – 29:56

    Labor markets are breaking: volume, automation, and the unified global marketplace thesis

    They connect Mercor’s original thesis—automating resume review and interviews—to today’s hiring reality: AI-enabled mass applying floods employers, forcing AI-driven filtering. Brendan describes a future of a global, unified labor market with near-perfect information flow, while work itself evolves.

    • Hiring is disaggregated and inefficient; companies see only a tiny slice of talent
    • AI makes applying easier, causing application volume explosions
    • Automation pushes the market toward AI screening and matching systems
    • Vision: a unified global labor market every candidate applies to and company hires from
    • Mercor’s dual identity: labor marketplace and data company for labs
  9. 29:56 – 34:54

    Pre-training vs post-training: how expert feedback upgrades real-world reasoning

    Using an X-ray diagnosis story, they clarify that Mercor’s expert work is largely post-training: teaching models what to prioritize, what’s correct, and how to reason. The quality of expert-created “stasis points” (diagnoses, rubrics) directly affects model decisions.

    • Pre-training loads broad knowledge; post-training shapes reasoning and priorities
    • Experts create reference points, rewards, penalties, and evaluation targets
    • Post-training helps models use the right context and reasoning chains
    • Scale of human effort: tens of thousands active, hundreds of thousands overall
    • Industry transition: fewer low-skill tasks, more high-skill expert input
  10. 34:54 – 38:58

    How Mercor operates: turnaround speed, top-10% talent leverage, and pay dynamics

    Brendan explains how Mercor fulfills specialized expert requests quickly and why a small fraction of contributors often drives most model improvement. They discuss project-based expert work, part-time vs full-time arrangements, and compensation that can reach elite professional rates.

    • 24-hour turnaround for sourcing specialized experts (e.g., award-winning writers)
    • Top 10% of experts often drive the majority of model improvement
    • Work is usually part-time, sometimes full-time, often alongside other jobs
    • Median pay ~$95/hr; can exceed ~$500/hr for deep expertise
    • Quality-focused economics differ sharply from legacy crowdsourcing rates
  11. 38:58 – 41:31

    Building Mercor’s hypergrowth engine: leading indicators, customer obsession, and values

    Lenny pushes on what enabled Mercor’s record growth, and Brendan emphasizes tracking leading indicators in fast markets and extreme customer obsession over sales/marketing. He shares the company’s operating values—can-do attitude, high standards, and intensity—and how they shape execution.

    • Find fast-moving demand pockets where customers pay “whatever it takes”
    • Customer obsession: minimal sales/marketing early, heavy product focus
    • Core values: can-do attitude, high standards, intensity
    • Ambitious goals as a forcing function for company trajectory
    • Hiring tradeoff: early patience for talent density, later speed once demand is proven
  12. 41:31 – 56:55

    Founder lessons and origin stories: xAI signal, incumbents failing talent, and PMF pull

    Brendan recounts pivotal moments that revealed the market’s size: early exposure to frontier labs and a painful incident with an incumbent that mishandled payments. The key takeaway is to listen for pull—customers that are surprisingly easy to sell—while staying flexible about how the thesis manifests.

    • Early meeting with xAI reinforced demand for high-quality expert pipelines
    • Incumbent failure created an opening to work directly with labs
    • Cutting middlemen while protecting expert dignity and incentives
    • PMF lesson: stop forcing; follow strong market pull and easy sales motion
    • Capital efficiency: bootstrapped early, lifetime profitable, rapid scale thereafter
  13. 56:55 – 1:07:07

    What’s next for model progress—and how Brendan personally uses AI

    They close on why evals remain evergreen, why superintelligence timelines may be longer, and why post-training data will matter more than simply scaling pre-training. Brendan shares his own AI workflows (writing, thought-partnering via voice), followed by a rapid lightning round and final advice: build things.

    • Evals are evergreen: improvement requires defining success across domains
    • Progress may slow vs hype; superintelligence fear is often overstated
    • Future gains likely driven by post-training and better reward design
    • Personal AI usage: document drafting and voice-mode reasoning partner
    • Final advice: take initiative—use AI to build and learn faster

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.