Skip to content
The Twenty Minute VCThe Twenty Minute VC

Mercor Head of Product on Revenue Concentration from Frontier Labs

Osvald Nitski is the Head of Product at Mercor, the AI-training and expert-data marketplace powering frontier-model development. Mercor last raised a $350 million Series C at a $10 billion valuation, and is reportedly in discussions for a new round at a $20 billion valuation. Mercor crossed $2BN in ARR in June; doubling from $1 billion in only four months. ----------------------------------------------- Timestamps: 00:00 Intro 01:12 Does Open-Source Cannibalize Mercor's Core Business? 02:29 Why 90% of Enterprise Workflows Can't Be Done With Open Models 07:47 Do We Have an Enterprise AI ROI Problem? 09:05 Balancing Token Spend vs Performance 10:07 Salesforce Spends $300M on Anthropic 12:13 AI Makes the PM Role Harder, Not Easier 14:55 The Biggest Product Mistake 22:42 Why RL Environments Are the Fastest-Growing Data Type Right Now 28:28 How Mercor's Product Teams Are Structured 33:54 The Three Secrets to Scaling Supply on the Marketplace 38:04 Mercor Has High Revenue Concentration 43:47 Why Data Projects for Enterprises Are So Operationally Intense 51:02 Why Cybersecurity Data Will Never Hit the 90% Sufficiency Ceiling 52:15 Hiring in SF: Brutal Talent War & What Mercor Looks For 54:15 Quick-Fire Round ---------------------------------------------------------------------------------------------- Subscribe on Spotify: https://open.spotify.com/show/3j2KMcZTtgTNBKwtZBMHvl?si=85bc9196860e4466 Subscribe on Apple Podcasts: https://podcasts.apple.com/us/podcast/the-twenty-minute-vc-20vc-venture-capital-startup/id958230465 Follow Harry Stebbings on X: https://twitter.com/HarryStebbings Follow Osvald Nitski on X: https://twitter.com/OsvaldNitski Follow 20VC on Instagram: https://www.instagram.com/20vchq Follow 20VC on TikTok: https://www.tiktok.com/@20vc_tok Visit our Website: https://www.20vc.com Subscribe to our Newsletter: https://www.thetwentyminutevc.com/contact ----------------------------------------------- #20vc #harrystebbings #founder #entrepreneur #mercor #ai

Osvald NitskiguestHarry Stebbingshost
Jul 25, 20261h 2mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:14

    Mercor’s demand surge and why data stays valuable at the frontier

    Osvald and Harry set the context: Mercor is seeing overwhelming demand and strong cash flow, and the discussion will focus on how data businesses evolve as model capabilities improve. The framing is that data retains value where models still have gaps—especially at the frontier of performance.

    • Mercor’s business health: demand outstrips ability to spend to meet it
    • Core question: what happens to data providers as models get better
    • Data value concentrates at the performance frontier, not in commoditized tasks
  2. 1:14 – 2:09

    Do open-source models cannibalize closed-model data demand?

    Harry challenges the common narrative that open-source will replace frontier closed models, potentially shrinking Mercor’s market. Osvald argues open models raise the baseline and reduce demand only for tasks they already solve, while frontier gaps continue creating new demand for eval and training data.

    • Open-source raises the floor; frontier gaps still drive purchases
    • Customers buy data to fill capability gaps tied to specific goals
    • Cannibalization happens only for tasks already solved by a given open model
  3. 2:09 – 3:23

    Why the “90% of enterprise workflows” framing is misleading

    Osvald disputes the claim that most enterprise workflows are already solvable by current models, arguing it ignores latent demand—especially long-horizon automation no one is attempting yet. He also reframes “can models do it?” into two categories: sufficiency-based tasks vs domains with uncapped improvement potential.

    • Latent demand: long-horizon agents (e.g., months-long procurement automation)
    • Apex benchmarks suggest ~50% on long-horizon workflows for top models
    • Sufficiency tasks (e.g., CRM updates) vs uncapped domains (law/medicine)
    • Binary % framing breaks for tasks with continuous rewards
  4. 3:23 – 6:12

    Enterprise data sensitivity: which workflows stay closed vs go open

    The conversation turns to enterprise skepticism about sharing sensitive data with frontier providers. Osvald explains sensitivity depends on how core and differentiating the workflow is; generic functions (HR/procurement) are easier to externalize, while core legal or advisory work triggers more caution.

    • Sensitivity varies by how core the workflow is to competitive advantage
    • Generic functions are less sensitive; differentiating work is more protected
    • Open weights can be safer depending on where inference runs (control over deployment)
  5. 6:12 – 7:29

    Specialized models per company and the resulting data flywheel

    Harry asks whether every company will end up with specialized models tailored to its priorities. Osvald agrees and notes this future is also “self-serving” for Mercor because specialization requires enterprise-specific evals and training data to define objectives and measure success.

    • Company-specific goals drive specialization (growth vs margin vs constraints)
    • Specialized models require proprietary eval + training data
    • ROI-gated adoption: specialization increases as value justification improves
  6. 7:29 – 8:31

    Is there an enterprise AI ROI problem—or just a waiting game?

    Osvald argues ROI pressure is not the dominant dynamic yet; enterprises are still exploring while token prices and model performance shift quickly. He acknowledges early tightening in some areas but says most organizations are still in a “let’s see what happens” phase.

    • Current phase: experimentation with higher tolerance for uncertainty
    • ROI math is unstable due to fast-moving token economics and performance
    • Some budget tightening is emerging, but not a broad ROI crisis
  7. 8:31 – 10:08

    Balancing token spend vs performance: use-case-driven budgeting

    Osvald gives a pragmatic framework: spend depends on whether tokens are COGS-like (e.g., customer service) or compounding productivity (e.g., coding agents for engineers). Founders should align spend with unit economics and growth impact rather than applying blanket caps.

    • Token spend should map to the economics of the specific workflow
    • Coding agent spend can be leverage, not pure COGS
    • Customer service agents must clear unit economics vs revenue served
    • Hypergrowth contexts tolerate more spend to meet demand
  8. 10:08 – 12:14

    Salesforce’s $300M Anthropic spend and where AI budgets are headed

    Harry brings up Salesforce’s reported $300M annual spend on Anthropic and asks if that’s the new normal. Osvald calls for better internal accounting by team type (R&D vs forward-deployed) and predicts the overall percentage of spend allocated to tokens will rise over time—potentially dramatically for hypergrowth companies like Mercor.

    • Different org functions warrant different spend profiles and tolerance
    • Need outcome-based accounting for token spend effectiveness
    • Macro trend: token spend share likely increases beyond today’s low single digits
    • Mercor’s reality: spending can already exceed salaries due to demand pressure
  9. 12:14 – 14:43

    AI makes product management harder: controlling surface area and boosting judgment

    Osvald argues faster engineering doesn’t mean shipping more features; it increases the risk of chaotic product sprawl. The PM bottleneck becomes choosing the right work, simplifying workflows, and maintaining strong judgment—leading to a higher PM-to-engineer ratio over time.

    • Core battle: reduce product surface area despite faster build velocity
    • Engineering is less bottlenecked; identifying revenue-driving workflows is harder
    • PM excellence shifts from tool mastery to judgment and business impact
    • Prediction: higher PM-to-engineer ratio as coding agents accelerate velocity
  10. 14:43 – 18:33

    Biggest product mistake: over-flexible tooling and the need for guardrails

    Osvald describes a key misstep: Mercor’s annotation platform became too flexible by trying to support every customer request across rapidly changing data formats. The result was operational chaos; the lesson was to introduce guardrails, align with ops, and invest only where enduring demand exists.

    • Annotation workflows diversified (SFT → preferences → rubric → environments, multimodal)
    • Maximal flexibility created complexity across hundreds of concurrent projects
    • Guardrails and best-practice standardization should have come earlier
    • Enduring-demand decisions rely on market proximity and leadership judgment
  11. 18:33 – 28:18

    Why services are rising in AI deployment (and why that may be temporary)

    The discussion shifts to the growth of services arms (e.g., Microsoft) and whether services indicate weak products. Osvald’s take: services fill a knowledge dissemination gap—AI deployment expertise is concentrated (notably in SF) and not yet broadly available—though over the long run it should become a standard internal capability.

    • Services are a short-term bridge while AI expertise spreads through industry
    • Talent scarcity: not enough in-house capability at every enterprise yet
    • Forward-deployed roles fit engineers strong in communication and simplification
    • Long-term arc: deployment becomes a common job function, less service-heavy
  12. 28:18 – 33:43

    How Mercor structures product teams: Marketplace vs Studio, pods, and cadence

    Osvald explains Mercor’s two core product areas—Marketplace (matching experts to jobs) and Studio (annotation/evals platform)—and how they support talent-only vs managed-service engagement models. He details pod-based sprint planning, weekly product-area meetings, and the scaling challenge of cross-team communication as headcount grows.

    • Two product teams: Marketplace and Studio (annotation/evals)
    • Two engagement modes: talent-only vs managed service (end-to-end datasets)
    • Small pods (PMs + eng + DS + flexible design) with sprint planning per pod
    • Weekly product-area sync for alignment; cross-product communication is the scaling pain
  13. 33:43 – 38:05

    Scaling marketplace supply: expert experience, referrals, and global sourcing

    Harry asks how Mercor scaled expert supply efficiently. Osvald attributes it to a dignified, reliable expert experience (pay, transparency, communication), which fuels referrals, plus a sourcing team that can find niche skills globally to meet spiky demand.

    • Reliable pay and strong treatment create a referral flywheel
    • Retention comes from sustained high-quality experience, not gimmicky bonuses
    • Clear instructions and communication reduce friction in unfamiliar online work
    • Global sourcing fills specialized skill gaps during demand spikes
  14. 38:05 – 39:27

    Revenue concentration risk and the push downmarket via self-serve human data projects

    Harry raises concerns about Mercor’s revenue concentration with frontier labs. Osvald frames Mercor’s biggest product challenge as moving downmarket: making human data projects self-serve and efficient for enterprises (e.g., AI project managers), enabling more customers, smaller heterogeneous projects, and reduced concentration.

    • Current lab projects are white-glove and ops-heavy
    • Strategic goal: democratize human data projects for enterprises
    • Product direction: self-serve workflows + AI project management assistance
    • Downmarket expansion diversifies revenue away from a few labs
  15. 39:27 – 41:11

    Why enterprise data projects are operationally intense: edge cases, paranoia, and rapid data evolution

    Osvald explains what makes human data projects hard to “simplify”: translating complex customer intent into precise guidelines, handling endless edge cases, and maintaining rigorous QC under time pressure. Complexity also comes from fast-changing data types across projects and modalities, demanding constant process adaptation.

    • Edge cases dominate and require rapid alignment among stakeholders
    • Ops must be highly paranoid: every datapoint must match evolving specs
    • Crisp communication is critical across customers, ops, and experts
    • Data types evolve quickly, increasing between-project complexity
  16. 41:11 – 51:59

    Fastest-growing data type: RL environments—and why cybersecurity never reaches “sufficiency”

    Osvald highlights RL environments as the fastest-growing data type: high-fidelity simulated apps/world states that resemble deployment conditions for agent training and eval. He also connects this to cybersecurity, where adversarial dynamics mean the goalposts always move—so performance is never “done,” sustaining continuous demand for evolving data.

    • RL environments: simulated tools/apps + rich start states + task suites
    • Training/evals increasingly mirror real deployment contexts (high-fidelity mocks)
    • Cybersecurity is adversarial with uncapped rewards; goalposts continually shift
    • Implication: security won’t hit a stable “90% sufficiency” ceiling
  17. 51:59 – 1:02:42

    Hiring in SF and Mercor’s interview focus: agency, experiments, and AI boundaries (plus quick-fire)

    The closing stretch covers SF’s brutal talent market and what Mercor prioritizes: agency/ownership, strong experimental thinking, stats literacy, and systems design—while avoiding over-delegating judgment to models. In quick-fire, Osvald discusses changing his mind on RL environments, advice to students (get internships), Mercor’s $200B bull case, and why robotics data may become a major revenue line.

    • SF talent war is intense; being on a “rocket ship” helps but raises hiring bar
    • Hardest trait to assess: agency and ownership
    • Hiring shifts: less tool-centric, more judgment/experimentation/systems thinking
    • Keep AI for execution, not decision-making; protect the “thinking muscle”
    • Quick-fire: RL environments surprise scale, internship advice, robotics data growth thesis

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.