The Twenty Minute VCMercor Head of Product on Revenue Concentration from Frontier Labs
CHAPTERS
- 0:00 – 1:14
Mercor’s demand surge and why data stays valuable at the frontier
Osvald and Harry set the context: Mercor is seeing overwhelming demand and strong cash flow, and the discussion will focus on how data businesses evolve as model capabilities improve. The framing is that data retains value where models still have gaps—especially at the frontier of performance.
- •Mercor’s business health: demand outstrips ability to spend to meet it
- •Core question: what happens to data providers as models get better
- •Data value concentrates at the performance frontier, not in commoditized tasks
- 1:14 – 2:09
Do open-source models cannibalize closed-model data demand?
Harry challenges the common narrative that open-source will replace frontier closed models, potentially shrinking Mercor’s market. Osvald argues open models raise the baseline and reduce demand only for tasks they already solve, while frontier gaps continue creating new demand for eval and training data.
- •Open-source raises the floor; frontier gaps still drive purchases
- •Customers buy data to fill capability gaps tied to specific goals
- •Cannibalization happens only for tasks already solved by a given open model
- 2:09 – 3:23
Why the “90% of enterprise workflows” framing is misleading
Osvald disputes the claim that most enterprise workflows are already solvable by current models, arguing it ignores latent demand—especially long-horizon automation no one is attempting yet. He also reframes “can models do it?” into two categories: sufficiency-based tasks vs domains with uncapped improvement potential.
- •Latent demand: long-horizon agents (e.g., months-long procurement automation)
- •Apex benchmarks suggest ~50% on long-horizon workflows for top models
- •Sufficiency tasks (e.g., CRM updates) vs uncapped domains (law/medicine)
- •Binary % framing breaks for tasks with continuous rewards
- 3:23 – 6:12
Enterprise data sensitivity: which workflows stay closed vs go open
The conversation turns to enterprise skepticism about sharing sensitive data with frontier providers. Osvald explains sensitivity depends on how core and differentiating the workflow is; generic functions (HR/procurement) are easier to externalize, while core legal or advisory work triggers more caution.
- •Sensitivity varies by how core the workflow is to competitive advantage
- •Generic functions are less sensitive; differentiating work is more protected
- •Open weights can be safer depending on where inference runs (control over deployment)
- 6:12 – 7:29
Specialized models per company and the resulting data flywheel
Harry asks whether every company will end up with specialized models tailored to its priorities. Osvald agrees and notes this future is also “self-serving” for Mercor because specialization requires enterprise-specific evals and training data to define objectives and measure success.
- •Company-specific goals drive specialization (growth vs margin vs constraints)
- •Specialized models require proprietary eval + training data
- •ROI-gated adoption: specialization increases as value justification improves
- 7:29 – 8:31
Is there an enterprise AI ROI problem—or just a waiting game?
Osvald argues ROI pressure is not the dominant dynamic yet; enterprises are still exploring while token prices and model performance shift quickly. He acknowledges early tightening in some areas but says most organizations are still in a “let’s see what happens” phase.
- •Current phase: experimentation with higher tolerance for uncertainty
- •ROI math is unstable due to fast-moving token economics and performance
- •Some budget tightening is emerging, but not a broad ROI crisis
- 8:31 – 10:08
Balancing token spend vs performance: use-case-driven budgeting
Osvald gives a pragmatic framework: spend depends on whether tokens are COGS-like (e.g., customer service) or compounding productivity (e.g., coding agents for engineers). Founders should align spend with unit economics and growth impact rather than applying blanket caps.
- •Token spend should map to the economics of the specific workflow
- •Coding agent spend can be leverage, not pure COGS
- •Customer service agents must clear unit economics vs revenue served
- •Hypergrowth contexts tolerate more spend to meet demand
- 10:08 – 12:14
Salesforce’s $300M Anthropic spend and where AI budgets are headed
Harry brings up Salesforce’s reported $300M annual spend on Anthropic and asks if that’s the new normal. Osvald calls for better internal accounting by team type (R&D vs forward-deployed) and predicts the overall percentage of spend allocated to tokens will rise over time—potentially dramatically for hypergrowth companies like Mercor.
- •Different org functions warrant different spend profiles and tolerance
- •Need outcome-based accounting for token spend effectiveness
- •Macro trend: token spend share likely increases beyond today’s low single digits
- •Mercor’s reality: spending can already exceed salaries due to demand pressure
- 12:14 – 14:43
AI makes product management harder: controlling surface area and boosting judgment
Osvald argues faster engineering doesn’t mean shipping more features; it increases the risk of chaotic product sprawl. The PM bottleneck becomes choosing the right work, simplifying workflows, and maintaining strong judgment—leading to a higher PM-to-engineer ratio over time.
- •Core battle: reduce product surface area despite faster build velocity
- •Engineering is less bottlenecked; identifying revenue-driving workflows is harder
- •PM excellence shifts from tool mastery to judgment and business impact
- •Prediction: higher PM-to-engineer ratio as coding agents accelerate velocity
- 14:43 – 18:33
Biggest product mistake: over-flexible tooling and the need for guardrails
Osvald describes a key misstep: Mercor’s annotation platform became too flexible by trying to support every customer request across rapidly changing data formats. The result was operational chaos; the lesson was to introduce guardrails, align with ops, and invest only where enduring demand exists.
- •Annotation workflows diversified (SFT → preferences → rubric → environments, multimodal)
- •Maximal flexibility created complexity across hundreds of concurrent projects
- •Guardrails and best-practice standardization should have come earlier
- •Enduring-demand decisions rely on market proximity and leadership judgment
- 18:33 – 28:18
Why services are rising in AI deployment (and why that may be temporary)
The discussion shifts to the growth of services arms (e.g., Microsoft) and whether services indicate weak products. Osvald’s take: services fill a knowledge dissemination gap—AI deployment expertise is concentrated (notably in SF) and not yet broadly available—though over the long run it should become a standard internal capability.
- •Services are a short-term bridge while AI expertise spreads through industry
- •Talent scarcity: not enough in-house capability at every enterprise yet
- •Forward-deployed roles fit engineers strong in communication and simplification
- •Long-term arc: deployment becomes a common job function, less service-heavy
- 28:18 – 33:43
How Mercor structures product teams: Marketplace vs Studio, pods, and cadence
Osvald explains Mercor’s two core product areas—Marketplace (matching experts to jobs) and Studio (annotation/evals platform)—and how they support talent-only vs managed-service engagement models. He details pod-based sprint planning, weekly product-area meetings, and the scaling challenge of cross-team communication as headcount grows.
- •Two product teams: Marketplace and Studio (annotation/evals)
- •Two engagement modes: talent-only vs managed service (end-to-end datasets)
- •Small pods (PMs + eng + DS + flexible design) with sprint planning per pod
- •Weekly product-area sync for alignment; cross-product communication is the scaling pain
- 33:43 – 38:05
Scaling marketplace supply: expert experience, referrals, and global sourcing
Harry asks how Mercor scaled expert supply efficiently. Osvald attributes it to a dignified, reliable expert experience (pay, transparency, communication), which fuels referrals, plus a sourcing team that can find niche skills globally to meet spiky demand.
- •Reliable pay and strong treatment create a referral flywheel
- •Retention comes from sustained high-quality experience, not gimmicky bonuses
- •Clear instructions and communication reduce friction in unfamiliar online work
- •Global sourcing fills specialized skill gaps during demand spikes
- 38:05 – 39:27
Revenue concentration risk and the push downmarket via self-serve human data projects
Harry raises concerns about Mercor’s revenue concentration with frontier labs. Osvald frames Mercor’s biggest product challenge as moving downmarket: making human data projects self-serve and efficient for enterprises (e.g., AI project managers), enabling more customers, smaller heterogeneous projects, and reduced concentration.
- •Current lab projects are white-glove and ops-heavy
- •Strategic goal: democratize human data projects for enterprises
- •Product direction: self-serve workflows + AI project management assistance
- •Downmarket expansion diversifies revenue away from a few labs
- 39:27 – 41:11
Why enterprise data projects are operationally intense: edge cases, paranoia, and rapid data evolution
Osvald explains what makes human data projects hard to “simplify”: translating complex customer intent into precise guidelines, handling endless edge cases, and maintaining rigorous QC under time pressure. Complexity also comes from fast-changing data types across projects and modalities, demanding constant process adaptation.
- •Edge cases dominate and require rapid alignment among stakeholders
- •Ops must be highly paranoid: every datapoint must match evolving specs
- •Crisp communication is critical across customers, ops, and experts
- •Data types evolve quickly, increasing between-project complexity
- 41:11 – 51:59
Fastest-growing data type: RL environments—and why cybersecurity never reaches “sufficiency”
Osvald highlights RL environments as the fastest-growing data type: high-fidelity simulated apps/world states that resemble deployment conditions for agent training and eval. He also connects this to cybersecurity, where adversarial dynamics mean the goalposts always move—so performance is never “done,” sustaining continuous demand for evolving data.
- •RL environments: simulated tools/apps + rich start states + task suites
- •Training/evals increasingly mirror real deployment contexts (high-fidelity mocks)
- •Cybersecurity is adversarial with uncapped rewards; goalposts continually shift
- •Implication: security won’t hit a stable “90% sufficiency” ceiling
- 51:59 – 1:02:42
Hiring in SF and Mercor’s interview focus: agency, experiments, and AI boundaries (plus quick-fire)
The closing stretch covers SF’s brutal talent market and what Mercor prioritizes: agency/ownership, strong experimental thinking, stats literacy, and systems design—while avoiding over-delegating judgment to models. In quick-fire, Osvald discusses changing his mind on RL environments, advice to students (get internships), Mercor’s $200B bull case, and why robotics data may become a major revenue line.
- •SF talent war is intense; being on a “rocket ship” helps but raises hiring bar
- •Hardest trait to assess: agency and ownership
- •Hiring shifts: less tool-centric, more judgment/experimentation/systems thinking
- •Keep AI for execution, not decision-making; protect the “thinking muscle”
- •Quick-fire: RL environments surprise scale, internship advice, robotics data growth thesis