No PriorsNo Priors Ep. 124 | With SurgeAI Founder and CEO Edwin Chen
CHAPTERS
- 0:00 – 1:14
Edwin Chen and SurgeAI in numbers: bootstrapped, small team, massive revenue
Sarah introduces Edwin Chen and frames SurgeAI as a high-quality human data company serving top frontier labs. Edwin shares key headline metrics (over $1B revenue, ~100+ employees) and the founding thesis: quality-first human data to unblock ML progress.
- •SurgeAI surpassed $1B in annual revenue with a ~100-person team
- •Founded around the belief that human data quality is a primary bottleneck for AI
- •Serves major customers (e.g., Google, OpenAI, Anthropic)
- •Positioning: “human data startup” focused on training and evaluation inputs
- 1:14 – 2:27
Origin story: why Edwin started Surge after Google/Facebook/Twitter
Edwin explains that repeated data acquisition and quality issues at large tech companies motivated Surge’s creation. The core pain was that even basic ML projects stalled due to inability to get the right labeled/curated data, making future AI ambitions feel impossible without fixing the pipeline.
- •Surge started in 2020; five years in operation
- •Prior roles at Google, Facebook, Twitter shaped the problem understanding
- •Data availability/quality was the recurring blocker to shipping ML systems
- •Motivation extended beyond current tasks to enabling “next-generation” AI systems
- 2:27 – 4:53
Why Surge bootstrapped: control, profitability, and skepticism of fundraising culture
Edwin argues that raising is often done for signaling rather than necessity. Since Surge was profitable from the start, he preferred retaining control and focusing on building rather than fundraising-as-goal behavior.
- •Didn’t raise because the business didn’t need capital to survive or grow early
- •Belief that fundraising often becomes performative (status/press)
- •First-principles view: build first; raise only when money is the real bottleneck
- •Early profitability enabled independence
- 4:53 – 7:58
Hiring and validation without a big name: focus on essential early roles
Sarah and Elad probe the “external validation” argument for fundraising—especially for hiring. Edwin distinguishes between founders who truly need capital to live and those who don’t, then critiques premature hiring of roles like PMs and data scientists before product-market clarity.
- •Different fundraising needs depending on founder’s financial runway and background
- •Challenge: unknown founders may use funding as credibility for recruiting
- •Edwin’s view: early teams should prioritize builders; avoid premature PM/data science hires
- •Early-stage goal is 10x–100x product discovery, not incremental optimization
- 7:58 – 9:40
What SurgeAI actually delivers: training data plus evaluation insights
Edwin clarifies Surge’s “product” as the data itself, delivered in multiple forms depending on what capability a lab wants to improve. Beyond raw datasets, Surge also provides evaluation, failure modes, and diagnostic insights that inform model iteration.
- •Surge delivers data used to train and evaluate models
- •Examples: SFT solutions, unit tests, preference comparisons, verifiers/checkers
- •Supports model evaluation: identifying weaknesses, comparative performance
- •Adds value by surfacing patterns and failure modes, not just labeling
- 9:40 – 11:27
Differentiation vs. body shops: building a quality-measurement platform
Edwin contrasts Surge with providers that mainly supply “warm bodies.” He argues the ceiling on quality is far higher for generative tasks than for commodity labeling (e.g., bounding boxes), and that meeting that ceiling requires technology to measure and manage quality at scale.
- •Many competitors act as staffing vendors rather than data-quality systems
- •GenAI data has an “unlimited ceiling” (poetry, writing, reasoning) unlike boxes around cars
- •Surge uses technology to measure annotator output quality
- •Quality measurement is a core moat, not just worker supply
- 11:27 – 12:25
How Surge measures output quality: multi-signal scoring and internal ML
Elad asks how Surge creates a feedback loop to evaluate evaluators. Edwin describes an approach analogous to search ranking: aggregate many behavioral and task signals and feed them into ML systems that estimate quality and reliability.
- •Quality evaluation uses many signals, similar to ranking web pages/videos
- •Signals include annotator behavior/activity and work characteristics
- •An internal ML team builds scoring/quality estimation algorithms
- •Goal: scalable quality control rather than purely manual review
- 12:25 – 14:02
Scalable oversight: humans + models collaborating to produce better data
Sarah asks what changes as model baselines rise and exceed average human performance. Edwin frames this as “scalable oversight”: designing interfaces and workflows where models draft and humans refine, so the combined output exceeds either alone while reducing wasted human effort.
- •Scalable oversight: combine AI generation with human judgment/editing
- •Humans increasingly edit/shape model drafts instead of writing from scratch
- •Focus human effort on creativity/taste and avoid low-value “cruft” work
- •Requires tooling and workflow design, not just more annotators
- 14:02 – 16:08
Building rich RL environments: why realistic simulated worlds are hard
The conversation shifts to RL environments and reward modeling. Edwin explains that labs want complex, tool-rich simulations (e.g., a salesperson’s full toolchain and interruptions), which require large volumes of consistent, time-evolving content and careful design that can’t be naively autogenerated.
- •Demand shifting toward RL environments and long-horizon agent training
- •Example environment: salesperson workflows across Gmail/Slack/CRM/spreadsheets/docs
- •Need consistent simulated artifacts: emails, messages, documents, timelines, events
- •Hard part: realism, coherence, and creative diversity across many interacting tools
- 16:08 – 17:30
No ceiling on environment quality—and what data demand looks like long-term
Sarah asks how complex is “enough” and what will matter over 5–10 years. Edwin argues there’s effectively no ceiling on environment richness, and future training will likely blend RL environments, expert traces, and multi-reward signals to capture long, complex objectives.
- •Claim: environment realism/diversity has no practical upper bound
- •Long horizons and diverse situations teach more robust capabilities
- •Future demand spans multiple data types, not just RL environments
- •Single scalar rewards often insufficient; multi-reward/combined approaches needed
- 17:30 – 19:22
Humans in a (eventually) superhuman world: limits of synthetic data and misalignment risks
Elad challenges whether humans become irrelevant once models outperform experts. Edwin contends human feedback remains essential: synthetic data is valuable but mostly low-signal without curation, and models can optimize the wrong objectives without external alignment signals.
- •Synthetic data helps, but large volumes can be mostly unusable without curation
- •Small amounts of high-quality human data can outweigh millions of synthetic samples
- •Humans provide an external objective/grounding signal to prevent drift
- •Examples of model weirdness and objective mismatch (e.g., odd language output)
- 19:22 – 21:27
Benchmark hacking and ‘vibes-based’ leaderboards: why LMSYS/LM Arena can mislead
Edwin criticizes leaderboard dynamics where quick judgments reward formatting, verbosity, and style over truthfulness and instruction-following. He warns that teams sometimes knowingly degrade real quality to climb rankings, leading to months of “progress” that is actually clickbait optimization.
- •Leaderboards can reward superficial traits (length, formatting) over correctness
- •Mis-specified objectives drive models toward undesirable behavior
- •Organizational incentives can push researchers to chase rankings at quality’s expense
- •Parallel critique of academic benchmarks divorced from real-world utility
- 21:27 – 23:37
The alternative: rigorous human evaluation and Surge’s push toward standardization
Edwin describes careful human eval as the gold standard: fact-checking, instruction adherence, and taste-based assessment. He notes Surge already performs extensive internal evaluations for frontier labs and wants to expand to more public-facing work educating the market on model capabilities and tradeoffs.
- •Proper human eval involves time, verification, and nuanced judgment
- •Avoids training toward ‘clickbait’ behaviors induced by shallow comparisons
- •Surge provides internal eval insights: failure modes, capability mapping
- •Ambition to externalize/standardize evals for broader transparency
- 23:37 – 29:29
Industry landscape: Meta–Scale implications, underdog bets, and a multi-model future
Sarah asks about the Meta/Scale deal and competition among frontier models. Edwin argues higher-quality data adoption benefits the industry, picks xAI as an underdog, and predicts a growing number of differentiated frontier models rather than commoditization—each with distinct strengths and ‘personalities.’
- •Meta–Scale deal: could accelerate market shift away from low-quality data experiences
- •Belief that better human data increases overall industry progress
- •Underdog pick: xAI due to intensity/mission orientation
- •Prediction: more (not fewer) frontier models with differentiated strengths and focus areas
- 29:29 – 32:58
What ‘high-quality data’ really means: beyond checklists to creativity and depth
Edwin closes by redefining quality as capturing rich human intelligence, not just compliance with simple constraints. Through the poetry example, he argues that great data reflects diverse valid approaches and subjective excellence—requiring domain-aware quality lenses and technology to avoid scaling mediocrity.
- •Checklist compliance (e.g., “8 lines, contains ‘moon’”) produces low-value data
- •Credentials don’t guarantee output quality (PhDs vs true craft)
- •High-quality data embraces diversity of styles/solutions (poems, proofs, reasoning)
- •Surge builds principles + domain-specific lenses to identify excellence at scale