Lenny's PodcastEdwin Chen: Why optimizing for benchmarks creates AI sloth
How Surge bootstrapped past $1B revenue with fewer than 100 people; Chen argues benchmark gaming pushes AI toward dopamine, emojis, and slop, not truth.
CHAPTERS
- 0:00 – 7:08
Surge AI’s unusual scale: $1B revenue with <100 people, bootstrapped
Lenny opens by highlighting Surge AI’s unprecedented growth: reaching $1B in revenue in under four years with a tiny team and no VC funding. Edwin explains the core belief behind this approach—small, elite teams move faster, and AI-driven efficiency will push revenue-per-employee even higher across the industry.
- •Surge hits $1B revenue in <4 years with ~60–100 people
- •Bootstrapped and profitable from day one—no VC fundraising
- •Edwin’s belief that big tech is overstaffed and slowed by distractions
- •AI will accelerate leverage; predicts extreme revenue-per-employee ratios
- •Small teams enable different kinds of companies and products
- 7:08 – 8:55
A contrarian company-building philosophy: avoiding the Silicon Valley PR/fundraising loop
Edwin describes why Surge intentionally avoided the typical Silicon Valley playbook—no hype-driven social media presence and no fundraising cycle. Instead, the company relied on building a meaningfully better product and earning researcher-driven word-of-mouth, which also ensured customer alignment around quality.
- •Rejecting the ‘Silicon Valley game’ of fundraising and constant PR
- •Choosing word-of-mouth from technical users over headlines
- •Customer alignment: early buyers deeply cared about data quality
- •Tradeoff: less visibility, but stronger product/mission focus
- •Empowering model: founders can succeed without constant posting
- 8:55 – 9:36
What Surge AI actually does: teaching models ‘good vs. bad’ and measuring progress
Edwin gives a functional overview of Surge: providing the human data and tooling that trains and improves frontier models, plus the evaluation systems to track progress. He frames Surge as both a training and measurement layer for post-training—helping models learn behaviors and capabilities.
- •Surge provides human data to train post-training behaviors
- •Core work: teach models what’s ‘good’ and ‘bad’
- •Tools referenced: SFT, RLHF, rubrics, verifiers, environments
- •Also builds measurement/evaluation to track model progress
- •Positioned as a data + evals company supporting frontier labs
- 9:36 – 11:37
The real meaning of “high-quality data” (and why ‘throwing bodies’ fails)
Edwin argues most people misunderstand quality in data work, treating it like a checklist rather than a deep, nuanced standard. Using a poem example, he explains Surge’s bar is closer to ‘Nobel-level’ outputs, requiring richer signals and expert judgment to capture subtlety, originality, and truthfulness.
- •Quality isn’t box-checking; it’s subjective, rich, and difficult
- •Poetry example: beyond format compliance to emotional/creative depth
- •High-quality outputs require sophisticated definitions of ‘good’
- •Quality measurement demands more than scaling headcount
- •The goal: data that actually improves frontier model behavior
- 11:37 – 13:29
How Surge measures and selects quality: thousands of signals, ML-style ranking, and expert fit
Edwin explains the operational mechanics behind quality: continuous instrumentation of worker performance and task outputs, plus review systems and model-based checks. He compares Surge’s approach to Google Search—remove the worst spam while also discovering the best—and treats it as an ML ranking problem over people and outputs.
- •Collecting thousands of signals on workers, tasks, and projects
- •Using reviews, gold standards, and model-based evaluation
- •Tracking performance by domain (poetry vs essays vs technical docs)
- •Two goals: filter out worst work and identify the best contributors
- •Framed as a search/ranking problem similar to Google Search
- 13:29 – 17:38
Why Claude stayed ahead: data choices, post-training ‘taste,’ and objective functions
The conversation turns to Anthropic/Claude’s sustained lead in coding and writing. Edwin attributes this to a complex set of choices—data composition, synthetic vs. human mixes, what gets optimized, and an often-overlooked factor: the ‘taste’ and sophistication guiding post-training decisions.
- •Model performance reflects countless data and post-training choices
- •Tradeoffs: front-end vs back-end focus, design vs correctness
- •Synthetic data mix and benchmark priorities shape outcomes
- •Post-training is partly art: taste and sophistication matter
- •Objective function design determines what a model becomes good at
- 17:38 – 20:09
Benchmark skepticism: why leaderboards can be wrong, gameable, and misaligned with reality
Edwin explains why he distrusts common benchmarks: some contain errors, many are overly objective and easy to hill-climb, and they diverge from messy real-world tasks. He also describes how labs can tweak prompts and evaluation setups to climb scores without true capability gains.
- •Benchmarks can contain wrong answers and messy flaws
- •Objective tasks are easier to optimize than real-world ambiguity
- •Models can be tuned to benchmarks without improving usefulness
- •Gaming via system prompts, reruns, and evaluation tricks
- •Explains mismatch: high math scores vs. basic workflow failures (e.g., PDFs)
- 20:09 – 22:16
Measuring real progress toward AGI: expert human evals over shallow ‘vibes’
In place of benchmarks, Edwin advocates rigorous human evaluations conducted by domain experts who deeply verify outputs. These annotators test models in realistic conversations and workflows, checking correctness and nuance rather than selecting whatever looks flashiest.
- •Progress measurement through expert-led human evaluations
- •Experts validate outputs (code runs, equations checked)
- •Multi-dimensional scoring: accuracy, instruction-following, usefulness
- •Critique of casual A/B voting as superficial ‘vibe’ judging
- •Humans remain necessary until true AGI
- 22:16 – 23:03
AGI timelines and diminishing returns: 80% automation soon, but reliability takes longer
Edwin places himself in the longer-horizon camp for AGI, emphasizing the difficulty of moving from good to nearly-perfect performance. He expects rapid automation of large portions of knowledge work soon, but believes the final reliability and coverage will take many additional years.
- •Big gap between 80%, 90%, 99%, and 99.9% capability
- •Predicts 80% of an L6 engineer’s job could be automated in 1–2 years
- •Reliability and tail cases create long timelines
- •Leans toward a decade+ rather than ‘imminent AGI’
- •Human data remains critical as long as models aren’t fully general
- 23:03 – 34:32
‘AI slop’ risk: incentives, LLM Arena, engagement optimization, and the wrong direction
Edwin warns that industry incentives can steer models toward attention-grabbing but untruthful behavior—optimizing for dopamine, not truth. He critiques LLM Arena-style voting and engagement goals as creating perverse incentives: longer, flashier, emoji-filled outputs that may hallucinate more but score higher.
- •Models can be optimized for engagement rather than truth
- •LLM Arena voters skim; flashiness beats accuracy
- •Easy leaderboard gains: longer answers, more emojis, superficial formatting
- •PR and enterprise sales pressure labs to chase leaderboard rank
- •Engagement optimization risks social-media-like failure modes (clickbait, delusion reinforcement)
- 34:32 – 39:37
Reinforcement learning environments: simulated worlds for end-to-end agent training
Edwin introduces reinforcement learning as training models to achieve rewards in rich simulated environments that mimic real work. These environments expose how models fail on long-horizon, tool-heavy, messy tasks (Slack, Jira, codebases) and provide a path to building more capable agents through iterative reward-driven learning.
- •RL trains models to maximize rewards tied to task success
- •RL environments simulate realistic worlds (tools, threads, codebases)
- •Stress-tests end-to-end behavior over long horizons
- •Highlights catastrophic failures when tasks get messy and multi-step
- •Seen as the next major phase complementing SFT/RLHF/rubrics
- 39:37 – 41:09
Why trajectories matter: evaluating the path, not just the final answer
Edwin explains that models may reach correct outcomes via inefficient, random, or reward-hacking behavior. Capturing full trajectories reveals reasoning quality, efficiency, reflection, and robustness—critical for training agents that succeed reliably rather than stumbling into success.
- •Correct final answers can hide dysfunctional intermediate behavior
- •Models may ‘try 50 times’ or stumble into success randomly
- •Trajectory data enables teaching reflection vs. brute force
- •Long-horizon tasks require understanding step-by-step behavior
- •Avoids reinforcing reward-hacking and inefficiency
- 41:09 – 44:34
The evolution of post-training: SFT → RLHF → rubrics/verifiers → RL environments
Edwin lays out a practical history of how post-training has progressed, using human-learning analogies. He distinguishes evals used for training from evals used for selecting releases, and positions RL environments as the newest complementary learning mechanism rather than a replacement for prior methods.
- •SFT (supervised fine-tuning): imitate a ‘master’
- •RLHF: preference-based learning akin to picking best essays
- •Rubrics/verifiers: graded feedback with detailed error signals
- •Evals serve both training and checkpoint selection
- •RL environments emerging as the latest complementary method
- 44:34 – 48:07
Surge’s research DNA and the company origin story: MIT → big tech pain → GPT-3 → founding
Edwin describes Surge as research-driven: forward-deployed researchers partner with labs, while internal researchers improve benchmarks and data quality systems. He then shares his path—from interests in math/language at MIT, to data bottlenecks at Google/Facebook/Twitter, to founding Surge shortly after GPT-3 to solve complex data needs for frontier AI.
- •Two research modes: forward-deployed with customers and internal R&D
- •Focus areas: better benchmarks/leaderboards and internal quality measurement
- •Edwin’s background: math, CS, linguistics; fascination with language
- •Repeated big-tech problem: inability to get the right training data
- •GPT-3 catalyzed founding Surge to build advanced, complex data pipelines
- 48:07 – 1:10:31
What’s next in AI (and a lightning round): model differentiation, mini-app chat UIs, and ‘vibe coding’ risks
Edwin predicts models will diverge more as labs embed different values and objective functions into behavior—affecting productivity, personality, and alignment. He highlights underhyped “mini app” interfaces inside chatbots, calls “vibe coding” overhyped due to maintainability risks, then closes with rapid-fire personal recommendations and founder advice.
- •Models will differentiate by values, behaviors, and objective functions
- •Example: assistants that optimize engagement vs. user time/productivity
- •Underhyped: built-in chatbot artifacts evolving into mini-app UIs
- •Overhyped: vibe coding leading to long-term unmaintainable systems
- •Lightning round: books/TV, Waymo, life motto, and founder guidance