Y CombinatorTom Brown: How Building GPT-3 Led to Founding Anthropic
Through GPT-3 scaling laws and then Claude Code architecture; Tom Brown traces the path from a B-minus in linear algebra to building frontier AI infrastructure.
CHAPTERS
- 0:00 – 1:12
From early startup uncertainty to building world-scale AI infrastructure
The conversation opens on how unlikely Anthropic’s early success looked compared to OpenAI’s funding and star power. It also foreshadows the coming “largest infrastructure build-out of all time” driven by AI compute demand.
- •Anthropic began as a small COVID-era team with unclear product direction
- •OpenAI’s scale and funding made success feel unlikely
- •AI compute spend is accelerating rapidly, implying historic infrastructure expansion
- •Sets up the episode’s throughline: product + infrastructure + long-term mission
- 1:12 – 2:30
Learning the “wolf” mindset at Linked Language (first employee at a YC startup)
Tom describes his first formative startup experience at Linked Language in 2009, contrasting it with big-tech career paths. The key takeaway is the mindset shift from executing assigned tasks to proactively owning survival-critical problems.
- •Joined friends’ startup straight out of MIT as first employee
- •Startups force self-direction: “company dies by default”
- •Big tech trains you for big tech; startups cultivate autonomy
- •The “wolves hunt” metaphor as a lasting career lesson
- 2:30 – 2:50
MoPub and the first engineering scaling lessons (while still leveling up as a programmer)
After returning to school, Tom joined MoPub as an early engineer to experience scaling systems firsthand. He frames this period as skill-building—wanting the startup ethos while still improving his engineering fundamentals.
- •Early engineer at MoPub (mobile advertising)
- •Learned what it takes to scale a product/system
- •Acknowledges struggling early as a programmer
- •Uses the role to grow into higher-leverage engineering work
- 2:50 – 4:06
Solid Stage in YC: an ambitious DevOps idea before Docker—and why it didn’t click
Tom recounts founding Solid Stage (a pre-Docker DevOps platform) and the difficulty of articulating what they were truly building. The experience highlights a common failure mode: novelty without crisp product clarity or durable mission fit.
- •Attempted “more flexible Heroku” before Docker existed
- •YC interview feedback centered on unclear product definition
- •Team itself lacked full clarity on the product path
- •Tom left mid-batch due to weak mission/product conviction
- 4:06 – 7:41
Grouper dating experiment: product thesis, culture, and why Tinder beat the mission
Tom joins Grouper, a group-dating product with a manually curated matching experience. He explains the personal motivation (making socializing safer for awkward people) and the market shift that undermined it when Tinder solved the core fear of rejection more cleanly.
- •Grouper’s format: group meetups + manual matching
- •Personal mission: help awkward people socialize safely
- •Recruiting and culture mattered; lots of smart people involved
- •Tinder’s mutual-interest swipe solved rejection anxiety better
- •Company momentum eventually flatlined/declined
- 7:41 – 9:16
Burnout and reset: rebuilding runway before retooling into AI
After Grouper, Tom describes burnout from pitching a dream he no longer believed in. He takes time off, then faces a practical constraint—money—leading him to rebuild runway and plan a serious transition into AI.
- •Emotional toll of startups during downturns
- •Took a multi-month break (art car, exercise, recovery)
- •Ran out of personal runway, forcing a pragmatic plan
- •Decided AI could be the most important work of a lifetime
- 9:16 – 11:10
Self-teaching ML in 2015: a concrete retooling playbook (and its limits)
Tom outlines a six-month self-study regimen designed to make him useful to top AI labs. He emphasizes that the exact plan is dated, but the structure—runway + curriculum + hands-on practice + compute access—was key.
- •Target labs at the time: DeepMind, Google Brain, MIRI
- •Runway first: short contract work (Twitch) to fund study
- •Curriculum: Coursera ML, Kaggle, linear algebra + stats texts
- •Hands-on GPU work via SSH; focus then was image classification post-AlexNet
- •Acknowledges the path would look different for people today
- 11:10 – 12:34
Making the leap to OpenAI: leveraging distributed systems + ML crossover
Tom gets in through his relationship with Greg Brockman and by offering a rare skill mix: distributed systems plus growing ML knowledge. His first work isn’t model training—it’s building environments (StarCraft), a foot-in-the-door step that later compounds.
- •Reached out to Greg immediately after OpenAI announcement
- •Value proposition: “paucity” of people with ML + distributed systems
- •Introductions/mentorship helped structure a learning path
- •Initial project: building a StarCraft environment, not ML research
- •OpenAI’s early feel: post-apartment, chocolate factory office, major committed funding
- 12:34 – 16:04
Scaling laws change the game: building GPT-3 infrastructure and betting on compute
Tom explains how scaling laws convinced him that capability would reliably increase with compute and the right recipe. He connects this to the engineering pivot from TPUs toward GPU training and to the cultural critique that scaling looked like ‘brute force’—yet it worked.
- •Timeline: OpenAI → Google Brain → back to OpenAI; GPT-3 buildup (2018–2019)
- •Scaling laws as a straight-line predictor over massive ranges
- •Compute scaling + algorithmic efficiency as a compounding trend
- •Controversy: “wasting money on GPUs” vs. pragmatic results
- •Framing: ‘do the stupid thing that works’ as an effective strategy
- 16:04 – 18:42
Why Anthropic spun out: combining safety urgency with scaling execution
Tom describes two groups under Dario/Daniela—safety and scaling—and how their tight collaboration and communication culture set the stage for Anthropic. The spinout is framed as building an institution capable of handling high-stakes transitions to transformative AI.
- •Safety org + scaling org worked closely with strong internal comms (public Slack culture)
- •Belief in an eventual ‘handoff’ to transformative AI raised urgency
- •Mission focus motivated the founding group to leave despite OpenAI’s advantages
- •Early Anthropic looked fragile: 7 cofounders in COVID, unclear product plan
- •Mission-first early hires helped preserve culture while scaling to ~2000 people
- 18:42 – 19:22
Early Anthropic execution: securing compute, training stack, and delaying serving readiness
The first year is described as intensely operational: building training infrastructure, getting compute, and doing basic company setup. Tom notes a key lesson—hesitating on product/serving infra because of uncertainty about whether launching was “good for the world.”
- •Primary early work: training infrastructure + compute procurement
- •Fast ramp: ~25 OpenAI alumni joined early, improving coordination
- •Pre-ChatGPT Claude existed as a Slack bot but wasn’t productized
- •Uncertainty about ‘theory of impact’ slowed serving infrastructure investment
- •Lesson: infra readiness matters even when product strategy feels unsettled
- 19:22 – 23:51
ChatGPT wake-up call to Claude 3.5 Sonnet: coding becomes the growth wedge
After ChatGPT, Anthropic relaunches its API and Claude product, but Tom says success didn’t feel assured until Claude 3.5 and coding. The chapter covers intentional investment in coding, rapid rollout, and surprise at how strongly the market responded—especially among startups.
- •ChatGPT catalyzed a faster product push (API + Claude.ai)
- •2023: default startup choice was OpenAI; 2024 shifted toward Sonnet for coding
- •Coding strength was partly intentional, then doubled down after market feedback
- •3.5 Sonnet as a turning point; 3.7 surprises by unlocking more agentic coding
- •Anecdote: Claude helps decompile/disassemble to usable C code
- 23:51 – 26:09
Why benchmarks mislead: avoiding ‘teaching to the test’ and focusing on dogfooding
Tom argues that public benchmarks can be gamed and don’t fully explain why developers prefer Claude for real coding work. Anthropic prioritizes internal benchmarks and extensive internal usage to improve practical capability, plus qualitative work on personality and interpretability.
- •Public benchmarks are easy to game; some labs optimize scores directly
- •Anthropic focuses on internal benchmarks (not published)
- •Heavy dogfooding: accelerating Anthropic’s own engineers is a top priority
- •Personality evals are complex; goal is broadly positive, culturally aware interactions
- •Interpretability is framed as a long-term safety bet for more powerful models
- 26:09 – 31:01
Claude Code’s secret sauce: an internal tool that treated Claude as a user (agent-first design)
Claude Code began as an internal hack to help Anthropic engineers, then unexpectedly became a standout product. Tom attributes part of the edge to a mindset shift: designing tools not only for developers but also for the model itself—giving it the right context and interfaces to act effectively.
- •Origin: internal developer tool built by Boris for Anthropic engineers
- •Surprise: became a market-leading agentic coding product despite API-first strategy
- •Key idea: treat Claude as a stakeholder/user; optimize tools for model effectiveness
- •This philosophy aligns with tool-calling standards (e.g., MCP) taking off
- •Advice to builders: there’s a rich vein in building “tools for models as users”
- 31:01 – 34:38
Compute arms race and multi-chip strategy: power bottlenecks, platforms, and software stack leverage
Tom zooms out to the macro trend: AI compute spend is tripling yearly and will surpass historic national-scale projects. He details infrastructure bottlenecks (especially power), Anthropic’s policy preference for US buildout, and the engineering trade-offs of running GPUs, TPUs, and Trainium with a strong software layer for rapid iteration.
- •AI compute spending trajectory: ~3× per year; unprecedented infrastructure build-out
- •Primary bottleneck: power (especially in the US), plus permitting/data center capacity
- •Solution set includes renewables + nuclear; desire for easier nuclear deployment
- •Multi-chip approach: GPUs, TPUs, Trainium—more flexibility but splits performance engineering
- •Core lesson from GPT-3 era: great software stacks enable faster iteration across platforms
- 34:38 – 35:56
Advice for the next generation: take risks and optimize for intrinsic ambition
In closing, Tom advises younger builders to take more risks and choose projects that their idealized selves would be proud of. The hosts reinforce the idea that traditional extrinsic credentials matter less in the current AI era than pursuing high-agency, meaningful work.
- •Take more risks earlier
- •Pick work that feels intrinsically meaningful and impressive to your own standards
- •De-emphasize credential-chasing as the world shifts quickly
- •Framing: pursue missions that would make ‘future you’ proud