Skip to content
Uncapped with Jack AltmanUncapped with Jack Altman

Andrew Feldman on Building Cerebras and the Future of Chips | Ep. 57

Andrew Feldman is the co-founder and CEO of Cerebras Systems, the AI chip company he founded in 2016 around a single radical insight: that winning in compute requires not incremental improvement but a fundamentally different architecture. Cerebras is the creator of the world's largest chip, the Wafer Scale Engine, and counts the US government, sovereign cloud providers, and OpenAI among its customers. Alongside Eric Vishria from Benchmark, we discussed why Andrew believes that if you are going to attack Goliath, being 10% or even twice as good is not an available strategy and you have to aim for 100x or 500x better. Andrew walked through Cerebras's near-death experience: 18 months of board meetings where the only thing to report was "still can't make it," spending $8 million a month, and what kept the team going. He explained how we go from sand to a ChatGPT answer and why the US semiconductor supply chain is in a precarious position. Andrew shared what most people get wrong about what makes Nvidia great (it’s not CUDA), and why he thinks of himself as a professional David in an ongoing battle with Goliath. Timestamps: (0:00) Intro (1:07) Why Andrew started Cerebras in 2016 (2:54) Eric on why he invested despite having no chip experience (4:10) Attacking Goliath (9:44) Near-death experiences and the Valley of Death (10:44) 18 months of "still can't make it" (12:16) Solving a 75-year-old compute problem (13:26) What comes after Wafer Scale (16:19) The chip supply chain explained (22:30) Why the US punted a strategic industry (25:51) How to be a good hardware board member (27:50) Hardware vs. software investing (29:05) The pivot from training to inference (27:00) Specialization vs. flexibility (35:22) Young product leaders and seasoned hardware engineers (39:14) External relationships and TSMC (42:20) The AI infrastructure buildout (44:12) The data center supply chain (53:14) Speed creates markets (54:10) Disaggregation with AMD and AWS (55:24) What actually makes Nvidia great (57:30) Near-death experiences and the DNA of a Goliath fighter (59:58) Andrew's childhood next to William Shockley Links: https://x.com/andrewdfeldman https://x.com/cerebras https://x.com/ericvishria https://x.com/jaltma https://uncappedpod.com/ friends@uncappedpod.com

Andrew FeldmanguestJack AltmanhostEric Vishriaguest
Sep 15, 20261h 2mWatch on YouTube ↗

CHAPTERS

  1. 0:24 – 2:30

    Starting Cerebras (2016): spotting an AI compute workload worth building for

    Jack asks why Andrew started Cerebras before the current AI boom. Andrew explains how chip architects look for new workloads where a new architecture can win, and why AI’s compute intensity made it a rare opportunity.

    • Two gating questions for new chips: “Can I build something better?” and “Should I?”
    • AI looked compute-bound (not just power-bound like mobile/Arm)
    • Belief that existing architectures were being adapted rather than designed for AI
    • Early conviction that a purpose-built system could be meaningfully better
  2. 2:30 – 4:10

    Why Benchmark invested anyway: naivete, scar tissue, and underwriting extreme difficulty

    Eric reflects on the psychological challenge of investing in hard hardware after you’ve seen how painful it gets. The conversation highlights how venture often requires optimism in the face of unknowable technical risk.

    • Experience creates “scar tissue” that can make investors overly cautious
    • Most people underestimate how much science/innovation is required
    • Hardware success often requires persistent financing through long uncertainty
    • Humor as truth: founders shouldn’t overshare how hard it will be
  3. 4:10 – 5:51

    “Attacking Goliath”: why incremental improvements can’t beat incumbents like Nvidia

    Andrew lays out Cerebras’ strategy for competing with powerful incumbents. The core idea: modestly better/cheaper loses; you need an order-of-magnitude leap that can’t be competed away with pricing or bundling.

    • Incumbents can neutralize “a bit better” via margins, bundling, and scale advantages
    • New entrants must target 10–500x+ improvements to matter
    • Radical innovation is required; incrementalism won’t create decisive advantage
    • Control the hard parts internally to make the advantage durable
  4. 5:51 – 9:44

    Building the whole stack: chip, board, system, software—and the vendor void

    They discuss why Cerebras had to own everything from hardware to API. Andrew explains that radical designs have no ready-made ecosystem, forcing the company to invent surrounding components (cooling, packaging, etc.) and become world-class through repeated failure.

    • Wafer-scale systems require end-to-end ownership (design through software/API)
    • Radical form factors break the catalog: no off-the-shelf cooling/packaging
    • “Hard stuff inside the building” is controllable; dependencies are not
    • Repeated failures built unique expertise—especially in packaging
  5. 9:44 – 11:54

    The Valley of Death: 18 months of “still can’t make it” and only new mistakes

    Andrew describes the darkest period: burning ~$8M/month while the wafer-scale approach repeatedly failed. Progress came from rigorous failure analysis and the discipline of never repeating the same failure mode.

    • Two axes: technical feasibility (“can you make it?”) vs. demand (“can you sell it?”)
    • 18-month stretch where wafer-scale wouldn’t work; board updates were brutal
    • Systematic failure analysis: learn, fix, and avoid repeating failures
    • “Only new mistakes” as a cultural mantra to sustain momentum
  6. 11:54 – 13:27

    Solving a 75-year compute problem: the moment wafer-scale stability finally held

    Andrew recounts the breakthrough day in 2019 when temperatures flattened and the system ran stably. The team realized they’d solved a long-standing compute scaling problem that others had failed at for decades.

    • Early failures were catastrophic (shattering wafers), then gradually improved
    • Breakthrough characterized by thermal stability and sustained runtime
    • Emotional payoff: solving something “nobody…had ever solved”
    • A milestone that validated the radical approach and unlocked the business
  7. 13:27 – 15:59

    What comes after wafer-scale: cores, memory, and IO as the next frontiers

    Andrew frames future R&D around three fundamentals: compute, memory, and moving results (IO). He describes efforts like stacking HBM onto SRAM-based wafers and optical switching close to compute to attack bandwidth limits.

    • Compute systems boil down to: calculate, store, move
    • Workstreams: faster AI-tuned cores, higher-capacity/faster memory, much faster IO
    • HBM-on-wafer approaches to combine HBM capacity with SRAM-like speed
    • Optical wafer stacking/switching to radically improve communication bandwidth
  8. 15:59 – 21:40

    From sand to answers: the chip supply chain and why fabs are modern pyramids

    Jack requests a simple explanation of the semiconductor pipeline. Andrew demystifies fabs, ASML lithography, wafer processing, and why the whole system is extraordinarily capital-intensive and slow to expand.

    • A cutting-edge fab is a ~$40–$50B, football-field-scale industrial system
    • ASML’s lithography tools are uniquely hard to replicate (near-monopoly)
    • Demand grows exponentially while fabs take ~5 years to build
    • Even after silicon works, packaging/system/software remain major challenges
  9. 21:40 – 23:53

    Why the US “punted” fabs: offshoring cascades into tools, packaging, and vulnerability

    They connect the fab bottleneck to decades of policy and industrial migration. Andrew argues that when fabs left the US, tool vendors and packaging capacity followed, creating strategic fragility that now shows up as supply constraints.

    • Fab-building expertise is scarce; scaling capacity isn’t “cookie cutter”
    • Offshoring fabs pulled ecosystems with it: tools, services, packaging (ASE/Amkor)
    • Strategic risk: disruptions ripple into everything from appliances to defense
    • Proposed remedy: long-term policy and permitting support for domestic fab buildout
  10. 23:53 – 27:48

    Being a good hardware board member: stay in your lane, finance the mission, ask long-game questions

    Eric and Andrew discuss what boards can realistically do in deep-tech hardware. The message: don’t pretend to solve engineering—help with financing, hiring, strategy tradeoffs, and thoughtful governance over multi-year cycles.

    • Board value comes from domain expertise and pattern recognition, not technical meddling
    • Hardware timelines: years to first real customer feedback; first chips are rarely great
    • Useful questions: hiring, roadmap endurance, specialization vs. flexibility
    • Patience and high-quality board discussions can materially improve execution
  11. 27:48 – 29:07

    Hardware vs software dynamics: irreversible decisions, customer concentration, and evolving ML paradigms

    They compare hardware company realities to software: longer time-to-revenue, chunkier deals, and fewer customers. They also note how the ML landscape changed dramatically from TensorFlow/ResNet to transformers, rewarding flexible architectural choices.

    • Hardware revenue arrives later and in large blocks; early customer sets differ
    • Customer concentration is higher (gov, sovereign cloud, frontier labs, hyperscalers)
    • AI shifted from ResNet-era simplicity to transformer-era scale
    • Designing for underlying algebra (not specific model types) preserved adaptability
  12. 29:07 – 34:48

    Training-to-inference shift: “hop-to-hop” strategy and why flexibility mattered

    Andrew explains how Cerebras navigated changing market demands without being able to pivot instantly like software. Their approach was iterative “hop-to-hop” reassessment, combined with early architectural decisions that avoided over-specializing for soon-to-age workloads.

    • Two styles of vision: fixed “circuit” vs adaptive “router hop-to-hop”
    • Inference became dominant as AI intelligence rose and usage exploded
    • Hardware pivots are discrete and slow, so early flexibility choices are critical
    • Avoiding convolution-specific accelerators helped them win when transformers emerged
  13. 34:48 – 42:23

    Teams and trust: young product leaders, seasoned engineers, and relationship-driven execution

    They dig into how Cerebras blended youthful product talent with veteran hardware experience. Andrew emphasizes rapid identification of extraordinary people and the importance of long-built external relationships (e.g., TSMC, manufacturers) in hard times.

    • Experience is essential in silicon/system-building; product insight can skew younger
    • Promotion-from-within and responsibility scaled quickly for exceptional performers
    • Extraordinary talent is evident quickly via clarity, prioritization, and delivery
    • External partnerships in hardware are numerous and trust-based over years
  14. 42:23 – 51:54

    The AI infrastructure buildout: data centers, grid limits, and the real bottlenecks

    Conversation shifts to the physical constraints behind AI growth: power, construction lead times, generators, transformers, and permitting. They argue AI moves at software speed while data centers move at real-estate and industrial speed, creating persistent shortages.

    • Data centers are a critical supply-chain link for both cloud and on-prem compute
    • Constraints include aging grid, long-lead electrical gear, and generator supply
    • Typical timeline: ~18 months for a 50MW block (faster with brownfield sites)
    • Industry missteps in community engagement (water/cost narratives) slowed deployment
  15. 51:54 – 57:29

    Speed creates markets: tokens-per-watt, disaggregation with AMD/AWS, and why Nvidia wins

    They discuss what matters under compute scarcity: throughput, efficiency, and especially speed. Andrew argues speed unlocks entirely new applications, describes disaggregation partnerships that boost throughput, and gives his view of Nvidia’s enduring advantage: grit forged in hard years.

    • Under constraints, maximizing “work per watt/facility” becomes central
    • Speed historically creates new markets (dial-up vs broadband; DVDs to streaming)
    • Disaggregation splits inference work across systems; cited ~5x throughput gains with AMD
    • Nvidia’s “greatness” attributed less to CUDA and more to relentless fight/operating intensity
  16. 57:29 – 1:02:45

    Goliath-fighter DNA and Andrew’s formative years beside Shockley on Stanford campus

    Andrew reflects on maintaining an underdog mindset through repeated skepticism and successive proofs (yield, packaging, customers, scale). He closes by describing his unusual upbringing among Stanford faculty and Nobel-caliber neighbors, and how he almost stayed in academia.

    • Milestones as rebuttals: wafer-scale feasibility → yield → packaging → customer diversity → frontier models
    • Identity as a “professional David” fuels motivation to do what others can’t
    • Childhood surrounded by intellectual achievement (Shockley, Tversky/Kahneman circle)
    • Nearly finished a PhD before being pulled into startups via Stanford business school

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.