Uncapped with Jack AltmanAndrew Feldman on Building Cerebras and the Future of Chips | Ep. 57
CHAPTERS
- 0:24 – 2:30
Starting Cerebras (2016): spotting an AI compute workload worth building for
Jack asks why Andrew started Cerebras before the current AI boom. Andrew explains how chip architects look for new workloads where a new architecture can win, and why AI’s compute intensity made it a rare opportunity.
- •Two gating questions for new chips: “Can I build something better?” and “Should I?”
- •AI looked compute-bound (not just power-bound like mobile/Arm)
- •Belief that existing architectures were being adapted rather than designed for AI
- •Early conviction that a purpose-built system could be meaningfully better
- 2:30 – 4:10
Why Benchmark invested anyway: naivete, scar tissue, and underwriting extreme difficulty
Eric reflects on the psychological challenge of investing in hard hardware after you’ve seen how painful it gets. The conversation highlights how venture often requires optimism in the face of unknowable technical risk.
- •Experience creates “scar tissue” that can make investors overly cautious
- •Most people underestimate how much science/innovation is required
- •Hardware success often requires persistent financing through long uncertainty
- •Humor as truth: founders shouldn’t overshare how hard it will be
- 4:10 – 5:51
“Attacking Goliath”: why incremental improvements can’t beat incumbents like Nvidia
Andrew lays out Cerebras’ strategy for competing with powerful incumbents. The core idea: modestly better/cheaper loses; you need an order-of-magnitude leap that can’t be competed away with pricing or bundling.
- •Incumbents can neutralize “a bit better” via margins, bundling, and scale advantages
- •New entrants must target 10–500x+ improvements to matter
- •Radical innovation is required; incrementalism won’t create decisive advantage
- •Control the hard parts internally to make the advantage durable
- 5:51 – 9:44
Building the whole stack: chip, board, system, software—and the vendor void
They discuss why Cerebras had to own everything from hardware to API. Andrew explains that radical designs have no ready-made ecosystem, forcing the company to invent surrounding components (cooling, packaging, etc.) and become world-class through repeated failure.
- •Wafer-scale systems require end-to-end ownership (design through software/API)
- •Radical form factors break the catalog: no off-the-shelf cooling/packaging
- •“Hard stuff inside the building” is controllable; dependencies are not
- •Repeated failures built unique expertise—especially in packaging
- 9:44 – 11:54
The Valley of Death: 18 months of “still can’t make it” and only new mistakes
Andrew describes the darkest period: burning ~$8M/month while the wafer-scale approach repeatedly failed. Progress came from rigorous failure analysis and the discipline of never repeating the same failure mode.
- •Two axes: technical feasibility (“can you make it?”) vs. demand (“can you sell it?”)
- •18-month stretch where wafer-scale wouldn’t work; board updates were brutal
- •Systematic failure analysis: learn, fix, and avoid repeating failures
- •“Only new mistakes” as a cultural mantra to sustain momentum
- 11:54 – 13:27
Solving a 75-year compute problem: the moment wafer-scale stability finally held
Andrew recounts the breakthrough day in 2019 when temperatures flattened and the system ran stably. The team realized they’d solved a long-standing compute scaling problem that others had failed at for decades.
- •Early failures were catastrophic (shattering wafers), then gradually improved
- •Breakthrough characterized by thermal stability and sustained runtime
- •Emotional payoff: solving something “nobody…had ever solved”
- •A milestone that validated the radical approach and unlocked the business
- 13:27 – 15:59
What comes after wafer-scale: cores, memory, and IO as the next frontiers
Andrew frames future R&D around three fundamentals: compute, memory, and moving results (IO). He describes efforts like stacking HBM onto SRAM-based wafers and optical switching close to compute to attack bandwidth limits.
- •Compute systems boil down to: calculate, store, move
- •Workstreams: faster AI-tuned cores, higher-capacity/faster memory, much faster IO
- •HBM-on-wafer approaches to combine HBM capacity with SRAM-like speed
- •Optical wafer stacking/switching to radically improve communication bandwidth
- 15:59 – 21:40
From sand to answers: the chip supply chain and why fabs are modern pyramids
Jack requests a simple explanation of the semiconductor pipeline. Andrew demystifies fabs, ASML lithography, wafer processing, and why the whole system is extraordinarily capital-intensive and slow to expand.
- •A cutting-edge fab is a ~$40–$50B, football-field-scale industrial system
- •ASML’s lithography tools are uniquely hard to replicate (near-monopoly)
- •Demand grows exponentially while fabs take ~5 years to build
- •Even after silicon works, packaging/system/software remain major challenges
- 21:40 – 23:53
Why the US “punted” fabs: offshoring cascades into tools, packaging, and vulnerability
They connect the fab bottleneck to decades of policy and industrial migration. Andrew argues that when fabs left the US, tool vendors and packaging capacity followed, creating strategic fragility that now shows up as supply constraints.
- •Fab-building expertise is scarce; scaling capacity isn’t “cookie cutter”
- •Offshoring fabs pulled ecosystems with it: tools, services, packaging (ASE/Amkor)
- •Strategic risk: disruptions ripple into everything from appliances to defense
- •Proposed remedy: long-term policy and permitting support for domestic fab buildout
- 23:53 – 27:48
Being a good hardware board member: stay in your lane, finance the mission, ask long-game questions
Eric and Andrew discuss what boards can realistically do in deep-tech hardware. The message: don’t pretend to solve engineering—help with financing, hiring, strategy tradeoffs, and thoughtful governance over multi-year cycles.
- •Board value comes from domain expertise and pattern recognition, not technical meddling
- •Hardware timelines: years to first real customer feedback; first chips are rarely great
- •Useful questions: hiring, roadmap endurance, specialization vs. flexibility
- •Patience and high-quality board discussions can materially improve execution
- 27:48 – 29:07
Hardware vs software dynamics: irreversible decisions, customer concentration, and evolving ML paradigms
They compare hardware company realities to software: longer time-to-revenue, chunkier deals, and fewer customers. They also note how the ML landscape changed dramatically from TensorFlow/ResNet to transformers, rewarding flexible architectural choices.
- •Hardware revenue arrives later and in large blocks; early customer sets differ
- •Customer concentration is higher (gov, sovereign cloud, frontier labs, hyperscalers)
- •AI shifted from ResNet-era simplicity to transformer-era scale
- •Designing for underlying algebra (not specific model types) preserved adaptability
- 29:07 – 34:48
Training-to-inference shift: “hop-to-hop” strategy and why flexibility mattered
Andrew explains how Cerebras navigated changing market demands without being able to pivot instantly like software. Their approach was iterative “hop-to-hop” reassessment, combined with early architectural decisions that avoided over-specializing for soon-to-age workloads.
- •Two styles of vision: fixed “circuit” vs adaptive “router hop-to-hop”
- •Inference became dominant as AI intelligence rose and usage exploded
- •Hardware pivots are discrete and slow, so early flexibility choices are critical
- •Avoiding convolution-specific accelerators helped them win when transformers emerged
- 34:48 – 42:23
Teams and trust: young product leaders, seasoned engineers, and relationship-driven execution
They dig into how Cerebras blended youthful product talent with veteran hardware experience. Andrew emphasizes rapid identification of extraordinary people and the importance of long-built external relationships (e.g., TSMC, manufacturers) in hard times.
- •Experience is essential in silicon/system-building; product insight can skew younger
- •Promotion-from-within and responsibility scaled quickly for exceptional performers
- •Extraordinary talent is evident quickly via clarity, prioritization, and delivery
- •External partnerships in hardware are numerous and trust-based over years
- 42:23 – 51:54
The AI infrastructure buildout: data centers, grid limits, and the real bottlenecks
Conversation shifts to the physical constraints behind AI growth: power, construction lead times, generators, transformers, and permitting. They argue AI moves at software speed while data centers move at real-estate and industrial speed, creating persistent shortages.
- •Data centers are a critical supply-chain link for both cloud and on-prem compute
- •Constraints include aging grid, long-lead electrical gear, and generator supply
- •Typical timeline: ~18 months for a 50MW block (faster with brownfield sites)
- •Industry missteps in community engagement (water/cost narratives) slowed deployment
- 51:54 – 57:29
Speed creates markets: tokens-per-watt, disaggregation with AMD/AWS, and why Nvidia wins
They discuss what matters under compute scarcity: throughput, efficiency, and especially speed. Andrew argues speed unlocks entirely new applications, describes disaggregation partnerships that boost throughput, and gives his view of Nvidia’s enduring advantage: grit forged in hard years.
- •Under constraints, maximizing “work per watt/facility” becomes central
- •Speed historically creates new markets (dial-up vs broadband; DVDs to streaming)
- •Disaggregation splits inference work across systems; cited ~5x throughput gains with AMD
- •Nvidia’s “greatness” attributed less to CUDA and more to relentless fight/operating intensity
- 57:29 – 1:02:45
Goliath-fighter DNA and Andrew’s formative years beside Shockley on Stanford campus
Andrew reflects on maintaining an underdog mindset through repeated skepticism and successive proofs (yield, packaging, customers, scale). He closes by describing his unusual upbringing among Stanford faculty and Nobel-caliber neighbors, and how he almost stayed in academia.
- •Milestones as rebuttals: wafer-scale feasibility → yield → packaging → customer diversity → frontier models
- •Identity as a “professional David” fuels motivation to do what others can’t
- •Childhood surrounded by intellectual achievement (Shockley, Tversky/Kahneman circle)
- •Nearly finished a PhD before being pulled into startups via Stanford business school