Uncapped with Jack AltmanAndrew Feldman on Building Cerebras and the Future of Chips | Ep. 57
At a glance
WHAT IT’S REALLY ABOUT
Cerebras’ wafer-scale bet: radical chip innovation amid AI supply bottlenecks
- Andrew Feldman explains why he started Cerebras in 2016: AI was emerging as a uniquely compute-intensive workload where a new architecture could outperform repurposed incumbents.
- Cerebras’ strategy to “attack Goliath” is to pursue radical, full-stack innovation (wafer-scale compute plus system and software) because incremental improvements can’t beat entrenched players like Nvidia.
- The company’s hardest period was an 18-month stretch where wafer-scale simply wouldn’t work despite an ~$8M/month burn, resolved through rigorous failure analysis and iterative engineering breakthroughs.
- Feldman demystifies the AI chip supply chain (ASML → TSMC → packaging → systems → data centers) and argues that exponential AI demand is colliding with slow-to-build industrial capacity.
- The conversation expands to inference-driven infrastructure: data centers, grid power, generators, and permitting are now key bottlenecks, making speed and throughput (e.g., via disaggregation partnerships) central to compute economics.
IDEAS WORTH REMEMBERING
5 ideasTo beat an incumbent in chips, incremental gains aren’t a strategy—only orders-of-magnitude improvements are.
Feldman argues incumbents can always respond to incremental improvements through pricing, bundling, and scale advantages. To win, a challenger needs a step-function advantage (often 10–1000x) that remains compelling even if the incumbent discounts heavily.
Radical hardware innovation forces you to own the whole stack—and that pain can become the competitive moat.
Cerebras chose wafer-scale and full-stack ownership (chip, board, system, software, API) because radical innovation has no ready-made ecosystem of parts or vendors. Building the “surrounding components” created hard-earned expertise (e.g., packaging) that became a durable moat.
In deep tech, survival is often an extended ‘can we make it?’ phase, not a ‘can we sell it?’ phase.
They spent ~18 months unable to make wafer-scale work while burning about $8M/month, reporting essentially “still can’t make it” in recurring board meetings. Their engineering discipline—deep failure analysis and “only new mistakes”—let them iterate toward a breakthrough moment in 2019 when thermal stability proved the core concept.
The next breakthroughs will come from co-optimizing compute, memory, and IO—not compute alone.
Feldman frames future computer architecture work as advancing compute cores, memory (capacity + bandwidth/latency), and IO. He highlights R&D directions like stacking HBM onto an SRAM-based wafer (to get HBM capacity with SRAM-like behavior) and optical switching/stacking to attack the “moving flops” bottleneck.
AI is constrained by industrial-scale time constants: fabs and data centers can’t scale at software speed.
He describes cutting-edge fabs as $40–$50B “modern pyramids” with multi-year build cycles, fed by ASML lithography and executed at scale by TSMC. The core constraint is that demand (exponential) moves faster than fabs and data centers (multi-year, real-estate-speed) can be built.
WORDS WORTH SAVING
5 quotesYou have board meetings every six weeks. All you've got to say is, "Still can't make it."
— Andrew Feldman
If you're gonna attack Goliath, if, if there's a, a, a giant standing in the market, l- like Nvidia was even at that time, that being a little bit better or a little bit cheaper is, is not an available strategy... What that means is you have to go out with something way better. 10, 100, 500 times faster.
— Andrew Feldman
Holy crap, we've solved this problem that nobody in 75 years of compute had ever solved.
— Andrew Feldman
A fab, and especially a, a fab that, that builds at cutting-edge geometries, is a, is a modern pyramid. It's one of the greatest things humans make.
— Andrew Feldman
AI's moving at the speed of software, and data centers are moving at the speed of real estate.
— Andrew Feldman
High quality AI-generated summary created from speaker-labeled transcript.