Skip to content
No PriorsNo Priors

The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile

Founder and CEO of full-stack AI chip company Fractile, Walter Goodwin, joins Sarah Guo to discuss the bets he’s made on the future of the chip market as other major players like Broadcom, NVIDIA, and AMD try to accelerate their workloads. They discuss the difference in Fractile’s newer approach on model architecture with a full-stack team in the current chip landscape and the technical bets they’re making in that direction. Walter also talks about compressing the gap between the chip design cycle and its payoff period, and making a generational leap in AI inference to realize the bet in volume against the value to be captured. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @goodwin_ml | @fractile_ai Chapters: 00:46 – Walter Goodwin and Fractile Introduction 02:29 – The Chip Landscape Now 04:56 – Common Handoffs From Architecture-Focused Players 07:18 – Full Stack Approach and Team Setup 09:49 – Fractile’s Most Important Technical Bets 15:32 – Workload Predictions and Compressing the Chip Design Cycle 23:03 – Architect Intent to Output Bottlenecks and Accelerating Trials 28:16 – Workload Bets on Model Architectural Shifts 31:20 – The Future of AI Chip Players and Market Structure 35:14 – Conclusion

Walter GoodwinguestSarah Guohost
Oct 2, 202635mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Fractile’s bet: bandwidth-first chips to win fast frontier inference

  1. Fractile is building ultra-fast inference chips aimed at serving the largest frontier models, with a central focus on maximizing effective memory bandwidth to accelerate real-time generation and agentic workloads.
  2. Goodwin argues the current accelerator landscape is less diverse than it looks because many hyperscaler and third-party ASICs converge on similar ingredients—HBM, tensor-core-like compute, and TSMC packaging—often delivered through a small set of ASIC houses like Broadcom.
  3. Fractile positions itself as “full-stack,” keeping architecture, RTL, physical design, packaging, and foundry interaction in-house to create a tighter feedback loop and reduce dependency on slow organizational handoffs.
  4. A major technical pivot discussed is moving from SRAM-heavy approaches (extremely fast but capacity-limited) toward architectures that can provide very high bandwidth access to higher-capacity, cheaper DRAM-class memory to better support long contexts and large-state inference.
  5. On market structure, Goodwin expects buyers to maintain multi-sourcing (NVIDIA/AMD/internal/accelerators), while cautioning frontier labs against deeply idiosyncratic silicon bets that could become existential if competitors discover hardware-specific algorithmic advantages.

IDEAS WORTH REMEMBERING

5 ideas

Inference speed (not training) becomes the decisive battleground as AI deployment scales.

Fractile’s core thesis is that as foundation models move from demos to ubiquitous products, the recurring (marginal) cost of serving inference dominates—and latency/speed becomes a primary axis of competitive advantage, especially for long-running agentic workloads.

The key architectural challenge is marrying extreme bandwidth with economical, scalable memory capacity.

Early SRAM-centric designs can deliver extreme bandwidth but hit capacity/scalability limits as models grow and contexts explode; Fractile shifted toward architectures that aim to provide very high bandwidth while leveraging higher-capacity, cheaper DRAM-class memory.

Many AI ASICs are more alike than they appear; real differentiation demands full-stack execution.

Goodwin argues that much of today’s “zoo” of accelerators is structurally similar (HBM + tensor cores + TSMC packaging), especially hyperscaler internal chips built with common ASIC houses; differentiated advantage requires going deeper across the stack.

Full-stack organization is a speed lever: it tightens iteration loops and reduces handoff friction.

Owning front-end, back-end, physical design, and packaging inside a ~150-person company is framed as a way to shorten the observe→design→ramp loop and avoid being “at the mercy” of partner timelines and priorities.

You won’t ship a new chip every few weeks, but faster cycles let you pick better ramp decisions.

Even if fabs and deployment economics impose hard lower bounds (e.g., 3–5 month fab cycles; multi-year amortization), compressing the design cycle increases “shots on goal” and improves the odds of ramping the right platform at the right time.

WORDS WORTH SAVING

5 quotes

The snappier chatbot is kind of the faster horses of kind of fast inference.

— Walter Goodwin

We’ve scaled FLOPS like a million fold in the last twenty years. Memory bandwidth has gone up about forty x in the same timeframe.

— Walter Goodwin

It’s the equivalent of the Frontier model for the chip space is if you can just find a way to structurally carve out a three to six months advantage, uh, you will be winning all of those deployments.

— Walter Goodwin

I think having deeply differentiated bets and going all in on those differentiated bets at the chip layer is, uh, sort of an irrational and a very dangerous move for anybody that is playing at the frontier.

— Walter Goodwin

There’s a bit of a joke today that the sort of first-party efforts, their primary purpose is to reduce the price that people pay NVIDIA, and there may be some truth to that as well because those efforts are kind of architecturally quite similar, right?

— Walter Goodwin

Full-stack AI chip developmentInference vs. training economicsMemory bandwidth vs. FLOPs scalingSRAM vs. DRAM/HBM capacity-bandwidth tradeoffsLong-context and KV-cache constraintsMoE sparsity and attention mechanism churnChip design cycle compression and EDA bottlenecks (GDS II, place-and-route, DRC/LVS)

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.