No PriorsFrontier Chips for Frontier AI Labs, with Walter Goodwin, Founder/CEO of Fractile
EVERY SPOKEN WORD
35 min read · 7,062 words- 0:00 – 0:46
Intro
- WGWalter Goodwin
Right now, we're trying to build a single chip. If you crack open an NVIDIA system, it has anywhere between kind of six and nine custom chips all built by NVIDIA to come together to build something that is really, really potent. So there's already a gap there. You know, it'd be great to be productive enough that we could start to build our own responses in that same kind of way. This idea of being able to have, at any given moment, a kind of rolling frontier of bets that you hope are deeply aligned, in which you're ready to trigger the ramp of. This is a huge advantage if you can kind of build that machine and build that engine. It's the equivalent of the frontier model for the chip space, is if you can just find a way to structurally carve out a three to six months advantage, you will be winning all of those deployments.
- SGSarah Guo
[upbeat music]
- 0:46 – 2:29
Walter Goodwin and Fractile Introduction
- SGSarah Guo
Hi, listeners. Welcome back to No Priors. Today, I'm here with Walter Goodwin, founder and CEO of full-stack AI chip company, Fractile. We talk about what it means to be full-stack as a company, the technical bets they're making, why he's so focused on memory bandwidth, his predictions for frontier model architectures in the future, and the structure of the chip market at the frontier today and tomorrow amongst NVIDIA, AMD, internal efforts, and this new class of accelerators. Welcome, Walter. Walter, thanks so much for doing this.
- WGWalter Goodwin
Very good to be here, Sarah.
- SGSarah Guo
You started this incredibly interesting company called Fractile. Can you give us an overview of what you guys do?
- WGWalter Goodwin
Fractile is a chip company. Uh, we build very, very fast inference chips for the world's largest models. This sort of bet on speed is something which we've had right from the outset. Um, started the company in summer of twenty twenty-two, and I think at the time we were starting to see two things. You know, one was obviously the arrival of foundation models, uh, trained on the internet and generalizing across everything. And the other was, you know, a few kind of wise people saying, you know, "We need to find a way to pour more compute into these models at test time. We need to find a way to take what we had for kind of AlphaGo, where you have a capable neural network, but then it becomes superhuman as you roll it out at scale and bring that to kind of language modeling," for example. Um, and so our big pursuit over the past four years has been to find a way to build chips that will allow for kind of simultaneously taking these huge models and running them much, much, much faster than the chips of today, um, but do that in a way that kind of scales. So scales to models beyond the frontier today, scales to extraordinarily long contexts. Um, and so that's the, the mission that we've
- 2:29 – 4:56
The Chip Landscape Now
- WGWalter Goodwin
been on.
- SGSarah Guo
It's become a much more interesting chip landscape in the, uh, four years since you started the company. Uh, where would you place yourself in the, uh, overall GPU and accelerator landscape for inference?
- WGWalter Goodwin
Yeah, I'd say, uh, there's a kind of extraordinary zoo now of options available. Um, and if you look across this kind of entire space of kind of AI ASICs, um, one of the things that I think is very striking is there's a lot of relatively identikit chips out there. Um, and this is kind of a structural-- It's a kind of structural industry, uh, property, where if you look especially across kind of hyperscaler efforts, and, you know, Google kind of kicked this off with the TPU more than ten years ago. Um, you see this paradigm where although there are a number of kind of proprietary internal chips, you know, Google with the TPU, Meta with MTIA, Microsoft with Maia, OpenAI now with Jalapeno, um, these are chips that are ultimately developed and delivered in partnership with kind of a relatively small number of what we might call ASIC houses, these kind of back-end delivery houses. Broadcom is, is the largest at, you know, two trillion dollar company, um, that help people kind of realize these designs. And so when you kind of zoom into this, this sort of seeming, uh, embarrassment of riches in the spectrum of what exists today, uh, you see that actually there's a lot of similarity across these platforms. Um, you know, the-- all of these chips, you have HBM, uh, this type of DRAM memory that is, you know, common to NVIDIA GPUs, AMD GPUs, and all of these other ASICs. You have the same kinds of bets on tensor cores to do your matrix multiplications, the same advanced packaging with TSMC. Um, and so I think one of the things that we see in the landscape is there's a relative dearth still of kind of efforts that go all the way up and down the silicon stack and try and build kind of fundamentally new capabilities. Um, and this is kind of a structural thing. There's just not that many teams that have decided we're gonna build from, you know, what you might call the kind of, you know, the architecture layer, the front-end design layer, which mostly looks like kind of writing code, but also all the way down to what we call kind of physical design, the process technology, the kind of foundry interactions, the stuff that essentially has you flying back and forth to Taiwan or Korea every week. Um, that's something which actually tends to still be left to a relatively small number of companies.
- 4:56 – 7:18
Common Handoffs From Architecture-Focused Players
- SGSarah Guo
For people who don't work in the chip industry, what is the common abstraction or handoff between what they're-- what, you know, architecture-focused players are doing and a player like Broadcom?
- WGWalter Goodwin
Yeah. So I think the-- If you're-- Let's say you're Google and you're designing the TPU, um, you have a number of people in your team who have a really strong understanding of the workloads that they're, they're trying to accelerate. Um, and so, you know, going back over ten years, this is, uh, a world that then sees you conceive of things like the Tensor Core, which is a de-designated circuit dreamt up by an architect, uh, that is exceptionally good at matrix multiplications because you have this insight that matrix multiplications are the overwhelming majority of the kind of flop count in these models that you're running. Um, that's kind of the preserve of the architect, I guess, is somebody who really understands the way that chips behave, but also ideally has a very strong understanding of the kind of target workloads. Um, what you'll then see inside those same organizations is a set of sort of, you know, very smart people that kind of lower that down into, I guess, you know, essentially a kind of circuit-level description of how this chip should behave. And so- A lot of this is what we call front-end design. Um, you know, mechanistically, it still looks like, like writing code on a computer, and it's still code that is then handed off to, uh, a player like Broadcom. Um, and so, you know, that kind of front-end design that describes the kind of underlying intent of the chip, the underlying logic of the chip, is then transformed ultimately into something that, you know, what gets shipped to TSMC in the end is literally a kind of bitmap. Uh, you know, we have this file type you call GDS II, and it's literally, you know, where do I put the metal layers? Where do I put each individual transistor? And so it's a, it's a full layout, and so you go through this process of synthesis of that RTL, uh, into a kind of set of circuits that then ultimately get laid out. And a lot of the complexity there is around, um, you know, in Broadcom's case, you know, it will be ownership of analog IP that is used for chip-to-chip connectivity, and it is this kind of physical placement that is specific to a given process node at TSMC. So, uh, you know, we talk about, like, three nanometers, five nanometers, et cetera. Um, so that piece is still today, you know, generally it is a kind of outsourced activity for all of these projects.
- 7:18 – 9:49
Full Stack Approach and Team Setup
- SGSarah Guo
You are describing what feels like a very large project to do end-to-end, um, new chip creation. Uh, from a workload perspective, you're working with several frontier players to sort of characterize that and make sure that you can serve them. How did you think about-- Like, what does the Fractile team look like today? How can you do that as a startup?
- WGWalter Goodwin
Yeah. I think there is a-- You know, it's funny, there's a lot of industries where there will be a kind of waterfall-style approach to doing things, and then suddenly, you know, along comes a new way of thinking, and it's kind of more agile, to borrow from, like, management philosophy, right? And I think for us as a very full-stack company, so we have a team that does span, you know, very deep workload understanding. I think we actually try to, you know, run ahead, uh, in many places and look at, you know, where would we change a model architecture to be more exceptionally aligned with our bet? What do we think is, like, kind of the scaling law vector for the particular bets that we are taking? Uh, and then we can advocate to our customers and our partners on that, and it also helps us inform a set of bets for a chip that, you know, isn't going to necessarily exist in volume for, you know, one or two years. Um, that becomes a really important sort of part of the muscle that we have to build. But then as we look across this kind of, you know, ability to spin a very kind of agile loop, it means that inside Fractile, we do have front-end designers. We have our own physical design team. We have our own kind of back-end implementation team. You know, we do our own work on advanced packaging, for example. Um, and that's not an enormous headcount, right? I mean, Fractile today is about a hundred and fifty people. Uh, we're sort of skinny in, in every single one of those, uh, sectors, but what it allows us to do is have this kind of much more agile closed loop. Um, and I think that's becoming increasingly essential as you look at the sort of cadence needs that are set by this industry. Um, you know, the chasing, you know, perennially of the sort of chasing of the tail of the workloads is, you know, you have to, in the AI chip space more than any other kind of chip play before, you have to really place your bets correctly. Uh, so you have to have a lot of, you know, skill and a lot of luck, and you also then need to strike very fast. And so this kind of structural setup where we have the entire story in-house, uh, it's very, very different to this kind of, I guess, prior approach where there's almost like a handoff point. And so you get to a certain level, and then you hand off to another partner. And at some-- to some extent, you're kind of at the mercy of, of how that other partner then behaves.
- 9:49 – 15:32
Fractile’s Most Important Technical Bets
- SGSarah Guo
What do you think are the single most important technical bets that the company has made?
- WGWalter Goodwin
So for us, I think, you know, one is the direction. You know, four and a half years ago, um, uh, you were obviously, uh, right in the heart of an inference story, but I still felt that we were educating the world about what inference meant. Um, for two years or so following that, some of the inference chips today, including one that recently went public, was a training chip until about eighteen months ago. You know, there was a real avoidance of this idea, this sort of small, you know, final marginal cost. You know, marginal cost I think has two senses, right? It either sounds like it's a small cost. What it in fact means is this is the cost that you pay every single time you deploy these models. Um, so I think firstly, just the, the sheer bet that we were going to slip into a kind of deployment era, um, and the bet on speed. That was something that kind of really informed a lot of what we then did as a kind of architectural response. On the architectural side, you know, we've been through a journey, I think, for the first kind of two years or so of the company's life. Um, we were working on a little like Groq or Cerebras, uh, working on an SRAM-based chip. So we'd sort of made the observation that, um, uh, SRAM is this super high bandwidth memory. It lives on the same piece of silicon as your logic, uh, and so you have this extremely high bandwidth between your compute, where you're running the maths as you run these models, and let's say the weights of your model or the KV cache of this model as you're rolling it out. Um, and that's something which then, then can drive you to thousands of tokens per second on these language models. Um, but I think one of the things that we started to worry about in kind of, uh, towards the end of 2023 and certainly in 2024 was the scalability of this approach. Um, and you know, I think there are two things that ... grow with AI today. One is obviously the parameters of the model. Um, but the other, and this was the one that kind of really got us nervous about that architectural approach was, um, uh, you know, this kind of growing context length that was becoming more and more part of the story for, for how we saw these models rolling out. Um, and it's that that has taken us over the past couple of years to a, a kind of what we see as a more exciting bet, I guess, in some ways, which is working more closely with memory vendors as well as our sort of logic foundry partners to find ways to get extremely high bandwidth access to higher capacity memories. So for the past couple of years, we've been involved in these kind of skunkworks projects to move away from SRAM, you know, looking at how we can get, for instance, much, much, much higher bandwidth to DRAM memories. Um, and this has been a very exciting bet for us because it's allowed us now to put together a platform which we'll be ramping in the second half of next year, um, which, uh, kind of unites, I guess, the scalability of these kind of higher capacity, lower cost DRAM memories that you have on a GPU, on a TPU, uh, with all of the speed advantages that you get from a Groq chip or a Cerebras chip. Um, and I think where we see this as especially vital is if you look at where it is that speed really moves the dial today, there is a kind of-- there's a version of this which is like a snappier chatbot. But I think this is like, you know, when, um, uh, Henry Ford sort of asked the rhetorical question of what people would've said they wanted and, you know, faster horses, uh, rather than a car. Uh, the snappier chatbot is kind of the faster horses of kind of fast inference. The place where the ability to take a multi-trillion parameter model and run it comfortably at many thousands of tokens per second, in our view, is a kind of fundamental elevator on capability for AI is in taking these kind of very long-running agents and making them radically faster. Um, and so that's kind of a-- it's a tantalizing mismatch today between the properties of the fast inference chips that we have, which have super high bandwidth memory, but incredibly low capacity. And if you look at the sort of technical detail of how they actually get employed and deployed today, they're not running long context attention. You still flip back to a GPU to do that. And so there's this sort of tantalizing mismatch where precisely the area where we are almost able to go and take these things and make them much faster, we also don't have that technical capability. Um, and so in a sense, this is, you know, like many technical challenges, it boils down to a slightly mundane technical observation, which is we need chips that have this kind of particular ineffable property, which is incredibly high bandwidth to memory, so we can load the weight, so we can load this state, you know, many thousands of times per second, uh, but also have like economical memory. Because when you look at what it looks like to run inference at data center scale for thousands of users, the economics actually collapse down to, uh, essentially a kind of cost per gigabyte of the memory that you're employing. Um, and so that's been the sort of other key focus for us, is unlocking this kind of fundamentally new building block, which is, uh, finding a path to get aggressively high bandwidth from this kind of, uh, you know, from the world's lower cost memory, uh, DRAM memory.
- SGSarah Guo
You've said a few things that I, I think are, um, not the conventional wisdom. One, in, in, in terms of making a bet yourselves on where the workloads are going and then, two, the idea that you can, um, you know, traditional chip people might think of delivery of chips in a generation-
- WGWalter Goodwin
Hmm
- SGSarah Guo
... you know, every year or a little faster than that. Um, and, you know, big architectural changes c-come slower than, um, that, that sort of tick rate. Uh, but you,
- 15:32 – 23:03
Workload Predictions and Compressing the Chip Design Cycle
- SGSarah Guo
you, you've aspirationally said within Fractile that you think the speed of like change in architecture, um, and change in, you know, the, the technical plays you're making can be much faster than that. Can you talk a little about like what would enable that to happen? Because the, you know... I, I, I think people have a better understanding of the physical limits of the supply chain today than they-
- WGWalter Goodwin
Yeah
- SGSarah Guo
... used to. So what is the flexibility that the industry could have?
- WGWalter Goodwin
Yeah. So I think there's a nuance on, uh, on this, but what, what you want at any given moment in time for an AI chip is you wish that you had a chip that was precisely targeting the workload that you care about, uh, and in your hand today in volume. But if you had that chip, uh, you know, you're gonna want something else again in six months' time. Like, these workloads move on very fast.
- SGSarah Guo
Well, today, you know, we have a model coming out about every, uh, two weeks. So it's a little tough.
- WGWalter Goodwin
That's right.
- SGSarah Guo
Yeah. [chuckles]
- WGWalter Goodwin
And so, so sanity comes from looking across those models and saying, "Well, okay, what's, what's common to all of these?" And like, thankfully, there's enough that is common. So, you know, it's always the case that a new LLM, uh, desperately wants to have orders of magnitude more bandwidth to memory in order to run faster. Uh, it remains the case today that, you know, these LLMs tend to be auto-aggressive, very low batch as you generate text, uh, with this sort of fundamental trade-off again between kind of throughput efficiency and cost and the speed that you can serve these models at. Uh, so there are certain things that these models are unified and kind of crying out for, and, and memory bandwidth is one of them. Um, but then, you know, as you say, there are these evolutions, and especially if you look at the kind of frontier of open source Chinese models, you know, the exact nature of an attention mechanism will churn every couple of weeks in terms of what-
- SGSarah Guo
Or the sparsity
- WGWalter Goodwin
... kind of at the frontier. The sparsity, the sparsity level of the MOEs that you have, the sparsity on the attention itself. Um, and you know, we may be in for like bigger shocks as well. Um, there may be kind of more fundamental shifts. And so I think there's this duality, uh, around sort of the limits of the physical and the kind of financial world, I guess, and then these demands of, you know, workloads that churn at the pace of essentially software, where although there is this sort of clear desire to be able to kind of ship new platforms, you know, aggressively faster, I think there are also some sort of fundamental breaks on this. So, so at Fractile, w-you know, we're extremely excited about being AI forward in how we do chip design. I think we've seen this ability to own the entire end-to-end gamut of the problem is really enabling, uh, a kind of fundamental rethink of how we do things. You know, if you think about, um, classic CS law is Amdahl's law, right? Like everything I can paralyze becomes very, very fast, and the part that I can't paralyze doesn't. And I think it's similar when you're in a, like, long chain of organizations, um, that You know, there isn't necessarily an overwhelming need to, uh, or return on radically changing your processes as a front-end-focused chip design startup adopting AI in aggressive ways to shorten your timelines from being, you know, maybe 12 months of front-end design work to a tiny fraction of that if there's then going to be kind of the conventional bottlenecks and the conventional pace on the next piece. Um, and I think that's sort of, you know, at some le- extent there will always be those bottlenecks. You know, we all have the same kind of fab cycle times with our foundry partners. Uh, so from the moment you send the chip to them to getting it back is, you know, three to five months even in a kind of super hot lot scenario. And so these are the kind of innate latencies, I guess, in the industry. And then when you get that chip back, I think it is vital to look at the, uh, the way that these chips actually become financializable, which is that they do have a kind of payoff period. You have to have really a kind of three-to-five-year amortization window for that chip to make it a, a sensible financial decision. Um, and so there's these sort of two things that are in tension in which I, I guess I believe two slightly conflicting things at the same time. One is there is a huge value in being able to take the overall chip design cycle and compress that and compress that. Uh, there is also, I think w- what we don't lose is the need to place really thoughtful and subtle bets architecturally because the chip you actually end up making does still need to have a, a three-year-plus, you know, useful lifespan. Um, and so I think for, for me, you know, what it means to be able to over time compress that kind of chip design cycle is essentially that you get to have more shots on goal. You want to be forever ready to take a given flagship platform and say, "Right, that's it, that is actually the flagship, and this is the one we're gonna ramp." But then it's a, you know, 12-month, 18-month volume ramp, and you're expecting it to have longevity. You're expecting it to actually continue to drive value to your customers. And so I think that the, the sort of subtlety here is, you know, I'm not a believer in the idea that we will ultimately get to a place where we are shipping a fundamentally new chip, uh, you know, every few weeks just because we've shortened that down because I think this is a, this is a physical world. There is a certain amount of constraint on, you know, data center power. There are latencies to installing these things, and you need to finance the actual underlying silicon, so it needs to have that payoff period. But I think what you can get to is a world where the shorter you can make that latency, that gap between an observation and realizing that bet in volume, that is where there's an a, you know, an enormous amount of value to be captured. Um, and I think it's also an area where, you know, Fractile is a-
- SGSarah Guo
If only because y- the decisions that you make about what to ramp are the correct ones.
- WGWalter Goodwin
That's right. I think you, you know, you want to be-- And maybe you get to have more irons in the fire as well. So, you know, one of the things that I think is so exciting about AI capability rising is, um, as we do become more productive in doing things, uh, you know, there's obviously these kind of two responses, you know, generally across the economy. One is, "Oh, no, we're not gonna have as much to do. We may-- Maybe there's jobs loss." The other is, uh, we get to do more things. Um, and I think, you know, standing here from, from Fractile, you know, right now we're trying to build a single chip. We know that we're up against competitors that have, uh, you know, if you crack open an NVIDIA system, it has anywhere between kind of six and nine custom chips all built by NVIDIA to come together to build something that is really, really potent. Um, so there's already a gap there. You know, it'd be great to be productive enough that we could start to build our own responses in that same kind of way. Um, and so I think this idea of being able to have at any given moment a kind of rolling frontier of bets that you hope are deeply aligned and which you're ready to trigger the ramp of, this is a huge advantage if you can kind of build that machine and build that engine. And the six-month gap that that might then perennially give you against your comp- competition, uh, that is the kind of, you know, wedge that allows you to drive in. And we see this in, in Frontier Labs, right? It's the equivalent of the Frontier model for the chip space is if you can just find a way to structurally carve out a three to six months advantage, uh, you will be winning all of those deployments.
- SGSarah Guo
And there's a meme in here somewhere of CEOs saying, "Wow, this AI stuff is great, and we structured the organization in a way that can consume it. Instead of one program, we're gonna have six, guys," and everybody cheers.
- WGWalter Goodwin
Yeah.
- SGSarah Guo
[chuckles]
- WGWalter Goodwin
I think there's some truth in that.
- SGSarah Guo
What,
- 23:03 – 28:16
Architect Intent to Output Bottlenecks and Accelerating Trials
- SGSarah Guo
uh-- I was talking to, um, one of the top three semi CEOs, and he wouldn't let me put this prediction, uh, you know, actually on air, but I, I asked him, like, "How soon until we can have essentially, like, intent from chief architect into a fully usable GDS II file?" And he's like, "Ten years."
- WGWalter Goodwin
Yeah.
- SGSarah Guo
It's not gonna happen, like, now. What, what is your view on that? Or, or where do you see the impact first in your organization?
- WGWalter Goodwin
You know, a, a heuristic that has been successful over the past couple of years is to just always question your, uh, your sort of logical assumption and then divide it by four on timescale.
- SGSarah Guo
[chuckles]
- WGWalter Goodwin
So, you know, there is a world in which 10 years seems like a very sensible timeframe for that, but I would divide it by four, and I'd probably subtract a bit from that. You know, I think that we will have, uh, a space where there is a kind of prototyping that goes end to end, um, uh, you know, in the next few years. I think one of the things about chip design that is very interesting is there are still these loops in the middle of it that are kind of conventional solutions to NP-hard problems. And so you have these sort of-- If you really wanna output that GDS II file, today you are still running through a bunch of layout tools that are doing place and route with conventional kind of, uh, um- Algorithmic approaches which will run for days on end. You know, I think this is a, a really interesting thing that we see, and actually to bring it back somewhat to the sort of workloads we're excited about as well. Um, a lot of hard problems in the world, I think, look a bit like this, where there is a lot of thinking that can now be automated. There's a lot of cerebral work. Um, but then there is also, uh, you know, this kind of intrinsic latency to something that you do. Um, so I think we see this in, uh, you know, AI for guiding, uh, AI model development today, so the RSI work. Um, you know, yes, we can now have extreme intelligence that is guiding our experiments, but now we're just experiment bottlenecked. You know, we're compute time bottlenecked. And so I think with, with chip design, it's a little bit like this. Maybe to take the kind of, you know, uh, architect's intent through to GDS II question. It's a little bit like the RSI question in that the things that will prevent us getting there extremely fast is the fact that actually in that flow, there are a number of things that are genuinely kind of computationally expensive in a more conventional way. I think what we will generally see for all of these sorts of problems and this-
- SGSarah Guo
And they won't be replaced by surrogate models? [laughs]
- WGWalter Goodwin
Well, I think, you know, there's definitely these sort of, uh, simulation models, I think driving approximations, uh, you know, uh, all of the models that we see today for kind of finite element analysis and, and sort of thermal estimation and so on. I think there is a huge return to that kind of work precisely because of this kind of Amdahl's law, that as we accelerate the intelligence, as we accelerate the period between these experiments, between these trials-
- SGSarah Guo
They become the bottleneck, yep
- WGWalter Goodwin
... it becomes much more important to accelerate those trials as well. And so, you know, we are looking internally at, you know, what is actually-- what are in the guts of some of those algorithms, and is there a crude approximate way that you could do some of this? You know, uh, the-- I think chip design is not going to change in the final sign-off for quite some time. It's incredibly valuable to have Cadence and Synopsys, who've been working with TSMC and the other foundries for decades, to build the sort of, you know, the final check mark that says, "This thing is, uh, what we call DRC and LVS clean. It conforms to the rules that that foundry has set." Um, that is actually where, you know, that sort of work is immensely valuable. But could you have some kind of fuzzy placement algorithm in the meantime that gets you almost all the way to your final design?
- SGSarah Guo
It helps you iterate faster against it, yeah.
- WGWalter Goodwin
Helps you iterate faster. You know, I think that is exactly where people should be, should be doing more work, because otherwise that is the bottleneck. Um, the other part there, and I think this is true of, you know, RSI for AI experiments as well, is as we do become, you know, as our intelligence levels on the, the thinking that then drives that experiment gets higher and higher, um, and as that experiment therefore becomes the bottleneck, I think we do more thinking. And this is just generally Fractile's workload bet. It's the reason we're so excited about hard problems, and it's kind of the reason that we see Fractile as the chip that takes us toward accelerating the solution to hard problems in general is I think it becomes absolutely-- it's like almost, uh, you know, a, a moral compunction to think harder before every single experiment that you fire off. Um, I think, uh, uh, it was maybe Baren Millage on Dwarkesh recently who said, you know, um, that perhaps you should think for a hundred years, you know, before you fire off an experiment and then think for another hundred years of human equivalent with these models, uh, about the results of that AI experiment, because the experiment in the middle is pretty expensive and takes, takes a certain amount of wall clock time. And so I think you get there also with things like chip design, where are the-- there are these kind of fundamental wall clock time bottlenecks for the more conventional algorithms. We're gonna be generating a lot of reasoning tokens, um, before we go and place a given circuit.
- 28:16 – 31:20
Workload Bets on Model Architectural Shifts
- SGSarah Guo
You are making workload bets and then collaborating closely with design partners, uh, on the model architectural shifts, hopefully a little ahead, so you can actually do some planning. What are you seeing?
- WGWalter Goodwin
The biggest thing that we're doing, of course, is this kind of speed chasing, right? Maximizing memory bandwidth, uh, in order to be able to run these models faster. Um, and one of the dualities that I think you have to span as a chip company is to build something that is both exceptionally good at the kind of gamut of workloads that are coming down the pipe under this idea of the kind of hardware lottery, where these will be trained on HBM-based GPUs or XPUs. Um, and so it's great that this is then the chip that's exceptional at taking kind of today's autoregressive MoE transformers and running them much faster and more efficiently. Um, but I think what's exciting about building chips with kind of fundamentally new capabilities is you then also get to explore whether there are kind of gravitational pulls that you yourself get to pull the new model landscape in because of those properties. And so for Fractile, you know, this idea of having, like, twenty-five times more bandwidth per chip than a, than an HBM-based chip, we get to explore, for instance, internally these kind of ideas of, like, scaling laws for bandwidth, basically. You know, the scaling laws, traditionally, we think of them as, like, flop scaling laws. You know, we, uh, we get better and better performance the more flops we pour into these models at train time and, and test time. Um, and if you look at just, like, the known landscape of ideas today, we already see some scaling laws for bandwidth. So one is, uh, these mixture of expert models. It's pretty well known that actually Ideally, we would make them sparser and sparser and sparser. So for like ISO intelligence, you will save a ton of FLOPS if you go from being like one in sixteen sparse on your MoE to one in a hundred and twenty-eight, one in two hundred and fifty-six. Um, but one of the challenges, one of the headwinds to doing that is actually that it becomes incredibly prohibitive on today's HBM-based GPUs, XPUs to serve those models efficiently. You end up often bandwidth bottlenecks. You end up, uh, serving at really, really low, uh, MFU, low FLOPS utilization. Um, and so there are these areas where even as you build for today's world, you get to kind of look at where you can unlock greater capability. And, you know, with LLMs today, it's very similar actually with attention. There are forms of attention that are less bandwidth hungry. They tend to be more FLOPS hungry for a certain level of intelligence. And so one of the things I think we're excited about is as we elevate this other quality, uh, this property of memory bandwidth, which hasn't really been scaled so much on chips recently. We've scaled FLOPS like a million fold in the last twenty years. Memory bandwidth has gone up about forty x in the same timeframe. Uh, so as you scale that frontier, you get to conserve more of the other thing. We get to reduce the number of FLOPS that we're using on these models to get to a certain level of intelligence. Um, and so that's the kind of, uh, really a multiplying factor on global throughput for these models.
- SGSarah Guo
I will resist asking you if that changes people's ideas about how they should
- 31:20 – 35:14
The Future of AI Chip Players and Market Structure
- SGSarah Guo
train for now.
- WGWalter Goodwin
[chuckles]
- SGSarah Guo
One last question for you just on, on market structure, uh, here. It's very unclear to me, you know, five years from now what a, uh, large AI player, um, or a hyperscaler buys-
- WGWalter Goodwin
Mm.
- SGSarah Guo
-in terms of or, or what, what chips they consume-
- WGWalter Goodwin
Mm.
- SGSarah Guo
-between the NVIDIA and AMDs of the world, um, the new accelerator class and, you know, f-four at least four of these players have their own internal efforts. Like, how do you make sense of this?
- WGWalter Goodwin
So I think that today we see, uh, you know, behavior where really everybody who is deploying at scale is kind of trying to deploy as many separate platforms as they can. Um, I think one thing that will remain the case is that there will be a real need for kind of a diversity of supply as something like compute becomes so existential to these players. There's a bit of a joke today that the sort of first-party efforts, their primary purpose is to reduce the price that people pay NVIDIA, and there may be some truth to that as well because those efforts are kind of architecturally quite similar, right? So it's a, it's a bet that is not on enabling a fundamental capability that those other chips cannot enable. And so there I think it is this kind of gamesmanship of, you know, can we build some of this over here or can I buy something from AMD so that maybe I then get a better price from NVIDIA and so on. Um, and a, and a play also for kind of aggregate capacity and a bit of control. I think as you then see this move into chips that actually give new capabilities, you know, that's a place where, you know, I suddenly see for what we're doing a real need for everybody that wants to deploy AI at the frontier. They have to have some solution to running these models orders of magnitude faster. Um, you know, for the frontier labs, there's this window of kind of premium capability, premium intelligence that is their entire raison d'être. You know, otherwise, we'd all be using Kimi models all the time. Um, and so as you do have this pressure from behind from open source, you know, it does become clearly even more important for folks that wanna deploy at the frontier to have every aspect of what that means. So the best model weights, but also then the fastest deployment, so you can do the most reasoning in the, in the shortest amount of time. And so I think this kind of premium end of, of, of the chip space where you're optimizing for speed over everything else becomes a kind of very important part of this, uh, part of this world. I think one thing that is notable about, you know, the dynamic between a chip vendor and a frontier lab, you know, if you look at Fractile, we try to be as aligned as possible with the needs of the frontier. We try to look a bit like some of those teams as well. We have people that really think deeply about these workloads and so on. So one of the cases that I think we sometimes need to make is explaining why we believe that over, you know, multiple decades the world sustains some kind of frontier third-party chip players. You know, why does this not just roll up into these, into these labs? Um, and I think that comes down to the kind of asymmetric game that all of these labs and frontier players have to play, uh, where they're exposed to an enormous risk if they go all in on a hardware bet. Um, suppose I'm, you know, I'm lab one, uh, and I've gone all in on some proprietary silicon, and then lab two discovers some new computational breakthrough that delivers far better computational efficiencies for the same level of intelligence, but it only works on the chip that they've decided to deploy. I could die in the nine months before I get to deploy enough of that chip that I've also now gained that kind of five x in computational efficiency. And so there is a really urgent need for these chips to actually deploy the same-- these companies to deploy the same platforms as one another. They're playing different games, right? They're trying to now compete, you know, in the model layer. Um, I think having deeply differentiated bets and going all in on those differentiated bets at the chip layer is, uh, sort of an irrational and a very dangerous move for anybody that is playing
- 35:14 – 35:36
Conclusion
- WGWalter Goodwin
at the frontier.
- SGSarah Guo
Well, I think that's a great place to end. Thanks so much, Walter.
- WGWalter Goodwin
Thanks very much, Sarah. [upbeat music]
- SGSarah Guo
Find us on Twitter @NoPriorsPod. Subscribe to our YouTube channel if you wanna see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.
Episode duration: 35:38
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode OpeCP4wCxkA