The Twenty Minute VCHow Many Will Actually Get Built & Is Energy AI's BIGGEST Bottleneck? | Positron AI Co-founder
EVERY SPOKEN WORD
70 min read · 13,975 words- 0:00 – 1:03
Intro
- TSThomas Sohmers
The scariest thing to me on the political spectrum is that it's now become almost unifying issue on left and right about being anti-data centers, and I think that is almost entirely a Chinese PSYOP. A single In-N-Out uses more water than the largest data centers in the United States.
- HSHarry Stebbings
We have a true expert of the space on the show today, Thomas Sohmers. He's the co-founder and chairman of Positron AI. They just raised an $875 million Series C at a $5 billion valuation. Thomas did not hold back in this episode. Incredible discussion about what we can expect in the next few months from the biggest players in this space.
- TSThomas Sohmers
I would say I'm overall opposed to basically the frontier direction it's going in.
- HSHarry Stebbings
Ready to go? [upbeat music] Thomas, I am so excited for this, dude. I said to you just before this [laughs] that I think there are some big questions that the world doesn't know, or things that they know, that I think we're gonna correct today. So thank you so much for joining me.
- TSThomas Sohmers
Yeah,
- 1:03 – 1:46
What Positron Is Building for AI Inference
- TSThomas Sohmers
great to be here. Thanks, Harry.
- HSHarry Stebbings
Not at all, but I'd love to just start, can we just have a brief description of what Positron is, and where does it sit in the stack?
- TSThomas Sohmers
Yeah. Um, so Positron's a fabulous semiconductor startup that's building, uh, hardware, uh, so really everything from the chips, the software directly running on top of that, all the way up to the full systems and rack scale deployments, um, to power, uh, generative AI inference. So, uh, effectively, everything in the hardware and sort of the low-level direct talking to hardware, software stack that powers all of the applications, uh, you know, everyone in the world's excited about right now. So, uh, everything from, um, you know, the, the likes of ChatGPT,
- 1:46 – 4:18
Why Inference Infrastructure Is Totally Different From Training
- TSThomas Sohmers
uh, Claude, et cetera.
- HSHarry Stebbings
How does the infrastructure stack required for inference, what you're working on, change compared to training?
- TSThomas Sohmers
So training, uh, I would say from, from the underlying compute level, uh, fundamentally, uh, you know, is a compute bound problem. So it's a, a workload that's the more FLOPS that you have, and, uh, if you look at, you know, be it from the regulatory and, and, uh, some of the, you know, export control frameworks are, are heavily focused on the just FLOPS, uh, required. So the, uh, how many floating point operations per second can be, can be done. And more or less, uh, you know, the, the amazing thing that the scaling laws of the past, uh, decade have shown is that, uh, the more parameters you, you add to a network, the, uh, and the more FLOPS you dedicate to that during training, uh, the better that that model is going to become. [lips smack] Uh, big difference with inference, and so the, the deployment of those models, is the fact that, um, uh, for, uh, uh, the, the actual math and the, the steps that you're doing is about half of what you're doing during, during, uh, training in terms of the, the steps not, shouldn't be thought of as, like, the actual compute involved. But the, uh, uh, what it turns out to be is that that forward pass, that, uh, inference portion of it, is heavily, heavily memory bound due to the fact that basically for every single token that's generated, every little bit of output, that requires going through the, the weights, the, uh, the, the parameters. You could kind of, you know, from a biological sense, think of the neurons. You have to read the values of that for every single individual token. And so, um, the way that, that, that's, I, I think very interesting about this and, and I, I don't know if it says anything about the, the, uh, value or, or kind of actually saying that the, the inference is, is somehow, um, I don't want to say more important 'cause of course you have to train, but, um, uh, fundamentally when you're training, you already have the corpus. You already have all the training data. And so with all that data, you can massively paralyze the token inputs, all of the, the sequences of words and sentences, paragraphs, et cetera, that are going into that. So that's something you can just crush through with a bunch of compute. But when you're inferring, because that's actually generative, you don't know what the token is, you know, five thi- five, uh, words down the line. And so you have to generate each and every one auto-aggressively or in order without, you know, foresight. And so that becomes a hugely memory bound problem that can't just be massively paralyzed
- 4:18 – 8:38
The Memory Wall: AI’s Next Infrastructure Bottleneck
- TSThomas Sohmers
like training.
- HSHarry Stebbings
So I totally get that in terms of the shift from compute bound to memory bound. Is that what people mean when they say about the memory wall with regards to what you're doing?
- TSThomas Sohmers
Partially. I mean, the memory wall as a phrase has, has been around for a long time before, you know, the, the hype around, um, [lips smack] uh, AI. And really what it's come down to is if you look at the past 50, 60 years of computing, um, we've been able to, uh, you know, have Moore's law giving us more transistors per, you know, square millimeter of silicon, um, uh, you know, consistently. Um, and while that has been able to result in, you know, greater raw compute, FLOPS, et cetera, um, uh, the, uh, improvement of the memory technology has not kept up at the same rate. So roughly speaking, you know, between, you know, 2014, you know, just very early, uh, innings of the new AI era, till 2024, um, you had about a, um, a 120x improvement in the FLOPS of, of GPUs. So a single NVIDIA GPU had about 120-fold improvement in, in, uh, FLOPS. You know, and that's what enabled, uh, a whole lot of the, the improvements over that decade. The improvement in memory bandwidth is only 17x. So I would say, like, the, the real embodiment of this is we had massive improvement on a per device basis of the FLOPS, uh, and then a whole bunch of, you know, elements on the periphery of improving the, uh, connectivity, et cetera, et cetera. But just like the ratio of the compute to memory bandwidth, uh, had this divergence. And so, um, you, you had cases where if a problem was memory bound and you couldn't just scale the compute, uh, linear with- li- linearly with that, you were getting more and more memory bound, um, as the decade progressed.
- HSHarry Stebbings
Why was there such a misalignment in the progression between the two? One's 100x, one's 17x. Why is that the case?
- TSThomas Sohmers
Well, it, it comes down to a lot of, uh, you know, technical, uh, uh, kind- ... implementation details, like the fact that, uh, if you, if you look at the lowest level, the type of memory that is used on the silicon itself is called SRAM, static RAM. And SRAM is made out of, you know, six transistors with, uh, a bit line and word line, some other, uh, control logic around it. Um, but that SRAM cell has not scaled in terms of the, the sizing of that, um, with Moore's Law over the past about 15 years. Um, so they have grown, or shrunk I should say, um, uh, very, very, mu- much slower than just a, a group of transistors that you'll use for, for other purposes. And I would say that there's just been a lot more, um, architectural advancements that could happen on the compute side, while, uh, an S60 SRAM more or less has not changed in 30 or 40 years, um, from, from like an architectural, uh, you know, primitive perspective. And so on the input side of, of like raw technical capabilities on fabrication, et cetera, have not been able to improve. Um, but, uh, uh, I would also say that there wasn't the right motivations for most of that, that decade. So with convolutional neural networks, so things that powered like AlexNets, which w- you know, really launched the deep learning revolution in 2012, and then ResNet and, uh, you know, all, all of the, I would say, the, the advancements during the, the 2010s, um, uh, was in the realm of, uh, machine learning models that were fundamentally compute bound. You could just throw more and more flops at CNNs and, uh, get better results. Um, and you didn't really need all that much, be it memory capacity or memory bandwidth. Um, but it was really with the transformer and even though the attention is all you need paper came out in, in 2017, I would say it did not really get the, um, uh, attention, you know, pun intended, uh, it deserved until, uh, uh, 2020. Well, uh, uh, GPT-1 and GPT-2 came out prior to that, 2018 and 2019, but it was really GPT-3 showing that, okay, you go from a billion-ish parameter up to 175 billion parameters, and you actually get this, this improv- this massive improvement in capability, and that's really where I would say the transformer revolution started. And most people didn't really catch on un- to that until the end of 2022
- 8:38 – 10:32
The Hidden Economics Behind AI Tokens
- TSThomas Sohmers
when, uh, uh, ChatGPT came out.
- HSHarry Stebbings
When you look at token economics and token efficiency today, what does, like, no one know or talk about that you think should be much more front and center?
- TSThomas Sohmers
I do find it, like, compared to a year or two ago, there are now different prices listed for cached versh- versus uncached tokens, but I don't think people realize how any providers that charge the same amount, and even with a, a lot of people's cached prices, how high margin that is. It's, it's, like, insane. You, you make all of your money on selling cached, uh, in- input, uh, uh, and output tokens. So, uh, as a, as a provider-
- HSHarry Stebbings
Wh- why is that, sorry? Just so I understand that.
- TSThomas Sohmers
Oh, because for, like when we discussed earlier that, uh, that a serve-- b- both processing a cached token is essentially free. It's, it's o- one one-thousandth of the cost. Yeah, I'm, you know, order of magnitude of, of, uh, uh, you know, actually ge- having to generate to, to recompute and, and generate that token. So, um, there, there is so much, uh, you, you can juice out of, out of, uh, selling those, those, uh, cached tokens. Basically all the providers, they charge you to cache a token. They charge a higher rate than just, um, you know, the normal processing fee for, like, an input token. Uh, and then they charge you a lower rate when you read from that, and it's great when you're paying that lower rate, but they're making obscene margin on that, uh, that, that, uh, cache read. Um, uh, and yeah, it's, um, uh, there, there's a reason why, you know, Anthropic is being, is reported to have, you know, 80 points of gross margin right now on, on their, uh, you know,
- 10:32 – 12:12
Why Anthropic Could Already Be an 80% Gross Margin Business
- TSThomas Sohmers
API business. Um, and-
- HSHarry Stebbings
Were you surprised by that 80 points? I was impressed.
- TSThomas Sohmers
Not really. I, I, well, uh, I'm impressed by 80 points of margin in basically any industry. It's, it's difficult to get that margin and great thing about capitalism is we will, uh, tho- those margins will compress with competition. So I'm, I'm confident and happy for that even though those people are, you know, theoretically my customers and my margin is sort of based on their margin. But I, I care more about, uh, you know, a healthy ecosystem long term. Um, I would say that the-- i- it's more surprising to me how many people still today think that these are horribly unprofitable businesses and that the, the whole market's going to zero. It's like, it, it's, it's absurd to me that the, the meme of these, you know, OpenAI, Anthropic, et cetera, are just burning cash and eventually they'll run out of cash that they can burn. Like, if they stop training, they'd be massively profitable overnight. Um, and there's a ton of other levers that they have without, you know, pacing the frontier as, as Dario just had in his, uh, hi- his essay. I, I mean [laughs] I, I ha- jokingly think that, you know, a, a little bit of the pacing the frontier discussion is, "Oh, this is a great way to, to reduce costs, uh, ahead of IPO." Um, but, uh, uh, but I, I, I don't think Anthropic or, or anyone needs to do that. I think they're, they're amazingly profitable businesses with their scaling rates and without-
- HSHarry Stebbings
And they would reduce costs, just so I understand, 'cause they would spend less on training because they would be slowing-
- TSThomas Sohmers
Yeah
- HSHarry Stebbings
... down the speed.
- TSThomas Sohmers
Yeah.
- HSHarry Stebbings
Right?
- TSThomas Sohmers
Yeah. So I, I don't think that's actually the intention or anything. But, uh, um, but, but yes,
- 12:12 – 13:59
Should Frontier AI Labs Actually Slow Down?
- TSThomas Sohmers
the, that's the bulk of their costs.
- HSHarry Stebbings
How did you analyze the pacing the frontier? You brought it up. How did you analyze it?
- TSThomas Sohmers
I have, I have mixed feelings on, on the safety topic. I, I am a human, so I, I, you know, would like to, uh, uh, live old age, [laughs] uh, and, and more, more so than that, um, have, have humanity, uh, uh, continue to the stars and beyond. Um, but, uh, I believe much more in the ability for this technology to revolutionize every part of, of, uh, humanity, and in, in a positive way. And, um, I do worry that pacing in, in a lot of the ways that's being talked about, not necessarily how it will be implemented, has two big risks. One, a major, um, pause, you know, at stopping and, and sort of that playing into, for lack of a better word, the Luddite, um, sentiment that exists. And that by pushing for pause, it's actually giving ammunition, giving, um, uh, better basis for those that actually just wanna stop the technology over completely, and I, I see that as a major risk for humanity. Um, the second piece is I'm very, very... Like, my actual P doom, my... The thing that I'm most worried about of any, you know, AI outcomes is that te- that, uh, technology and capability being concentrated to relatively few people. Um, and having a lot of what's being discussed from a regulatory framework and, and limitations on, on technology, e- et cetera, I, I think is, like, the modern,
- 13:59 – 17:07
Why AI Regulation Could Create a New Class of Technology Monopolies
- TSThomas Sohmers
you know, road to serfdom. It's like the concentration of technological capability and, like, making legal to do matrix multiplications is, like, the thing that will set us back to, uh, you know, pre, uh, uh, uh, n- not just industrial revolution, [laughs] it's like pre, um, uh, Enlightenment, you know, uh, capabilities. Like, that, that is like the biggest attack on, uh, classical, like, liberal freedom, uh, concepts that I, I can, I can think of. Because while I do think that the vast majority of tokens are going to be produced by the big, you know, players, if the technology itself is restricted to just those, then they're going to be the new, you know, uh, lords and, and kings, and everyone else is, is back to serfs.
- HSHarry Stebbings
With the greatest of respects, is it not just lip service? Great, we'll stick Meta in the corner. They can do their compliance, and then we can IPO. Sam can have a reason not to IPO 'cause his numbers aren't as good as Anthropic's. It plays into both our, uh, desires.
- TSThomas Sohmers
And Elon wants time to catch up as well, and there's, there's plenty-
- HSHarry Stebbings
So, so it plays into everyone's.
- TSThomas Sohmers
Yeah. I, I, I completely agree, and I, I think that is the biggest internal reason for probably everyone other than Dario. I, I... Like, Dario and, and I would say the vast majority of people in Anthropic are true believers both in all of the promise and capabilities of the technology and the risks. And, uh, I mean, if I were in their shoes, I would also be, uh, take it, the, the massive amount of responsibility and, and, and, you know, for that. Yeah, there, there's a lot of strategic reasons of saying, "Okay, by having these auditors, et cetera, that removes some potential responsibility, culpability from, like, a legal perspectives, et cetera." The risk that I just don't think that any of... Like, I think this is a problem with a lot of very smart people, especially when they've amassed large wealth and power, et cetera, is they think that they're going to be able to keep that. And the, the scariest thing and kind of my, my point, like, the centralization of technology, like, if it gets concentrated with companies, governments, et cetera, um, is you've got people that think that they're the smartest people in the room, not realize that they are not going to be the ones to actually control it when they put these measures in place. Like Dario basically, on one hand, verbally begging for government, governments to take over Anthropic. He's in part saying that 'cause he doesn't think it will actually happen, and I would love to see his reaction if and when that actually happens and he realizes, "Oh, shit. Um, I thought that if I, you know, was begging for regulat- regulations, they would then make me the regulator in some form or fashion." And when that doesn't happen and it just becomes a, um, bureaucracy that halts all progress, and the capabilities that currently exist basically get squandered to select bureaucrats, um, that, uh, that's, that's the
- 17:07 – 22:44
What Happens if the US Paces AI and China Does Not?
- TSThomas Sohmers
worst outcome I can imagine.
- HSHarry Stebbings
I'm very naive. I'm a podcaster and a venture capitalist for a living, so one of the lowest IQ on the spectrum.
- TSThomas Sohmers
[laughs]
- HSHarry Stebbings
Uh, when you consider the advancements that China are making, especially with their open ecosystem, but they are an incredibly talented ecosystem right now moving forwards. Um, if we pace and they don't, what happens then?
- TSThomas Sohmers
I, I guess when I said that, "The worst possible outcome," I wasn't counting that as... Uh, I, I wasn't counting the Terminator outcome, and I wasn't counting that. So, um, I, I think of course everyone can agree Terminator or similar, or similar is, is very bad, but I think is extremely low probability and, and just... I, I'm, I'm not a believer in that, that sort of doom scenario. For the vast majority of people, it would result in the same level of serfdom that I worry about with the scenario that I, I described, would be significantly worse for some number of people in a, you know, Chinese, uh, CCP-controlled, you know, super intelligent AI scenario. Um, I think that, you know, on one hand, their strategic angle right now is- ... per have technology proliferate through open source, et cetera. I think as soon as they get into, you know, pole position, the, the ladder gets pulled up with them in some way. I d- I don't think they actually want the technology to be, you know, easily accessible to everyone. Now, I don't know if they will decide that it's okay if the rest of the world has some access to the technology, but they definitely will not let the billion people, uh, that are not CCP party members, uh, you know, benefit equally from, from the technology. So-
- HSHarry Stebbings
So just so I understand, d- do you agree with it? 'Cause to me, I just didn't get it.
- TSThomas Sohmers
Oh, yeah.
- HSHarry Stebbings
You can't, you can't pace the frontier unless the global AI community paces the frontier, and I don't see Putin signing up.
- TSThomas Sohmers
[laughs] Agreed. Um, and y- yeah, I, I think th- this is a little bit the same naivety that, that I described by these company leaders and, and in general people in the Western world thinking, "Oh, we're so great, we're so advanced, so far ahead that we can't get caught up to." I mean, on paper, is the US the greatest military force in the world? Yes. Um, if we had to all of a sudden have a drone incursion the same level of, you know, what's happening from, from, uh, uh, in, in Ukraine, Russia, you know, coming up from Mexico. And if you take Mexico, let's say that they developed, you know, very naive drone technology e- and et cetera, like on the level of what's happening in Russia, Ukraine, and, and Iran. How would we respond to that as a country if we had that coming up across our border? Like, I don't... It doesn't matter our amazing military might. I, we, we've, we, we built our military to fight the last war, and I think geopolitically our thinking is, "Oh, we're the big dog still" and that, uh, um, when it comes to AI technology, there's not the acceptance that, um, export controls and all of the other elements that theoretically would allow us to pace and have people keep paces behind us just aren't good long-term solutions.
- HSHarry Stebbings
Do you think we should have export controls?
- TSThomas Sohmers
I am very much a strong believer in free trade and, and free exchange of ideas. I... The, the exception to that is, uh, you know, a little bit I think China has been a free rider of all of the benefits of, you know, a liberal free war- free trade order for the rest of the world, while they get to keep everything closed off. So I am very, very happy and th- and think that any government societies, people that want to embrace free exchange of ideas and trade and everything else, we should have a very vibrant, uh, e- economy and ecosystem. Um, but totalitarian regimes should not be able to participate with that, especially in the, the case where they get all of the benefits of that and get to export themselves, uh, you know, things that make them better able to, to have that totalitarian, um, uh, you know, system keep up.
- HSHarry Stebbings
Can I ask, we, we mentioned, you know, the pacing the frontier and, you know, the different people who supported it. You had Zuck and, um, Jensen say nothing. Well, Zuck actually come out in opposition to it saying that we should continue as planned. What should we take from those two seemingly silence in opposing it?
- TSThomas Sohmers
Based on what I, my, my overall beliefs right now, as probably evidenced by the, the conversation so far, I would say I'm overall opposed to, you know, the pacing the frontier direction it's going in. I appreciate anyone that is adding to the discussion that is, is I think being realist about the, uh, benefits and risks and... But you always have to take that with a, a grain of salt of what are the motives of [laughs] uh, of anyone that's, that's discussing in it. And I would say I probably appreciate Zuck or, or Dario's comments infinitely more than a random politician. Um, and not just random, uh, the, the quote unquote "leading politicians" that, that don't actually understand the technology. And I would, I... The scariest thing to me on, on the, the political spectrum and the way all, all of this is being treated is that it's now become a almost unifying issue on left
- 22:44 – 28:24
Is Anti-Data Center Sentiment Becoming a Strategic Risk?
- TSThomas Sohmers
and right about being anti-data centers, and I think that is e- entire, like almost entirely a Chinese psyop.
- HSHarry Stebbings
Can I ask, why is it a Chinese psyop being anti-data centers? 'Cause like it does increase... I, I, I'm totally with you in the benefits of them, and, you know, Gavin Baker said it, how they're the greatest economic kind of needle mover for large parts of the country. I guess they see increased electricity prices, increased water prices, and ugly data centers in their backyard. Why, why, why is it a Chinese psyop? What, what am I not seeing?
- TSThomas Sohmers
Um, well, just on the ugly and all that, I, I totally am in support of a, uh, a, we, we, we need to have beautification campaigns and, and really turn them into, uh, centerpieces of, of our society. I think of thousands of years in that from now, um, you know, a future history looks and they should see these massive data centers, like the, the really, really massive impressive ones, uh, should have a, you know, be like the great pyramids or, or [laughs] something like i- i- they... We, we need to dress them up to be the, the world wonders that, that they, they are technologically. The thing from water usage and, and, um, like the amount of power they consume, et cetera, um, like so much of the early information that went out by unsophisticated writers or things that are just patently false, like the, you know, a single In-N-Out uses, you know, more, more water than, than, uh, you know, the largest data centers, uh, in the United States. And like it's, you know, go- all golf courses are orders of magnitude more. Like these are closed loop, you know, liquid cooled systems that, that, you know, uh, you don't even want to use water in a lot of these cases. And from the Chinese psyop perspective is they're not, to, to your, your point, they're not p- pacing the frontier. They're adding, you know, gigawatts of, uh, um, you know, new electricity generation capacity, most of it being dirty. Um, uh, they're, they're- Building massive new data centers, horribly displacing people. Like, it just, it, it irks me so much that we have the freedom in the Western world to criticize companies, governments, everything, and slow down and based on false information, and I lo- I love the freedom elements of that. But it is a strategic disadvantage when China can just say, "Yeah, we're going to just bulldoze all these people's homes and, and do rolling blackouts, um, wherever in order to, uh, uh, serve, uh, the greater good of, of, uh, new, new training capacity."
- HSHarry Stebbings
I mean this with the greatest of respect, but I don't understand how anyone thinks the US or Europe can beat China when they have no regulatory or policy re- restrictions. And I mean, the UK, you can't, you know, put up a paper airplane without getting a permit. So, like, we're, we're fucked. But you are getting there, and you're getting, you're becoming a European state in terms of the regulation and policy requirements. Am I wrong? Am I being overly negative? I, I didn't get it.
- TSThomas Sohmers
No, I, I, you're right. Um, I think the greatest advantage the US has in that regard is that there's still a lot of land, a lot of places that, um, do not have all the same levels of restrictions. I, I don't agree with a lot of things of most pol- you know, administrations of my lifetime. The current administration gets attacked for saying, like, they're destroying our, our environments and, and, you know, destroying national parks, et cetera. Like, the vast, vast majority, 90-plus percent, I don't know the exact numbers, of national federal land is just open, empty desert in, in the West that is not part of a national park or anything. And the fact that there are so many restrictions to utilizing, you know, BLM land for building data centers where it's literally hundreds of miles from, you know, a pers- from, from any, you know, populated area. My great state of Nevada, um, has plentiful geothermal, solar, um, all these green energy technologies, and we could, you know, build, you know, nuclear and other things in the middle of the desert where it won't impact anyone, and that there's restrictions to that is completely absurd to me. And I will say, there, there is the, there, there has been some political will and, and push to solve these things, but literally just in the past year, you have, you know, Republican governors and, and other politicians that at least had part of their platform to be pro-growth and all of these things backing away because they see from their own, you know, political base being anti-data center based on completely false, uh, premises. And, and one of the points I wanna go back to that you brought up was, like, that people would have higher electricity costs. If we incre- Like, this is the most basic supply and demand. If we increase generation capacity, and no one's saying we wanna be taking energy from what's reserved for people's homes. Like, one of the regulatory problems I see is, is power companies have to have this, um, buffer of, of energy availability, uh, that, that's is baked into the cost and, and capabilities for everyone. Um, there, there is absolutely zero cases where a data center could be potentially pulling power from anything that's already been allocated. So that, that's just impossibility. And all these data centers that are getting built right now are coming with generation capacity that covers their own use and b- beyond that. And we're just not allowing them to hook up to the grid [laughs] uh, where they could actually be lowering the prices for everyone. And then you've got people on the power company side, they're lobbying to, to, against new generation capacity, 'cause that will actually, you know, market forces will, more, more capacity will, will decrease prices, which would be good
- 28:24 – 30:46
How Much of the AI Data Center Boom Actually Gets Built?
- TSThomas Sohmers
for consumers.
- HSHarry Stebbings
What percentage of data centers that are planned will be completed, do you think?
- TSThomas Sohmers
From the major providers, I would say the, the capacity that they have planned, they may be in different locations. Um, I mean, you, you've had some local communities that have, uh, successfully, you know, stopped, um, facilities going in there, but then tho- those data centers just move. I don't think a year ago the major data center builders and operators were thinking that the political problems were as bad as they, they were. Um, and so there is a lot more effort being put into education in, in those communities now, which I think will s- you know, turn the tide a bit. But I mean, it's also just gonna mean, mean that those data centers move to locales that aren't going to have those problems as well. And like I said, we've got, uh, uh, large tracts of land that, uh, can, can support it. So, um, so I, I, I'm, I'm not too worried that it's going to be, like, an existential threat in, in capacity build-out. And then of course there's space if, uh, Elon's successful.
- HSHarry Stebbings
Do you believe that space is a viable alternative truly, or is it conference talk and lip service to justify a, a market cap?
- TSThomas Sohmers
I think something can start as one thing and turn into something else. Um, I, I would never, ever bet against Elon. I, I mean, I, I, I primarily bet, I bet for Elon. If you asked me a year ago, I just would not have thought that there would be a good reason for it in the near term because it's going to be cheaper, easier, et cetera, to build on land. I also think there's great alternative technol- technologies. Company we're partnered with, and I'm good friends with, uh, the, the CEO, um, is, is a company called Pantalaossa that's building ocean-based data centers. Basically a very interesting pumped hydro solution in the middle of the ocean. So there are alternatives that don't require going to space, I think long term. Part of the, the reason I'm long-term big believer in space data centers is I just think we're going to need to have a space economy for humanity to live up to its long-term potential.
- HSHarry Stebbings
Love that. Totally agree on never bet against Elon. If we think about, like, the cost of intelligence being tied to the cost of energy, how should we think about energy as a bottleneck moving forwards? To what extent is it, we mentioned policy and regulation being a core bottleneck. Is energy a bottleneck moving forwards or
- 30:46 – 32:41
Is Energy Really the Biggest Bottleneck to AI?
- HSHarry Stebbings
less than people consider?
- TSThomas Sohmers
I think there's two pieces to it. I mean, one, you know- Positron is trying to deliver, you know, more compute, more capabilities per watt, per megawatt. Um, and so, uh, you know, sort of on our base case, um, if, if we can turn a, what you would've spent 500 megawatts with Nvidia equipment and do that on 100 megawatt, I don't think that's actually going to, um, mean that you're only going to build a 100-megawatt facility. You're still going to build the maximum amount of compute that you can. You're just getting more tokens, more intelligence per, per, uh, joule. If, if I go back to the long-term thinking, [laughs] you know, assuming humanity continues for hundreds, thousands of years, everything turns into an energy problem, where we're, we're ga- And, and you can go back thousands of years and, and just look at the progression of mankind. Fundamentally, that is a perfect track of our ability to produce and use energy, you know, discovery of fire up to, uh, you know, nuclear power plants. The, the simple, um, uh, tongue in cheek answer to your question is everything, uh, all, all progress is gated by, by energy. And even if there's energy available, it may not be economical, and so it won't be done. So a- actually, I would say the bigger limiter than just saying that energy, like our ability to build and produce energy, we've got plenty of technologies and capability to do it. I would say we have got way more economic limitations. It's like how much, how much debt is the world [laughs] willing to take on to build all out everything over the next couple years, is, ties into energy. It ties into the infrastructure itself, et cetera. So I, I think economics is a much, um, uh, easier sort of scapegoat. [laughs]
- HSHarry Stebbings
I mean, people are already very concerned by the levels of debt being taken out and the debt cycle. Do you think their concerns are justified,
- 32:41 – 43:16
Why Sovereign Debt Is a Bigger Risk Than AI Infrastructure Debt
- HSHarry Stebbings
and then do you share them?
- TSThomas Sohmers
I think we've got a major sovereign debt problem that, um, masks a huge amount of, um, second and third, third order, uh, elements in the financial system. Just the, the inflationary consequences of, you know, [laughs] the government that, that can print, you know, infinite amounts of its own currency. Um, and the fact that we are, as we're already seeing the treasuries and, you know, the greater bond, bond markets, that there is greater and greater perceived risk of the most, quote-unquote, "risk-free asset," um, I think will, uh, trickle down to all elements of the financial system. Um, and so like when people worry about Oracle's, uh, debt and, and credit rating, um, I'm like, I, I believe in Oracle's business model and, and ability to execute and, and do everything a whole lot more than United States government. It's just the United States government can, uh, issue its own currency and, um, also has guns and nukes to, uh, take tax revenue. So my biggest economic concern there is that there will be a, a more acute, um, specific, uh, crisis that, that arises out of the compounding of, of national debt leading to, uh, devaluation of the currency that has all of the, uh, consequences downstream. Um, rather than like I, I, I, I'm really not worried about any of the companies in the AI debt stream, like not hitting their revenue targets. Like if the past three, four years have shown, we're accelerating every aspect of these businesses in terms of revenue profits and, uh, uh, how they are ac- improving the productivity and value down- downstream.
- HSHarry Stebbings
I'm jumping around, but fuck it. When I was doing the research, I was reading about KV caching and compression as part of this, and I was honestly getting lost. But I was intrigued and digging deeper and deeper, and I was like, "Why did I not know this before?" And so I don't think many will know this. What should we know about KV caching? Why is it important? Can you explain it to me a little bit?
- TSThomas Sohmers
There's, um, always a, a, I would say a, a, you know, give and take, uh, relationship with, with innovation. I guess one element I'll, I'll just have to explain to, to make all this, this clear is the concept of a, a sequence. I kinda already talked about a, a, a token. Um, but, uh, you know, just to, to define things. You know, a token, um, effectively as part of the training process when any of these big, you know, model apps are, are developing a new model, they have a vocabulary that they define. So they take their big, giant corpus and they do some statistical work, um, to figure out what is the best encoding method to take all of the text in this and, uh, break it into chunks that get reused frequently to, to, uh, have things be more efficient. And so what you end up doing is if, if you o- if you took a, a English dic- dictionary, you'll find that there are common, um, prefixes and suffixes and, and, you know, groupings of words. And if you just try to think of how would I best, um, uh, compress th- this, uh, if, if I just had symbols for these prefixes, suffixes, et cetera, um, compress this into a thing. And basically what, what ends up happening is a token, um, you know, if you're using ChatGPT and you see text streaming out, if that's going particularly slow or you, you've quite, uh, a keen eye, you'll see that it's portions of words that come out at times. Sometimes a full word, sometimes a, a small fraction of word, and each of those little flashes that you see is a token. And, uh, roughly speaking, it's between half and, and, uh, a, a token is equivalent to a half to, like, 75% of a word on average in, in large English, uh, corpuses. So that's token. A sequence or, you know, the thing that builds up to be in context in, in a model is the grouping of all those, those tokens in, in an, an order. And what happens when you're, you know, running an inference, you, you're given a prompt. You have, uh, you know, what is the capital of France as, as your input prompt. Um, that is tokenized, you know, that is, you know, four or five, six tokens, uh, and that, that go into the model. And when you do inference, it's going to say, "The capital of France is Paris," and, you know, the, the City of Lights, you know, some, some, you know, uh, uh thing after that. And so when you have that entire sequence, when transformers originally came out, for every token that, that was generated, you were doing the computation for generating all of those tokens, including the ones that you've already processed. And the, you know, clever, I would say, kind of obvious based on, you know, all the de- developments in the past of, of computer history but wasn't done initially, was that, well, you don't actually have to redo the compute of the things that you've already, you know, had as inputs and, uh, what you've already generated in this turn. And so, um, the KV cache was born, where within the model, there's these two matrices called K- K and V, keys and values, and those matrices, um, are, are fully based on the, uh, prompt and whatever is generated during a turn. And so by actually storing those two matrices, you can avoid having to do-- redo computation, um, at the cost of now having to store this thing in memory. And that's, you know, the simple example, uh, you know, very, very small, you know, kilobytes, uh, of, of hundreds of kilobytes of data. Um, but the, the thing is that these things grow with, with the sequence length. And so, um, but the, the interesting thing is that for the attention mechanism, the, the compute for, um, uh, per token grows quadratically with the, uh, with the sequence length. So you're having to, to spend more and more compute quadratically, so it's, you know, an exponential curve, um, uh, for a- as sequence length grows. But, uh, when you store that a- a- as just your K and V, that's just a linear growth. And so you're, you're really trading off, like, what KV caching does is means that you don't have to do that compute, which gets very expensive very quickly, quickly, at the expense of needing to store these things. And, uh, storing that, uh, is, is a complexity in itself because that's a unique KV cache for every single user that you're serving. And it comes questions of how long do you wanna keep that for, how long, uh, you know, and how do you manage all of that in a large system?
- HSHarry Stebbings
So what does it mean then when we hear about compression and uncompression of KV caching and potential entropy within the system?
- TSThomas Sohmers
There's two different forms of compression. I'll, I'll... I have a couple more than that, but the two main ones. So one is quantization. So, um, one is-- So the short form of quantization is if you've got, um, each of your values, be it your weights, your KV caches, um, activations, um, stored in a particular data type. So before the machine learning at w- uh, revolution, you know, most of the world's, you know, com- computation was done in FP32. So you have 32 bits to represent a floating point number, and that's broken into, you know, uh, mantissa and exponent. Well, it's pretty quickly realized that having 32 bits of precision was super overkill for the things that you're wanting to represent and, you know, it's both s- costs more from a storage and co- computational perspective than lower precision. So we went to FP16. You know, Google developed BF16, you know, a little rejiggering of those bits. Went to FP8, now we're at FP4, you know, in, in popular, um, systems. And so we've been reducing the precision quite quickly, um, but that does lead to, you know, for lack of better words, some brain loss, um, when, when these models, uh, run, just because you are now trying to encode the, the same information into fewer bits. And so there's been a lot of interesting schemes to say, "Okay, I'm going to take this group of 16, um, FP16 values or BF16 values, and I'm going to quantize those. So I'm going to, um, you know, use a, uh, a truncate and rounding that, that down to, let's say, uh, int4 values." So now you, you actually saved 75% of your, uh, uh, total, uh, size of, of that group of values. You shrunk that down from 16 bits to 4 bits. Um, but just doing that naively will mean that on a lot of benchmark scores, you'll have them go, you know, get 20, 30% worse. So you, you get that 75% savings in, in space, but, um, uh, you know, p- uh, y- you kind of lobotomize the model. Um, now, but, you know, advanced quantization techniques actually say, "Okay, these 16 values, I'm able to have a, you know, shared, um, uh, mult- a bias or a mult- and a multiplier for it." So y- let's say for those 16 now int4 or FP4 values, um, you store one new FP16 value that gets applied to all of those at comp- compute time. So you get a 75% compression on all those values at the cost of now adding to that one new FP16. And basically, the state-of-the-art here is you're able to get things compressed from, you know, FP16, 16 bits per, per value, down to, like, four and a half bits per value. And that can be applied to weights, the, the, you know, actual parameters and model, that could be applied to the KV caches. Um, uh, but, uh, you know, there is, there's no such thing as free lunch. You, you do, uh, uh, still have some lobotomy, but, uh, thankfully it's kept within, like, 1%, um, of a unquantized model.
- HSHarry Stebbings
Is, is KV caching the hardest element of building that inference infrastructure or is it, uh, you name it, uh, latency SLOs or load spikes or anything else that we could come up with? Is that the hardest? Like, what
- 43:16 – 49:32
Why KV Caching Matters So Much to AI Economics
- HSHarry Stebbings
do we not see that we should see?
- TSThomas Sohmers
You, you can run a service and do something without having KV caching at all. You're, you're going to economics and performance and everything else why it's gonna be much worse. Um, the, the dark arts and magic with it is the workloads that the industry so far has found the most valuable happen to be very, very highly cacheable. So, um, you know, semi-analysis has their, uh, Agent X, uh, benchmark, um, and, and, you know, suite of, of test data based on taking a whole lot of, uh, Claude code sessions and, and having dozens to hundreds of turns in, in those Cl- Claude sessions with sub-agents and everything else. And what they found is over these massive number of interactions of these, like real traced, um, uh, code generation, you know, agentic- agentic coding sessions, about 96% of all the tokens that go through these entire sessions are cached. If you know your workload is going to have this extremely high caching rate, where you're going to be reusing the same tokens again and again, that drastically shifts the, um, importance of, um, how you can retrieve those caches because it's, it's not just-- Like, and these things get to be very, very large. Like we, we, you know, have gone into trillions of parameters. So if we just take, um, you know, the GPT-4, you know, ki- got leaked as, you know, 1.8 trillion parameter model. Now, assuming that that is int4 quantized and, and rounding down a little bit, you know, that's 900 gigabytes of, of data size for, for the model weights. Um, if we take like the high expectations of like Claude Fable, um, you know, that's a 10 trillion parameter model, so around five terabytes of model weights. But the crazy thing is, at these long context links for these size models, you have the individual user sessions being in the, uh, let's say in the 100 gigabyte range. So with just 50 users on your service, the, the user context, they're just those individual sessions end up being greater than the model weights that you're trying to store. So that's, that's, uh, uh, you know, Claude and OpenAI have a, have a wh- whole lot more than 50 users. And so it becomes a really interesting, um, trade-off of, okay, how much of the, uh, you know, accelerator memory do you wanna dedicate to, uh, weights which you need to process every single token generated, and you want that to be as fast as possible 'cause that sets your SLO, that sets the, the token latency. But if you don't have their KV caches persistent, you're actually losing a huge amount of efficiency because that was work that you didn't have to actually repeat. And so it, it saves you as an operator money more than anything. So at, at some level, you know, having, uh, users' KV caches be persistent will give some level of speed improvements to-- that the user perceives, but it's mostly an economics thing for the service provider, where if you can return to them and, and use those, those tokens again and again, um, that saves you money as an operator massively.
- HSHarry Stebbings
Totally get that. It saves us money 'cause we don't have to use as much compute, but then it's harder from a memory challenge perspective. How do you think about the right logical next step then? If you appreciate the importance of saving on compute, but the challenge of memory with KV caching, uh, what's the answer then? That we just have bigger and bigger memory stacks on chip? What does that look like?
- TSThomas Sohmers
Most common deployed solution and, you know, the vast majority of, of inferences out there are taking place on GPUs, you know, followed up by TPUs and, you know, a couple other devices. But, um, most common paradigm today is you've got your GPU accelerator memory that is primarily responsible for holding the weights, and you will keep some number of user sessions on there, the ones that you're actively processing. But, you know, the, the larger group of users has a tiered hierarchy. So you'll have users that were around, say, in the past, uh, you know, couple of seconds, uh, that but haven't returned, don't have an active request. That's residing in host memory. And let's say that's on the order of, you know, um, anywhere 4 to 10 X more memory on the host than in the accelerators, uh, so you'll be able to store more, uh, of those there. And then if someone hasn't been around in a couple of minutes, maybe a couple hours, that's going to be stored in even further away memory. So that could be in, in NVMe, so, so flash storage, so a lot slower, but a lot larger capacity on that host. It could be in flash storage on a network attached, you know, drive. And eventually, like I bet the ChatGPT sessions that I had, uh, you know, six months ago, um, you know, somewhere residing on a, on a, you know, disk, you know, slow SSD or, or something, uh, um, uh, you know, somewhere [chuckles] in a, in a data center. Um, but it would be dumb for them to use expensive memory to, to store that. So, um, that tiering is, is the norm, but that introduces a huge amount of complexity of how do you decide when and where you're going to store something, um, for your massive number of users. I would say our solution, kind of how we're trying to go about it, you know, both from our expectation that model sizes are going to drastically in- increase, the number of users for all these things are going to drastically increase, and the context themselves. Like two, three years ago, you know, the, the, you know, typical context lengths were on the order of 8,000 to like 64,000 tokens. Um, then it got up to 128, 256. You know, a million token context lengths are the norm now in terms of what the models support. Um, but a million token context lengths- Can only hold, you know, a portion of some of, like, our internal company's, like, largest, uh, code repositories. Like, it will, it will be a fraction of that. And so if you really want agent that can, can take over, you know, the, the capabilities of a whole team of programmers, I think the main limiter today isn't, like, the model capabilities itself and, and scaling the model size. It's on how much context can that, that model have of all, all of the, uh, uh, data it needs
- 49:32 – 51:14
Will Frontier Models Keep Getting Bigger?
- TSThomas Sohmers
to make smart decisions.
- HSHarry Stebbings
I just wanna break some of the things you said out there. You said that you think model sizes will increase. I thought we were all moving to owning our own intelligence, every enterprise having their own smaller model with proprietary data. Does that go against what you think in terms of model sizes increasing? Can you help me understand?
- TSThomas Sohmers
Yeah. I, I think you can kind of break it into two tiers, uh, again. So there's going to be the frontier models and capabilities that are being really at the, the forefront of the development by OpenAI, Anthropic, maybe Google, uh, you know, SpaceX AI, et cetera. And I, I still think there's a long road to go in terms of getting to, you know, uh, pushing the frontier of, of model capabilities, and those, those will continue to grow, continue to get, to get better. And there are a lot of workloads where, uh, let's say internal, just speaking for how Positron uses LLMs, I don't today care that much about the cost. I'm, I'm ... If, if I can get 10 times the, the output, uh, you know, value out of a model today, I very gladly pay 10 times more, you know, per token, and I really want the frontier to push that. Um, I think, you know, some of it is cost saving, some of it is just owning, you know, truly owning your proprietary data. Um, there is a push, you know, uh, from, from enterprises to, to have inference on site and, and it's, it's a lot more difficult to provide that for largest models. And most companies, you know, if, if they're adapting from open source or developing their own model, don't have the resources to, you know, be pushing the frontier. And so, um, that, that is kind of what's
- 51:14 – 54:04
Will Enterprises Really Own Their Own AI Models?
- TSThomas Sohmers
gone to smaller models.
- HSHarry Stebbings
And do you not buy that reality?
- TSThomas Sohmers
I would say, like, right now, somewhere around 80, 85% of all tokens consumed and produced are done by just the, the, uh, top four, uh, models, uh, model companies. Um, and so, um, and I would say the next, you know, 5 or 10% is done by the three or four after them. So I can totally buy, believe that 5% of all tokens consumed will be done by things on-prem, you know, not, you know, locked into, to the big guys. But, I mean, I'm both from Positron's business perspective and just, like, how I see the world evolving, I, I'm going to care more about the, the high volume set of things. That, that being said, so I actually think the, the small model stuff is actually much more interesting for everything happening on, on your phone. Um, and, and the amazing thing about that is I think, like, there's this misconception that, oh, if, um, questions could be answered by your phone or, or, you know, any prompts can be, can be done locally, um, that actually is meaning that there's less tokens going to be used with the big guys in the cloud. Um, I, I think it's the opposite. The, the reality is that if I have an LLM running on my laptop or phone or in my enterprise's, you know, secure on-prem cloud, whatever, um, that is going to be consuming data at such a fast rate of everything coming into it, uh, and it's going to be generating, uh, you know, analysis based on that. And it will decide, okay, what is it that I'm going to actually return to the user, uh, you know, locally, um, and what is it actually requires more intelligence from a better model that's isn't self-hosted. And so, uh, I think for, you know, any of the tokens that are being, you know, quote-unquote saved by running locally, that's actually going to generate more things. 'Cause, I mean, in, in some ways, for, for a simple naive use case of a, a personal user of LLMs, they're only going to prompt ChatGPT or Claude every so often. Like, they're, they're sort of limited by, um, their thoughts of, of when to actually ask an LLM something. But if they have a local LLM that is constantly checking their email, their calendar, messages, et cetera, and deciding to do these, uh, lookups to, to cloud-hosted models frequently, that's now on, on a per person basis a massive increase in the nu- number of tokens being consumed and generated by the cloud models, even though there was, you know, the, the naive view is that, oh, there's this shift to this on-device, uh, LLM.
- HSHarry Stebbings
So can I just understand? So I completely hear you in terms of maybe 5% of them will be in this smaller model, enterprise-owned kind of model landscape. Why are you so bullish then on much larger models
- 54:04 – 55:08
Why Scaling Laws May Still Have a Long Way to Run
- HSHarry Stebbings
and the sizes increasing?
- TSThomas Sohmers
I, I'll, I'll break it into portions. So there's the si- increase of the model sizes which, um, I think of [laughs] the, the, you know, thing that would be shared by a lot of people in the AI space is it's kind of a gut feel, um, where, uh, based on the fact that we have seen these scaling laws. Like, we, we call them scaling laws, the fact that going from a, a, you know, 100 million to a billion to 10 billion, 100 billion, one trillion parameter models, we've seen this amazing increase in capabilities with that. We see that also, uh, like still to this day, going from a trillion to 5 to 10 trillion at, at the largest end right now. And there's no ... We, like, we call it a law because we've observed it, but there's no, like, actual mathematical proof that this will continue. So it's a sort of on vibes that, okay, this has continued scale. There's no sign of it slowing down, so is that going to continue to 50 trillion, 100 trillion, and, and beyond? And- Um, I don't see any indication that that's going to stop, so, uh,
- 55:08 – 57:16
What Comes After AGI?
- TSThomas Sohmers
I'll, I'll be bullish on that
- HSHarry Stebbings
What, what does scaling laws look like at three times what it is now? If AGI has been declared now by Jansen-
- TSThomas Sohmers
Yeah
- HSHarry Stebbings
... forgive me, but what is three times this?
- TSThomas Sohmers
That's a good question. I mean, there's, it is still, um... I, I, GPT-6 Astra is, uh, my, my first, you know, 24 hours with it were basically as magical as when, uh, my, my first experience with ChatGPT, with, with GPT-3.5 in, in, uh, November of '22. Uh, I w- I was at the, uh, uh, ChatGPT launch at NeurIPS in, in 2022, and it was so funny because, um, [lips smack] uh, you know, Sam and, and Ilya were there and, and they, you know... It was a party in, in New Orleans for NeurIPS conference. And, um, uh, basically at the end they just said, "Hey, we launched this little fun, uh, uh, you know, experiment, uh, uh, called ChatGPT. Uh, go check it out." And, like, zero fanfare. You know, uh [chuckles] it was, it was, uh, uh, really just a side mention, and I don't think anyone really gave it a thought at the event. But when I went back to the hotel, I loaded it up, and I got back at, you know, 10:00, 11:00 PM or whatever, and I was up for four or five hours straight just giving random prompts and, and just being... This, that this was the most magical experience that I've ever had with a computer. And I would say I got very close when Sora 2 came out, that I had similar experience. Short amount of time, but just mind blo- blown by the quality of the, the videos, and especially the f- weekend that Sora 2 launched when there being no restrictions on what you could generate. Um, [lips smack] uh, but, uh, yeah, GPT-6 Astra I do think is, is AGI. And to, to your question of, like, what does that mean going forward, uh, I think my guess is as [chuckles] good as basically anyone's, but I, I-
- HSHarry Stebbings
Why, why, why was, why was-
- TSThomas Sohmers
Yeah
- HSHarry Stebbings
... GPT Astra so good for you? Why was it comparably such a breakthrough? 'Cause I, I have it and it's great, but honestly
- 57:16 – 1:01:33
Why GPT Astra Feels Like a Step Change
- HSHarry Stebbings
kind of the same as before.
- TSThomas Sohmers
Oh, um, in terms of the things that I've found LLMs to fail the most at in the past... So I'll, I'll give a case where it, it is more linear improvement. So just in terms of general coding capabilities, performance, and analyzing problems, et cetera, um, it is a step function improvement, but not mind-bogglingly so. There are a bunch of things that other models have not been able to fix or, um, [lips smack] kind of went in circles and kind found inelegant solutions and it's still, like, a human software architect that really understands the problem, um, uh, is able to come up with a better solution. With Astra, they're just initially giving it a couple really hard problems that I've n- not been able to solve with other LLMs, was able to do it one shot. Um, having it go through a code base and find both performance improvements, bugs, bugs, et cetera, and just solve them without, like... Basically discovering new spaces that I didn't know existed in our, our bunch of portions of our code base. So that's one element. Step function, but not mind-boggling. The second case that was mind-boggling just from a, like, who- like, wow, is the, the computer usabilities with a lot of set of generic tools. So, like, being able to do blender animations. Like, this, like, you know, it's become... There, there's a bunch of memes online of, of it recreating different videos, et cetera. But just the fidelity of that and where that was basically impossible with GPT 5.6 Soul, um, was massive increase in capability. And, like, I had it design, you know, do interior design of my house just based on a couple pictures and just like, wow. Like, I did not think that a, what's fundamentally a text model could, could do that. Um, and then finally, like the, the biggest thing for Positron was, um, I've been trying with every single new model release to have these models be able to, like, actually take a relatively simple logic design problem, uh, you know, implementing a, a encryption block in this case, um, and, uh, being able to take that through the full RTL to GDS flow. So from basically the specification of do this encryption function, um, implement the Verilog, so the hardware description language for that. Um, uh, so write that code, and then be able to take that code and go through all the way until you have got a chip design that theoretically you could go to tape or tape out. LLMs could do different portions of that and could, like, write the scripts and, you know, fail at a lot of different midpoints on, on the way. But a big problem i- with the, um, electronic design automation tools, the EDA tools for doing chip design, is that they were designed in the '90s, early 2000s. They're really unintuitive. None of the documentation exists out, like, in the public web. Um, and so the, you know, training, th- these models don't have, like, a real good innate view of them. But GPT-6 with [chuckles] uh, uh, both combination of computer use and just an ungodly amazing, uh, uh, scripting ability has been able to take this, this, uh, tech spec block and implement it, you know, with the TSMC N3 PBKs and take that all the way to GDS and do that in, like, a little over, like, 50 something hours. And so, um... And meet timing at, you know, over a gigahertz, et cetera. And, like, that, that as a task, if I was giving to someone similarly new to a, a, a, uh, a thing, like, getting the flow mostly working, I would say would take on the order of a week, and getting it optimized to the point that Astra is at with, with that design would maybe be one or two, two additional weeks depending on the person. So compressing that two to three weeks down to two days and change when the model- I- it's still mind-boggling. It like, it shouldn't be this good at this, [laughs] uh, just as I would naively think about, uh, uh, its training sets. But obviously with OpenAI's own chip development, um, and, you know, in-house, they've... Uh, I'm, I'm glad that those capabilities are getting added to the, the models they're releasing to the public
- 1:01:33 – 1:04:09
What Happens When Every AI Lab Builds Its Own Chips?
- TSThomas Sohmers
and not just being kept in- inside.
- HSHarry Stebbings
We see Jalapeno, terrible name, I think, personally.
- TSThomas Sohmers
[laughs]
- HSHarry Stebbings
But, you know, their own chip development, Anthropic are de- de- are, are developing their own chips. DeepSeek is supposedly developing their own chips. We see the commoditization of the chip layer with everyone building their own chips. How should we think about that?
- TSThomas Sohmers
As a consumer of all these things, if I take my Positron hat, shirt off, um, uh, I would say, I, that that's a great thing for the industry, having, um, uh, fundamentally that's gonna bring costs down and, you know, capabilities up and, and bring it to, to more people. Um, I think it's such an interesting world where when I got started in the semiconductor space, you know, 13 years ago, uh, you know, silicon was a dirty word in Silicon Valley, and now you have all the biggest companies in the world being, you know, somehow connected to the semiconductor industry and the most interesting, exciting applications and the, the companies building them, having, you know, vertically integrating down to the silicon layer. The interesting thing with all the, the ones that you mentioned and, and, you know, the broader set is the companies have the same macro goals. Um, the implementation details are all unique though, and that's, uh, just as an engineer and, and technologist is exciting to me that, that, um, there are a lot of different ways to skin a cat. And, uh, people can have their own, um, architectural view and, and go about implementing it and, and get different results. And, um, you know, Hot Chips, the biggest conference for, for this design space and was where, um, OpenAI un- in- unveiled Jalapeno, uh, last month. And, um, uh, it's, it's still really good, like, I'm, I'm happy that, um, the industry is still pretty open and willing to share, uh, uh, not as much details as people would've shared, you know, five, six years ago, but, uh, um, there's still, still a good amount of, um, in the open, uh, discussion of, of, of things. And so, um, I think how that applies to Positron is, you know, we have our particular, um, architectural views and, and, uh, way that, that we've decided to do things and that will, you know, evolve in, in the future as will everyone else's. But, you know, there's still plenty of space to make bets and, um, you know, go in, you know, different directions. And the great thing about the market is that the market gets to, to decide what is, is valuable, and those that, uh, create, create value [laughs] will receive a reward
- 1:04:09 – 1:09:16
How Far Can Context Windows Really Expand?
- TSThomas Sohmers
for that.
- HSHarry Stebbings
We spoke about context window length earlier and the expansion of it. How much does that expand? Is there infinite expansion capability of context window length, and what does that mean we can do that we can't do today? I'm just fascinated.
- TSThomas Sohmers
I think with traditional linear or, uh, uh, quadratic attention, there is going to be limits of scale what it can, to what the hardware can provide. Now, one of the big things for Positron is we're trying to massively increase the memory capacity per device. So, you know, with our, our upcoming, uh, generation, we're going to have, you know, eight times more memory capacity than the highest memory SKU from Nvidia, and Nvidia's actually decreasing the amount of memory per device, uh, you know, based on the, the m- market memory conditions. For, I frankly think that context length going from, like, a million tokens that it is today, going to 10 or even great- much more than that is really, really hard with that quadratic expansion of, of memory cost. The algorithmic advancements that have happened over the past year have been very, very interesting in terms of being able to further, uh, reduce the amount of, of storage and, and compute necessary for that context with linear and sparse attention mechanisms. Um, and those have really, really been innovated by the Chinese model labs, and it's, it's, this is a great example of when you have constraints of, you know, we had export controls on, on, you know, the, the chips with the highest memory capacity, um, and, and FLOPS. And so they innovated on not needing that. And so, um, you know, DeepSeek beginning of, of 2025 with, uh, DeepSeek V3 had, um, uh, you know, made a lot of waves because they were able to, to get massive decrease in KV cache size with, um, [lip smack] uh, multi-head latent attention. Uh, so you were actually spending more FLOPS to be able to have a smaller KV cache, and that's advanced a lot over the past year and a half. And probably the, the most interesting or my, my personal favorite right now is, uh, you know, gated Delta Nets and, and its derivative, uh, versions where you can have, like, a 75 de- percent decrease in, in, um, total time, uh, you know, you're spending on, on the attention portion, um, you know, with, with this mechanism, so.
- HSHarry Stebbings
How important then is new hardware if DeepSeek without it, just on architecture innovation alone, can cut costs by 80%?
- TSThomas Sohmers
Yeah, I would say that that's, like, there's no such thing as a free lunch. [laughs] So when, when they have that MLA compression, it does co- come at the cost of, um, model, um, uh, capabilities in, in some form. Um, and, like, the, uh, you know, there, there's a reason why the Chinese labs have really heavily embraced, um, you know, MLA while none of the US labs have. I should say based on rumors, but I also have, uh, on, on very good information and belief that, uh, you know, none, none of the major US companies are, are, uh, uh, they, they definitely are not using MLA. And, uh, you know- Uh, not using some of its, you know, brethren. Um, I think that, that, that will evolve and change in the future, but, um, uh, it... Yeah. Ba-basically the short version of that is, um, there's not, uh, it's not just a, a, you know, pure savings on that, on that side. I think that's... But the, the reality is to, and to go back to your previous question of, like, everyone does want greater context length. Like, if I had 10 million token context length, I think that would be enough for, um, holding multiple of our largest code bases and really have that cross-pollinization happen between them for an agentic coding model. It's not just having the full context. Like, there were some early models that had, that advertised a million token context length, but as soon as you went above, like, 64,000 tokens, that, you know, that's two years ago, um, its recall ability just went garbage. So just saying that something has this maximum context length is one thing. It's can it actually use that context length effectively is an entirely different thing. And that's actually going back to the Astra thing. Amazing thing about it, the, there, there's a couple different benchmarks measuring, like, the long context perf- performance. Uh, one of them's called Ruler, and there's, like, these find a needle in a haystack. So you just flood the, the context window with a bunch of junk basically, like, just passages from books and all this. And you just put somewhere randomly in the middle of all of that, um, you know, a, a, a hash, or, you know, some, some value that looks out of place. Um, and, uh, you prompt the model, say, "What, what, what's the, what's the secret value?" And a lot of models have done really poorly on this. GPT 5.6, which is only, like, six or seven weeks old, um, uh, you know, only could do this about 70% of the time. GPT 6 Astra does it, like, over 95% of the time correctly. So that's... There, there's a lot of room for
- 1:09:16 – 1:11:58
The Cost of AI Tokens Has Collapsed 60x
- TSThomas Sohmers
improvement in these things.
- HSHarry Stebbings
Speaking of room for improvement, can I ask you, when I was doing the research for the show, I saw the, the Silicon data token price index drop below $1 per million tokens this month, and five years ago it was $60 per million. $60 to one. What does a million tokens cost in 2028, say, two years from now, do you think?
- TSThomas Sohmers
I, I think the much, the thing I care a lot more about than just that 60 to one is the fact that a $60 token five years ago, no one would pay a cent for today. Like, that was a, a complete garbage token, relatively speaking, five years ago. And the level of quality of a, a token that you pay a dollar per million tokens for now, um, is, uh, so much astronomically more valuable. And so-
- HSHarry Stebbings
And that's because the... And sorry, that's because of token efficiency and what can be done?
- TSThomas Sohmers
No, I'm speaking just in model capabilities. If you, if you say, okay, so, so it's 2026, so in the best model in the world in, in 2021, uh, was GPT III. Um, it's kind of crazy at the rate models get released today that GPT III was the best in the world basically from tw- uh, from 2020, I think it was August 2020 when it released, all the way up to they, they didn't have a new release until ChatGPT in November of, of, uh, tw- of 2022. So it was, you know, two, two and a half years [laughs] between, between model releases. Uh, and really GPT 3.5 was just doing reinforce- reinforcement learning with human feedback on the same base model. If you remember how bad [laughs] uh, GPT 3.5 was and, like, what was the, the, um, value of that in terms of economic, economic productivity value of GPT 3.5 versus GPT 6 today or, you know, pick whatever comparison points you want. Um, the value per token in terms of what it can improve a person's life, you know, a company's, you know, business practices, et cetera, et cetera, is orders of magnitude, I would say 100 or 1,000 fold. So I, I think there's actually two points to, to your, your axis of going from $60 to $1. Yes, that's a decrease in cost, but that token today is, let's just say, I think conservatively 100 times more valuable. So I would actually be saying that there needs to be some multiplier there as well, where, like, the, the value per, you know, unit of intelligence is probably closer to 1,000 fold, not, not just the 60 fold
- 1:11:58 – 1:13:49
Will AI Stop Being Priced Per Token?
- TSThomas Sohmers
you're, you're talking about.
- HSHarry Stebbings
What does that mean then if we extrapolate that out to 2028? What does that mean that... Like, d- does the cost of a token then actually matter? Like, is that the primary unit that we should measure? 'Cause everyone talks about cost of token. Is there actually a different metric that we should measure?
- TSThomas Sohmers
It is interesting that with the GPT 6 launch, Greg Brockman had said that, um, that he doesn't think that cost per token, that, that they're going to be pricing things in tokens much longer and that they, they want to be moving and, uh, to, you know, cost per useful result, like, uh, to that. And I don't think that that is, I don't think that's where it'll end up because that's really difficult to price and, um, uh, you know, qualia, et cetera. But I think that the price per token is really great because you can easily calculate the cost, like the, the cost to generate a token. So determining a margin on that and pricing it in, as in bulk volume to generic customers is really easy, and I think that's going to stick around in large form because that is so easy. We'll see for the largest providers of tokens how they potentially evolve their business models in terms of, um, if you have a GPT 7 or 8 that is superhuman and, and can fully function as an employee in, in an amazing capacity- And OpenAI calculates through whatever method that, um, running at full tilt, et cetera, it's only going to cost them, you know, however many hundreds of thousands of dollars to produce tokens continuously with that. They may decide that it's actually easier and they'll be able to get more adoption if they just charged a million dollars a year just, you know, using a, a random number to have full unlimited usage of that, uh, of, of that virtual agent worker.
- 1:13:49 – 1:16:31
What Would Actually Burst the AI Bubble?
- TSThomas Sohmers
So that, that may be how things evolve.
- HSHarry Stebbings
Can I ask you, I, I'm always very careful of being, like, the young, naive one. I'm not that young anymore, but, like, being the naive one who ha- who's not seen cycles. But Gavin Baker says it well when he says, "I can't speak to a company that don't have numbers that are parabolically up and to the right and just, uh, everything is, is better than it's ever been." What would be the first signs of a crack in the chasm? It's like a shift from frontier models to open-weight models. Anthropic and OpenAI are not continuing in the same level, not quite growth rate 'cause it's impossible to say that, but level of growth, missing numbers next year, and then the bubble getting burst a little bit whether the two core leaders are having some form of strife.
- TSThomas Sohmers
The reason I, I, I agree that that's a possibility, the reason I don't think it's likely is I think that the development of open source models and, and things happening locally, et cetera, will actually drive greater usage, greater, you know, token volumes for, for the big guys. I, I really do think that that's-
- HSHarry Stebbings
Sorry, how does that work? I, I thought they were competitive. I thought you... Yeah.
- TSThomas Sohmers
Yeah, no, I, I think that, you know, the smarter and more capable that Siri is on my phone, um, is going to result on in it doing a whole bunch of background tasks and things that remove me from having to be the one that instigates, um, it going, having requests and, and data be processed by even smarter models. I, I, I really do think that in, in a lot of AI applications right now, the bottleneck is actually a human, um, making some sort of decision, and different tasks have different levels of autonomy that, that will result in, in things getting sent to be processed by, by a model, you know, by OpenAI or Anthropic. Um, but I think the, the, the next really big, you know, order of magnitude, couple orders of magnitude increase in, um, token volumes is going to come when, when us humans trust a local LLM bec- that, that has access to all of our data all the time to have it decide to do things on its own that it is not cap- smart enough to do. It's just like if it ... Right, right now, I trust Astra a lot more than myself on a whole lot of different things, but I still prompt it to do things, and maybe it will go run autonomously for 12 hours or three days. I think the next big leap is when I trust a model running on my laptop or on my phone to prompt the smarter models to do even more wider, uh, set of tasks.
- HSHarry Stebbings
When we look at, like, the data economy that powers the larger models
- 1:16:31 – 1:19:43
How Big Can the AI Data Economy Become?
- HSHarry Stebbings
which you believe in, we, we, we see Merqur, we see Surge hitting 3 billion in revenue. How big do these companies become? Because Anthropic and OpenAI are $4 to $5 trillion businesses, say, feasible. It, it's wholly feasible, isn't it, that Merqur and Surge are $200 billion businesses which serve both frontier labs and some of the world's biggest enterprises bu- building their own models?
- TSThomas Sohmers
I think my only skepticism there is, um, on there being vertical integration by, by the frontier labs. I think the, the reason that hasn't happened is because Anthropic and OpenAI have better uses of their capital and mental power, et cetera, than, uh, doing all of that, the, you know, scale AI Merqur, et cetera, work. Um, but, uh, the, uh, I, I think the, the main ... That won't always exist. And let, let's say that GPT-7 or 8 could be an effective replacement for Sam Altman in terms of being able to manage, you know, a, a large business, why they wouldn't just have agents be taking over those tasks.
- HSHarry Stebbings
I like to finish on a tone of optimism. What are you most excited about today that you think the world does not spend enough time on that we should spend time on?
- TSThomas Sohmers
I'm very sympathetic to the, um, problems that I think the, the smarter set of the AI alignment, AI safety community, um, are when it, when it comes to thinking about how do you align incentives. Like, uh, a lot of people just talk about AI alignment being a problem. Um, I think there's, uh, a, a key part of that is human alignment. It's like how do we as a society, you know, human race align ourselves to have a good outcome that I think will be empowered by artificial intelligence. And so we discussed all, a, a bunch of the different problems that we're facing geopolitically and, and socially and, and, and how these different things are handled. And, um, I think a lot of smart people are doing good work on the AI alignment problem and, and thinking through how we solve that, but they may be gated in what can be done there if we don't get better human alignment on, um, regulatory frameworks and energy production, where we're going to put the data centers, et cetera, et cetera. And it, I, I think framing it as this being a similar sort of technical problem that smart people can work and reason through will hopefully get more people thinking about it that way. And I think a core element that gets discounted by, I think, a lot of people in, in, um, that sphere, there are a lot of economic factors, and I, I think it's the economic factors that, that actually will drive real decision-making and, and, and actions. And, um, if, if we don't look at it from, you know, a rational, um, you know, self-interested actors and, and all these different things, um, you're just not going to make progress.
- HSHarry Stebbings
Thomas, this has been the most varied discussion ever from educational and, you know, uh, u- unbelievable infrastructure ev- evolution to, uh, Dario and Sam. You are a star. Thank you so much for joining me today.
- TSThomas Sohmers
Thank you so much, Harry.
Episode duration: 1:19:54
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode 6ohZuFkq-aU