AcquiredNvidia Part III: The Dawn of the AI Era (2022-2023) (Audio)
EVERY SPOKEN WORD
150 min read · 30,002 words- 0:00 – 1:30
Nvidia Part III setup: why an extra episode was necessary
- BGBen Gilbert
You like my Buck's T-shirt?
- DRDavid Rosenthal
I love your Buck's T-shirt.
- BGBen Gilbert
I went for the first time, what, two weeks ago, when I was down for a meeting at Benchmark, and the nostalgia in there was just unbelievable.
- DRDavid Rosenthal
I can't believe you hadn't been before. I know Jensen is a Denny's guy, but I feel like he would meet us at Buck's if we asked him.
- BGBen Gilbert
Or at the very least, we should figure out some Nvidia memorabilia to get on the wall at Buck's.
- DRDavid Rosenthal
Totally.
- BGBen Gilbert
Fit right in. All right, let's do it.
- DRDavid Rosenthal
Let's do it.
- SPSpeaker
Who got the truth? Is it you? Is it you? Is it you? Who got the truth now? Is it you? Is it you? Is it you? Sit me down, say it straight. Another story on the way. Who got the truth?
- BGBen Gilbert
Welcome to Season 13, Episode 3 of Acquired, the podcast about great technology companies and the stories and playbooks behind them. I'm Ben Gilbert.
- DRDavid Rosenthal
I'm David Rosenthal.
- BGBen Gilbert
And we are your hosts. Today, we tell a story that we thought we had already finished: Nvidia. But the last eighteen months have been so insane, listeners, that it warranted an entire episode on its own. So today is a part three for us with Nvidia, telling the story of the AI revolution, how we got here, and why it's happening now, starting all the way down at the level of atoms and silicon. So here's something crazy that I did a transcript search on to see if it was true. In our April 2022 episodes, we never once said the word "generative." That is how fast things have changed.
- DRDavid Rosenthal
Unbelievable.
- 1:30 – 4:36
From bleak 2022 to AI’s ‘Netscape/iPhone moment’
- BGBen Gilbert
Totally crazy. And the timing of all of this AI stuff in the world is unbelievably coincidental and, uh, very favorable. So recall back to eighteen months ago. Throughout 2022, we all watched financial markets, from public equities to early-stage startups to real estate, just fall off a cliff due to rapid rise in interest rates. The crypto and Web3 bubble burst, banks fail. It seemed like the whole tech economy, and potentially a lot with it, was heading into a long winter.
- DRDavid Rosenthal
Including Nvidia.
- BGBen Gilbert
Including Nvidia, who had that massive inventory write-off for what they thought was over-ordering.
- DRDavid Rosenthal
Yep. Wow, how things have changed.
- BGBen Gilbert
[chuckles] Yeah. But by the fall of 2022, right when everything looked the absolute bleakest, a breakthrough technology finally became useful after years in research labs. Large language models, or LLMs, built on the innovative transformer machine learning mechanism, burst onto the scene, first with OpenAI's ChatGPT, which became the fastest app in history to a hundred million active users, and then quickly followed by Microsoft, Google, and seemingly every other company. In November of 2022, AI definitely had its Netscape moment, and time will tell, but it may have even been its iPhone moment.
- DRDavid Rosenthal
Well, that is definitely what Jensen believes.
- BGBen Gilbert
Yep. Well, today, we'll explore exactly how this breakthrough came to be, the individuals behind it, and of course, why the entire thing has happened on top of Nvidia's hardware and software. If you wanna make sure you know every time there's a new episode, go sign up at acquired.fm/email. You'll also get access to two things that we aren't putting anywhere else: one, a clue as to what the next episode will be, and two, follow-ups from previous episodes from things that we learned after release. You can come talk about this episode with us after listening at acquired.fm/slack. If you want more of David and I, check out our interview show, ACQ2. Our next few episodes are about AI, with CEOs leading the way in this world we are talking about today, and a great interview with Doug Demuro, where, uh, we wanted to talk about a lot more than just Porsche with him, but, uh, you know, we only had eleven hours or whatever we had-
- DRDavid Rosenthal
[laughing]
- BGBen Gilbert
... in Doug's garage, so a lot of the, uh, car industry chat and learning about Doug and his journey and his business we saved for ACQ2, so go check it out. One final announcement, many of you have been wondering, and we've been getting a lot of emails: When will those hats be back in stock? Well, they're back. For a limited time, you can get an ACQ embroidered hat at acquired.fm/store. Go put your order in before they, uh, go back into the Disney vault forever.
- DRDavid Rosenthal
[laughing] This is great. I can finally get Jenny one of her own, so she stops stealing mine.
- BGBen Gilbert
[chuckles] Yes. Well, without further ado, this show is not investment advice. David and I may have investments in the companies we discuss, and this show is for informational and entertainment purposes only. David, history and facts.
- 4:36 – 7:59
Revisiting Nvidia’s old $1T TAM—and the accidental prophecy
- DRDavid Rosenthal
Oh, man. Well, on the one hand, we only have eighteen months to talk about.
- BGBen Gilbert
Except that I know you're not gonna start eighteen months ago. [chuckles]
- DRDavid Rosenthal
On the other hand, we have decades and decades of foundational research to cover. So when I was starting my research, I went to the natural first place, which was our old episodes from April 2022, and I was listening to them, and I got to the end of the second one, and, uh, man, I had forgotten about this. I think Jensen maybe wishes we all had forgotten about this. In one of Nvidia's earnings slides in 2021, they put up their total addressable market, and they said they had a one trillion dollar TAM, and the way that they calculated this was that they were gonna serve customers who provided $100 trillion worth of industry, and they were gonna capture just 1% of it. And there was some stuff on the slide that was fairly speculative, you know, like autonomous vehicles and the omniverse, and I think robotics were a big part of it.
- BGBen Gilbert
And the argument is basically like, well, cars plus factories plus all these things added together is $100 trillion, and we can just take 1% of that, 'cause surely their compute will amount to 1% of that, which I'm not arguing is wrong, but it is a very blunt way to analyze that market.
- DRDavid Rosenthal
Yeah, it's usually not the right way to, um, think about starting a startup. You know, "Oh, if we can just get 1% of this big market," blah, blah, blah.
- BGBen Gilbert
It's the toppiest-down way I can think of to size a market.
- DRDavid Rosenthal
So you, Ben-... rightly so, called this out at the end of Nvidia part two, and you're like, "You know, I think to justify where Nvidia is trading at the moment, you kinda actually gotta believe that all of this is gonna happen, and happen soon: autonomous cars, robotics, everything."
- BGBen Gilbert
Yeah. Importantly, I felt like the way for them to become worth what they were worth at that time literally had to be to power all of this hardware in the physical world.
- DRDavid Rosenthal
Yep. I kinda can't believe that I said this, because it was unintentional and uninformed, but I was kinda grasping at straws trying to play devil's advocate for you. And we'd just spent most of that whole episode talking about how machine learning powered by Nvidia ended up having this incredibly valuable use case, which was powering social media feed recommenders, and that Facebook and Google had grown bigger than anyone ever imagined on the internet with those feed recommendations, and Nvidia was powering all of it. And so I just sorta idly proposed, "Well, maybe, but what if you don't actually need to believe any of that to still think that Nvidia could be worth a trillion dollars? What if maybe, just maybe, the internet and software and the digital world are gonna keep growing, and there will be a new foundational layer that Nvidia can power? Is that possible?" And I think we were both like, "Yeah, I don't know. Let's end the episode."
- BGBen Gilbert
Yeah, sure. We shrugged it off, and we were like, "All right, carve-outs."
- DRDavid Rosenthal
But the crazy thing is that, of course, at least in this timeframe, most things on Jensen's trillion-dollar TAM slide have not come to pass, but that crazy question just might have come to pass, and from Nvidia's revenue and earnings standpoint, definitely has. It's just wild.
- BGBen Gilbert
All right, so how did we get here?
- 7:59 – 10:30
2012’s AlexNet: the Big Bang for GPU-accelerated AI
- DRDavid Rosenthal
Let's rewind and tell the story. So back in 2012, there was the Big Bang moment of artificial intelligence, or as it was more humbly referred to back then, machine learning, and that was AlexNet. We talked a lot about this on the last episode. It was three researchers from the University of Toronto who submitted the AlexNet algorithm to the ImageNet computer science competition. Now, ImageNet was a competition where you would look at a set of 14 million images that had been hand-labeled with what the pictures were of, like of a strawberry or a cat or a dog or whatever.
- BGBen Gilbert
And David, you were telling me it's the largest-ever use of Mechanical Turk up to that point, was to label the ImageNet dataset?
- DRDavid Rosenthal
Yeah, it's wild. I mean, until this competition and until AlexNet, there was no machine learning algorithm that could accurately label images. So thousands of people on Mechanical Turk got paid however much, two bucks an hour, to label these images.
- BGBen Gilbert
Yeah, and if I'm remembering from our episode, basically what happened is the AlexNet team did way better than anybody else had ever done, the complete step change better. I think the error rate went from mislabeling images twenty-five percent of the time to suddenly only mislabeling them fifteen percent of the time, and that was, like, a huge leap over the tiny incremental progress that had been made along the way.
- DRDavid Rosenthal
You are spot on. And the way that they did it, and what completely changed the fortunes of the internet, of Google, of Facebook, and certainly of Nvidia, was they actually used old algorithms, a branch of computer science and artificial intelligence called neural networks, specifically convolutional neural networks, which had been around since the '60s, but they were really computationally intensive to train. And so nobody thought it would be practical to actually train and use these things, at least not anytime soon or in our lifetimes. And what these guys from Toronto did is [chuckles] they went out, probably to their local Best Buy or equivalent in Canada. They bought two GeForce GTX 580s, which were the top-of-the-line cards at the time, and they wrote their algorithm, their convolutional neural network, in CUDA, in Nvidia's software development platform for GPUs, and by God, they trained this thing [chuckles] on, like, a thousand dollars' worth of consumer-grade hardware.
- 10:30 – 12:39
Why GPUs win: parallelism as a lever on Moore’s Law
- BGBen Gilbert
And basically, the algorithm that other people had been trying over the years just wasn't massively parallel the way that a graphics card sort of enables. So if you actually can consume the full compute of a graphics card, then perhaps you could run some unique, novel algorithm and do it on, you know, a fraction of the time and expense that it would take in these supercomputer laboratories.
- DRDavid Rosenthal
Yeah, everybody before was trying to run these things on CPUs. CPUs are awesome, but they only execute one instruction at a time. GPUs, on the other hand, execute hundreds or thousands of instructions at a time. So GPUs, Nvidia graphics cards, accelerated computing, what Jensen and the company likes to call this, you can really think of it like a giant Archimedes lever. Whatever advances are happening in Moore's Law and the number of transistors on a chip, if you have an algorithm that can run in parallel, which is not all problem spaces, but many can, then you can basically lever up Moore's Law by hundreds of times or thousands of times, or today, tens of thousands of times, and execute something a lot faster than you otherwise could.
- BGBen Gilbert
And it's so interesting that there was this first market called graphics that was obviously parallel, where every pixel on a screen is not sequentially dependent on the pixel next to it. It literally can be computed independently and output to the screen, so you have however many tens of thousands or now hundreds of thousands of pixels on a screen that can all actually be done in parallel. And little did Nvidia realize, of course, that AI and crypto and all this other linear algebra, matrix math-based things that turned into accelerated computing, pulling things off the CPU and putting them on GPU and other parallel processors-... was an entire new frontier of other applications that could use the very same technology they had pioneered for graphics.
- DRDavid Rosenthal
Yeah, it was pretty useful stuff, and this AlexNet moment, and these three researchers from Toronto kicked off, you know, Jensen calls it, and he's absolutely right, the Big Bang moment for AI.
- 12:39 – 26:02
The people behind the breakthrough—and the path to OpenAI
- BGBen Gilbert
So David, the last time we told this story in full, we talked about this team from Toronto. We did not follow what this team of three went on to do afterwards.
- DRDavid Rosenthal
Yeah. So basically what we said was, it turned out that a natural consequence of what these guys were doing was, "Oh, actually, you can use this to surface the next post in a social media feed, on, like, an Instagram feed or the YouTube feed or something like that." And that unlocked billions and billions of value, and those guys and everybody else working in the field, they all got scooped up by Google and Facebook. Well, that's true, and then as a consequence of that, Google and Facebook started buying a lot of Nvidia GPUs. But turns out there's also another chapter to that story that we completely skipped over, and it starts with the question you asked, Ben: Who are these people?
- BGBen Gilbert
Yes.
- DRDavid Rosenthal
So the three people who made up the AlexNet team were, of course, Alex Krizhevsky, who was a PhD student under his faculty advisor, the legendary computer science professor Jeff Hinton. I have an amazing piece of trivia about Jeff Hinton.
- BGBen Gilbert
Hmm.
- DRDavid Rosenthal
Do you know who his great-great-grandparents were?
- BGBen Gilbert
No, I have no idea.
- DRDavid Rosenthal
He is the great-great-grandson of George and Mary Boole. You know, like Boolean algebra and Boolean logic?
- BGBen Gilbert
This guy was born to be a computer science researcher. Oh, my God!
- DRDavid Rosenthal
Right? [laughing] Foundational stuff for computation and computer science.
- BGBen Gilbert
I also didn't know there were people named Boole, that that's where that came from. That's hilarious.
- DRDavid Rosenthal
Yeah. You know, the AND, OR, XOR, NOR operators, that comes from George and Mary. Wild. So he's the faculty advisor, and then there was a third person on the team, Alex's fellow PhD student in this lab, one Ilya Sutskever, and if you know where we're going with this, you are probably jumping up and down right now in your seat. Ilya is the co-founder and current chief scientist of OpenAI.
- BGBen Gilbert
Yes.
- DRDavid Rosenthal
So after AlexNet, Alex, Jeff, and Ilya do the very natural thing: they start a company. I don't know what they were doing in the company, but, uh, it made sense to start one.
- BGBen Gilbert
And whatever they did, it was gonna get acquired real fast.
- DRDavid Rosenthal
By Google, within six months. So they get scooped up by Google. They join a bunch of other academics and researchers that Google has been monopolizing, really, in the field. Three specifically, Greg Corrado, Jeff Dean, and Andrew Ng, the famous Stanford professor. The three of them had just formed the Google Brain team within Google to turbocharge all of this AI work that has been unleashed by AlexNet, and of course, to turn it into huge amounts of profit for Google.
- BGBen Gilbert
Turns out, individually serving advertising that's perfectly targeted on the internet through Facebook or Google-
- DRDavid Rosenthal
Or YouTube
- BGBen Gilbert
... is an enormously profitable business, and one that consumes a whole lot of Nvidia GPUs.
- DRDavid Rosenthal
Yes. So about a year later, Google also acquires DeepMind, famously, and then right around the same time, Facebook scoops up computer science professor Yann LeCun, who also is a legend in the field, and the two of them basically establish a duopoly on leading AI researchers. Now, at this point, nobody is mistaking what these companies and these people are doing for true human-level intelligence or anything close to it. This is AI that is very good at narrow tasks, like we talked about, social media feed recommendations. So the Google Brain team and Jeff and Alex and Ilya, one of the big projects they work on is redoing the YouTube algorithm, and this is when YouTube goes from, like, money-losing, you know, crazy thing that Google acquired to the just absolute juggernaut that it is today. I mean, back then in, like, twenty thirteen, twenty fourteen, we did our YouTube episode not that long after. The majority of views of YouTube videos were embeds on other web pages. This is when they build it into a social media site, they start the feed, they start autoplay. All this stuff is coming out of AI research. Some of the other stuff that happens at Google, famously, after they acquired DeepMind, DeepMind built a bunch of algorithms to save on cooling costs. And Facebook, of course, they probably had the last laugh in this generation because they're using all this work, and Yann LeCun is doing his thing and hiring lots of researchers there. This is just a couple years after they acquired Instagram. Man, we need to, like, go back and redo that episode, because Instagram would have been a great acquisition anyway, but it was AI-powered recommendations in the feed that made that into a hundred, two hundred, five hundred billion dollar asset for Facebook.
- BGBen Gilbert
And I don't think you're exaggerating. I think that is literally what Instagram is worth to Meta now. By the way, I have bought a lot of things on Instagram ads, so the, the targeting works.
- DRDavid Rosenthal
It absolutely does. There's this amazing quote from Astro Teller, who ran Google X at the time, and still does, in a New York Times piece, where he says that the gains from Google Brain during this period, I don't think this even includes DeepMind, just the gains from the Google Brain team alone in terms of profits to Google more than funded everything they were doing in Google X.
- BGBen Gilbert
Which, has there ever been anything profitable out of Google X?
- DRDavid Rosenthal
Google Brain. [laughing]
- BGBen Gilbert
Yeah, I mean, yeah.
- DRDavid Rosenthal
We'll leave it at that.... So this takes us to 2015 when a few people in Silicon Valley start to realize that this Google, Facebook, AI duopoly is actually a really, really big problem. And most people had no idea about this. This is really visionary of these two people.
- BGBen Gilbert
And not just a problem for, like, the other big tech companies, 'cause you could make the argument it's a problem 'cause, like, Siri's terrible, all the other companies that have lots of consumer touchpoints have pretty bad AI at the time, but the concern is for a much greater reason.
- DRDavid Rosenthal
I think there are three levels of concern here. One, obviously, is the other tech companies. Then there's the problem of startups. This is terrible for startups! How are you gonna compete with Google and Facebook when this is the primary value driver of this generation of technology? I mean, there really is another lens to view what happened with Snap, what happened with Musical.ly, and having to sell themselves to ByteDance and becoming TikTok and going to the Chinese. Maybe it was business decisions, maybe it was execution or whatever, that prevented those platforms from getting to independent scale. Snap's a public company now, but, like, it's no Facebook. Maybe it was that they didn't have access to the same AI researches that Facebook and Google had.
- BGBen Gilbert
Hmm. That feels like an interesting question. It's probably a couple steps too far on the conclusion, but still sort of a fun straw man to think about.
- DRDavid Rosenthal
A fun straw man. Nonetheless, this is definitely a problem. The third layer of the problem is just, like, this sucks for the world, [chuckles] that all these people are locked up in Google and Facebook.
- 26:02 – 30:42
Pre-transformer ambitions: early language-model intuition and next-word prediction
- DRDavid Rosenthal
Turns out, uh, no. [laughing] So as we were talking about a little bit, AI at this point in time, super good for narrow use cases. Looks nothing like GPT-4 today. The capabilities that it had were pretty limited, and one of the big reasons was that the amount of data that you could practically train these models on was pretty limited. So the AlexNet example, you're talking about fourteen million images. In the grand scheme of the internet, fourteen million images is a drop in the bucket.
- BGBen Gilbert
And this was both a hardware and a software constraint. On the software side, we just didn't actually have the algorithms to sort of suppose that we could be so bold to train one single foundational model on the whole internet. Like, it wasn't a thing.
- DRDavid Rosenthal
Yeah, that was a crazy idea.
- BGBen Gilbert
Right. People were excited about the concept of language models, but we actually didn't know how we could algorithmically get it done. So in 2015, Andrej Karpathy, who was then at OpenAI and went on to lead AI for Tesla and is actually now back at OpenAI, writes this seminal blog post called The Unreasonable Effectiveness of Neural Networks. And David, I don't think we're gonna go into it on this episode, but note that recurrent neural networks are a little bit of a different thing than convolutional neural networks, which was the 2012 paper.
- DRDavid Rosenthal
The state-of-the-art had evolved.
- BGBen Gilbert
Yes, and right around that same time, there is also a video that hits YouTube, or a little bit later in 2016, that is actually on Nvidia's channel and has two people in this very short one minute and forty-five second video. One is a young Ilya Sutskever, and two is Andrej Karpathy, and here is a quote from Andrej from that YouTube video: "One algorithm I'm excited about is a language model. The idea that you can take a large amount of data and you feed it into the network, and it figures out the pattern in how words follow each other in sentences. So, for example, you could take a large amount of data on how people talk to each other on the internet. You can train basically a chatbot, but you can do it in a way that the computer learns how language works and how people interact. Eventually, we'll use that to talk to computers just like we talk to each other."
- DRDavid Rosenthal
Wow, this is 2015?
- BGBen Gilbert
This is two years before the transformer, while Karpathy is at OpenAI.
- DRDavid Rosenthal
Wow!
- BGBen Gilbert
He both comes up with the idea or espouses the idea of a chatbot, so that sort of had already been discussed. But even before we had the transformer, the method to actually pull this off, he sort of had the idea that... And there's an important part here: It figures out the pattern in how words follow each other in sentences. So there's this idea that the very structure of language and the way to interpret knowledge is actually embedded in the training data itself, rather than requiring labeling.
- DRDavid Rosenthal
This is so cool. So at Spring GTC this year, Jensen did a fireside chat with Ilya, and it's amazing. You should go watch the whole thing. But in it, this question comes up. Jensen kind of poses as a straw man, like: Hey, some people say that GPT-3, 4, ChatGPT, e- everything going on, all these LLMs, they're just probabilistically predicting the next word in a sentence. They don't actually have knowledge. And Ilya has this amazing response to that. He says: Okay, well, consider a detective novel.
- BGBen Gilbert
Yes.
- DRDavid Rosenthal
At the end of the novel, the detective gathers everyone together in a room and says, "I am now going to tell you all the name of the person who committed the crime, and that person's name is" blank. [chuckles] The more accurately an LLM predicts that next word, i.e., the name of the criminal, ipso facto, the greater its understanding, not only of the novel, but of all general-... human-level knowledge and intelligence, because you need all of your experience in the world and as a human to be able to guess who the criminal is, and the LLMs that are out there today, GPT-3, GPT-4, Llama, Bard, these others, they can guess who the criminal is.
- 30:42 – 44:10
2017 Transformer paper: attention, context windows, and GPU-friendly parallelism
- BGBen Gilbert
Ooh, yeah, put a pin in that. Understanding versus predicting, the hot topic du jour. So David, is now a good time to fast-forward two years to 2017 to the transformer paper?
- DRDavid Rosenthal
Absolutely. Ben, tell us about the transformer.
- BGBen Gilbert
Okay, so Google, 2017, transformer paper. Paper comes out. It's called Attention Is All You Need.
- DRDavid Rosenthal
And it's from the Google Brain team, right?
- BGBen Gilbert
Yes.
- DRDavid Rosenthal
That Ilya just left.
- BGBen Gilbert
Just left, two years before, to start OpenAI. So machine learning on natural language, just to set the table here, had long been used for things like autocorrect or foreign language translation. But in 2017, Google came out with this paper and discovered a new model that would change everything for these fields and unlock another one. So here is the scenario: You're translating a sentence from English to French. You could imagine that a way to do this would be one word at a time, in order. But for anyone who's ever traveled abroad and, uh, tried to do this, you know that words are sometimes rearranged in different languages, so that's a terrible way to do it. You know, United States in Spanish is Estados Unidos, so failure on the very first word in that example. So enter this concept of attention, which is a key part of this research paper. So this attention, this fairly magical component of the transformer paper, it literally is what it sounds like. It is a way for the model to attend to different areas of the input text at different times. You can look at a large amount of context while considering what word to pick next in your translation. So for every single word that you're about to output in French, you can look over the entire set of inputted words to figure out what words you should weight heavily in your decision for what to do next.
- DRDavid Rosenthal
This is why AI and machine learning was so narrowly applicable before. If you anthropomorphize it, and you think of it like a human, it was like a human with a very, very short attention span. [chuckles]
- BGBen Gilbert
Yes. Now, here's the magical part. While it does look at the whole input text to consider what the next word should be, it doesn't mean that it throws away the notion of position entirely. It uses a technique called positional encoding, so it doesn't forget the position of the words altogether. So it's got this cool thing where it weights the important part relevant to your particular word, and it still understands position. So remember I said the attention mechanism looks over the entire input every time it's picking what word to output?
- DRDavid Rosenthal
That sounds very computationally hard.
- BGBen Gilbert
Yes. In computer science terms, this means that the attention mechanism is O of N squared.
- DRDavid Rosenthal
Oh, that's giving me the heebie-jeebies-
- BGBen Gilbert
[chuckles]
- DRDavid Rosenthal
... back to my intro CS classes in college.
- BGBen Gilbert
Oh, just wait till we get through this episode. It gets deeper. So obviously, yes, traditionally, you'd say this is very, very inefficient, and it actually means that the larger your context window, AKA token limit, AKA prompt length, gets, the more computationally expensive it gets on a quadratic basis. So doubling your input means quadrupling the cost to compute an output, or tripling your input means nine times the cost.
- DRDavid Rosenthal
It gets real gnarly.
- BGBen Gilbert
Yeah, it gets real expensive, real fast. But GPUs to the rescue. The amazing news for us here is that these transformer comparisons can be done in parallel. So even though there are lots of them to do, if you have big GPU chips with tons of cores, you can do them all at exactly the same time, and previous technologies to accomplish this, like recurrent neural networks or LSTMs, long short-term memory networks, which is a type of recurrent neural network, et cetera, those required knowing the output of each step before beginning the next one, before you picked the next word. So in other words, they were sequential, since they depended on the previous word. Now, with transformers, even if your string of text that you're inputting is a thousand words long, it can happen just as quickly in human measurable time as if it were ten words long, supposing that there were enough cores in that big GPU. So the big innovation here is you could now train sequence-based models in a parallel way. You couldn't train models of this size at all before, let alone cost-effectively.
- DRDavid Rosenthal
Yeah. This is huge, and probably for all listeners out there, starting to sound very familiar to the world that we live in today.
- BGBen Gilbert
Yeah, I sort of did a sleight of hand there, morphing translation to using words like context window and token length. You can kind of see where this is going.
- DRDavid Rosenthal
Yep. So this transformer paper comes out in 2017. The significance is huge, but for whatever reason, there's a window of time where the rest of the world doesn't quite realize it. So Google obviously knows how important this is, and there's, like, a year where Google's AI work, even though Ilya has left and OpenAI is a thing now, accelerates again beyond anybody else in the field. So this is when Google comes out with Smart Compose in Gmail, and they do that thing where they have an AI bot that'll call local businesses for you. Remember that, uh, demo from I/O-
- BGBen Gilbert
Yeah
- DRDavid Rosenthal
... that they did?
- BGBen Gilbert
Did that ever ship?
- DRDavid Rosenthal
I don't know. [laughing] Maybe it did, maybe... I mean, this is Google here. Like, the capabilities are there. The product sense, eh, not as much. This is when they really start investing in Waymo, but again, where it really manifests is just back to serving ads in Search and recommending YouTube videos. Like, they're just crushing it in this period of time.... OpenAI and everyone else, though, they haven't adopted transformers yet. They're kind of stuck in the past, and they're still doing these really research-y computer vision projects. So, like, this is when they build a bot to play Dota 2, Defense of the Agents 2, the video game.
- BGBen Gilbert
And super impressive stuff, like they beat the best Dota players in the world at Dota by literally just consuming computer vision, like consuming screenshots and inferring from there. And that's a really hard problem, 'cause Dota 2 is not a game where you get to see the whole board at once, so it has to do a lot of, like, really intelligent construction of the rest of the game based on just a single player's worth of input. So it's unbelievably cutting-edge research.
- DRDavid Rosenthal
For the past generation. It's a faster horse, basically.
- BGBen Gilbert
Maybe. Yeah. I mean, they were also doing stuff like, uh, Universe, which was the 3D-modeled world to train self-driving cars. You don't really hear anything about that anymore, but they built this whole thing. I think it was using Grand Theft Auto as the environment, and then it was doing computer vision training for cars using the GTA world. I mean, it was crazy stuff, but it was kind of scattershot.
- DRDavid Rosenthal
Yeah, it was scattershot, and I guess what I'm saying is it was still in this narrow use-case world. They weren't doing anything approaching GPT at this point in time. Meanwhile, Google had kind of moved on.
- BGBen Gilbert
Yep.
- DRDavid Rosenthal
Now, one thing I do want to say in defense of OpenAI and everybody else in the field at the time, they didn't just have their heads in the sand. To do what transformers enabled you to do, which, Ben, you're gonna talk about in a sec, cost a lot in computing power. GPUs and Nvidia and the transformer made it possible, but to work with the size of models you're talking about, you're talking about spending an amount of money that certainly for a nonprofit, and anybody really except Google, was untenable.
- 44:10 – 51:24
OpenAI’s pivot and capitalization: from nonprofit lab to Microsoft-backed compute engine
- DRDavid Rosenthal
So with this, it starts to look like maybe this whole OpenAI boondoggle didn't actually accomplish anything, and the world's AI resources are more than ever just locked back into Google. So in 2018, Elon gets super frustrated by all this, basically throws a hissy fit and quits, and pieces out of OpenAI. There's a lot of drama around this that we're not gonna cover now. He may or may not have given an ultimatum to the rest of the team that he would either take over and run things or leave. Who knows? It's Elon. But whatever happened, this turns out to be a major catalyst for the rest of the OpenAI team and truly a history-turning-on-a-knifepoint moment. It was also a probably super bad decision by Elon, [chuckles] but again, story for another day.
- BGBen Gilbert
So there's this great explanation of what happened in the Semafor piece that, uh, we'll link to in our sources. The author says: "That fall, it became even more apparent to some people at OpenAI that the costs of becoming a cutting-edge AI company were going to go up. Google Brain's transformer had blown open a new frontier where AI could improve endlessly, but that meant feeding endless data to train it, a costly endeavor. OpenAI made a big decision to pivot toward these transformer models. On March 11, 2019, OpenAI announced it was creating a for-profit entity so it could raise enough money to pay for all the compute power necessary to pursue the most ambitious AI models. 'We want to increase our ability to raise capital while still serving our mission, and no preexisting legal structure that we know of strikes the right balance,' the company wrote at the time. OpenAI said it was capping profits for investors, with any excess going back to the original nonprofit. Less than six months later, OpenAI took a one billion dollar investment from Microsoft."
- DRDavid Rosenthal
Yeah, and I believe this is mostly, if not all, due to Sam Altman's influence and taking over here. So, you know, on the one hand, you can look at this sort of skeptically and say, "Okay, Sam, you took your nonprofit, and you converted it into an entity worth thirty billion dollars today." On the other hand, knowing this history now, this was kind of the only path they had. They had to raise money to get the computing resources to compete with Google, and Sam goes out and does these landmark deals with Microsoft.
- BGBen Gilbert
Yeah, truly amazing, and their opinion at the time of why they're doing this is basically, "This is gonna be super expensive. We still have the same mission to ensure that artificial general intelligence benefits all of humanity, but it's gonna be ludicrously expensive to get there, and so we need to basically be a for-profit enterprise and a going concern and have a business that funds our research eventually to pursue that mission."
- DRDavid Rosenthal
Yep, so 2019, they do the conversion to a for-profit company. Microsoft invests a billion dollars, as you say, and becomes the exclusive cloud provider for OpenAI, which is going to become highly relevant here for Nvidia. More on that in a minute. June of 2020, GPT-3 comes out. In September of 2020, Microsoft licenses exclusive commercial use of the underlying model for Microsoft products. 2021, GitHub Copilot comes out. Microsoft invests another two billion dollars in OpenAI, and then, of course, this all leads to November 30, 2022, in Jensen's words, "The AI heard around the world," OpenAI comes out with ChatGPT. As you said, Ben, the fastest product in history to reach a hundred million users. In January 2023, this year, Microsoft invests another ten billion dollars in OpenAI, announces they're integrating GPT into all of their products, and then in May of this year, GPT-4 comes out, and that basically catches us up to today. We eventually need to go do a whole 'nother episode [chuckles] about all the details here of OpenAI and Microsoft, but for today, the salient points are, one, thanks to all this, generative AI as a user-facing product emerges as this enormous opportunity. Two-... to facilitate that happening, you needed enormous amounts of GPU compute, obviously benefiting Nvidia. But just as important, three, it becomes obvious now that the predominant way that companies are gonna access and provide that compute is through the cloud. And the combination of those three things turns out to be basically the single greatest moment that could ever happen for Nvidia.
- BGBen Gilbert
Yes. So you're teeing all of this up, and so far I'm thinking, "So this is like the OpenAI and Microsoft episode? Like, what does this have to do with Nvidia?" And, God, there's a great Nvidia story here to be told. So let's get to the Nvidia side of it. But first, we wanna thank our friends at Statsig. So we've been talking this episode about names in AI that you know: Nvidia, OpenAI, Google, et cetera. But there's a name you probably don't know that's powering a lot of the AI wave behind the scenes, Statsig. A ton of big AI companies, OpenAI, Anthropic, Character.ai, rely on Statsig to test, deploy, and improve their models and applications. And how this happened is crazy because Statsig did not start as an AI company.
- DRDavid Rosenthal
Yeah. If you listened to our ACQ2 interview with founder and CEO Vijay Raji, you know that Statsig is a platform that combines feature flagging, experimentation, and product analytics. They help teams run experiments in their products, automate the analysis, launch new features, and analyze product performance. Their focus was on taking these pretty traditional product workflows and making them easier by giving teams one connected tool to move fast and make data-driven product decisions. So how does that relate to AI?
- BGBen Gilbert
Well, if you've ever built anything with these AI APIs, you know there are a ton of things to test, like the model version, the prompt, or the temperature, and adjusting these can have huge impact on the AI application's performance. So these AI companies have started using Statsig to measure the impact of changes to their models and the customer-facing applications using real user data. Even non-AI companies, like Notion and Figma, have been using Statsig to launch their AI features, ensuring that these new features drive successful outcomes for their businesses.
- DRDavid Rosenthal
In today's generative AI world, product decisions aren't just the product features anymore. They're literally, like, the weights and temperatures of the models underlying the products.
- BGBen Gilbert
Yep.
- DRDavid Rosenthal
So whether you're building with AI or not, Statsig can help your team ship faster and make data-driven product decisions. If you're a startup, they have a super generous free tier and a special program for venture-backed companies. If you're a large enterprise, they have clear, transparent pricing with no seat-based fees. Acquired community members can take advantage of a special offer, too, including five million free events a month and white-glove onboarding support. Just go visit statsig.com/acquired. That's S-T-A-T-S-I-G.com/acquired to get started on your data-driven journey.
- 51:24 – 53:28
Why Nvidia was uniquely prepared: re-architecting the data center into ‘the computer’
- BGBen Gilbert
Okay, so Nvidia.
- DRDavid Rosenthal
Okay, so we just said these three things that we've painted the picture of on the first part of the episode here, that, A, generative AI is, like, possible, a thing, and it's now getting traction, B, it requires an unbelievably massive amount of GPU compute to train, and, three, it looks like the predominant way that companies are going to use that compute is gonna be in the cloud. The combination of these three things is, I think, the most perfect example we've ever covered on this show of the old saying about luck being what happens when preparation meets opportunity [chuckles] for Nvidia here. So obviously, the opportunity is generative AI, but the preparation front, Nvidia has literally just spent the past five years working insanely hard to build a new computing platform for the data center, a GPU-accelerated computing platform to, in their minds, replace the old CPU-led, Intel-dominated x86 architecture in the data center. And for many years, I mean, they were getting some traction, right? And the data center segment was growing for Nvidia, but people were like: "Okay, you want this to happen, but, like, why is it gonna happen?"
- BGBen Gilbert
Right. There's these little workloads here and there that we'll toss you, Jensen, that we think can be accelerated by your cool GPUs, and then, you know, crazy things like crypto happened, and there was, like, AI researchers in academic labs that are using it as, you know, supercomputers. But for the longest time, the data center segment of Nvidia, it just wasn't clear that organizations had enormous parts of their software stack that they were gonna shift to GPUs. Like, why? What's driving this? And now we know what could be driving it, and that is AI.
- DRDavid Rosenthal
Uh, not only could be, but if you look at their most recent quarter, absolutely freaking is.
- 53:28 – 1:00:49
Computer architecture detour: von Neumann bottleneck and why memory/networking now dominates
- BGBen Gilbert
Okay, so now it begs the question, why is it driving it? And David, are you open to me giving a little computer science lecture on computer architecture?
- DRDavid Rosenthal
Ooh, please do.
- BGBen Gilbert
All right, I need to do my best professor impression here.
- DRDavid Rosenthal
Dude, I loved computer science in college. They were my favorite classes.
- BGBen Gilbert
I will say, doing these episodes, this TSMC, it really does bring back the thrill of being in a CS lecture and being like: "Oh, that's how that works!" Like, it's just really fun. So let's take a step back and consider the classic computer architecture, the von Neumann architecture. Now, the von Neumann architecture is what most computers, most CPUs, are based on today, where they can store a program in the computer's memory and run that program. You can imagine why this is the dominant architecture. Otherwise, we'd need a computer that is specialized for every single task.... The key thing to know is that the memory of the computer can store two different things: the data that the program uses and the instructions of the program itself, the literal lines of code. And in this example we're about to paint, all of this is wildly simplified, because I don't want to get into caching and speeds of memory and, you know, where memory is located and not located, so let's just keep it simple. So the processor in the von Neumann architecture executes this program, written in assembly language, which is the language that compiles down to the byte code that the processor itself can speak. So it's written in an instruction set architecture, an ISA, from Arm, for example.
- DRDavid Rosenthal
Or Intel before that.
- BGBen Gilbert
[chuckles] Yes. And each line of the program is very simplistic. So we're gonna consider this example where I'm gonna use some assembly language pseudocode to add the numbers two and three to equal five.
- DRDavid Rosenthal
Ben, are you about to program live on Acquired?
- BGBen Gilbert
[laughing] Well, it's pseudo-assembly language code. So the first line is we're gonna load the number two from memory. We're gonna fetch it out of memory, and we're gonna load it into a register on the processor. So now we've got the number two actually sitting right there on our CPU, ready to do something with. That's line of code number one. Two, we're gonna load the number three in exactly the same fashion into a second register. So we've got two CPU registers with two different numbers. The third line, we're gonna perform an add operation, which performs the arithmetic to add the two registers together on the CPU and store the value in some either third register or into one of those registers. So that's a more complex instruction, since it's arithmetic that we actually have to perform, but these are the things that CPUs are very good at, doing math operations on data fetched from memory. And then the fourth and final line of code in our example is we are going to take that five that has just been computed and is currently held temporarily in a register on the CPU, and we are gonna write that back to an address in memory. So the four lines of code are load, load, add, store.
- DRDavid Rosenthal
This all sounds familiar to me.
- BGBen Gilbert
So you can see each of those four steps is capable of performing one and only one operation at a time, and each of these happens with one cycle of the CPU. So if you've heard of gigahertz, that's the number of cycles per second. So a one gigahertz computer could handle the simple program that we just wrote two hundred and fifty million times in a single second. But you can see something going on here. Three of our four clock cycles are taken up by loading and storing data to memory. Now, this is known as the von Neumann bottleneck, and it is one of the central constraints of AI, or at least it has been historically. Each step must happen in order and only one at a time. So in this simple example, it actually would not be helpful for us to add a bunch more memory to this computer. I can't do anything with it. It's also only incrementally helpful to increase the clock speed. If I double the clock speed, I can only execute the program twice as fast. If I need, like, a million X speed up for some AI work that I'm doing, I'm not gonna get it there with just a faster clock speed. That's not gonna do it. And it would, of course, be helpful to increase the speed at which I can read and write to memory, but I'm kind of bound by the laws of physics there. There's only so fast that I can transmit data over a wire. Now, the great irony of all of this is that the bottleneck actually gets worse over time, not better, 'cause the CPUs get faster and the memory size increases, but the architecture is still limited, so this one pesky single channel, known as a bus, I don't actually get to enjoy the performance gains nearly as much as I should, 'cause I'm jamming everything through that one channel, and it only gets to sort of be used one time per every clock cycle. So the magical unlock, of course, is to make a computer that is not a von Neumann architecture, to make programs executable in parallel and massively increase the number of processors or cores. And that is exactly what Nvidia did on the hardware side and all these AI researchers figured out how to leverage on the software side. But interestingly, now that we've done that, David, the constraint is not the clock speed or the number of cores anymore. For these absolutely enormous language models, it's actually the amount of on-chip memory-
- DRDavid Rosenthal
Yeah
- BGBen Gilbert
... that concerns us.
- DRDavid Rosenthal
I thought you were going there.
- BGBen Gilbert
[chuckles]
- DRDavid Rosenthal
And this is why the data center and what Nvidia has been doing is so important.
- BGBen Gilbert
Yes. There's this amazing video that we'll link to on the Asianometry YouTube channel that we linked to also on the TSMC episode, but the constraint today is actually in how much high-performance memory is available on the chip. These models need to be in memory all at the same time, and they take up hundreds of gigabytes. So while memory has scaled up, I mean, we're gonna get flashing all the way forward, the H100's on-chip RAM is, like, eighty gigabytes. The memory hasn't scaled up nearly as fast as the models have actually scaled in size. The memory requirements for training AI are just obscene, so it becomes imperative to network multiple chips and multiple servers of chips and multiple racks of servers of chips together into one single computer, and I'm putting computer in air quotes there, in order to actually train these models. It's also worth noting we can't make the memory chips any bigger. Due to a quirk of the extreme ultraviolet photolithography that we talked about, the EUV, on the TSMC episode, chips are already the full size of the reticle.... It's a physics and wavelength constraint. You really can't etch chips larger without some new invention that we don't have commercially viable yet. So what it ends up meaning is you need huge amounts of memory, very close to the processors, all running in parallel with the fastest possible data transfer. And again, this is a vast oversimplification, but you kind of get the idea of why all of this becomes so important.
- 1:00:49 – 1:11:52
Nvidia’s data center ‘three-legged stool’: Mellanox/InfiniBand, Grace CPU, Hopper + CoWoS
- DRDavid Rosenthal
Okay, so back to the data center, and here's what Nvidia is doing that I don't think anybody else out there is doing, and why it's so important for them that all of this new generative AI world, this new computing era, as Jensen dubs it, runs in the data center. So Nvidia has done three things over the last five years. One, and probably most importantly related to what you're talking about, Ben, they made one of the best acquisitions of all time back in 2020, and nobody had any idea. They bought a quirky little networking company out of Israel called Mellanox.
- BGBen Gilbert
Well, it wasn't little. They paid seven billion dollars for it.
- DRDavid Rosenthal
Okay. Yeah, [chuckles] and it was already a public company, right?
- BGBen Gilbert
It was, yep.
- DRDavid Rosenthal
Yeah, but it was definitely quirky. Now, what was Mellanox? Mellanox's primary product was something called InfiniBand, which we talked about a lot with Chase Lochmiller on our ACQ2 episode with him from Crusoe.
- BGBen Gilbert
And actually, InfiniBand was an open-source standard, or managed by a consortium. There were a bunch of players in it, but the traditional wisdom was, while InfiniBand is way faster, way higher bandwidth, a much more efficient way to transfer data around a data center, at the end of the day, Ethernet is the lowest common denominator, and so everyone had to implement Ethernet anyway, and so most companies actually exited the market, and Mellanox was kind of the only InfiniBand spec provider left.
- DRDavid Rosenthal
Yeah. So you said, wait, what is InfiniBand? It is a competing standard to Ethernet. It is a way to move data between racks in a data center. And back in 2020, everybody was like: Ethernet's fine. Why do you need more bandwidth than Ethernet between racks in a data center? What could ever require thirty-two hundred gigabits a second of bandwidth running down a wire in [chuckles] a data center? Well, it turns out, if you're trying to address hundreds, maybe more than hundreds, of GPUs as one single compute cluster to train a massive AI model, yeah, you want really fast data interconnects between them.
- BGBen Gilbert
Right. People thought, "Oh, sure, for supercomputers, for these academic purposes, but what the enterprise market needs in my shared cloud computing data center is Ethernet, and that's fine, and most workloads are gonna happen right there on one rack, and maybe, maybe, maybe things will expand to multiple computers on that rack. But certainly, they won't need to network multiple racks together." And Nvidia steps in, and you got Jensen saying, "Hey, dummies, the data center is the computer. Listen to me when I tell you the whole data center needs to be one computer." And when you start thinking that way, you start thinking, "Geez, we're really gonna be cramming huge amounts of data through wires that are going between these ra-- Like, how can we sort of think about them as if it's all sort of on-ship memory, or as close as we can make it to on-ship memory, even though that's in a box located three feet away?"
- DRDavid Rosenthal
Yep. So that's piece number one of Nvidia's grand data center plan over the last five years. Piece number two is in September 2022, Nvidia makes a quite surprising announcement of a new chip. Not just a new chip, an entirely new class of chips that they are making called the Grace CPU processor. Nvidia is making a CPU! This is, like, heretical.
- BGBen Gilbert
"But Jensen, I thought all computing was gonna be accelerated. What are we doing here on these Arm CPUs?"
- DRDavid Rosenthal
Yeah, these Grace CPUs are not for putting in your laptop. [laughing] They are for being the CPU component of your entire data center solution that is specifically from the ground up designed to orchestrate with these massive GPU clusters.
- BGBen Gilbert
This is the end game of a ballet that has been in motion for thirty years. Remember when the graphics card was subservient to the PCIe slot in Intel's motherboard? And then eventually, you know, we fast-forward to the future, Nvidia makes these GPUs that are these beautiful standalone boxes in your data center, or perhaps these little workstations that sit next to you while you're doing graphics programming, while you're directly programming your GPU. And then, of course, they need some CPU to put in that, so they're using AMD or Intel, or they're licensing some CPU, and now they're saying, "You know what? We're actually just gonna do the CPU, too. So now we make a box, and it's a fully integrated Nvidia solution with our GPUs, our CPUs, our NVLink between them, our InfiniBand to network it to other boxes, and, you know, welcome to the show."
- DRDavid Rosenthal
One more piece to talk about, the third leg of the stool of their strategy, before we get to what it all means, that I think you're about to go to. Spoiler alert: you say solution, I hear gross margin. [laughing] The third part of it is the GPUs. Up until Nvidia's current GPU generation, the Hopper generation of GPUs for the data center, there was only one GPU architecture at Nvidia, and that same architecture and those same chips from the same wafers made at TSMC-... Some of them went to consumer gaming graphics cards, and some of those dies went to A100 GPUs in the data center. It was all the same architecture. Starting in September of 2022, they broke out the two business lines into different architectures. So there's the Hopper architecture, named after great computer scientist Grace Hopper. I think Rear Admiral in the US Navy, Grace Hopper. Get it? Grace, CPU, Hopper-
- BGBen Gilbert
Heyo
- DRDavid Rosenthal
... GPU, Grace Hopper, the H100s, that was for the data centers. And then on the consumer side, they start a whole new architecture called Lovelace, after Ada Lovelace, and that is the RTX 40XX. So you buy, uh, you know, top-of-the-line RTX 40 what-have-you gaming card right now, that is no longer the same architecture as the H100s that are powering ChatGPT. It's got its own architecture. This is a really big deal, because what they do with the Hopper architecture is they start using what's called chip-on-wafer-on-substrate, COWOS.
- BGBen Gilbert
COWOS. When you start talking to the real semi nerds, that's when they start busting out the COWOS conversation.
- DRDavid Rosenthal
This is when a certain segment of our listeners are gonna get really excited. So essentially what this is, back to this whole concept of memory being so important for GPUs and for AI workloads, this is a way to stack more memory on the GPU chips themselves, essentially by going vertical in how you build the chips. This is the absolute bleeding-edge technology that is coming out of TSMC. And by Nvidia bifurcating their chip architectures into a gaming segment that does not have this latest COWOS technology, this allows them to monopolize, like, a huge amount of TSMC's capacity to make the COWOS chips specifically for these H100s, which allows them to have way more memory than other GPUs on the market.
- BGBen Gilbert
Yes. So this gets to the point of why can't they seem to make enough chips right now? Well, it's literally a TSMC capacity problem. So there's these two components that are extremely related, that you're talking about, the COWOS, chip-on-wafer-on-substrate, and the high-bandwidth memory. So there's this great post from SemiAnalysis, where the author points out a 2.5D chip, which is basically how you assemble this COWOS stuff to get the memory really close to the processor. And of course, 2.5D, it is literally 3D, but 3D means something else. It's even more 3D, so they came up with this 2.5D denominator. Anyway, the 2.5D chip packaging technology from TSMC is where you take multiple active silicon dies, like the logic chips and the stack of high-bandwidth memory, and they stack them on one piece of silicon. And there's more complexity here, but the important thing is, COWOS is the most popular technology for GPUs and AI accelerators for packaging these chips, and it's the primary method to co-package high-bandwidth memory... Again, remember, think back to the thing that's most important right now, is get as much high-bandwidth memory as you can, closest to the CPU, next to the logic, to get the most performance for training and inference. So COWOS represents, right now, about ten to fifteen percent of TSMC's capacities, and many of the facilities are custom-built for exactly these types of chips that they're producing. So when Nvidia needs to reserve more capacity, there's a pretty good chance that they've already reserved some large part of the ten to fifteen percent of TSMC's total footprint, and TSMC needs to, like, go make more fabs in order for [chuckles] Nvidia to have access to more COWOS-capable capacity.
- DRDavid Rosenthal
Yeah, which, as we know, it takes years for TSMC to do this.
- BGBen Gilbert
Yep. There are more experimental things that are happening. Like, I would be, uh, remiss not to mention, there are actually experiments of doing compute in-memory. Like, as we shift away from von Neumann, and sort of all bets are off now that we're open to new computing architectures, there are people exploring, Well, what if we just process the data where it is in memory, instead of doing the very lossy, expensive, energy-intensive thing of moving data over the copper wire to get it to the CPU? All sorts of trade-offs in there, but it is very fun to sort of dive into the academic computer science world right now, where they really are rethinking, like, "What is a computer?"
- DRDavid Rosenthal
So these three things that Nvidia has been building, the dedicated Hopper data center GPU architecture, the Grace CPU platform, the Mellanox-powered networking stack, they now have a full suite solution for generative AI data centers. [chuckles] And then, when I say solution-
- BGBen Gilbert
I hear margins. But let's be clear, you don't need to offer some sort of solution to get high margins if you're Nvidia. Price is set where supply meets demand, and they're adding as much supply as they possibly can right now. Like, believe me, for all sorts of reasons, Nvidia wants everyone who wants H100s to have H100s. But for now, the price is kind of like a, "I'll write you a blank check, and Nvidia, you write whatever you want on the check." So their margins are crazy right now, just literally because there's way more demand than supply for these things.
- 1:11:52 – 1:27:40
Productization and monetization: H100 economics, DGX systems, SuperPODs, and DGX Cloud
- DRDavid Rosenthal
Yes. Okay, so let's break down what they're actually selling. So like you were saying, Ben, of course, you can, and lots of people do, just go buy H100. So you're like, "I don't care about the Grace CPU. I don't care about this Mellanox stuff. I'm running my own data center. I'm really good at it."
- BGBen Gilbert
And the people who are most likely to do this are the hyperscalers, or as Nvidia refers to them, the CSPs, the cloud service providers.
- DRDavid Rosenthal
This is AWS, this is Azure, this is Google, this is Facebook for their internal use.
- BGBen Gilbert
Like, "Nvidia, don't give me one of these DGX servers that you assemble. Just give me the chip, and I will integrate it the way that I want to integrate it."
- DRDavid Rosenthal
"I am a world-class data center architect and operator. I don't want your solution, I just want your chips." So they sell a lot of those. Now-... Nvidia, of course, has also been seeding new cloud providers out there in the ecosystem, like our friends at Crusoe, also CoreWeave, and Lambda Labs, if you've heard of them. These are all new GPU-dedicated clouds that Nvidia is working closely with. So they're selling H100s and A100s before that to all these cloud providers.
- BGBen Gilbert
But let's say you are an arbitrary company in the Fortune 500 that is not a technology company, and, my God, do you not want to miss the boat on generative AI, and you've got a data center of your own. Well, Nvidia has a DGX for you.
- DRDavid Rosenthal
Yes, they do. Full GPU-based supercomputer solution in a box that you can just plug right into your data center, and it just works. There's nothing else on the market like this.
- BGBen Gilbert
And it all runs CUDA. It is all speaking the exact language of the entire ecosystem of developers that know exactly how to write software for this thing.
- DRDavid Rosenthal
Which means that whatever developers you already had who were working on AI or anything else, everything they were working on is just gonna come right over and run within your brand-new, shiny AI supercomputer 'cause it all runs CUDA.
- BGBen Gilbert
Amazing.
- DRDavid Rosenthal
More on CUDA in a minute, but as we said, [chuckles] you say solution, I hear gross margin. Nvidia sells these DGX systems for, like, a hundred and fifty to three hundred thousand dollars a box. That's wild. And now, with all these three new legs of the stool, Hopper, Grace, and Mellanox, these systems are just getting way more integrated, way more proprietary, and way better. So if you wanna buy a new top-of-the-line DGX H100 system, the price starts at five hundred thousand dollars for one box. And if you wanna buy the DGX GH200 SuperPOD, this is the AI wall that Jensen recently unveiled, the huge, like, room full of AI.
- BGBen Gilbert
And it's, like, twenty racks wide. Imagine an entire row in a data center.
- DRDavid Rosenthal
Yes, this is two hundred and fifty-six Grace Hopper DGX racks all connected together [chuckles] in one wall. They're billing this as the first turnkey AI data center that you can just buy and can train a trillion-parameter GPT-4 class model. The pricing on that is: call us. [laughing]
- BGBen Gilbert
[laughing] Of course it is.
- DRDavid Rosenthal
But I'm imagining, like, hundreds of millions of dollars. Like, I doubt it's a billion, but hundreds of millions, easily.
- BGBen Gilbert
Wild. Well, let's talk about the H100. I've got a baseball card right here on this, uh, insane thing that they've built. So they launched it in September 2022. It's the successor to the A100. One GPU, one H100, costs forty thousand dollars. So that's how you get to that price point you're talking about.
- DRDavid Rosenthal
That's what they're selling to Amazon and Google and Facebook.
- BGBen Gilbert
Right. And you mentioned that five hundred thousand dollar price point. The five hundred thousand dollars is the eight forty thousand dollar H100s in a box with the gray CPU and, you know, the nice bow around it.
- DRDavid Rosenthal
Yep, which do the math on that. So eight times forty thousand, that's three hundred and twenty thousand dollars. So that's essentially an extra hundred and eighty thousand dollars of margin that Nvidia is getting out of selling the solution. It's an Arm CPU. It doesn't cost them anything to make that. [chuckles]
- BGBen Gilbert
And these forty thousand dollar H100s have margin of their own. So, like, every time they bundle more, there's more margin [chuckles] in the fully assembled... I mean, that's literally bundle economics. You are entitled to margin when you bundle more things together and provide more value for customers. But just to, like, illustrate the way that this pricing works, so the reason you want an H100 is they're thirty times faster than an A100, which, mind you, is only, like, two and a half years older. It is nine times faster for AI training. The H100 is literally purpose-built for training LLMs, like the full self-driving video stuff. It's super easy to scale up. It's got eighteen and a half thousand CUDA cores. Remember when we were talking about the von Neumann example earlier? Like, that is one computing core that is able to handle, you know, those four assembly language instructions. This one H100, which they're calling a GPU, has eighteen and a half thousand cores that are capable of running CUDA software. It's got six hundred and forty tensor cores, which are highly specialized for matrix multiplication. They have eighty streaming multiprocessors. So what are we up to here? Close to twenty thousand unique cores on this thing. It's got meaningfully higher energy usage than the A100. I mean, a big takeaway here is that Nvidia is massively increasing the power requirement every time they come out with the next generation. They're both figuring out how to push the edge of physics, but they're also constrained by physics. Some of this stuff is only possible with way more energy. This thing weighs seventy pounds. This is one H100.
- DRDavid Rosenthal
Jensen makes a big deal about this every keynote that he gives, like, "Oh, I can't lift it."
- BGBen Gilbert
It's got a quarter trillion transistors across thirty-five thousand parts. It requires robots to assemble it. Not only does it require physical robots to assemble it, it requires AI to design it. They're actually using AI to design the chips themselves now. I mean, they have completely reinvented the notion of what a computer is.
- DRDavid Rosenthal
Totally. And this is all part of Jensen's pitch here to customers. "Yes, our solutions are very expensive." However, he uses the line that he loves, "The more you buy, the more you save."
- BGBen Gilbert
If you could get your hands on some.
- DRDavid Rosenthal
Right. But what he means by that is, like, okay, say, you're McDonald's, and you're trying to build a generative AI, so that, I don't know, customers can order or something. You're using it in your business. If you were gonna try and build and run that in your existing data center infrastructure-... It would take so much time and cost you so much more over the long run in compute than if you just went and bought my SuperPOD here. You can plug and play and have it up and running in a month.
- BGBen Gilbert
Yep, and by the fact that this is all accelerated computing, the things you're doing on it, you literally wouldn't be able to do otherwise, or might take you a lot more energy, a lot more time, a lot more cost. There is a very valid story to buying and running your workloads here, or renting from any of the cloud service providers, and running your workloads here is more performant because the results just happen much faster, much cheaper, or at all.
- DRDavid Rosenthal
Yep. You mentioned energy here. Like, this is also Jensen's argument. He's like: "Yes, these things take a ton of energy, but the alternative takes even more energy." So we are actually saving energy if you assume this stuff is going to happen. Now, there's a bit of a caveat here in that it can't happen except on [chuckles] these types of machines, so he enabled this whole thing, but he has a point.
- BGBen Gilbert
Oh, I totally buy it, though. I mean, I think there's a very real case around, look, you only have to train a model once, and then you can do inference on it over and over and over again. I mean, the analogy I think makes a lot of sense for model training is to think about it as a form of compression. LLMs are turning the entire internet of text into a much smaller set of model weights. This has the benefit of storing a huge amount of usefulness in a small footprint, but also enabling a very inexpensive amount of compute, again, relatively speaking, in the inference step for every time that you need to prompt that model for an answer. Of course, the trade-off you're making there is once you encode all of the training data into the model, it is very expensive to redo it, so you better do it right the first time or figure out little ways to modify it later, which a lot of ML researchers are working on. But I always think a reasonable comparison here is to compress a zillion-layer Photoshop file. For anybody that's ever dealt with, "Oh, I've got a three-gigabyte Photoshop file," well, that's not a thing you're gonna send to a client. You'm gonna compress it into a JPEG, and you're gonna send that, and the JPEG is, in many ways, more useful as a compressed facsimile of the original layers comprising the Photoshop file, but the trade-off is you can never get from that compressed little JPEG back to the original thing. So I think the analogy here is, like, you're saving everyone from needing to make the full PSD every time because you can just use the JPEG the vast, vast majority of the time.
- DRDavid Rosenthal
So hopefully, we've now painted a relatively coherent picture of both the advances that made the generative AI opportunity possible, that it has truly become a real opportunity, and why Nvidia, even above the obvious reasons, was just so well-positioned here, particularly because of the data center-centric nature of these workloads and that they had been working so hard for the past five years to fundamentally re-architect the data center.
- BGBen Gilbert
Yep.
- 1:27:40 – 1:37:01
2023 financial shockwave and the new $1T narrative: data center CapEx as the real TAM
- DRDavid Rosenthal
So all of this brings us to 2023. In May of this year, NVIDIA reported their Q1 fiscal twenty-four earnings. NVIDIA's on this weird January fiscal year-end thing, so Q1 twenty-four is essentially Q1 twenty-three, but anyway, in which revenue was up nineteen percent quarter over quarter to seven point two billion, which is great, 'cause remember, they had a terrible end of 2022, with the write-offs and crypto falling off a cliff and all that.
- BGBen Gilbert
Yes, it's amazing that in that Stratechery interview, when was that? In, uh, March of 2023, Jensen said, "Last year was unquestionably a disappointing year." This is the year ChatGPT was released! It is wild, the roller coaster this company has been on.
- DRDavid Rosenthal
The timeframe is so compressed here.
- BGBen Gilbert
And part of that, of course, is Ethereum moving to proof of stake, the end of the crypto thing for NVIDIA, which I'm sure they're actually thrilled about. But part of it was they also put in a ton of pre-orders for capacity with TSMC that then they thought they weren't gonna need, so they had to write down. So from an accounting perspective, it looks like a big loss, like a really big blemish on their finances last year. But now, oh, my God, are they glad that they reserved all that capacity.
- DRDavid Rosenthal
Yep, it's actually going to be quite valuable. So speaking of, you know, this Q1 earnings is, like, great, up nineteen percent quarter over quarter, but then they dropped the bombshell. Due to unprecedented demand for generative AI compute in data centers, NVIDIA forecasts Q2 revenue of eleven billion dollars, which would be up another fifty-three percent quarter over quarter over Q1 and sixty-five percent year over year. The stock goes nuts.
- BGBen Gilbert
Twenty-five percent in after-hours trading.
- DRDavid Rosenthal
Yep.
- BGBen Gilbert
This is a trillion-dollar company, or at least this made them a trillion-dollar company, but, like, a company that was previously valued at around eight hundred billion dollars popped twenty-five percent after earnings.
- DRDavid Rosenthal
Well, and it's even crazier than that. Back when we did our episodes last April, NVIDIA was the eighth-largest company in the world by market cap, had about a six hundred and sixty billion dollar market cap. That was down slightly off the highs, but that was kind of the order of magnitude back then. It crashed down below three hundred billion, and then within a matter of months, it's [chuckles] now back up over a trillion. Just wild. And then, all of this culminates last week, at the time of this recording, when NVIDIA reports Q2 fiscal twenty-four earnings. And this earnings release, we usually don't talk about, like, individual earnings releases on Acquired, 'cause, like, in the long arc of time, who cares? This was a historic event. I think this was one of, if not the, most incredible earnings release by any scaled public company ever. Seriously, no matter what happens going forward, last week was a historic moment.
- BGBen Gilbert
The thing that blows my mind the most is that their data center segment alone did ten billion dollars in the quarter. That's more than doubling-... off of the previous quarter. In three months, they grew from four-ish billion to ten billion of revenue in that segment. And revenue only happens when they deliver products to customers. This isn't pre-orders, this isn't clicks, this isn't wave your hands around stuff. This is, "We delivered stuff to customers, and they paid us an additional six billion dollars this quarter than they did last quarter."
- DRDavid Rosenthal
So here are the full numbers. For the quarter, total company revenue of thirteen point five billion, up eighty-eight percent from the previous quarter and over a hundred percent from a year ago. And then, Ben, like you said, in the data center segment, revenue of ten point three billion. So ten point three out of thirteen point five for a segment that basically didn't exist five years ago for the company. That's up a hundred and forty-one percent from Q1 and a hundred and seventy-one percent from a year ago. This is ten billion dollars. That kind of growth at this scale, I've never seen anything like it.
- BGBen Gilbert
No.
- DRDavid Rosenthal
Neither has the market.
- BGBen Gilbert
That's right.
- DRDavid Rosenthal
And so this, this is the first time I noticed it. Jensen had talked about this in Q1 earnings, so it wasn't the first time, but he brings back the trillion-dollar TAM. Not in a slide, I think this time, he just talks about it.
- BGBen Gilbert
No, but in a new way that I think is a better way to slice it.
- DRDavid Rosenthal
This time, it's different. You know, look, we'll spend a while here now talking about what we think about this, but this is very different. This time, he frames Nvidia's trillion-dollar opportunity as the data center, and this is what he says: There is one trillion dollars worth of hard assets sitting in data centers around the world right now.
- BGBen Gilbert
Growing at two hundred and fifty billion a year.
- DRDavid Rosenthal
Annual spend on data centers to update and add to that CapEx is two hundred and fifty billion dollars a year, and Nvidia has certainly the most cohesive, fulsome, and coherent platform to be the future of what those data centers are gonna look like for a large amount of compute workloads. This is a very different story than like, "Oh, we're gonna get one percent of this hundred trillion dollars of industry out there." [chuckles]
- BGBen Gilbert
And the thing you have to believe now, 'cause whenever someone paints a picture, you say, "Okay, what do I have to believe?" The thing you have to believe is there is real user value being created by these AI workloads and the applications that they are creating. And there's pretty good evidence. I mean, ChatGPT made it so OpenAI is rumored to be doing over a billion-dollar run rate now, maybe multiple single-digit billions, and still growing meaningfully. And so that is like the shining example. Again, that's the Netscape Navigator here of this whole boom. But the bet, especially with all these Fortune five hundreds, is that there are going to be GPT-like experiences in everyone's private applications, in a zillion other public interfaces. I mean, Jensen frames it as, in the future, every application will have a GPT front end. It will be a way that you decide that you wanna interact with computers that is more natural. And I don't think he means like, versus clicking buttons. I think he means everyone can kinda become a programmer, but the programming language is English. And so when you're sort of like: Well, why is everyone spending all of this money? It is that the world's executives, with the purchasing power to go write a ten-billion-dollar check last quarter to Nvidia for all this stuff, wholeheartedly believes from the data they've seen so far, that this technology is gonna change the world enough for them to make these huge bets. And the thing that we don't know yet is: Is that true? Is the GPT-like experiences going to be an enduring thing for the far future or not? There's pretty good evidence so far that people like this stuff, and that it's quite useful in transforming the way that, you know, everyone lives their lives and goes about day-to-day, and does their jobs, and goes through school, and, you know, on and on and on. But that is the thing you have to believe.
- DRDavid Rosenthal
So we have a lot to talk about with regard to that in analysis. But before we move to analysis, I think we should talk about another one of our very favorite companies here at Acquired, Blinkist from GoOne.
- BGBen Gilbert
Yes, absolutely, and listeners, you know we're doing something very cool with them this season. As you know, Blinkist takes books and condenses them into the most important points, so you can read or listen to the summaries.
- DRDavid Rosenthal
It's almost like a large language model compression for books.
- BGBen Gilbert
There you go. So a couple cool things we're doing, one of which is David and I have made a Blinkist page that represents our bookshelf. So if you wanna read the books that influence us, you can go to blinkist.com/acquired. And for this particular Nvidia episode, Blinkist has made a special collection for us. Amazingly, there are not really books about the history of Nvidia itself, at least not yet.
- DRDavid Rosenthal
Which is unbelievable.
- BGBen Gilbert
Yeah, but there are plenty on AI through the years. So if you go to blinkist.com/nvidia, you can find books by Garry Kasparov, Kai-Fu Lee, and Cade Metz, who we already mentioned earlier on the show.
- DRDavid Rosenthal
Who wrote the amazing Wired article.
- BGBen Gilbert
Yep. You'll, of course, get free access to that Nvidia Blinkist collection, and anyone who signs up through that link or uses the coupon code NVIDIA will then get a fifty percent off premium subscription to all sixty-five hundred titles in their library.
- DRDavid Rosenthal
And for those of you who are leaders at companies, check out Blinkist for Business. This gives your whole team the power to tap into world-class knowledge right from their phones, anytime they need it, available at blinkist.com/business. Blinkist is a great way for your team to master soft skills, which, if you believe that this new AI world is coming, [chuckles] is going to be even more important for the humans in your workforce.
- 1:37:01 – 2:53:08
Moats and competitive dynamics: CUDA as platform, TSMC capacity as cornered resource, and the bull/bear debate
- BGBen Gilbert
Yes. Our huge thanks to Blinkist and their parent company, GoOne, where David and I are both huge fans and angel investors. GoOne and Blinkist are both amazing ways for your company to get access to the most engaging and compelling content in the world. Our thanks to both of them and links in the show notes.... Okay, so David, analysis. We gotta talk about CUDA before we even start analyzing anything else here. Talked about a lot of hardware so far on this episode, but there's this huge piece of the Nvidia puzzle that we haven't talked about since part two, and CUDA, as folks know, was the initiative started in 2006 by Jensen and Ian Buck, and a bunch of other folks on the Nvidia team, to really make a bet on scientific computing, that people could use graphics cards for more than just graphics, and they would need great software tools to help them do that. It also was the glimmer in Jensen's eye of, "Ooh, maybe I can build my own relationship with developers," and, you know, there can be this notion not of a Microsoft or an Intel developer who happens to be able to, you know, have a standard interface to my chip, but I can have my own developer ecosystem, which has been huge for the company. So CUDA has become the foundation that everything that we've talked about, all the AI applications, are written on top of today. So, you know, you hear Jensen in these keynotes reference CUDA, the platform, CUDA, the language, and I spent some time trying to figure out, like, when I was watching developer sessions and, like, literally learning some CUDA programs, what is the right way to characterize it?
- DRDavid Rosenthal
And what is the right way to characterize it today? Because it has evolved a lot.
- BGBen Gilbert
Yes. So today, CUDA is, starting from the bottom and going up, a compiler, a runtime, a set of development tools, like a debugger and a profiler. It is its own programming language, CUDA C++. It has industry-specific libraries. It works on every card that they ship and have shipped since 2006, which is a really important thing to know, and if you're a CUDA developer, your stuff works on everything, anything Nvidia, all this unified interface. It has many layers of abstractions and existing libraries that are optimized, so these libraries have code that you can call to keep your development work short and simple instead of reinventing the wheel. So, you know, there are things that you can decide that you wanna write in C++ and just rely on their compiler to make it run well on Nvidia hardware for you, or you can write stuff in their native language and try to implement things yourself in CUDA C++. The answer is, it's incredibly flexible, it is very well-supported, and there's this huge community of people that are developing with you and building stuff for you to build on top of. If you look at the number of CUDA developers over time, it was released in 2006. It took four years to get the first hundred thousand people. Then, by 2016, thirteen years in, they got to a million developers. Then just two years later, they got to two million. So thirteen years to add their first thirteen million, then two years to add their second. 2022, they hit three million developers, and then just one year later, in May of 2023, CUDA has four million registered developers. So at this point, there's a huge moat for Nvidia, and I think when you talk to folks there, and frankly, when we did talk to folks there, they don't describe it this way. They don't think about it like, "Well, CUDA is our moat versus competitors." It's more like, "Well, look, we envisioned a world of accelerated computing in the future, and we thought, there are way more workloads that should be parallelized and made more efficient, that we want people to run on our hardware, and we need to make it as easy as possible for them to do that. And we're going to go to great lengths and have one, two thousand people that work at our company that are gonna be full-time software engineers building this programming language, and compiler, and foundation, and framework, and everything on top of it, to let the maximum number of people build on our stuff." That is how you build a developer ecosystem. It's different language, but the bottom line is, they have a huge reverence for the power that it gives them at the company.
- DRDavid Rosenthal
This is something we touched on on our last episode, but has really crystallized for me in doing this one. Nvidia thinks of themselves as, and I believe is, a platform company, especially this week [chuckles] after the blowout earnings and everything that happened this quarter, and the stock and whatnot. Sort of a popular take out there that you've been seeing a lot is, "Oh, we've seen this movie before." This happened with Cisco. You could say over a longer timescale, this happened with Intel. Yeah, these hardware providers, these semiconductor companies, they're hot when they're hot, and people wanna, you know, spend CapEx, and then when they're not hot, they're not hot. But I don't think that's quite the right way to characterize Nvidia. They do make semiconductors, and they do make data center gear, but really, they are a platform company. The right analogy for Nvidia also is Microsoft. [chuckles] They make the operating system, they make the programming environment, they make many of the applications.
- BGBen Gilbert
Right. Cisco doesn't really have developers. Intel never had developers. Microsoft had developers, and Intel had Microsoft, but Intel didn't have developers. Nvidia has developers. I mean, they've built a new architecture that is not a von Neumann computer. They've bucked fifty years of progress, and instead, every GPU has a stream processor unit, and as you'd imagine, you need a whole new type of programming language, and compiler, and everything to deal with this new computing model. And that's CUDA, and it freaking works, and there's all these people that develop their livelihood in it.
- DRDavid Rosenthal
You talk to Jensen, and you talk to other people at the company, and they will tell you, "We are a foundational computer science company. We're not just slinging hardware here."
- BGBen Gilbert
Yeah, I mean, it's interesting. They're a platform company, for sure. They're also a systems company. They're effectively selling mainframes. I mean, it's not that different than IBM way back when. They're trying to sell you a, you know, a hundred million dollar wall [chuckles] that goes in your data center, and it's all fully integrated, and it all just works.
- DRDavid Rosenthal
... Yeah, and maybe IBM actually is a really good analogy, like old school IBM here. They make the underlying technology, they make the hardware, they make the silicon, they make the operating system for the silicon, they make the solutions for customers, they make everything, and they sell it as a solution.
- BGBen Gilbert
Yep. Okay, so a couple other things to catch us up here as we're starting analysis. One big point I wanna make is, let's look at a timeline, 'cause I didn't discover this until, like, two hours before we started recording. In March of 2019, Nvidia announced they were acquiring Mellanox for seven billion dollars in cash, and I think Intel was considering the purchase, and then Nvidia came in and kinda blew them out of the water. And it is fair to say nobody really understood what Nvidia was going to do there, and why it was so important, but the question is, why? Well, Nvidia knew that these new models coming out would need to run across multiple servers, multiple racks, and they put a huge level of importance on the bandwidth between the machines. And, of course, how did they know that? Well, in August of 2019, Nvidia released what was at the time, the largest transformer-based language model called Megatron. Eight point three billion parameters trained on five hundred and twelve GPUs for nine days, which at the time, at retail, would have cost something like half a million dollars to train, which at the time was a huge amount of money to spend on model training, which is, what? Only four years ago, but-
- DRDavid Rosenthal
Today, that's quaint.
- BGBen Gilbert
Nvidia did that because they do a huge amount of research at the company, and they work with every other company doing AI research, and they were like: "Oh, yes, this stuff is gonna work, and this stuff is gonna require the fastest networking available." And I think that has to do with why no one else saw how valuable the Mellanox technology could be.
- DRDavid Rosenthal
Yep.
- BGBen Gilbert
Another thing that I wanna talk about for Nvidia's business today is this notion of the data center is the computer, and Jensen did a great interview with Ben Thompson last year, where he talks about the idea that they build their systems full-stack. Like, their dream is that you own and operate a DGX SuperPOD. And he says, "We build our systems full-stack, but we go to market in a disaggregated way, integrating into the compute fabric of the industry." So I think that's his sort of way of saying: "Look, customers need to use us in a bunch of different ways, so we need to be flexible on that. But we wanna build each of our components such that if you do assemble them all together, it's this unbelievable experience, and we'll figure out how to provide the right experience to you if you only wanna use them in piecemeal ways, or you wanna use us in the cloud, or the cloud providers wanna use us." Again, it's build the product as a system, build the system full-stack, but go to market in a disaggregated way.
- DRDavid Rosenthal
And I think if I remember right in that interview, Ben picked up on this and was like: "Wait, are you building your own cloud?" And Jensen was like, "Well, maybe. We'll see." And of course, then they launched DGX Cloud in a, "Well, maybe we'll see," sort of way. [chuckles]
- BGBen Gilbert
Yeah, you could imagine there are more Nvidia data centers likely on the way that are, uh, fully owned and operated. Speaking of all of this, we gotta talk some numbers on margin. This last quarter, they had a gross margin of seventy percent, and they forecasted for next quarter to have a gross margin of seventy-two percent. I mean, if you go back pre-CUDA, when they were a commoditized graphics card manufacturer, it was twenty-four percent. So they've gone twenty-four to seventy on gross margin, and with the exception of a few quarters along the way for these strange one-time events, that's basically been a linear climb, quarter over quarter, as they've deepened their moat and as they've deepened their differentiation in the industry. We're definitely at a place right now that I think is temporary, due to the supply shortage of the world's enterprises, and in some cases, even governments. You look at the UK or some of the Middle Eastern countries, like blank check, I just need access to Nvidia hardware. That's gonna go away, but I don't think this very high, you know, sixty-five percent-plus margin is gonna erode too much.
- DRDavid Rosenthal
Yes, I mean, I think two things here. One, I really do believe what we were talking about a minute ago, that Nvidia is not just a hardware company. They're not just a chips company. They are a platform company, and there is a lot of differentiation baked into what they do. If you wanna train GPT or a GPT-class model-
- BGBen Gilbert
There's one option.
- DRDavid Rosenthal
You're doing it on Nvidia. There's one option, and yes, we should talk about-- There's lots of less than GPT-class stuff out there that you can do, and especially inference is more of a wide-open market versus training, that you can do on other platforms, but they're the best, and they're not just the best because of their hardware. They're not just the best because of their data center solutions. They're not just the best because of CUDA. They're the best because of all of those. [chuckles] So the other sorta illustrative thing for me that shows how wide their lead is, we haven't talked about China yet.
- BGBen Gilbert
The land of A800s.
- DRDavid Rosenthal
Yes. So what's going on? Last year, China was twenty-five percent, or sales to mainland China was twenty-five percent of Nvidia's revenue, and a lot of that is they were selling to the hyperscalers, to the cloud providers in China: Baidu, Alibaba, Tencent, others.
- BGBen Gilbert
And by the way, Baidu has potentially the largest model of anyone. Their GPT competitor is over a trillion parameters and may actually be larger than GPT-4.
- DRDavid Rosenthal
Wow, I didn't know that.
- BGBen Gilbert
Yep.
- DRDavid Rosenthal
Ah, that's wild. So then, I believe also in September of 2022, [chuckles] last year, the Biden administration announced pretty sweeping regulations and bans on sales of advanced computing infrastructure.
- BGBen Gilbert
David, they're export controls. Don't say bans. [chuckles]
- DRDavid Rosenthal
[chuckles] I mean, yes, that's a fine line, and this is pretty close to bans, what the administration introduced.... As part of that, Nvidia can no longer sell their top-of-the-line H100s or A100s to anybody in China. So they created a nerfed SKU, essentially, that meets the regulations, the performance regulations, the A800 and H800s.
- BGBen Gilbert
Which I think they basically just cranked down the NVLINK's data transfer speeds. So it's like buying a top-of-the-line A100, but not with as fast of data connections as you need, which basically makes it so you can't train large models.
- DRDavid Rosenthal
Right, or you can't train them as well or as fast as you could with the latest stuff. The incredibly telling thing is that those chips and those machines are still selling like hotcakes [chuckles] in China. They're still the best hardware and platform that you can get in China, even a crippled version, and I think that's true anywhere in the world.
- BGBen Gilbert
And there's been a, even a more recent spike of them because a lot of Chinese companies are reading the tea leaves and saying, "Ooh, export controls might get even more severe, so I should get them while I still can, these A800s."
- DRDavid Rosenthal
Yep. So I mean, I can't think of a better [chuckles] illustration of just how wide their lead is.
Episode duration: 2:54:09
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode nFB-AILkamw