Dwarkesh PodcastJensen Huang on Dwarkesh Patel: Why CoWoS Is Nvidia's Moat
CoWoS and HBM commitments placed years early lock supply before rivals can react; no challenger has posted inferencemax results matching Nvidia tokens per watt.
CHAPTERS
- 0:00 – 4:28
Electrons-to-tokens: why Nvidia sees itself as a full-stack “token factory”
Dwarkesh opens by framing Nvidia as largely ‘software sent to manufacturers’ and asks whether AI commoditization could commoditize Nvidia. Jensen responds with his core mental model: transforming electrons into valuable tokens is an end-to-end engineering challenge that’s far from commoditized. He emphasizes partnering wherever possible, while owning the hardest parts across the compute stack.
- •Nvidia’s core input/output framing: electrons → tokens
- •Commoditization is unlikely because the hardest parts are still poorly understood and evolving
- •Strategy: do ‘as much as needed, as little as possible’ via an ecosystem of partners
- •AI as a multi-layer stack where Nvidia participates across layers
- •“Software tools” may grow because agents will massively increase tool usage
- 4:28 – 8:31
Supply-chain moat: purchase commitments, ecosystem alignment, and manufacturing ‘velocity’
Dwarkesh points to Nvidia’s enormous purchase commitments and suggests Nvidia’s moat could be locking up scarce upstream components. Jensen agrees supply commitments matter, but argues the deeper advantage is the ability to align upstream investment with downstream demand. He describes GTC and his keynotes as mechanisms to coordinate the entire ecosystem around a shared growth narrative.
- •Large upstream commitments (explicit and implicit) help secure supply
- •Upstream invests because Nvidia can reliably absorb and resell at massive scale
- •GTC as a ‘360°’ meeting point: upstream, downstream, startups, and AI natives
- •Keynotes as deliberate education to synchronize expectations and planning
- •Business ‘turns’ and demand visibility determine whether supply chains will build for an architecture
- 8:31 – 13:54
Can upstream capacity keep doubling? Bottlenecks, CoWoS, and pre-emptive ‘prefetching’
Dwarkesh challenges how AI compute can keep growing when Nvidia is already a major share of leading-edge nodes. Jensen argues bottlenecks are temporary if demand signals are clear and the industry ‘swarms’ constraints. He uses CoWoS as an example of a bottleneck that received massive coordinated investment and claims Nvidia increasingly prefetches future constraints years ahead.
- •Instantaneous demand can exceed global supply, creating short-lived choke points
- •CoWoS/HBM shifted from ‘specialty’ to mainstream via rapid scaling
- •Bottlenecks attract capital and attention; the industry reallocates quickly
- •Nvidia tries to ‘prefetch’ bottlenecks years in advance
- •Examples: investing in silicon photonics ecosystem, workflows, testing, capacity
- 13:54 – 16:28
Scaling fabs vs. scaling society: EUV, demand signals, and the real constraint—energy
Dwarkesh presses on EUV and the feasibility of doubling logic output; Jensen insists most hardware constraints can be scaled within a few years given credible demand. He argues the longer-term limiting factor is downstream: energy availability and policy. Alongside capacity expansion, Nvidia relies heavily on efficiency gains through architecture and algorithmic improvements.
- •EUV/fab scaling framed as replicable with sufficient demand signal
- •Influencing pinch points sometimes indirect (convince TSMC → ASML follows)
- •Claim: bottlenecks rarely last more than 2–3 years
- •Compute efficiency improvements (10–50×) matter alongside capacity growth
- •Energy and permitting/policy are harder, slower constraints than chip capacity
- 16:28 – 24:53
TPUs vs GPUs: accelerated computing’s breadth and programmability as the hedge
Dwarkesh notes top models trained on TPUs and asks what that implies for Nvidia. Jensen argues Nvidia is selling ‘accelerated computing’ broadly, not a narrow AI ASIC, and that flexibility plus an operator-friendly platform wins across clouds and industries. He stresses programmability as essential because AI progress depends on rapid algorithmic invention, not just matrix-multiply throughput.
- •Nvidia positions itself as general accelerated computing, not a single-purpose TPU
- •GPU programmability enables fast experimentation with new model architectures
- •AI gains come from algorithmic leaps beyond Moore’s Law improvements
- •Extreme co-design across chip, fabric (NVLink), networking, and libraries
- •Nvidia aims to be operable by anyone, unlike many in-house hyperscaler ASICs
- 24:53 – 29:16
Is CUDA still the moat if hyperscalers write their own kernels?
Dwarkesh argues hyperscalers can afford custom kernels and stacks (e.g., Triton), potentially weakening CUDA lock-in. Jensen responds that CUDA is an ecosystem and Nvidia actively contributes to higher-level frameworks; the bigger advantage is install base, reliability, and ubiquity across clouds and edge. The argument: developers optimize first where the ecosystem and deployed hardware are largest.
- •CUDA’s value is ecosystem depth, not just a programming language
- •Nvidia contributes to frameworks like Triton and supports many stacks
- •Trust and debugging: developers prefer a stable, ‘wrung out’ foundation
- •Install base and portability are decisive for framework/model builders
- •Ubiquity across every cloud/on-prem/robotics makes CUDA strategically sticky
- 29:16 – 36:25
Performance, margins, and why Nvidia claims better TCO than ASIC alternatives
Dwarkesh questions whether CUDA-based differentiation can sustain margins if big customers can port away. Jensen argues Nvidia’s in-house optimization expertise is critical—GPUs are ‘F1 cars’ that need specialized tuning—and claims Nvidia delivers the best performance per total cost of ownership. He challenges competitors to show proof via benchmarks and emphasizes tokens-per-watt as a revenue driver for data centers.
- •Nvidia embeds large engineering teams with AI labs to unlock major speedups
- •GPU optimization requires deep architectural knowledge and tooling
- •Claims: best perf/TCO and tokens-per-watt; skepticism toward rival cost claims
- •Benchmarking pressure: invites TPUs/Trainium to show results publicly
- •Hyperscaler demand is mostly external customers, not just internal workloads
- 36:25 – 43:29
Why some frontier labs still pursue ASICs: Anthropic as an ‘investment-driven’ exception
Dwarkesh points out Anthropic’s TPU/ASIC deals and asks why they diversify if Nvidia’s TCO is best. Jensen argues Anthropic is a special case driven by early supplier investment (Google/AWS) that Nvidia couldn’t match at the time. He also notes that making an ASIC isn’t enough—you must beat Nvidia’s rapid yearly leaps, and the economics of ASIC margins may be less compelling than assumed.
- •Anthropic described as the main driver behind TPU/Trainium growth
- •Early compute access tied to strategic investments from hyperscalers
- •Nvidia claims it wasn’t positioned (or didn’t realize) it needed to invest then
- •Skepticism: many ASIC projects get canceled; beating Nvidia is hard
- •ASIC vendors also take high margins, reducing ‘savings’ vs Nvidia
- 43:29 – 51:13
Why Nvidia doesn’t become a hyperscaler: ecosystem-first strategy and ‘not picking winners’
Dwarkesh asks why Nvidia doesn’t directly rent compute at scale given its cash and GPU scarcity. Jensen reiterates the philosophy of doing only what wouldn’t happen otherwise: Nvidia must build the platform stack, but clouds would exist without Nvidia becoming one. He describes selective ecosystem financing (neo-clouds) as bootstrapping, not a desire to be a permanent financier, and explains why Nvidia avoids picking a single foundation-model winner.
- •‘As much as needed, as little as possible’ guides vertical integration decisions
- •Nvidia focuses on platform creation; cloud operation is not uniquely dependent on Nvidia
- •Neo-cloud support (e.g., CoreWeave) framed as catalyzing ecosystem diversity
- •Nvidia avoids being a long-term financier; prefers partnering with finance specialists
- •Doesn’t pick winners: humility from Nvidia’s own unlikely survival story
- 51:13 – 57:37
GPU allocation and pricing: purchase orders, readiness, FIFO, and ‘no highest bidder’
Dwarkesh probes how Nvidia allocates scarce GPUs and whether it strategically fragments the market. Jensen denies preferential allocation in the way described, emphasizing forecasting, purchase orders, and first-in-first-out with adjustments for datacenter readiness. He also rejects surge pricing as bad practice, arguing dependability strengthens long-term trust with customers and suppliers like TSMC.
- •Allocation depends on forecasting, then actual POs; otherwise nothing to allocate
- •General rule: FIFO, with scheduling based on customer readiness to deploy
- •Explicit refusal to allocate by ‘highest bidder’ or to dynamically raise prices
- •Dependability as a strategic asset and industry foundation
- •Longstanding trust relationship with TSMC (even described as contract-light)
- 57:37 – 1:31:34
Should we sell AI chips to China? Cyber risk, compute realism, and ecosystem geopolitics
Dwarkesh raises the national-security case: compute enables cyber-offensive models and scaling inference could magnify harm. Jensen argues China already has substantial chips, energy, and talent; attempting to starve compute is unrealistic and may backfire by splitting ecosystems. He frames the strategic goal as keeping global developers—especially open source—on the American tech stack and warns against conceding major markets that shape standards and optimization targets.
- •Jensen claims China already has enough compute and manufacturing capacity to progress
- •Energy abundance + parallelism can offset older nodes via scale-out
- •Algorithmic innovation can compensate for hardware gaps; cites DeepSeek-like advances
- •Risk of bifurcated ecosystems: open models optimized for non-American stacks
- •Argues conceding China harms US tech leadership and echoes telecom-policy lessons
- 1:31:34 – 1:36:40
What matters more than nodes: architecture, packaging, networking—and why ‘7nm vs 1.6nm’ isn’t the full story
Dwarkesh challenges Jensen’s ecosystem argument using node-gap logic (EUV constraints) and asks about using older nodes if leading-edge capacity saturates. Jensen argues Blackwell’s large gains over Hopper are mostly architectural and stack-driven, not lithography alone, and says going backward is expensive unless forced by extreme constraints. He reiterates that AI performance is shaped by the full system stack, not transistor size in isolation.
- •Blackwell’s gains attributed to architecture/system/stack more than node scaling
- •Node differences are real but not the dominant driver of end-to-end AI gains
- •‘Going back’ to older nodes is possible but would require extreme circumstances
- •Networking/packaging/energy and software stack co-determine capability
- •China won’t remain ‘stuck’ forever; manufacturing and ecosystem adapt over time
- 1:36:40 – 1:39:36
Why Nvidia doesn’t run radically different architectures in parallel (and when it might)
Dwarkesh asks why Nvidia doesn’t pursue multiple fundamentally different chip architectures simultaneously (wafer-scale, Dojo-like, etc.). Jensen says Nvidia evaluates alternatives but doesn’t see better options; they appear worse in simulation. He notes they may add complementary accelerators when markets segment—pointing to integrating Groq for low-latency inference tiers as token value rises.
- •Alternative architectures considered but judged ‘provably worse’ in Nvidia’s simulations
- •Focus resources on improving the main architecture rather than diversifying blindly
- •Workload/market shape (not just algorithms) could justify additional accelerators
- •Inference segmentation: latency-sensitive ‘premium tokens’ create new design space
- •Example: folding Groq into the CUDA ecosystem to expand the Pareto frontier
- 1:39:36 – 1:43:12
If deep learning hadn’t happened: Nvidia’s enduring thesis on domain-specific acceleration
Dwarkesh closes by asking what Nvidia would do without the deep learning revolution. Jensen says the mission would be the same: accelerated computing to overcome the limits of general-purpose CPUs. He lists scientific and industrial domains where GPUs and CUDA accelerate workloads dramatically and frames GTC’s non-AI content as evidence of a broader computing strategy.
- •Core thesis: general-purpose scaling has run its course; acceleration is the path forward
- •GPU+CUDA offload kernels from CPUs for 100×–200× application speedups
- •Broad domains: physics, fluid dynamics, molecular dynamics, seismic, graphics, data processing
- •GTC includes major non-AI advances (e.g., computational lithography cuLitho)
- •Mission: democratize high-end computation for scientists, students, and industry