Skip to content
All-In PodcastAll-In Podcast

Jensen Huang: Why cheap chips still produce expensive tokens

Via Groq, Nvidia routes each inference step to the right chip. Vera Rubin targets agentic racks; Huang's key metric is token cost, not the datacenter price tag.

Jason CalacanishostJensen HuangguestDavid FriedberghostChamath Palihapitiyahost
Mar 19, 20261h 6mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 3:32

    Nvidia’s “AI factory OS” and why Groq fits: disaggregated inference + heterogeneous compute

    Huang frames Nvidia’s strategy as building an “AI factory” rather than selling GPUs, anchored by Dynamo—the operating system for AI factories. He explains disaggregated inference (splitting inference stages across different GPUs/chips) and how adding Groq expands the stack for the right workload on the right processor.

    • Dynamo as the operating system of the AI factory; analogy to industrial-era dynamos
    • Disaggregated inference: separating prefill/decode and other stages across devices
    • Shift from a GPU company to a full AI-factory platform (GPU, CPU, networking, DPUs, switches)
    • Rationale for adding Groq processors into Nvidia data center architectures
  2. 3:32 – 6:36

    Agentic workloads drive the “inference explosion”: storage, memory, and multi-model complexity

    The conversation moves from LLM inference to agentic systems, which stress different parts of the stack—especially storage and memory—while orchestrating many models and tools. Huang argues the new workload diversity is why next-gen systems (e.g., Vera Rubin) are designed to be heterogeneous and rack-scaled.

    • Agents access working + long-term memory and use tools, heavily stressing storage/IO
    • Multi-agent collaboration plus diverse model types (diffusion, autoregressive, small/large)
    • Vera Rubin designed for highly diverse workloads; rack-scale expansion
    • TAM expands as more system components (BlueField, networking, CPUs, Groq) become essential
  3. 6:36 – 8:52

    Tokens, not capex: why an expensive AI factory can still produce the cheapest inference

    Sacks challenges the idea that Nvidia’s higher-priced factories will lose to cheaper ASIC alternatives. Huang argues factory sticker price is the wrong metric; what matters is cost per token, and throughput/efficiency can overwhelm capex differences once land/power/shell and the rest of the stack are accounted for.

    • Don’t confuse data center price with token cost; optimize for lowest $/token
    • Large portions of “factory cost” are non-GPU basics (land, power, cooling, servers)
    • Throughput gains (order-of-magnitude) dominate relatively small capex deltas
    • “Even free chips aren’t cheap enough” if they can’t keep up with the pace of progress
  4. 8:52 – 10:46

    How Jensen decides what to build: choosing ‘insanely hard’ problems that fit Nvidia’s superpowers

    Calacanis asks how Huang makes strategic decisions at extreme scale. Huang describes a CEO filtering mechanism: avoid easy markets with many competitors, seek problems that are uniquely hard, unprecedented, and aligned with Nvidia’s capabilities—accepting that pain and “suffering” are part of the process.

    • CEO role: define vision/strategy, informed by deep technical teams
    • Criteria: ‘insanely hard,’ never done before, and matches company superpowers
    • Easy problems attract too many competitors; hard problems create durable advantage
    • Embracing difficulty as inherent to breakthrough invention
  5. 10:46 – 12:10

    Physical AI as a $50T opportunity + digital biology’s ‘ChatGPT moment’

    Huang lays out long-tail bets and why they’re now inflecting: physical AI as tech’s first real shot at digitizing massive real-world industries and digital biology nearing a step-change in representation and prediction. He places both on multi-year curves but argues the payoff is increasingly visible now.

    • Physical AI targets a largely non-digitized ~$50T real-world industry footprint
    • Nvidia’s 10-year investment now inflecting; physical AI approaching ~$10B/year scale
    • Digital biology nearing a ‘ChatGPT moment’ for genes/proteins/cells representation
    • Healthcare and agriculture highlighted as near-term beneficiaries
  6. 12:10 – 15:59

    From generative → reasoning → agents: why OpenClaw makes agents mainstream (and looks like an OS)

    Chamath brings the discussion to desktops and hobbyists; Huang describes three recent AI inflection points and argues agents are the next major shift. He explains why OpenClaw matters culturally and technically, and claims agent frameworks increasingly resemble a modern operating system for “personal AI computers.”

    • Three inflections: ChatGPT (generative UI), reasoning wave, then agentic systems
    • Claude Code as early useful agentic system; OpenClaw popularizes agents broadly
    • OpenClaw-like systems include memory, scheduling, IO, APIs/skills—OS-like primitives
    • Implication: a ‘personal AI computer’ blueprint that can run everywhere (including desktop)
  7. 15:59 – 18:23

    Governing agents + AI policy: pushing back on doomer narratives and regulatory mismatch

    Huang argues policymakers need an accurate mental model: AI is software, not conscious life, and the industry should avoid catastrophic rhetoric that fuels fear. He emphasizes governance/security constraints for agentic systems and warns that overreactive regulation could slow domestic adoption versus competitors abroad.

    • Agent risks require governance: secure access to tools, sensitive data, external comms
    • AI is ‘computer software,’ not alien/biological consciousness; ‘we understand a lot’
    • Doomerism can distort policy; don’t get regulation ahead of fast-moving tech
    • National security concern: other countries adopt AI faster if the US hesitates
  8. 18:23 – 21:32

    Anthropic comms as a case study: warning vs scaring, and why leaders must be moderate

    Asked what he’d advise Anthropic during a defense-related PR flare-up, Huang praises their technology and safety culture but critiques extreme messaging. He argues tech leaders now carry social responsibility, and catastrophic predictions without evidence can harm trust and slow beneficial adoption.

    • Strong endorsement of Anthropic’s capability, security, and safety orientation
    • Distinction: warning is good; scaring is counterproductive
    • Need humility about forecasting; avoid extreme claims without evidence
    • AI’s public sentiment is fragile; communications shape adoption and policy outcomes
  9. 21:32 – 23:32

    The million‑X compute thesis: agents do ‘work,’ and compute needs jump 100× per paradigm shift

    Huang explains why demand keeps compounding: reasoning and agents each multiply compute needs dramatically, while users are willing to pay more when AI completes real work. He reframes AI’s revenue and infrastructure growth as broad-based beyond OpenAI/Anthropic, with open models as a major category.

    • Generative→reasoning ≈100× compute; reasoning→agentic ≈another 100× (≈10,000× in 2 years)
    • People pay more for ‘work done’ than for information/chat answers
    • AI ecosystem is broader than frontier labs; open models are a major category
    • Claim: the industry is still early; ‘we are absolutely at a million X’ trajectory
  10. 23:32 – 27:46

    Token allocation as a productivity strategy: engineers with ‘hundreds of agents’ and superhuman output

    The group discusses internal AI usage economics—how many tokens employees should consume and why. Huang argues high-value engineers should spend large token budgets (like CAD tools for chip design), changing what feels ‘too hard’ and pushing work toward creativity, specs, and evaluation.

    • Tokens as tooling: under-spending on AI is like refusing CAD tools in chip design
    • Heuristic: a $500k engineer should consume substantial token budget (order ~$250k)
    • Agents remove constraints: ‘too hard,’ ‘too long,’ ‘need more people’ thinking fades
    • Future workflow: write specs/architectures, define evaluation, manage agent teams
  11. 27:46 – 30:50

    Auto‑research and the future of enterprise software: agents amplify tools instead of killing them

    Friedberg describes rapid auto-research breakthroughs and rebuilding software stacks with agents; Huang responds that agentic systems expand tool usage dramatically. Rather than destroying enterprise software, agents become ‘100× more butts’ operating existing systems (SQL, design tools, creative apps), with outputs returned in controllable formats.

    • Auto-research enables thesis-level results in minutes with local/desktop workflows
    • Timing + tool-use breakthroughs make OpenClaw-like systems powerful now
    • Counterpoint to ‘enterprise software dies’: agents massively increase tool utilization
    • Results must flow back into trusted tools (e.g., Synopsys/Cadence) for control/ground truth
  12. 30:50 – 33:40

    Open source, open weights, and decentralizing training: why the future is ‘A and B’

    Calacanis asks about open models and decentralized training efforts; Huang argues both proprietary frontier models and open models are essential. He suggests startups can route between best-in-class closed models and cost-reduced specialized open models, enabling rapid capability on day one plus later optimization.

    • Open and proprietary models are complementary, not substitutes
    • Models are ‘technology,’ not merely a service; consumers prefer using top hosted models
    • Industries need controllable specialization—often enabled by open models
    • Routing strategy: start with best model, then fine-tune/specialize and cost-reduce over time
  13. 33:40 – 36:48

    Global diffusion, China access, and supply-chain resilience: what ‘winning’ looks like

    Sacks presses on US export policy and global diffusion; Huang argues US leadership requires broad global adoption of the American tech stack. He discusses losing China share under restrictions, restarting licensing/supply, and contrasts desired outcomes with past strategic dependencies (telecom, rare earths).

    • Goal: American tech stack (chips→systems→platforms) broadly used worldwide
    • Export restrictions previously drove Nvidia from major markets; licensing path to re-enter
    • National security risk when critical industries become foreign-dependent (telecom, rare earths)
    • Preferred outcome: many countries build their own AI on US infrastructure rather than a single global model winner
  14. 36:48 – 39:45

    Geopolitical shocks and manufacturing strategy: Iran/Middle East concerns, Taiwan risk, and reindustrialization

    Chamath asks about conflicts and supply risks (Taiwan, helium); Huang addresses employee/family safety and Nvidia’s commitment to regions like Israel. He outlines three priorities: accelerate US reindustrialization with Taiwan partners, diversify manufacturing globally, and practice restraint to avoid unnecessary escalation.

    • Human impact: employee family anxiety in Iran/Middle East; ongoing support
    • Taiwan: urgency of US reindustrialization (fabs, computer manufacturing, AI factories)
    • Diversify supply chain across geographies (Korea, Japan, Europe) for resilience
    • Helium could be a constraint but supply chains often carry buffers
  15. 39:45 – 1:06:05

    Autonomy everywhere, plus space, healthcare, robotics—and advice to young people in the AI era

    The conversation ranges across self-driving strategy (enable everyone, not build cars), competing with customers building ASICs, and speculative frontiers like space data centers. It closes with healthcare and robotics timelines, then Huang’s guidance: build deep math/science and become truly expert at using AI—jobs change, but purpose expands.

    • Self-driving: Nvidia supplies training + simulation + in-car compute; ‘reasoning AV’ approach
    • Competition with hyperscaler ASICs: Nvidia differentiates via full-stack systems + portability (cloud→on-prem→edge→space)
    • Space data centers: energy abundance but cooling constraints; near-term focus is on-orbit compute for imaging
    • Healthcare: AI biology, clinical agents, and agentic medical instruments; Robotics: 3–5 year path from proof to products
    • Career advice: deep math/science + language skills; become an expert AI user; radiology example—tasks automate, purpose expands

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.