Skip to content
No PriorsNo Priors

No Priors Ep. 54 | With Sarah Guo & Elad Gil

Host-only episode discussing NVIDIA, Meta and Google earnings, Gemini and Mistral model launches, the open-vs-closed source debate, domain specific foundation models, if we’ll see real competition in chips, and the state of AI ROI and adoption. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil Show Notes: 0:00 Introduction 0:27 Model news and product launches 5:01 Google enters the competitive space with Gemini 1.5 8:23 Biology and robotics using LLMs 10:22 Agent-centric companies 14:22 NVIDIA earnings 17:29 ROI in AI 20:43 Impact from AI 25:45 Building effective AI tools in house 29:09 What would it take to compete with NVIDIA 33:23 The architectural approach to compute 35:42 the roadblocks to chip production in the US 38:30 The virtuous tech cycles in AI

Elad GilhostSarah Guohost
Mar 7, 202442mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 2:39

    Rapid-fire model and product updates: Gemini, Sora, and Mistral’s shipping velocity

    Sarah and Elad open by surveying recent headline launches across the model ecosystem. They contrast Google’s Gemini 1.5 context window advances with OpenAI’s Sora video model and Mistral’s unusually fast cadence from founding to near frontier performance.

    • Gemini 1.5’s million-token context window as a meaningful product capability
    • Sora’s leap in video generation quality and the broader generative video field (e.g., Pika)
    • Mistral Large and Le Chat positioning; rapid execution from startup to near GPT-4 class
    • Why serving efficiency/latency and multilingual coverage are strategic differentiators
    • Microsoft–Mistral Azure distribution as a major go-to-market milestone
  2. 2:39 – 3:41

    Is RAG “dead”? Larger context windows vs retrieval trade-offs

    They debate claims that retrieval-augmented generation becomes unnecessary once context windows are big enough. Sarah argues bigger context expands the design space rather than eliminating retrieval; Elad expects both paradigms to coexist.

    • Bigger context changes—rather than removes—the retrieval vs reasoning trade-off
    • RAG remains valuable for grounding in specific datasets without retraining
    • Stuffing context vs building sophisticated retrieval pipelines as an engineering choice
    • Efficiency and latency implications of different approaches
    • Practical view: hybrid systems are likely the norm
  3. 3:41 – 5:18

    Inference-time reasoning, feedback loops, and continuous model upgrading

    Elad highlights a less-discussed trajectory: shifting more ‘intelligence’ to inference-time computation (search, planning, tool use) and using it to create feedback loops. They connect this to future agent behavior and to the idea that models may evolve continuously rather than in discrete annual training runs.

    • Historical precedent: strong AI systems often do extra work at inference (e.g., search)
    • Agents and advanced reasoning will increase inference-time compute demands
    • Moving from “train a model file yearly” to continuous upgrading/retraining
    • Feedback loops from deployment data as a central driver of improvement
    • Why this shift matters for product architecture and cost structure
  4. 5:18 – 8:29

    Google’s position in the model race: research strength vs organizational will

    They assess whether Gemini changes the competitive outlook for Google. Sarah believes research capability is proven; the open question is whether Google can stay focused and competitive amid internal constraints, while Elad argues Gemini 1.5 signals ‘sleeping giant’ momentum and renewed urgency.

    • Gemini as evidence Google can still ship competitive frontier models
    • Google’s structural advantages: distribution, proprietary data, and compute
    • Internal execution risk: focus, incentives, and organizational priorities
    • Elad’s view: competitive pressure has restored Google’s “will” to move fast
    • Expectation of increased launch velocity and external accessibility of products
  5. 8:29 – 9:50

    Where general LLMs still struggle: biology as a distinct data and reasoning regime

    Sarah and Elad discuss biology as a domain where today’s LLMs are not yet sufficient and where specialized models and long contexts may matter. They point to protein/cell design, drug discovery, and target identification as areas seeing promising transformer and diffusion progress.

    • Examples of current limitations (e.g., designing functional DNA/CRISPR sequences)
    • Protein and cell design require different datasets and evaluation constraints
    • Longer context windows can be important for biological sequence coverage
    • Transformers/diffusion models improving predictions for drug discovery and targets
    • Why biology likely drives differentiated model architectures and tooling
  6. 9:50 – 10:20

    Robotics and the coming proliferation of domain-specific models

    They broaden from biology to robotics and other sciences, framing 2024–2025 as a period of model proliferation beyond general chat. Data constraints remain a gating factor for robotics, but synthetic and structured data generation may unlock progress.

    • Robotics is earlier-stage than language due to data and environment constraints
    • Synthetic data and simulation as potential unlocks
    • Expectation of expanding model types: chemistry, materials, physics, math, robotics
    • Domain models may outperform general LLMs for specialized tasks
    • A shift from a few frontier models to many fit-for-purpose systems
  7. 10:20 – 11:46

    Agent-centric companies: reinforcement learning ideas return in product form

    Elad describes a wave of agent-focused startups applying lessons from games (Go, Diplomacy, poker) to sequential decision-making systems. They argue this approach differs from simply scaling LLMs and may yield new product categories over the next year.

    • Agents require sequential action under changing information
    • Self-play and reinforcement learning as inspiration for agent training
    • Different allocation of work: training vs inference-time computation
    • Expectation: products emerge over 6–12+ months, not instantly
    • A “new wave” distinct from ‘just build a bigger LLM’
  8. 11:46 – 14:10

    Making agents actually useful: constrained domains, validation, and paid feedback data

    Sarah emphasizes that agent success comes from system design and constraints, not generality. They discuss narrowing environments to enable sampling/validation (like coding) and using paid human feedback data to turn an open-ended problem into a costed, tractable one.

    • Early agent attempts failed by being too broad and compounding errors
    • Constrained environments (games, structured web apps, code) support evaluation
    • System-level agent design vs single-agent ‘do everything’ prompting
    • Post-training with paid human feedback to map data needs and costs
    • Framing: ‘How much does it cost to make this task work?’
  9. 14:10 – 17:28

    NVIDIA earnings and the GPU upgrade cycle: demand, constraints, and hyperscaler spend

    They interpret NVIDIA’s results as evidence demand remains far ahead of supply and likely stays constrained. Sarah highlights the economic incentive to upgrade GPU generations; Elad notes most spending comes from hyperscalers, with enterprise adoption still early but accelerating.

    • Supply constraints expected to persist; demand keeps expanding
    • Upgrade cycle (A100→H100→H200→B100) driven by training efficiency gains
    • Startups are not the main spenders; hyperscalers dominate CapEx
    • Azure revenue uplift attributed to AI signals early monetization
    • Enterprise adoption is early, implying further compute demand growth
  10. 17:28 – 19:25

    The ROI debate: ‘AI’s $200B question’ and Meta’s market-cap proof point

    They connect the scale of AI CapEx to the need for tangible returns. Using Meta’s earnings and market-cap jump as an example, they argue AI-driven improvements to targeting, recommendations, and ad tooling can justify large infrastructure investments.

    • ROI framing: massive compute spend demands massive economic output
    • Meta’s CapEx guidance and scale (H100-equivalent compute counts)
    • AI impact on ads: targeting, conversion, engagement, and tooling improvements
    • Market reaction as validation that CapEx can be rational when returns show up
    • Implication: incumbents may capture disproportionate near-term value
  11. 19:25 – 21:38

    Who wins from AI? Incumbent value capture, services automation, and adoption cascades

    They discuss how AI may disproportionately benefit existing large companies in the short run while disrupting labor-heavy services sectors. Sarah proposes an ‘AI impact’ lens for investing; Elad cites early enterprise examples and argues adoption will move slowly—then all at once.

    • Historical pattern: large tech often adds more market cap than startups collectively
    • Services firms and labor-heavy businesses face automation pressure
    • Adoption cascades: once one player proves value, the sector must follow
    • Early signals of business impact (e.g., enterprise vendors seeing AI effects)
    • Strategic question: which management teams can invest through the transition
  12. 21:38 – 24:05

    Klarna’s AI assistant as a concrete case study in operational leverage

    Elad walks through Klarna’s reported results: large volumes of customer chats handled with high satisfaction, faster resolution, and reduced repeats—equating to hundreds of FTEs. They use it to illustrate how quickly a single use case can deliver measurable ROI and trigger broader adoption.

    • 2.3M chats in 4 weeks; ~two-thirds of inquiries handled by the assistant
    • Comparable CSAT to humans; higher accuracy and fewer repeat requests
    • Faster resolution time (minutes vs tens of minutes) and always-on multilingual support
    • Equivalent output of ~700 agents; meaningful workforce and cost implications
    • Societal and labor-market questions emerging from real deployments
  13. 24:05 – 29:16

    Build vs buy in enterprise AI: internal tools, incumbents, and startup openings

    They explore why some companies build AI tools in-house (due to technical capability and direct ROI) while others will rely on vendors. Elad outlines a ‘human capital wave’ view of startup formation—from researchers to infra to application builders—arguing the biggest app wave is still ahead.

    • Three paths: in-house builds, incumbent platforms adding AI, or new startups
    • Feedback loops (ratings, thumbs) make customer support a self-improving product
    • Big-tech SMB-facing platforms (Meta, Shopify, Square, Klarna) as leading indicators
    • Human capital waves: research → infra → applications; app wave still early
    • Large white spaces remain where ‘obvious companies’ still don’t exist
  14. 29:16 – 42:14

    Competing with NVIDIA and the future of compute: moats, manufacturing, and geopolitics

    They close by analyzing what it would take to challenge NVIDIA: silicon performance, CUDA-like software ecosystems, and interconnect scale, plus manufacturing capacity at TSMC and beyond. The discussion expands to US chip-production roadblocks, cultural/human-capital arguments, and the broader virtuous cycle accelerating AI innovation.

    • NVIDIA moat components: chip performance, CUDA ecosystem, and interconnect (e.g., Mellanox)
    • Potential challengers: incumbents (AMD/Intel) and startups (Cerebras, others) plus hyperscaler custom silicon
    • Manufacturing constraints: TSMC capacity, yields, pricing, and time-to-scale
    • US fab roadblocks (permitting/reviews) vs alternative geographies (Japan as hedging location)
    • AI’s virtuous cycle: startups push big tech; big tech funds compute; competition accelerates everything

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.