Skip to content
Y CombinatorY Combinator

The 10 Trillion Parameter AI Model With 300 IQ

Earlier this month, OpenAI raised the largest venture round ever at $6.6 billion. The company’s CFO says AI is now at the point where orders of magnitude matter and the next generation of models will be capital intensive. In this episode of the Lightcone, the hosts consider what a world with ultra intelligent models would look like and what potential unlocks could be made possible. Apply to Y Combinator: https://ycombinator.com/apply Chapters (Powered by https://bit.ly/chapterme-yc) - 0:00 Coming Up 0:54 What models get unlocked with the biggest venture round ever? 5:35 Some discoveries take a long time to actually be felt by regular people 9:53 Distillation may be how most of us benefit 14:26 o1 making previously impossible things possible 21:17 The new Googles 23:47 o1 makes the GPU needs even bigger 25:44 Voice apps are fast growing 27:05 Incumbents aren’t taking these innovations seriously 31:52 Ten trillion parameters 33:15 Outro

Garry TanhostDiana HuhostHarj Taggarhost
Nov 1, 202433mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:51

    o1’s impact on founders: value capture vs better product execution

    The hosts frame the central question: if o1-level reasoning is “magical,” does OpenAI capture most of the value, or does it simply remove prompt-wrangling so startups can focus on UX and classic software execution? They set up the episode’s tension between platform power and builder opportunity.

    • Fear: a supermodel could absorb downstream value creation
    • Optimistic view: higher determinism reduces human-in-the-loop prompt debugging
    • Builders shift effort to UX, reliability, distribution, and sales
    • As models improve, “prompt craft” may become temporary knowledge
  2. 0:51 – 1:36

    OpenAI’s $6.6B round and the scaling-law bet

    Garry connects OpenAI’s record fundraise to a compute-driven roadmap where “orders of magnitude matter.” The discussion moves from capital intensity to what it would mean to train and run much larger models.

    • $6.6B raise framed as compute-first, talent-second spending
    • Scaling laws imply successive models are order-of-magnitude larger
    • Bigger frontier models increase capex and operating intensity
    • Sets up the hypothetical jump to 10T parameters
  3. 1:36 – 3:38

    What a 10-trillion-parameter jump could resemble (GPT-2 → GPT-3 déjà vu)

    Diana contextualizes today’s frontier parameter counts and argues that a two-order-of-magnitude jump could feel like the GPT‑2 to GPT‑3 transition. The group ties model-scale leaps to new waves of company formation and capability unlocks.

    • Frontier models are roughly in the ~500B-parameter range (speculated)
    • 10T parameters is a ~100x leap from today’s frontier scale
    • Historical analogy: GPT‑2 (~1B) to GPT‑3 (~170B) changed the ecosystem
    • A similar leap could trigger another startup boom and capability step-change
  4. 3:38 – 5:31

    From “AGI for knowledge work” to 200–300 IQ ASI: what gets unlocked?

    Garry argues current models already rival “normal intelligence” for many knowledge-worker tasks, especially with chain-of-thought style reasoning. The question becomes what superhuman reasoning could enable, including frontier science and new ways of modeling reality.

    • Claim: models can cover 90–98% of many knowledge-worker workflows
    • Tools like Cursor make software engineers dramatically more productive
    • 10T-scale models are framed as potentially “200–300 IQ” systems
    • Examples of superhuman assistance: Terence Tao using ChatGPT for math
    • Potential unlock: deeper modeling of reality akin to past scientific leaps
  5. 5:31 – 8:57

    Why breakthroughs take time to reach everyday people (Fourier transform analogy)

    Diana explains that society may not “feel” AI evenly yet, just as foundational discoveries historically took decades to translate into consumer impact. Fourier transforms are used as a case study in long-lag diffusion from theory to mass utility.

    • AI capabilities exist, but are unevenly distributed in daily life
    • Fourier transform discovered in the 1800s; mass impact felt ~150 years later
    • Enablers: signal processing, compression, telecom, imaging, information theory
    • Consumer technologies (e.g., color TV) emerged after long incubation
    • Raises the question: when does the AI clock start—research decades or ChatGPT era?
  6. 8:57 – 9:53

    Software distribution and consumer “feel the AI” moments (glasses + voice)

    The hosts contrast slow physical-device adoption with fast software rollout via platforms like Google and Facebook. They argue that always-available voice interfaces and smart glasses could make AI’s presence unmistakable for mainstream users.

    • Software can ship globally faster than hardware-dependent innovations
    • Existing distribution (Meta/Google-scale) accelerates adoption
    • Meta Ray-Bans cited as a consumer-facing wedge
    • Voice-first, always-on assistants could be the “feel the AI” tipping point
  7. 9:53 – 12:11

    Distillation as the path for most users: teacher models and cheap students

    Garry and Diana predict that most benefits of massive frontier models will come through distillation into smaller, cheaper models. They discuss how large “teacher” models improve smaller “student” models and how OpenAI is productizing this inside its API.

    • 10T-scale inference may be reserved for high-value frontier work
    • Most users benefit via distilled, cheaper models
    • Meta’s 405B used as a teacher to improve smaller variants
    • OpenAI enables distillation (e.g., using o1/4o to create cheaper internal models)
    • Distillation also acts as a lock-in and platform strategy
  8. 12:11 – 14:25

    Developer model choice is diversifying: YC batch market-share signals

    The conversation shifts to empirical YC batch data showing developers increasingly use multiple model providers. Claude’s growth and LLaMA adoption are highlighted, alongside early uptake of o1 despite its recent release.

    • Founders often choose smaller/faster models over largest models
    • Developer stack has diversified beyond “ChatGPT wrappers”
    • YC S24 stats: Claude jumps ~5% → ~25% usage; LLaMA 0% → ~8%
    • o1 adoption appears quickly (~15% of batch) despite being very new
    • YC usage patterns are framed as a predictor of future winning products
  9. 14:25 – 15:53

    Hackathon proof points: o1 enabling previously impossible builds

    Diana reports live observations from a YC hackathon with early o1 access, claiming teams built capabilities that didn’t work with prior models. An example shows o1 reasoning over docs/code to generate working apps with minimal prompting.

    • YC hackathon with early o1 access; OpenAI team present
    • Claim: o1 enables demos not previously possible with other models
    • Example: Freestyle building a Replit-Agent-like workflow using o1
    • o1 can ingest docs/code and produce functioning applications (with higher latency)
  10. 15:53 – 18:36

    Does o1 reduce moat-building or expand the market? Accuracy as the unlock

    The hosts revisit the builder dilemma: o1 could centralize value at OpenAI, or it could lower friction so more teams compete on product fundamentals. They emphasize accuracy improvements as the barrier remover, citing legal and workflow automation examples.

    • Pessimistic: OpenAI captures a “light cone” of downstream value
    • Optimistic: determinism shifts competition to UX and execution
    • CaseText example: huge effort required historically to reach high accuracy
    • DryMerge example: swapping GPT‑4o → o1 boosts ~80% → ~99–100% accuracy
    • Higher accuracy unlocks mission-critical deployments previously blocked
  11. 18:36 – 20:50

    Profitability shock in the YC portfolio: automation as margin expansion

    Garry shares a portfolio story where support automation materially improved unit economics and removed the need for fundraising. The group connects this to broader enterprise narratives like Klarna replacing internal systems with AI-built tools.

    • Example: company automates ~60% of support tickets and reaches break-even
    • Automation enables growth without additional capital (compounding effect)
    • AI can rescue overcapitalized companies struggling post-2021 valuations
    • Klarna anecdotes: replacing systems like Workday with internal/LLM-built apps
    • AI adoption reframes founder outcomes: from survival to profitability
  12. 20:50 – 23:47

    The “new Googles”: vertical agents and workflow takeovers (TaxGPT)

    They discuss how AI-native companies can act like the next generation of Google: acquire users via informational queries, then expand into high-ACV workflow products. TaxGPT is used as a wedge-to-enterprise example.

    • Analogy: OpenAI as a platform like Google, enabling many downstream winners
    • Vertical agents as “new Googles” for specific domains
    • TaxGPT starts as RAG-based tax Q&A to earn trust and distribution
    • Expands into enterprise document workflows with $10k–$100k ACV
    • Value comes from workflow replacement and time savings, not just answers
  13. 23:47 – 27:00

    Compute and infrastructure implications: o1 raises inference costs; voice hits cost parity

    The hosts note that o1-style reasoning increases inference-time compute, shifting infrastructure economics. They also highlight OpenAI’s real-time voice API pricing as a sign that AI agents are approaching call-center cost parity, driving rapid adoption of voice apps.

    • o1 increases inference compute/latency, impacting infrastructure demand
    • Hybrid approach: use o1 for hard cases, distill for cheap repetitive work
    • Enterprise can tolerate higher latency and pass through costs; consumers may not
    • Voice API priced around $9/hour suggests call-center displacement pressure
    • YC trend: voice startups (debt collection, logistics coordination) now “work” due to latency/interrupt handling improvements
  14. 27:00 – 33:44

    Incumbent complacency, founder advantage, and a final return to 10T/ASI futures

    They argue many incumbents underestimate AI because past tech shifts (e.g., cloud) took a decade, while AI improves in months. The episode closes by returning to 10T-parameter ASI possibilities—accelerated scientific discovery via near-infinite intelligence applied to massive knowledge.

    • Many established organizations have minimal AI initiatives and underestimate pace
    • AI improvement is framed as faster than processors or cloud adoption curves
    • Founder tooling shifts: Cursor adoption outpaces Copilot within YC, showing startup agility
    • Competitive cycles: OpenAI breakthroughs vs fast replication by Claude/LLaMA/Gemini
    • Bull case for 10T/ASI: models synthesize all human knowledge to unlock major discoveries (fusion, superconductors, etc.)

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.