Skip to content
The Twenty Minute VCThe Twenty Minute VC

The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao

Lin Qiao is the Co-Founder and CEO of Fireworks AI, the leading specialized intelligence and AI inference platform that last week raised $1.5BN at a whopping $17BN valuation. With just 200 people, the company has hit $1BN in ARR and expects to hit $2BN before the end of the year. Prior to Fireworks, Lin spent several years at Meta including on the founding team of PyTorch. ----------------------------------------------- Timestamps: 0:00 Intro 02:00 - Why Starting a Company at 48 Was an Advantage 03:47 - The AI Layer Everyone Is Overlooking 09:49 - Is AGI Really the End Goal? 11:40 - Open Source vs Frontier Models: Who Wins? 14:30 - Are AI Giants Massively Overvalued? 17:48 - Why Open Models Could Beat Closed AI 19:23 - Should We Trust Chinese AI Models? 22:18 - Do AI Startups Need to Build Their Own Models? 26:49 - Will AI Model Breakthroughs Ever Slow Down? 29:03 - Why One Company Should Never Control Intelligence 32:47 - The Secret Behind Cursor's Explosive Growth 37:33 - Is AI Coding Already Yesterday's Biggest Trend? 41:36 - The AI Infrastructure Race Is Just Getting Started 46:32 - How Cheap Will AI Become? 54:07 - Hypergrowth vs Profit: Why Margins Can Wait 59:30 - Can the West Keep Up With China's Infrastructure Speed? 01:01:04 - Why AI Hardware Depreciates Faster Than Ever 01:04:45 - Why AI Will Create More Jobs, Not Fewer 01:08:35 - The Biggest Mistakes AI Founders Are Making 01:16:55 - Why Great Leaders Stay Close to the Work 01:18:44: Quick-Fire Round ---------------------------------------------------------------------------------------------- Subscribe on Spotify: https://open.spotify.com/show/3j2KMcZTtgTNBKwtZBMHvl?si=85bc9196860e4466 Subscribe on Apple Podcasts: https://podcasts.apple.com/us/podcast/the-twenty-minute-vc-20vc-venture-capital-startup/id958230465 Follow Harry Stebbings on X: https://twitter.com/HarryStebbings Follow Lin Qiao on X: https://twitter.com/lqiao Follow 20VC on Instagram: https://www.instagram.com/20vchq Follow 20VC on TikTok: https://www.tiktok.com/@20vc_tok Visit our Website: https://www.20vc.com Subscribe to our Newsletter: https://www.thetwentyminutevc.com/contact ----------------------------------------------- #20vc #harrystebbings #ceo #ai #linqiao #fireworksai #ceo #ai #founder

Lin QiaoguestHarry Stebbingshost
Jul 20, 20261h 28mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:53

    Fireworks, the core thesis: cheaper tokens (10X) will unlock 100X more usage

    The conversation opens with Lin Qiao’s central worldview: intelligence should not be monopolized, and the AI market is moving from last year’s “coding” wave into a broader “co-work” (knowledge work) wave. He previews Fireworks’ focus on specialized intelligence and his forecast that token costs will drop dramatically, triggering a usage explosion.

    • Fear of a single company “owning intelligence” and shaping society’s taste/values
    • Shift from “year of coding” to “year of co-work” apps and agents
    • 10X token cost reduction in ~3 years as a key macro driver
    • 100X usage growth as prices fall and AI becomes utility-like
  2. 1:53 – 3:47

    Why founding at 48 helped: immigrant journey, systems depth, and learning people at Facebook

    Lin reflects on starting Fireworks later than the stereotypical young founder path. He connects his distributed systems and database background to building complex AI infrastructure, and explains why he intentionally spent years at Facebook to learn organizational and people leadership skills before founding.

    • PhD in distributed systems/databases; experience across the full data-processing stack
    • Immigrated to the US in 2000; early desire to build a tech business
    • At LinkedIn felt technically ready but lacked “people/company-building” skills
    • Joined Facebook intending 1–2 years; stayed 7 to learn operating at scale
  3. 3:47 – 5:35

    The overlooked AI layer: inference + specialized intelligence between chips and model labs

    Harry frames Fireworks as sitting between hardware (NVIDIA) and model providers; Lin argues this layer isn’t commodity because enterprises need tailored intelligence. The big idea: most valuable data is private and will never be in frontier training corpora, so competitive advantage shifts to activating proprietary data with specialized models.

    • AI stack framing: chips → inference/platform → models → applications
    • Public internet data is small vs the world’s private enterprise data
    • Private data can’t be shared; companies need “specialized/private intelligence”
    • Fireworks’ mission: activate private data into deployable intelligence
  4. 5:35 – 11:40

    AGI skepticism and the case for many intelligences (Jensen’s ‘no specialized general company’)

    Lin challenges the assumption that one AGI model will solve everything best, arguing diversity of culture, policy, and taste demands specialization. He positions frontier labs as vital “power lines” but not replacements for differentiated businesses, reinforcing that each company’s uniqueness is encoded in product, systems, and user interaction data.

    • AGI framing implies no need to specialize—Lin rejects that premise
    • Diversity of values/taste makes a single standardizing model undesirable
    • Frontier labs provide foundational capability (“power lines”) for everyone
    • Jensen conversation: every company is inherently specialized; data encodes that
  5. 11:40 – 14:30

    Open-source vs closed models: control, tunability, and crossing the quality threshold

    Lin explains Fireworks’ early bet to build on open models, rooted in PyTorch/open-source culture. He argues both open and closed models have crossed a usability threshold, but open-weight models win on control and ease of steering—especially with small amounts of proprietary data—often outperforming general models for a specific enterprise eval.

    • Founding debate: build proprietary models vs bet on open models
    • Openness gives users weight-level control and customization rights
    • Open/closed both improved; both now solve many real workflows
    • Steering + tuning with small private datasets enables eval-driven “hill climbing”
  6. 14:30 – 19:23

    Are frontier labs overvalued? PMF vs durability when scaling can bankrupt you

    The discussion turns to economics: unlike SaaS, in AI you can have PMF but still fail due to inference costs. Lin describes “scaling to bankruptcy” for startups and incumbents, where CFOs block AI rollouts because unit economics don’t work—pushing organizations toward open models they can control and optimize.

    • In SaaS: PMF ≈ durable business; in AI they diverge
    • Inference/COGS can prevent scaling even with strong demand
    • Incumbents with huge traffic may be unable to afford broad AI feature rollouts
    • Control via open weights enables cost optimization and sustainable scaling
  7. 19:23 – 22:18

    Trust, China, and supply chain resilience: guardrails, alignment, and ‘sovereign’ concerns

    Harry raises security concerns around high-quality Chinese open models. Lin argues every model provider embeds judgments and tastes, so enterprises need their own guardrails and tuning regardless of origin; he also predicts a future with millions of specialized models and says an open ecosystem is resilient even if one country restricts access.

    • Security debate: Chinese open models are strong but raise national concerns
    • Need enterprise-specific guardrails for all models (open or closed)
    • Provider values/taste may misalign with a company’s needs; tuning is required
    • Open ecosystems reduce dependence on any single provider; US can build its own
  8. 22:18 – 26:48

    Do startups need to build their own models? Workflow harnesses, tool orchestration, and legal AI

    Using legal AI as an example, Lin argues defensibility isn’t only about building a base model—it’s about owning the workflow intelligence: tool selection, orchestration harnesses, and co-training systems for high accuracy. He notes coding tools pioneered model tuning early (e.g., Cursor), and many app categories will follow as requirements tighten.

    • Enterprise apps need bespoke orchestration, not just a general model API
    • Legal domain: conservative users, low error tolerance, many case “flavors”
    • Workflow harnesses (tool calling/routing) can be a moat and may need co-training
    • Coding tools led the way; tuning is becoming standard across app categories
  9. 26:48 – 30:23

    Will breakthroughs slow? Step-function base model leaps vs accelerating specialization

    Lin separates progress into two tracks: occasional step-function jumps in general capability, and rapid proliferation of specialized branches once the base trunk improves. He expects specialization to accelerate faster than base-model IQ gains, reshaping how companies build and maintain competitive intelligence.

    • Base-model capability improves in step functions (major vs minor releases)
    • Examples of step changes: improved reasoning/thinking processes
    • Specialization grows like branches from a stronger trunk
    • Net: specialized model innovation outpaces general model progress over time
  10. 30:23 – 32:47

    Routing and self-evolving systems: multi-model stacks, automation, and where value accrues

    The pair explores the “many models” world where tasks route to different models by complexity and cost. Lin sees routing as part of the frontier—especially if it becomes automatic and coupled with tuning—creating self-evolving production systems, though he notes companies with strong evals may build this themselves over time.

    • Future stacks: expensive models for hard judgments; smaller tuned models for sub-tasks
    • Routing mechanisms can be a competitive frontier alongside model quality
    • Vision: automatic routing + automatic tuning → self-evolving systems
    • OpenRouter-type layers may be transient; long-term value may shift to builders with evals
  11. 32:47 – 37:52

    Cursor case study: distributed RL infrastructure across regions to train efficiently as a startup

    Lin details how Fireworks collaborated deeply with Cursor, including building reinforcement-learning rollout infrastructure. The key innovation is decoupling training and rollout, running across multiple regions with scattered GPUs, and solving the ‘freshness’ problem of syncing weights fast enough to keep rewards numerically sound—without hyperscaler-scale clusters.

    • Early adopters want deep control; later markets want accessibility
    • Cursor partnership as boundary-pushing co-development
    • RL training decoupled into trainer + rollout (environment interaction + rewards)
    • Distributed multi-region approach avoids needing massive InfiniBand clusters; weight-sync latency is the core challenge
  12. 37:52 – 47:19

    Demand explosion and the infrastructure bottleneck: 40T tokens/day, energy/chips limits, multi-vendor reality

    Lin shares Fireworks’ scale (tens of trillions of tokens/day) and predicts massive growth, while noting supply chain constraints at the bottom of the “AI cake” (chips, energy, manufacturing parts). He clarifies that customers often keep multi-vendor strategies for safety, but Fireworks differentiates on customized model tuning and workload-specific deployments.

    • Fireworks processes 40+ trillion tokens/day; many from customized models
    • Next-year token volume could grow 20–100X as adoption hits an S-curve
    • Bottlenecks: energy, chips, manufacturing throughput, and even tiny components
    • Customers hedge with multi-vendor; Fireworks positions as specialized intelligence platform, not generic inference
  13. 47:19 – 50:41

    How cheap will AI get? Token-per-task economics, verbosity, tuning, and a 10X cost drop thesis

    Lin argues token pricing must be evaluated per task, because models vary in verbosity and efficiency. He outlines three levers for cost collapse: fewer tokens needed via precision and tuning, cheaper per-token inference via platform optimization, and easing GPU supply constraints—culminating in his 10X cost reduction / 100X usage prediction.

    • ‘Tokens aren’t equal’: cost should be measured per completed task
    • Verbosity can erase apparent price advantages between models
    • Customization/tuning improves precision and reduces tokens required
    • Infrastructure costs should compress in 2–3 years as supply chain improves; predicts 10X cost drop and 100X usage increase
  14. 50:41 – 58:46

    Hypergrowth vs profit: why Fireworks prioritizes quality and speed now, margins later (but never apps)

    The conversation shifts to business strategy: Lin says current margins reflect hypergrowth and experimentation, where heavy optimization can slow innovation. Fireworks won’t enter the application layer, may consider data centers depending on timing, and insists differentiation comes from quality (including bit-equivalent training-to-inference) rather than commoditized serving.

    • Margins are a ‘constraint’ choice; over-optimizing early can slow growth
    • Fireworks focuses on quality-first customization and inference optimization
    • Technical differentiation: ‘zero KLD’/bitwise equivalence across training→inference to preserve trained quality
    • Clear boundary: no move into applications; data centers possible later depending on scale/timing
  15. 58:46 – 1:15:06

    China’s infrastructure speed, data center specialization, and the new reality of fast hardware depreciation

    Lin explains why data centers aren’t commodity: power, cooling, liquid cooling, operations, and heterogeneous architectures create deep specialization opportunities. He also highlights a new dynamic—hardware and model release cadence accelerates depreciation assumptions, changing build-vs-buy decisions for compute and potentially compressing the advantage of owning fleets too early.

    • Data centers require deep expertise (construction, power, cooling, operations)
    • China’s physical infrastructure build speed is a meaningful strategic advantage
    • Heterogeneous systems (GPU + SRAM-heavy ASICs) may outperform single-chip approaches but complicate ops
    • Depreciation cycles are disrupted: rapid SKU/model cadence can shorten useful economic life of hardware
  16. 1:15:06 – 1:28:42

    Leadership, founder mistakes, and what Fireworks is building organizationally (George Hu, Jensen, marketing, ROI discipline)

    In the closing stretch and quick-fire, Lin discusses scaling the team without losing velocity, and why he hired veteran operator George Hu after building early trust. He highlights Jensen’s ‘stay close to the work’ operating model, admits Fireworks waited too long on marketing (education), and predicts the next wave will emphasize ROI measurement and every company owning its intelligence.

    • Hiring George Hu: waited until the company was ready to scale GTM with leverage
    • Jensen lesson: leadership is judgment; staying close to details preserves velocity
    • Founder regret: delayed marketing/education even though product was strong
    • Under-invested area: ROI monitoring/attribution; shift from ‘token maxing’ to ‘ROI maxing’
    • Prediction: every company will ‘own’ its intelligence as a non-optional capability

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.