The Twenty Minute VCThe Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao
CHAPTERS
- 0:00 – 1:53
Fireworks, the core thesis: cheaper tokens (10X) will unlock 100X more usage
The conversation opens with Lin Qiao’s central worldview: intelligence should not be monopolized, and the AI market is moving from last year’s “coding” wave into a broader “co-work” (knowledge work) wave. He previews Fireworks’ focus on specialized intelligence and his forecast that token costs will drop dramatically, triggering a usage explosion.
- •Fear of a single company “owning intelligence” and shaping society’s taste/values
- •Shift from “year of coding” to “year of co-work” apps and agents
- •10X token cost reduction in ~3 years as a key macro driver
- •100X usage growth as prices fall and AI becomes utility-like
- 1:53 – 3:47
Why founding at 48 helped: immigrant journey, systems depth, and learning people at Facebook
Lin reflects on starting Fireworks later than the stereotypical young founder path. He connects his distributed systems and database background to building complex AI infrastructure, and explains why he intentionally spent years at Facebook to learn organizational and people leadership skills before founding.
- •PhD in distributed systems/databases; experience across the full data-processing stack
- •Immigrated to the US in 2000; early desire to build a tech business
- •At LinkedIn felt technically ready but lacked “people/company-building” skills
- •Joined Facebook intending 1–2 years; stayed 7 to learn operating at scale
- 3:47 – 5:35
The overlooked AI layer: inference + specialized intelligence between chips and model labs
Harry frames Fireworks as sitting between hardware (NVIDIA) and model providers; Lin argues this layer isn’t commodity because enterprises need tailored intelligence. The big idea: most valuable data is private and will never be in frontier training corpora, so competitive advantage shifts to activating proprietary data with specialized models.
- •AI stack framing: chips → inference/platform → models → applications
- •Public internet data is small vs the world’s private enterprise data
- •Private data can’t be shared; companies need “specialized/private intelligence”
- •Fireworks’ mission: activate private data into deployable intelligence
- 5:35 – 11:40
AGI skepticism and the case for many intelligences (Jensen’s ‘no specialized general company’)
Lin challenges the assumption that one AGI model will solve everything best, arguing diversity of culture, policy, and taste demands specialization. He positions frontier labs as vital “power lines” but not replacements for differentiated businesses, reinforcing that each company’s uniqueness is encoded in product, systems, and user interaction data.
- •AGI framing implies no need to specialize—Lin rejects that premise
- •Diversity of values/taste makes a single standardizing model undesirable
- •Frontier labs provide foundational capability (“power lines”) for everyone
- •Jensen conversation: every company is inherently specialized; data encodes that
- 11:40 – 14:30
Open-source vs closed models: control, tunability, and crossing the quality threshold
Lin explains Fireworks’ early bet to build on open models, rooted in PyTorch/open-source culture. He argues both open and closed models have crossed a usability threshold, but open-weight models win on control and ease of steering—especially with small amounts of proprietary data—often outperforming general models for a specific enterprise eval.
- •Founding debate: build proprietary models vs bet on open models
- •Openness gives users weight-level control and customization rights
- •Open/closed both improved; both now solve many real workflows
- •Steering + tuning with small private datasets enables eval-driven “hill climbing”
- 14:30 – 19:23
Are frontier labs overvalued? PMF vs durability when scaling can bankrupt you
The discussion turns to economics: unlike SaaS, in AI you can have PMF but still fail due to inference costs. Lin describes “scaling to bankruptcy” for startups and incumbents, where CFOs block AI rollouts because unit economics don’t work—pushing organizations toward open models they can control and optimize.
- •In SaaS: PMF ≈ durable business; in AI they diverge
- •Inference/COGS can prevent scaling even with strong demand
- •Incumbents with huge traffic may be unable to afford broad AI feature rollouts
- •Control via open weights enables cost optimization and sustainable scaling
- 19:23 – 22:18
Trust, China, and supply chain resilience: guardrails, alignment, and ‘sovereign’ concerns
Harry raises security concerns around high-quality Chinese open models. Lin argues every model provider embeds judgments and tastes, so enterprises need their own guardrails and tuning regardless of origin; he also predicts a future with millions of specialized models and says an open ecosystem is resilient even if one country restricts access.
- •Security debate: Chinese open models are strong but raise national concerns
- •Need enterprise-specific guardrails for all models (open or closed)
- •Provider values/taste may misalign with a company’s needs; tuning is required
- •Open ecosystems reduce dependence on any single provider; US can build its own
- 22:18 – 26:48
Do startups need to build their own models? Workflow harnesses, tool orchestration, and legal AI
Using legal AI as an example, Lin argues defensibility isn’t only about building a base model—it’s about owning the workflow intelligence: tool selection, orchestration harnesses, and co-training systems for high accuracy. He notes coding tools pioneered model tuning early (e.g., Cursor), and many app categories will follow as requirements tighten.
- •Enterprise apps need bespoke orchestration, not just a general model API
- •Legal domain: conservative users, low error tolerance, many case “flavors”
- •Workflow harnesses (tool calling/routing) can be a moat and may need co-training
- •Coding tools led the way; tuning is becoming standard across app categories
- 26:48 – 30:23
Will breakthroughs slow? Step-function base model leaps vs accelerating specialization
Lin separates progress into two tracks: occasional step-function jumps in general capability, and rapid proliferation of specialized branches once the base trunk improves. He expects specialization to accelerate faster than base-model IQ gains, reshaping how companies build and maintain competitive intelligence.
- •Base-model capability improves in step functions (major vs minor releases)
- •Examples of step changes: improved reasoning/thinking processes
- •Specialization grows like branches from a stronger trunk
- •Net: specialized model innovation outpaces general model progress over time
- 30:23 – 32:47
Routing and self-evolving systems: multi-model stacks, automation, and where value accrues
The pair explores the “many models” world where tasks route to different models by complexity and cost. Lin sees routing as part of the frontier—especially if it becomes automatic and coupled with tuning—creating self-evolving production systems, though he notes companies with strong evals may build this themselves over time.
- •Future stacks: expensive models for hard judgments; smaller tuned models for sub-tasks
- •Routing mechanisms can be a competitive frontier alongside model quality
- •Vision: automatic routing + automatic tuning → self-evolving systems
- •OpenRouter-type layers may be transient; long-term value may shift to builders with evals
- 32:47 – 37:52
Cursor case study: distributed RL infrastructure across regions to train efficiently as a startup
Lin details how Fireworks collaborated deeply with Cursor, including building reinforcement-learning rollout infrastructure. The key innovation is decoupling training and rollout, running across multiple regions with scattered GPUs, and solving the ‘freshness’ problem of syncing weights fast enough to keep rewards numerically sound—without hyperscaler-scale clusters.
- •Early adopters want deep control; later markets want accessibility
- •Cursor partnership as boundary-pushing co-development
- •RL training decoupled into trainer + rollout (environment interaction + rewards)
- •Distributed multi-region approach avoids needing massive InfiniBand clusters; weight-sync latency is the core challenge
- 37:52 – 47:19
Demand explosion and the infrastructure bottleneck: 40T tokens/day, energy/chips limits, multi-vendor reality
Lin shares Fireworks’ scale (tens of trillions of tokens/day) and predicts massive growth, while noting supply chain constraints at the bottom of the “AI cake” (chips, energy, manufacturing parts). He clarifies that customers often keep multi-vendor strategies for safety, but Fireworks differentiates on customized model tuning and workload-specific deployments.
- •Fireworks processes 40+ trillion tokens/day; many from customized models
- •Next-year token volume could grow 20–100X as adoption hits an S-curve
- •Bottlenecks: energy, chips, manufacturing throughput, and even tiny components
- •Customers hedge with multi-vendor; Fireworks positions as specialized intelligence platform, not generic inference
- 47:19 – 50:41
How cheap will AI get? Token-per-task economics, verbosity, tuning, and a 10X cost drop thesis
Lin argues token pricing must be evaluated per task, because models vary in verbosity and efficiency. He outlines three levers for cost collapse: fewer tokens needed via precision and tuning, cheaper per-token inference via platform optimization, and easing GPU supply constraints—culminating in his 10X cost reduction / 100X usage prediction.
- •‘Tokens aren’t equal’: cost should be measured per completed task
- •Verbosity can erase apparent price advantages between models
- •Customization/tuning improves precision and reduces tokens required
- •Infrastructure costs should compress in 2–3 years as supply chain improves; predicts 10X cost drop and 100X usage increase
- 50:41 – 58:46
Hypergrowth vs profit: why Fireworks prioritizes quality and speed now, margins later (but never apps)
The conversation shifts to business strategy: Lin says current margins reflect hypergrowth and experimentation, where heavy optimization can slow innovation. Fireworks won’t enter the application layer, may consider data centers depending on timing, and insists differentiation comes from quality (including bit-equivalent training-to-inference) rather than commoditized serving.
- •Margins are a ‘constraint’ choice; over-optimizing early can slow growth
- •Fireworks focuses on quality-first customization and inference optimization
- •Technical differentiation: ‘zero KLD’/bitwise equivalence across training→inference to preserve trained quality
- •Clear boundary: no move into applications; data centers possible later depending on scale/timing
- 58:46 – 1:15:06
China’s infrastructure speed, data center specialization, and the new reality of fast hardware depreciation
Lin explains why data centers aren’t commodity: power, cooling, liquid cooling, operations, and heterogeneous architectures create deep specialization opportunities. He also highlights a new dynamic—hardware and model release cadence accelerates depreciation assumptions, changing build-vs-buy decisions for compute and potentially compressing the advantage of owning fleets too early.
- •Data centers require deep expertise (construction, power, cooling, operations)
- •China’s physical infrastructure build speed is a meaningful strategic advantage
- •Heterogeneous systems (GPU + SRAM-heavy ASICs) may outperform single-chip approaches but complicate ops
- •Depreciation cycles are disrupted: rapid SKU/model cadence can shorten useful economic life of hardware
- 1:15:06 – 1:28:42
Leadership, founder mistakes, and what Fireworks is building organizationally (George Hu, Jensen, marketing, ROI discipline)
In the closing stretch and quick-fire, Lin discusses scaling the team without losing velocity, and why he hired veteran operator George Hu after building early trust. He highlights Jensen’s ‘stay close to the work’ operating model, admits Fireworks waited too long on marketing (education), and predicts the next wave will emphasize ROI measurement and every company owning its intelligence.
- •Hiring George Hu: waited until the company was ready to scale GTM with leverage
- •Jensen lesson: leadership is judgment; staying close to details preserves velocity
- •Founder regret: delayed marketing/education even though product was strong
- •Under-invested area: ROI monitoring/attribution; shift from ‘token maxing’ to ‘ROI maxing’
- •Prediction: every company will ‘own’ its intelligence as a non-optional capability