Skip to content
Y CombinatorY Combinator

Open Models Change The Economics of AI

Ollama (YC W21) is used by 9 million developers and 85% of the Fortune 500, giving co-founder and CEO Jeffrey Morgan a unique view into which AI models people are actually using and how that’s changing. Right now, the biggest shift he sees is toward open models, driven by coding agents, falling costs, and capabilities that are rapidly catching up to the frontier labs. On Ollama Cloud, that shift has driven a 150x increase in token usage since the start of the year. In this episode of the Lightcone, Jeff joins us to talk about the future of open models and the story behind Ollama, from two years of searching for the right idea to building one of the most widely used AI developer tools in the world. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs 00:00 — Intro 01:16 — Who's Actually Using Open Models? 02:13 — Is It All About Cost? 03:31 — The Token Usage Explosion 05:31 — Fine Tuning: Hype Cycle or Here to Stay? 07:27 — Where Open Models Beat Claude 08:26 — Launching Models at Scale 11:31 — Olama as the OS Layer 13:58 — Hidden Layers Between Model and App 17:28 — Open vs. Closed: The Steady State 20:41 — Local vs. Cloud Models 23:12 — Why Chinese Models Dominate Cloud 24:00 — NVIDIA's Open Source Play 26:03 — The Return to Local 27:15 — Getting GPUs Is Hard 29:16 — The Flash Model Revolution 31:35 — God Model vs. Orchestration 33:37 — The Geopolitics Question 36:14 — From Docker to Ollama 37:12 — Applied to YC With the Wrong Idea 40:36 — Lost in the Wilderness for Two Years 42:15 — The Pivot That Changed Everything 44:49 — 100K GitHub Stars, No Revenue 47:40 — How Do You Monetize Open Source? 49:43 — Why Do YC as a Second-Time Founder? 51:49 — What Seeing "Good" Actually Does for You 53:38 — Old DevOps Rules That Don't Apply Anymore

Jeffrey MorganguestGarry TanhostJared FriedmanhostHarj Taggarhost
Sep 4, 202657mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Open AI models drive cheaper tokens, faster tools, enterprise adoption

  1. Ollama is seeing a pronounced enterprise shift toward open models, driven first by coding agents and then by long-running “co-worker” agents that dramatically increase per-user token consumption.
  2. Cost is the immediate adoption catalyst, but enterprises’ longer-term goal is control: the ability to customize, govern, and run models in secure environments aligned with internal requirements.
  3. Open-model usage is bifurcating by environment: cloud usage on Ollama skews heavily toward Chinese-origin models, while local usage is a more even mix of US/European/Chinese models.
  4. The operational challenge is no longer just model quality; it’s day-zero launch readiness across inference engines, harnesses/tool-calling, capacity, and hardware optimization—an “OS-layer” integration problem Ollama aims to standardize.
  5. The likely steady state is hybrid: most tokens inside companies flow through open models (80–90%), while a smaller share of spend remains with frontier closed models reserved for the hardest tasks and orchestration.

IDEAS WORTH REMEMBERING

5 ideas

Open models are winning enterprise volume primarily on cost, then control.

Jeffrey Morgan argues cost is the biggest near-term pain open models solve, but the “North Star” for businesses is control—customization, governance, and deployment choices that closed labs can’t always provide.

Agents—not chat—are the main driver of the token demand surge.

Ollama’s per-developer token usage spiked first with coding agents and later with “co-worker” agent workflows (e.g., OpenClaw/Hermes) that run longer, use tools, and exploit larger context windows.

The open-model release cadence makes fine-tuning harder, but tooling is catching up.

Rapid iteration (e.g., multiple DeepSeek Flash releases in a summer) can “stomp” bespoke fine-tunes, yet improving post-training tooling is making it feasible for teams that truly need specialization.

Open models can outperform closed models in security testing because they refuse less.

For pen-testing and security research, closed models may block requests; open and specialized “security researcher” variants can be more usable, creating a clear niche where openness is a functional advantage.

The biggest platform value is integration across many hidden layers.

Running models well requires coordination across inference engines, harnesses/SDKs, tool-calling mechanics, benchmarks, capacity planning, and hardware drivers—an “OS-like” combinatorial integration problem Ollama targets with a standardized runtime.

WORDS WORTH SAVING

5 quotes

Cost is by far the largest pain point that open models can jump in and solve. But, you know, every business has a vision of getting better control over AI and customizing it for their business, and that's really their North Star.

Jeffrey Morgan

This is on an individual user ba- user basis, how many tokens are they using a week?

Jeffrey Morgan

As a whole through Ollama's cloud, we saw 150x since the start of the year.

Jeffrey Morgan

The super majority of tokens, and this is our take, it will be open models within a business. Call it 80, 90%. That doesn't mean 80, 90% of the, the budget will go to open models.

Jeffrey Morgan

And there's this concept that if you're a layer on top of something else, that you're in kind of a vulnerable position as a startup Which is absolutely not true in the AI world.

Jeffrey Morgan

Enterprise adoption of open modelsToken-usage explosion from agentsFine-tuning cadence vs rapid model releasesSecurity testing and safety tradeoffsDay-zero model launch playbookOllama as an OS/runtime layerHybrid local–cloud execution and routingChinese vs US/European model share in cloudGPU scarcity and inference-provider partnershipsFlash/ultra-cheap workhorse modelsOrchestration vs “God model” debateGeopolitics, provenance, and verificationOpen-source monetization and cloud offeringFounder pivots and product-market fit lessons

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.