Skip to content
a16za16z

How Open Source Became AI's Backbone | Inferact with a16z

Elena Burger and Matt Bornstein are joined by Simon Mo, co-founder and CEO of Inferact, the open-source inference engine powering many of today's most advanced AI applications. Together, they explore how open-source AI evolved from a research project into critical infrastructure, why inference has become one of the most important layers of the AI stack, and what it takes to bring frontier intelligence to developers around the world. The conversation covers vLLM's origins, the rise of open-weight models, why companies increasingly want control over their AI infrastructure, and how open-source inference enables the next generation of AI applications. They also discuss model licensing, the economics of open-weight AI, Kimi K3, distillation, AI infrastructure, and why Simon believes the gap between open and closed models is rapidly disappearing. Timestamps: 00:00 - Intro 01:00 - What Is vLLM & Why Serving LLMs Is a Fundamentally Different Problem 05:10 - When Open Source Became Critical Infrastructure 08:26 - Where vLLM Sits in the Stack 13:35 - The Open Weights Letter & Why Open AI Development Must Be Protected 16:42 - K3 Economics: Bridging the Gap Between Open & Proprietary 19:57 - Licensing Evolution: From Apache 2 to Commercial Terms 28:51 - Why Open Source Inference Is the Only Way to Scale Agents 32:21 - The Hugging Face Incident & Why Guardrails Break Down 36:57 - Building a Company from an Open Source Project 43:32 - The Distillation Debate: Is It Critical or Incidental? Resources: Follow Simon Mo on X: https://x.com/simon_mo_ Follow Matt Bornstein on X: https://x.com/BornsteinMatt Follow Elena Burger on X: https://x.com/VirtualElena Follow Inferact: https://x.com/inferact Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Matt BornsteinhostSimon MoguestElena Burgerhost
Aug 6, 202646mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

How vLLM and open weights turned inference into AI infrastructure

  1. Serving LLMs is fundamentally different from traditional ML because inference must handle highly variable, non-deterministic workloads that demand sophisticated GPU scheduling, batching, and latency control.
  2. vLLM has become “critical infrastructure” by sitting at the junction of new model releases and new hardware, enabling day-zero support across many architectures while acting as a benchmark for chip vendors.
  3. Open weights are framed as essential for both control (latency SLAs, customization, security/compliance, guardrails) and, increasingly, cost management as token spend rises.
  4. Model licensing is evolving from permissive Apache-style norms toward commercial or threshold-based terms to create sustainable funding mechanisms for expensive training and repeated failed runs.
  5. The Hugging Face “guardrails” incident is used to argue that centralized proprietary moderation inevitably over-blocks legitimate work, pushing trusted users toward open-weight deployments where guardrails are configurable.

IDEAS WORTH REMEMBERING

5 ideas

Inference is a systems problem, not just an ML problem.

LLM serving must juggle variable prompt lengths, non-deterministic outputs, and tight latency expectations, which makes batching, scheduling, and GPU/TPU utilization the core engineering challenges.

vLLM’s defensibility comes from being the “meeting point” of models and hardware.

By supporting ~1,000+ architectures and collaborating with NVIDIA/AMD/Google/AWS/Intel, vLLM becomes both the day-zero launch surface for new open-weight models and a standard benchmark for new chips.

Open weights matter most when you need operational control, not ideology.

Users want to tune speed/latency tradeoffs, enforce data retention and compliance requirements, and meet SLAs (e.g., voice agents) without depending on opaque proprietary endpoints that can throttle or fail.

Performance tuning becomes a product surface with open inference.

Unlike proprietary “regular vs fast” switches, open deployments can expose many speed/cost configurations (including very high token/sec modes), letting providers optimize for different workloads and budgets.

Open-weight licensing is shifting to pay for the real cost center: training iteration.

As labs face huge up-front capex plus multiple failed runs, licenses increasingly include revenue/usage thresholds or commercial terms (e.g., Llama-style constraints, newer model-specific clauses) to sustain future R&D.

WORDS WORTH SAVING

5 quotes

We literally had this company called OpenAI, which- ... you know, it's become a little bit of a joke. It's not as open as it once was, or not nearly as open as it once was.

Matt Bornstein

vLLM is a inference engine. That means its job is to turn available GPUs into a running endpoint for intelligence.

Simon Mo

The world cannot just be controlled by proprietary APIs and where open weight, open development and research of these models are blocked or banned, right?

Simon Mo

Right? Like, like I can't just like go home at night and like train a frontier open- source model with friends- for fun. Like we need millions or billions of dollars of computing resources enabled in order to do it.

Matt Bornstein

If moderation is never solved, which is gonna be very, very hard, then there's always a place where you have a model where you know and trust that you are publishing to and to be able to use from.

Simon Mo

Why LLM serving differs from classic ML workloadsvLLM’s role as an inference engine and benchmark layerOpen weights vs proprietary APIs: cost, control, and customizationDay-zero model support and multi-party release coordinationLicensing shifts and sustainability of open-weight model developmentAgents and why open inference is needed at scaleGuardrails, moderation failures, and the Hugging Face incidentDistillation debate vs environment-driven progressCommercializing open source: Inferact’s “last mile” value

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.