No PriorsBaseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud
CHAPTERS
- 0:05 – 1:55
Baseten’s 30× growth: inference demand explodes as AI becomes “everywhere”
Sarah and Elad open by framing inference as a uniquely strategic market, then Tuhin explains Baseten’s rapid growth and why demand has accelerated over the last 24 months. He attributes the surge to better open-source baselines and mainstream post-training/RL techniques that let companies increasingly “own” their inference stack.
- •AI adoption broadens: “you can put AI everywhere”
- •Open-source model quality crosses a capability threshold
- •Post-training/RL becomes common enough to operationalize
- •Long-tail of specialized models increases infrastructure demand
- •Baseten benefits by indexing on the growing application layer
- 1:55 – 5:57
Why the independent app layer wins: moats come from proprietary workflow signals
Tuhin argues the app layer persists because durable value comes from unique user signals and workflow integration, not just access to frontier models. He uses healthcare and support workflows to illustrate how proprietary feedback loops enable differentiated, long-horizon agentic systems.
- •Moats are built from user signals only the app owner can collect
- •Encoding advantage in workflows is safer than encoding it solely in a model
- •Example: Abridge’s deep EMR integration creates defensibility
- •Long-horizon agents improve via reward signals from real user edits/actions
- •Support workflows are multi-step, enabling specialized models over time
- 5:57 – 7:55
Serving frontier AI-native companies today—and translating their needs into enterprise readiness
The discussion shifts to who is actually using inference at scale: mostly AI-native application companies, not enterprises directly (yet). Tuhin explains how serving companies that sell into enterprises effectively “imports” enterprise requirements into Baseten’s product roadmap.
- •By inference volume, ~99% of demand is still from AI-native app companies
- •Enterprise adoption is early: closed-source APIs first, custom models later
- •Baseten builds for the highest-scale, most demanding customers (Stripe analogy)
- •AI-native vendors relay enterprise needs: retention, SLAs, transparency, latency
- •Serving Abridge/OpenEvidence/Decagon etc. prepares Baseten for regulated verticals
- 7:55 – 9:22
Open-source model mix in production: capability first, then cost optimization
Elad asks how model preferences have evolved; Tuhin says customers prioritize capability to unlock value, then optimize for cost and reliability. Baseten sees wide diversity—text, speech, and frontier open models—reflecting a fast-moving model ecosystem.
- •Customers start with best capability; cost sensitivity comes later
- •Production mix spans many model families and modalities
- •Examples mentioned: GPT OSS, Moonshot, DeepSeek, specialized TTS models
- •Baseten’s advantage is operational knowledge of “how to run them well”
- •The key change: open models are now “good enough” for serious workloads
- 9:22 – 13:07
Chinese models, security fears, and geopolitics: how to reason about risk and advantage
The hosts probe concerns about Chinese-origin open models (security, bias, trojans) and broader national strategy. Tuhin downplays practical “network boundary” risks in typical deployments, argues the models are excellent, and emphasizes that the US needs a strong open-source footing to avoid strategic loss.
- •Security concern is often overstated if models are network-bounded in data centers
- •Little evidence of embedded “agenda” beyond early, quickly-noticed issues
- •US open-source competitiveness matters; relying solely on China would be a loss
- •Reframe: evaluate models as if they came from a US lab and build pragmatically
- •Economic angle: cheaper open models accelerate innovation by lowering intelligence cost
- 13:07 – 14:23
Custom inference dominates Baseten: dedicated vs shared endpoints and why “vanilla weights” are rare
Sarah asks what fraction of tokens are custom vs off-the-shelf; Tuhin says it’s overwhelmingly custom. He breaks down Baseten’s business lines and explains why customers nearly always modify, compile, or optimize models for quality and performance.
- •~90–95%+ of tokens are on dedicated/custom inference
- •Baseten offerings: dedicated inference, shared inference, and training
- •Customers typically fine-tune/post-train with their own data
- •Even without quality tuning, performance compilation/optimization is common
- •Running “vanilla open-source weights” in production is increasingly unusual
- 14:23 – 17:31
Post-training acquisition (Parsed): bringing research capability closer to customers
Tuhin explains Baseten’s acquisition of Parsed, a post-training team that was already a customer. The goal was to add research expertise, accelerate customer success earlier in the lifecycle, and build tighter coupling between post-training and inference performance.
- •Baseten started as product/infrastructure-heavy; lacked deep research bench
- •Market shift: post-training demand (expertise and tooling) is surging
- •Parsed wanted to become an inference company; Baseten needed their specialization
- •Post-training and inference are intertwined (e.g., quantization choices)
- •Vision: close the loop—production inference generates data → evals → post-training
- 17:31 – 18:35
When to invest in custom models: don’t post-train before product–market fit
Tuhin gives a practical adoption playbook: prove value with best-in-class models first, then optimize. Custom models make sense once there’s real scale, a clear user signal, and an understanding of what “better/faster/cheaper” means for a specific workflow.
- •Start with frontier models to validate the use case and ROI
- •Rule of thumb: “no post-training pre-PMF”
- •Custom models require user signal and a reward function to optimize
- •Specialization can beat general coding ability for domain tasks (e.g., support)
- •Lifecycle: capability → scale → customization → cost/performance optimization
- 18:35 – 21:39
Supply crunch reality: multi-cloud fabric, high utilization, and unreliable suppliers
The conversation turns to compute scarcity: Tuhin says the shortage is worse than most narratives suggest. Baseten’s strategy relies on a runtime fabric across many clouds, but operationally strong suppliers are limited and many new entrants struggle with true inference-grade SLAs.
- •Compute slack is minimal; Baseten runs clusters at mid-90s utilization
- •Baseten spans ~18 clouds and ~90 clusters globally
- •Fabric abstracts latency, reliability, failover, and portability for customers
- •Rapid onboarding of new providers (hours) helps unlock supply opportunistically
- •Supplier quality is uneven; only a small set of clouds meet “gold tier” ops/SLAs
- 21:39 – 24:10
Longer GPU contracts and capital dynamics: prepay, term risk, and IPO pressure
Sarah asks about buying supply years out; Tuhin notes it’s possible but requires big bets amid fast hardware/software change. He describes longer contract terms, substantial prepay requirements, and how working-capital needs can push companies to seek cheaper capital—potentially via public markets or structured financing.
- •Forward capacity can be purchased, but it’s a bet in a fast-moving market
- •H100 longevity is unusual; prices can still rise even years into lifecycle
- •Large allocations (e.g., B200s) often require 3–5 year terms
- •Typical deals may include 20–30% TCV prepay
- •Capital strategy matters: low cost of capital becomes a competitive lever
- 24:10 – 26:07
What makes an inference-cloud winner: sticky software + strategic compute access
Tuhin distinguishes commodity GPU rentals from sticky inference platforms that embed software, reliability, and performance tooling. He argues winners need both: a differentiated software layer and privileged access to scarce compute, noting strong retention dynamics among Baseten’s largest customers.
- •“GPUs as a service” is commoditized; inference software layers create stickiness
- •Baseten reports high retention and expansion among top customers
- •Software moat: deployment, optimization, SLAs, reliability primitives
- •Compute access becomes strategic as labs and hyperscalers lock supply
- •Winning requires excellence across software, supply, operations, and demand
- 26:07 – 28:19
Multi-chip future vs Nvidia gravity: ecosystem, supply chain, and inference-specific chips
Sarah probes whether the world diversifies beyond Nvidia; Tuhin expects more chip variety, including inference/decode-specialized hardware. Still, he emphasizes Nvidia’s supply chain execution, CUDA ecosystem, and the speed advantages of building on the dominant platform—especially in the next few years.
- •Long-term: diversification and inference-specific chips are likely
- •Near-term: Nvidia remains hard to displace due to CUDA and ecosystem momentum
- •Speed of execution matters most for infrastructure companies right now
- •Ecosystem formation is critical; exclusive supply deals can prevent it
- •Strategic risk: proprietary chip deals may concentrate benefits to a few buyers
- 28:19 – 31:08
Runtime roadmap for new workloads: agents, sandboxes, prefill/decode, and routing optimizations
Tuhin outlines where Baseten invests as workloads change—especially code agents, long-horizon tasks, and new modalities. He highlights runtime-level optimizations and new primitives that improve throughput, latency, and utilization, while tying these to the broader inference↔post-training loop.
- •Support for new workloads: coding agents need secure sandboxes
- •Speculative decoding and other techniques to accelerate inference
- •KV-cache-aware routing; separating prefill vs decode for better performance
- •Scaling diffusion transformers and emerging modalities (e.g., video)
- •Product thesis: integrate evals, training APIs, and batch/async inference to deepen the loop
- 31:08 – 33:49
Scaling edge cases and reliability: kernel-level failures and immature LLM runtimes
At high scale, Baseten encounters unexpected failure modes—some coming from logging, kernels, and infrastructure primitives rather than the models themselves. Tuhin notes that LLM runtime stacks are still immature, revealing new needs in security, performance, and systems engineering.
- •Hyperscaler components can hit surprising limits at real scale
- •Example: kernel panic triggered by excessive logging concurrency
- •Many “edge cases” are systems-level, not model-level
- •LLM runtime primitives (e.g., KV-cache usage) are still evolving
- •Scale exposes next-generation needs in performance, security, and primitives
- 33:49 – 38:21
Hiring, leadership, and pager culture: building an ops-first company
Tuhin describes shifting from a flat org to a strong leadership team that can own whole problems. He also emphasizes that cloud/inference requires relentless operational discipline—incidents, paging, and a culture that rapidly filters for people comfortable with always-on reliability expectations.
- •Leadership lesson: hire trusted leaders who can own entire problem spaces
- •Founder trap: “I must be in everything” often signals wrong team structure
- •Clear hiring rubric: first-principles thinking, kindness, low ego, no hero culture
- •Ops reality: inference can’t go down; paging and incidents are part of daily life
- •Reliability culture self-selects; people who avoid oncall won’t last
- 38:21 – 42:57
Efficiency drives more demand (Jevons) and the “concierge everything” future
Sarah raises Jevons paradox; Tuhin agrees that cheaper inference leads developers to embed more intelligence, especially via longer-running agents. They close by imagining a world of pervasive personalized copilots and warning companies that failing to adopt intelligence-infused workflows could be existential.
- •Lower inference costs increase usage—developers add more intelligence, not less
- •Agents run longer and do more work as price/performance improves
- •Inference remains “the last market”—even with AGI, inference persists
- •Consumer future: personalized concierge-like agents across health, education, life
- •Company imperative: not embracing AI-infused workflows risks extinction