Skip to content
YC Root AccessYC Root Access

The First Dedicated YC GPU Cluster - With Together AI

YC and Together AI are partnering to bring the first dedicated YC GPU cluster online, giving YC startups easier access to the compute they need to build and scale. In this episode of Founder Firesides, YC's Ankit Gupta and Together AI co-founder and CEO Vipul Ved Prakash dig into why compute has become one of the biggest bottlenecks for modern AI companies. They’ll discuss how Together AI is helping more than 8,000 customers, from early-stage research teams to companies like Cursor, Cognition, and ElevenLabs, train, fine-tune, & run inference on AI models, and why flexible access to GPUs is becoming a competitive advantage for the next generation of founders. Chapters: 00:00 — YC and Together AI Are Partnering on a GPU Cluster 00:26 — What Together AI Does 01:22 — From Research Labs to Cursor: Together's 8,000 Customers 01:58 — The Landscape of AI Native Startups at YC 03:24 — How Building an AI Company Has Changed Since 2018 04:56 — Why the Cost of Compute Keeps Going Up 05:29 — YC as the Biggest Seed Funder of Research Companies 07:07 — Why YC Chose Together AI 08:47 — Flash Attention, Mamba, and the Science of Production AI 10:43 — Compute Planning Advice for Early Stage Companies 12:39 — When Your Compute Bill Is Bigger Than Your Cash Balance 13:31 — How YC Companies Are Using the Cluster Today 14:24 — What's Next for the Partnership Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostVipul Ved Prakashguest
Jul 20, 202615mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

YC and Together AI launch dedicated GPU cluster for startups

  1. YC and Together AI are partnering to provide the first dedicated YC GPU cluster, aimed at improving capacity access, pricing, and support for YC’s AI-native portfolio.
  2. Together AI positions itself as an end-to-end generative AI cloud platform covering model building, post-training of open models, and large-scale serving, with a customer base ranging from research groups to major AI startups.
  3. The conversation highlights how AI startups’ compute constraints have shifted from merely cost to actual availability, with long reservation commitments increasingly incompatible with seed-stage cash realities.
  4. YC frames compute access as a strategic advantage for funding and enabling research-heavy startups that need substantial GPU resources before meaningful commercialization.
  5. Together AI emphasizes “production AI” systems research (e.g., Flash Attention, Mamba, compilers) as a lever to improve workload efficiency and unit economics for both early and scaled companies.

IDEAS WORTH REMEMBERING

5 ideas

Compute availability is now a primary bottleneck, not just compute price.

The speakers note that earlier it was relatively easy to spin up large numbers of GPU instances on public clouds, whereas now startups struggle to secure capacity at all—especially without long-term reservations.

Shorter commitments materially change seed-stage survival and strategy.

YC’s cluster model aims to avoid forcing startups into 24-month capacity bets that can exceed their cash balance, letting them commit for weeks or months and scale usage as needs evolve.

AI-native startups have highly variable workloads that require flexible infrastructure.

Within YC’s portfolio, some companies need a single node quickly while others start with hundreds of GPUs; needs also differ between training-heavy and inference-heavy setups, making a one-size plan inefficient.

Operational support and best practices are part of the product for founders.

Many founders haven’t managed clusters directly; Together’s engineering/research support helps them adopt effective workflows, realistic configurations, and scaling practices without a painful “zero-to-one” ops leap.

“Production AI” optimization is a competitive advantage because unit economics matter early.

Together frames systems work—efficient attention, new architectures like Mamba, and compiler-level improvements—as essential to lowering token costs and accelerating workloads, benefiting both experimentation and scaled serving.

WORDS WORTH SAVING

5 quotes

YC and Together are partnering to bring online the first dedicated YC GPU cluster.

Ankit Gupta

What we've built at Together AI is, uh, really a cloud service that's designed for the whole life cycle of generative AI, which is everything from building models to, uh, post-training open models, to serving them at scale.

Vipul Ved Prakash

I think now what we see is as people's compute needs, needs go up, actually just having access to capacity is a really big problem, let alone great pricing.

Ankit Gupta

Our, our probably very first real machine learning seed investment was literally OpenAI.

Ankit Gupta

The upfront they would need to pay in order to secure the capacity for their next two years of compute was greater than their current cash balance.

Ankit Gupta

Dedicated YC GPU cluster partnershipTogether AI’s model lifecycle platformAI-native startup compute heterogeneityGPU capacity scarcity and pricing pressuresSeed-stage compute planning and commitmentsYC as a funder of research-driven companiesProduction AI research: Flash Attention, Mamba, compilers

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.