Skip to content
YC Root AccessYC Root Access

The First Dedicated YC GPU Cluster - With Together AI

YC and Together AI are partnering to bring the first dedicated YC GPU cluster online, giving YC startups easier access to the compute they need to build and scale. In this episode of Founder Firesides, YC's Ankit Gupta and Together AI co-founder and CEO Vipul Ved Prakash dig into why compute has become one of the biggest bottlenecks for modern AI companies. They’ll discuss how Together AI is helping more than 8,000 customers, from early-stage research teams to companies like Cursor, Cognition, and ElevenLabs, train, fine-tune, & run inference on AI models, and why flexible access to GPUs is becoming a competitive advantage for the next generation of founders. Chapters: 00:00 — YC and Together AI Are Partnering on a GPU Cluster 00:26 — What Together AI Does 01:22 — From Research Labs to Cursor: Together's 8,000 Customers 01:58 — The Landscape of AI Native Startups at YC 03:24 — How Building an AI Company Has Changed Since 2018 04:56 — Why the Cost of Compute Keeps Going Up 05:29 — YC as the Biggest Seed Funder of Research Companies 07:07 — Why YC Chose Together AI 08:47 — Flash Attention, Mamba, and the Science of Production AI 10:43 — Compute Planning Advice for Early Stage Companies 12:39 — When Your Compute Bill Is Bigger Than Your Cash Balance 13:31 — How YC Companies Are Using the Cluster Today 14:24 — What's Next for the Partnership Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostVipul Ved Prakashguest
Jul 20, 202615mWatch on YouTube ↗

CHAPTERS

  1. 0:05 – 0:26

    YC + Together AI launch a dedicated GPU cluster for YC startups

    Ankit introduces Vipul and announces a partnership to bring the first dedicated YC GPU cluster online. The goal is to give AI-native YC companies easier, faster access to compute for building and scaling.

    • YC and Together AI are partnering on a dedicated GPU cluster
    • Cluster is designed to serve YC’s AI-native startup portfolio
    • Focus on easier access to capacity needed to scale generative AI workloads
  2. 0:26 – 1:09

    Together AI’s platform: end-to-end infrastructure for generative AI lifecycle

    Vipul explains Together AI’s origin and mission to keep genAI innovation participatory rather than concentrated. Together is positioned as a cloud service spanning model building, post-training of open models, and large-scale serving.

    • Founded ~4 years ago by academic/research-oriented founders
    • Concern about capital intensity driving concentration in genAI
    • Product covers training/building, post-training, and scalable inference/serving
  3. 1:09 – 1:58

    Who uses Together: from research groups to scaled AI products (8,000+ customers)

    The conversation shifts to Together’s customer base and usage patterns. Vipul notes the breadth from early experimentation to major AI-native companies running training and serving workloads on Together.

    • Together focuses on AI-native startups and teams
    • Customer base includes small research groups and large companies
    • Examples mentioned: Cursor, Cognition, ElevenLabs
    • Supports both model training and serving at scale
  4. 1:58 – 3:24

    Mapping YC’s AI-native landscape and its diverse compute needs

    Ankit outlines the spectrum of AI-native startups at YC—from training foundation models to building on open-source models to data and application-layer companies. This diversity drives highly variable compute requirements, making a flexible compute partner valuable.

    • Startups range from training-from-scratch to app-layer builders
    • Open-source models and hosted inference enable new product types
    • Compute needs are heterogeneous across the YC portfolio
    • Partnership value: multiple ways to engage depending on workload
  5. 3:24 – 4:04

    How building AI companies changed since 2018: more industries now “software-first”

    Ankit contrasts today’s environment with his 2018 experience in AI-native pharma. More sectors now behave like software companies, expanding what founders can tackle—but also introducing new bottlenecks that weren’t previously common.

    • Software-first thinking now common in insurance, healthcare, pharma, etc.
    • Broader set of viable startup ideas and categories at YC
    • New bottlenecks emerge as AI capabilities expand into more domains
  6. 4:04 – 4:56

    Why GPU access is harder now: capacity constraints and longer commitments

    They discuss the shift from relatively easy GPU provisioning in prior years to today’s capacity shortages and pricing pressure. Ankit emphasizes the importance of securing capacity with good terms without forcing startups into multi-year reservations.

    • Previously: easier to spin up large GPU fleets via AWS/spot
    • Now: capacity access itself is a major problem (not just price)
    • YC/Together aim: secure capacity, good pricing, strong support
    • Short commitments (weeks/months) vs 2-year reservations
  7. 4:56 – 5:30

    Compute costs rise as model value rises; YC’s role funding research-heavy startups

    Vipul frames a dynamic where better models increase the value of tokens, affecting the economics of FLOPs and driving compute costs up. Ankit adds that YC increasingly funds research-style companies, where compute is a core constraint to progress.

    • Improving models changes the economic value of compute and tokens
    • Compute costs may keep expanding with capability demands
    • YC is a major seed funder of research-style AI companies
    • Compute access is crucial for technical breakthroughs and scaling
  8. 5:30 – 7:33

    Examples of YC research modalities: voice, edge models, robotics, biology

    Ankit gives concrete examples of the kinds of frontier or research-driven companies YC backs. Many won’t commercialize immediately, so their ability to iterate hinges on reliable, affordable compute access.

    • Historical note: OpenAI as an early ML seed investment via YC Research
    • Voice AI examples: Deepgram; downstream/distribution examples like VAPI
    • New model work across language, edge devices, and novel architectures
    • Research spans biology/healthcare and also hard tech/robotics
    • Compute is a bottleneck for non-immediate-commercialization “research bets”
  9. 7:33 – 9:01

    Why YC chose Together: speed, scaling range, and hands-on best practices

    Ankit explains selection criteria: rapid deployment of the cluster, support for everything from single-node users to 256+ GPU deployments, and an engineering/research team that can teach operational best practices. The intent is to help founders avoid a painful ‘zero-to-one’ in cluster operations.

    • YC prioritized getting a cluster online quickly
    • Must support tiny and massive workloads (single node to thousands)
    • Together provides support and guidance on operational best practices
    • Many founders haven’t managed clusters directly before
    • Together’s research track record contributed to trust
  10. 9:01 – 11:19

    Production AI as a science: Flash Attention, Mamba, compilers, and unit economics

    Vipul describes Together’s research focus on “production AI”—the systems and efficiency work that makes training and inference economical at scale. He highlights attention innovations, long-context needs, architecture work like Mamba, and compiler/system improvements to drive better unit economics.

    • Production AI is an understudied discipline focused on efficiency
    • Research targets attention mechanisms as contexts grow longer
    • Together contributed to/worked on Mamba-style architectures
    • Compiler/systems work helps accelerate workloads on accelerators
    • Benefit: improved unit economics for both early and scaled companies
  11. 11:19 – 12:52

    Compute planning for seed-stage teams: scaling without overcommitting capital

    They discuss how today’s tooling and open-weight models enable efficiency, but compute planning remains complex. Vipul and Ankit stress that compute can become the largest expense, so pricing/packaging and planning mechanisms are essential to help startups scale responsibly.

    • More tools, data, and open models make early execution more efficient
    • Compute planning is complex and changes with company stage
    • Compute may be the biggest development/product expense for AI startups
    • Goal: products and packaging that let companies scale to the next stage
  12. 12:52 – 13:45

    When reservations exceed runway: solving the ‘compute bill bigger than cash’ problem

    Ankit explains a key motivation: some startups would need to prepay or reserve two years of capacity at a cost exceeding their cash balance. The YC cluster model uses YC’s aggregate scale so individual startups can commit in smaller increments and preserve flexibility.

    • Two-year capacity commitments can exceed a startup’s cash on hand
    • Without a solution, startups may need to raise just to secure compute
    • YC scale enables shorter commitments and better access terms
    • Startups can scale continuously without locking up all cash
  13. 13:45 – 14:37

    How YC companies use the cluster today: support workflows + advanced booking

    Ankit shares early feedback: companies value Together’s support in building performant workflows and understanding practical cluster setups. A notable feature is the ability to book capacity in advance, letting teams plan small usage now and large training runs later.

    • Different companies need training vs inference clusters
    • Support team helps design workflows and practical compute setups
    • Advance booking enables multi-month compute planning
    • Companies can ramp usage up/down based on upcoming training runs
    • Cluster runs at full utilization overall while remaining flexible per company
  14. 14:37 – 15:47

    What’s next: expanding the cluster as a shared leverage tool for founders

    They close by framing compute as another ‘founders union’ lever YC can provide—akin to capital access, advisors, and other shared resources. Both express enthusiasm about scaling the partnership and making the cluster larger over time.

    • Plans/intent to expand capacity and deepen the partnership
    • YC aims to give founders leverage similar to large companies
    • Compute joins other shared resources (capital support, advisors, etc.)
    • Mutual optimism about compute becoming essential for AI startups

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.