Skip to content
YC Root AccessYC Root Access

Building AI That Optimizes AI

Wafer (S25) is building a fast AI inference cloud that runs open-source models at industry-leading speeds by using agents to optimize GPUs at every level of the stack, from custom kernels to speculative decoding. The result is the low cost of open source with the latency of a much smaller model. In this episode of Founder Firesides, YC Managing Director Jared Friedman sat down with Wafer co-founders Emilio Andere and Steven Arellano to talk about the college side project that became the company, how they got GLM 5.2 running two to three times faster than anyone else in the market, and why speed and not just cost is pushing enterprises off commercial models. Chapters: 00:00 - What Wafer is 01:01 - Who's using it 01:35 - The $40 million Series A 01:53 - Zero to $8M ARR in four months 02:15 - The GLM 5.2 launch that broke everything 04:21 - Onboarding from a hotel room 05:29 - Neon Health and the latency problem 06:18 - The YC Office Hour Simulator 07:55 - What the A/B test showed 10:03 - AI that optimizes AI 11:44 - Starting as cursor for CUDA 12:55 - The night the agent beat NVIDIA's libraries 14:41 - Open source adoption is going exponential 16:07 - Speed, not cost, is the real reason to switch 18:36 - The side project that became the company 20:50 - Build things for fun, not as startups 21:53 - Turning down the jobs 22:20 - "You don't need a business co-founder" 23:33 - How to pick a rocket ship 24:46 - Hiring at Wafer 25:38 - A day in the life as a member of staff Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Jared FriedmanhostEmilio AndereguestSteven Arellanoguest
Sep 1, 202626mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
September 1, 2026
Duration
26m
Channel
YC Root Access
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

Wafer (S25) is building a fast AI inference cloud that runs open-source models at industry-leading speeds by using agents to optimize GPUs at every level of the stack, from custom kernels to speculative decoding. The result is the low cost of open source with the latency of a much smaller model. In this episode of Founder Firesides, YC Managing Director Jared Friedman sat down with Wafer co-founders Emilio Andere and Steven Arellano to talk about the college side project that became the company, how they got GLM 5.2 running two to three times faster than anyone else in the market, and why speed and not just cost is pushing enterprises off commercial models. Chapters: 00:00 - What Wafer is 01:01 - Who's using it 01:35 - The $40 million Series A 01:53 - Zero to $8M ARR in four months 02:15 - The GLM 5.2 launch that broke everything 04:21 - Onboarding from a hotel room 05:29 - Neon Health and the latency problem 06:18 - The YC Office Hour Simulator 07:55 - What the A/B test showed 10:03 - AI that optimizes AI 11:44 - Starting as cursor for CUDA 12:55 - The night the agent beat NVIDIA's libraries 14:41 - Open source adoption is going exponential 16:07 - Speed, not cost, is the real reason to switch 18:36 - The side project that became the company 20:50 - Build things for fun, not as startups 21:53 - Turning down the jobs 22:20 - "You don't need a business co-founder" 23:33 - How to pick a rocket ship 24:46 - Hiring at Wafer 25:38 - A day in the life as a member of staff Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

SPEAKERS

  • Jared Friedman

    host

    Y Combinator partner and host/interviewer on YC Root Access.

  • Emilio Andere

    guest

    Co-founder of Wafer, discussing product, fundraising, growth, and hiring.

  • Steven Arellano

    guest

    Co-founder of Wafer, focused on ML-systems performance, GPU optimization, and kernel-level work.

EPISODE SUMMARY

In this episode of YC Root Access, featuring Jared Friedman and Emilio Andere, Building AI That Optimizes AI explores wafer uses agents to deliver fastest AI inference for customers Wafer provides a high-speed AI inference cloud and wins customers primarily by delivering the lowest latency and highest throughput through agent-driven GPU optimization.

RELATED EPISODES

Senator Scott Wiener Press Conference at YC

Senator Scott Wiener Press Conference at YC

Building And Structuring An AI Native Company

Building And Structuring An AI Native Company

Improving Small Language Model Reasoning With A* Search

Improving Small Language Model Reasoning With A* Search

LeanAgent: Lifelong Learning for Formal Theorem Proving

LeanAgent: Lifelong Learning for Formal Theorem Proving

Zero-Shot Predictive Models for Relational Databases

Zero-Shot Predictive Models for Relational Databases

Interpretability and Safety for Robot Foundation Models

Interpretability and Safety for Robot Foundation Models

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.