Skip to content
YC Root AccessYC Root Access

Building AI That Optimizes AI

Wafer (S25) is building a fast AI inference cloud that runs open-source models at industry-leading speeds by using agents to optimize GPUs at every level of the stack, from custom kernels to speculative decoding. The result is the low cost of open source with the latency of a much smaller model. In this episode of Founder Firesides, YC Managing Director Jared Friedman sat down with Wafer co-founders Emilio Andere and Steven Arellano to talk about the college side project that became the company, how they got GLM 5.2 running two to three times faster than anyone else in the market, and why speed and not just cost is pushing enterprises off commercial models. Chapters: 00:00 - What Wafer is 01:01 - Who's using it 01:35 - The $40 million Series A 01:53 - Zero to $8M ARR in four months 02:15 - The GLM 5.2 launch that broke everything 04:21 - Onboarding from a hotel room 05:29 - Neon Health and the latency problem 06:18 - The YC Office Hour Simulator 07:55 - What the A/B test showed 10:03 - AI that optimizes AI 11:44 - Starting as cursor for CUDA 12:55 - The night the agent beat NVIDIA's libraries 14:41 - Open source adoption is going exponential 16:07 - Speed, not cost, is the real reason to switch 18:36 - The side project that became the company 20:50 - Build things for fun, not as startups 21:53 - Turning down the jobs 22:20 - "You don't need a business co-founder" 23:33 - How to pick a rocket ship 24:46 - Hiring at Wafer 25:38 - A day in the life as a member of staff Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Jared FriedmanhostEmilio AndereguestSteven Arellanoguest
Sep 1, 202626mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Wafer uses agents to deliver fastest AI inference for customers

  1. Wafer provides a high-speed AI inference cloud and wins customers primarily by delivering the lowest latency and highest throughput through agent-driven GPU optimization.
  2. The company’s breakout came after pivoting from selling optimization tooling to offering an inference cloud for optimized open-source LLMs, highlighted by making GLM 5.2 run 2–3x faster than alternatives.
  3. Real-time use cases like voice agents and interactive video avatars amplify the value of speed, with YC’s A/B test showing faster responses changed user behavior and increased session length.
  4. The underlying moat is a year of ML-systems “deep work,” where agents write custom kernels, tune runtimes, and apply techniques like quantization and speculative decoding per workload.
  5. Founders emphasize building for technical curiosity first, taking “rocket ship” opportunities, and hiring exceptional generalist problem-solvers rather than narrow GPU specialists.

IDEAS WORTH REMEMBERING

5 ideas

Speed is a primary product differentiator in AI, not just cost.

Wafer argues that for real-time experiences (voice, interactive avatars), shaving milliseconds materially improves UX; YC’s test showed faster models kept users talking longer.

A targeted pivot can unlock distribution when the tech is already mature.

They spent months building optimization agents, then switched from selling optimization as a service to deploying it internally on open-source models—instantly broadening the market.

“AI that optimizes AI” is a practical workflow, not a slogan.

Wafer takes customer workload characteristics and uses agents to produce a tuned engine/runtime by writing kernels, tuning the stack, quantizing models, and training speculative decoding components.

Open-source models can win on latency because you can optimize the entire stack.

Unlike closed APIs, open models allow deep systems-level tuning; Wafer claims this is why GLM 5.2 on their platform beat proprietary options even when using a much larger model.

Operational scaling becomes the bottleneck when product-market fit hits in infra.

Going from 0 to $8M ARR in four months created immediate GPU supply and reliability challenges, forcing rapid node onboarding and prompting the Series A to scale compute.

WORDS WORTH SAVING

5 quotes

So Wafer is a fast AI inference cloud. We run AI models at the best speeds in the market, and the way we do this is by having agents optimize GPUs.

Emilio Andere

It's kinda like revenue will take care of itself if you can actually achieve the best speed in the market.

Emilio Andere

You know, the headline here is AI that optimizes AI.

Steven Arellano

We started as cursor for CUDA.

Steven Arellano

If you're offered a seat in a rocket ship, you just don't ask what exactly you're gonna do within the rocket ship.

Emilio Andere

Fast AI inference cloudLatency-sensitive voice and real-time agentsThroughput for coding agents (e.g., Vercel)Agent-driven GPU/runtime optimizationCustom kernels, quantization, speculative decodingOpen-source LLM acceleration (GLM 5.2)Pivot strategy, scaling GPUs, and hiring approach

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.