Skip to content
YC Root AccessYC Root Access

Building AI That Optimizes AI

Wafer (S25) is building a fast AI inference cloud that runs open-source models at industry-leading speeds by using agents to optimize GPUs at every level of the stack, from custom kernels to speculative decoding. The result is the low cost of open source with the latency of a much smaller model. In this episode of Founder Firesides, YC Managing Director Jared Friedman sat down with Wafer co-founders Emilio Andere and Steven Arellano to talk about the college side project that became the company, how they got GLM 5.2 running two to three times faster than anyone else in the market, and why speed and not just cost is pushing enterprises off commercial models. Chapters: 00:00 - What Wafer is 01:01 - Who's using it 01:35 - The $40 million Series A 01:53 - Zero to $8M ARR in four months 02:15 - The GLM 5.2 launch that broke everything 04:21 - Onboarding from a hotel room 05:29 - Neon Health and the latency problem 06:18 - The YC Office Hour Simulator 07:55 - What the A/B test showed 10:03 - AI that optimizes AI 11:44 - Starting as cursor for CUDA 12:55 - The night the agent beat NVIDIA's libraries 14:41 - Open source adoption is going exponential 16:07 - Speed, not cost, is the real reason to switch 18:36 - The side project that became the company 20:50 - Build things for fun, not as startups 21:53 - Turning down the jobs 22:20 - "You don't need a business co-founder" 23:33 - How to pick a rocket ship 24:46 - Hiring at Wafer 25:38 - A day in the life as a member of staff Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Jared FriedmanhostEmilio AndereguestSteven Arellanoguest
Sep 1, 202626mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:001:01

    What Wafer is

    1. JF

      [upbeat music] Steven and Emilio, welcome back to YC. Thank you so much for coming back to join us today.

    2. EA

      Thank you for having us.

    3. SA

      Thank you. Thank you.

    4. JF

      Yeah. So you guys just graduated from the University of Chicago about one year ago. You are 23 years old, and today you are announcing that you just raised a $40 million Series A for your company. Uh, congratulations.

    5. EA

      Thank you.

    6. SA

      Thank you.

    7. EA

      Thank you so much, Jared.

    8. JF

      Um, that is a pretty wild story, even by, like, the standards of Silicon Valley and the AI revolution that we are having in right now. Um, and so I'm excited for you guys to tell everyone the story of how you guys went from being University of Chicago undergrads to running this company that's on fire in basically a year. Why don't you start by just telling everyone, like, what is Wafer? What are you guys doing?

    9. EA

      So Wafer is a fast AI inference cloud. We run AI models at the best speeds in the market, and the way we do this is by having agents optimize GPUs.

    10. JF

      What kind of companies are using Wafer? What, what would

  2. 1:011:35

    Who's using it

    1. JF

      I use it for?

    2. EA

      Yeah, totally. So we have a lot of customers that use us for, um, voice agents, because it happens to be that when you're running a voice agent in production, you want the lowest latency possible, 'cause I'm speaking to an agent over the phone and it takes two to five seconds to respond, it kinda like breaks the-- It's a bad customer experience. So it ends up being a lot of what we do, like low latency applications. The other side of what we do is high throughput applications. You can think of, like, coding as one example of this, where we, um, work with a lot of customers like Vercel to make sure their coding agents go as fast as humanly possible on, on the GPUs

  3. 1:351:53

    The $40 million Series A

    1. EA

      they're using.

    2. JF

      Let's talk a little bit about the Series A that you guys are announcing today. Can you guys tell us about it?

    3. EA

      Yeah. So we raised $40 million, um, and this was co-led by, um, Marathon and Chemistry. There were a bunch of, like, awesome angels followed on with, like, Jeff Dean as existing investors, obviously YC. Yeah, I mean, we're super excited

  4. 1:532:15

    Zero to $8M ARR in four months

    1. EA

      about it.

    2. JF

      One reason why you guys raised this crazy oversubscribed Series A is that Wafer, in the last four months, has really blown up. You went from essentially zero to $8 million ARR in four months, which is like bananas, and as a result, you like ran out of GPUs and had to, like, immediately raise money in order to, like, keep scaling. Um, that's insane. Can you guys talk about what happened? Like, what caused you guys to blow up

  5. 2:154:21

    The GLM 5.2 launch that broke everything

    1. JF

      like this?

    2. EA

      Basically, around four months ago, we decided to shift from selling the AI that optimized GPUs to other companies to just use the agent internally to grab open source LLMs and make them run really fast. So we did this, for example, like most famously when we began with GLM 5.2. It was a huge announcement when it came out, and we basically grabbed our agent, we rewrote some kernels, we optimized all levels of the stack of the GPU, and we made it run 2 to 3x the performance that everybody else in the market was running it at. This caused like a huge spike in demand, 'cause suddenly you didn't need to go to Cerebras to get these massive speeds, you can just get them on Open Router. And that was the moment where we sort of projected where our GPU spend and our, where our GPU costs were going, and we were like, "Okay, we cannot sustain this demand [laughs] if it continues," and it indeed has continued. So that was the moment where we were like, "Okay, it's time to go for the next funding round to grow faster."

    3. JF

      And I remember we actually did office hours. I don't know if you guys remember this, but we actually did office hours about two weeks before you launched your, like, fastest GLM 5.2 on the market. And I remember you were like, "Jared, I think when we launch this thing, it might take off really fast, and we might run out of money really quickly."

    4. EA

      Yeah.

    5. JF

      And your question was like, "Should we raise money now?" And I was like, "No, you should probably, like, launch and let it take off and run out of money," and like that's a great story to go to investors and be like, "Our thing is growing so fast we're like running out of money 'cause, like we can't, we can't onboard compute fast enough." And I, I remember being like, "Wow, I really hope that that's what happens," but you know, who knows-

    6. EA

      Yeah

    7. JF

      ... because like, like launches are unpredictable. And then it was pretty wild because you like, a few days after you launched, you were like, it literally played out exactly the way we, we like we thought it was gonna play out.

    8. EA

      Yeah. Yeah. I think the demand for fast inference is, is just extremely high.

    9. SA

      [laughs]

    10. EA

      Um, and if you can actually... It's kinda like revenue will take care of itself if you can actually achieve the best speed in the market. That's kinda like something we follow internally. It's just the technical difficulty of being the number one is just very high and, um, yeah, that's the thing we focus

  6. 4:215:29

    Onboarding from a hotel room

    1. EA

      on.

    2. SA

      Mm.

    3. JF

      I feel like this experience of like having customers signing up for your product so fast you can't keep up with them and having to like scrounge around to find GPUs any place that you can find them, like this is kinda the promised land for a startup. And like Steven, you, you were the guy in charge of like finding the GPUs and like keeping the site up while it was like melting. What, what was the experience of the last few months like for you?

    4. SA

      It's been very chaotic. I, I think back to a team offsite we took to Santa Cruz, which was supposed to be a s- a fun weekend on, on the beach. Uh, we got a nice hotel right, right next to the beach. Everybody had rooms. However, what it turned out to be was actually constant onboarding new Blackwell nodes. Um, everybody just packed in the same hotel room, uh, trying to onboard these nodes, also ensure they're reliable, and then also ensure that, you know, they're performant. We're trying to debug, you know, networking issues a- and so on. So this, you know, this specific instance is maybe a, a taste of how chaotic it's been, but, uh, at the very least, I, I look back and think that this time was very worth it. And also, you know, with the team also for some reasons pretty fun as well, so we're, you know, bonding, bonding along the way.

  7. 5:296:18

    Neon Health and the latency problem

    1. JF

      Tell us about some of the customers. Like you guys went zero to $8M ARR. There must be like some big customers paying you all this money. Who's, uh, who's using it and how are they using it?

    2. EA

      Yeah, so one of, one of the best examples is a company called Neon Health. So they do, um, automation for hospitals. So for example, let's say a customer calls and they want to process, um, some sort of medicine and they want to get the medicine. Usually, there would be another human on the line which fills out a form for, for these humans. What they have done is they basically do this entire process with AI agents- And they really, really care about having low latency. So Neon Health was previously on another larger inference provider, but they decided to switch to us because we were able to give them 30 to 50% better performance, um, on a per call basis. So that's, that's, that's a good example of like the sort of latency improvements that you can achieve with, with Wafer.

  8. 6:187:55

    The YC Office Hour Simulator

    1. JF

      Yeah. I have direct experience with these latency improvements because YC is not just an investor in Wafer, we are also a customer. So perhaps we can talk about how YC is actually using you guys. A while back, I built this product that we call the Office Hour Simulator. It's this pretty crazy idea. It kind of looks like an arcade machine. It's designed to look like that, and you like walk into it, and you hit go, and you can talk to an AI clone of a YC partner. Unlike Neon Health, which does voice agents over the phone, this is more like a Zoom call where there's like, it's like a full like video experience. You're talking to this like AI version of a, of a YC partner, and you can both see and hear them. But similar to Neon Health, we care a lot about the latency because like the, the whole point is like you're supposed to feel like you're actually talking to a YC partner, and so shaving off milliseconds of latency is really important. Um, and I ran this benchmark to try to find the lowest latency LLM in the market that's actually smart enough to still do a good job with the, with the conversation, and you guys actually beat the benchmark in like a, like a clean head-to-head against basically like every model and every inference provider I could find, like I could get my hands on.

    2. EA

      Yeah. No, we're super proud of that, and I think the coolest thing about that is that even for the, for the comparisons that you were doing for other providers, we are using a model that's like 20 times larger. So the fact that we can actually get that model to be much, much faster than the, than the ones you were trying out, um, is, is, is super awesome to us, and that came from just like deep work with like the agents we built to, to enable that latency for

  9. 7:5510:03

    What the A/B test showed

    1. EA

      you.

    2. JF

      Yeah. Before Wafer, I was experimenting with Gemini models and OpenAI models. They both have like pretty light like mini models that are g- that are good for voice agents, but since you guys got GLM 5.2 running super fast, I switched to your GLM 5.2 like super fast mode.

    3. EA

      Mm-hmm.

    4. JF

      And it's like both smarter and better than the commercial models, which is pretty wild.

    5. EA

      Yeah.

    6. JF

      And I can actually share, um, another cool stat. I-- We haven't released this publicly before, but when, when we made the switch, we were actually able to A/B test it, and so I actually ran a three-way A/B test, um, between your model and the two proprietary models. We could not only measure a decrease in the like latency of the LLM, we were actually able to see that the users changed their behavior when the calls got faster, and the users who got the Wafer GLM model actually stayed and talked to the avatar significantly longer than the ones who had the commercial models.

    7. EA

      This is pretty awesome.

    8. JF

      [laughs] Yeah.

    9. EA

      This is pretty awesome. Yeah. It's-- I think having-- We have all these examples too of like where people can ditch other parts of sort of like their, their scaffolding around the LLMs because it's so fast now that it can just do more things live. Um, this is an example where it's li- literally affecting user behavior. So yeah, speed, speed is the mo- as we like to say. [laughs]

    10. JF

      Yeah. Yeah. I mean, it super matters for real-time conversation. Like, you know, even with the fastest LLMs, no LLM voice agent is quite at human level yet in terms of like latency.

    11. EA

      Mm-hmm.

    12. JF

      And so you really feel it when you're talking to a voice agent that it's just like a little slow to respond compared to a human. And so until we can get it to like absolutely human level, which is probably not like imminent, I think it's like really hard to get it like fully human level, like there's gonna be so much gains to like anyone who can shave off some milliseconds. When I saw these benchmark results, I was like pretty stunned, guys. I was like, "Holy shit, what magic are they cooking up in the lab that was like able to make this the fastest provider?" And so I don't know how much of your secret sauce you can share, but like I, I'd love for you to talk a bit about the technology behind it and how you, h-how you made it faster.

  10. 10:0311:44

    AI that optimizes AI

    1. SA

      You know, the headline here is AI that optimizes AI. But, but what does this really mean? I mean, when we went through our, um, benchmarking period with, with Y Combinator, you know, we asked for some information on our end, you know, like the, uh, you know, specifics on the specific like workload that, that YC was running. This is all very valuable information for performance engineers. Um, and that being said, the general flow at, at which we use AI that optimizes AI is we take the workload characteristics of any given customer, um, throw this into, you know, let's say an AI black box, and this AI black box does a ton of things, whether it's, you know, goes and writes custom kernels, uh, goes and trains a speculative decoding model, uh, quantizes the model, a-and so on. In the end of this whole process comes out with, let's say a engine or a runtime that is hyper-optimized for a specific user, allowing us to get 2X, 3X and so on speedups over other inference runtimes. Um, I, I think some cool examples in this specific situation are, you know, maybe looking back to where, you know, let's say this YC example, um, we also have examples of our agent going off and writing custom kernels, um, bringing up AMD hardware to find perf per dollar parity with NVIDIA hardware, um, a-and so on and so forth.

    2. JF

      Even though this new product that you guys launched, the, the Inference Cloud, took off just four months ago when you launched it, this is a little bit of like an overnight success, like a year in the making because all the technology that you just talked about, you'd basically been working on that for like a while, and the reason that it was able to take off so quickly is that you'd spent months doing all of this like very hard, deep technical work, right?

  11. 11:4412:55

    Starting as cursor for CUDA

    1. SA

      Yeah. I think back to our time in YC and, you know, not, not a ton of, let's say-

    2. JF

      ... progress was really made on [laughs] on let's say the, the revenue side. Um, but on, on the tech side, I mean, I'll-- I can give a flashback. We started as cursor for CUDA. From the very beginning, we're using agents to optimize kernels themselves. Um, as we progress more in the ML performance space, we moved up the stack. Uh, we were able to do contracts with folks like AWS writing, you know, kernels for them, uh, do work with, uh, DigitalOcean, optimizing models for them. And, and so we, let's say, developed this muscle, you know, or just develop agents, um- Mm-hmm ... for doing performance engineering. And, and so when we made the switch to become the cloud, it was very easy to get AI to optimize AI because this is something that we had been doing, right, since, since the start of our YC batch last year. Interesting. So the, the first idea was basically customers would give you their, like, their proprietary model and you would use agents to optimize their model and then the big unlock was just switching to, like, open source models that you could just optimize and then just, like, give the optimized version, like, like for anyone.

  12. 12:5514:41

    The night the agent beat NVIDIA's libraries

    1. JF

      Yeah. So there, there's a, there's a funny story that, um, that's, that's fun to talk about here where Steven basically-- We had a big demo with one of the inference providers, one of the large inference providers that has bro-broken out in the last, like two years and we were getting the agent to get results on like what performance improvements it could get over like base LLMs to show them, um, the next day. And basically Steven stays up all night 'cause he just often does that [laughs] . [laughs] And he, he just running the agent, like orchestrating and then he shows me the results at-- in, in the morning and he's like, "We beat like some of these like core NVIDIA libraries with just like the agent running overnight and me just like making sure that, um, we're just, just coordinating like very, very high level coordinating these agents." And I was like, "What?" And he just shows me these results and I'm like... There was this moment where we were just like, "Okay, let's cancel this meeting and let's just meet and like talk about my-- about this more because it just seems that there's so much value in these optimizations that I don't even know if we want to give them away." Interesting. And that was kind of like the start basically four months ago that we talked about where we were like, "Oh wow. Like we really have something here that we might just like verticalize and just compete with the inference clouds directly instead of selling them these, these optimizations." Was another unlock just that the open source models like got really good? Like last summer when you guys started the open source models just like weren't that good. Yeah. This is, this is true. I think GLM 5.2 was a big moment. Yeah. I think GLM 5.2, um, brought out a ton of use cases. Many folks started to really use open source models for their coding workflows. Um, and, and this of course brought a ton of traffic to, to our platform and, um, to my understanding, uh, you know, brought a ton of traffic to just serverless business as a whole for selling tokens. Yeah.

  13. 14:4116:07

    Open source adoption is going exponential

    1. JF

      We were talking with the Ollama and OpenCode founders recently and the sense that I get from them is that with the new wave of open source models that came out in the last few months that like adoption of open source models is just like on an insane exponential right now. I'm curious what you guys are seeing when you talk to like large companies. Like is there like a huge push across the board to like switch from commercial to open s- open source models and, um- Yeah. Yeah. Personally from the like GTM side we've seen like these massive companies start doing inbound to us and I think there's like a very special moment going on where like people are realizing that they can actually run the same workloads that they run with OpenAI and Anthropic and save let's say like 80% of the cost. So on the cost axis there's just like whatever like lands on the CFO desk and like big, big enterprises want to cut those costs and they're seeing the token cost rise and then on the other front they're seeing how much is possible when you can actually get like 5X the speed with open source models. Yeah. So if you can actually have an Opus level model that will respond five times faster- Faster ... that's a huge value add for some products. Like think of your usage of like Granola. They're not one of our customers but like when I'm in a sales meeting and I ask Granola something I want it to respond as fast as humanly possible because it's like something that the customer is asking me for, right? I would pay a lot of money to just-- I would pay extra per token money to get a faster response. There's many applications of s- like the performance sensitive applications that we really like are seeing, um, trickle down.

  14. 16:0718:36

    Speed, not cost, is the real reason to switch

    1. JF

      So I think this is a really, um, insightful point because I think a lot of people, um, think that the, like the reason to switch to open source models is that they are cheaper and that like-- which like fair enough is like a good reason to adopt a thing but I think a lot of people think that's the only reason to switch to open source models is sort of like, ah it's like almost as good as Anthropic and OpenAI but it's like a lot cheaper and that's probably driving some of the demand but what I didn't realize until a couple months ago when I started using Wafer is that it's actually-- there's a whole other reason to switch to them which is they are actually the best models if you have a latency sensitive use case. They are actually the fastest models. Like for my office hour simulator I was not price sensitive like YC is in a fortunate position where like we don't care about the LLM cost like at all and so I was actually happy to pay a huge premium for the fastest model that money could buy and it turns out that the fastest money the model could buy was GLM 5.2 running on Wafer. Yeah. And there's many situations where people's core product depends on how fast they can deliver, um, the, the actual response so we're very excited about all the applications that were just like the core product depends on how fast you can generate a response. Totally. I would actually-- I, I actually think for my office hour simulator that like basically switching to GLM 5.2 that was actually the key turning point where it found product market fit. Like I actually think like before GLM 5.2, it was just like a little too slow or a little too stupid, and so users would churn. GLM 5.2 in Wafer is like just fast enough that people are like, "Oh, actually this is great. I'm just gonna like talk to it all day."

    2. EA

      Mm-hmm.

    3. JF

      [chuckles]

    4. EA

      Have you seen people talk to it all day?

    5. JF

      Yes.

    6. EA

      Wow. [laughing] I mean, that's great for us. Please do.

    7. SA

      We'll start having folks schedule-

    8. EA

      [laughing]

    9. SA

      ... time blocked.

    10. EA

      That's awesome though. I mean, having access to infinite Jared is obviously very valuable for the world.

    11. JF

      Well, it was cool because I feel like because we're sort of like building at the cutting edge and sort of getting access to these products like before the rest of the world, that I have like a little peek into the future that like other people didn't have.

    12. SA

      Mm-hmm.

    13. JF

      Yeah. Maybe let's rewind back to the very beginning and like talk about how you guys ended up starting this company. I actually met you guys before you had started Wafer when you were undergrads at the University of Chicago. Maybe you guys can like tell us the story of like how you guys met each other in school and what you were doing at the University of Chicago and like how you ended up deciding to start a company together.

  15. 18:3620:50

    The side project that became the company

    1. EA

      Yeah, I mean, there was a point where Steven showed me, um, a tool he was building, and it was called Yakka, um, which was actually what we demoed at YC sometime later, like an expanded version of Yakka. And this was a thing that basically grabbed some piece of C++ code and asked an LLM to just optimize the code and just loop through it. Like, hey, optimize it, measure if it's better, then do it again, do it again, do it again. And I just thought that it was very interesting to think of this new idea of having the AI as a sort of, maybe as a sort of compiler where you can just grab some piece of code and through the LLM you are able to find how to map it to the hardware so that it runs more efficiently, right? So basically there was one point where we were like, okay, you were gonna go to New York. You were gonna do your job as like, um, you engineering at Two Sigma, obviously like really good job. I was gonna come to a Series A startup to do like AI research. First of all, it would be way more fun to just like keep working together [laughs] in something. So there's like one reason for that where it's just like very fun. And also this exact problem that like the thing that we've been sort of working on here, um, of the Yakka that you showed me, um, it has the like thing that like if you do solve this problem it's like obviously very impactful 'cause like a bunch of energy is being wasted on just not running GPUs as efficiently as they should just because the software is not optimized.

    2. JF

      Mm-hmm.

    3. EA

      And the other aspect was like we're both like uniquely suited and very interested in this challenge. Like it's something that we see ourselves working like for the next hopefully like 20 years on, right? So we were like, "Yeah, why not?" And we-- and I mean, we applied to YC, moved to San Francisco and, and that was kinda like the start of, the start of Wafer.

    4. JF

      What made you decide to build Yakka in the first place?

    5. SA

      When you think about a lot of folks in quant finance space, they, they really care about how good or how fast their code is. Um, you know, HFC is, uh, a very, very competitive field and most oftentimes you, you win at, at let's say the, the nanosecond, right? And so I, I was really curious on, um, at least during, during my time in, in these experiences, it seemed that folks really cared about, you know, really fast C++, really fast Python, so on. How can I get an AI to, to solve this, uh, solve this problem of generating this really fast code? 'Cause it seemed relatively au-automatable.

  16. 20:5021:53

    Build things for fun, not as startups

    1. JF

      So it really did come out of your like quant experience 'cause like, yeah, like quant firms are obsessed with like squeezing like, you know, like microseconds of performance out of, out of code. But like it wasn't initially it wasn't to be a startup, right? It was like, it sounds like it was just like you were just building it just to see if it was possible to build just 'cause like you were interested in like exploring it, right?

    2. SA

      Yeah.

    3. JF

      When we talk to college students who are interested in starting companies, one of the pieces of advice that we often give them is that they should just build projects just for fun that aren't like trying to be startups just 'cause like they find them like technically interesting just to explore like, "Oh, I wonder what the models are capable of." And it sounds like you, like that, that like pretty much perfectly describes what happened here where you had this like side project that you were working on just 'cause it was interesting and then you guys began talking about it and you're like, "Wait, this side project, this could be an actual startup." Let's talk quickly about the decision to do the startup though 'cause when, when I met you guys at the University of Chicago it was your senior spring and you both had great jobs lined up to do after you graduate and you were not planning to do a startup.

  17. 21:5322:20

    Turning down the jobs

    1. JF

      You were planning to go do these jobs at great companies. Um, and then like something changed in your mind in the next couple of months where you're like, "Actually, maybe this project that Steven built before, maybe we should just like go do that as a startup." Like I'm curious like what that decision process was like for you guys. Like was, was it like a difficult decision where you went back and forth? Did your parents and your friends try to talk you out of it? Was it emotional or was it just like, "Yeah, sure why not?"

  18. 22:2023:33

    "You don't need a business co-founder"

    1. EA

      My thing was very simple. It was like when, when YC came and gave the talk, I think we've talked about this before, but you guys had a, a line that I still remember which is like you don't, you don't need a, like a business co-founder. Steven and I like just we were both engineers. Um, I've never had a real job. [laughing] So it was kinda shocking to hear that directly from you guys and I think on my side this is what like started conversations more seriously but yeah, curious on Steven.

    2. SA

      Yeah, I, I, I think personally that held a lot of, how to say, weight also in my decision primarily from, from a perspective of now it, it flipped a switch of all I have to do is look around at, at many of my peers and many of the folks who I hang out on a daily basis and, and really, um, just ask should we, should we do this? Um, I, I personally, you know, throughout life have always wanted to start a company or, or at least have a lot of ownership in something that I was to pursue. Um, and it seemed that the opportunity was, was definitely there to one, pursue something worth a lot of let's say, um, you know, value to society while also, you know, maintaining this, you know, let's say ownership and, and being able to do it, uh, on, on my own.

    3. JF

      Do you guys have

  19. 23:3324:46

    How to pick a rocket ship

    1. JF

      advice for other people who are interested in starting companies in, like, a similar space or, like, people maybe with, like, similar backgrounds to yourselves? Like, imagine rewinding back to, like, a year, a year or a year and a half ago. Like, what is, what is advice, like, you wish other people had told you?

    2. EA

      If you're interested in startups at all and you're sort of looking into what startups are, um, the best, there's a lot of people that I've talked to recently that tell me a lot about the role. But I'm a very big believer on, like, if you're offered a seat in a rocket ship, you just don't ask what exactly you're gonna do within the rocket ship.

    3. JF

      Hmm.

    4. EA

      What you should be evaluating in a company as, like, more and more people are evaluating, like, joining, like, startups and, and YC and many others, like, after grad or starting one, but specifically for the people that are evaluating, um, joining one, what you should really ask is, like, how could this company go from, like, where it is today to, like, 1000x? And I would specifically focus on, like, the founders and, like, the founding engineers. Like, what is the culture of the company, and will that enable us to, like, grow massively? Because, like, the job itself is gonna change super fast. We've never hired for specific roles because the roles just change super fast. So that's kinda something I've been thinking about, of just, like, make sure that you're investing in the right company and founders instead of, like, your specific role within that company.

    5. JF

      I feel like that's

  20. 24:4625:38

    Hiring at Wafer

    1. JF

      a great segue. Um, speaking of joining rocket ships, you guys just raised $40 million. Are you hiring? Uh, can you tell us a bit about, like, who's on the team now and who you're, who you're planning to add?

    2. EA

      Yeah. Um, I can talk about, like, we're, we're, we're definitely hiring. Um, we're hiring across the board. We're hiring... Basically, the engineering role that we hire for is member of technical staff. This is a person that doesn't need to have any, like, GPU experience. What we really care about is, has this person been exceptional at anything before? This is basically our recruiting guideline that has worked, uh, super well in the past and that we're, like, doubling down on. So if you've done anything exceptional in the past, it doesn't have to be computer-related, doesn't have to be anything, please, like, reach out because we're... Yeah, we're hiring as, as quickly as we can the, the, the best people at, at ML Systems, um, performance stuff.

    3. JF

      Yeah. And can you, like, tell us a little bit about, like, I don't know, like, a day in the life as a member of technical staff at Wafer? Like-

    4. SA

      You know, it's,

  21. 25:3826:39

    A day in the life as a member of staff

    1. SA

      it's variable from day to day. [laughs] Um, you know, what we did last month is very different than what we do this month. But at the end of the day, I think, uh, I wanna put that we're problem solvers first. Most often, you know, whether, uh, it's a customer or, you know, a specific issue we're having with GPUs we, we, we're operating with, um, we're, we're being fed, like, technical problems from the top, and most oftentimes it's, it's just a matter of, of solving these problems. You know, what this actually looks like, um, could be, you know, things from, uh, how do we get an agent to write this specific kernel, or how do we solve this specific reliability issue we're seeing with these GPUs? Um, uh, there's a lot of variance, um, but at the end of the day, I, I very much so view it as, uh, solving problems within the domain of a ML performance or ML system. Yeah.

    2. JF

      I feel like that's a great note to end on. Um, thank you guys so much for joining us today, and huge congratulations on, uh, the Series A.

    3. EA

      Thank you so much, Jared.

    4. SA

      Thank you.

    5. EA

      Yeah. Thanks for having us. [outro music]

Episode duration: 26:41

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode 7JoqmM5EPXo

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.