Skip to content
Stanford CS153 Frontier Systems | Mati Staniszewski from ElevenLabs on The Future of Voice Systems
This video isn’t embeddableWatch on YouTube →
Stanford OnlineStanford Online

Stanford CS153 Frontier Systems | Mati Staniszewski from ElevenLabs on The Future of Voice Systems

For more information about Stanford's online Artificial Intelligence programs, visit: https://stanford.io/ai Follow along with the course schedule and syllabus, visit: https://cs153.stanford.edu/ In week two of CS153 ("AI Coachella"), Anjney Midha interviews Mati Staniszewski, founder and CEO of ElevenLabs, tracing the company’s origins from an early Discord text-to-speech bot to a fast-growing frontier audio and speech platform. Mati explains ElevenLabs’ initial focus on solving AI dubbing inspired by Poland’s single-voice film narration, the shift to prioritizing emotional, natural-sounding text-to-speech for creators, and the evolution from cascaded pipelines (transcription, translation/LLM, and speech generation) toward real-time voice agents. They discuss tradeoffs between cascaded versus fused multimodal systems, efforts to detect and convey emotion, safety and voice authentication limits, on-device model deployment, collaboration with teams like Sesame, and business lessons on PLG plus enterprise deployment, team structure, pricing from customer value, and growth to over $430M revenue with ~450 employees. Guest Speaker: Mati Staniszewski is the CEO and co-founder of ElevenLabs, the AI voice/audio platform. Born in 1995 in a town outside Warsaw, Poland, he attended Copernicus Bilingual High School in Warsaw before earning a degree in mathematics from Imperial College London. While at Imperial, he organized Mathscon, a UK student-led mathematics conference. His earlier career included roles at Opera Software, BlackRock (where he worked in the Portfolio Analytics Group and helped launch the Aladdin Wealth platform), and Palantir Technologies (as a Deployment Strategist managing large-scale public- and private-sector implementations). In 2022, he co-founded ElevenLabs with his high school friend Piotr Dabkowski. He has raised hundreds of millions from investors, including Sequoia, Andreessen Horowitz, and Salesforce Ventures, with the company valued at $11 billion as of February 2026. He joined the board of Klarna in 2025 and was named to Forbes 30 Under 30 Europe in 2024 and TIME's 100 Most Influential People in AI in 2025. Follow the playlist: https://youtube.com/playlist?list=PLoROMvodv4rN447WKQ5oz_YdYbS74M5IA&si=DOJ5amlyRdyMJBhG

Anjney MidhahostMati Staniszewskiguest
May 4, 20261h 6mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:073:06

    CS153 welcome + early ElevenLabs origin story on Discord

    1. AM

      Welcome to week two of CS153, also known as AI Coachella. We are super lucky to be kicking off this week with Mati. Mati is the founder and CEO of ElevenLabs. How many people here have heard of ElevenLabs? All right, so pretty much everybody. Mati and I go back a ways. About three years ago, I think, when I was still running platform at Discord, um, a friend said, "You know, Anj, there's a, a little bot, like a, a text-to-speech bot on Discord that's blowing up. Um, you should check it out." Um, and, you know, we had a lot going on at the time at Discord, and so I, I actually didn't, and I should have. And then a month later, somebody pinged me again a- and said, "You really should check out this bot." And I, I checked it out. It was called ElevenLabs, and it was quite an extraordinary bot. It, it, it was a, um, a Discord bot that allowed you to generate audio clips with just a text prompt. Uh, and within 24 hours, I'd asked one of our mutual friends, Nat Friedman, to introduce us. Mati was gracious enough to explain what they were working on. I had-- You let me come on as an angel investor, so thank you. Um, and since then, Mati has gone on to build one of the most, um, the fastest growing, uh, one of the most widely used, and I would say trusted brands and services in frontier audio and speech. Um, so thank you for joining us, Mati.

    2. MS

      Thank you.

    3. AM

      Thank you so much.

    4. MS

      Thank you so much. Good mor- good morning, everyone. It was also a crazy thing. Anj, Anj, uh, when we met for the first time, it was me and my co-founder, Piotr, uh, we both came from Google and Palantir before that. So we were trying to, like, redo the company setup from scratch of, like, what not to do, and we tried to, like, go against some of the lessons from those days. Um, so we were allergic to meetings. We were allergic to, um, to, like, any email-based communication internally. But we also want- wanted to not do any of the internal communication the standard way. So when we started, we actually ran the company on Discord.

    5. AM

      I did not know that.

    6. MS

      So in that conversation, you were, you were helping us, A, on, on, on, on the text-to-speech. And we were trying to, like, figure out, is that the right play for us to base all the company on Discord? We swapped from s- to Slack-

    7. AM

      I know. Sad times

    8. MS

      ... uh, which, which was-

    9. AM

      I'm aware

    10. MS

      ... which was, uh, easier for Freddie. But that was, uh, that was an interesting few, few, few first months of trying to build all the bots on Discord to, like, make it easy and quick for us.

    11. AM

      Th-th-this was a bit of a theme we talked about last year, too, which is that often gaming ends up being this petri dish for innovation. Some of the hardest infra, product, design experience problems that are solved in gaming then become sort of, uh, um, leading indicators for the rest of the world. And the stuff you were doing, and a bunch of other, our friends were doing on Discord at the time, have ended up becoming indicative of, of, y-you know, value in AI a few years later. Is that-- Do you feel like that's an, a true assessment, or, uh, am I overfitting?

  2. 3:065:13

    Why community-driven PLG mattered: finding real problems and unexpected use cases

    1. MS

      Yeah. No, I think the, the, the true part there, which, you know, we, we were following the journey model at the time of, like, how they've built that community piece on Discord. And for us at ElevenLabs, when we started, we knew that we want to fix two things. We want to fix the research and foundational models around audio and voice, and then build product around that to, to bring that AI into more of an applied AI setting and fix the problems that our customers are facing. We started on a very PLG-driven motion, so working on the product-led growth with a lot of the, the, the, the creators in the space of developers in the space. And we thought that the best way to do it is close the, the, the loop as close as possible to the people that are using those tools. And, like, Discord at the time and, and, and, and generally keeping access open to a lot of those creators and developers was the best way for us to learn, is it good enough? Is the quality finally there to, to serve the needs they have? To what are the use cases we might not predict that people might want to build so we can bring that, uh, quicker and then free? Um, and that's still a big tissue today of our work across that is, um, we want to work with the community to find a ways for them to contribute back to the product development. And of course, whether that's using the models to refine based on the data of, of how you use the model all the way through that can you contribute. In our case, we've created a voice marketplace where k- people can con-con-contribute their voice to, to be used by others. So that community aspect was very important, and it's always, uh, true where I feel the technology adopted by the community will show you use cases-

    2. AM

      Right

    3. MS

      ... that might, like, diffuse to the rest of the world six, 12, 18 months later. So being, like, close is, is super valuable. But more so, more so than not in that early days, just, uh, I think you need to be, like, extremely problem obsessed. What is the problem that they are having? And, and the variation of what you think the problem is to what the customer actually thinks is a problem is slightly, is slightly different.

    4. AM

      Okay. Let, let's actually stop. Uh, take, take a beat there. Can you go back, take us back in time. What was the problem you guys were obsessed with when you started Eleven? What is it today? How might it evolve? Give people a bit of a ElevenLabs 101. How do, how do we get here?

  3. 5:137:07

    ElevenLabs 101: the Poland dubbing pain that sparked the mission

    1. MS

      Cool. Cool. The whole chronology. I will, I will, I'll give you a zoom-in into that first day and then, and then, and accelerate over last few years. But when we started-- So I'm from Poland, my co-founder is from Poland. A very peculiar thing that happens in Poland is that if you watch a foreign movie in Polish, all the voices, whether that's a male voice or a female voice, get narrated with one single character. So you have one voice reading every character. As you can imagine, a pretty, pretty terrible experience. And you would think that with the modern technology, this is a problem that would have fix- been fixed, and no, it's still the case. So most of the content is delivered this way. So that was the-

    2. AM

      Wh- Whose voice was this? Who did the-

    3. MS

      They have five characters. There's like five of those voices, usually monotone, male, deep, old voices. Uh, uh, and the, the-- It's also crazy because the part of the thing is they are kind of encouraged to deliver the, the movie in a flat delivery, so-

    4. AM

      Ah

    5. MS

      ... the audience can interpret the emotions for themselves, uh, which is like another-

    6. AM

      Wow

    7. MS

      ... another, another [chuckles] level.

    8. AM

      They're expecting a lot from the audience.

    9. MS

      They do expect a lot. Um, so if you, uh, like any Polish person, if you ask them, they will, like, account how-Like, not good experience, that is. And when you learn English, you finally get to learn everything in original, and that's a, an extremely positive one. So that was, like, the first piece and inspiration for us. We know the future is different. The future will be where you can access all types of content in any language, uh, with that incredible tonality, incredible emotions. So, so, uh, so we left Google, we left Palantir at the time, and, um-

    10. AM

      Were you guys both, both in the Bay at the time?

    11. MS

      We are both, uh, between Warsaw and London. So-

    12. AM

      So-

    13. MS

      At the time when we started, we started in, in, in London.

    14. AM

      Right.

    15. MS

      Then moved to Warsaw for a little bit, then moved back to London.

    16. AM

      Yep.

    17. MS

      So from Europe, uh, at the time, which actually was, uh, for that part was pretty useful because the whole language thing is, is a big problem there, less of a problem in US, so kind of the inspiration might, might have not occurred, uh, otherwise.

    18. AM

      Hmm.

  4. 7:078:19

    Decomposing AI dubbing into a pipeline: transcription → translation → speech generation

    1. MS

      And then we knew that we wanna fix that AI dubbing problem, and as we started going deeper, we knew that we need to understand two parts. We need to understand whether there is a, uh, a research potential for us fixing it, and two, whether there's an actual product need, actual problem for the customer. So, so Piotr, my co-founder and a amazing researcher, started diving in and quickly realized that the current models could potentially create a Frankenstein version of the dubbing. They're not good enough, but they are almost giving you an experience where you can switch from language A to language B, preserve the voice, preserve the intonation emotions, but still not good enough. So there will be research, but possible research to fix.

    2. AM

      Hmm.

    3. MS

      Now, in that research, a very important part is, as you think about the dubbing problem, from speech, there are three models that need to play a, a, a role. You need to transcribe what's happening, who is the speaker, remove the background sounds, background noise, then you bring it to text. You translate it to the, another language. Uh, you might need to do additional set of corrections, and then you recast it back on the other side of text-to-speech to produce audio on the other side borrowing from the original performance. So there are those three key models: transcription, translation, and, and, and the text-to-speech on the other side.

    4. AM

      Hmm.

  5. 8:1913:09

    The pivot: from full dubbing to best-in-class text-to-speech for creators

    1. MS

      And you need to fix all the components to make it good. While at the same time, I was trying to figure out, does anybody need that? So we would call the email of creators, lot of studios, saying like, "Hey, if dubbing was a problem," if-- Sorry, "If dubbing was possible automatically, and you could take your movie and bring it to all international language with your own voice, would you be interested?" Every so often we would give, uh, give them, like, a early version of the samples and, um, and frequently the reply we got was, "Yes, interested. But if you can do that, could you also do just a simpler voiceover corrections for me? Because when I record," let's say even as we record now, "after that, some of the parts are n- not recorded properly, and I want to fix them." Or, "Could you replace my voice from the, from the script just in my original language so I can narrate that without appearing in the script and screen?" Um, and that was, like, a very clear piece where as he dived into the space on research and we dived into the user problem space, it was clear that there's other problems that we can fix first.

    2. AM

      Hmm.

    3. MS

      And we can focus research on one of the components instead of all of those components and, and, and actually bring that to the market and bring that to, to there, to users.

    4. AM

      So a couple of things. I mean, the class is called Frontier Systems, right? Um, if you notice what Mati just talked about is the anatomy of what i- is their intelligence pipeline. Remember last, last week we talked about the anatomy of, of how, um, intelligence is manufactured, but he just gave you a breakdown into the system of, of that Eleven was trying to build at the time, which was this cascaded workflow, right? You had the TTS.

    5. MS

      Exactly.

    6. AM

      You had an LLM in the middle to do some reasoning, essentially, right? And then you have a, had a speech-to-text model.

    7. MS

      Exactly.

    8. AM

      No, actually, the, the other way around. You had transcription, you had an LLM, and then you had-

    9. MS

      Exactly

    10. AM

      ... the generative model.

    11. MS

      Exactly. And, like, it was still 2022, so this was a year where the topics of the day were crypto and metaverse, so the, it was still before the GPT moment.

    12. AM

      The good old days. Yeah. [laughs]

    13. MS

      [laughs] The different days. And LLM translation was still relatively poor, so we were, like, using both, uh, uh, attempts. So, like, the whole pipeline didn't produce the right result, which was a great piece for us being close to the users trying to get, like, what is the actual problem, because we then shifted, okay, let's fix the text-to-speech, which is the most common denominator of cr- of all those problems, and let's fix the voiceover for the creators. And then we've noticed that there's other parts of, like, just reading the scripts, reading articles, reading books that we can deliver. So in 2022, we decided that the first set of, of the biggest potential will be the more basic version, which is just bringing text into audio, making it sound human-

    14. AM

      Hmm

    15. MS

      ... making it sound emotional, um, in just English.

    16. AM

      So you decided not to innovate on the transcription part or the LLM part, just the last mile-

    17. MS

      Exactly

    18. AM

      ... with the generation. And you, you said the mission there was let's try to improve the state-of-the-art in, uh, na- like, it had to sound natural? Is that-

    19. MS

      It needs to sound natural. So the, the main things that were not possible at the time is you couldn't really replicate a voice, um, with the same characteristics, and two, you couldn't make it sound and follow the, the, the entire delivery in the way you would expect. So if you have a fragment of text, it's a, say, a happy, happy sentence. If you are reading it, you know it is happy, you deliver that in a happy way. If it's a dialogue sequence in a book, you know it's a dialogue, so when you read it or a voice actor would read it, you know you need to read it as a dialogue.

    20. AM

      'Cause you have context of the-

    21. MS

      You have the context of the entire thing.

    22. AM

      Right.

    23. MS

      And, um, and at the time, at the same time, the LLM breakthrough started occurring where you knew that you could predict the next token based-

    24. AM

      Right

    25. MS

      ... on a previous token, so you could bring that context into account in a smarter way. So that was the first big thing that we knew we could solve, and two, we could recreate the voice characteristics of, uh, um, a lot better. So instead of... The common approach at the time was effectively hard coding parameters of a voice.

    26. AM

      Right.

    27. MS

      The, the gender, the accent, the age of the voice, and trying to predict those. Instead, we can keep them more abstracted and let the model try to define what those parameters are. And, and, and those two innovations we knew will fix the research side potentially, and then on the product side, we know- knew that a lot of those people will want to... create their longer forms with books and audiobooks, create, uh, scripts and turn them into audio, do the corrections on the, on the voiceover for a video. So we knew that we'll need to build a product for making that easy and, and go all in on helping the creators and, of course, the API to developers to build, to build with that.

    28. AM

      So at, at this-- It's 2022. You've had this clarity now that, okay, the, the state-of-the-art, the frontier we're gonna push is the, the, this, this fle-flexibility of the voice, the generation part, right? Um, but it was just you and Peter. You hadn't raised any money. I don't think so. Um-

    29. MS

      Not yet, yeah. No

    30. AM

      ... what did you do first? Did you, did you go and look for an open model? Did you try to go call somebody up and use an API? What, what was step one?

  6. 13:0915:29

    Early research approach: open source, papers, and Tortoise as a catalyst

    1. MS

      Step one, we, we, of course, were drawing from our savings to, to, to, to get the models off the ground and, um, and, and get that first, like, are we solving the right problem? And, uh, uh, you know, that, that-- I think the, the clear thing for, for, for anyone here is it's-- it was so valuable for us to be as close to the users as possible to keep that interactive loop pretty quick. Um, but then to your point, like, the, the main thing as we were, like, actually trying to do the research was, of course, look what's available on open source, what's available in closed source, and, um, and then look for the papers in the space. Are there other innovations in those papers that might have not been applied to audio?

    2. AM

      Right.

    3. MS

      And that was the case. Uh, there was, uh, the closed source was still lagging, not nothing really good, but at the time, the hyperscalers, so the Googles of the world were still kind of leading on the best text-to-speech. Open source was ahead, so you had some, uh, inklings of incredible voice models. There was a, a famous model, uh, called Tortoise from, from an incredible guy, uh, James Betker, who's-

    4. AM

      Uh, we're gonna have James come do office hours in the class for everybody. Yeah.

    5. MS

      Oh, okay. Amazing. But crazy thing about James, so he was, he was out of Google at the time. Uh, he was working I think on infra or something unrelated, and in his spare time, he was exploring audio and voice and effectively created the best open source model of that time w-

    6. AM

      On nights and weekends. Basically-

    7. MS

      On nights and weekends

    8. AM

      ... as a side project. Yeah.

    9. MS

      Uh, and, uh, and it was, it was, it was finally the model that could produce something that sounded human-like on short fragments. So you could have amazing delivery, the right prosody, the right intonation, right emotions. However, two things were impossible at, with, with that model. One, it took extremely long time to generate anything, and then two, it was very unstable below the, the, um, above that short sentence. So that was a, a, a tricky limitation. Uh, and then on the papers, there was, of course, a lot of good innovations coming from the space.

    10. AM

      Yeah.

    11. MS

      The Diffusion in 2021 came out. Um, uh, the, I mean, the Transformer was like, uh, four years, uh, prior to that. But still, some of the ideas from Transformer were just getting through-

    12. AM

      Yep

    13. MS

      ... to the audio space. So we knew that combining some of those ideas, we can, we can potentially take inspiration from open source, take inspiration from those papers, and try a slightly different architecture for text-to-speech and, um, and for voice creation of how you create those voices.

  7. 15:2917:23

    Compute constraints, startup pragmatism, and why patents didn’t matter

    1. AM

      Do you remember by any chance how much compute you spent on the first Eleven checkpoint that was inspired by Tortoise?

    2. MS

      It was tiny amounts in comparison. It was tiny. So one thing that I recommend, uh, if you are looking to start a side project or a company is, uh, is look through all the, uh, accelerator, in quotes, programs from big companies that give you free compute and free credits.

    3. AM

      Oh, that world is gone. There's no free compute anymore.

    4. MS

      Okay. Okay. So-

    5. AM

      Ex- except for maybe f- in, for this class of students, actually.

    6. MS

      I should know that before saying this, I guess. But there was a great-- Maybe they still do something. They-- So NVIDIA Inception program and a few others, like, gave us-

    7. AM

      Yeah. Yeah, I remember that

    8. MS

      ... the-- So we, we-

    9. AM

      Back when GPUs were still available.

    10. MS

      Exactly.

    11. AM

      Yes. Yes.

    12. MS

      So, like, the first ones for us were in, like, in tens of thousands l- of dollars, and that felt, that felt, that felt big. Maybe it was, like, approaching the hundred dol- thousand dollar category, but, like, still tiny. Um, I remember like, you know, you, like, in, in that days, you, the budgeting and, and the sayings, uh, like, um, uh, like optimizing your budget is so important. It's not the compute piece, but I remember us arguing whether we should do a patent for our work, which we decided against. Uh, but I remember the, the, the lawyer quoted us, uh, for the entirety of the patent work, $6,000. And we said, "There's no way we are paying that amount of money for, for, for anything at this stage," and decided not to do it.

    13. AM

      Well, it seems like not having a patent has not held ElevenLabs back.

    14. MS

      No, I don't think. I, I think, like, we, like, ultimately, like, is it, is it even valuable? The, the, the sh- you're innovating so quickly that it becomes-

    15. AM

      Right

    16. MS

      ... obsolete. Two, like, would you ever want to stop innovation from the other side? So, so likely not. Do you do it as a defensive measures? So, like, then we learn about the whole patent trolls industry-

    17. AM

      Exactly

    18. MS

      ... that other people that will try to attack you with their own patents that aren't great, but just waste your time. Um, so, like, do you do that? Then we decided, like, no, we'll just fight those cases anyway.

    19. AM

      Yeah.

  8. 17:2321:59

    From TTS to full audio stack: transcription, agents, and beyond (2022–2026 roadmap)

    1. MS

      So, so, so it didn't stop us. But, um, but bottom line, so small models in comparison. They were, like, h- in hundreds of millions, uh, parameter models at the time, if not lower. Uh, um, so that was relatively small text-to-speech models. And you asked earlier, just to give you, like, a quick-- So that was 2022, and then the kind of the common facet across ElevenLabs is we continued across research over last years. So the models, of course, g-got bigger, and that applied across the entirety of audio. So we did text-to-speech model to start, then built a transcription model or speech-to-text model to understand what's happening on the audio side. Um, the wider set of how you bring those models together in AI dubbing, how you bring those models in interaction, so how you bring a conversational tissue around the speech-to-text, the LLM, the text-to-speech together, and even expand it to music. So entirety of audio and how voice models or audio models work together with other modalities. And then, uh, alongside build a platform that helps businesses and developers, um-Uh, uh, um, or solo creators transform how they interact with their audience or how they interact with their, uh, their, um, uh, uh, their people, um, through agents and support and sales, uh, with creative tools and marketing and storytelling. So how you can create that interactive tissue between, between an entity and the audience it serves. Um, and that's like a, the common, common piece as a company.

    2. AM

      I, I forget when it was, 'cause these timelines are so blurry now, but there was a moment where I woke up, um, to, like, a, like, six different people texting me a link to a tweet by, I think, um, of, of, of a speech by Javier Milei-

    3. MS

      Mm-hmm

    4. AM

      ... which, which had been completely dubbed with ElevenLabs.

    5. MS

      Yes.

    6. AM

      When was-- Was that a year ago?

    7. MS

      That was, uh, n- yeah, two years. Year and a half ago.

    8. AM

      Two and a half. Okay.

    9. MS

      Yeah, exactly. Yeah.

    10. AM

      I feel like everything changed after. Some- something changed. What, what happened?

    11. MS

      Yeah, so it was, uh... So we-- Uh, to give you a rough timeline year by year of, like, how audio models developed. So 2022, first breakthrough, uh, which, which my, my, uh, my co-founder is, like, the smartest researcher in the space I, I got to know, was finally able to bring into the field of how you, like, get that context, get that, that tonality out. So that was 2022. 2023 is where you started seeing expansion around the wider voiceover, wider narration space. So you could use text-to-speech model across languages, create more voices. Uh, we created the ability for people to recreate their own voice in a high quality, created a marketplace around that, and a wider set of creative tooling to help you with the audiobooks, uh, for authors, uh, creative tooling for authors. Then 2024 finally brought good transcription models combined with the LLM for translation, combined with the speech generation. So you finally had the AI localization version.

    12. AM

      Yes. Right.

    13. MS

      So that was the Javier Milei speech. So the project was, uh, there was, like, few versions of that. He delivered his speech on UN and, and we brought it from Argent- uh, from Argentinian into English so people could, could, could, could listen to that while still having his iconic delivery. And then in that same year, we worked with Lex Fridman on, on the full conversations that he had with different world leaders. So Javier Milei of Argentina, all the way through to President Zelenskyy of Ukraine, um, uh, to later on with Narendra Modi of India, where you could hear both of them speak that language, which was, which was a new opening. So that was 2024. Um, and then 2025, what happened is you could finally, like roughly, finally have those models act in a real-time basis. So you could create more of the interactive voice agent experiences where you could, uh, you know, you mentioned, uh, the kind of the frontier system architecture of combining the stack. The same kind of stack applies in the voice agent, uh, uh, uh, com-co-combo, where you can have this cascaded architecture of having speech-to-text transcription. Then you have LLM to generate responses back and text-to-speech to, to narrate it. So if you're speaking a voice agent, then it can predict, are you stopping the sentence, start generating the response, um, and all of those components can work in tandem to create that, uh, uh, that experience. So that was 2025, and I think 2026, we'll see an extension of that, of how maybe cascaded, I'm going too much into detail, but cascade go to fused or cascaded continual, uh, uh, cutting the latency to be great. But, uh, back to your question, uh, the first great, uh, AI dubbing experiences in, were in, in that static context in 2024 and, um, they were really good.

    14. AM

      Yeah.

    15. MS

      They were really good.

  9. 21:5925:37

    Cascaded vs fused ‘omni’ voice systems: tradeoffs in quality, reliability, and latency

    1. AM

      So thank you for the, the, the cheat sheet on, on how we got here. About a year ago, you and I started talking about, okay, where does the space go next? And I remember you saying, Anj, I think, you know, we've had this cascaded system so far, but now we, we- we're starting to see what people wanna do with it. And a big, a big part of what people want from a capabilities perspective is, is sort of a more, uh, is deeper reasoning from these systems, right? Especially an, an audio agent that can understand the tone, the voice, the, the inflection, the accent of what's coming in, because you lose all that context when you transcribe just pure audio. Um, I remember you saying, um, that there were going to be more and more, uh, either some, some forms of unification of these modalities, um, or s- some new breakthrough where we try to combine these into omni models. Um, where, where are we right now in terms of how far are we from these... Like, the-- When, when I talk to, uh, you know, James Betker is a good example. Mati just talked about James. Uh, shortly after open sourcing Tortoise TTS, James went to OpenAI and worked on ChatGPT Advanced Voice Mode. And when I talk to the ChatGPT Advanced Voice Mode, it still doesn't understand if I'm angry or sad. If, if, if I said the same word or the same sentence in a really angry way or in a sad way, it just transcribes it. And the audio understanding, the semantic understanding of the audio is very limited. Why is that, and what is it going to take to make the systems more capable of understanding what's coming in? Does that make sense?

    2. MS

      Yeah. Makes sense completely. And, and it's, you know, one of the huge questions, and we, we kind of changed our own perspective over the last couple of years. Like, what is the, what's the right, uh, approach across the different, um, uh, uh, domains across other fields? Um, so to, to Anj's point, uh, you know, when you have a cascaded architecture, frequently what would happen in the past, you would transcribe just the text element, pass it over to the LLM to generate, and then, of course, you would narrate that on the other side. Um, now of course, that's a relatively poor version of that. Uh, what, um, what, what you could theoretically do is, is still try to capture a lot more of that information. Um, and now generally, you have those two approaches as you think about the future. One is you continue on the cascaded side. So you continue with the three different models and, and potentially think about how they can be improved-To continue bringing the right context, optimize latency, continue reliability, um, or you can think about training a model together where you kind of combine all of them together. And then you get things in between where maybe you'll try to create a, a transcription and LLM model at the same time, but keep the other side separate. So there's, like, everything in the, in the middle, so it's not a truly binary choice. But that's the, the big question many people in the audio side will, will be, um, alluding to. Like, do you go the fused approach, where you train them all together and generate effectively, when I say something, the voice agent automatically generates a speech talk and it doesn't go through text, or do you go through text? Like, where we are today as we think about our approach to the space, as we think about the enterprise, the business use cases, where the reliability and the smartness and the intelligence of the model is one of the key things, we think the cascaded approach is the right thing for the next, for the next few years. And if you are thinking about places where maybe the reliability isn't as essential, but the speed is, um, we think the fused approach will be. And as in so many of those cases, probably for,

  10. 25:3730:43

    Making agents emotionally aware: labeling data, sentiment signals, and controllable delivery

    1. MS

      for within a given customer, you'll probably want to blend and use a little bit of one model and the other, depending on what you are solving for. Now, to the actual question, we think, so there's, like, kind of those three key parameters. The quality or the emotionality of the speech, uh, and the experience of the interaction. Two is the reliability. Does it interact, not hallucinate, calls the right tools behind the scenes in the right way, and then the latency. We think emotionality is fixable in both approaches. So we hope that later this year already, when you are speaking a voice agent, you're excited, the agent can respond in an excited way. If someone is calling and stressed, we-- the agent can respond in a reassuring way. And we actually just released a new version in, in our voice agent work, which is trying to do that, trying to detect the emotions on the transcription side, pass that over to the LLM, uses that as the context, and then generates response accordingly. And, um, and that was one of the big, like, big breakthroughs on our side to, like, create this expressivity, um, to be able to have this expressive mode finally work. What happened behind the scenes to make this possible and was the hard thing is you didn't have that much data that could tell you, is that delivery happy, sad, stressed? So over last year, we are effectively investing a lot to create this whole labeling exercise to be able to create a data to train that model and control it. So finally, we think you have the expressivity, and the second part is you can actually make it controllable, so you can define what type of experiences should happen based on those emotions. So we think that expressivity or the emotionality quality-

    2. AM

      Mm

    3. MS

      ... is largely fixable in both, uh, or, like, will be very close in par. Our friends at Sesame are doing, like, a great example of how you can-

    4. AM

      Yeah. Bre-

    5. MS

      You can-

    6. AM

      Brendan will be speaking in the class. Yeah.

    7. MS

      Amazing. So he will probably tell you of, like, how they are trying to, like, finally get through, um, that, uh, that, that cusp as well. So it's almost a race, who will be first between us to pass that, like, emotional voice Turing test across other, any of those cir-circumstances. Then the second lev-level is the reliability level, and here, in general, you want the smartest model. You want the smartest model to work in the agent and, and of course, we are seeing so much of the innovation happening across LLMs. Our customers, the businesses, whether it's Deutsche Telekom or Revolut or Klarna, all of them will have a slightly different models they will want to use-

    8. AM

      Right

    9. MS

      ... um, depending on their use case that we can empower them to, to, to use. Um, and then the second thing in that reliability, as you think about voice agent, let's say you are calling in to customer support to rebook your ticket, um, for, for flying to, to Poland. Um, the-- You will want it to authenticate your account. You will want it to pull your own details. If you are to process the payment, you want to make sure it's your, uh, details that are attached. So it needs to be reliable, it needs to follow that flow, and here you are calling so many different tools in the, in the, in the, in, in those steps. So you will, uh, uh, likely go through some two-factor authentication to get the code. Uh, you will likely want your, your email and the database to be pulled up from the information, uh, from that, uh, from that email. So all of that needs to work reliable. It will take time. Um, so how you orchestrate that part is important. Cascaded models work already really well on that, given the intelligence layer will be fixing that. In the fused, you sacrifice that. You are kind of-- You need to then bring all that tooling into the fused model, um, and it, and, and becomes very tricky to track what happened in that, each of the steps. What happened on the transcription step, what happened on the LLM step, what happened on text-to-speech side. And then similarly, you cannot give the guardrails or the safeguards-

    10. AM

      Mm

    11. MS

      ... to bring it. And then latency, the hardest. I think here, fused models are winning, where you can make it very quick. You can make it, like, respond in roughly 300 millisecond response.

    12. AM

      Mm.

    13. MS

      Uh, but you sacrifice on the reliability piece. And what we've seen for, for, for, for, like, our customers, which is the businesses, you don't want that. You want the reliability over, over latents. However, in other use cases, um, let's say like a companion use case that we are not solving, in that equation, maybe that's slightly different.

    14. AM

      Right. I see.

    15. MS

      Uh, and that's, that's how I would think about that stack. Um, and maybe, like, last piece, maybe in the next few years, the way we think this will transform is if you are interacting with, let's, let's take that booking airline, um, example. Maybe if you're just trying to get information about what are the products and services and what, where can I travel, like, where it doesn't have to execute those actions, maybe for that part of the interaction, there is a version of fused approach. But then the moment you need to go into account and authenticate it, you go to cascaded.

    16. AM

      Yeah.

    17. MS

      So on our side, what we are exploring internally is trying the both approaches, um, from research of, of, of, of continually doing, doing that. And the reason I mentioned we swapped, we, like, almost, uh, the good set of, of examples in the field, there's different companies that will explore both of those approaches. So those are very unclear whether the emotionality is fixable with cascaded approaches.

    18. AM

      Right.

    19. MS

      Uh, but, but a lot of the, the recent innovation, whether that'sGreat work from Sesame or other companies prove that it, it is, and, and hopefully the recent release that we did brings it to another, another level.

  11. 30:4335:10

    Ecosystem culture: collaborating with ‘competitors’ and accelerating the frontier together

    1. AM

      So I, I wanna pause for two seconds to kind of, uh, sort of overlay something a- and highlight something Mati may not have realized he did. But did you notice that twice in his last, you know, two, three minutes of exposition, he gave a shout-out to multiple other teams, right? Including Sesame. Sesame, how many people have heard of Sesame? Okay, so about half. You know, you hear-- You'll-- We have Brendan, who's the CEO of Sesame, and he was the former CEO of Oculus. How many people have heard of Oculus? Okay, so almost everybody. Yeah, so, so this is an important sort of cultural leadership point w- that I learned slowly over time, which is there's two types of leaders in the systems world, right? There's those who see what they're doing as one part of a greater sort of collective collaborative project as an industry, right? We're at the frontier of this new space called voice AI and audio AI. Mati's got one point of view on how that system should be developed for the customers he cares about, the mission of Eleven. But at the same time, he's happy to collaborate with other people who have different perspectives, right? And one of those people, teams, was Sesame. And I think it was maybe three years ago, um, shortly after I invested in, in, uh, um, in Eleven, that I decided, you know, Brendan, Ankit, and I... Ankit was my former co-founder and CTO of Ubiquity6. We had been observing at Discord that there was an, this ex- really urgent need to have better and better voice models in Discord. You know, you spend like 60% of your day in, in Discord and, and voice. And so I remember calling Mati up for advice and saying, "Hey, we're thinking about building a new kind of product that has AI, sort of a, a real-time voice companion built into it. What do you think?" Right? And he could've said, "Anj, I don't have time. I've got 1,000 other things going on, not something I think about." But instead, he took the time to really break down for Brendan, for myself, for, uh, you know, for Ankit, like what, what his perspective would be for Sesame's needs. And, and he had the maturity to say, you know, even though he was a-- had to fundraise a bunch of money, and there was a lots of competitors and so on, to say, "You know what? We're all in this together. You're not really a competitor. In fact, maybe at some point we could help each other out." And so I think actually around that time, Brendan ended up in angel investing in Eleven as well. You subsequently became an angel in, in Sesame. There's a level of collaboration between now the two teams that is quite rare, I would say, in the ecosystem. I, I wish more teams had that approach that Mati took and Brendan took. Um, actually, I think about a year ago, based on some of the insights that Mati had given and Peter had given, um, Eleven, uh, sorry, Sesame, they open sourced a speech model called CSM, uh, a conversation speech model, which is just different from any of the models that Eleven was working on. And, uh, you know, that was awesome because that, it, that model is now on GitHub. Many of you can use it for your own projects. In fact, some of the students in last year's project used it. And that, I think, is the through line, one of the through lines I want you to take away from this class too. You know, you had, you heard Anj's life scaling lessons number one last time. I would say a, a scaling law from my experience working with Mati has been, you know, you can go further together, especially in a new space like this, where often what seems like a competitive project just because the VCs or the business ecosystem is trying to create some nice like landscape slide that says, "Here's audio AI, and here are the logos," and then, "And here's like, you know, visual AI." I mean, these, these categories and labels are largely artificial constructs, you know, created by non-technical people to try to make progress legible to the world. But I, I, I've, I-- What I find is really it's people that are driving a lot of this progress, and collaborations between people is what drives fr- the frontier. And so I know I'm hyping you up a little bit too much maybe, but thank you for the-

    2. MS

      No, no, no. This is great.

    3. AM

      [laughs]

    4. MS

      Keep going. It's, uh-

    5. AM

      You take as much of it as you can get.

    6. MS

      No, that's, that's, that's very true. It's, it's very easy to get, uh, stuck in this loop of like, you know, like other startups are your competition, and ultimately it, like, it does not matter for the mission you are solving. It's a long-term game, and, and, and many of those people will come through different intersections of your path in the future. So, so keeping that partnership, uh, or, or working, exchanging ideas together is, is crucial. And if you are to pick competition, if you are to start a company, you probably want it to be one of the hyperscalers or legacy companies. It also, like, will pull you up a little bit. So yeah, I think it's, I think it's-

  12. 35:1039:31

    ElevenLabs business scale: revenue growth, team structure, and execution model

    1. AM

      Well, okay, so on that point, let's talk a little bit for a few minutes about the business, 'cause proof point that you can actually succeed being a nice guy like Mati. So when we first met, I think the company had less than a million or two-

    2. MS

      Yeah

    3. AM

      ... million ARR. Today, is it public? W- what, what is it public about where you guys are today?

    4. MS

      So we crossed 2025 at, uh, 330 million in revenue, and this quarter was our, our biggest quarter, and we added 100 m- over 100 million in additional ARR. Um, so now we are over, over, over 430, um, in, in-

    5. AM

      Over 430 million in revenue in 36 months. Can we just, like, pause for a second and acknowledge that? [clapping]

    6. MS

      Oh.

    7. AM

      This is insane. How big is the team?

    8. MS

      You know, the, uh, it's insane, uh, uh, uh, of course, and, and it's crazy that we get the chance to be building and, and all of you are building or, or developing and learning at like the frontier of the biggest change, uh, in the, in the world, uh, uh, uh, where it's like maybe bigger change than internet or electricity. Uh, the team is, uh, although the... In perspective, of course, Anthropic, as maybe many of you have seen, it's, it's-

    9. AM

      Oh, they were, they were-- They had a little bit of a head start, I would say.

    10. MS

      Yeah. They, they, they added 11 milli- uh, 11 billion in additional ARR over last month.

    11. AM

      Yes. Yes.

    12. MS

      Which is crazy.

    13. AM

      Well, they're, they're in the outer years.

    14. MS

      Which is crazy. Um-

    15. AM

      But we'll get there

    16. MS

      ... but, uh, well, the, the, the common thread is like-That there's such a clear value that you can provide to, to your customers across, in their case, knowledge work, in our case, how you transform how any business can interact with their customers. And, um, we are 400, uh, over 450 people now, so r- relatively one-to-one. Um, we, we have, um, maybe just to give you, like, a quick perspective, we have, uh, a good amount of people between US and Europe. So our biggest bases are in, in London, then New York, and then Warsaw and SF are fighting for the third spot. Um, we are all the time actively hiring, I need to say this. And, um, uh, we-- I think the, the, the big piece that maybe is slightly different in how we've set up a lot of that, that 400, uh, people is we keep it an extremely small team. So each team is less than 10 people, have a, a big, uh, ownership and mandate to run ahead, make their own independent decisions. It's okay to be wrong. The speed in understanding the customer, understanding the problem is much more important than-

    17. AM

      Right

    18. MS

      ... going through this like, uh, uh, hierarchical process. And, and that helped us across that journey so often and so much. That's why we can do across so many different models and then bring them to customers, whether that's in the marketing department or whether that's in the, um, helping on the customer support side, or whether it's growing revenue and helping company figure out how to, uh, reinvent that with new version of sales or fan engagement. So kind of the common thread across all of the product work we do is how you can really iterate-

    19. AM

      Right

    20. MS

      ... in a different way.

    21. AM

      Well, so, uh, again, this is my last question because I, I think, um, you know, the, the final project is the one-person frontier team, o- one-person frontier lab, right? Where the idea is they could now, using many of the tools available to them today that weren't around four years ago, including Eleven, can push the frontier or, or create a state-of-the-art system that, uh, would have previously needed many tens of people. And on the business side, what, what is it that has allowed you with, um, such a small group, relatively speaking, to create a repeatable engine where new capability, you know, results in repeatable revenue? I think we talked about Anthropic last time, where in, in that case, you know, uh, revenue scales quite predictably with compute. In your case, what I've observed is revenue has scaled quite predictably with deployment, you know, as you built out the deployment team. I would love for you to talk about that for a second. What, what's going on there? Why, why is-- Why have you somehow been able to tell people, "I'm gonna do this much in revenue," and just hit it year over year with extraordinary growth, which is quite rare? Um, and two, pricing and... Let's talk about pricing and packaging for a second, because these numbers like ARR, where, again, the, the, the accounting of these things was developed in a pre-AI world for software that looked quite different.

    22. MS

      Mm-hmm.

    23. AM

      Now people are going, should we do token-based pricing, annual subscription? What is the right way to meter intelligence, right? And meter capabilities of the kind. So can you talk for a second about those two things?

  13. 39:3142:32

    Predictable growth via deployment + pricing by value (not cost)

    1. MS

      Hundred percent. Yeah, the, there's, uh... And I, and I like this ambition of building, yeah, effectively your final project and, uh, and it sounds like a fun one too, especially now you can just do so, so, so much if you, if you have the agency and keenness to just, to just go after it. Um, I think the two, on the two questions, so on the predictability side, uh, it's still hard. I, it's like actually very... I think we, you know, we, we are, uh, um, happily trying to set a target. We are currently overshooting our targets. Um, but the main part that is contributing to us to figure out where roughly this will place is, is the deployment side of, uh, we get a chance to work with some of the iconic and the biggest companies in the, in the world. We combine a lot of the forward deployed engineering to work alongside them, which is the team that effectively figure out how to take the AI and, and kind of bring it into the applied AI side, be the lab version of a lot of those companies and transform their work together. And, um, and in those cases, you know, it's, it's, um, we roughly know how much value we can deliver each year, and then that really the bi- main bottleneck is can you bring incredible people that are passionate, that have that, um, that level of IQ and EQ, are striving for excellence but stay humble. Like, how do you bring those people that are really keen to innovate and, and, and, and be incredible at their craft? Um, so if you, if you are able to predict the growth of, of, of the, of the people, of the value of how much you can deliver, then you can start getting back to the, the predictable part of the business. That's more on our sales side, enterprise side, which is more than 50% of the revenue, and then 50% is still, uh, PLG and, and continue growing on the self-serve side. And this one is a little bit harder to predict because it's effectively a big part of can we continue our stride and innovation across the models. And we took the approach, um, of all the product, all the research we launch is thus available to the, to the biggest companies in the world, but we try to make it available for everyone. So if you are building your one, uh, person project in the future, you have access to a lot of the superpower that the biggest companies will. There is some concurrency limits and of course the, the, the, the closer compliance elements, but ultimately you get the same, uh, capacity as the, as the top lab. Um, but what this means for, uh, frequently for revenue, it's that part is less predictable. Um, however, we know where our initiatives are lying and roughly can, can, can estimate, like, what's the value of the world. And happily, frequently, we see the value for the world is bigger than, than we expected. Um, that's the, the first one. On pricing, always think about the value you deliver to the customer and work backwards from there, never from the cost of how much it runs to do something. If you deliver the value, you roughly want to capture one-tenth of what you delivered as a value in your packaging. And any pricing and packaging that accomplishes that, which is the hard piece of, like, how you calculate the value, how you drive that good, uh, metric for that is, is the problem frequently more than anything else. So never

  14. 42:3259:29

    Safety, fraud, and governance: voice cloning, watermarking, and nation-state dynamics

    1. MS

      start from the cost, start from the value and work backwards from there.Thank you for the question. The big question is, uh, voices, uh, y- with a lot of the technological advancement, you can replicate voices relatively easily. What this means for the future of the, of the security and safety in the space. So I think there are two parts. The one, we as a company, uh, given we, we develop our own models, we can actually build and, and bake in a lot of the safety inside of the models, and that's something that we've done from, from, from the start, whether that's being able to, like, trace back the content that's generated back to who generated it and take action as needed. Two, moderate whether you are abusing and doing a fraud or a scam, and stop it before it's generated or, or flag it internally. And then three, contribute to the wider space, which is where we think is headed, is like, can you have a publicly available system where you can, uh, drop an audio sample and get information? Is it AI-generated or not? Have a watermark piece. And that applies as much to the negative version of the use case, but also on the, on the positive side, if people will want to license their voice, how do you make sure that it was licensed in the right way? Which we, you know, we work with people like Sir Michael Caine or Matthew McConaughey, and you want to have that information in there. Two, to your second part of the question, like, a lot of systems today will rely on voice authentication in their banking systems and other parts. We think this is not the future, and you should step away from this and not use that as an authentication side. We think that's, uh, from the security perspective, it's the, it's the, it's the, it's the wrong approach. Um, and every so often we see an interesting way of using voice agents against abusers still. We had this amazing charity we worked with who, based on IP of a caller, could detect whether they can be a likely scammer. And if it was a scammer, so kind of the opposite of the usual, they would serve it a voice agent, and the voice agent would then speak with the scammer, and the whole intention was to waste their time. And some of the most fun conversations happened from that, from that part. So there are some ways-

    2. AM

      Love it

    3. MS

      ... to use that against-

    4. AM

      Counter-offensive security

    5. MS

      ... counter-offensive security.

    6. AM

      Love it. Troll the trollers.

    7. MS

      So the question, biggest bottlenecks for ElevenLabs and maybe for wider audio space. You know, beyond the obvious ones, which is, which is, um, in- incredible people, incredible research, so they can continue innovating on the architecture side. Um, and, and apart from compute, like the main thing from the research we would love to fix this year is how you combine... So we spoke a little bit about that cascaded architecture of how you create the speech-to-text, the text-to-speech, how you can tr- make it truly interactive, um, where, where, where it understands your emotion and understands, can pull up any of the knowledge from your systems and become that personalized extension of yourself. So maybe the non-obvious one, it's, um, you know, every s- every time you interact with different service, you interact with different experience, you will have your own preference of, of what's, um, what's good for you and what's good for everybody else. Uh, and we are trying now to figure out how to make that interaction, like, really custom, really personalized, and really good for all those cases. Uh, and just to recap the obvious ones, so hiring is just the architecture side. Compute, if we had more compute, you can run more expense, more models. There's a version where maybe too much compute is also harmful, like the necessity is the mother of invention is also helpful in some cases.

    8. AM

      The optimal amount of compute.

    9. MS

      The optimal amount of compute. Um, and the preference piece of what really works is, is going to be helpful. And, like, maybe a, a good example of that, in healthcare space, if you are deploying an agent that is taking appointment but then follows up with the person, asks about how they are feeling, um, maybe recommends, uh, uh, additional appointments, you need to transcribe very different nomenclature and make it perfect. Um, there's also different people and how they will communicate. Some people want a, a slower delivery, some people want a faster speed. So how you cater and make that experience unique and different and deliver the most value is, is, uh, is not so much of a bottleneck, but more like thing that we want to solve, and we know we need to more, collect more data, collect what works, and then deploy to the industry. About the com- uh, the difference, um, and, and the other complications, um, on the training side between the cascaded approach, uh, and the fused approach. So the, you know, when you work on a cascaded side, of course you, uh, you get to train the models independently to a large extent. Um, uh, but then you need to figure out how they work in tandem together aft- after that. So you spend, uh, kind of quite a bit of time on, on trying to figure out what will the interaction look like when I combine them together, and when we actually combine them together, you might need to fine-tune and tweak some of those experiences. So we spoke about how you finally got the con- controllability and emotionality of that work. The pipeline for passing off, uh, on like transcribing the sentiment, is it a, a stressed speech, and then passing that as a parameter and making that d- delivery emotional. You need to bake that in as you are training the model before you combine that in the, in the, in the end pipeline. So you need to bake that in in the training step before, before you bring that across, and maybe there is a version of how you do the language control or how you specify the, um, pronunciation in different c- c- setups, which you all want to make sure it's, it's clear. Um, on the fused approach, there's more of the emergent behavior, of course. So, so you get that for free to some extent. But the main thing you need to do in the training side, which is the usual complexity, you need a good open source intelligence, um, model for intelligence, an LLM, that you can bake into that model. Um, and then there are two complexities. One is how you fuse the tokens from the text space-

    10. AM

      Mm

    11. MS

      ... to the audio tokens. Super hard, uh, and most people cannot figure that step out. And then the second part, uh, even if you do figure that out, you do rely on the open, open models that are out there in the field, and today at least, uh, the open source will lag behind the closed source in the intelligence space. The question was around, uh, the future of ElevenLabs and what, what's, what's going, how is it going to look like in five years between the models, between the platform, between the, the, the application side of our work. Um, at the core-Five years is an interesting timeline because there's so much research that we know will need to happen in two, three years, and we think there will be still, still improvements you can make beyond that. But at the core, we know we want to continue leading the space in all foundational research around audio. Um, so whether that's the conversational models in, uh, uh, uh, on, on the future, whether that's how you fuse it with other modalities, uh, is going to be, uh, one of the big pilings that we want to be the best at. So, like, truly be able to pass that voice during test in any conversation, um, and potentially extend that to any interaction. So on the research side, what this means for the tech space already to potentially visual avatar space. We want to be that lead, leading frontier, uh, uh, uh, continuously, specifically figuring out the conversational or the interactive side. So step beyond voice on that side. Um, the platform, it's effectively where we see the, the biggest value we can provide in that five-year timeframe. So as the models start becoming more incremental, you need to really understand in depth what the business, what the, what the, what the creator or the developer is trying to solve and give them the tools to, to do that. And for us, it's going to be the, the effectively the go-to platform. And in some ways we imagine in the future, the same ti- the same way you have three or four clouds that serve all the compute needs on the, on the cloud stack, we imagine there will be likely three to five platforms that's, that help on that conversational setup between any business and their audience, and we want to be one of those, those platforms, whether that's in the, in the, um, in the conversational setup, in the support, in sales, in, in, in, in that, um, in that part of the business or whether that's the marketing and how you engage your customers on, on, on, on, on being able to deliver the better story, all the way through to the internal of how you hire people to scale and train them to then, uh, to then let them, uh, continue get that expertise. So if we can be that for the, for, for the businesses, and in some way everybody will be that maybe one-person business, we, we, we, we, we will do that. And maybe last part on your question because that's definitely something that's ba-back in our mind, it's where does the platform and application start and end? We think this will blur in the future of AI, where you will be able to create those applications on the platform, uh, much easier than you were in the past, and, and we hope to provide all the different modules for the builders of the future to be able to build, to build that seamlessly. One of the proudest work we do at ElevenLabs is actually working with people that, uh, lost their voice, and we can bring it back, so people with ALS or throat cancer. Uh, and, and so far we've been able to work with, uh, almost ten thousand people that lost it, and we could synthesize it back so they could communicate, uh, um, uh, naturally.

    12. SP

      That's cool.

    13. MS

      Uh, which is, uh, which is, we hope, like something that can be available for anyone, anyone in, in need. Um, and similarly on the question around, around Ukraine, so maybe to give you a quick context what happened there. So of course, in that, when the war started, the government need to figure out how to provide a lot of the usual services to people everywhere around the country. And as you can imagine, a lot of the people will just not have the usual access points. They will not be able to go to, uh, to, uh, uh, a local administrative office to get help on, um, on, you know, where can you travel? What benefits can I have because something got displaced? How can I get additional support for, uh, for food or benefits? So it's, like, a lot of different problems that you need to face, and that stretches across the entire economy. In education, the same thing. People cannot go to a school in the usual way. How do you deliver that education to the people that, that need it? How does the government send messages around the country efficiently if you don't have that access in the same way? How does it send it outside of the country so people know what's happening in Ukraine? So the way one of the many initiatives they've done, uh, was creating a central citizen app called Diia, where every person can access a lot of those services going directly through their mobile device and, and, and gets access to information what's happening, uh, some of the educational courses or, or some of the guidance. And, um, and one thing that, that was missing in that equation is how can you open up even further where people, uh, that might not have that technical acumen to, to, to be able to get that? How can people, uh, that might not have the internet call a number and get that information, um, in an easier, in an easier way? Um, so we worked with them effectively on adding that voice side inside of, um, uh, inside of the app and outside of the app, so people can make it a lot easier. So we traveled to, to, to, to Kyiv at the time to, to work and understand those problems and, and, and work with different set of their ministries on, on that work. And, you know, one very clear thing that came out was an, an incredible model of how they were running it. So every ministry had their own technical resources to try to innovate and bring that across, not go through the red tape, go independently, go quickly and bring that across. And so far, we've seen an incredible way of people being able to engage with the government and something we think might be a version of the future of government services. How incredible if everybody here could, could, could, um, could open an app and have access to your, to your passport, to your driving license, all the way through to accessing, uh, the best educational courses, uh, uh, maybe like the Coachella course, uh, directly from the app too. Um, now to your second part of the question, which is important on, like, how we think about approaching our work with, um, with governments in the, in the war zone. Ultimately, as a company, we are choosing to be Western allied and, and, um, Western, uh, um, uh, countries plus their allies and support their work in, in, in a way that's, uh, of course following the, the, the legal, the legal, uh, guidances and something that, that we'll continue, continue to do.

    14. SP

      Can we-

    15. AM

      Can we talk for a second about, since we're in the zone of sort of sovereign scale deployments, um, how are you thinking about China?

    16. MS

      Mm.

    17. AM

      Are, are, are there distillation attacks you're seeing? Um, you know, there's a bunch of news this week about how, uh, a number of nation states have been trying to run distillation attacks on Western companies that produce models like Eleven. Um, is that a concern? How do you reason about the ecosystem in China, which historically has been very collaborative with the Western ecosystem, but there's tensions on various fronts? How, how should they think about that topic?

    18. MS

      Yeah. I, I think in general there's a clear race in AI of, of, you know, uh, the Western world and the, and the developments on the, on the China front. Uh, from our side, we, we try to stop all types of distillation attacks, uh, period. But of course, uh, with, with any, any, any of the kind of IP coming from that region, we, we are taking that as additional strong signal of how we can build that into the, the, the protection. Uh, the truth is there's-- and I'm thinking about some of the other questions that were asked, that, um, that there's a lot of great models coming from that region too.

    19. AM

      Totally.

    20. MS

      And, um, in audio voice, especially as you need to optimize for that language layer too, so you need to, you need to really get the nuance of the dialect, of the accent, of the different voices. Uh, there will be models in the region that will be better than ElevenLabs', uh, uh, models, um, for their, their use case and their remit. Um, uh, and, and to large extent, um, we try to out-compete them on that, on that level and provide a better service to, to the companies on the other side.

    21. AM

      And, you know, as it relates to the open ecosystem, because one of the things as we've talked about is having a good open base model allows you to then customize it for various enterprises, for very specific regions, deployments. What we're starting to see, at least in some of the other modalities coming out of China, like video models, a year ago, lots and lots of innovation, lots of open video models. Now not so much. As the-- It seems like as several labs start to catch up to the frontier, you know, Seed Dance is a great example. It was a SOTA video model that came out. It's not open source. Uh, or it's not open weight. Is that a trend that you think is going to continue? What does that imply for a team like Eleven? And what do you wish more participants in the ecosystem, like the students here, other labs in the space, would be thinking more about or doing that would collectively help?

    22. MS

      Yeah. Great question. I, I think, uh, so there are two parts. Uh, like as we look at the models in, in that space, frequently we take slightly different approach in general because there's the, the research element, we spoke about the product element, and there is the ecosystem of how you become a trusted brand that people can, can work with. Um, we briefly skimmed across one of those examples of like how you can create and give ability for people to create a voice, share that voice, and earn possibly on that voice, and figure out how people can participate in that model innovation too. And that's like a stark different approach between like some of the work we do on that front versus some of the, the, the, the wider models that, um, uh, uh, players in China will take.

    23. AM

      Right.

    24. MS

      And the same will apply to video, of course. You've probably seen the, the work on like how people approach the IP from Disney or Netflix on that side, and how they approach it on, on the, on the Western side. So I think that's got the big dichotomy, and I wish they didn't.

    25. AM

      Right.

    26. MS

      And I think, um, on the security level question, it's going to be a, like almost a big combination of the work that everybody needs to do, where you, uh, you, you can contribute your work, but then you need to figure out the watermarking system, and you need to figure out how to have the, the, the models coming from China follow that parting.

    27. AM

      Right.

    28. MS

      Or at least if not, then you don't serve the content or don't give the same content, um, uh, permissioning on some of the platforms. So there's a, a, a, I think there are two different parts. One is how you think about the wider ecosystem participating in a model. Two, how you bring, um, uh, bring the safety parameters around-

    29. AM

      Ah, I see

    30. MS

      ... around, a-around, around the models in, in tandem. And three, uh, in general, of course, very helpful that open source ecosystem continues, and I think you're doing phenomenal work to help nurture that across the, the companies out there, where, um, where ultimately you will have always a incredible set of builders that need the access to the weights, whether it's to fuse the models, whether it's to fine-tune the models for different cases. Um, and I hope our open source, our meaning Western open source models, are at least at par or better-

  15. 59:291:06:15

    Creative adoption + on-device models: controllability, economics, and the next platform layer

    1. MS

      ... from Australian Blackfirst Labs that will speak about this. Effectively, why do studios not adopt, um, still, uh, like why they're hesitant of, of switching to a lot of the AI voiceovers, and is it, is it because of the fear of AI slop or the backlash itself? In general, as we think about a lot of the creative tooling that we provide for, for studios or creators, uh, we heavily believe in this concept where, um, you want to use the tools effectively middle to middle rather than end to end, meaning you will have the story you want to tell, then you will use, use the tool to create a narration in this case. Then you'll want to refine it, then maybe recreate, and then you can get a great output at the end. So there's this, this, um, big iterative step that needs to happen, and why we think about it as like this kind of mid- middle to middle AI part that it will replace. Uh, as I think about AI slop version of that, I think about the kind of the end-to-end version, like can I type a prompt and get a voiceover or video or any, any, any of that work? And that doesn't carry-- it doesn't have either the initial input of the story or doesn't have this great iterat- i-iterative spirit o-of that. And I think studios realize that. So, so I think studios realize that, but at the same time, you'll have very different set of studios. And I think, you know, the, the high end of the version of Hollywood version is just getting there, where the quality became good enough where you could go through those iterations and finally get what you intended. So to make it m-more specific, um-Until very recently, you were, in the speech side at least, you would give a model a text, you would rely on the model to read out the text in the way the model, uh, thought is best, and you could regenerate it or understand the context. But ultimately, it would be down to the model to decide how to narrate it. Until six months ago, roughly, we finally figured out how to control it, like the director would do in those studios, where you tell it, "Redeliver this in a slightly more dramatic way while slowing down a little bit." That was a big br-breakthrough and bottleneck for that to happen. And, and in the last six months, as that started happening, we started seeing more studios finally adopt that technology in their work to, to bring that. Of course, too, there is definitely understanding from the studios that if you are to bring that technology, you need to figure out how the wider economic model will, will work. Like, you want to respect the IP of the people you work with, um, and the economics of that haven't been figured out. Like, how much do you charge for AI voiceover of a person that would not otherwise go to the studio? It's, um... So I think studios, from our conversations and, and ourselves, are trying to figure out what's the right balance in those economics. Um, but frequently what we see, like, actually happening is you will have AI re-replace parts of the work that, um, that, um, that you wouldn't want to do in the first place. So frequently in a studio you'll have scratch work that you read and want to listen to. Post-production's almost where it started, where you co-co-compare the lines, and then ultimately the actual delivery you want the, the additional art coming, coming, coming out in the, in the top end. Um, and then of course there will be, like, other things that, that they will start with, which is AI localization or interactive experience of how you can bring experience and, and, and, and, and interact with the, with the movie. So that's where we see, like, most of the, the use case. And as we think about the future, the, the two pieces that will be stopping it is figuring out the economics model that's backlash-related if you don't figure it out in the right way. And then two, um, those step changes that weren't possible to make, that really high-quality versions are just becoming possible. Uh, but I think th-there still needs, more needs to happen as we think are like-

    2. AM

      Mm

    3. MS

      ... high enough that con- of that content. So will the models be on device and, and, and how we think about ElevenLabs' future in that context. By the way, when does this, uh, when do we pu-put it on, uh, stream?

    4. AM

      Uh, I think tomorrow.

    5. MS

      Okay, tomorrow. So quick. Um, I guess I can still say it, but we finally figured out how to bring our models on device. So we will, we will, um... So we, we, we found a way to constrain it to a given language and, and potentially bring it to any device out there.

    6. AM

      That's cool.

    7. MS

      Uh, which we'll be selectively working with, with bringing that and opening it up to, to wider side of the audiences, which, which is the big innovation there. Um, but it's still the quality... There, there's the, the, definitely a quality difference between the on-device version and of course on-cloud version and the wider set of things you wanna do. So the on-device version will do text-to-speech, but you still won't have the wider transcription interactivity, how you transfer the emotions from one side to the other, um, how you make it, uh, uh, uh, with additional kind of reliability p- elements bu-built, built in. So there is a gap definitely that will, will, will exist between those for, for a l-long while, and you still haven't fixed them on that side. Um, and maybe as a quick side note here, our approach in general as we think about on-device was we need to fix quality first. We want to make it as good as possible. Only when we do that we'll consider bringing that on-device or on-prem. So, like, instead of trying to go on-device with lower quality, we want to deliver the best experience to everybody there in that context. And I think that's now starting to happen, but still not happening in everywhere. We think it happens on the text-to-speech narration, less so on the interaction. Um, the role of ElevenLabs in that future will continue being the, the platform of how you really go deeper in the problem that the customer is facing or the enterprise is facing, and delivering not only the models, but delivering the wider tooling that's required for you to bring that technology in the field. So how you bring that applied AI to your local... Sorry, to your, to your com-

    8. AM

      Customer.

    9. MS

      Exactly.

    10. AM

      Your context.

    11. MS

      To your custom context work. So, and, you know, if you are deploying that in a, um, in a, in a, any, any interactive support or sales context, you will need the entire knowledge base of how you want the, the interaction to work with your customer. You want the piping of how it calls the phone number, or it goes through chat or WhatsApp or email. Um, you will want additional tool calling to get and pull data from the database, um, from Salesforce or ServiceNow to, to deliver that experience. And then the whole framework of how you evaluate, monitor, test that is essential. You know, it's like, is it working? Can I self-improve based on that? Um, what is the domain tests that I bring? So, uh, we spoke about the, the, the, the route scheduling for airlines. How do I know that there is additional sp- seats available in that route? You need the, the, the, the test for always checking for that parameter being true, and if you left customer happy. So in the future where models become more available, we know there is still a huge gap of what the models can do and what you need to bring that in that, um-

    12. AM

      To manage it

    13. MS

      ... enterprise context or even creator context for the question earlier, earlier on.

    14. AM

      We are out of time. But thank you, Mati.

Episode duration: 1:06:25

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode vfF011ko89o

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.