Skip to content
a16za16z

How Real-Time AI Video Is Changing How Creators Work

a16z General Partner Jennifer Li sits down with fal co-founder Gorkem Yurtseven and Head of Engineering Batuhan Taskaya to discuss what changes when generative video becomes fast enough to run in real time. They unpack the technical work behind H3 Max, fal’s post-trained version of MiniMax’s open-weight video model, and how combining model post-training with systems and hardware optimization significantly reduced generation time while maintaining quality. That speed has enabled experiments with continuous video, including streams that can remember previous scenes and respond to new directions while they’re running. They also discuss why the next challenge may be less about speed and more about control, from camera movement and lighting to characters, motion, and lip sync. And they explore what those capabilities could mean for professional creative workflows, where artists and studios need predictable tools rather than simply generating a video from a prompt. Timestamps: 00:00 - Intro 00:47 - Meet fal & the H3 Max Launch 02:45 - From Inference Platform to Post-Training: Why fal Made This Bet 06:33 - The Secret Sauce: System Work, Architecture & Cost/Latency Wins 10:44 - The Twitch Moment: Real-Time Video the Day After Launch 13:11 - How the Model Actually Remembers What Happened in a Scene 15:47 - The Economics: Serving Costs, Chip Footprint & Streaming Experiences 22:02 - Beyond Consumer: Unlocking Hollywood-Grade Controllability 28:37 - Prompting a Director: Camera Angles & Scene Understanding 35:43 - Where Video Models Go Next & the Creator Economy Shift Resources: Follow Gorkem Yurtseven on X: https://x.com/gorkem Follow Batuhan Taskaya on X: https://x.com/isidentical Learn more about fal: https://fal.ai Follow Jennifer Li on X: https://x.com/JenniferHli Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Gorkem YurtsevenguestJennifer Lihost
Sep 17, 202638mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:000:47

    Intro

    1. GY

      Generative media is along with coding agent market, what we call is token market fit. Everyone's waiting for a large consumer moment in AI. I believe H3 Max makes it possible

    2. JL

      Were you surprised by the speed up and the gain you could get from post-training this model?

    3. SP

      We have a version called H3 Max Turbo that's public, that can generate, like, a five-second video in, like, 1.5 seconds. From a cost standpoint, it's also, like, 2X level.

    4. GY

      And all of a sudden it unlocked a whole new workflow for Hollywood and professional people.

    5. SP

      We have been very, very focused towards speed, performance, quality, and now we have a really good base model. The next month or two is gonna be fully focused on-

    6. JL

      Welcome, Gorkem, Batuhan, to our podcast again. Uh, we did the last one last year. This is long overdue,

  2. 0:472:45

    Meet fal & the H3 Max Launch

    1. JL

      and we have such an exciting model to talk about, which is fal's, um, H3 Max. Um, the day when it came out, I was calling it, it's really in the league of its own. Like, it's so funny to see the benchmarks where you have, like, you know, the dot of this model on the far left or far right, and then everything else is, like, on the other half.

    2. GY

      And, and that graph is actually log scale.

    3. JL

      [laughs]

    4. GY

      So it's actually further, but we had to, to, to fit it in, we had to do log scale. Yeah.

    5. JL

      That is hilarious. Um, and, and it, you know, the internet-

    6. GY

      The time, the time portion, the quality is not, yeah.

    7. JL

      For sure.

    8. GY

      Yeah.

    9. JL

      The internet noticed, for sure. There were so many viral tweets about it. Like, people really played around with this model. Maybe just give us the backstory of what inspired you to, uh, post-train this open weight model from MiniMax, and how did you get the quality and speed to-

    10. GY

      Yeah

    11. JL

      ... where it is.

    12. GY

      F- first of all, the MiniMax H3 model is, is the first truly open source, very capable, like, latest generation video model out there. So even though we work with some of the other model labs to run inference for them, we never had the, had, had this, this capability, like, had the right to add this capability on top of it. So when MiniMax came up with their open, very capable open source model that is, is truly last generation, can take references, like, very familiar architecture to, to any other video model. We thought this is a great opportunity to go all in and, and see what, like, we can do. And again, we did, like, many different things that we are gonna talk about, like, that, that combined gave the results that, that you show on the graphs. But yeah, the, the biggest reason why everything came together for this particular moment was because H3 was the first truly next generation video model that's open source.

    13. JL

      What is the idea, like, given like, you know, fal has been known to be, like, a generative media inference serving platform.

  3. 2:456:33

    From Inference Platform to Post-Training: Why fal Made This Bet

    1. GY

      Mm-hmm.

    2. JL

      Like, what is the idea to get into post-training open weight model? Like, you know, you, you talked quite, quite a bit about, uh, about it in the, in the blog of, like, combining the system work with, you know, the, the model, uh, i- itself. Like, maybe talk more about, um, the work behind that.

    3. GY

      Like, generative media is, is, I would say, along with the coding agent market, uh, what we call is, is token market fit. And the way we define it is as can a single person productively spend a lot of tokens and, and the, the amount is, like, 10K a month, something like that. So there is incredible amount of demand in the market to, to generate video, to, to generate many things at the same time, and a person who is doing this for their daily job, they spend in front of a computer and do this all day long, and they spend thousands of dollars, lots of tokens. Um, and since around April, the whole industry and fal itself, we've been compute constraint. We, we are growing as much as we are adding compute. Like, there are things we do here and there, but the whole industry has been compute constraint, and we've, we've always been looking for efficiencies where we can relieve that a little bit so people can, can use this more. So that has been the, the idea behind everything we've been doing since April, and this just came at the right time because this makes everything maybe an order of magnitude more efficient, so it gives more compute for, for other, other models or even, like, more tokens can be generated using, using H3 Max. I think Batuhan would, would agree on, on that. Like-

    4. SP

      Yeah. Like, just from a system-wide optimizations, which is what we have been doing for the past three, four years, you can maybe make the model 2X, 3X faster while producing the same quality, right? Because it's at the end of the day, same model, same architecture, you have the same constraints. You're just trying to optimize what, uh, what, what you can get out of and what you can get out of the chip itself. And there, there is a roof line there. And like, you know, we, we have been approaching that roof line more and more, uh, especially like, you know, lately because our entire team has been focusing on how do we get out more video pixels out from a single chip, uh, as much as possible. And this, this new set of like post-training related optimizations with like, you know, system/model co-design enables us to go beyond that roof line by an order of magnitude. And like we, we just like felt the pressure. We have been working on it on top of open source image models before the video models. We did one version with Ideogram, we did one version with Flux. So we have been, we have been like experimenting with how do, how can we build post-training infrastructure to take an existing model, build kernels and systems design around it to run it very, very fast for a specialized version that can beat anything else that we would get just by running the model itself. And, you know, combination of that plus, you know, just like getting a frontier video model on our hands and all this expertise, we were able to, you know, go by like an order of magnitude in terms of speed.

    5. JL

      Incredible. Let's, let's dig into that. Um, I may get some of number, numbers wrong, but, um, uh, the-

    6. SP

      There's like efficiency numbers, there's cost numbers, there's speedup numbers.

    7. JL

      Right.

    8. SP

      Like not, not everything means efficiency, but it all adds up to be very efficient. Yeah.

    9. JL

      Yeah. I guess what, what is stunning to me is like there is like magnitude lower cost and also like much faster, I think it was like 35X speed up, right?

    10. SP

      Yeah, compared to the original MiniMax H3 endpoint. Yeah.

    11. JL

      Compared to the original. Well, at ELO score, you didn't really sacrifice quality. So yeah, just [laughs] maybe

  4. 6:3310:44

    The Secret Sauce: System Work, Architecture & Cost/Latency Wins

    1. JL

      reveal a bit more of the, the, the secret sauce behind of like is this more of like, um, the, the, the type of system work you have done that like... Did you have to do like model architecture change?

    2. SP

      Mm-hmm.

    3. JL

      Like is it some system work that really brought down the, the, um, uh, the cost and, and latency? And like how about like the, the next generation of chips, like, uh, GB200 fits into the, the, the whole story?

    4. SP

      Yeah. It, it, it's just a compounding effect of like multiple, multiple different optimization variables that we have been targeting. The first one is obviously, okay, you go from, you go from like the base model to a model that's like post-trained to be like more efficient. Uh, for, for diffusion models, this is just essentially how do you go from like running 50 steps to running something like 20 steps, right? Like you're, you're just like trying to optimize that pipeline. But as, as soon as you go from 50 steps to 20 steps, you lose quality.

    5. JL

      Right.

    6. SP

      So you, you need to target in the optimization scene, okay, I wanna improve the quality, and then I wanna apply the optimization. So we have like checkpoints of this that are significantly higher quality, but obviously slower. So what we initially did was, okay, let's run our post-training and RL pipelines so that we can improve the model's quality, and then apply the optimization stack on top of it so that the end result gets you to the same quality or like even like higher quality than the original model. But at the same time, you're like an order of magnitude faster. So most of the gains come from like, you know, post-training this model, uh, to be like, you know, compatible, that you can run this on like less amount of steps. But on top of that, you add like all the kernels and systems engineering work that you do that brings your like hardware utilization from like 30, 40%, which is like standard in like many inference workloads, to like 70, 80%. Like you're essentially trying to... And 70, 80% on theoretical MFU, which is like impossible to reach, so you're essentially at the roof line of what you can get out. And then these models are not just like a single, oh, you just give a prompt and you get a video back.

    7. JL

      Right.

    8. SP

      There's like, there are actually pipelines underneath. You need to get, like take a prompt. You need to, you know, like use a, run a LLM, like a very large LLM, to go expand that prompt to a format that the model was initially trained at, uh, run the vi- generate the video in the latent space, and then decode those latents in, back into pixels.

    9. JL

      Right.

    10. SP

      And then depending on the workload, there might be an upscaling component involved, so there's like multiple components. And every single component by default is unoptimized. There is still like lots to be gained there, and like we just like looked at it from a perspective of we are gonna get the maximum out of every single component. This made us go run LLMs at super high speeds, right? Like it, there, there, there is that component. But for a different workload, this is not like something like an agent decoding LLM workload where you have very high cache rates, where you have higher sessions. It's a single shot. You ta- you give a prompt, you get a prompt back, and there's no caching. You're operating at low batch sizes. So there's a completely different set of optimizations on, on the prompt expansion side, completely different set of optimizations on the diffusion model, completely different set of optimizations on the VAE that you take from latents to pixels, and you just combined all of these to, to have an effect that compounds. From a hardware standpoint, going from something like Hoppers to Blackwells, you see something like two to 3X improvement by, by itself. But from, from, from a cost standpoint, it's like pretty comparable because the cost is also like in, in that league. So, mm, it-- I would say it only reduces your like wall clock time, but like not just the efficiency itself. But it, it obviously helps if you wanna go significantly beyond real time. Like if you wanna generate, you know, five seconds of video in less than, you know, three, two seconds, then you need like some of these like latest generation hardware, uh, today, uh, to, to unlock that possibility.

    11. JL

      Maybe this is a detailed question, like is, is the model being served on single GPU or like it's just like a-

    12. SP

      It is ru- Like majority of the video models today run on a single node configuration, which is eight GPUs, because once you start scaling beyond eight GPUs, the efficiency gets less and less because of the communication overhead. Uh, and com- existing, both like the existing MiniMax H3 endpoints, as well as like other video models are probably getting served at like, you know, single node configuration. Same with this. It's like running on pa- in parallel across eight GPUs.

    13. JL

      And do you think there will be more efficiency gains in there that you can either, you know, optimize more of the steps-

  5. 10:4413:11

    The Twitch Moment: Real-Time Video the Day After Launch

    1. SP

      Mm-hmm

    2. JL

      ... in between by sacrificing maybe some of the, um, like narrow down the user experiences, let's say like the different type of inputs and outputs? Or like as you're thinking of parallelism, like i- is there more juice to squeeze? Maybe that's the, that's the question too.

    3. SP

      Yeah. We, we, we released a turbo version of H3 Max. So initial idea was calling this H3 Turbo, and like we were like, "We don't wanna call this Turbo because the quality is like better than the original one," right? Like that's-

    4. JL

      [laughs]

    5. SP

      ... like this, this needs to signify how good of an achievement it is. So we released H3 Max, but like a week later, we had like, you know, we, like our team was like, "We can run this two X faster at like 97th percentile of quality." Like we run evals, they're like almost the same, right? Like there's still like, there's like a notice-- like there's a small noticeable, uh, loss in quality. But we, we have a version called H3 Max Turbo that's public that can generate like a five-second video in like 1.5 seconds, which is like insane. And that's also like two X the, two X the co- like the, from a cost standpoint, it's also like two X less. So there, like depends on like how, how okay you are with like losing quality, you know, you can go down. And like today, these models are so cheap and so fast that I don't think people need any faster or like any cheaper. Like it's already like at, at a point where, uh, the, the, from a cost standpoint, compared to the frontier itself- It's an order of magnitude cheaper compared from like a speed perspective. It's more than an order of magnitude faster and like you just like enable all the experiences. I think we would need to see, like, I think we would need to see what else levers that people would need, but I-- my bet today is we just need to improve quality more than like the speed at these speeds, right? Like let's fix the speed and let's try push for quality and controllability of these models, which is like, you know, what we have been pushing in the past two, three weeks.

    6. GY

      I think controllability is, is key. Like when we first did it, we did, um, text to video and then image to video, and then references came later, which, which adds a ton of controllability, and it's basically the default mode how people use these models these days, references. And then we are now adding different, uh, LoRAs fine tunes of the, of the base model as well. We are working on like a lip syncing version. We are working on, uh, a different, different camera angle LoRA, different style LoRAs. So a-again, open source adds a whole ecosystem around the model, and it really, really helps.

    7. JL

      Were you surprised by the speed up and the gain you could get from post-training this model? Like, like,

  6. 13:1115:47

    How the Model Actually Remembers What Happened in a Scene

    1. JL

      I, I, I saw it as a little bit of a surprise that like one day, I think it was a Saturday, you launched the model, and the Sunday people put it on Twitch, it become a real time model.

    2. GY

      Yeah. [chuckles]

    3. JL

      Like-

    4. GY

      I mean-

    5. JL

      ... that's the interesting part of like why you released-

    6. GY

      We, we, we, we did evals. Like we spent a ton of money doing evals on our own, I don't know, like tens of thousands of dollars even. And like the results were unbelievable. And, and then like the plan was to just release the model without doing external evals and then, okay, we decided let's, let's hold off. Let, let's not, let's not tell people that this is like so much faster and so much better before we have some external validation. So, so we waited like three, four days to, uh, all these other like eval platforms to actually run the eval. So we, we matched the results that we have externally as well, and that's, that's how we launched it because as you said, the results were a little too good to be true.

    7. JL

      Yeah.

    8. GY

      And, and it was. [chuckles]

    9. JL

      [laughs] I guess were you taken by surprise the, the real time, uh, use case that came out of it? Or, or what are some examples that you think this model has like unlocked of the experiences-

    10. GY

      Like this-

    11. JL

      ... the prior models couldn't?

    12. GY

      This, this happens at fal once in every couple of months where like the whole company gets, gets hold of something and, and the c-creativity just explodes and everyone is just working on a, a, a new little app or, or a, a, a different optimization LoRA, whatever it might. Like the whole company gathered around this, this model and like some, some front end engineers started working on like interesting applications. Uh, we can talk about our like world model accelerator team, which is, which is brand new. They, they started working on the, the live, uh-

    13. JL

      Yeah. The-

    14. GY

      ... experience

    15. JL

      ... RTC, uh-

    16. GY

      The web RTC live experience. So like there were five, six different parallel little projects within the company and like I, I, I think we broke a record on Slack that day how many messages were, were sent-

    17. JL

      [laughs]

    18. GY

      ... uh, in the company because, like, and, and, and like we have a distributed team. We have, we have people all around the world, like m-mostly in San Francisco. But, uh, it's, it's like incredible when you see like the 24-hour development, like when people like w-work 16, 17 hours and then someone else wakes up and picks up that like, and that went on for like three, four days and that's, that's when we released all these projects.

    19. JL

      Take me into that. It's so interesting-

    20. GY

      [chuckles]

    21. JL

      ... 'cause, um, like you imagine like a model or product launch being like planned out,

  7. 15:4722:02

    The Economics: Serving Costs, Chip Footprint & Streaming Experiences

    1. JL

      like having, again, like having all these like eval vendors being ready, lined up and, you know, ship something out and then like you let the, the world, um, or the external users take it and then experiment and like build experiences, put online. Yeah, it, it seems like, you know, people internally who are very creative just like took this, dropped everything they were doing, like launched a experience that got really popular on Twitter. Do you wanna tell us about that one?

    2. GY

      Yeah, of course. O-one of our engineers, Rehan, um, just, just by himself completely, um, started streaming a live stream of continuous generations of H3 Max from his laptop. Like he was just-

    3. JL

      His computer.

    4. GY

      His computer, exactly. He, he was doing some like prompt tricks, trying to keep like a coherent story, uh, and then like he started live streaming that on, on Twitch. In parallel, Levelsio, a famous Twitter influencer at this point-

    5. JL

      Yep

    6. GY

      ... had a, had a similar idea and he, he reached out to us that he has a website ready already. He wants to like host the streaming himself and, and have, have a website that does infinite streaming. Internally also, we had another team who was working on a continuous version of H3 Max. So H3 Max is like the, the Rehan's version and, and Levelsio version were independent clips. It's still very fast, but, uh, the clip starts, it ends, and then you take the last frame of the, of the clip, try to put it in the next one, and try to create a continuous-

    7. JL

      You need to put some work into like-

    8. GY

      Yes

    9. JL

      ... the last frame prompt gen.

    10. GY

      And there, there's no memory. Like the, the second clip doesn't really remember anything from the first clip other than, uh, the last frame. But internally we w- the, the ML team was working on a version where it's the, the transition is more seamless. There is like two minutes of memory. So like y-you're in a scene and when, when you direct the, the model or someone else enters the room, it actually like everyone looks at that person entering and the s- the scene is continuous. So internally we were working on that, and then another team was working on an experience we, we called fal Live for the continuous version. So we had three parallel efforts going on, uh, that were all independently going viral on Twitter by the way.

    11. JL

      And these were all like spontaneous. Like you didn't like-

    12. GY

      We didn't plan for it at all

    13. JL

      ... plan for any of them.

    14. GY

      Yes. Exactly. Yeah.

    15. JL

      And they just became products and experiences. In the next-- in the following days that-

    16. GY

      Like, B-Batuhan, let's talk about the, how, how we made the model more continuous.

    17. SP

      Mm.

    18. GY

      That was very surprising to me because I've, I've never seen, uh, that actually work on, on a video model before.

    19. SP

      So yeah. We, so going back, we have been like very, very focused towards role models and essentially like action controlled or like, you know, action-driven, real-time continuous streams of video. And the problem till like, you know, something like H3 Max was quality was not good enough at all. It was just like, you know, you, it, it degraded a lot. It didn't remember the past before. But we built the infrastructure, we built the infrastructure that we can like go stream video, have people control it in real time, being able to like multiplex it to multiple people, very low latency. And at the same time, our ML team was essentially trying to take every single video model and try to apply this set of like optimizations and tricks to, okay, how can we make this generate instead of a five-second video, 15-second video, 30-second video? Uh, but you were always like below the real-time factor, uh, where, you know, you, you, you were always like, you know, you, like you, you never could generate like five seconds under five seconds. Once H3 Max unlocked it, the ML team was like, "This is insane." Which are like separate teams internally. We have a research team, we have an inference team, we have an ML team. They're like, they, they saw this and like, "This is insane." We can apply all these like set of learnings that we had in previous models where we attempted to do this, where instead of ge- trying to generate a five-second chunk, let's try to generate, you know, like a 50-- like 10-second video. Uh, and then the five seconds from previous one is still attended. We still remember it. And like as the video goes up, we can like extend that memory up to two minutes. And you need to do extremely clever optimizations because attending to a two-minute video-

    20. JL

      Mm-hmm

    21. SP

      ... is just extremely, extremely compute intensive and just like it goes up exponentially be-- uh, uh, from, from like a compute standpoint. So like we, we, we did like lots of optimizations there, but at the end of the state we were able to, okay, we can remember back to two minutes, which is like generally good enough from a memory perspective. And then obviously with like prompt tricks, you can still like continuously, uh, remember more finer, uh, like d- grain details above the two-minute mark. And you can essentially stream infinitely. We capped it at an hour, uh, from that perspective, and then that team just like released that model under H3 Max Director, which is public for people to use. And I think it's the only model that can generate like, you know, up to 60 minutes continuous videos that is action controlled. You can like, you know, start with a prompt, say, uh, like there's like an office setting and someone is like, you know, working and then like 30 seconds later it just imagines by itself. 30 seconds later you can like say, "A woman walks in through the door." Like it, it can take the prompt and reflect it immediately, which is the most fun part.

    22. GY

      And the office is still the same office.

    23. SP

      Same office.

    24. GY

      The camera can pan back to the original person, and the original person is still there in the same state. Yeah.

    25. SP

      So, you know-

    26. JL

      Yeah

    27. SP

      ... we, we, we, we released that and it, it got like we, we, we did this like fal live website to just like demonstrate it because it's like people need to see how cool this is, right? This is a new technology. I, I don't think people are like really aware and it, it got also like very viral very immediately because we also let people vote on what the next section is. It was like, you know, uh, like a form of like a-

    28. GY

      Crowdsource. Yeah [laughs]

    29. SP

      ... crowdsource to like the chat was controlling the what-

    30. JL

      [laughs]

  8. 22:0228:37

    Beyond Consumer: Unlocking Hollywood-Grade Controllability

    1. JL

      to me. It's like I, I, I found it interesting in the, in the Gen Media market that you-- it, it's not like, you know, like language model, you have like this linear graph-

    2. SP

      Mm-hmm

    3. JL

      ... of like just continuously up-- compounding on like, you know, intelligence capability and so on. Like feels like in the field you're operating in, it's always like a few month of like sort of quiet time, but like a lot of things are bubbling, but like in a very short period of time, like everything bursts. Like all these things come in com-in combination come together of like the, the, um, uh, base model being good enough. Like you can get the latency down to the point where you can like-

    4. GY

      References. Yeah.

    5. JL

      Yeah.

    6. GY

      Yeah.

    7. JL

      Um, like get the real-time experience, but also like apply controllability on top of that real-time experience. Like this just opens so many, you know, opportunities of like live experiences where like end user can control what's happening on the screen, which is incredible. Like, uh, we have imagined a lot of these experiences, but never been able to like really pr- play around with it. Maybe just like tell us more about what you're seeing from the market of like, how are people using like, um, uh, the director, uh-

    8. SP

      Mm

    9. JL

      ... capability? Like h- what are you, uh, seeing creators are creating that you haven't seen before? Uh, and what do you think that unlocks as far as, you know, what people can do with the, the, this, this medium?

    10. GY

      Yeah. It's been like almost three weeks since we, uh, released H3 Max, and already it is the most popular video model on the platform, on, on the fal platform by like double almost, like-

    11. JL

      Wow

    12. GY

      ... little more than double, um, in terms of like volume. Um, so in, in a lot of other platforms, it's also becoming, uh, the default model that people interact with because it's so fast, so cheap, it just makes sense. If, if you, if you come to a platform, this is the experience that, th- that you wanna see. Um, so in terms of like popularity and volume, it's, it's taking over, at least from our, uh, vantage point. And for, for Max Director, again, there has been, I don't know, tens of different versions of these live streams. Some of them are, uh, still going on and like becoming more and more popular. We are trying to work with-

    13. SP

      Some, like AI IP holders, people who have like AI shows on Instagram and TikTok and, and do train a LoRA on, on their style and do a live version of their show. So, uh, we have couple lined up already, so that's gonna be very exciting. Um, and like the way people-- like if you talk to, uh, a creative technologist, prompting with voice has already become, uh, something that, like they use all the time is like using Whisper flow or the ChatGPT voice mode.

    14. JL

      Yep.

    15. SP

      And, and now like you can keep talking to the model and it, it like almost as if it's a real director in a real movie set, uh, directing like the camera, directing people where to go. You, you can do that. And like our creative engineers started using these models like that. So we'll see, like a lot of interesting experiences are built as we speak.

    16. JL

      Very interesting. As in like the video is playing, um-

    17. SP

      The video is playing and you are-

    18. JL

      ... on the screen and then-

    19. SP

      You are, you are like talking to the video and, uh, what, what's, what's being displayed changes accordingly. Yeah.

    20. JL

      That's incredible. Um, and talking about like how the, the memory piece holds now, like are-- again, this may be a technical detail, like the capability of remembering what happened in the last-

    21. SP

      Mm-hmm

    22. JL

      ... scene or in the last couple minutes of scene. Like are you remembering that through like the, the, the frames, the images, or is it like through text-

    23. SP

      It, it essentially-

    24. JL

      ... condensed, um-

    25. SP

      No, it's, it's essentially the, like it remembers the raw video, obviously very, very compressed because you can't attend to the full video, but it essentially knows like most of the, like the details happened in the past two minutes from its own generations and above the two-minute mark, it has like, think of it as like a, it has this evolving system prompt on top of the two-minute mark from two to 60 minutes where it knows like the overall structure, overall detail. So it, it remembers like the last few scenes. If you think a scene is like 15, 30 seconds, then it remembers like the last four to eight scenes. And then on top of that, there, there's like a continuously evolving, gradually evolving system prompt that like keeps remembering the, like the overall, uh, coherence of the, of the world. Like e- everyone's waiting for a, a large consumer moment in AI. Now it's like good enough and, and cheap enough that like a truly novel social AI experience can be built on top of it.

    26. JL

      Maybe let's talk more about the e-economic side of this. Like what is the, um, uh, I guess one, just like talking about serving cost for-

    27. SP

      Yeah

    28. JL

      ... like same, uh, minutes of video, uh, with H-H3 Max, um, and how has it changed your thinking around like your footprint of like inventory of chips? Like how do you wanna have, um, like different, uh, steps of-

    29. SP

      Yeah

    30. JL

      ... experiences serving to the end user?

  9. 28:3735:43

    Prompting a Director: Camera Angles & Scene Understanding

    1. JL

      Um, and it seems like Batuhan is happy with all the-

    2. SP

      [chuckles]

    3. JL

      ... efficiency, uh, like squeezed out of the GPUs. Now we're talking more about how do we like improve quality and controllability of these models so that like, you know, the high end of the market, the Hollywood creators, uh, directors can, can take this to the next level. Um, I saw some demos, uh, coincidentally, like, you know, this model came out of, uh, came out the same week or a week prior to Astra. People were combining the Blender experience-

    4. SP

      Yes

    5. JL

      ... with, uh, H3, uh, Max from, from fal. Like, um, talk about how it's gonna impact the, the, the, the Hollywood world.

    6. SP

      U-using Blender, uh, with one, one of these AI models together is a extremely popular workflow for, for professional work. Basically, you, you, you render a low resolution of your scene, what you wanna do using, using Blender, previous like non-AI technology. And then once you add that video as a reference to an AI, AI model, you basically get close to 100% controllability. Um, and th-this, this is an in-in-incredibly popular, um, workflow for VFX artists, people who are doing this professionally because they want exact-- they, they, they wanna get exactly what they put in into, into the model. And a-as, as you mentioned, uh, a week after we, we launched, uh, H3 Max, uh, people starting creating, generating, uh, these beautiful scenes using an LLM model, GPT-Astra in, in Blender, and all of a sudden it unlocked a, a whole new pipeline using LLM to create a Blender scene and then passing that to the H3 Max model or, or any video model. But it, it works very well with H3 Max because it's extremely fast, and you can like try many things all at once in parallel. Um, and that, that unlocked a whole new workflow for Hollywood and professional people, and it gets you to like close to 100% controllability. As, as I said, we have been very, very focused towards speed, performance, quality, and now we have a really good base model. I think the next month or two is gonna be fully focused on, okay, w- how much controllability we can add to these models so that professionals at studios, professionals who want to actually produce, like, produce content that fits their use case perfectly can leverage these models. Uh, the team has been working on an amazing, you know, like, uh, a lip synchronization model where, you know, you can just supply the audio, you can supply your, your, uh, you can supply, like, a video or an image reference, and then it can, like, synchronize the lips perfectly. Same with, like, motion controls. Uh, you can just take a motion of someone dancing and apply it to, like, your, uh, AI-generated character, and it fits perfectly. And this, like-- You can, like, get these results with, like, basic prompting, and you're gonna get, like, 80%, 90% reliability. What we are targeting is, like, 99.9% reliability in the outputs so that you can actually trust the model did every single, uh, aspect of this generation perfectly, and that's, like, what we have been pushing. Uh, one big launch that we had last week was the camera controls, uh, which is essentially you can direct h- where the camera is going within the video perfectly to the, to, to the degree. Uh-

    7. JL

      And this is by, like, describing in the prompt or, like-

    8. SP

      It-

    9. JL

      ... generating the, the scene?

    10. SP

      You just essentially, like, underneath you give a JSON of, like, "I want camera at, like, zero, zero, zero at T0. I want camera at, like, 90 degrees angle at T1." Like, you essentially supply a, a structured, uh, structured description of where your camera needs to be at any point in time, and then the model is, like, perfectly conditioned to, to regard it as, like, the only source of truth, and it doesn't, like, hallucinate, uh, um, like, where the camera should go. And it just, like-- You can essentially reconstruct 3D scenes from a single input, like, because the model is itself is a very good video model, but at the same time, you know, it's, like, perfectly adheres to, to the camera itself.

    11. JL

      And this is because, uh, th-the base model itself already has the understanding of the camera angle that you can-

    12. SP

      It doesn't respect it. It just under-- Like, you need to tune the model. You need to tune the model to a sig-significant degree. And this is, like, what enables, like, at large scale, you know, post-training infrastructure. We now have the infrastructure to take H3 Max, add any capability to it. Same applies for any new model, right? If there's a new video model, we, we, we, we essentially spend most of the time building it as an infrastructure than just, like, one-off training runs so that we can add t- we can build, like, services around this for not just, like, you know, open source models, but for, like, frontier closed source models as well. Because we see in the market this is, like, the biggest gap, is just how controllable these models are. Even, like, first we start with text to video, where you put a prompt, you get a video back. It was good, but, like, you never could describe the perfect character for you.

    13. JL

      Mm-hmm.

    14. SP

      Then we had image to video where, you know, you use an image editing model and then generate, like, you know, the first scene, and then the model was, like, obviously much more fitting. But you still couldn't, like, say, "Oh, I want this new character appear at, like, second three." You need to put it to your first frame or, like, you, you can't, like, you, you can prompt it, but it was never perfect. And then we added reference to video where, you know, uh, you can provide, like, an initial starting frame, and you can also provide, "I want these characters with these voices." Like, you, you know, that's also, like, a big, big unlock where you can essentially say, "This is the, this is the voice for this character." And now, like, you know, we are adding, oh, within the scene, I want camera to look at this degree at, like, T0. I want camera to look at this degree at, like, T3. And then we are adding lighting controls where you essentially say where the light is coming from. These are all, all compounding on top of each other, and, like, we just have the unified infrastructure to just apply this to any model at this point.

    15. JL

      That's incredible.

    16. GY

      Hollywood is our fastest growing segment, and, um, there, there's a lot of noise about how AI might disrupt Hollywood, but it-- Hollywood usage was nonexistent a year ago. And in, in the past year it grew, and now it's, it's the fastest growing segment. Like, um, Amazon MGM Studios in, in their conference, they, they released their Nara, uh, tool. It's, it's mostly backed by, uh, file infrastructure behind the scenes. And we are seeing incredible, incredible pull coming from Hollywood. And it's exactly what they need, these, like, small point solutions rather than generating everything from scratch. They wanna be able to extend the video a little bit. They wanna be able to change the, the camera controls. They wanna change the lighting. And someone has to build these solutions for them. What, what Hollywood needs and what the creators actually need and what the research labs are working on, there's a little bit of a disconnect there, and we believe we can come in and do these little post-training projects to, to close that tab- gap, because we work with, uh, all the Hollywood studios, and we hear from them what, what they need, and these are exactly the things they need, these small point solutions that actually make them more efficient, push out more, more video, and, uh, AI can actually close that gap very nicely.

    17. JL

      Maybe say it in a little bit different way. Like, we have been staring at this problem for the last three years as well. Like, we see, like, companies trying

  10. 35:4338:41

    Where Video Models Go Next & the Creator Economy Shift

    1. JL

      to, like, build a, you know, a movie director, uh, like, video model, like, by, uh, either pre-train or post-train on, on the video side. But what I'm hearing is, like, different people expressing the way they want the output to come out very differently. Consumers talk about it and then, like, write the prompt and generate the results very differently from a Hollywood director-

    2. SP

      Professionals

    3. JL

      ... which is obvious, right?

    4. SP

      Yeah. Yeah.

    5. JL

      Like, professionals wanna talk about, like, you know, these camera angles. They wanna talk about the lighting. Like, you sort of have built a library or, like, a, a collection-

    6. SP

      Toolkit. Yeah, yeah

    7. JL

      ... of, of, um, post-training, like, I would call it data and toolkits that can apply these on any model that you can-

    8. SP

      Mm-hmm

    9. JL

      ... like, grab the weights on so that they are adapted to, like, a different audience where they can express their creativity in a bit different, uh, fashion to control the model-

    10. SP

      100%

    11. JL

      ... when they It unlocks a lot of the capability underneath

    12. GY

      And h- half the problem was capabilities of the, of these models. We are solving that. The other half of the problem was legal and data residency, things like that.

    13. JL

      Yep.

    14. GY

      We, we, we made a ton of progress there as well. Uh, we now have, have a system of, uh, people can apply with, with their own IP, and we unlock their own IP in, in the models. Uh, we are gonna grow that, and that's gonna be, uh, a, a very powerful thing we do with Hollywood studios. Also, we, we now have CDance US-hosted as well. We already had previously other Chinese models. CDance was the missing part. Every Hollywood studio wanted us to have it US-hosted. Now that's available, um, so there are no, no obstacles i- in front of these Hollywood studios now. Everything is ready, and we, we believe they are gonna 10X, 100X their AI usage in the coming months.

    15. JL

      It's s- such a exciting world for, uh, movie lovers, consumers, people who consume a, a lot of video and, and creative content.

    16. GY

      And, uh, uh, we have our conference, Gener- Generative Media Conference next, next week. Uh, this is our second time we are doing it. Last year, it was mostly consumer AI. There were like maybe c- couple Hollywood executives here and there just curious about it, and now it's dominated by studios. New AI studios who are like offshoots of the bigger studios trying to do like only AI, AI shows, but also like the, the biggest of the Hollywood studios are also there because now they have big plans integrating AI into their workflows, into their, uh, existing systems. So you can see the, the change in the attendance of the conference as well.

    17. JL

      That's awesome. Well, for the audience, check out, uh, the content coming out of the, the Gen Media Conference. Uh, it's gonna be very, very exciting. And thank you so much-

    18. GY

      Of course

    19. JL

      ... Gorkem and Batuhan coming onto our show.

    20. GY

      Thank you for having us. Yeah.

    21. JL

      It's super exciting time for Gen Media.

    22. GY

      Thank you.

Episode duration: 38:56

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode SDbRJXQrYGY

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.