Skip to content
a16za16z

How Real-Time AI Video Is Changing How Creators Work

a16z General Partner Jennifer Li sits down with fal co-founder Gorkem Yurtseven and Head of Engineering Batuhan Taskaya to discuss what changes when generative video becomes fast enough to run in real time. They unpack the technical work behind H3 Max, fal’s post-trained version of MiniMax’s open-weight video model, and how combining model post-training with systems and hardware optimization significantly reduced generation time while maintaining quality. That speed has enabled experiments with continuous video, including streams that can remember previous scenes and respond to new directions while they’re running. They also discuss why the next challenge may be less about speed and more about control, from camera movement and lighting to characters, motion, and lip sync. And they explore what those capabilities could mean for professional creative workflows, where artists and studios need predictable tools rather than simply generating a video from a prompt. Timestamps: 00:00 - Intro 00:47 - Meet fal & the H3 Max Launch 02:45 - From Inference Platform to Post-Training: Why fal Made This Bet 06:33 - The Secret Sauce: System Work, Architecture & Cost/Latency Wins 10:44 - The Twitch Moment: Real-Time Video the Day After Launch 13:11 - How the Model Actually Remembers What Happened in a Scene 15:47 - The Economics: Serving Costs, Chip Footprint & Streaming Experiences 22:02 - Beyond Consumer: Unlocking Hollywood-Grade Controllability 28:37 - Prompting a Director: Camera Angles & Scene Understanding 35:43 - Where Video Models Go Next & the Creator Economy Shift Resources: Follow Gorkem Yurtseven on X: https://x.com/gorkem Follow Batuhan Taskaya on X: https://x.com/isidentical Learn more about fal: https://fal.ai Follow Jennifer Li on X: https://x.com/JenniferHli Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Gorkem YurtsevenguestJennifer Lihost
Sep 17, 202638mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

H3 Max makes real-time AI video creation and live directing practical

  1. fal’s H3 Max is a post-trained version of the open-weight MiniMax H3 video model that achieves dramatic speed/cost gains while retaining near-frontier quality.
  2. The team attributes the leap to compounding optimizations across the entire inference pipeline—reduced diffusion steps plus high-utilization kernels and systems tuning for low-batch, single-shot video workloads.
  3. Real-time performance quickly spawned new creator products (Twitch-style continuous generation), and fal extended this into H3 Max Director with longer continuity via minutes of scene memory and live mid-stream steering.
  4. The roadmap shifts from “faster/cheaper” toward “more controllable,” adding references, LoRAs (lip-sync, styles), and structured camera/lighting controls targeting 99.9% reliability for professional workflows.
  5. Hollywood is described as fal’s fastest-growing segment, driven by point solutions (camera, lighting, extensions) and enabled by US-hosted deployments and IP/data-residency tooling.

IDEAS WORTH REMEMBERING

5 ideas

H3 Max’s “secret sauce” is system/model co-design, not just a better base model.

fal combined post-training (to reduce diffusion steps without losing quality) with aggressive kernel/system optimization across the full pipeline (prompt expansion LLM + diffusion + VAE decode + optional upscalers), producing compounding latency/cost gains.

Post-training unlocked an order-of-magnitude speed jump without the usual quality cliff.

They report ~35× speedup versus the original MiniMax H3 endpoint while maintaining comparable ELO/quality, plus a public “Turbo” variant that trades a small quality drop for extreme latency/cost improvements.

Real-time video is achieved by squeezing efficiency from the standard 8-GPU serving unit.

Most state-of-the-art video inference is still served on a single 8-GPU node; scaling beyond that often reduces efficiency due to communication overhead, so they focus on maximizing utilization within the node (raising MFU toward roofline).

Speed made real-time streaming inevitable—and then memory made it feel continuous.

Within a day of launch, users and internal teams created Twitch-like “infinite” streams by chaining clips; fal then shipped a more seamless approach with ~2 minutes of native scene memory plus a rolling “system prompt” summary for longer coherence.

The creator workflow shifts from “render then wait” to “direct live while it plays.”

H3 Max Director extends generation length (up to ~60 minutes) and supports live, action-controlled updates (e.g., injecting new events mid-scene), enabling chat/voice-directed experiences that resemble directing a live set.

WORDS WORTH SAVING

5 quotes

We have a version called H3 Max Turbo that's public that can generate like a five-second video in like 1.5 seconds, which is like insane.

Batuhan (fal)

We decided let's, let's hold off. Let, let's not, let's not tell people that this is like so much faster and so much better before we have some external validation.

Gorkem Yurtseven

This happens at fal once in every couple of months where like the whole company gets, gets hold of something and, and the c-creativity just explodes and everyone is just working on a, a, a new little app or, or a, a, a different optimization LoRA, whatever it might.

Gorkem Yurtseven

There is like two minutes of memory. So like y-you're in a scene and when, when you direct the, the model or someone else enters the room, it actually like everyone looks at that person entering and the s- the scene is continuous.

Gorkem Yurtseven

What we are targeting is, like, 99.9% reliability in the outputs so that you can actually trust the model did every single, uh, aspect of this generation perfectly, and that's, like, what we have been pushing.

Batuhan (fal)

Post-training open-weight video modelsSystem/model co-design and kernel optimizationLatency, cost, and GPU utilization (MFU/roofline)Real-time streaming video experiences (Twitch/WebRTC)Longer-horizon memory and continuous generationControllability: references, LoRAs, camera/lighting controlsHollywood workflows, IP/legal and US-hosted deployments

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.