How I AIUsing Veo 3 to create AI-generated music videos, like a Tiny Desk concert with Notorious B.I.G.
CHAPTERS
- 0:00 – 6:09
Why AI audio/video tools feel like a new era of remix culture
Anish and Claire frame the episode around AI as an expansion of creative possibility rather than a replacement for human artistry. Anish explains how AI breaks long-standing constraints in music production (like being stuck with a single mixed track), and connects today’s tools to decades of remix and sampling culture.
- •AI enables new forms of audio manipulation (e.g., isolating vocals/drums from a finished track)
- •Creative constraints historically shaped genres (mixtapes, sampling, hip-hop)
- •AI as the next step in remix culture—extending audio remixing into video
- •Claire’s perspective: AI unlocks dormant creativity for busy adults
- 6:09 – 7:40
Concept: resurrecting a Tiny Desk-style performance with Notorious B.I.G.
Anish introduces the core creative project: generating an AI Tiny Desk Concert for artists who can’t perform anymore, starting with Notorious B.I.G. He emphasizes doing it respectfully while capturing the format’s distinctive constrained aesthetic.
- •Tiny Desk as a constraint-driven format that’s uniquely compelling
- •Motivation: “artists you’d want to see” who are no longer alive
- •Ethical intent: respectful recreation vs. derivative mimicry
- •A quick demo clip establishes the target output quality
- 7:40 – 8:05
Workflow overview: still image + target audio + lip-synced animation
Anish lays out the basic recipe behind the Tiny Desk creation: generate a believable still frame, procure or prepare the right audio, then use a tool to animate and lip-sync. Hedra is positioned as a convenient all-in-one that can do frame-to-video plus audio-driven lip sync.
- •Inputs needed: a still image and a clean audio segment
- •Hedra’s role: animate from a still frame and align lip sync to custom audio
- •Alternatives noted (e.g., Sync Labs) but Hedra chosen for simplicity
- •This workflow generalizes beyond music (speeches, characters, storytelling)
- 8:05 – 11:29
Generating the “Tiny Desk” still image with GPT-4o Image Gen
They generate a Tiny Desk-style image of Kurt Cobain to demonstrate the process live, discussing why GPT-4o Image Gen works well for this. Prompt adherence and fine-grained control are highlighted, along with the practical reason for starting with an image: it becomes the anchor frame for animation.
- •GPT-4o praised for prompt adherence and controllability
- •Image gen improvements: reliable text/logos and finer manipulation
- •Adjusting the image (e.g., removing guitar for a cappella) to fit the plan
- •Still image serves as the base asset for frame-to-video animation
- 11:29 – 12:55
Sourcing the right performance audio (and why live aesthetics matter)
Anish explains how the Biggie Tiny Desk effect depends on authentic-sounding live instrumentation. He describes pulling live performance audio from YouTube (e.g., a cover band for Biggie; Nirvana Unplugged as a proxy for Tiny Desk acoustic vibe) and prepping it for downstream editing.
- •Tiny Desk sound target: live, acoustic, in-room feel
- •Biggie approach: live band backing + extracted Biggie vocals layered on top
- •Nirvana approach: leverage existing Unplugged concert audio/video as source
- •Pragmatic tooling: YouTube download utilities (4K Video Downloader)
- 12:55 – 15:37
Editing and syncing audio quickly in Adobe Audition
They move into a practical editing step: trimming dead space and selecting a short workable segment to fit current model clip-length limits. The conversation also reframes limitations as creative constraints, comparing today’s short-generation caps to historical sampler constraints in hip-hop production.
- •Drop video into Audition to access and visualize the audio track
- •Trim silence, select a short segment (e.g., ~15 seconds) for generation
- •Current AI clip-length limitations acknowledged (short segments)
- •Constraints can increase creativity—parallels to early sampler limits
- 15:37 – 16:37
Separating vocals with Demucs (stem extraction)
Anish introduces Demucs as the key utility to split a track into stems, enabling a cappella or re-layered mixes. He demonstrates looking up the command quickly and explains the outcome: clean vocal extraction that unlocks entirely new remixes.
- •Demucs extracts vocals vs. instrumentation from an audio file
- •Fast workflow: look up the CLI command when needed (e.g., via Perplexity)
- •Result: novel outputs (e.g., a cappella that never existed historically)
- •Stem separation enables more believable or creatively reimagined performances
- 16:37 – 22:35
Generating the lip-synced performance in Hedra (simple prompts, strong results)
They combine the generated still image with the chosen audio in Hedra and generate an animated, lip-synced clip. Claire notes Anish’s minimal prompting style, and they discuss why leaving room for the model can produce more delightful outputs—especially in creative work.
- •Hedra inputs: upload start frame + upload audio + short descriptive prompt
- •Prompting style: extremely concise prompts can work well
- •Creative prompting mindset: don’t over-constrain; allow exploration
- •Output review: convincing lip sync and performance micro-gestures
- 22:35 – 27:57
Making a ’90s-style Nirvana music video with Veo 3 (prompt iteration + editing)
Anish shows a separate project: a grunge, camcorder-like ’90s music video assembled from Veo 3 clips. He describes using GPT-4o to refine prompts until the aesthetic matches the intended era, then stitching clips together in an editor (Kapwing), and they discuss typical generation artifacts.
- •Veo 3 used to generate multiple short clips with coherent “physics”
- •GPT-4o helps iterate prompts to dial in grunge/Seattle/camcorder aesthetics
- •Editing pipeline: assemble clips in Kapwing (alternatives like CapCut)
- •Artifacts to watch: duplicated motion/characters, object inconsistencies
- 27:57 – 32:38
Multimodal app building: catalog a bookshelf from video with Gemini Flash
Shifting from art to practical utility, Anish demonstrates building a video-to-structured-data app in Google AI Studio using Gemini Flash. The app takes a video of someone flipping through books, extracts key frames, and returns titles/authors with images—illustrating a “native” multimodal workflow.
- •Gemini Flash highlighted as underused but strong for multimodal/video tasks
- •Google AI Studio positioned as a streamlined surface for Gemini models
- •App logic: video → key frames → vision extraction → sequential list of books
- •Fast prototyping: a few prompts can yield a working demo
- 32:38 – 35:36
From prototype to shareable tool: deployment, costs, and “personal software”
They discuss the gap between a quick personal prototype and a polished public app. Anish shows how AI Studio can deploy via Cloud Run, notes API cost considerations, and both reflect on a future where individuals routinely build bespoke tools for themselves rather than waiting for products.
- •Personal demo can take minutes; productionizing/sharing takes longer
- •Deploy option: Cloud Run makes the app accessible via a link
- •Practical constraint: API usage costs require deliberate sharing
- •Broader trend: rise of personal software built on-demand
- 35:36 – 37:35
Comet browser as an AI agent for personal finance (RPA on websites)
In a lightning round, Anish explains why he uses Perplexity’s Comet browser: it can operate websites on the user’s behalf. He describes using it to analyze his Robinhood portfolio and ask higher-level questions without manual exporting, reframing the browser as an action layer rather than a passive viewer.
- •Comet enables agentic browsing (RPA): models operate the browser for you
- •Finance workflow: portfolio performance summaries and comparative analysis
- •Eliminates manual clicking, exporting, and spreadsheet work
- •Illustrates consumer AI value: make every website “smarter” via an assistant
- 37:35 – 41:24
AI for kids: interactive stories, play, and social-emotional learning
Claire asks what consumer AI demos would change a non-technical person’s mind; Anish answers through parenting. He describes interactive bedtime storytelling and imaginative play scenarios, then looks ahead to AI’s potential role in classrooms—especially observing dynamics and supporting social-emotional development.
- •Interactive bedtime stories: infinite Q&A and personalized creativity
- •Kids use AI for imaginative debates (e.g., “who would win?” scenarios)
- •Longer-term: AI in classrooms for social-emotional learning, not just homework
- •Claire’s example: kids adopting new hardware/interaction patterns naturally
- 41:24 – 43:00
Getting better outputs: embrace surprises and restart without sunk cost
They close with tactical advice for when AI outputs are poor. Anish recommends a mindset shift: follow unexpected directions when they’re interesting, and abandon failing approaches quickly because the model’s iterations aren’t the same as invested human labor.
- •“Go with it” when outputs are strange—unexpected can be delightful
- •Avoid sunk cost fallacy: abandon broken approaches and restart fast
- •Treat failed generations like disposable branches, not hard-earned work
- •Wrap-up and episode sign-off