CHAPTERS
- 0:00 – 0:19
AI video floods social feeds: why this week felt like an inflection point
The hosts set the tone: AI video went from “cool demos” to mainstream social content seemingly overnight. They frame the episode around major consumer AI releases and what they unlock for creators and everyday users.
- •AI video suddenly dominates TikTok/Instagram-style feeds
- •Veo 3 positioned as a ‘ChatGPT moment’ for video
- •Creators and entrepreneurs are becoming ‘AI-assisted’ by default
- •Preview of topics: Veo 3, ChatGPT voice, Apple, ElevenLabs, revenue data, demo workflow
- 0:19 – 2:06
Meet the hosts & the show format: “This Week in Consumer AI”
Justine and Olivia introduce themselves as a16z investing partners (and identical twins) and explain the show’s goal: highlight standout consumer AI product updates and practical workflows. They outline the exact lineup for the episode.
- •Hosts: Justine and Olivia Moore (identical twins)
- •New series: tracking weekly consumer AI breakthroughs
- •Episode roadmap: Veo 3, ChatGPT Advanced Voice, Apple AI, ElevenLabs V3, revenue growth data, Flux Context demo
- •Promise of a hands-on tutorial at the end
- 2:06 – 3:47
What Veo 3 is—and why native audio changes everything
They explain Google DeepMind’s Veo 3 and what separates it from Veo 2: generating audio and video together from a single text prompt. This enables one-shot creation of realistic ‘talking human’ clips, vlog snippets, and podcast-like scenes without external voiceover tools.
- •Veo 3 is Google DeepMind’s latest flagship video model
- •Key leap: native audio generation alongside video (from text prompts)
- •Enables multi-character dialogue scenes in one generation
- •Reduces workflow steps (no separate TTS/voiceover pipeline)
- •Explains why formats like ‘stormtrooper vlogs’ took off
- 3:47 – 4:32
Constraints, hacks, and virality: making longer stories from 8-second clips
Despite the hype, Veo 3 has important limitations: eight-second generations and audio only from text-to-video (not image-to-video). Creators work around character-consistency issues by using masked/non-human characters (stormtroopers, yetis, capybaras) where viewers tolerate small variations across stitched clips.
- •Hard limit: ~8-second clips per generation
- •Audio doesn’t work in image-to-video mode
- •Longer narratives require stitching many clips
- •Character consistency is easier with known or masked characters
- •Virality driven by repeatable formats and recognizable archetypes
- 4:32 – 5:24
How to access Veo 3: pricing, platforms, and why it’s still expensive
They clarify confusion about availability: initial access required Google’s $250/mo AI Ultra plan via Flow, but now Veo 3 is accessible via API. Consumer tools (Hedra, Krea) and developer platforms (Fal, Replicate) offer paid access, though usage costs remain high and prompting mistakes can be costly.
- •Initial distribution: Flow + Google AI Ultra ($250/month)
- •Now available through API integrations
- •Access via consumer apps (e.g., Hedra, Krea) and dev tools (Fal/Replicate)
- •Approx pricing discussed: ~75 cents per second
- •High cost means careful prompting and selective usage
- 5:24 – 6:34
What Veo 3 signals for creators: faceless channels, new storytelling, and cost pressure
They discuss near-term creator behavior shifts—especially the rise of ‘faceless’ channels powered by AI characters. On the model side, they expect push toward longer generations, but note coherence and economics are major constraints; the future likely includes optimized/distilled models to reduce costs.
- •Rise of AI-driven ‘faceless’ creator channels
- •AI characters can narrate stories without filming yourself
- •Audiences latch onto recurring comedic narratives (e.g., incompetent stormtrooper)
- •Model providers face economics: video+audio is expensive to run
- •Expect longer video generation and cheaper/distilled variants over time
- 6:34 – 10:14
ChatGPT Advanced Voice Mode gets more human: what changed and why now
They review OpenAI’s Advanced Voice Mode improvements, which add more natural expressiveness—intonation, filler words, and human-like rhythm. They compare OpenAI’s early lead in real-time voice with newer competitors (Sesame, open source, Gemini, Grok) that surpassed it on realism, and speculate delays were influenced by ‘Her’ controversy and shifting lab priorities.
- •Update rollout: first for paid users, then broader availability
- •Improvements: more natural prosody, pauses, ‘ums/uhs,’ expressiveness
- •Competitive context: Sesame/open source/Gemini/Grok felt more human earlier
- •Live demo highlights realism and conversational cues
- •Possible reasons for delay: ‘Her’ controversy + many competing OpenAI priorities
- 10:14 – 12:17
Apple’s AI announcements: useful features, but Siri still falls short
They react to Apple’s developer conference AI updates and public disappointment with Apple Intelligence. The core critique: Apple appears to offload ‘true AI’ to ChatGPT integrations while Siri remains unable to handle basic assistant tasks; Apple focuses instead on incremental features like Genmoji, transcription, and real-time translation.
- •Apple Intelligence seen as underwhelming vs expectations
- •Siri anecdote: fails a simple calendar-style question, offers to query ChatGPT
- •Apple appears to ‘outsource’ advanced AI to ChatGPT on-device
- •Past notification summaries caused confusion and may have spooked Apple
- •Highlights: Genmoji traction, call transcription, FaceTime real-time translation
- 12:17 – 13:43
ElevenLabs Eleven V3: emotion, interruptions, and richer control via text tags
They explain why ElevenLabs’ third-gen voice model is notable: it moves nuanced delivery (emotion, inflection, accents) into text-based prompting using tags. This reduces the need for recording performances just to transfer style, and adds narrative controls like interruptions for more realistic back-and-forth dialogue.
- •ElevenLabs releases Eleven V3 (next-gen TTS/voice model)
- •New capability: emotion/inflection/accents via text tags (no STT-to-TTS workaround)
- •Interface acts like an editable script with tagged direction
- •Supports sound effects and conversational dynamics (e.g., interruptions)
- •Big unlock for storytelling, ads, and marketing dialogue
- 13:43 – 15:48
Eleven V3 demo: multi-character conversation, accents, sound effects, and cut-offs
Justine shares a short clip demonstrating two characters, prompted accents, and even a cow moo—showing how tags can orchestrate realistic timing and interruptions. They emphasize how these controls make AI voice finally feel like a natural conversation, not isolated lines stitched together.
- •Demo showcases two distinct characters in one scene
- •Prompted accent behavior (including intentionally ‘bad’ accents)
- •Added sound effects (cow moo) through prompting
- •Tag-driven interruptions/cut-offs improve realism
- •Positions V3 as a leap for narrative building and ad creative
- 15:48 – 18:07
a16z data: consumer AI startups are hitting revenue milestones faster than ever
Olivia walks through an analysis of gen-AI-era companies they’ve met, focusing on how quickly revenue ramps after monetization. Compared to pre-AI norms (especially consumer companies delaying monetization), the median consumer AI startup reaches $4.2M ARR run-rate by month 12, with top quartile near $8.7M—outpacing even AI-era B2B benchmarks.
- •Dataset: companies a16z met over the last ~22–24 months
- •Pre-AI benchmark: $1M ARR year one was best-in-class for B2B; consumer monetized much later
- •AI-era shift: consumer companies monetize early via subscription
- •Findings: median consumer ARR run-rate ~$4.2M at month 12; top quartile ~$8.7M
- •Consumer revenue ramp now exceeds B2B benchmarks in the AI era
- 18:07 – 22:17
Why subscription works in AI: inference costs, higher willingness to pay, and retention reality
They explain structural drivers behind faster monetization: AI has real marginal inference costs, pushing companies to charge from day one, and the products are valuable enough that consumers pay more (around $22/month on average). They address skepticism about churn by distinguishing ‘tourism’ among free users from solid paid retention, plus new revenue expansion mechanics via credit packs.
- •AI changes unit economics: real marginal inference costs per user/query
- •Average consumer AI pricing discussed: ~$22/month (higher than pre-AI)
- •Value props: creative superpowers, companionship, language learning, tutoring, coaching, nutrition via vision models
- •Retention: free-user ‘tourism’ exists, but paid retention is comparable to pre-AI
- •Revenue expansion: credit add-ons/overages create upsell dynamics more like enterprise (or games)
- 22:17 – 23:14
Bottoms-up to enterprise: how consumer AI products convert into high-ACV deals
They note an accelerating pattern: consumer/prosumer adoption increasingly leads to enterprise deployment faster than in prior eras (e.g., compared with Canva’s long path). ElevenLabs is cited as an example where hobbyist usage can turn into enterprise contracts when users bring tools into media and entertainment workplaces.
- •Consumer adoption can quickly become enterprise usage
- •Faster consumer-to-enterprise path than historical examples (e.g., Canva)
- •ElevenLabs: starts as $10/mo experimentation, upgrades to enterprise contracts
- •Parallels to early Midjourney usage inside agencies/entertainment
- •Highlights the power of bottoms-up distribution in AI tools
- 23:14 – 27:56
Demo of the week: building a froyo brand with AI (ChatGPT → Ideogram → Krea/Flux Context)
Justine demonstrates an end-to-end AI workflow to create a modern frozen yogurt brand (‘Melt’) including naming, logo direction, packaging colors, product imagery, and even store concepts. The centerpiece is Flux Context (hosted on Krea), an image editing model that preserves identity/consistency far better than typical prompt-based edits—making it viable for real brand collateral.
- •Workflow: ChatGPT for naming/brand direction → Ideogram for logo/typography → Krea with Flux Context for consistent edits
- •Flux Context described as ‘Photoshop with natural language’ and high consistency retention
- •Examples of edits: change environments, adjust colors/borders, modify product attributes (e.g., ube purple froyo)
- •Generated product photos and store signage concepts quickly
- •Sets up next step: turning stills into video ads using Veo 3/Higgsfield
- 27:56 – 29:35
The bigger implication: full-stack AI brands and AI-assisted entrepreneurship
They zoom out from the demo to predict a wave of ‘full stack’ AI brands: AI-designed logos, product photos, websites, ads, and even AI influencers marketing on social platforms. The key shift is accessibility—people can iterate with text prompts instead of mastering complex tools like Photoshop, lowering the barrier to launching a product or small business.
- •AI enables end-to-end brand building: identity, creative, web, marketing
- •Potential for AI-designed products + vibe-coded websites + dropshipping
- •AI-generated social ads and AI influencer-style promotion
- •Natural language editing replaces steep tool learning curves
- •Expectation: many more individuals can launch brands/businesses faster
