Skip to content
a16za16z

What You Missed in AI This Week (Google, Apple, ChatGPT)

Things in consumer AI are moving fast. In this episode, Justine and Olivia Moore, investing partners (and identical twins!) at a16z, break down what’s real, what’s overhyped, and what’s next across the consumer AI space. They cover: - Veo 3: how Google's video model unlocked a new genre of content - OpenAI’s Advanced Voice Mode: upgrades, realism, and... um, human-like hesitation - Apple's AI announcements - ElevenLabs' V3: expressive voice tags, real-time interruptions, and narrative tools for creators - New data from a16z: AI consumer startups are ramping revenue faster than ever—and they show you how - Justine walks through how she used ChatGPT, Ideogram, and Krea to launch a fully AI-assisted brand prototype (store photos and all) It’s exhausting (in the best way) to be a creative in the age of AI. Timecodes: 00:00 Introduction 00:28 Meet the Hosts: Justine and Olivia 00:45 Veo 3: The Game-Changer in AI Video 06:34 ChatGPT's Advanced Voice Mode Updates 10:22 Apple's AI Announcements and Siri's Shortcomings 12:18 ElevenLabs' New Voice Model: 11 V3 15:50 Report from a16z: AI Revenue Growth 23:14 Demo of the Week: AI in Brand Creation Resources: Read ‘What “Working” Means in the Era of AI Apps’: https://a16z.com/revenue-benchmarks-ai-apps/ Find Justine on X: https://x.com/venturetwins Find Olivia on X: https://x.com/omooretweets Tools Discussed: Veo 3: https://gemini.google/overview/video-generation OpenAI: https://openai.com/chatgpt ElevenLabs (V3 voice model) – https://elevenlabs.io/ Ideogram (logo/image generation) – https://ideogram.ai/ Black Forest Labs/Flux Context (image editing via Krea) – https://www.krea.ai/ Flux Context demo (Krea launch post) – https://www.krea.ai/blog/flux-context Hedra: https://www.hedra.com/ Stay Updated: Let us know what you think: https://ratethispodcast.com/a16z Find a16z on Twitter: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Subscribe on your favorite podcast app: https://a16z.simplecast.com/ Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Olivia MoorehostJustine Moorehost
Jun 13, 202529mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:19

    AI video floods social feeds: why this week felt like an inflection point

    The hosts set the tone: AI video went from “cool demos” to mainstream social content seemingly overnight. They frame the episode around major consumer AI releases and what they unlock for creators and everyday users.

    • AI video suddenly dominates TikTok/Instagram-style feeds
    • Veo 3 positioned as a ‘ChatGPT moment’ for video
    • Creators and entrepreneurs are becoming ‘AI-assisted’ by default
    • Preview of topics: Veo 3, ChatGPT voice, Apple, ElevenLabs, revenue data, demo workflow
  2. 0:19 – 2:06

    Meet the hosts & the show format: “This Week in Consumer AI”

    Justine and Olivia introduce themselves as a16z investing partners (and identical twins) and explain the show’s goal: highlight standout consumer AI product updates and practical workflows. They outline the exact lineup for the episode.

    • Hosts: Justine and Olivia Moore (identical twins)
    • New series: tracking weekly consumer AI breakthroughs
    • Episode roadmap: Veo 3, ChatGPT Advanced Voice, Apple AI, ElevenLabs V3, revenue growth data, Flux Context demo
    • Promise of a hands-on tutorial at the end
  3. 2:06 – 3:47

    What Veo 3 is—and why native audio changes everything

    They explain Google DeepMind’s Veo 3 and what separates it from Veo 2: generating audio and video together from a single text prompt. This enables one-shot creation of realistic ‘talking human’ clips, vlog snippets, and podcast-like scenes without external voiceover tools.

    • Veo 3 is Google DeepMind’s latest flagship video model
    • Key leap: native audio generation alongside video (from text prompts)
    • Enables multi-character dialogue scenes in one generation
    • Reduces workflow steps (no separate TTS/voiceover pipeline)
    • Explains why formats like ‘stormtrooper vlogs’ took off
  4. 3:47 – 4:32

    Constraints, hacks, and virality: making longer stories from 8-second clips

    Despite the hype, Veo 3 has important limitations: eight-second generations and audio only from text-to-video (not image-to-video). Creators work around character-consistency issues by using masked/non-human characters (stormtroopers, yetis, capybaras) where viewers tolerate small variations across stitched clips.

    • Hard limit: ~8-second clips per generation
    • Audio doesn’t work in image-to-video mode
    • Longer narratives require stitching many clips
    • Character consistency is easier with known or masked characters
    • Virality driven by repeatable formats and recognizable archetypes
  5. 4:32 – 5:24

    How to access Veo 3: pricing, platforms, and why it’s still expensive

    They clarify confusion about availability: initial access required Google’s $250/mo AI Ultra plan via Flow, but now Veo 3 is accessible via API. Consumer tools (Hedra, Krea) and developer platforms (Fal, Replicate) offer paid access, though usage costs remain high and prompting mistakes can be costly.

    • Initial distribution: Flow + Google AI Ultra ($250/month)
    • Now available through API integrations
    • Access via consumer apps (e.g., Hedra, Krea) and dev tools (Fal/Replicate)
    • Approx pricing discussed: ~75 cents per second
    • High cost means careful prompting and selective usage
  6. 5:24 – 6:34

    What Veo 3 signals for creators: faceless channels, new storytelling, and cost pressure

    They discuss near-term creator behavior shifts—especially the rise of ‘faceless’ channels powered by AI characters. On the model side, they expect push toward longer generations, but note coherence and economics are major constraints; the future likely includes optimized/distilled models to reduce costs.

    • Rise of AI-driven ‘faceless’ creator channels
    • AI characters can narrate stories without filming yourself
    • Audiences latch onto recurring comedic narratives (e.g., incompetent stormtrooper)
    • Model providers face economics: video+audio is expensive to run
    • Expect longer video generation and cheaper/distilled variants over time
  7. 6:34 – 10:14

    ChatGPT Advanced Voice Mode gets more human: what changed and why now

    They review OpenAI’s Advanced Voice Mode improvements, which add more natural expressiveness—intonation, filler words, and human-like rhythm. They compare OpenAI’s early lead in real-time voice with newer competitors (Sesame, open source, Gemini, Grok) that surpassed it on realism, and speculate delays were influenced by ‘Her’ controversy and shifting lab priorities.

    • Update rollout: first for paid users, then broader availability
    • Improvements: more natural prosody, pauses, ‘ums/uhs,’ expressiveness
    • Competitive context: Sesame/open source/Gemini/Grok felt more human earlier
    • Live demo highlights realism and conversational cues
    • Possible reasons for delay: ‘Her’ controversy + many competing OpenAI priorities
  8. 10:14 – 12:17

    Apple’s AI announcements: useful features, but Siri still falls short

    They react to Apple’s developer conference AI updates and public disappointment with Apple Intelligence. The core critique: Apple appears to offload ‘true AI’ to ChatGPT integrations while Siri remains unable to handle basic assistant tasks; Apple focuses instead on incremental features like Genmoji, transcription, and real-time translation.

    • Apple Intelligence seen as underwhelming vs expectations
    • Siri anecdote: fails a simple calendar-style question, offers to query ChatGPT
    • Apple appears to ‘outsource’ advanced AI to ChatGPT on-device
    • Past notification summaries caused confusion and may have spooked Apple
    • Highlights: Genmoji traction, call transcription, FaceTime real-time translation
  9. 12:17 – 13:43

    ElevenLabs Eleven V3: emotion, interruptions, and richer control via text tags

    They explain why ElevenLabs’ third-gen voice model is notable: it moves nuanced delivery (emotion, inflection, accents) into text-based prompting using tags. This reduces the need for recording performances just to transfer style, and adds narrative controls like interruptions for more realistic back-and-forth dialogue.

    • ElevenLabs releases Eleven V3 (next-gen TTS/voice model)
    • New capability: emotion/inflection/accents via text tags (no STT-to-TTS workaround)
    • Interface acts like an editable script with tagged direction
    • Supports sound effects and conversational dynamics (e.g., interruptions)
    • Big unlock for storytelling, ads, and marketing dialogue
  10. 13:43 – 15:48

    Eleven V3 demo: multi-character conversation, accents, sound effects, and cut-offs

    Justine shares a short clip demonstrating two characters, prompted accents, and even a cow moo—showing how tags can orchestrate realistic timing and interruptions. They emphasize how these controls make AI voice finally feel like a natural conversation, not isolated lines stitched together.

    • Demo showcases two distinct characters in one scene
    • Prompted accent behavior (including intentionally ‘bad’ accents)
    • Added sound effects (cow moo) through prompting
    • Tag-driven interruptions/cut-offs improve realism
    • Positions V3 as a leap for narrative building and ad creative
  11. 15:48 – 18:07

    a16z data: consumer AI startups are hitting revenue milestones faster than ever

    Olivia walks through an analysis of gen-AI-era companies they’ve met, focusing on how quickly revenue ramps after monetization. Compared to pre-AI norms (especially consumer companies delaying monetization), the median consumer AI startup reaches $4.2M ARR run-rate by month 12, with top quartile near $8.7M—outpacing even AI-era B2B benchmarks.

    • Dataset: companies a16z met over the last ~22–24 months
    • Pre-AI benchmark: $1M ARR year one was best-in-class for B2B; consumer monetized much later
    • AI-era shift: consumer companies monetize early via subscription
    • Findings: median consumer ARR run-rate ~$4.2M at month 12; top quartile ~$8.7M
    • Consumer revenue ramp now exceeds B2B benchmarks in the AI era
  12. 18:07 – 22:17

    Why subscription works in AI: inference costs, higher willingness to pay, and retention reality

    They explain structural drivers behind faster monetization: AI has real marginal inference costs, pushing companies to charge from day one, and the products are valuable enough that consumers pay more (around $22/month on average). They address skepticism about churn by distinguishing ‘tourism’ among free users from solid paid retention, plus new revenue expansion mechanics via credit packs.

    • AI changes unit economics: real marginal inference costs per user/query
    • Average consumer AI pricing discussed: ~$22/month (higher than pre-AI)
    • Value props: creative superpowers, companionship, language learning, tutoring, coaching, nutrition via vision models
    • Retention: free-user ‘tourism’ exists, but paid retention is comparable to pre-AI
    • Revenue expansion: credit add-ons/overages create upsell dynamics more like enterprise (or games)
  13. 22:17 – 23:14

    Bottoms-up to enterprise: how consumer AI products convert into high-ACV deals

    They note an accelerating pattern: consumer/prosumer adoption increasingly leads to enterprise deployment faster than in prior eras (e.g., compared with Canva’s long path). ElevenLabs is cited as an example where hobbyist usage can turn into enterprise contracts when users bring tools into media and entertainment workplaces.

    • Consumer adoption can quickly become enterprise usage
    • Faster consumer-to-enterprise path than historical examples (e.g., Canva)
    • ElevenLabs: starts as $10/mo experimentation, upgrades to enterprise contracts
    • Parallels to early Midjourney usage inside agencies/entertainment
    • Highlights the power of bottoms-up distribution in AI tools
  14. 23:14 – 27:56

    Demo of the week: building a froyo brand with AI (ChatGPT → Ideogram → Krea/Flux Context)

    Justine demonstrates an end-to-end AI workflow to create a modern frozen yogurt brand (‘Melt’) including naming, logo direction, packaging colors, product imagery, and even store concepts. The centerpiece is Flux Context (hosted on Krea), an image editing model that preserves identity/consistency far better than typical prompt-based edits—making it viable for real brand collateral.

    • Workflow: ChatGPT for naming/brand direction → Ideogram for logo/typography → Krea with Flux Context for consistent edits
    • Flux Context described as ‘Photoshop with natural language’ and high consistency retention
    • Examples of edits: change environments, adjust colors/borders, modify product attributes (e.g., ube purple froyo)
    • Generated product photos and store signage concepts quickly
    • Sets up next step: turning stills into video ads using Veo 3/Higgsfield
  15. 27:56 – 29:35

    The bigger implication: full-stack AI brands and AI-assisted entrepreneurship

    They zoom out from the demo to predict a wave of ‘full stack’ AI brands: AI-designed logos, product photos, websites, ads, and even AI influencers marketing on social platforms. The key shift is accessibility—people can iterate with text prompts instead of mastering complex tools like Photoshop, lowering the barrier to launching a product or small business.

    • AI enables end-to-end brand building: identity, creative, web, marketing
    • Potential for AI-designed products + vibe-coded websites + dropshipping
    • AI-generated social ads and AI influencer-style promotion
    • Natural language editing replaces steep tool learning curves
    • Expectation: many more individuals can launch brands/businesses faster

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.