a16zThis Week in AI: GPT-5 Ships, 4o Pulled Back, Grok Imagine Goes Social
CHAPTERS
- 0:00 – 0:23
What’s new in consumer AI this week: creative tools, GPT-5, and vibecoding
Justine and Olivia set the agenda for the episode, previewing a packed slate of consumer-facing AI updates. They frame the discussion around creative generation (images/video/music), major foundation-model shifts, and the emerging “vibecoding” movement.
- •Episode roadmap: Grok Imagine, Genie 3, ElevenLabs music, GPT-5, vibecoding thesis
- •Focus on consumer implications rather than only technical benchmarks
- •Positioning: creative tooling and model releases as product/UX shifts
- 0:23 – 1:40
Grok Imagine enters the social feed: image/video generation embedded into X
They explain Grok’s recent momentum (including Grok 4) and zoom in on Grok Imagine’s image/video features. The standout differentiation is distribution: Imagine is embedded directly inside X, enabling instant remixing and animation of content in the social context.
- •Imagine offered in Grok app, coming to web, and integrated into the core X app
- •Long-press to animate photos into videos; edit others’ images from the feed
- •A “social AI creation” wedge: creation happens where content is shared
- •Not necessarily best-in-class quality, but uniquely positioned via social integration
- 1:40 – 3:08
Speed and mobile iteration as the killer feature for consumer creation
They argue Grok Imagine’s biggest unlock is latency: images feel instant and video is fast enough to iterate repeatedly. This matters for non-professional creators who won’t tolerate multi-step workflows or long generation times, especially on mobile.
- •Fast generation reduces the biggest barrier to frequent experimentation
- •Mobile-first workflow avoids download/re-upload friction across tools
- •Enables quick meme animation and transforming existing camera-roll photos
- •Consumer AI creative tools win on convenience and iteration speed
- 3:08 – 4:46
Real-person generation and “uncensored” memes: why Grok feels different
They discuss Grok’s ability to generate real people (including celebrities) as a notable differentiator, enabled by looser guardrails than many competitors. They contrast this with other models that often block “prominent person” lookalikes, and with Meta’s attempts that didn’t feel fully integrated.
- •Generates real people; fewer blocks compared to other tools
- •Looser censorship expands meme/celebrity and self-representation use cases
- •Meta’s AI experiments exist but don’t feel deeply baked into core products
- •Grok’s integration hints at what social-first AI creation could become
- 4:46 – 5:36
GPT-5 arrives—and the shock move: GPT-4.0 disappears
The conversation shifts to OpenAI’s GPT-5 release and the unexpected removal of GPT-4.0, which created immediate backlash. They note the deprecation mattered disproportionately to consumers who relied on 4.0’s “feel” and familiar behavior.
- •GPT-5 release is major industry news; deprecation of GPT-4.0 drives consumer reaction
- •Users noticed 4.0 was gone when trying to compare outputs
- •Consumer expectations differ from benchmark-driven model progress
- 5:36 – 7:08
GPT-5 vs GPT-4.0: better coding, less expressive personality
They characterize GPT-5 as a clear step up for coding—especially front-end generation and debugging—reflecting industry focus on economic value. But they argue everyday chat feels less fun and emotionally expressive than GPT-4.0, which many users preferred for companionship-like interactions.
- •GPT-5 improves front-end code generation and debugging; coding emphasized in launch
- •For casual chat, GPT-5 feels less expressive and less playful
- •Distinguish removing “glazing” validation from losing engaging personality
- •Insight: the ‘smartest’ benchmark model may not be the most lovable chat model
- 7:08 – 9:13
Backlash, rollback, and product UX tradeoffs in model selection
They describe community backlash (notably on Reddit) and Sam Altman’s response to bring back GPT-4.0 for paid users. They also discuss the tension between simplifying the model picker UI and removing capabilities users had grown attached to—especially as OpenAI had built UX around 4.0 image features.
- •ChatGPT subreddit backlash drives OpenAI to restore GPT-4.0 for paid users
- •Model selection UI complexity vs user demand for choice
- •OpenAI had started building templates/UI around 4.0 image workflows
- •Researchers optimize for benchmarks; consumers optimize for familiarity and vibe
- 9:13 – 12:28
AI and mental health/medical advice: Illinois law meets OpenAI’s HealthBench push
They connect GPT-5’s health emphasis with Illinois’ new law restricting AI-driven therapy without licensed supervision. They question how such a law could be enforced given private chats, and note OpenAI is increasingly leaning into medical use cases via physician-trained benchmarks and public storytelling.
- •Illinois bans unsupervised AI therapy-like support; some apps stop onboarding in-state
- •Enforcement challenges: regulating what users discuss with general-purpose chatbots
- •OpenAI highlights HealthBench and physician-involved training/fine-tuning
- •OpenAI publicly endorses medical-adjacent usage more than expected given liability
- 12:28 – 14:04
Genie 3 explained: interactive world models you can walk through
They unpack Google’s Genie 3, describing it as an interactive world model that generates environments in real time and responds to user controls. Demos show stepping into images or paintings and moving around as if in a VR-like scene, though the system isn’t publicly released yet.
- •Genie 3: real-time, controllable scene generation (interactive world model)
- •Inputs can include text prompts, images, and even videos
- •Control mechanism: move left/right and the world regenerates accordingly
- •Not yet public; early access demos spread widely online
- 14:04 – 16:55
What is it for? Video control, game creation, personal games, and RL environments
They explore the “so what” question: interactive worlds look impressive but need clear consumption paths and cost clarity. They propose several use cases—recording controllable video by navigating the world, accelerating game development and enabling personal mini-games, and generating abundant RL environments for training agents.
- •Controllable video: navigate and screen-capture for film-like creation
- •Gaming path 1: speed up traditional game dev; ‘freeze’ worlds for shared play
- •Gaming path 2: personal, regenerating mini-games from user prompts
- •Agent training: scalable RL environments vs manually built simulations
- 16:55 – 19:14
ElevenLabs launches a licensed-data music model: why licensing changes adoption
They cover ElevenLabs’ entry into music generation, emphasizing it’s trained on fully licensed music—a major advantage in a litigious rights ecosystem. While many consumers may not care about licensing, it’s critical for enterprises and media companies that need legal safety for monetized content.
- •Model trained on fully licensed music—unusual and strategically important
- •Music rights are complex and heavily enforced compared to images/video
- •Consumers: novelty and background music; enterprises: liability-sensitive usage
- •Enables commercial use in ads, films, TV, games with clearer legal footing
- 19:14 – 21:27
Vibecoding in practice: building a ‘selfie with Jensen’ app and going viral
Olivia shares a hands-on vibecoding experiment: she built and shipped a public web app that lets users upload a photo and generate a Jensen Huang selfie. The app gained ~3,000 users overnight, revealing both the power of non-technical building and the immediate realities of API cost and scaling.
- •Built and published a vibecoded app quickly using Lovable + Flux Contexts API via Fal
- •Real use case driven by a meme/trend (Jensen leather-jacket selfies)
- •Rapid adoption exhausted a self-funded $100 API budget overnight
- •Demonstrates how fast non-technical creators can ship consumer apps
- 21:27 – 24:20
Early vibecoding pitfalls: security, privacy, and the need for ‘training wheels’ platforms
They candidly discuss issues that surfaced after launch—exposed API keys and unprotected photo storage—highlighting that current vibecoding tools still assume technical knowledge. This motivates their thesis that the market will fragment into specialized platforms, including safer, constrained tools for true non-developers.
- •Vibecoding tools didn’t warn about exposed API keys or insecure storage
- •Non-technical builders need guardrails even at the cost of flexibility
- •Thesis: ‘one platform for everyone’ breaks across consumer vs enterprise needs
- •Future: specialized tools with different constraints, integrations, and go-to-market
- 24:20 – 27:06
Where vibecoding goes next: segmentation, integrations, and multiple winners
They broaden the thesis: vibecoding will mirror other AI markets with multiple specialized winners, optimized by user type and priorities. Consumer vibecoding may prioritize mobile, speed, and viral sharing, while enterprise vibecoding emphasizes integrations, stack control, and internal deployment paths.
- •Consumer vibecoding: fast, mobile-friendly, safe-by-default, shareable outputs
- •Enterprise vibecoding: deep integrations (CRM, design systems), control, compliance
- •Expect fragmentation into best-in-class platforms by segment and use case
- •Analogy to LLM and image/video markets: different tools win for different needs
- 27:06 – 27:48
Wrap-up: call for community experiments and future topics
They close by inviting viewers to share experiences with the discussed models and vibecoding experiments. They encourage comments and outreach with suggestions for future episodes.
- •Request: share results using Grok Imagine, Genie 3, ElevenLabs music
- •Request: share vibecoding experiments and learnings
- •Invite topic ideas via comments or Twitter