a16zWhere does consumer AI stand at the end of 2025?
CHAPTERS
- 0:00 – 0:20
Market reality check: single-assistant behavior and early winner-take-most signs
The discussion opens with usage data suggesting consumers rarely multi-home across AI assistants, hinting at a winner-take-most dynamic. The panel frames why leadership in general assistants matters and how quickly shares can shift late in the year.
- •Only ~9% of consumers pay for more than one of ChatGPT/Gemini/Claude/Cursor
- •For much of the year, <10% of ChatGPT users visited other major LLM providers
- •Why distribution + habit formation may create a durable lead
- •Set-up: recap 2025 winners and predict what changes in 2026
- 0:20 – 2:22
Who’s leading consumer AI in 2025: scale, growth rates, and shifting momentum
Olivia quantifies the current market: ChatGPT’s dominant active-user scale versus Gemini’s smaller base but faster growth. The team notes the last 3–6 months have been unusually volatile, with some challengers carving out specialized user segments.
- •ChatGPT estimated at ~800–900M weekly active users; Gemini materially smaller
- •Gemini growth accelerating (desktop up ~155% YoY) while ChatGPT grows slower (~23% YoY)
- •Challengers (Claude/Grok/Perplexity) trail at ~8–10% usage each
- •Emerging specialization: e.g., Anthropic leaning into hyper-technical users
- 2:22 – 2:57
2025 launch timeline: OpenAI vs Google consumer strategy (one super-app vs many surfaces)
The group compares how OpenAI and Google shipped consumer features in 2025. OpenAI concentrated additions inside ChatGPT (with Sora as an exception), while Google spread experiments across Gemini, AI Studio, and Labs-style standalone products.
- •OpenAI: feature accretion inside ChatGPT (Pulse, group chats, shopping, research, tasks)
- •Google: many standalone launches across multiple surfaces (Gemini, AI Studio, Labs)
- •Tradeoff: unified interface vs purpose-built UI per workflow
- •Sora stands out as OpenAI’s primary standalone consumer app
- 2:57 – 4:16
The viral heart of 2025: image and video models that broke through mainstream
Justine argues the most viral consumer moments came from image/video—not text chat. They highlight OpenAI’s Ghibli-style image wave and Sora, alongside Google’s Veo and NanoBanana lines that drove widespread sharing and experimentation.
- •OpenAI’s standout moments: ChatGPT 4.0 image ‘Ghibli’ trend; Sora 2 video
- •Google’s viral engines: Veo 3/3.1; NanoBanana & NanoBanana Pro
- •Best-in-class creative models create trend cycles and pull users into new apps
- •Viral creative capability can temporarily override assistant “default” behavior
- 4:16 – 6:39
From aesthetics to realism + reasoning: what actually improved in multimodal quality
The panel explains how progress moved beyond style into realism, consistency, and multi-input reasoning. They discuss details like background physics, coherent scene logic, and the ability to combine multiple references into a cohesive output.
- •Realism gains: fewer artifacts, better background motion/consistency in video
- •Multi-input reasoning: combining text + multiple images to design cohesive outputs
- •Anecdote shift: from “letters render correctly” to “infographics/market maps”
- •Persistent character/style consistency enables storyboarding workflows
- 6:39 – 7:39
What’s still hard: multi-step image edits, evaluation gaps, and the ‘accuracy’ layer
Despite big leaps, advanced multi-step transformations still break models. Anish introduces ‘accuracy’ as a distinct capability enabled by search integration—critical for things like product photography or historically correct imagery.
- •Multi-step edit benchmark: replacing Monopoly property names with AI labs reveals gaps
- •Common failure modes: placement errors, repeats, overlaps, missing key entities
- •Accuracy differs from realism/reasoning; often needs retrieval/search integration
- •NanoBanana’s search integration seen as an under-appreciated breakthrough
- 7:39 – 10:20
Under-hyped productivity primitives: Pulse, connectors, and the ‘everything app’ ambition
The conversation shifts to productivity and proactive assistants. Pulse and connectors are viewed as important primitives, but their current execution and reliability limit habitual usage—even among power users.
- •ChatGPT’s high frequency (~3–4x/day for some users) enables proactive nudges in theory
- •Pulse seen as under-hyped conceptually but weak in execution/retention
- •Connectors (calendar/email/docs) are powerful when they work, unreliable today
- •Prosumer opportunity: assistants that truly ingest and act on personal/work data
- 10:20 – 11:45
Prosumer workflows that actually stick: Perplexity’s Comet browser as an ‘AI-native workspace’
Olivia highlights Perplexity’s Comet as her most-used standout product, emphasizing agentic browsing and repeatable workflows. The group notes surprising launch traction relative to ChatGPT’s browser efforts, despite distribution disadvantages.
- •Comet: agent + browser UX with reusable workflows (scheduled or page-triggered)
- •Launch spike and sustained traffic outperformed ChatGPT’s Atlas browser launch
- •Perplexity expanding ambition: email assistant, acquisitions, more dedicated interfaces
- •Takeaway: workflow design can beat raw distribution in prosumer niches
- 11:45 – 16:08
Gemini vs ChatGPT: distribution, brand gravity, and onboarding nuance
The panel debates whether Gemini’s viral creative models plus Google distribution can close the gap with ChatGPT’s brand dominance. They argue onboarding UX and ‘what do I do now?’ prompts materially shape mainstream adoption.
- •Google distribution advantage shows on Android (Gemini much closer to ChatGPT than on iOS)
- •ChatGPT as the ‘Kleenex’ of AI: brand becomes the default verb/noun
- •Onboarding contrast: Gemini’s blank prompt vs ChatGPT’s trend/templates-first UI
- •Character consistency and guided prompts increase creation loops and retention
- 16:08 – 21:02
Social AI attempts: group chats, Sora 2 feed, and why the status game is tricky
BK is skeptical that ChatGPT’s productivity core naturally extends into social needs like identity, connection, and entertainment. Sora 2 is framed as a strong creator tool that struggles as a native consumption network because outputs travel better on existing platforms.
- •Product ‘job’: ChatGPT = ‘help me be better’; social apps = entertain/connection/status
- •Group chat works for small-group planning but may not become a broad social graph
- •Sora 2 resembles CapCut more than TikTok: creation-heavy, in-app consumption weaker
- •Status shifts from “me” to “promptcraft/humor,” and viral distribution happens off-platform
- 21:02 – 26:51
The challengers and their personalities: Claude, Meta’s SAM 3, and Grok’s rapid slope
They assess non-duopoly players: Claude’s strength with technical workflows but weaker mainstream accessibility; Meta’s powerful segmentation models that lack consumer packaging; and Grok’s unusually fast multimodal shipping pace with entertainment templates.
- •Claude: opinionated, strong artifacts/skills/workflows; still ‘too technical’ for mainstream
- •Teen adoption anecdote: Character AI far more used than Claude, highlighting distribution/positioning gaps
- •Meta: SAM 3 segmentation across video/image/audio; standout consumer feature is IG AI translations/voice cloning
- •Grok: steepest image/video progress; fast feature shipping (text/video/audio/lip-sync/longer clips) + viral templates
- 26:51 – 36:25
Predictions for 2026: enterprise pull-through, apps directories, multimodality, and compute constraints
The group forecasts a 2026 shaped by enterprise adoption feeding consumer habits, and by ‘apps’ as a new channel inside assistants. They also predict broader multimodality (‘anything in, anything out’) while emphasizing compute tradeoffs that labs must manage.
- •Enterprise usage as a consumer growth flywheel (work mandates can set defaults)
- •ChatGPT ‘apps SDK/directory’ could reshape both consumer discovery and SaaS workflows
- •Multimodality: video↔image↔text editing loops; convergence toward mega-model behavior
- •Compute tension: training vs inference; entertainment virality vs serious coding/intelligence use cases
- 36:25 – 43:14
What to try right now: standout products, stacks, and practical recommendations
They close by recommending concrete tools across multimodal agents, creative suites, reading/listening, note-taking, and coding. The throughline is that trying many products quickly builds intuition about where AI is already delivering real value.
- •Pomeli (Google Labs): agent + brand understanding + campaign generation demo of ‘agentic multimodal’
- •Krea: multi-model creative interface with reusable elements/characters/styles
- •ElevenLabs Reader: convert reading backlog into audio to fit modern consumption habits
- •Gamma (decks), Granola (meeting notes), Comet (AI-native browsing), Codex/Cursor for coding + knowledge work