Skip to content
a16za16z

What Would Make an AI Assistant Worth Paying For?

a16z General Partner Anish Acharya sits down with Assistant Benchmark creator David Pawlan to unpack the sudden explosion of personal AI agents and what it will take for one to become part of everyday life. David has been testing dozens of assistants across real-world tasks, from managing email and booking travel to handling financial admin. They discuss why the most useful agents may become increasingly invisible, proactively checking you into flights, finding refunds, filing reimbursements, or simply handling the small tasks that pile up across everyday life. They also explore whether the winning interface is an app, text thread, voice, or wearable; how much autonomy consumers will actually give their agents; and what happens when agents start interacting with other agents. From commerce and restaurant reservations to entirely new agent-native services, Anish and David ask what the internet looks like when software starts acting on our behalf. Timestamps: 00:00 - Intro 00:54 - Poke, Instinct, Muse and the agent boom 03:41 - What Assistant Bench actually tests 09:14 - Cost savers beat time savers 14:53 - Muse charm as Meta's data play 20:03 - Silent agents in group chats 26:01 - Proactivity is the real moat 32:21 - Assistant vs agent, defined 40:22 - Amazon blocks Muse, Shopify opens the door 47:35 - The $20/day agent economics Resources: Follow David Pawlan on LinkedIn: https://www.linkedin.com/in/david-pawlan/ Follow Anish Acharya on X: https://x.com/illscience Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

David PawlanguestAnish Acharyahost
Sep 29, 202650mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 3:42

    Why personal agents are suddenly booming (Poke → OpenClaw → Instinct/Muse wave)

    The conversation opens with how consumer AI assistants shifted from a novelty to a fast-moving market, driven by a few recent “magic moments” in the ecosystem. David traces the rapid launch cadence from early pioneers like Poke to OpenClaw, then the explosion of mainstream excitement around Instinct, Muse, and a growing long tail of alternatives.

    • •ChatGPT as the first mass-market “synthetic person” moment
    • •Coding agents as the second big inflection in agent adoption
    • •Poke as an early consumer agent with a memorable pricing/negotiation gimmick
    • •OpenClaw as the capability primitive that later products packaged for consumers
    • •Instinct/Muse and many new entrants creating a ‘Wild West’ experimentation phase
  2. 3:42 – 5:17

    AssistantBench: comparing assistants by real tasks, not hype

    David explains AssistantBench as a consumer-facing benchmark that tests assistants on practical prompts like booking flights or finding nearby restaurants. The goal is to evaluate actual outcomes, follow-up questions, and performance across multiple dimensions so users can understand which assistants work for which jobs.

    • •Same prompt across assistants to compare outcomes apples-to-apples
    • •Focus on one-shot task completion and practical reliability
    • •Measures behavior like asking clarifying questions and response speed
    • •Designed as a user decision tool more than a deep technical benchmark
    • •Rapid adoption: large traffic and inbound interest from many founders
  3. 5:17 – 6:50

    Market map: horizontal consumer agents vs vertical specialists vs B2B workflow assistants

    They break down the crowded landscape into major buckets: generalized B2C assistants, specialized consumer agents (e.g., travel/email), and B2B workflow tools that resemble executive assistants for teams. Despite many categories, most products are still trying to be horizontal and do everything.

    • •Two main groupings: B2C generalized vs B2B workflow assistants
    • •Specialized consumer agents exist (travel, email) but are less dominant
    • •Examples of long-tail consumer entrants (Caddie, Ollie, Season, etc.)
    • •B2B assistants focus on workflow/executive-assistant patterns
    • •The competitive focus is mostly horizontal ‘do-everything’ agents
  4. 6:50 – 9:15

    What people actually use agents for: daily admin beats flashy travel demos

    David shares insights from large group chats of agent power users: travel is viral but not the most common daily use. Real demand clusters around tedious life admin, plus meta-use-cases like agent orchestration, with finance as an aspirational—but often impractical—area people love to explore.

    • •Travel is highly shareable but infrequent for most users
    • •Top use: daily admin (email cleanup, forms, repetitive tasks)
    • •Second: agent orchestration (agents coordinating with other agents/tools)
    • •Third (biased by tech users): development/coding and design iteration help
    • •Finance is popular conceptually but hard to execute reliably in practice
  5. 9:15 – 12:24

    Cost savers beat time savers: ‘free money’ workflows that drive adoption

    They argue most consumers don’t care about marginal productivity gains, but they do care about removing painful admin and saving money. The strongest hooks feel like “free money” or invisible cost reduction, where the assistant quietly delivers value without requiring management overhead.

    • •Most users don’t value being ‘10% more efficient’
    • •Winning agents feel invisible: do work, then report results
    • •Examples: auto-filing HSA reimbursements from receipts
    • •Example: monitor airfare price drops and request airline credits
    • •Example: agent-controlled sprinklers tied to weather cutting water bills ~50%
  6. 12:24 – 14:38

    Where assistants should live: iMessage vs dedicated apps vs widgets, plus hardware surfaces

    The discussion turns to interaction surfaces and consumer preference. They compare messaging-based experiences (iMessage) to standalone apps and emerging surfaces like widgets and ambient hardware, emphasizing that preferences may fragment by generation, task type, and even visual vs functional workflows.

    • •iMessage feels ‘privileged’ and personal; less friction than new apps
    • •Preferences may differ by generation (text-first vs app-first) and by use case
    • •Some users want visual interfaces for planning, goals, and travel inspiration
    • •New surfaces emerging: widgets (e.g., Skye) and ambient devices
    • •Hardware form factors may diversify like jewelry—no single ‘ultimate’ device
  7. 14:38 – 17:35

    Muse Charm and Meta’s angle: ambient capture and data flywheels

    They analyze Meta’s Muse Charm hardware, debating whether it’s primarily a mass-market assistant device or a strategic data collection play. David’s hot take is that always-on, real-world capture could generate training data that supports Meta’s longer-term ambitions around mapping the real world.

    • •Muse Charm introduces a new interaction form factor beyond phone apps
    • •Skepticism about mass adoption beyond the early-adopter ‘Twitter bubble’
    • •Thesis: Charm may be more about real-world data capture than hardware sales
    • •Comparison to glasses: Charm could be more consistently ‘always on’
    • •Implications for privacy norms and social acceptability in public spaces
  8. 17:35 – 20:03

    Voice assistants as the breakthrough: hands-busy, invisible execution (ChatGPT Voice connectors)

    David describes a second “magic moment”: using ChatGPT Voice with Gmail/Calendar connectors while biking to reach inbox zero. They argue voice becomes most valuable when users are occupied, making assistants feel truly invisible and supportive rather than another screen-bound workflow.

    • •ChatGPT Voice + connectors enables real work: email labeling, replies, invites
    • •Voice shines when hands/eyes are busy (biking, cooking, gardening)
    • •Full-duplex voice and thread awareness enable higher-context assistance
    • •Assistants should execute in the background, then update the user
    • •Mass adoption likely depends on solving painful real-life moments, not tech demos
  9. 20:03 – 23:23

    Multiplayer agents in group chats: intrusive bots vs silent listeners (Dock pattern)

    They explore why agents added to group chats often feel socially intrusive, and what designs might work. David highlights three approaches, with the most promising being ‘silent’ agents that listen, take notes, and privately message individuals only when action is needed.

    • •Group-chat agents can diminish social capital by ‘taking the mic’
    • •Three patterns: (1) bot in chat, (2) agent-to-agent behind-the-scenes network, (3) silent listener
    • •Dock approach: agent listens, notes, then pings individuals in separate threads
    • •Forum/Discord-style posting may feel less intrusive than messaging threads
    • •Agents should be utilitarian in social spaces unless they truly add value
  10. 23:23 – 25:24

    Character and personality: cute avatars vs uncanny humans, and what’s actually defensible

    They discuss whether agents should feel human and how design affects comfort. While communication style can be personalized, David argues personality alone isn’t a moat; the deeper differentiator is how agents behave under ambiguity—especially how proactive or presumptuous they are.

    • •Human-like agents can feel uncanny; Muse’s ‘adorable yeti’ feels safer
    • •Poke’s sass is polarizing; some users won’t share data due to tone
    • •Personality is easy to configure and store as user preference
    • •More durable differentiation comes from behavior policies and risk tolerance
    • •Key axis: presumptuous/proactive vs cautious/permissioned action
  11. 25:24 – 32:04

    Proactivity as the real moat—and the trust line you can’t cross

    They argue the best agents will act without being asked, but only within safe bounds. The challenge is calibrating proactive behavior to maximize delight and usefulness while avoiding irreversible mistakes that shatter trust—illustrated by flight check-in mishaps and the ‘break up with her’ joke.

    • •Proactivity differentiates leaders from the pack
    • •Some domains require permission (insurance switches, sensitive decisions)
    • •Other domains are ‘safe proactive’ (draft email, get credits/refunds)
    • •Trust is fragile: cross the line once and users may churn permanently
    • •Virality often comes from edgy presumptuous actions—good and bad
  12. 32:04 – 38:06

    Assistant vs agent: definitions, social framing, and the long-run human impact

    They clarify terminology: assistants follow instructions, while agents have agency and can initiate actions. The discussion expands to how agent framing may change social comfort and how, over time, assistants could reduce bureaucratic burden and expand what people can do—ultimately shifting attention away from screens.

    • •Assistant: executes what you ask; Agent: has autonomy to create outcomes
    • •Language matters for anthropomorphism and social acceptance
    • •Agents could mediate awkward interactions via non-human ‘social indirection’
    • •Near-term win: reduce life admin and bureaucratic pressure (legal/financial access)
    • •Long-term hope: more time for relationships and exploration, less screen fixation
  13. 38:06 – 47:35

    Agent-to-agent infrastructure and agentic commerce: Amazon blocks, Shopify welcomes

    They predict a new layer of infrastructure built specifically for agent-to-agent interactions: agent emails, phone numbers, security, and new service workflows. In commerce, they analyze why Amazon might block agents (ad/impulse economics) while Shopify benefits, and they explore how reservations, auctions, and recommendation engines may be reshaped.

    • •New paradigm: agent-to-agent booking, service requests, and coordination
    • •Security becomes a major follow-on industry as agents gain capabilities
    • •Amazon’s ads depend on human eyeballs; agents threaten that model
    • •Shopify benefits from more transactions and merchant empowerment
    • •Restaurants/reservations may shift to loyalty/LTV allocation or bidding markets
    • •Recommendation becomes central: agents must infer taste beyond generic search
  14. 47:35 – 50:14

    The hard economics: free leaders vs paid long tail, and the path to $1,000/month agents

    They close on business models and cost structure: many assistants charge despite leading products being free, but the underlying compute/browser-use costs are high. They outline two paths forward—cost deflation and premium products with extreme value—arguing the winning paid agent must be compelling enough that users happily pay large monthly fees.

    • •Landscape snapshot: majority paid or freemium; few fully free
    • •Leaders (Muse/Instinct) are free, making long-tail monetization difficult
    • •Estimated costs can be very high (order of ~$20/user/day for ambitious agents)
    • •Bet: browser use and model costs should deflate quickly, enabling broader free tiers
    • •Long-term opportunity: premium agents users *want* to pay $1,000/month for

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.