Skip to content
a16za16z

What Would Make an AI Assistant Worth Paying For?

a16z General Partner Anish Acharya sits down with Assistant Benchmark creator David Pawlan to unpack the sudden explosion of personal AI agents and what it will take for one to become part of everyday life. David has been testing dozens of assistants across real-world tasks, from managing email and booking travel to handling financial admin. They discuss why the most useful agents may become increasingly invisible, proactively checking you into flights, finding refunds, filing reimbursements, or simply handling the small tasks that pile up across everyday life. They also explore whether the winning interface is an app, text thread, voice, or wearable; how much autonomy consumers will actually give their agents; and what happens when agents start interacting with other agents. From commerce and restaurant reservations to entirely new agent-native services, Anish and David ask what the internet looks like when software starts acting on our behalf. Timestamps: 00:00 - Intro 00:54 - Poke, Instinct, Muse and the agent boom 03:41 - What Assistant Bench actually tests 09:14 - Cost savers beat time savers 14:53 - Muse charm as Meta's data play 20:03 - Silent agents in group chats 26:01 - Proactivity is the real moat 32:21 - Assistant vs agent, defined 40:22 - Amazon blocks Muse, Shopify opens the door 47:35 - The $20/day agent economics Resources: Follow David Pawlan on LinkedIn: https://www.linkedin.com/in/david-pawlan/ Follow Anish Acharya on X: https://x.com/illscience Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

David PawlanguestAnish Acharyahost
Sep 29, 202650mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Paid AI assistants win by saving money, not minutes

  1. The conversation maps the sudden boom in consumer AI assistants/agents (Poke, Instinct, Muse, ChatGPT Voice) and argues the category is still in a “Wild West” phase with unclear winners.
  2. David Pawlan explains Assistant Bench, a consumer-oriented benchmark that tests assistants on identical practical prompts and scores outcomes across multiple dimensions to help users choose what works.
  3. They predict the strongest adoption drivers will be cost-saving, “feels like free money” workflows (refunds, reimbursements, bill reduction) more than minor productivity boosts.
  4. They debate interaction surfaces—iMessage vs dedicated apps vs hardware—highlighting voice + connectors as a major breakthrough for hands-busy moments and hardware (e.g., Meta’s Muse Charm) as both interface and potential data-collection strategy.
  5. They argue proactivity is the core defensible moat, but it requires careful trust boundaries, and they anticipate new agent-to-agent infrastructure plus major disruption to commerce and marketplaces as agents become the new ‘front door.’

IDEAS WORTH REMEMBERING

5 ideas

“Worth paying for” assistants will feel invisible, not like another app to manage.

They argue most people don’t value marginal productivity gains; they value having annoying life-admin disappear. The strongest early “killer apps” are workflows where the agent quietly saves money or effort (HSA reimbursements, fare-drop credits, smart sprinklers) and reports back after the fact.

A consumer-facing benchmark is emerging because the market is too noisy to trial manually.

Assistant Bench compares assistants by giving them identical real-world prompts (e.g., book flights, find nearby restaurants) and grading one-shot outcomes, follow-ups, speed, and performance across 16 dimensions. It’s positioned as a consumer decision aid rather than a deep technical eval, and its traction signals widespread confusion about what actually works.

Cost-saving workflows may beat pure time-saving as the mainstream wedge.

They distinguish “time savers” from “cost savers,” predicting broader adoption from experiences that feel like “free money.” Examples include retroactive reimbursements, price-drop monitoring, and automations that reduce recurring bills—value that’s tangible and easy to understand.

Voice + connectors creates “hands-busy” utility that messaging/app UIs can’t match.

They see voice as a breakout interaction mode because it fits moments when people are busiest (biking, cooking, gardening). A highlighted ‘magic moment’ is using ChatGPT Voice with Gmail/Calendar connectors to triage email and schedule meetings hands-free to reach inbox zero on a commute.

Proactivity is the moat, but trust boundaries are the product.

They expect the biggest moat to be proactivity—agents acting without explicit prompts—while warning that one overstep can permanently damage trust. A useful framing is whether the action requires user behavior change/commitment (needs permission) versus purely additive outcomes (drafting, refunds/credits).

WORDS WORTH SAVING

5 quotes

I think the general population does not care about being ten percent more efficient.

— David Pawlan

And so I think these like invisible type agents that are gonna be doing these cost-saving type workflows, um, fall within that finance category, but, but they're really gonna lean into these like cost-saving activities-

— David Pawlan

But it will be a very fine line because if you, if you cross that line once, I think you immediately, you lose all trust with your user.

— David Pawlan

We were joking internally, like we're days away from an agent messaging someone saying, "I noticed you weren't that into her, so I went ahead and broke up with her for you."

— Anish Acharya

I have a hot take thesis that this, uh, Muse Charm is actually less about trying to win the hardware game, and it's more about data collection in the real world for the Meta team in terms of like how the real world actually functions, given it has a camera and multiple microphones, and it's primary... I don't really see it being a mass adoption thing.

— David Pawlan

Assistant Bench methodologyConsumer agent market landscape (horizontal vs vertical)Cost-saving automation vs productivityVoice assistants with Gmail/Calendar connectorsHardware surfaces and ambient capture (Muse Charm)Group chat dynamics and “silent” agentsProactivity, trust, and permissioning models (assistant vs agent)

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.