Skip to content
a16za16z

Aaron Levie and Steven Sinofsky on the AI-Worker Future

What exactly is an AI agent, and how will agents change the way we work? In this episode, a16z general partners Erik Torenberg and Martin Casado sit down with Aaron Levie (CEO, Box) and Steven Sinofsky (a16z board partner; former Microsoft exec) to unpack one of the hottest debates in AI right now. They cover: - Competing definitions of an “agent,” from background tasks to autonomous interns - Why today’s agents look less like a single AGI and more like networks of specialized sub-agents - The technical challenges of long-running, self-improving systems - How agent-driven workflows could reshape coding, productivity, and enterprise software - What history — from the early PC era to the rise of the internet — tells us about platform shifts like this one The conversation moves from deep technical questions to big-picture implications for founders, enterprises, and the future of work. Timecodes: 0:00 Introduction: The Evolution of AI Agents 0:36 Defining Agency and Autonomy 1:39 Long-Running Agents and Feedback Loops 4:27 Specialization and Task Division in AI 6:04 Anthropomorphizing AI and Economic Impact 9:10 Predictions, Progress, and Platform Shifts 11:31 Recursive Self-Improvement and Technical Challenges 13: 13 Hallucinations, Verification, and Expert Productivity 16:16 The Role of Experts and Tool Adoption 22:14 Changing Workflows: Agents Reshaping Work Patterns 45:55 Division of Labor, Specialization, and New Roles 48:47 Verticalization, Applied AI, and the Future of Agents 54:44 Platform Competition and the Application Layer Resources: Find Aaron on X: https://x.com/levie Find Martin on X: https://x.com/martin_casado Find Steven on X: https://x.com/stevesi Stay Updated: Let us know what you think: https://ratethispodcast.com/a16z Find a16z on Twitter: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Subscribe on your favorite podcast app: https://a16z.simplecast.com/ Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details, please see a16z.com/disclosures.

Erik TorenberghostMartin CasadohostSteven Sinofskyguest
Aug 25, 202556mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 2:22

    From chat UI to background autonomy: what “agentic” really means

    The conversation opens by reframing AI’s end-state: not a back-and-forth chatbot, but background processes doing real work with minimal human intervention. The hosts introduce a practical measure of agency—how much value an agent produces without needing you to step in.

    • Agents as autonomous background tasks rather than conversational interfaces
    • A simple metric for agency: less user intervention per unit of work completed
    • Early skepticism: today’s agents feel like “bad interns,” but are improving
    • The trajectory from interactive prompting to delegated execution
  2. 2:22 – 4:00

    Defining agency: long-running tasks plus feedback loops (and why that’s hard)

    They distinguish “long-running” from true agency. Agency implies the system can generate output, feed it back as input, and continue coherently—an ability constrained by distribution shift and lack of reliable self-reflection.

    • Long-running inference is easier than genuine autonomy
    • Feedback loops (output feeding back as input) are a stronger test of agency
    • Distribution shift makes self-consumption of outputs risky and unstable
    • Near-term reality: agents will need periodic check-ins to avoid wasted work
  3. 4:00 – 5:49

    Orchestrating many specialized agents vs one monolith

    The group argues the ecosystem is moving away from a single monolithic AGI concept toward multi-agent systems. Specialization helps agents avoid getting lost, while orchestration becomes a critical separate capability.

    • Task subdivision reduces failure modes and drift
    • Unix-style philosophy: small tools, narrow scope, composability
    • A “system of agents” needs both deep specialists and an orchestrator layer
    • Human-in-the-loop remains central in most high-performing systems
  4. 5:49 – 7:34

    Stop anthropomorphizing AI: AGI as a term that does ‘infinite work’

    They push back on human-like narratives around AGI and job fears, arguing it obscures economics and practical deployment. The discussion centers on productivity gains, not magical autonomy, and on grounding claims in feasibility and equilibrium effects.

    • Anthropomorphization derails clear thinking and policy discussions
    • AGI labels often substitute for analysis of costs, incentives, and constraints
    • AI can write strong artifacts (e.g., case studies) but lacks situational intent
    • The discourse is becoming more concrete and economically grounded
  5. 7:34 – 11:27

    Predictions and platform shifts: why date-based forecasts fail on exponential curves

    Asked about “AI 2027” style timelines, they argue that years and milestones devolve into metric disputes. Instead, they frame AI as an exponential platform shift—hard to predict but clearly continuing to advance across layers like past compute trends.

    • Skepticism of timeline certainty; forecasting becomes OKRs for an industry
    • Exponential improvement breaks traditional prediction intuition
    • Analogy to prior exponential shifts: storage, bandwidth, connectivity
    • Focus on capabilities and constraints rather than specific calendar dates
  6. 11:27 – 13:12

    Recursive self-improvement: control theory reality vs sci‑fi narratives

    They unpack recursive self-improvement as a feedback-loop claim that sounds decisive but is technically under-specified. Nonlinear control systems can converge, diverge, or asymptote, and without distributional understanding, “self-improve” predicts little.

    • “Box with an arrow back to itself” is not a proof of runaway intelligence
    • Control theory: adaptive feedback loops are notoriously hard to analyze
    • Recursive improvement can plateau; improvement doesn’t imply unbounded takeoff
    • Anthropomorphic leaps turn a technical concept into exaggerated conclusions
  7. 13:12 – 20:09

    Hallucinations, verification, and why experts get disproportionate gains

    The group describes how enterprise attitudes toward hallucinations matured: models improved and teams learned to treat outputs as probabilistic, requiring review. The net effect is that experts—who can evaluate and steer—often see the biggest productivity lift.

    • Two shifts: lower hallucination rates and better organizational literacy
    • AI deployment depends on verification cost vs doing the work from scratch
    • Experts can safely exploit “slot machine” generation for large gains
    • Non-experts may not know what to ask or how to validate outputs
  8. 20:09 – 21:58

    Tool adoption dynamics: prompting, jargon, and why ‘one prompt’ won’t happen

    They argue prompting won’t disappear because context from the user’s head is essential, and long prompts often yield better results. The conversation links this to why formal languages and jargon exist: experts compress intent into efficient shared codes.

    • Prompting persists because intent and constraints must be communicated
    • Long, detailed prompts can significantly improve outputs
    • Formal languages evolved to convey precise meaning efficiently
    • Jargon is domain-expert compression; AI use will mirror that pattern
  9. 21:58 – 28:45

    When tools reshape work: historical parallels from accounting, Word, email, and the web

    Sinofsky traces how early tools mimic old workflows (e.g., printing pre‑formatted expense reports) before flipping the process entirely. They map the same pattern onto AI: enterprises initially graft generative AI onto old “AI-shaped holes,” then reorganize around new behaviors.

    • Early automation often preserves legacy process artifacts before inversion
    • Examples: expense reports → full digital workflows; agendas → email bullets
    • Generative AI is being forced into prior centralized enterprise AI models
    • Adoption is shifting toward individual/tool-level experimentation and reuse
  10. 28:45 – 35:43

    Abdicating logic and shifting abstractions: AI as a deeper platform change than the internet

    They debate whether AI represents a consumption-layer change like the internet or something more: software delegating portions of logic to third-party models. Sinofsky compares it to past abstraction shifts (drivers, clipboard, browser constraints) that reset what apps are built on.

    • AI changes interaction (agents/characters) and also how programs are authored
    • Delegating logic feels new, but parallels exist in surrendering control to platforms
    • Windows printer drivers/clipboard enabled new app ecosystems and disrupted incumbents
    • Web constraints (e.g., ‘Submit’ button era) forced radical UI simplification
  11. 35:43 – 36:06

    Agents dictating workflows: parallelization, serialization, and the rise of PR-level management

    They observe senior engineers running multiple background coding agents and interacting at the pull-request layer. The key idea is workflow parallelization: many tasks were serialized by human bandwidth and tooling, and agents let work proceed concurrently with human gating points.

    • Senior devs manage fleets of agents via PR review rather than direct coding
    • Workflows become less linear as agents handle independent sub-tasks in parallel
    • Human check-ins become explicit gating points to prevent compounded wrong turns
    • Analogy to templated enterprise workflows (e.g., duplicating an ‘Event’ folder)
  12. 36:06 – 45:26

    The counter-AGI pattern: context rot drives narrower tasks, more agents, and more complex prompts

    They highlight an emergent pattern that contradicts “bigger context + higher-level tasks”: large contexts can degrade performance (“context rot”), pushing teams toward partitioning. This yields many specialized agents with scoped READMEs, often mapping to microservices or subdomains.

    • Context windows can become lossy; partitioning improves reliability
    • Agent-per-microservice pattern: scoped ownership + agent-specific documentation
    • Specialization works because base models are strong, enabling composability
    • Trendlines: more agents and more detailed instructions, not fewer
  13. 45:26 – 50:44

    Division of labor and economic outcomes: AI accelerates specialization and creates new roles

    They argue AI will expand specialization rather than collapse it, drawing parallels to the evolution of job functions in computing and construction. New roles emerge (e.g., AI productivity leads), while existing professions may subdivide into finer specialties empowered by tools.

    • Historical pattern: growing capability creates more specialized roles and tooling
    • AI makes individuals more capable, enabling finer task segmentation
    • New organizational roles focused on AI-enabled productivity and workflow design
    • Professional domains (e.g., medicine) likely see increased specialization
  14. 50:44 – 56:04

    Verticalization and the application layer: applied agents, domain data, and platform competition

    The conversation turns to startups vs model providers: the biggest risk was early “ChatGPT ate the wrapper,” but most value now lies in applied, domain-specific workflows and proprietary data. They discuss pre-training’s broad generalization giving way to post-training/RL and vertical execution advantages.

    • Applied AI relies on domain workflows, permissions, and enterprise data access
    • Pre-training was a ‘head fake’—broad generalization is giving way to domain tuning
    • Model providers can’t credibly go deep in dozens of vertical categories
    • Economic reality: a minority of inferences drive most cost—apps optimize what to run

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.