Skip to content
ClaudeClaude

Building secure agents for knowledge work

Katelyn Lesse, Head of Platform Engineering at Anthropic sat down with Dan Shipper, Co-Founder and CEO at Every and Willie Williams, Head of Platform at Every to talk about the Every Agent, the AI coworker they built on Claude Managed Agents. They discuss how knowledge work is changing, testing new AI models on their own day-to-day work, agent security, and why they built on Claude Managed Agents. Learn more about Claude Managed Agents: https://platform.claude.com/docs/en/managed-agents/overview

Katelyn LessehostDan ShipperguestWillie Williamsguest
Oct 6, 202633mWatch on YouTube ↗

CHAPTERS

  1. 0:07 – 2:18

    How AI is reshaping workflows: from benchmarks to “extending” knowledge workers

    Katelyn and Dan frame the core problem: work is changing fast, and the only way to adapt is to rebuild workflows by actually using models in real work. Dan contrasts benchmark progress (fear of job loss) with the more practical goal of amplifying human capability and influence.

    • •AI impact is best understood by experimenting in real workflows, not theorizing
    • •Benchmarks often map to job-loss anxiety, but don’t capture how tools extend workers
    • •Engineering adopted agentic workflows first due to verifiable feedback loops
    • •Knowledge work lags because quality is harder to measure and norms differ from software
  2. 2:18 – 2:47

    What Every is: media + tools to keep companies at the edge of AI

    Dan explains Every’s dual identity as a media and technology company. They teach people how to work with new models and also build tools that operationalize those practices inside organizations.

    • •Every positions itself as an “edge of AI” subscription: education + tooling
    • •Education includes newsletters, courses, and hands-on model “vibe checks”
    • •Tooling includes an in-company agent that brings Every’s workflows to teams
  3. 2:47 – 4:33

    Why build the Every Agent: spreading AI workflows across a company in public

    Dan introduces the Every Agent as a Slack-based AI coworker intended to “AI-pill” an entire company. The key insight is that prompting in public (in Slack) helps best practices spread and supports internal champions who otherwise struggle to convince teammates.

    • •Early internal experiments (e.g., OpenClaude setups) were exciting but hard to maintain
    • •A single shared company agent helped spread workflows more effectively than many siloed setups
    • •“AI ambassadors” need a way to show value rather than repeatedly explain it
    • •Slack makes prompting visible, enabling shared learning and adoption
  4. 4:33 – 6:26

    High-leverage ‘small tasks’ example: automating open enrollment guidance

    Dan emphasizes that the biggest value often comes from eliminating frequent, time-consuming micro-tasks. They used the Every Agent to guide employees through benefits open enrollment by combining plan knowledge with employee context.

    • •Open enrollment is information-dense and cognitively costly for employees
    • •Centralizing the workflow avoids repeated one-off copy/paste into private LLM chats
    • •Company-authored guidance builds confidence and consistency in decisions
    • •Agent skills can be created once by leaders and disseminated org-wide
  5. 6:26 – 8:27

    Editorial automation that helps experts scale: “Kate Bench” top edits

    Willie and Dan describe “Kate Bench,” a workflow that captures their editor-in-chief’s editing patterns and makes them available as a reusable skill. The agent can suggest edits directly in Google Docs, while Kate reviews and the system learns from deltas.

    • •Rote editorial tasks can be routed to an agent via Slack instead of flooding inboxes
    • •A skill was built from ~30,000 historical edits to replicate editorial “taste”
    • •Agent proposes Google Doc suggested edits; humans accept/reject and supervise quality
    • •Tracking additional human edits feeds improvements back into the skill over time
  6. 8:27 – 9:47

    Reframing automation: compounding an expert’s impact instead of replacing them

    Dan argues the real opportunity is shifting experts from doing every request to building systems that embed their judgment. As organizations scale, agents can propagate expert taste without requiring proportional increases in expert time.

    • •Automation becomes a way to scale expert judgment across a growing organization
    • •Agents reduce repetitive interruptions while preserving expert oversight
    • •The value accrues twice: experts get time back, and others gain access to expertise
    • •Continuous improvement loops turn “one-off help” into an evolving system
  7. 9:47 – 11:31

    From vibe checks to personalized evals: measuring AI by your real work

    They explain why generic benchmarks are insufficient for knowledge work and how Every operationalizes “taste-based” evaluation. Manual vibe checks evolve into codified checks—like unit tests—that measure model performance on tasks Every actually cares about.

    • •The meaningful question is: ‘How good is it at the work I do?’ not generic scores
    • •Manual model testing (“vibe checks”) is useful but doesn’t scale
    • •Work examples can be captured and converted into repeatable checks/rules
    • •This creates a benchmark-like system that reflects personal/org preferences
  8. 11:31 – 13:02

    Why Slack won: consolidating into one shared agent that everyone improves

    Willie describes their shift from “one agent per person” to a single shared agent. The shared approach concentrates investment and learning so improvements propagate across the team, and Slack becomes the natural surface for collaborative work.

    • •Multiple personal agents led to isolated gains that didn’t spread to teammates
    • •A single agent improves faster when many people contribute skills and guidance
    • •Slack reduces copy/paste overhead and keeps humans+agent in the same workspace
    • •Collaborative visibility makes adoption and norms easier to establish
  9. 13:02 – 16:37

    Building on Claude Managed Agents: avoiding infra distractions and gaining primitives

    Willie details why they chose Claude Managed Agents after painful lessons running their own infrastructure. Managed primitives like sandboxes, memory, and session control freed the team to focus on product design and experimentation.

    • •Self-managed infra became unwieldy when trying to scale an ‘agent for everyone’
    • •Claude Managed Agents provides sandboxes, memory, and session control out of the box
    • •The team can focus on interaction design rather than orchestration and reliability
    • •Faster experimentation: clone state, test edge cases, and manage token costs
  10. 16:37 – 18:39

    Security, isolation, and fast model updates: harness and model are now coupled

    They discuss security as an engineering and product concern, emphasizing isolation and controlled access. Dan adds that modern agent performance depends on integrated “computer use” and harness capabilities, and managed systems simplify breaking model changes.

    • •Built-in isolation helps control tool calls, sandbox boundaries, and memory access
    • •Session management supports nuanced access control per user/request
    • •Integrated harness + model improves performance vs. stitching together components
    • •Managed Agents can absorb breaking prompting changes via simple upgrades
  11. 18:39 – 20:09

    Identity & authorization model: separate coworker in Slack, acting on your behalf for tools

    Willie explains the nuanced identity approach: the agent is its own entity conversationally, but typically uses the user’s permissions when taking actions via tool calls. This mirrors how people expect a coworker to behave while preserving access boundaries.

    • •Users naturally treat in-company agents like humans and ask for gray-area actions
    • •In Slack, the agent is a distinct entity with its own isolated memory structures
    • •For tool calls (docs/lookups), it generally acts on the requesting user’s behalf
    • •The boundary between internal conversation and external action is key to trust
  12. 20:09 – 23:14

    Continuous improvement through ‘taste nudges’ + the hiring analogy for agents

    They describe how day-to-day corrections (“that joke was bad”) become training signals to refine skills. Dan likens benchmarks to SAT scores—useful but insufficient—and argues companies need work trials and reference-check-like evaluations to make agents fit local norms.

    • •Small corrections accumulate into a usable representation of personal/organizational taste
    • •Skills improve from many examples (edits, emails, decks), not a single spec
    • •Agents should get better from day 1 to day 10 by compounding experience
    • •Work-trial evals beat generic benchmarks for selecting and refining model behavior
  13. 23:14 – 27:01

    What’s next: personal benchmarks, computer-use experiments, and social interaction design

    Dan and Willie outline upcoming directions: a compounding system where anyone can create a personal benchmark and optimize skills that spread across the org. They also explore computer-use agents and the emerging challenge of social norms and personality in shared channels.

    • •Build ‘personal benchmarks’ to compare models and optimize skills over time
    • •Experiments like ‘Hands’ enable the agent to operate a user’s computer when needed
    • •Public vs. private work modes should be seamless (team Slack vs. solo workflows)
    • •Agents in group chats need new norms: when to speak, how to interject, introvert/extrovert styles
  14. 27:01 – 30:51

    Looking ahead + asks from Anthropic: fidelity, usability for knowledge workers, and personality knobs

    Dan predicts knowledge work will adopt manager-like oversight loops once agents have “fidelity” to user intent and taste. They ask Anthropic to focus not just on frontier intelligence, but also comprehension, speed/cost, and better personality/configurability tooling.

    • •Future value comes from high ‘fidelity’ to what users want, not generic competence
    • •Knowledge workers will shift from doing tasks to managing systems and interventions
    • •Requests to Anthropic: usability at the user’s level, clearer responses, speed/cost improvements
    • •Need better controls for consistent yet configurable agent personality
  15. 30:51 – 33:46

    Hot take and ‘magic moments’: more agents can mean more work—and more delight

    Dan argues that as agents reduce friction, teams often discover more work worth doing rather than running out of tasks. They close with playful examples of delight: the agent tracking social debts (beers owed) and generating announcements in a colleague’s signature voice.

    • •Using agents changes the kind of work; it doesn’t necessarily eliminate it
    • •Automation can enable higher-quality output that still feels personal
    • •Agent delight comes from context-aware social behaviors (e.g., reminding about beer debts)
    • •Voice-based skills can capture organizational culture (e.g., announcements in a coworker’s style)

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.