Skip to content
ClaudeClaude

Building secure agents for knowledge work

Katelyn Lesse, Head of Platform Engineering at Anthropic sat down with Dan Shipper, Co-Founder and CEO at Every and Willie Williams, Head of Platform at Every to talk about the Every Agent, the AI coworker they built on Claude Managed Agents. They discuss how knowledge work is changing, testing new AI models on their own day-to-day work, agent security, and why they built on Claude Managed Agents. Learn more about Claude Managed Agents: https://platform.claude.com/docs/en/managed-agents/overview

Katelyn LessehostDan ShipperguestWillie Williamsguest
Oct 6, 202633mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Every’s Slack agent scales expert taste with secure managed infrastructure

  1. Every describes building a Slack-based AI coworker (“Every Agent”) to make AI workflows visible, shareable, and scalable across an organization.
  2. They argue generic benchmarks are becoming less informative for knowledge work and propose personalized evals based on real tasks and “taste,” akin to work trials and reference checks.
  3. The product’s core value is compounding: capturing repeated workflows and expert standards (like an editor’s “Kate Pass”) into reusable skills that improve through tracked human edits.
  4. They chose Claude Managed Agents to avoid becoming an infrastructure team, leveraging built-in sandboxes, memory, isolation, and session controls for faster experimentation and safer deployment.
  5. They see the next frontier as agents with higher fidelity, better social norms in group contexts (when to interject), and configurable personality that matches a company’s working style.

IDEAS WORTH REMEMBERING

5 ideas

A single shared company agent compounds value faster than one-agent-per-person.

Every found that when many people contribute improvements (skills, guidance, examples) to a shared agent, the quality increases faster and the benefits spread across the org—unlike isolated “personal agents” where gains stay siloed.

Benchmarks matter less than work-specific “taste” and real-task evaluation.

They emphasize that generic model benchmarks are like SAT scores: useful as a rough filter, but insufficient once models are broadly “smart enough.” Real adoption depends on how well the agent performs your actual tasks, to your standards.

Capture expert behavior once, then scale it via skills that improve from feedback loops.

By turning an expert’s historical work into a reusable skill (e.g., 30,000 editor-in-chief edits), the org can apply that expertise on demand, while tracking what the human still changes and feeding it back to improve the skill.

Slack as the interface enables “show, don’t tell” AI adoption.

Running the agent inside Slack makes prompting visible, allowing colleagues to learn workflows by watching and reusing them—turning AI adoption into a social, collaborative process rather than an individual productivity hack.

Managed agent infrastructure reduces operational drag and speeds iteration.

Managed agent primitives—sandboxes, memory, session control, isolation—let the team focus on interaction design and workflow experiments instead of building and maintaining complex infrastructure.

WORDS WORTH SAVING

5 quotes

What ends up happening is you have to reinvent your workflow from scratch, and the only way to do that is to, like, try it on the problems that you're, you're actually working on.

— Dan Shipper

The beauty of a Slack agent, the Every Agent in particular, is that you can show instead of tell... how to use this stuff to actually do really, really great work because you're doing it in Slack. You're prompting in public instead of in private.

— Dan Shipper

I just, uh, took like 30,000 of her historical edits and turned it into a skill that is in the Every Agent.

— Dan Shipper

If you're hiring a job candidate and you're looking at their SAT scores, that's a little bit like a benchmark score.

— Dan Shipper

What we've found is the more we use agents, the more work there is to do.

— Dan Shipper

Slack-native AI coworkerShared agent vs personal agentsSkills built from expert history (Kate Bench)Workflow dissemination and “prompting in public”Personalized evals/vibe checks vs generic benchmarksClaude Managed Agents primitives (sandbox, memory, sessions)Security, isolation, auth, and identity modeling

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.