Skip to content
ClaudeClaude

Building secure agents for knowledge work

Katelyn Lesse, Head of Platform Engineering at Anthropic sat down with Dan Shipper, Co-Founder and CEO at Every and Willie Williams, Head of Platform at Every to talk about the Every Agent, the AI coworker they built on Claude Managed Agents. They discuss how knowledge work is changing, testing new AI models on their own day-to-day work, agent security, and why they built on Claude Managed Agents. Learn more about Claude Managed Agents: https://platform.claude.com/docs/en/managed-agents/overview

Katelyn LessehostDan ShipperguestWillie Williamsguest
Oct 6, 202633mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:00 – 0:07

    Intro

    1. KL

      [upbeat music]

  2. 0:07 – 2:18

    How AI is reshaping workflows: from benchmarks to “extending” knowledge workers

    1. KL

      Super excited to be chatting with you guys. I have said this to you before, but I think you guys are some of the most AGI-pilled people out there with the things that you build and the way that you talk about what's happening in AI, and so very exciting for me to get to spend some time chatting with you guys about what's going on and what you're building. So maybe actually just to start, I would love to talk about the fact that the way everybody is doing work is changing so much right now, and that means the tools that people need to be able to use to get work done are changing a lot. Um, how do you guys think about that?

    2. DS

      We think about it a lot. Um, and it's, it's one of those, one of those questions that I, I-- it's on everybody's mind. Everybody is super afraid of it, or they're like, "Oh, it's gonna be a utopia, and, like, no one's gonna have to work at all." The space we try to occupy is, "Well, why don't we just use the models and try to figure it out?" What ends up happening is you have to reinvent your workflow from scratch, and the only way to do that is to, like, try it on the problems that you're, you're actually working on. And I think for us, we've been seeing over the last couple years, you, you, like, watch all the benchmark progress, and when benchmark progress happens, you're like, the benchmarks are sort of measuring is it, is it gonna take my job away?

    3. KL

      Yeah.

    4. DS

      Like, the, the higher it gets on a benchmark. And our perspective is if you use these tools well, they, uh, are good at extending you and your power and your influence, and there needs to be ways to both measure that and to help people understand, uh, how to do that, because yes, if you look at benchmarks, they can do pre-AI tasks, like, at or better than humans in certain cases. But if you're thinking about how do I use this for my work-

    5. KL

      Mm-hmm

    6. DS

      ... that, like, opens up a whole se-set of options. We start to see that in engineering first because there's verifiable rewards in programming, and so y- now you have, like, 10X, 100X engineers who are orchestrating teams of agents to pull forward their roadmap by, like, two years.

    7. KL

      Yeah.

    8. DS

      And I think that is coming to the rest of knowledge work, but, uh, it's taken longer, uh, A, because it's, it's just harder to know whether or not it's doing a good job.

    9. KL

      Mm-hmm.

    10. DS

      And B, knowledge work, it's a different way of thinking that maybe developers are a l- a little bit more used to.

    11. KL

      Yeah.

    12. DS

      Um, but we think that's coming. We're trying to, we're trying to, uh, both discover it and disseminate it so everyone can work that way.

  3. 2:18 – 2:47

    What Every is: media + tools to keep companies at the edge of AI

    1. KL

      So for people who don't know, I'm curious, uh, what is Every? Tell us about what you guys do.

    2. DS

      Every is a media and technology company. We build the only subscription you need to stay at the edge of AI. So for our subscribers, we, uh, do a series of education and equipment that helps them work at the edge. Uh, on the education side, we have everything from a newsletter to courses. When new models come out, we do vibe checks. And on the equipment side, now we have an agent that allows you to work in the way that we do in public with your company.

  4. 2:47 – 4:33

    Why build the Every Agent: spreading AI workflows across a company in public

    1. KL

      Okay, so, um, you guys built an AI coworker, um, to help solve this problem. Tell us about it. What were your goals? How's it going?

    2. DS

      So we built the Every Agent. It is an AI coworker. It's an agent that is the easiest way to AI-pill your whole company. We started basically, like, in the OpenClaude craze. We started using OpenClaude, and we were like, "Oh my God, this is the coolest thing," and it spread throughout the whole company. And then-

    3. KL

      You had your Mac Minis.

    4. DS

      We had all the Mac Minis.

    5. KL

      Yeah.

    6. DS

      Um, and then unfortunately, all the OpenClaudes died off because, like, they're just really hard to maintain.

    7. KL

      Yeah.

    8. DS

      And we started just using one company agent, and we were like, "Wow, this is actually really, really great," and, and noticed how easy it was to spread workflows that we discovered throughout the whole company because our job inside of Every is to work at the edge, but you'd be surprised how hard that is to make working at the edge even across the entire company.

    9. KL

      Yeah.

    10. DS

      And so that started to happen. We started to see these workflows start to spread. What we've also realized is there's a certain type of person that loves Every who we've been calling an AI ambassador-

    11. KL

      Mm-hmm

    12. DS

      ... who is, like, uh, has that feeling about AI where they're like, "Oh my God, I have this Codex set up or this Claude set up or whatever that is, uh, absolutely 10X-ing the way that I do my work, and I can't imagine working without it. But none of my coworkers have any idea what this is, and I'm trying to, like, explain to them all the time how they should use it, but they're, like, tuning me out." And the beauty of a Slack agent, the Every Agent in particular, is that you can show instead of tell-

    13. KL

      Yeah

    14. DS

      ... how to use this stuff to actually do really, really great work because you're doing it in Slack. You're prompting in public instead of in private. And so we built the Every Agent to have all the things you would expect in an, in an agentic coworker. You know, it can help you do recurring tasks like gather reports or do social posts or-- and any of the kinds of things you might expect, but it happens inside of your company Slack-

    15. KL

      Yeah

    16. DS

      ... so that your coworkers just automatically pick it up, and it starts to spread your best practices. We try to empower the ambassadors as much as we can with a tool like this.

  5. 4:33 – 6:26

    High-leverage ‘small tasks’ example: automating open enrollment guidance

    1. KL

      Awesome. What's one of your favorite examples of something you've seen somebody do with the Every Agent that maybe, maybe they could do it before, it would've taken them longer, maybe they couldn't have even really accomplished that thing before?

    2. DS

      I really think that having an-- having the Every Agent, having an agent in your Slack, there's-- I mean, there's the sort of like crazy top-end stuff, but the, the stuff that I wanna focus on is actually there's, there are all these small little things that you're doing day-to-day that take up a lot of people's time and, and having a company agent makes really easy. So a simple example, we just had open enrollment. We did the entire open enrollment process through an agent, uh, through the Every Agent. The open enrollment, I-I'm sure, I'm sure you know, you get like a PDF that's like, "Here are all these plans," and, like-

    3. KL

      Yeah

    4. DS

      ... doing the algebra to figure out the plans is actually very hard, and as part of that process, the Every Agent both knew all the plans and knew who you were and, um, and helped you kind of figure out which one you should select for your particular situation, and that's something that would've been very hard to do and coordinate without a central company agent, and the Every Agent makes it really easy for, A, leaders of the company to create a set of skills that, that work in that way, and then B, get that disseminated to everyone in the org and make sure that the open enrollment process actually happens smoothly.

    5. KL

      Yeah, like the amount of time you probably saved with not every single person pasting separately into Claude-

    6. DS

      Oh my God

    7. KL

      ... like, "This is the PDF."

    8. DS

      Yeah.

    9. KL

      "What do I pick?"

    10. DS

      Yeah.

    11. KL

      That saves you a lot of time.

    12. DS

      And, and honestly, like, that's some- even for me, like, I-- as I-- anytime I get one of those things, I just, like, call my dad, and he's like-

    13. KL

      [chuckles]

    14. DS

      ... "I can't believe you're, like, thirty-five. I can't believe I have to still explain this to you."

    15. KL

      [chuckles]

    16. DS

      'Cause it is actually really complicated, and yes, I could just individually paste it into Claude, but, but knowing that this is coming from the company and the company has thought about for each person, like, "Here's what you might wanna pick," it really helps me feel confident in my choice.

    17. KL

      Yeah, now you can call your dad for more fun reasons.

    18. DS

      Yes, exactly. Sorry, Dad, I automated that task. [laughing] Um, it's not one that, not one that he relishes.

    19. KL

      He didn't like it at all. [laughing]

    20. DS

      [laughing]

  6. 6:26 – 8:27

    Editorial automation that helps experts scale: “Kate Bench” top edits

    1. WW

      Well, we also, like, we have a lot of pretty rote tasks that go along pretty with the editorial arm. And we've spent a lot of time, like, building out something we call Kate Bench, which is named after our editor-in-chief, which is mostly like how do we do top edits before things go out? And there's still a question of like when to trigger these top edits, and so it's like we flood Kate's inbox. It's now much easier to just be- to tag the agent in Slack and say like, "Hey, just review this." And the nice part is Kate can see that-

    2. KL

      Mm-hmm

    3. WW

      ... and she can kind of like judge like, okay, is the agent, like, doing the job or not?

    4. KL

      Yeah.

    5. DS

      Yeah, this is, this is a really important thing that I think, I think Kate Bench gives or gives you a, a, an, a view into how this stuff is gonna be used in knowledge work going forward. So l- like Willie said, Kate's our editor-in-chief. Like, I've been trying to automate her for many years-

    6. KL

      [laughs]

    7. DS

      ... but like, but, like, in the most loving way possible.

    8. KL

      I was gonna say-

    9. DS

      Um-

    10. KL

      ... maybe don't have her watch. [laughs]

    11. DS

      Yeah, like, like, uh, automating I think sounds scary-

    12. KL

      Yeah

    13. DS

      ... but when you actually see it in practice, you're like, "Oh, this is actually helpful for her"-

    14. KL

      Yeah, for sure

    15. DS

      ... because we publish so much. Um, and when, when Kate started, there was like four people at Every. She was like the fourth person, and now there's, now there's 30 people. Um, and she has incredible taste for copy edits, for like making sure that every single period, every single verb, every single thing is in the right place. But, um, we have 10 times the amount of things going out today than we did when she started, but she still has to do it all.

    16. KL

      Yeah.

    17. DS

      So what I did was I just, uh, took like 30,000 of her historical edits and turned it into a skill that is in the Every Agent. So now every time we are publishing something, whether it's a piece or a landing page or an email, anyone in the organization can just say like, "@Every," like, "do a Kate Pass on this," and it will take her taste and apply it, like literally make edits in the Google Doc, like with suggested changes as her, and then, uh, they can accept and reject it. And what she does is she just goes in and looks at it, and it has most of the things that she would do and even catches things she might not. She has to spend less time on each piece 'cause-

    18. KL

      Yeah

    19. DS

      ... it has done a lot of the work. And then what we do is we track what is the extra work that she added to this-

    20. KL

      Mm

    21. DS

      ... and then we compound it back into the skill so it gets better and better over time.

    22. KL

      Oh, nice.

  7. 8:27 – 9:47

    Reframing automation: compounding an expert’s impact instead of replacing them

    1. DS

      And she gets to see a graph of here's how, here's how much time it's saving you, and here's also a rule we might wanna add to this so that it catches more things you don't, you didn't necessarily know about. And I think what that does is it, it flips the sort of like automation debate from, like, is it gonna take away my job to actually I'm an expert in an organization. My time's taken up all day by having people ask me to do things for them that only I can do.

    2. KL

      Yeah.

    3. DS

      And as the organization scales, in order to scale my impact, previously I would've had to spend more time. And if you have an agent that compounds like the Every Agent, the opportunity for experts inside of organizations is you can spread your impact and spread your, the way that you do work without having to spend more time because you're working on a system that does the work for you instead of doing all the work yourself.

    4. KL

      Yeah, it's doubly valuable. Kate gets more time to do other things, um, maybe the hardest copy edits-

    5. DS

      [laughs]

    6. KL

      ... so Kate, and then everybody else gets to take advantage. I actually would love to talk about, um, Kate Bench-

    7. DS

      Mm

    8. KL

      ... for an ex- for an example. You guys talk a lot about being really, um, you know, bench pilled, right?

    9. DS

      Mm.

    10. KL

      And being very on top of here's how we do some, like, really cool use case, and then measure how well we're performing on that really cool use case. Tell us more about that mentality-

    11. DS

      Yeah

    12. KL

      ... and how, like, as you have people going out there and starting to use the Every Agent, they can kind of bring this mentality in on, like, continuous improvement.

  8. 9:47 – 11:31

    From vibe checks to personalized evals: measuring AI by your real work

    1. DS

      Yes. We, we have been- really started to get, to get good at this, get, get good at evals and start trying to measure these things. I think, you know, we started with everyone who's in AI and, and follows us closely knows that that doesn't really mean that much.

    2. KL

      Yeah.

    3. DS

      Um, and the question is not how good is it at a generic benchmark, it is how good is it at the kind of work I do?

    4. KL

      Right.

    5. DS

      And that's a whole- that's a totally different way of thinking about, um, whether AI is good because it's much more personal, and that's where we started with our vibe checks, where before a new model comes out, we get access to it early, and we actually test it by hand across a lot of our workflows, and then on the day it launches, we can tell you, "Here's the kind of work we do. Here's our taste. Here's what we think of this model." And people rely on that to be like, "Well, I trust Every and I trust their taste," or like, "There's a certain writer at Every that I know has similar taste to me, so like, if they like this model, I'll try it." Problem is that that's very manual and takes a long time, and so what we've started to do is capture those tasks that we do. So like anytime one of us is doing work with a model day to day and we, we notice, ooh, it did something I like, or ooh, it did something I didn't like, we can sort of capture that in, in one place, and then talk to this platform that we've built that will eventually be part of the Every Agent, talk to this platform that we built about why we liked it and why we didn't, and that turns it into this set of checks, which are basically like unit tests or rules for here's the kind of thing you look for. So for Kate Pass or Kate Bench, it would be is the headline sentence case would be an example. And, um, what, what that allows us to do is run many different models on the same scenario and then be, be able to see for our real work, did this model do a good job or not according to my taste? We've started to be able to add this sort of like quantified layer to vibe checks that I, I think looks like a benchmark, but actually measures things that you care about.

  9. 11:31 – 13:02

    Why Slack won: consolidating into one shared agent that everyone improves

    1. KL

      Awesome. Um, going back to form factor a bit, you talked about obviously being in Slack is really powerful. All the knowledge workers can have really easy access to this. Um, how did you land on Slack? How are you thinking about that form factor?

    2. WW

      Well, we s- we actually started with sort of a model of, uh, every one agent per person, you know, available in whatever surface you really wanted, uh, that agent to be there, if it was Telegram, if it was Slack. Um, and part of the reason we over time went from like a multi-agent setup to really a singular agent setup is we found that agents got better via the same dynamic the more you invested in them.

    3. KL

      Yep.

    4. WW

      And so you had a lot of folks who were like, "Oh, I invest- I'm investing in my agent," but those gains weren't really like spreading out to the rest of the team. And the places where we, we really saw like agents being able to do work, real work, was when there were multiple people investing in them, making better skills, giving it better guidance, and it just made so much more sense for everyone to be investing in like a single agent rather than, uh, everyone having their own agent sort of spread around. And Slack was like- The very natural form factor for like where are we when we're having these conversations with individual agents and, and when we're doing, like, collaborative work. There's a lot of copy and paste.

    5. KL

      Mm-hmm.

    6. WW

      Right? I'm like, I'm like, we're gonna have a conversation. I'm gonna, like, copy this over to my, like, Claude coworker, and I'm gonna copy the, the response back, and then you're gonna do that. And instead it's much easier to just have that conversation in the same place that the, your coworkers are and the same place your agent is.

    7. KL

      Perfect and easily accessible. Let's talk a bit about how you built Everyagent.

    8. WW

      Okay.

  10. 13:02 – 16:37

    Building on Claude Managed Agents: avoiding infra distractions and gaining primitives

    1. KL

      So you built this product on Claude Managed Agents-

    2. WW

      Yeah

    3. KL

      ... uh, which we're really excited about. Tell us just starting out, what were some of the main reasons why you chose to start building on Claude Managed Agents?

    4. WW

      Well, we learned a lot of pain- painful lessons with this first version of One Agent for Everyone.

    5. KL

      Okay.

    6. WW

      Um, one of the main ones is, uh, infrastructure's hard, you know? Um, and a lot of what we want to focus on is like the future of work and like how to make a better product for like that future that we see coming. We don't necessarily want to spend that time like, how do we orchestrate sandboxes?

    7. KL

      Yep.

    8. WW

      Right? And so we, [chuckles] we had this very painful arc where we're like, "Okay, we're going to, uh, we're gonna build one agent for everyone," and then it just got unwieldy very quickly because you're sort of like trying to improve a product and have a, have a strong, durable like infra layer underneath. And so w- we took those lessons and turned them into the Everyagent, and one of our first things was we don't want to be an infra team.

    9. KL

      Mm-hmm.

    10. WW

      Right? We want to focus our energy in the place where we think, uh, the future is going. And so Claude Managed Agents was great. It wa- it was, it gave us, you know, it gave us sandboxes, it gave us memory, it gave us session control, all these primitives that really allowed us to, um, focus on the next layer of how do we do interaction design for, for an agent. You know, we can kind of just like shuffle it off.

    11. KL

      Yeah.

    12. WW

      Just be like, okay, this is being taken care of. You know, I'm not getting alerted w- uh, when, when orchestrating.

    13. KL

      I-

    14. WW

      Yeah, yeah. [laughs]

    15. KL

      [laughs]

    16. DS

      Thank you. Thank you for your service.

    17. WW

      Yeah, yeah, yeah, yeah. [laughs]

    18. DS

      [laughs]

    19. WW

      Uh, and it's been great.

    20. KL

      Nice.

    21. DS

      Yeah. That's, that's one of those like classic things in startups where you're like, "Okay, we're gonna build an agent platform. Should we do the infra layer or not?" And you're like, "Oh, infra's not that hard anymore. Like, you can just... Like, we'll just have Claude spin up a bunch of servers and like manage the servers for us and like blah, blah, blah." I was saying that at least, and Willie was like, "No, no, it's, it's hard."

    22. WW

      [laughs] Yeah.

    23. DS

      And I was like, "We're gonna, we're, we're going to try it." And then a month later I was like-

    24. WW

      Yeah.

    25. DS

      Yeah.

    26. WW

      You're like-

    27. DS

      [laughs] This is hard.

    28. KL

      [laughs]

    29. DS

      And yeah, Claude Managed Agents, it takes something that you don't know is hard, but when you do it, you're like, "This is so complicated," and it makes it very easy for us.

    30. KL

      Nice. That's awesome. So there's infra management, that's obviously helpful. What are some of the other features? You mentioned memory.

  11. 16:37 – 18:39

    Security, isolation, and fast model updates: harness and model are now coupled

    1. KL

      Yeah, that makes a lot of sense. And, uh, there's a lot going on right now in terms of security-

    2. WW

      Yeah

    3. KL

      ... with agents and, um, a lot of energy is going into how do we make sure that agents are doing the things we want them to do-

    4. WW

      Yeah

    5. KL

      ... and not the things that we don't want them to do.

    6. WW

      Yeah.

    7. KL

      How do Claude Managed Agents help you think about security with Everyagent?

    8. WW

      Yeah. A, a lot of it is just the built-in isolation. It, it makes it very easy to, um, isolate out, uh, the sandbox, the tool calls, like what's, what's happening agent in the loop, um, and what are the environments and where are the memories and like we can, we can, because we can control all that, we can make choices around like, uh, who accesses what when. You know, session management is great for this, and that really gives us the, the nuanced tools to say like, okay, you know, if Dan requests something, does he get, uh, the same thing that I would get if I requested it?

    9. KL

      Mm-hmm.

    10. WW

      Right? And underneath the hood there's, there's a lot of playing with the primitives of like what memories do we attach to that request? Like is the, are the sandboxes the same or different? Um, and so yeah, just having those primitives, uh, is just helpful from an engineering standpoint.

    11. DS

      One thing that I do want to add, too, I think this relates to s- security, but it also relates to performance, which is there's not a clear separation between the model and the harness anymore.

    12. KL

      Mm-hmm.

    13. WW

      Mm-hmm.

    14. DS

      And a model that doesn't come with a computer that it knows how to use, like it's just not as performant.

    15. KL

      Yeah.

    16. DS

      And having all of that built in one place by one provider, uh, ensures that the model knows how to use that stuff really well, and you know, uh, the, the day we're filming this, there's a new, a new, new model launch from you guys.

    17. WW

      Yeah.

    18. DS

      And it's one of those things Willie, Willie was telling me, if you look at the prompting guide, there are some breaking changes-

    19. KL

      Mm-hmm

    20. DS

      ... that teams would have to like care about.

    21. KL

      Yeah.

    22. DS

      But in, in Claude Managed Agents, it's just like a one-line update.

    23. KL

      Nice.

    24. DS

      And that's the kind of thing that you get for free that I think is really powerful.

    25. KL

      Yeah, that makes a lot of sense. Um, some of those breaking changes and just some of the how do we get the best out of Claude, um, and make the best use of all the things that it can do- As part of re-releasing the model, we're baking all those things-

    26. WW

      Mm-hmm

  12. 18:39 – 20:09

    Identity & authorization model: separate coworker in Slack, acting on your behalf for tools

    1. KL

      ... into the Claude Managed Agents harness, and so that's cool to hear. How did you think about, uh, auth and identity? Like, identity for the Everyagent, is it acting on behalf of humans? Is it acting on behalf of itself? Um, how did you think about that?

    2. WW

      Yeah, there's a lot of nuance to this because as soon as you sort of put an agent inside of a, a, you know, basically a company in-in-internally, people think of it like a human, right? And they ask it to do things that are, that fall again into these gray zones of like, well, a human would know that, like, there's no harm in me being able to access this, like, report, you know? Um, but that's, that's hard to sort of encode into an, a piece of agent logic. Um, and so we... But, but that we settled on that as our North Star. Let-let's try and make it act as close to a human as it can within, within s- you know, Slack. And the way we do it is when it's in Slack, it acts as a separate entity, where you're interacting with it, and, uh, like, underneath the hood it has y- you know, isolated memories and I- an understanding of who you are and what you have access to within the, the, within the company. It's a conversation between two separate entities. When you ask it to go out and grab something via, like, a tool call or a connection or whatnot, it, for the most part, acts on your behalf, where it's, it's as an extension of you when it goes into, um, go grab a document or look up something. Um, and we found that model to be pretty accurate to the way people want to use it, um, where they're fine with this, this boundary around sort of like internal versus external.

  13. 20:09 – 23:14

    Continuous improvement through ‘taste nudges’ + the hiring analogy for agents

    1. KL

      Nice. That makes a lot of sense. Um, we touched a little bit on continuous improvement, memory. Um, how are you guys thinking about that? Like, what is most important to you when it comes to the agent can do certain things today really well, maybe not as well as you want. Um, how are you thinking about get that agent really great over time at all of those things better and better?

    2. WW

      Yeah. The-- One of the nice things is the agent is able to see all of the little moments where you are sort of adding your taste into requests. These little nudges of like, ah, this... You know, you didn't write this well. Uh, one today was like, it told a bad joke. I was like, "This is not a great joke. Give me another joke," you know? [laughs]

    3. DS

      We take humor seriously.

    4. WW

      Yeah, yeah, yeah. [laughs]

    5. DS

      Of course. [laughs]

    6. WW

      Um, and those little nudges are really the foundation of like how... what is your taste? Like, it's hard to describe your taste-

    7. KL

      Mm-hmm

    8. WW

      ... uh, you know, all in one big go, but it's easy to take it from a thousand examples and sort of distill it out. And so those form the foundation of like, how do we build skills for, as Dan was saying, like particular jobs?

    9. KL

      Mm-hmm.

    10. WW

      Right? If we're doing-- If you're a person who does a lot of editing, the way the agent gets better at editing is looking at your 30,000 edits and con- using that as the, the foundation for like, oh, this is, these, these, this is the, the skill, the taste I have when it comes to editing.

    11. KL

      Mm-hmm.

    12. WW

      And, you know, you can think about this for like PowerPoint generation, you can think about this for email, you can think about this... It happens in coding as well, but, like, this is a new frontier for a lot of other non-coding knowledge work, where you're sort of accumulating these examples. Um, and the agent is really the repository for all of those.

    13. DS

      I think a good way to, to think about this is... I've been talking about, a lot about benchmarks, but I think it's so, I think it's so important, and, and the way that we think about it is starting to shift, where think about hiring or, or using an agent as hiring for a job.

    14. KL

      Mm-hmm.

    15. DS

      If you're hiring a job candidate and you're looking at their SAT scores, that's a little bit like a benchmark score.

    16. KL

      Yeah.

    17. DS

      And directionally, that actually helps a lot. So if I'm hiring a job candidate and they got a 1600 on their SATs versus a 300, I would probably take the 1600 person.

    18. KL

      Hope so. [laughs]

    19. DS

      Probably. But if everyone got a 1600, which is where we are with, with these models, um, there's a lot more I need to do. What I actually want is not your SAT score. Uh, I want a re- a reference check. I want a work trial, right? And for, for these models, uh, when they come into your company on the first day, it's their first day of work. Like, they're, they're a new grad who's really smart-

    20. WW

      Mm-hmm

    21. DS

      ... has studied a lot, but they don't know how you work.

    22. KL

      Yeah.

    23. DS

      And what you wanna try to do is create a system to, as they get experience working with you, to compound their knowledge of how you work and what you care about, so that on day 10 they're better than on day one. And what, what we are starting to build is a way to actually, uh, prove that. So like, to allow you to gather cases where the model did, did good work or didn't do good work, pull out your taste, and then allow, allow it to provably improve at those things over time. That won't be in the agent on day one, but I think that's where we're going, and I think generally where work is going, where-

    24. KL

      Yeah

    25. DS

      ... instead of doing each thing manually every single time, you're working on the system, and hopefully your agent, whether that's the Everyagent or Claude or, or any agent you're working with, enables you to do that, enables you to see how it's getting better.

    26. WW

      Mm-hmm.

  14. 23:14 – 27:01

    What’s next: personal benchmarks, computer-use experiments, and social interaction design

    1. KL

      Yeah. Makes a lot of sense. What are your next couple ideas for how you wanna improve the product?

    2. DS

      It's a good question. [laughs] I, I think, uh, so the, the big one is this, this sort of like compounding system-

    3. KL

      Yeah

    4. DS

      ... um, that allows anyone to basically create their own little personal benchmark for here's how I work and what I care about, and what I think is good about this model's response versus another model's response, and allows you to see, uh, which one you might want to use over time, and allows you to build and optimize skills that are for you and based on your taste to do your work better over time, and allow that to spread throughout the org so that other people can use those skills and work like you can to enable more experts to get more work done without spending more time. So I think, I think that's, that's where we're going. We do tons and tons of experiments, so I have like, I have this little experiment called Hands, which is like a computer use agent on my, on my desktop that, like, the Everyagent can then go use my computer.

    5. KL

      Oh, nice.

    6. DS

      'Cause I, 'cause you're always getting-

    7. WW

      Yeah

    8. DS

      ... like, requests where it's like, "Oh, can you do this thing?" And I'm just like, " @every use Hands" to, like, do it on my computer.

    9. KL

      Nice.

    10. DS

      And that's really cool. But yeah, we're experimenting with a lot of things.

    11. WW

      This is one of the, uh, like one of the number one questions we get from our customers is just like, how does this compare to sort of like working with Claude Code or Claude Coworker agent? And we see them as like very, um- Like, they work together really well because of experiments like Hans, where it's like sometimes you wanna be with coworkers and, and doing this thing and sort of doing it in public.

    12. KL

      Yeah.

    13. WW

      And then sometimes you wanna just shift that work very seamlessly over to, like, "I just wanna do it alone." You know, just like you would, you know, as if you're a human.

    14. KL

      Yeah.

    15. DS

      I mostly wanna do it alone.

    16. WW

      Yeah, yeah. [laughs]

    17. DS

      [laughs]

    18. WW

      Yeah. Um, and, and, and, and really it's this whole, you know, like everyone just sort of figuring it out, right? Like, where do I wanna spend my time with, with this sort of amplified intelligence? Um, and, and how do I go back if I don't? Like, like doing- like, for us, like doing context- how you do context management is both, like, gotten more invisible but also more visible as we've moved into, like... Personally, it's like I, I don't really think about context management anymore. It's- [laughs]

    19. DS

      [laughs] Yeah, yeah, yeah. Not thank you so much.

    20. KL

      [laughs]

    21. WW

      Um, uh, but with coworkers it, it really is. It's like, okay, we're trying to... Uh, uh, again, if you model it as a, as a coworker, it's like, oh, there was this conversation over here. Were you, were you part of it? And do you have, uh, enough of an understanding of human conversational norms to be like, you should, you should, you know, interject over here, or maybe you should take these two ideas from separate threads and, like, they're relevant even though one was a brainstorm and one was something a little bit more directional. And so a lot of, like, the interaction design and, like, personality design-

    22. KL

      Mm-hmm

    23. WW

      ... I think is, like, a huge thing that's coming.

    24. DS

      That is such a big thing-

    25. WW

      It's coming

    26. DS

      ... 'cause Claude so far, and I think this is starting to change, but so far every time you- it's used to every time it sees a message, it's supposed to respond-

    27. KL

      Mm

    28. DS

      ... because you're prompting it-

    29. WW

      Yeah

    30. DS

      ... one-on-one usually.

  15. 27:01 – 30:51

    Looking ahead + asks from Anthropic: fidelity, usability for knowledge workers, and personality knobs

    1. KL

      Well, as very AGI pill people, I'm sure you're thinking a lot about where things go from here, where are we gonna be in three months, six months, a year, um, just in terms of how people are gonna do work, right?

    2. DS

      Mm-hmm.

    3. KL

      What do you think is gonna happen, and where do you think that Every Agent's gonna need to be? What are you gonna need to build into it so that it's ready for that moment of how work is gonna change?

    4. DS

      It's a very good question, and I think it goes back to exactly what we were talking about in the beginning, where, um, you see developers starting to work in loops, starting to instead of, uh, you know, writing every line of code by hand, you're, you're, you're directing sometimes a whole organization of, of Claudes to do your work for you. Um, and your job is to make sure that even if you're not doing every unit of work, you've created a system that feels like it- it's an extension of you that... I've been thinking about words for this. I think one word is, like, fidelity. Like, the agents have fidelity-

    5. KL

      Hmm. Yeah

    6. DS

      ... for what you want, what you care about-

    7. KL

      Love that

    8. DS

      ... so that you don't have to be involved in every single thing. That's starting to be the case in how we do programming. Obviously, like, with s- with all the security stuff going on, like, sometimes it's not the case-

    9. KL

      Yeah

    10. DS

      ... and that's an interesting one.

    11. KL

      Yeah.

    12. DS

      It's also something that benchmarks don't really measure that well because they have- they're, by, by definition, they are more generic. They're not about you and what you want.

    13. KL

      Right.

    14. DS

      Um, but I think over the next year or so, uh, in order to take that next leap in, in knowledge work, um, we're going to have to make agents that have fidelity to what you want and what you care about, and I think what that will unlock is for individual knowledge workers, the same kind of thing that you see in coding where you're extracting yourself out of, out of the, uh, the, the loop. You're working on the loop. You're working on the system, and making sure the system operates according to what you care about and, and what your taste is and what your vision is, and you're intervening at certain, at certain times, just like a manager would, to make sure things are on track. And what that will do is allow you, A, as an individual person, to do way more than you could have before, and then, B, allow other people, whether it's in your organization or your clients, to use that, use your taste, do things like you would, without taking up more of your time.

    15. KL

      Yeah, that makes a lot of sense. And what are you gonna need from us-

    16. DS

      [laughs]

    17. KL

      ... from Anthropic, from the Claude platform, in order to make Every Agent what it needs to be at that time?

    18. DS

      We're starting to reach the point where successive gains in model intelligence are not necessarily able to be consumed-

    19. WW

      Mm-hmm

    20. DS

      ... even by power users.

    21. WW

      Mm-hmm.

    22. DS

      I think it's so amazing that there are in- these advances in frontier math or, or now, like, now biology, and those are super, super important, but, um, I don't want you guys to forget about us knowledge workers.

    23. KL

      Yeah. [laughs]

    24. DS

      Um, because I truly do think ev- uh, e- even as those gains in those other areas are happening, there's s- there's still so much low-hanging fruit-

    25. KL

      Yeah

    26. DS

      ... in terms of knowledge work. Like, when I get a response back from Claude, like, do I understand it?

    27. KL

      Yeah.

    28. DS

      Does it reference things that, like, are at my level of knowledge-

    29. WW

      Mm

    30. DS

      ... or that I, I would remember? All the, like, all those kinds of little things I think are so critically important to making a tool that people can use every day-

  16. 30:51 – 33:46

    Hot take and ‘magic moments’: more agents can mean more work—and more delight

    1. KL

      Um, awesome. What's your hot take?

    2. DS

      My hottest take is there's all this discussion about how AI is changing work, and there's all this fear of it, like, making work totally go away. And m- my experience or our experience, I think of, I think of Every as being like a little bellwether for how knowledge work might, might change because everyone is trying to use agents as much as possible to get as much of their work done as, as possible. And what we've found is the more we use agents, the more work there is to do. And, um, certainly, the work changes, and certainly, like, the way this will play out in different parts of the economy and different companies is going to be different, but I, I think that that is a sort of general truth that when you reach this horizon of work, you expect there to be nothing, but actually there's ... It just opens up a whole new horizon in front of you. And, um, and, and that provides an opportunity for everyone to actually just do better work at a higher quality that feels like them. Like, automating your work doesn't mean that it's done in a robotic way or that you don't pay attention to it at all. It just means that where you're working in the system has changed.

    3. KL

      Okay, last fun question. I think a lot of people have this, like, magic fun moment that they sometimes feel with agents-

    4. DS

      Mm-hmm, mm-hmm

    5. KL

      ... that can be really powerful. One of my favorite ones is the emoji that I love to use the most is the blue colored heart. And one time, I was working, and I was working with Claude, and Claude showed up and reacted to one of my messages-

    6. DS

      Mm

    7. KL

      ... with a blue heart.

    8. DS

      Mm-hmm.

    9. WW

      Mm-hmm.

    10. KL

      I was like, "Whoa."

    11. DS

      Mm-hmm.

    12. KL

      "How did Claude know?" I'm so curious for you guys with Every Agent, what's been just the most fun, magic moment that you've experienced in working with it?

    13. WW

      Our CO, uh, keeps track of, uh, basically beer COs folks.

    14. KL

      Yeah. [laughs]

    15. WW

      Like, "Oh, my bad. Uh, I... My bad. I did this, you know. Like E- Every, remind me that, that-

    16. KL

      [laughs]

    17. WW

      ... I owe Dan a beer, uh, for this." Uh, and the agent will insert that back in occasionally, like, in, in the appropriate conversation where it's just like-

    18. KL

      Oh, that's really cool

    19. WW

      ... "Oh, yeah, remember when Dan, like, hooked you up? Like, are you still... Is that... Have you, have you reclaimed that beer? Have you claimed that beer yet?"

    20. KL

      [laughs]

    21. DS

      The lead on the Every website, his name is Andre, uh, and he does these, like, really, really funny announcements-

    22. WW

      Mm-hmm

    23. DS

      ... um, that are just, like, they're in all caps, and they're like, "I- You must turn on two-factor authentication-

    24. WW

      Yeah, yeah, yeah

    25. DS

      ... in all of your logins or whatever." But he, like, spends a lot of time, like, making them just, like, really funny, and I just made a skill that was like, have-

    26. WW

      Yeah

    27. DS

      ... the Every Agent, like, do announcements in his voice.

    28. KL

      Nice.

    29. DS

      And so now people are just using his announcement skill to, like, do all of their announcements-

    30. KL

      Nice

Episode duration: 33:46

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode z7cNbsr3b5s

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.