ClaudeBuilding secure agents for knowledge work
EVERY SPOKEN WORD
40 min read · 7,758 words- 0:00 – 0:07
Intro
- KLKatelyn Lesse
[upbeat music]
- 0:07 – 2:18
How AI is reshaping workflows: from benchmarks to “extending” knowledge workers
- KLKatelyn Lesse
Super excited to be chatting with you guys. I have said this to you before, but I think you guys are some of the most AGI-pilled people out there with the things that you build and the way that you talk about what's happening in AI, and so very exciting for me to get to spend some time chatting with you guys about what's going on and what you're building. So maybe actually just to start, I would love to talk about the fact that the way everybody is doing work is changing so much right now, and that means the tools that people need to be able to use to get work done are changing a lot. Um, how do you guys think about that?
- DSDan Shipper
We think about it a lot. Um, and it's, it's one of those, one of those questions that I, I-- it's on everybody's mind. Everybody is super afraid of it, or they're like, "Oh, it's gonna be a utopia, and, like, no one's gonna have to work at all." The space we try to occupy is, "Well, why don't we just use the models and try to figure it out?" What ends up happening is you have to reinvent your workflow from scratch, and the only way to do that is to, like, try it on the problems that you're, you're actually working on. And I think for us, we've been seeing over the last couple years, you, you, like, watch all the benchmark progress, and when benchmark progress happens, you're like, the benchmarks are sort of measuring is it, is it gonna take my job away?
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
Like, the, the higher it gets on a benchmark. And our perspective is if you use these tools well, they, uh, are good at extending you and your power and your influence, and there needs to be ways to both measure that and to help people understand, uh, how to do that, because yes, if you look at benchmarks, they can do pre-AI tasks, like, at or better than humans in certain cases. But if you're thinking about how do I use this for my work-
- KLKatelyn Lesse
Mm-hmm
- DSDan Shipper
... that, like, opens up a whole se-set of options. We start to see that in engineering first because there's verifiable rewards in programming, and so y- now you have, like, 10X, 100X engineers who are orchestrating teams of agents to pull forward their roadmap by, like, two years.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
And I think that is coming to the rest of knowledge work, but, uh, it's taken longer, uh, A, because it's, it's just harder to know whether or not it's doing a good job.
- KLKatelyn Lesse
Mm-hmm.
- DSDan Shipper
And B, knowledge work, it's a different way of thinking that maybe developers are a l- a little bit more used to.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
Um, but we think that's coming. We're trying to, we're trying to, uh, both discover it and disseminate it so everyone can work that way.
- 2:18 – 2:47
What Every is: media + tools to keep companies at the edge of AI
- KLKatelyn Lesse
So for people who don't know, I'm curious, uh, what is Every? Tell us about what you guys do.
- DSDan Shipper
Every is a media and technology company. We build the only subscription you need to stay at the edge of AI. So for our subscribers, we, uh, do a series of education and equipment that helps them work at the edge. Uh, on the education side, we have everything from a newsletter to courses. When new models come out, we do vibe checks. And on the equipment side, now we have an agent that allows you to work in the way that we do in public with your company.
- 2:47 – 4:33
Why build the Every Agent: spreading AI workflows across a company in public
- KLKatelyn Lesse
Okay, so, um, you guys built an AI coworker, um, to help solve this problem. Tell us about it. What were your goals? How's it going?
- DSDan Shipper
So we built the Every Agent. It is an AI coworker. It's an agent that is the easiest way to AI-pill your whole company. We started basically, like, in the OpenClaude craze. We started using OpenClaude, and we were like, "Oh my God, this is the coolest thing," and it spread throughout the whole company. And then-
- KLKatelyn Lesse
You had your Mac Minis.
- DSDan Shipper
We had all the Mac Minis.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
Um, and then unfortunately, all the OpenClaudes died off because, like, they're just really hard to maintain.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
And we started just using one company agent, and we were like, "Wow, this is actually really, really great," and, and noticed how easy it was to spread workflows that we discovered throughout the whole company because our job inside of Every is to work at the edge, but you'd be surprised how hard that is to make working at the edge even across the entire company.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
And so that started to happen. We started to see these workflows start to spread. What we've also realized is there's a certain type of person that loves Every who we've been calling an AI ambassador-
- KLKatelyn Lesse
Mm-hmm
- DSDan Shipper
... who is, like, uh, has that feeling about AI where they're like, "Oh my God, I have this Codex set up or this Claude set up or whatever that is, uh, absolutely 10X-ing the way that I do my work, and I can't imagine working without it. But none of my coworkers have any idea what this is, and I'm trying to, like, explain to them all the time how they should use it, but they're, like, tuning me out." And the beauty of a Slack agent, the Every Agent in particular, is that you can show instead of tell-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... how to use this stuff to actually do really, really great work because you're doing it in Slack. You're prompting in public instead of in private. And so we built the Every Agent to have all the things you would expect in an, in an agentic coworker. You know, it can help you do recurring tasks like gather reports or do social posts or-- and any of the kinds of things you might expect, but it happens inside of your company Slack-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... so that your coworkers just automatically pick it up, and it starts to spread your best practices. We try to empower the ambassadors as much as we can with a tool like this.
- 4:33 – 6:26
High-leverage ‘small tasks’ example: automating open enrollment guidance
- KLKatelyn Lesse
Awesome. What's one of your favorite examples of something you've seen somebody do with the Every Agent that maybe, maybe they could do it before, it would've taken them longer, maybe they couldn't have even really accomplished that thing before?
- DSDan Shipper
I really think that having an-- having the Every Agent, having an agent in your Slack, there's-- I mean, there's the sort of like crazy top-end stuff, but the, the stuff that I wanna focus on is actually there's, there are all these small little things that you're doing day-to-day that take up a lot of people's time and, and having a company agent makes really easy. So a simple example, we just had open enrollment. We did the entire open enrollment process through an agent, uh, through the Every Agent. The open enrollment, I-I'm sure, I'm sure you know, you get like a PDF that's like, "Here are all these plans," and, like-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... doing the algebra to figure out the plans is actually very hard, and as part of that process, the Every Agent both knew all the plans and knew who you were and, um, and helped you kind of figure out which one you should select for your particular situation, and that's something that would've been very hard to do and coordinate without a central company agent, and the Every Agent makes it really easy for, A, leaders of the company to create a set of skills that, that work in that way, and then B, get that disseminated to everyone in the org and make sure that the open enrollment process actually happens smoothly.
- KLKatelyn Lesse
Yeah, like the amount of time you probably saved with not every single person pasting separately into Claude-
- DSDan Shipper
Oh my God
- KLKatelyn Lesse
... like, "This is the PDF."
- DSDan Shipper
Yeah.
- KLKatelyn Lesse
"What do I pick?"
- DSDan Shipper
Yeah.
- KLKatelyn Lesse
That saves you a lot of time.
- DSDan Shipper
And, and honestly, like, that's some- even for me, like, I-- as I-- anytime I get one of those things, I just, like, call my dad, and he's like-
- KLKatelyn Lesse
[chuckles]
- DSDan Shipper
... "I can't believe you're, like, thirty-five. I can't believe I have to still explain this to you."
- KLKatelyn Lesse
[chuckles]
- DSDan Shipper
'Cause it is actually really complicated, and yes, I could just individually paste it into Claude, but, but knowing that this is coming from the company and the company has thought about for each person, like, "Here's what you might wanna pick," it really helps me feel confident in my choice.
- KLKatelyn Lesse
Yeah, now you can call your dad for more fun reasons.
- DSDan Shipper
Yes, exactly. Sorry, Dad, I automated that task. [laughing] Um, it's not one that, not one that he relishes.
- KLKatelyn Lesse
He didn't like it at all. [laughing]
- DSDan Shipper
[laughing]
- 6:26 – 8:27
Editorial automation that helps experts scale: “Kate Bench” top edits
- WWWillie Williams
Well, we also, like, we have a lot of pretty rote tasks that go along pretty with the editorial arm. And we've spent a lot of time, like, building out something we call Kate Bench, which is named after our editor-in-chief, which is mostly like how do we do top edits before things go out? And there's still a question of like when to trigger these top edits, and so it's like we flood Kate's inbox. It's now much easier to just be- to tag the agent in Slack and say like, "Hey, just review this." And the nice part is Kate can see that-
- KLKatelyn Lesse
Mm-hmm
- WWWillie Williams
... and she can kind of like judge like, okay, is the agent, like, doing the job or not?
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
Yeah, this is, this is a really important thing that I think, I think Kate Bench gives or gives you a, a, an, a view into how this stuff is gonna be used in knowledge work going forward. So l- like Willie said, Kate's our editor-in-chief. Like, I've been trying to automate her for many years-
- KLKatelyn Lesse
[laughs]
- DSDan Shipper
... but like, but, like, in the most loving way possible.
- KLKatelyn Lesse
I was gonna say-
- DSDan Shipper
Um-
- KLKatelyn Lesse
... maybe don't have her watch. [laughs]
- DSDan Shipper
Yeah, like, like, uh, automating I think sounds scary-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... but when you actually see it in practice, you're like, "Oh, this is actually helpful for her"-
- KLKatelyn Lesse
Yeah, for sure
- DSDan Shipper
... because we publish so much. Um, and when, when Kate started, there was like four people at Every. She was like the fourth person, and now there's, now there's 30 people. Um, and she has incredible taste for copy edits, for like making sure that every single period, every single verb, every single thing is in the right place. But, um, we have 10 times the amount of things going out today than we did when she started, but she still has to do it all.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
So what I did was I just, uh, took like 30,000 of her historical edits and turned it into a skill that is in the Every Agent. So now every time we are publishing something, whether it's a piece or a landing page or an email, anyone in the organization can just say like, "@Every," like, "do a Kate Pass on this," and it will take her taste and apply it, like literally make edits in the Google Doc, like with suggested changes as her, and then, uh, they can accept and reject it. And what she does is she just goes in and looks at it, and it has most of the things that she would do and even catches things she might not. She has to spend less time on each piece 'cause-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... it has done a lot of the work. And then what we do is we track what is the extra work that she added to this-
- KLKatelyn Lesse
Mm
- DSDan Shipper
... and then we compound it back into the skill so it gets better and better over time.
- KLKatelyn Lesse
Oh, nice.
- 8:27 – 9:47
Reframing automation: compounding an expert’s impact instead of replacing them
- DSDan Shipper
And she gets to see a graph of here's how, here's how much time it's saving you, and here's also a rule we might wanna add to this so that it catches more things you don't, you didn't necessarily know about. And I think what that does is it, it flips the sort of like automation debate from, like, is it gonna take away my job to actually I'm an expert in an organization. My time's taken up all day by having people ask me to do things for them that only I can do.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
And as the organization scales, in order to scale my impact, previously I would've had to spend more time. And if you have an agent that compounds like the Every Agent, the opportunity for experts inside of organizations is you can spread your impact and spread your, the way that you do work without having to spend more time because you're working on a system that does the work for you instead of doing all the work yourself.
- KLKatelyn Lesse
Yeah, it's doubly valuable. Kate gets more time to do other things, um, maybe the hardest copy edits-
- DSDan Shipper
[laughs]
- KLKatelyn Lesse
... so Kate, and then everybody else gets to take advantage. I actually would love to talk about, um, Kate Bench-
- DSDan Shipper
Mm
- KLKatelyn Lesse
... for an ex- for an example. You guys talk a lot about being really, um, you know, bench pilled, right?
- DSDan Shipper
Mm.
- KLKatelyn Lesse
And being very on top of here's how we do some, like, really cool use case, and then measure how well we're performing on that really cool use case. Tell us more about that mentality-
- DSDan Shipper
Yeah
- KLKatelyn Lesse
... and how, like, as you have people going out there and starting to use the Every Agent, they can kind of bring this mentality in on, like, continuous improvement.
- 9:47 – 11:31
From vibe checks to personalized evals: measuring AI by your real work
- DSDan Shipper
Yes. We, we have been- really started to get, to get good at this, get, get good at evals and start trying to measure these things. I think, you know, we started with everyone who's in AI and, and follows us closely knows that that doesn't really mean that much.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
Um, and the question is not how good is it at a generic benchmark, it is how good is it at the kind of work I do?
- KLKatelyn Lesse
Right.
- DSDan Shipper
And that's a whole- that's a totally different way of thinking about, um, whether AI is good because it's much more personal, and that's where we started with our vibe checks, where before a new model comes out, we get access to it early, and we actually test it by hand across a lot of our workflows, and then on the day it launches, we can tell you, "Here's the kind of work we do. Here's our taste. Here's what we think of this model." And people rely on that to be like, "Well, I trust Every and I trust their taste," or like, "There's a certain writer at Every that I know has similar taste to me, so like, if they like this model, I'll try it." Problem is that that's very manual and takes a long time, and so what we've started to do is capture those tasks that we do. So like anytime one of us is doing work with a model day to day and we, we notice, ooh, it did something I like, or ooh, it did something I didn't like, we can sort of capture that in, in one place, and then talk to this platform that we've built that will eventually be part of the Every Agent, talk to this platform that we built about why we liked it and why we didn't, and that turns it into this set of checks, which are basically like unit tests or rules for here's the kind of thing you look for. So for Kate Pass or Kate Bench, it would be is the headline sentence case would be an example. And, um, what, what that allows us to do is run many different models on the same scenario and then be, be able to see for our real work, did this model do a good job or not according to my taste? We've started to be able to add this sort of like quantified layer to vibe checks that I, I think looks like a benchmark, but actually measures things that you care about.
- 11:31 – 13:02
Why Slack won: consolidating into one shared agent that everyone improves
- KLKatelyn Lesse
Awesome. Um, going back to form factor a bit, you talked about obviously being in Slack is really powerful. All the knowledge workers can have really easy access to this. Um, how did you land on Slack? How are you thinking about that form factor?
- WWWillie Williams
Well, we s- we actually started with sort of a model of, uh, every one agent per person, you know, available in whatever surface you really wanted, uh, that agent to be there, if it was Telegram, if it was Slack. Um, and part of the reason we over time went from like a multi-agent setup to really a singular agent setup is we found that agents got better via the same dynamic the more you invested in them.
- KLKatelyn Lesse
Yep.
- WWWillie Williams
And so you had a lot of folks who were like, "Oh, I invest- I'm investing in my agent," but those gains weren't really like spreading out to the rest of the team. And the places where we, we really saw like agents being able to do work, real work, was when there were multiple people investing in them, making better skills, giving it better guidance, and it just made so much more sense for everyone to be investing in like a single agent rather than, uh, everyone having their own agent sort of spread around. And Slack was like- The very natural form factor for like where are we when we're having these conversations with individual agents and, and when we're doing, like, collaborative work. There's a lot of copy and paste.
- KLKatelyn Lesse
Mm-hmm.
- WWWillie Williams
Right? I'm like, I'm like, we're gonna have a conversation. I'm gonna, like, copy this over to my, like, Claude coworker, and I'm gonna copy the, the response back, and then you're gonna do that. And instead it's much easier to just have that conversation in the same place that the, your coworkers are and the same place your agent is.
- KLKatelyn Lesse
Perfect and easily accessible. Let's talk a bit about how you built Everyagent.
- WWWillie Williams
Okay.
- 13:02 – 16:37
Building on Claude Managed Agents: avoiding infra distractions and gaining primitives
- KLKatelyn Lesse
So you built this product on Claude Managed Agents-
- WWWillie Williams
Yeah
- KLKatelyn Lesse
... uh, which we're really excited about. Tell us just starting out, what were some of the main reasons why you chose to start building on Claude Managed Agents?
- WWWillie Williams
Well, we learned a lot of pain- painful lessons with this first version of One Agent for Everyone.
- KLKatelyn Lesse
Okay.
- WWWillie Williams
Um, one of the main ones is, uh, infrastructure's hard, you know? Um, and a lot of what we want to focus on is like the future of work and like how to make a better product for like that future that we see coming. We don't necessarily want to spend that time like, how do we orchestrate sandboxes?
- KLKatelyn Lesse
Yep.
- WWWillie Williams
Right? And so we, [chuckles] we had this very painful arc where we're like, "Okay, we're going to, uh, we're gonna build one agent for everyone," and then it just got unwieldy very quickly because you're sort of like trying to improve a product and have a, have a strong, durable like infra layer underneath. And so w- we took those lessons and turned them into the Everyagent, and one of our first things was we don't want to be an infra team.
- KLKatelyn Lesse
Mm-hmm.
- WWWillie Williams
Right? We want to focus our energy in the place where we think, uh, the future is going. And so Claude Managed Agents was great. It wa- it was, it gave us, you know, it gave us sandboxes, it gave us memory, it gave us session control, all these primitives that really allowed us to, um, focus on the next layer of how do we do interaction design for, for an agent. You know, we can kind of just like shuffle it off.
- KLKatelyn Lesse
Yeah.
- WWWillie Williams
Just be like, okay, this is being taken care of. You know, I'm not getting alerted w- uh, when, when orchestrating.
- KLKatelyn Lesse
I-
- WWWillie Williams
Yeah, yeah. [laughs]
- KLKatelyn Lesse
[laughs]
- DSDan Shipper
Thank you. Thank you for your service.
- WWWillie Williams
Yeah, yeah, yeah, yeah. [laughs]
- DSDan Shipper
[laughs]
- WWWillie Williams
Uh, and it's been great.
- KLKatelyn Lesse
Nice.
- DSDan Shipper
Yeah. That's, that's one of those like classic things in startups where you're like, "Okay, we're gonna build an agent platform. Should we do the infra layer or not?" And you're like, "Oh, infra's not that hard anymore. Like, you can just... Like, we'll just have Claude spin up a bunch of servers and like manage the servers for us and like blah, blah, blah." I was saying that at least, and Willie was like, "No, no, it's, it's hard."
- WWWillie Williams
[laughs] Yeah.
- DSDan Shipper
And I was like, "We're gonna, we're, we're going to try it." And then a month later I was like-
- WWWillie Williams
Yeah.
- DSDan Shipper
Yeah.
- WWWillie Williams
You're like-
- DSDan Shipper
[laughs] This is hard.
- KLKatelyn Lesse
[laughs]
- DSDan Shipper
And yeah, Claude Managed Agents, it takes something that you don't know is hard, but when you do it, you're like, "This is so complicated," and it makes it very easy for us.
- KLKatelyn Lesse
Nice. That's awesome. So there's infra management, that's obviously helpful. What are some of the other features? You mentioned memory.
- 16:37 – 18:39
Security, isolation, and fast model updates: harness and model are now coupled
- KLKatelyn Lesse
Yeah, that makes a lot of sense. And, uh, there's a lot going on right now in terms of security-
- WWWillie Williams
Yeah
- KLKatelyn Lesse
... with agents and, um, a lot of energy is going into how do we make sure that agents are doing the things we want them to do-
- WWWillie Williams
Yeah
- KLKatelyn Lesse
... and not the things that we don't want them to do.
- WWWillie Williams
Yeah.
- KLKatelyn Lesse
How do Claude Managed Agents help you think about security with Everyagent?
- WWWillie Williams
Yeah. A, a lot of it is just the built-in isolation. It, it makes it very easy to, um, isolate out, uh, the sandbox, the tool calls, like what's, what's happening agent in the loop, um, and what are the environments and where are the memories and like we can, we can, because we can control all that, we can make choices around like, uh, who accesses what when. You know, session management is great for this, and that really gives us the, the nuanced tools to say like, okay, you know, if Dan requests something, does he get, uh, the same thing that I would get if I requested it?
- KLKatelyn Lesse
Mm-hmm.
- WWWillie Williams
Right? And underneath the hood there's, there's a lot of playing with the primitives of like what memories do we attach to that request? Like is the, are the sandboxes the same or different? Um, and so yeah, just having those primitives, uh, is just helpful from an engineering standpoint.
- DSDan Shipper
One thing that I do want to add, too, I think this relates to s- security, but it also relates to performance, which is there's not a clear separation between the model and the harness anymore.
- KLKatelyn Lesse
Mm-hmm.
- WWWillie Williams
Mm-hmm.
- DSDan Shipper
And a model that doesn't come with a computer that it knows how to use, like it's just not as performant.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
And having all of that built in one place by one provider, uh, ensures that the model knows how to use that stuff really well, and you know, uh, the, the day we're filming this, there's a new, a new, new model launch from you guys.
- WWWillie Williams
Yeah.
- DSDan Shipper
And it's one of those things Willie, Willie was telling me, if you look at the prompting guide, there are some breaking changes-
- KLKatelyn Lesse
Mm-hmm
- DSDan Shipper
... that teams would have to like care about.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
But in, in Claude Managed Agents, it's just like a one-line update.
- KLKatelyn Lesse
Nice.
- DSDan Shipper
And that's the kind of thing that you get for free that I think is really powerful.
- KLKatelyn Lesse
Yeah, that makes a lot of sense. Um, some of those breaking changes and just some of the how do we get the best out of Claude, um, and make the best use of all the things that it can do- As part of re-releasing the model, we're baking all those things-
- WWWillie Williams
Mm-hmm
- 18:39 – 20:09
Identity & authorization model: separate coworker in Slack, acting on your behalf for tools
- KLKatelyn Lesse
... into the Claude Managed Agents harness, and so that's cool to hear. How did you think about, uh, auth and identity? Like, identity for the Everyagent, is it acting on behalf of humans? Is it acting on behalf of itself? Um, how did you think about that?
- WWWillie Williams
Yeah, there's a lot of nuance to this because as soon as you sort of put an agent inside of a, a, you know, basically a company in-in-internally, people think of it like a human, right? And they ask it to do things that are, that fall again into these gray zones of like, well, a human would know that, like, there's no harm in me being able to access this, like, report, you know? Um, but that's, that's hard to sort of encode into an, a piece of agent logic. Um, and so we... But, but that we settled on that as our North Star. Let-let's try and make it act as close to a human as it can within, within s- you know, Slack. And the way we do it is when it's in Slack, it acts as a separate entity, where you're interacting with it, and, uh, like, underneath the hood it has y- you know, isolated memories and I- an understanding of who you are and what you have access to within the, the, within the company. It's a conversation between two separate entities. When you ask it to go out and grab something via, like, a tool call or a connection or whatnot, it, for the most part, acts on your behalf, where it's, it's as an extension of you when it goes into, um, go grab a document or look up something. Um, and we found that model to be pretty accurate to the way people want to use it, um, where they're fine with this, this boundary around sort of like internal versus external.
- 20:09 – 23:14
Continuous improvement through ‘taste nudges’ + the hiring analogy for agents
- KLKatelyn Lesse
Nice. That makes a lot of sense. Um, we touched a little bit on continuous improvement, memory. Um, how are you guys thinking about that? Like, what is most important to you when it comes to the agent can do certain things today really well, maybe not as well as you want. Um, how are you thinking about get that agent really great over time at all of those things better and better?
- WWWillie Williams
Yeah. The-- One of the nice things is the agent is able to see all of the little moments where you are sort of adding your taste into requests. These little nudges of like, ah, this... You know, you didn't write this well. Uh, one today was like, it told a bad joke. I was like, "This is not a great joke. Give me another joke," you know? [laughs]
- DSDan Shipper
We take humor seriously.
- WWWillie Williams
Yeah, yeah, yeah. [laughs]
- DSDan Shipper
Of course. [laughs]
- WWWillie Williams
Um, and those little nudges are really the foundation of like how... what is your taste? Like, it's hard to describe your taste-
- KLKatelyn Lesse
Mm-hmm
- WWWillie Williams
... uh, you know, all in one big go, but it's easy to take it from a thousand examples and sort of distill it out. And so those form the foundation of like, how do we build skills for, as Dan was saying, like particular jobs?
- KLKatelyn Lesse
Mm-hmm.
- WWWillie Williams
Right? If we're doing-- If you're a person who does a lot of editing, the way the agent gets better at editing is looking at your 30,000 edits and con- using that as the, the foundation for like, oh, this is, these, these, this is the, the skill, the taste I have when it comes to editing.
- KLKatelyn Lesse
Mm-hmm.
- WWWillie Williams
And, you know, you can think about this for like PowerPoint generation, you can think about this for email, you can think about this... It happens in coding as well, but, like, this is a new frontier for a lot of other non-coding knowledge work, where you're sort of accumulating these examples. Um, and the agent is really the repository for all of those.
- DSDan Shipper
I think a good way to, to think about this is... I've been talking about, a lot about benchmarks, but I think it's so, I think it's so important, and, and the way that we think about it is starting to shift, where think about hiring or, or using an agent as hiring for a job.
- KLKatelyn Lesse
Mm-hmm.
- DSDan Shipper
If you're hiring a job candidate and you're looking at their SAT scores, that's a little bit like a benchmark score.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
And directionally, that actually helps a lot. So if I'm hiring a job candidate and they got a 1600 on their SATs versus a 300, I would probably take the 1600 person.
- KLKatelyn Lesse
Hope so. [laughs]
- DSDan Shipper
Probably. But if everyone got a 1600, which is where we are with, with these models, um, there's a lot more I need to do. What I actually want is not your SAT score. Uh, I want a re- a reference check. I want a work trial, right? And for, for these models, uh, when they come into your company on the first day, it's their first day of work. Like, they're, they're a new grad who's really smart-
- WWWillie Williams
Mm-hmm
- DSDan Shipper
... has studied a lot, but they don't know how you work.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
And what you wanna try to do is create a system to, as they get experience working with you, to compound their knowledge of how you work and what you care about, so that on day 10 they're better than on day one. And what, what we are starting to build is a way to actually, uh, prove that. So like, to allow you to gather cases where the model did, did good work or didn't do good work, pull out your taste, and then allow, allow it to provably improve at those things over time. That won't be in the agent on day one, but I think that's where we're going, and I think generally where work is going, where-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... instead of doing each thing manually every single time, you're working on the system, and hopefully your agent, whether that's the Everyagent or Claude or, or any agent you're working with, enables you to do that, enables you to see how it's getting better.
- WWWillie Williams
Mm-hmm.
- 23:14 – 27:01
What’s next: personal benchmarks, computer-use experiments, and social interaction design
- KLKatelyn Lesse
Yeah. Makes a lot of sense. What are your next couple ideas for how you wanna improve the product?
- DSDan Shipper
It's a good question. [laughs] I, I think, uh, so the, the big one is this, this sort of like compounding system-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... um, that allows anyone to basically create their own little personal benchmark for here's how I work and what I care about, and what I think is good about this model's response versus another model's response, and allows you to see, uh, which one you might want to use over time, and allows you to build and optimize skills that are for you and based on your taste to do your work better over time, and allow that to spread throughout the org so that other people can use those skills and work like you can to enable more experts to get more work done without spending more time. So I think, I think that's, that's where we're going. We do tons and tons of experiments, so I have like, I have this little experiment called Hands, which is like a computer use agent on my, on my desktop that, like, the Everyagent can then go use my computer.
- KLKatelyn Lesse
Oh, nice.
- DSDan Shipper
'Cause I, 'cause you're always getting-
- WWWillie Williams
Yeah
- DSDan Shipper
... like, requests where it's like, "Oh, can you do this thing?" And I'm just like, " @every use Hands" to, like, do it on my computer.
- KLKatelyn Lesse
Nice.
- DSDan Shipper
And that's really cool. But yeah, we're experimenting with a lot of things.
- WWWillie Williams
This is one of the, uh, like one of the number one questions we get from our customers is just like, how does this compare to sort of like working with Claude Code or Claude Coworker agent? And we see them as like very, um- Like, they work together really well because of experiments like Hans, where it's like sometimes you wanna be with coworkers and, and doing this thing and sort of doing it in public.
- KLKatelyn Lesse
Yeah.
- WWWillie Williams
And then sometimes you wanna just shift that work very seamlessly over to, like, "I just wanna do it alone." You know, just like you would, you know, as if you're a human.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
I mostly wanna do it alone.
- WWWillie Williams
Yeah, yeah. [laughs]
- DSDan Shipper
[laughs]
- WWWillie Williams
Yeah. Um, and, and, and, and really it's this whole, you know, like everyone just sort of figuring it out, right? Like, where do I wanna spend my time with, with this sort of amplified intelligence? Um, and, and how do I go back if I don't? Like, like doing- like, for us, like doing context- how you do context management is both, like, gotten more invisible but also more visible as we've moved into, like... Personally, it's like I, I don't really think about context management anymore. It's- [laughs]
- DSDan Shipper
[laughs] Yeah, yeah, yeah. Not thank you so much.
- KLKatelyn Lesse
[laughs]
- WWWillie Williams
Um, uh, but with coworkers it, it really is. It's like, okay, we're trying to... Uh, uh, again, if you model it as a, as a coworker, it's like, oh, there was this conversation over here. Were you, were you part of it? And do you have, uh, enough of an understanding of human conversational norms to be like, you should, you should, you know, interject over here, or maybe you should take these two ideas from separate threads and, like, they're relevant even though one was a brainstorm and one was something a little bit more directional. And so a lot of, like, the interaction design and, like, personality design-
- KLKatelyn Lesse
Mm-hmm
- WWWillie Williams
... I think is, like, a huge thing that's coming.
- DSDan Shipper
That is such a big thing-
- WWWillie Williams
It's coming
- DSDan Shipper
... 'cause Claude so far, and I think this is starting to change, but so far every time you- it's used to every time it sees a message, it's supposed to respond-
- KLKatelyn Lesse
Mm
- DSDan Shipper
... because you're prompting it-
- WWWillie Williams
Yeah
- DSDan Shipper
... one-on-one usually.
- 27:01 – 30:51
Looking ahead + asks from Anthropic: fidelity, usability for knowledge workers, and personality knobs
- KLKatelyn Lesse
Well, as very AGI pill people, I'm sure you're thinking a lot about where things go from here, where are we gonna be in three months, six months, a year, um, just in terms of how people are gonna do work, right?
- DSDan Shipper
Mm-hmm.
- KLKatelyn Lesse
What do you think is gonna happen, and where do you think that Every Agent's gonna need to be? What are you gonna need to build into it so that it's ready for that moment of how work is gonna change?
- DSDan Shipper
It's a very good question, and I think it goes back to exactly what we were talking about in the beginning, where, um, you see developers starting to work in loops, starting to instead of, uh, you know, writing every line of code by hand, you're, you're, you're directing sometimes a whole organization of, of Claudes to do your work for you. Um, and your job is to make sure that even if you're not doing every unit of work, you've created a system that feels like it- it's an extension of you that... I've been thinking about words for this. I think one word is, like, fidelity. Like, the agents have fidelity-
- KLKatelyn Lesse
Hmm. Yeah
- DSDan Shipper
... for what you want, what you care about-
- KLKatelyn Lesse
Love that
- DSDan Shipper
... so that you don't have to be involved in every single thing. That's starting to be the case in how we do programming. Obviously, like, with s- with all the security stuff going on, like, sometimes it's not the case-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... and that's an interesting one.
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
It's also something that benchmarks don't really measure that well because they have- they're, by, by definition, they are more generic. They're not about you and what you want.
- KLKatelyn Lesse
Right.
- DSDan Shipper
Um, but I think over the next year or so, uh, in order to take that next leap in, in knowledge work, um, we're going to have to make agents that have fidelity to what you want and what you care about, and I think what that will unlock is for individual knowledge workers, the same kind of thing that you see in coding where you're extracting yourself out of, out of the, uh, the, the loop. You're working on the loop. You're working on the system, and making sure the system operates according to what you care about and, and what your taste is and what your vision is, and you're intervening at certain, at certain times, just like a manager would, to make sure things are on track. And what that will do is allow you, A, as an individual person, to do way more than you could have before, and then, B, allow other people, whether it's in your organization or your clients, to use that, use your taste, do things like you would, without taking up more of your time.
- KLKatelyn Lesse
Yeah, that makes a lot of sense. And what are you gonna need from us-
- DSDan Shipper
[laughs]
- KLKatelyn Lesse
... from Anthropic, from the Claude platform, in order to make Every Agent what it needs to be at that time?
- DSDan Shipper
We're starting to reach the point where successive gains in model intelligence are not necessarily able to be consumed-
- WWWillie Williams
Mm-hmm
- DSDan Shipper
... even by power users.
- WWWillie Williams
Mm-hmm.
- DSDan Shipper
I think it's so amazing that there are in- these advances in frontier math or, or now, like, now biology, and those are super, super important, but, um, I don't want you guys to forget about us knowledge workers.
- KLKatelyn Lesse
Yeah. [laughs]
- DSDan Shipper
Um, because I truly do think ev- uh, e- even as those gains in those other areas are happening, there's s- there's still so much low-hanging fruit-
- KLKatelyn Lesse
Yeah
- DSDan Shipper
... in terms of knowledge work. Like, when I get a response back from Claude, like, do I understand it?
- KLKatelyn Lesse
Yeah.
- DSDan Shipper
Does it reference things that, like, are at my level of knowledge-
- WWWillie Williams
Mm
- DSDan Shipper
... or that I, I would remember? All the, like, all those kinds of little things I think are so critically important to making a tool that people can use every day-
- 30:51 – 33:46
Hot take and ‘magic moments’: more agents can mean more work—and more delight
- KLKatelyn Lesse
Um, awesome. What's your hot take?
- DSDan Shipper
My hottest take is there's all this discussion about how AI is changing work, and there's all this fear of it, like, making work totally go away. And m- my experience or our experience, I think of, I think of Every as being like a little bellwether for how knowledge work might, might change because everyone is trying to use agents as much as possible to get as much of their work done as, as possible. And what we've found is the more we use agents, the more work there is to do. And, um, certainly, the work changes, and certainly, like, the way this will play out in different parts of the economy and different companies is going to be different, but I, I think that that is a sort of general truth that when you reach this horizon of work, you expect there to be nothing, but actually there's ... It just opens up a whole new horizon in front of you. And, um, and, and that provides an opportunity for everyone to actually just do better work at a higher quality that feels like them. Like, automating your work doesn't mean that it's done in a robotic way or that you don't pay attention to it at all. It just means that where you're working in the system has changed.
- KLKatelyn Lesse
Okay, last fun question. I think a lot of people have this, like, magic fun moment that they sometimes feel with agents-
- DSDan Shipper
Mm-hmm, mm-hmm
- KLKatelyn Lesse
... that can be really powerful. One of my favorite ones is the emoji that I love to use the most is the blue colored heart. And one time, I was working, and I was working with Claude, and Claude showed up and reacted to one of my messages-
- DSDan Shipper
Mm
- KLKatelyn Lesse
... with a blue heart.
- DSDan Shipper
Mm-hmm.
- WWWillie Williams
Mm-hmm.
- KLKatelyn Lesse
I was like, "Whoa."
- DSDan Shipper
Mm-hmm.
- KLKatelyn Lesse
"How did Claude know?" I'm so curious for you guys with Every Agent, what's been just the most fun, magic moment that you've experienced in working with it?
- WWWillie Williams
Our CO, uh, keeps track of, uh, basically beer COs folks.
- KLKatelyn Lesse
Yeah. [laughs]
- WWWillie Williams
Like, "Oh, my bad. Uh, I... My bad. I did this, you know. Like E- Every, remind me that, that-
- KLKatelyn Lesse
[laughs]
- WWWillie Williams
... I owe Dan a beer, uh, for this." Uh, and the agent will insert that back in occasionally, like, in, in the appropriate conversation where it's just like-
- KLKatelyn Lesse
Oh, that's really cool
- WWWillie Williams
... "Oh, yeah, remember when Dan, like, hooked you up? Like, are you still... Is that... Have you, have you reclaimed that beer? Have you claimed that beer yet?"
- KLKatelyn Lesse
[laughs]
- DSDan Shipper
The lead on the Every website, his name is Andre, uh, and he does these, like, really, really funny announcements-
- WWWillie Williams
Mm-hmm
- DSDan Shipper
... um, that are just, like, they're in all caps, and they're like, "I- You must turn on two-factor authentication-
- WWWillie Williams
Yeah, yeah, yeah
- DSDan Shipper
... in all of your logins or whatever." But he, like, spends a lot of time, like, making them just, like, really funny, and I just made a skill that was like, have-
- WWWillie Williams
Yeah
- DSDan Shipper
... the Every Agent, like, do announcements in his voice.
- KLKatelyn Lesse
Nice.
- DSDan Shipper
And so now people are just using his announcement skill to, like, do all of their announcements-
- KLKatelyn Lesse
Nice
Episode duration: 33:46
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode z7cNbsr3b5s