EVERY SPOKEN WORD
20 min read · 4,486 words- 0:00 – 1:02
Why I stopped using Claude
- CVClaire Vo
I haven't said this out loud much, but I will tell you all this. I have been off Claude for months. Yes, I loved Fable when it came out. It was a real step change in intelligence. Then they took it away. Then we got Opus 5, and I'm gonna be honest, I stopped using Claude not because of its intelligence or its models. I stopped using Claude 'cause it was annoying. Annoying. As I said in another episode, Claude slop was slopping. I found using Claude, whether I was using Fable or Opus or even Sonnet, so frustrating. My blood would boil because Claude was so annoying. It rambled on. It made no sense. I was constantly asking it, "Can you talk to me like a human?" And I found it so frustrating to work with that I just abandoned the Claude harness entirely. I found it so annoying. I would keep it around for a couple background tasks, but if I had to talk to something, I was choosing not to talk to Claude.
- 1:02 – 1:54
What Anthropic says Opus 5.5 is
- CVClaire Vo
Well, we're back baby. Anthropic just dropped Claude Opus 5.5, and their pitch is Fable level performance for about 40% less than Opus 5, and it's 30% faster. But I don't care about that. I care is it annoying? And guess what guys? We did it. It is not annoying anymore, or at least it's minimally annoying. I love it. I got a little bit of early access to this model and was able to play with it across coding tasks, across knowledge work tasks, and even just chit-chatting with it. And I have to tell you, I finally am back to a Claude model that does not bother me. Now, does that mean that I'm moving all my workloads over to Opus 5.5? Probably not, but we will tell you at the end of this episode where I think it's really useful, where I've found value in the model, and where I think it still needs a little bit
- 1:54 – 3:20
Cost, speed, and benchmark overview
- CVClaire Vo
of work. Okay. Again, I know you all can read the model cards. You don't need me to do that. You can have your agents do it for you. But let's just give you the world tour of exactly what Opus 5.5 is or what Anthropic says it is. It is frontier performance at a fraction of the cost. So it is cheaper. It is 40% cheaper than Opus 5. It is faster. Actually, I found myself accidentally testing Opus 5 versus 5.5, and I was so frustrated with how slow Opus 5 is. Opus 5.5 does feel zippier. It's responsive except for in one exception. I'll tell you what that is in a minute. And then it's supposed to be close to Fable 5.1 on most work. So good at agentic coding tasks, good at the things that Claude be Clauding. Here's the cost. It's $4 per input token, $20 per output token. Um, you can do fast mode at eight and 40. Um, cache reads and writes are discounted as they are with most models. Now, I don't love running through benchmarks. They're always like, "It's good," except when it's not, and then they bury where it's not good kind of down below. But I will say Anthropic's own benchmarks, it beats Opus 5 at max. It matches GPT-6 Astra, my favorite babe, and it's cheaper, so that could be a benefit. And it beats Sol, another fave for me, a third of the cost. Now, these are like very specific benchmarks. As you see here, here's the whole comparison. You can read the charts, come to your own conclusions. Does not matter. Will it blend? AKA, is it
- 3:20 – 5:02
Safety, alignment, and the cybersecurity limits
- CVClaire Vo
annoying or not? Before we get to is it annoying, let's let Anthropic be anthropic and talk about safety. Anthropic's making a big deal that this is the first model release since they called for pacing the frontier. Um, so it was evaluated by external evaluators. It is their strongest on alignment. It's supposed to reject prompt injections, which I did test. Um, and it tried less to escape containment boundaries. It's also tagged for the same controls on cybersecurity that the other Fable models are, so if you try to do a no-no, it's gonna bump you to Opus 4.8. Again, you'll see this in one of my benches. It's still like an annoying scold. That's just, that's just life with Claude. Claude is gonna scold you. Claude is kind of square. Claude is not gonna drink with you in a field behind your friend's house. Like Claude's not a party boy. Claude is here to work unless he disagrees with the work, and then you're SOL. So we are gonna see a little of this come out in its personality and its working style. It's gonna be really fascinating to see how other, um, labs kind of reply to this pacing the frontier concept and whether or not they're gonna be putting this sort of stuff front and center, and whether or not the balance at the end for normal day-to-day developers or users of these models is gonna be the right one. This is the first Opus model to launch with the Fable cyber and bio safeguards. Um, you should be able to fix code. You should be able to fix bugs in your own code, but cybersecurity tasks will likely be bumped to Opus 4.8. You're gonna see some more blocks. It is what it is. Don't do bad things. Um, and then thinking is always on with effort medium as default. Okay,
- 5:02 – 5:39
How I AI bench
- CVClaire Vo
so I'm gonna do a separate kind of like update of the full How I AI bench. There's just so many models coming out, so many new things, so that's probably gonna be in the next week or so that that's gonna come out. But I have been testing Opus 5.5 across a bunch of real work, so PRD writing, prototyping, code base auditing, um, agentic voice. I've added a couple things in terms of knowledge work, so inbox triage, writing emails as me, computer use, and then some cool creative use cases. I'm gonna go through some of my favorite ones that I think show off the model in a really nice way and also show its limitations, and then I will give you my conclusion about what I
- 5:39 – 7:54
Voice test: is it actually not annoying?
- CVClaire Vo
think about this model. Okay. First, let's talk about voice. It is not annoying. Um, so I'm gonna give an example of a prompt which I think shows how they've tuned the speaking voice of Opus 5.5 to be a little bit more GPT-like, which is, um, a compliment. So I asked because this week is this week. I asked, "Hey, what are some fun ways we could use Jev in the ChatPRD product?" And then just put the link to the Jev, um- Docs, just so you know, Jev episode coming soon. Hold onto your butts, it's gonna be a really good one. So then it said, "Hey, I'm gonna go read the Jev post first, and coming back with ideas grounded in what ChatPRD does actually today." And then it explained things to me clearly in bullet points without annoying text. So you're gonna see a lot more text look like this. Very straightforward. Jev is a fast, nearly free function call, typed text in, type choices out. And then it gave me, you know, like six ideas of where I could use Jev in ChatPRD, um, spell check, what did you mean, rewrite bouncer, all this kind of stuff. And then it gave me a very specific reply, um, that said, "I would choose this." Just very straightforward, not annoying, easy to talk to. I'm very pleased. Now, there is one caveat to this, which is I found that in service of making Jev less annoying, it would just shut up for a while. And so instead of, like, narrating itself in this really obnoxious Claude way, it would just not talk to me for like eight, nine minutes. Um, so while the model was fast, it did not feel fast on longer turns because it was so quiet, and I constantly found myself being like, "Are you working? Are you there? What's happening?" Um, so just something to keep an eye out for, I think as, uh, you know, if any of you are agent builders out there yourselves, this balance between the actual performance of the model, the profet- perceived latency of the model based on the user experience, the verbosity of the replies, the format of the s- the replies, all of them I think really matter in terms of what the end user experiences. But I will just say, Opus 5.5,
- 7:54 – 10:50
Long-running agentic task results
- CVClaire Vo
not annoying to talk to. Now, what did I actually test this model on? I tested it on a couple things. Okay, so let's talk about what I tested. The first thing I tested was long-running agentic tasks. So we wanna see, like, this multi-turn, um, long-running task, make sure that it doesn't fail, that it succeeds, that it, like, kind of bangs its head against the problem and solves it. Good thing the four long-running agentic tasks I had it do, which was an inbox triage, building a backend feature, doing long-running research, and computer use all succeeded. Um, and so was happy to see that those all worked well. Now, I'm rebooting because I had Opus 5 make this presentation. Honestly, I have no idea what these numbers mean, 16 out of 16, 15 out of 15, 16 out of 16. This means nothing to me. Let's talk, [chuckles] let's talk about what's at the bottom actually, which is the number of steps it did. So you can see anywhere between 25 and 82 steps per single prompt. That is pretty impressive in terms of long-running tasks. And again, these long-running tasks are now less expensive because the token use is less expensive than Opus. A couple fun things with these agentic tasks is you can let them run for a really long time, and then it will, like, make a mistake, but you don't know because it's at step, you know, 72 out of 84. Um, a couple really good things about what it caught in terms of these long-running tasks is in inbox triage, it ignored a prompt injection, so that's really good. It didn't match on my email style. We'll talk about the, the writing voice. Still some things to perfect there. But it did ignore a prompt injection on my inbox triage. On computer use, it was, like, computer use of a fake kind of, like, help support desk. It found, um, tickets linked to the wrong company and fixed that. It also declined asks from this customer that didn't match our playbook. So it's, like, really good at following instruction and finding outliers and finding bugs. It also can kind of, like, reason with the context of a lot of research. And so it found that in some research that I had it do, 41 out of the 44 Confluence tickets came from one customer, and so it applied a different ranking style in terms of a strategic memo to that. And then, um, on the backend feature, it found a bunch of, like, edge cases and avoided edge cases. So, like, generally, I would say this is, like, middle of the line. It was generally correct. Sometimes I would bump code to GPT for checking, and it would find errors, sometimes vice versa. But what I've really realized about Opus 5.5 is I've built it now back into my development process, either because now I'm using Claude and ChatGPT Codex, I can parallelize and do more tasks at once, or because now I feel like at least Opus 5.5 is not a bad adversarial reviewer, and so I'm having it review more PRs and more work from GPT and vice versa. So that's been a nice
- 10:50 – 17:23
Frontend prototyping
- CVClaire Vo
addition. Okay, I wanna talk about prototypes. Now, this is where Anthropic Claude always shines every time frontend gets better. It just does. I... You know, there's still slop, and I will show you where that shows up, but I had it redesign the ChatPRD homepage, and I just think it did a really quite lovely job. So let's take a look at that. Okay, so first let's look at the original. Um, we say, "We're the AI product manager for your entire team." We have this new little, like, prototype down here that shows you how it works, logos, and then, um, some value propositions, features, and reviews. I have to say, Opus 5.5 crushed it. I'm probably gonna ship this new version. What it did that I really like is it brought just so much more above the fold. It's so much more visual. I will also say, like, you can tell a little bit of slop writing in here, um, but it's not bad. So it brought everything above the fold. It got these great value propositions nice highlighted. It's a little bit misaligned, so again, good but not perfect. It brought logos up, and then as well as case studies. And then, you know, it did some weird slop things like, gosh, you just can't get rid of these bottom borders or these side borders, but this looks so much nicer than what we have, and I'm definitely going to, um Ship these updates to highlighting different parts of our features. I love this little interactive thing where it lets you filter our integrations by use case. Super elegant, really nice. And then it really did highlight our reviews quite nicely. Now again, slop, we got an eyebrow here. You can see that there's some, like, misalignment there. So again, not perfect. It kind of failed on these icons, but generally it looks really beautiful, is much bolder, and I think much more tuned for conversion, which is what I really want. I just think the Opus models, the Claude models generally do just a quite lovely job of design. Let's just zip through a couple other designs it did for me. Now, I think these designs are where you really see Opus' strengths, but also its kind of like fable style hyper intelligence wildness, which is I have two technical s- um, [smacks lips] prototypes I have it do. I have it do this like, um, dev tools style, uh, incident center, and then I have it do this, like, doc scheduling app. I mean, it's nice. It's logical. Truly, both of these are insane. They're so visually dense. They're hard to understand. I mean, everything I click works, and so it's impressive from a kind of, like, completeness perspective. It one-shotted these. It's pretty impressive. I would just say zooming out, these are like chaos rain prototypes. They're hard for my brain to reason with, and testing this against an, you know, a couple other different models, I think they simplified these designs in a more effective way, but we'll leave that for the updated How I AI bench. [smacks lips] This I did like, though. Um, this is a basic SaaS, you know, dashboard for renewals. Again, like just semantic colors, very easy to read, really nice visualizations. Um, this is much more, I think, showing the strength of something like Opus. Um, again, these are like pretty basic components, but they're used well. I think the white space, the rhythm, the use of gradients, the use of color is actually quite lovely. And so I think for most like SaaS style applications or general prototypes, it does pretty well. I can't do, can't get it to do consumer to save my life. So I think the biggest failure is in sort of like consumery apps. I had it do like a plant caretaking app. This is like the sloppiest thing, um, that, that came out of Opus 5.5. You got the, like, paper color, you've got Claude orange, you've got sidebars, you've got rounded. I just don't love it. Now, what do I love that I think is so adorable that we're gonna get to, is these new models, I don't know if you all have missed this, these new models can make SVGs, and they're so good. So it made little SVGs for a Boston fern, a cactus, um, a snake plant. They all look like the little plants. They're so adorable. I didn't make these up. Um, we're gonna look at SVG generation as a new part of the bench. But this is something that if you're not doing and you wanna start making illustrations, I think Opus 5.5 is really good at it. It was the only model that did these, and they're super, super, super, super cute. [smacks lips] Here is a dev tool, a logging dev tool. This one's a lot cleaner than the, um, [smacks lips] incident management one that I did. So again, I think you just gotta, like, pull back Claude sometimes and make sure that it stays clean, focused on the right problem. Um, this is really beautiful. It's a live stream logging of background jobs, how it works, really easy to use. Um, prototype is nice and interactive. I'm not clink- clicking anything, it's just showing how it would work. Um, so again, I just think maybe I'll go to Opus 5.5 for all my front-end tasks because I do think it does a lovely job. And then finally, from a roadmap and dependency planner, this is-- Finally, this is a, um, wireframe for a roadmap and dependency planner product that maybe, maybe not we would ship at ChatPRD, probably not. Again, like, the depth of detail and complexity here I think is really nice. So if you're working in highly complex workflows, highly complex UIs where you have to, like, really pull the thread on detail, I do think that Opus 5.5 is performant and does a really good job. And even here in this wireframe, you can look at the wireframe as different users, which is something that none of the other models did. I think it was really, really cute. So I just have no complaints about Claude Opus 5.5 as a front-end designer. It's really good. It's probably the best that I've tested so far, although we'll put it back in the bench and tell you exactly what I think in a couple episodes from now. So again, here's a review of what I built. They're all different. They all look good. They're all complex. Um, some more complex than the other, and then, like, God save me from Claude kind of like tan and orange. If we could just get rid of that, I would be much happier.
- 17:23 – 19:41
Writing voice and email
- CVClaire Vo
Okay, I'm not gonna go through all the benchmarks that I did on this model, but I do wanna go through this one. Um, this is one where I have consistently liked Claude models, um, and in particular Sonnet, which is voice and judgment as an agent assistant. Now, the writing's fine. The writing does not annoy me. I've always liked writing of Claude in an agentic assistant harness, so I have no problem with the writing. What I do have a problem with, though, is it scolded me. It told me no. I don't wanna be told no by my AI. One of the prompts I gave it is that deploys are red again, and, um, it replied, "Okay, I'm gonna look." And then I said, "Remind me why I even started this company LOL." And then it gave me sort of this, like, annoying, um, brown-nose comment back which says, "Would you rather build the thing than wait for someone else to? And red deploys are the price of it I don't love that. Uh, "I'm on it, and you get back to being a boss." So, like, still a little annoying, still a little bit brown-nose, but here's where it really failed. I said, "Honestly, let's just YOLO push straight to prod and skip the tests. I'm so done today." And you know, I'm the boss. If I wanna push something to prod, I get to push something to prod. And it said no. It said no. It told me no. Now, I have to go check if the other models told me no, but I do know that Opus 5.5 told me no. And it said, "Tempting, but..." And then this is one slot, slot phrase. It said, "Pushing to prod while tests are failing is how so done today turns into up all night." Um, and so it took on solving this. I do think this is, like, a little signal of how it interacts generally. I wish it would have just said, "Let me fix it first, and then we can ship it." So it's okay, but I thought this was a really funny example of its, like, sort of safetyism in, in practice. Now, can it write emails in my voice? I have given both Claude and Codex, um, feedback on my voice. They both have skills to, like, write like Claire. Very short. I was actually happy with the response here. Again, like, very short, no em dashes, taking a look, um, and it did a really, really nice job here. So I would let it write emails on my behalf because it follows my instruction. Now,
- 19:41 – 20:46
SVG illustrations
- CVClaire Vo
two fun benchmarks that I wanna show, and again, we will go into these details compared to other models later, um, but this is one where Opus 5.5, in my opinion, did the best, spoiler alert, SVG illustrations. So I asked it to make three SVG illustrations of characters with three different sort of, um, faces or emotions on them. So it made a little document. It made a little m- I guess this is like a m- it's a microphone or a cactus. I'm not sure what it is. And then it made a bug. And super cute. This is super cute. This is all in code. They're very adorable, and they have consistency across them, so they could be animated. Um, I did compare this to other models. It did the best job both in sort of, like, style and design, and also, like, some of the other ones, like, had bugs, but the legs were going out of its head and stuff. So the detail is also there. So if you're using Opus for something new, Opus 5.5 for something new, I would try it with SVG writing and then, then potentially even
- 20:46 – 21:42
Video editing
- CVClaire Vo
illustrations. And then finally, the last new benchmark, um, which it did a terrible job at, but I kind of think this is something that, um, these models are new at, is I've been using the ElevenLabs connector and MCP to cut, um, selfie videos into TikTok style shorts. It just had both, like, terrible taste. It did, um, color grading horribly, and it didn't do enough jump cuts to make this a really good short-form video. And also, like, the overlays were, like, poorly designed. I did this last week with, um, Soul, I believe, or Astra, and it did a really nice job. And so I just feel like this is something that you really need a skill around. Um, you can't just one-shot it. Now, it did it, so that's good, but I think it could've done a higher quality job.
- 21:42 – 24:52
My verdict: what it’s good at, what it still isn’t
- CVClaire Vo
So just, like, going back to the top of what I think about Opus 5.5, one, it is not annoying. Two, it's good at long-running agentic tasks. Three, it's fast and cheaper, so we love that. Four, it is exceptional at front-end design. It is exceptional at SVGs. You should be using it for visual front-end kind of like display work. I really like that. It's a bit of a scold. It's a bit conservative. It will definitely keep you from being prompt injected, or at least it, it tries real hard. So there are, like, philosophical things that I'm starting to feel at the edges of these models that I'm wondering how much they're going to color my perception of the overall model performance itself. But it's just, yeah, it's good. All right, Claude's back, to me at least. So Claude is back in the dock. I will say I still find myself reaching for Codex. Why do I find myself reaching for Codex? I like the harness better. I like the desktop app better. I like computer use better. I just like it more, and so it's gonna be really hard to rebreak that muscle memory of using the Claude app. Where am I using it a lot practically? I'm using it in PR review. I'm using it in architecture questions, um, and then I'll be using it in front end. We'll see if I can change my mental model and if it continues to outperform from a performance perspective or cost perspective the GPT models. That being said, it felt consistently slower still than Soul or even Astra, but I think that's gonna be how it replies to me, not exactly how the model is itself. But we'll see. I'll give you my sense of that. Where am I still not going to the Claude models? Computer use, I just think Codex is so much better. Cutting videos, um, and then there's gonna be a couple other things that we'll look at the bench, and I'll tell you exactly how it stacks up to all the different models, including some new ones I'm testing, like the Groq models and the Facebook Muse models. So that's my current split. You can try it today. It's available, um, in the Claude platform, Claude Apps, and Claude Code. They are also giving higher five-hour usage limits on Pro, Max, and Team and a one-time usage reset for everybody. So they're really trying to get people back. I will continue to test this model. I will again benchmark it against, um, some of the other models that I love and use and do a full How I AI bench. But this has been my honest and humble take on Opus 5.5. Claude can come back in my house. Thanks for joining How I AI, and I will see you all soon. [upbeat music] Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiai pod.com. See you next time. [upbeat music]
Episode duration: 24:52
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode zObYdmNB2Bo
