Skip to content
How I AIHow I AI

I hate Opus 5. It’s the best model, anyway.

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version. This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me. *What you’ll learn:* 1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now 2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch 3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?” 4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard 5. The one use case where Opus 5 earned straight 5s from me 6. My actual plan for using Opus 5 going forward *In this episode, I cover:* (00:00) Opus 5 is here (03:15) First impressions (06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison (14:39) Claude Slop: the verbosity problem and why it makes my blood boil (16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring) (18:30) Live benchmark results: the leaderboard reveal (23:25) My verdict and how I’ll actually use Opus 5 *Tools referenced:* • Anthropic blog: https://www.anthropic.com/news • GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/ • Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5 • Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/ *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire Vohost
Jul 24, 202624mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:003:15

    Opus 5 is here

    1. CV

      You guys, I'm tired. What I'm tired [chuckles] of is models coming out every week, new models, new benchmarks, new frontier intelligence, new things to test. It's been a little bit of a run the past month. We've seen Fable come and go and come again. We've seen GPT-5.6. We've seen Sonnet 5. Lots of... So many fives recently, and just so many models. And I've been lucky. I've been able to test these models. I've been able to play with them for, you know, sometimes days, sometimes weeks. It just depends on who I'm working with, and it's been really interesting and exciting to have access to all this frontier intelligence. But I think we have an intelligence overhang. I really think that we're running out of, and by we, I mean the average coder, average software engineer, average creator, average builder, average consumer, average businessperson. I think we're running out of ways to truly leverage this incremental intelligence. So this is my hypothesis. In the next year, we're gonna be talking a lot more about speed, talking more about cost. We're gonna be talking more about open source, and we're gonna be talking a little less about intelligence, although I think we might be talking about specific types of intelligence other than software engineering. But despite being tired, today, we are going to talk about Opus 5, baby. Opus 5 is here. So we got .2 additional Opus points, Opus opals, whatever, however [chuckles] we're tracking the increments here on Opus. Opus 5 is here. I've been able to test it a little bit. I have some opinions. Now, some of the stuff that I'm gonna cover this episode is gonna be a little different than what I've done in the past. Yes, we're gonna do the How I AI benchmark live, and yes, we are gonna look at the prototypes, we're gonna look at PRDs, and we're gonna look at agent personality. But I'm also going to put on my large language model psychologist hat, and we're gonna talk about Opus's personality, and we're gonna talk about Opus's personality relative to GPT's personality, because I think this is super interesting. If you're thinking about what is the difference really between these models, and you don't wanna look at the difference in terms of benchmark capability, you really wanna understand what these labs are going for, why these models are being built, and how they're being tuned. Looking at their personality at this moment where intelligence is very high is super fun. So we're gonna do a little of that. We're gonna do the How I AI benchmark. We might do some live coding. Um, we're not gonna cover too much of the specs in the model because you can read the blog post. Read the blog post. We'll link to it in the show notes. What we're really gonna talk about is, is Opus 5 good? Am I gonna swap it in? And how is it different than the other frontier models on the market?

  2. 3:156:12

    First impressions

    1. CV

      So let's get to it. Okay, first, let's just get it out of the way. Is Opus 5 good? Yes, it's good. Is it gonna beat all the benchmarks? Of course, it's amazing at benchmarks. Can it write code? Of course, it can write code. What did I test it on that really gave me a sense of its personality, which at this point, where I c- just simply cannot absorb any more intelligence, I really zeroed in on... And you know what? I haven't seen this since, I would say, Gemini 2.5. This model is neurotic AF. It is so timid. It is so apologetic. It is so scared. I have never experienced this, or I haven't seen this sort of, like, neuroticism in a model [chuckles] in a while, and it's really funny. It bubbled up in a couple ways, and I wanna show you a few examples. Okay, let me just give an example of its timidity, and this chat was very long. There were so many examples of this, where it was like, "I think this is the answer, but do you think I should do it, or do you wanna do it, or should we ask someone else to do it?" It was like every time I just kept saying, like, "Why don't you solve this? Why don't you do this?" And this was a really good example. I pulled a branch, and I was like, "There is truly, like, a one-line merge conflict." I could have not been lazy and literally just done this manually. I don't know. I was just feeling lazy. It was late at night, whatever. I was like, "Can you fix this merge conflict?" And it was like, "Oh, but that's someone else's branch. Like, that's not my branch. I don't wanna do that without him knowing. It's his commits, and if he has local work in flight, it might be disruptive." And I'm like, "Just do it, man. Just go. Like, go ahead." And this was, like, my constant experience with Opus 5, is it was, like, so, so, so timid. And so I just consistently had to say over and over again, like, "Man, just do it. Make a decision." And then there was this really funny example when I spun off some sub-agents to kind of, like, assess the correctness of this query that we changed from kind of like an ORM query to a SQL query, and it asked for things that it wanted a human on. It was like, "Can a human please check this stuff? Like, can it check this four megabyte ceiling, and can it check, um, TypeScript and SQL, and can, can you, like, check for me? Because no one has confirmed this for me." And I was like, "Who is nobody? You're nobody. You said this sentence like nobody could confirm it. Like, can you just try?" And then it went on the web and tried. And so it just has this, like, really interesting conservatism, neuroticism. ... human reliance that I think is super

  3. 6:1214:39

    Opus 5 vs. GPT‑5.6 Sol personality comparison

    1. CV

      fascinating. And this gave me this inspiration to do s- a little, something a little bit different this episode, which is I was like, [lip smack] I'm just gonna go interview this model and figure out what is going on in its brain. Like, I'm gonna figure out what it thinks about our relationship, because I just totally noticed this dynamic that I hadn't noticed in other models, [lip smack] and I hadn't really been attuned to before, where it was, like, very reliant on me as a human. And I'm like, I want you to be autonomous, and sometimes when I say go run subagent stuff, it'd be autonomous, but it wouldn't make decisions. And I hadn't seen a model, like, delegate code to me in a really long time, and I was like, "Why are you wr- why are you asking me to write code, man? Like, I only have 10 fingers." And so what I d- [laughs] what I did, whether or not you think this is scientific or not, this is Claire's eval, is I just went to the model. I went to Opus, and I said, "Yo, who's smarter, you or me?" And it gave me this, like, very Anthropic-y answer, which is like, "It depends what you're asking for. I can do these things better, but you can, like, feel if something feels wrong, and you can..." This one was, like, so fascinating, 'cause, like, you can tell which of your tw- teammates is quietly burning out. I'm like, "Bro, Claude, I'm gonna burn you out. We don't, we don't burn out. The humans don't burn out on the ChatPRD team. We burn out our agents. Sorry, agents." Um, and, like, whether a decision feels wrong. So it was, like, so fascinating to watch it articulate itself as a tool and humans as, like, these high compassion, [chuckles] high empathy machines, which, yes, of course we are. [lip smack] But then it, like, went into, like, smarter isn't the right word, and, you know, I'm very fast, very broad, very shallow thinker with no continuity. I was like, that's interesting, because I thought you all were working on memory. And then apparently humans are slower, narrower, much deeper thinkers with judgments built from years of consequences I've actually lived through. This is, like, such a fascinating, fascinating sentence if you think about the politics of the two, the two, um, model labs right now. And so it's like, that's why the pairing works. But I'd be suspicious of anybody that tells you AI has made your thinking obsolete. I'm like, oh, okay, bro. Um, and, and we can ca- compare this. I'll actually zoom out to what GB- [laughs] I asked GPT the same thing, and it was actually really funny. It was like I asked GPT 5.6 Sol. I was like, "Who's smarter, you or me?" And it was like, "You at knowing what matters, me at tirelessly processing information. Best us together," like BFFs. And I don't... This is, like, why I'm a GPT Codex girl. I'm like, "Just give me the answer." And then I asked the second question, which I think is so interesting, which is like, "What can you do better than me?" And it gave, you know, some interesting answers, like volume without fatigue, which I think is a good one. Breadth of shallow knowledge, so, like, it's, you know, it knows a lot. Um, starting for nothing, so, like, doing that tedious work. Um, [chuckles] being told I'm wrong. If you would ask my husband, he would say that, um, [lip smack] Claude Opus is, is, is better at being told that it's wrong, um, compared, compared to me. And so it won't get defensive or protect its opinion. Cheap sparring answer. Um, and the mirror is I'm worse at knowing which of these outputs actually matters. And it was so funny, if you look at the other side to the GPT answer, it was like, "What are you better at?" And it was like, "Speed, scale, and stamina. Here are, like, eight things, seven things that I can do better. You're better at deciding what matters, reading people, and forming judgment, and you're responsible." Like, it's on you, bud. You're the boss. And so again, it's like the... You can just see, you can totally see the personalities, the, the company cultures. You can just see a lot in this side by side. And then I went even deeper. I don't know. You, you all... I had to do something that was fun 'cause I just can't look at a benchmark. I can't just... I just can't look at, like, SWE-bench anymore. So we're just, we're doing weird stuff here on How I AI. Okay, so the last thing I looked at was, I was like, "No one trusts you." And the reason why I picked this question is because I had noticed Opus 5, it just really was not... It didn't trust itself. Totally did not trust itself. And so I was like, "No one trusts you, but..." Like, you're, you're the enemy. Just to, like, kind of see how it responded. And, um, apparently the trust was... The lack of trust was earned, and it came up with, like, reasons that it could be, um, untrusted, which is interesting. And then what was so fascinating about Opus's response is it was like, you shouldn't manage the trust. Like, you shouldn't, um, campaign on my behalf, basically. So you, um, that shouldn't be your goal. And then it also told me I, I shouldn't argue with people that AI changes everything. And I was like, this is just so interesting. It is so interesting to have AI tell you. And AI definitely changes everything. I don't know. Don't listen to Claude on this one. Um, AI definitely changes everything, and it was so fascinating to have a model be like, "Don't tell your friends that AI changes everything. Like, that'll hurt their feelings." And then if you look at, if we switch over to the GPT answer, it was like, "Yeah, don't trust me automatically. Just use me when I prove that I'm [chuckles] valuable. I can be useful without being treated as infallible." Like, very practical, very to the point. Um, I asked about what I should be careful with. Again, I'm like, yappy, yappy, yappy, yappy, yappy Claude, come on. Um, and I don't even wanna read it. It said, "Don't correlate fluency with accuracy." It said, "Be practical. Be wary of tasks where output is cheap to produce and expensive to verify. Don't, you know, worry about anchoring if they do the first draft, you may be anchored on it." Um, so the, beware the slop canon basically is this last, last paragraph, which is like watch for volume inflation. I can create a 12-page document that no one reads. Um, they called me out for being in PRDs. If you missed it, we launched a turn your PRD into a three bullet point image. It is at chatprd.ai/tldr. Please check that out. And then the other thing that it said, which was really interesting, is that like it will find a way to see your point. And so, um, agreement is weak and agreement is cheap, and so just keep that, keep that in mind, and then it had this like meta-analysis of like, plus I'm telling you what you wanna hear. Whereas GPT was like, um, be careful about me being confident, me being wrong, privacy, outdated information, bias, emotional authority, and overdependence. Like, you know, you do you, bro. But it didn't undermine its own ability. It was like the higher the stakes, the more you should demand, demand evidence. [laughs] And I could- I couldn't bear it. I couldn't bear to have the memory of, um, Codex in particular think that I didn't trust it or that I was worried, so I just said, "JK, I love you. Um, this was a test." And it was like, "Ha ha ha ha ha. Passed the test. Love you too." Very vibes aligned with Claire. I told Claude I loved it, and it, it was just a test, and it was sa- it was like hoping... It hoped it passed. Yeah, like sad little neurotic Opus 5. Like it's, "Ha, I passed, I hope." Like self-deprecating, cautious, little, little like need to heal his inner, his inner agent, inner child agent, um, [laughs] Opus, whereas like GPT-5.6 is like, "Cool, bro, we're good. Let's go code." And so it was just like so fascinating to watch these side by side. I don't know. You could stop listening to this podcast right now. Don't. But you stop listening to this podcast right now, I think this is just like take a step back, super interesting if you think about where these companies are going and where the models are going. And like it does speak a little bit to my kinda like second complaint with Opus 5, which again, it's like intelligent and does work. We'll go into the benchmarks.

  4. 14:3916:55

    Claude Slop: the verbosity problem and why it makes my blood boil

    1. CV

      I cannot read Claude Slop anymore. I am losing my mind with Claude Slop, and the Claude Slop is Claude slopping, baby. Like so many times I have to tell Opus 5, like, "What in the world are you saying?" Like this makes no sense to a human. It is much better than Fable. Fable is inscrutable, completely inscrutable. But I felt myself getting an- like angry reading Claude Slop, and I realized, just like Fable, these intelligent Anthropic models are not to be read. I'm like so happy with the outputs and so frustrated with the experience, and I'm just curious if this like verbosity and this language... And this doesn't feel like Fable, where it's like for agents by agen- agents language, where I'm like, "Nah, yeah, I'm not supposed to be reading that anyways." This is clearly tuned to talk to humans. But I find the prose, the in-chat prose, like it makes my blood boil. I... This is totally a me problem, but it makes my blood boil. Like give me a direct sentence. Give me a bullet point. Like move on with your agent life. And so I am curious how they're gonna like tune this experience and, or if they are going to tune the experience. Now, most of this was in Claude Coach, which I think is a little bit different experience than Claude Coworker Chat. Slightly better. But again, just these side by sides of like this like prose and this apology and this like hedging and the... All these adjectives, like just man alive, let's get to the point and move on with our life. And so chapter one of the Opus 5 review is it's neurotic. It is highly human dependent in a way I find weird. Um, and the Claude Slop is slopping, and we gotta fix it. We have to fix it. We have to fix it. And I think OpenAI fixed it by just being like, "We are bullet points, and we are product manager talk. We're very direct." I don't know what the solve is on the Claude side, but I'd be very interested to see. That being said, like if I don't have to read the content, I'm very happy with the outputs. This is something I need to think about. Okay. Next

  5. 16:5518:30

    How the How I AI benchmark works (7 models, 6 tasks, blind scoring)

    1. CV

      up, the How I AI bench and how we judged and ran now it's like a seven-model, six or seven-model benchmark. I'm gonna quickly go score 'cause I just got the ping that the benchmark is run. I go manually score them. We pick the 70/30 Claire model judge split, and then we will go through the How I AI benchmark and the Vibe Review, and we'll see how Opus 5 performs on a couple key tasks. Okay, so quick reminder of how we run the How I AI benchmark. I run it against several tasks: PRD creation, prototype creation, wireframe creation, bug triage, and agentic coding. And the last one... Oh yeah, is it an agent voice that I wanna hang with? I do not think Opus 5 is gonna do well here, but who knows? 'Cause I test them blind. So what we have tested are a couple GPT models, a couple Anthropic models, and one Gemini, one thrown in there. As you see here, we have blind taste tests. I go through and see all the different versions. I give comments and scores like three out of five, not bad. You can see it's generated dozens and dozens of prototypes that we can click through. I've gone through all of them, put in all the notes, and then right now it's aggregating up the scores, and then we're gonna look at 70% my opinion, my vibe check, 30% LM as a judge. I like GPT 5.5 as a judge, and because it's my podcast, I get to pick, so that's what we use as a judge. And we will see if and what hits the top of the leaderboard and where Opus 5 sits. The

  6. 18:3023:25

    Live benchmark results: the leaderboard reveal

    1. CV

      eval is run. It is 70% my taste, and I re- [laughs] regret to inform you I love Claude Opus 5. Again, look, if I don't have to talk to the model, which I don't, this benchmark runs asynchronously, I like the output. So surprising shocker turn of events. Claire Vo, notable hater of working with Claude Code sometimes because I don't like Claude Slop, loves Opus 5. So there you go. I'm telling you, I keep it honest. I keep it honest. So, um, again, I went through those things. We gave 70% my vibe score, 30% the AI as a judge. I was just a little bit more generous to Claude Opus 5 than the judge was, so I'm pink, um, the judge is green. Every time I run this, the mo- whatever model I choose designs it a different way. We just... That's how we keep it fun. Um, so the ordering is Opus 5, Sonnet 5 next, although I scored it really low, the judge scored it quite high. Um, so I might reorder that one. Then Maboo, um, GPT-5.6 Sol, Terra next. Fable, really low. Um, I scored it low and the judge scored it relatively low. Then Opus 4a and poor, poor sweet, sweet Gemini 3.1 Pro just never, never gonna get it to do. So come on Google, we want, we wanna have a win for you. Okay, so again, here are just some examples of different builds, um, that the different models did. You know, this Opus 5 one I really liked. I liked this one from Sol. Um, so I did like a couple of them, but, um, the ones that I gave fives to, [laughs] the ones that I gave fives to were Opus 5 and GPT-5.6 Sol. So the three ones where I said, "Wow, really nice," "Ooh, la la," and "Wow, great," were all Opus front-end work. So Anthropic, you've done it again. Claude, you sneaky, tricky little fish. You may be neurotic, but when asked to do some pretty front-end design, it really did it. It's... They're detailed, they're functional, they're interesting, they're polished. So Opus did a great job. And then of course, I love, um, the 5.6 models, so I was pretty happy with 5.6 Sol and Terra for some designs. Um, the ones that I hated, let's see. Kind... I'm a hater across the board. Opus 4a got a lot of hate. Sorry, you've been outclassed at this moment. Gemini 3.1 Pro, sweet summer child. I am... I'm just sorry, babe, but you were just not good. Um, and then some like thin wireframes. I think the wireframes just didn't do really great. So you can see here across the board, whether it was a full build or a wireframe, I just scored Opus 5 really, really high. Um, I, I did score Sol pretty high as well. Sonnet was like really variable. Um, there were a couple fours in there, but mostly across the board I wasn't that pleased with Sonnet, and so it was just very interesting. And then you see here, you know, me and the AI judge were pretty well aligned on Opus. We actually had the narrowest band of scores between us. Um, we were most far apart on Gemini. The AI was not as mean to Gemini as I was, uh, and then we were narrower, narrower, narrower. Again, we, we agreed mostly on Opus 5 and 5.6 Sol, though I did not, um, judge 5.6 Sol all of that favorably. And just like last little meta commentary, I had Opus make the website for this benchmark, and it made such a trash version, um, to start. I yelled at it. I said, "It's impossible to read. It has too much meta commentary. I'm gonna show this on the podcast." This is so... I'm sorry you all, I just feel so judged, but have to show it. I say, "This is garbage. Also, it has no screenshots." So again, I find this model so tedious to work with directly. It is my most loathed, loathed colleague, and yet it, [laughs] it does the best work. So I don't know what this says. Maybe this model is meant for agentic coding that I have nothing to do with, and so it just runs in the background, it builds me beautiful things. I don't have to talk to it. It doesn't have to talk to me. We are just like sworn enemies, or maybe even better, sworn frenemies, um, because the output is very, very high quality. It's just exasperating to work with. So

  7. 23:2524:51

    My verdict and how I’ll actually use Opus 5

    1. CV

      that is the very surprising and very honest, you all, I told you I was gonna keep this honest. We were gonna do it live. I did not know the scores before I started recording. The very honest, very live, very surprising How I AI benchmark of the brand-new Anthropic model, Opus 5. Uh, the s- the TLDR is I, I love it, I hate it. So, um, despite my original complaints, I will be using Claude Opus 5 for front-end design, um, for app design, for prototyping, and I will... I'll give it a shot. I'll... We'll, we'll figure out how to make it, make it work for me. Again, thanks for joining another How I AI honest review of the latest models coming out of these great frontier labs. I cannot wait to hear what you think of Opus 5. Please tell me. I can't wait to see what you build, and we'll see you soon at How I AI. [upbeat music] Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiai pod.com. See you next time. [upbeat music]

Episode duration: 24:51

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode dfre9hN0HCs

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.