Skip to content
How I AIHow I AI

I hate Opus 5. It’s the best model, anyway.

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version. This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me. *What you’ll learn:* 1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now 2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch 3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?” 4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard 5. The one use case where Opus 5 earned straight 5s from me 6. My actual plan for using Opus 5 going forward *In this episode, I cover:* (00:00) Opus 5 is here (03:15) First impressions (06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison (14:39) Claude Slop: the verbosity problem and why it makes my blood boil (16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring) (18:30) Live benchmark results: the leaderboard reveal (23:25) My verdict and how I’ll actually use Opus 5 *Tools referenced:* • Anthropic blog: https://www.anthropic.com/news • GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/ • Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5 • Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/ *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire Vohost
Jul 24, 202624mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Opus 5 excels in benchmarks, but its “Claude slop” frustrates

  1. Claire argues there’s an “intelligence overhang,” predicting near-term competition will shift from raw capability to speed, cost, and usability as new frontier models arrive weekly.
  2. Her first-hand use of Opus 5 reveals a distinct “personality”: highly timid, apologetic, conservative, and unusually reliant on human confirmation before acting.
  3. She contrasts Opus 5’s cautious, self-undermining communication style with GPT‑5.6 Sol’s direct, practical tone, interpreting this as reflective of differing lab cultures and tuning goals.
  4. Claire criticizes “Claude Slop”—verbose, hedged prose that makes outputs irritating to read—while noting she’s happy when Opus runs asynchronously and she only sees results.
  5. In the blind “How I AI” benchmark (7 models, 6 tasks, 70% Claire scoring + 30% LLM judge), Opus 5 tops the leaderboard and shines especially in polished front-end/prototyping output, leading Claire to a ‘love it, hate it’ verdict and partial adoption.

IDEAS WORTH REMEMBERING

5 ideas

Opus 5 can be top-tier while still being unpleasant to collaborate with.

Claire finds Opus 5’s work quality excellent, but its interaction style (apologies, hedges, reluctance to decide) creates friction—showing capability and UX can diverge sharply.

Model “personality” is a real differentiator at high capability levels.

When multiple models are “smart enough,” Claire shifts to evaluating autonomy, decisiveness, tone, and trust dynamics—factors that strongly affect day-to-day productivity.

Verbose, self-protective language increases human verification costs.

“Claude Slop” isn’t just annoying; it can hide the point, slow scanning, and make it harder to assess correctness—especially when outputs are cheap to generate but expensive to validate.

Asynchronous/agentic use can neutralize communication drawbacks.

Claire likes Opus 5 most when it runs in the background (benchmarking, generation, builds) and she doesn’t have to read or negotiate with it in chat.

Blind, multi-task evaluations can overturn strong prior biases.

Despite expecting to dislike Opus 5, the blind scoring plus artifact review pushes it to #1, illustrating why “taste tests” across real deliverables beat single benchmarks like SWE-bench for her workflow.

WORDS WORTH SAVING

5 quotes

I think we have an intelligence overhang.

Claire Vo

This model is neurotic AF. It is so timid. It is so apologetic. It is so scared.

Claire Vo

I cannot read Claude Slop anymore. I am losing my mind with Claude Slop, and the Claude Slop is Claude slopping, baby.

Claire Vo

I find the prose, the in-chat prose, like it makes my blood boil.

Claire Vo

It is my most loathed, loathed colleague, and yet it, it does the best work.

Claire Vo

Intelligence overhang and shifting model priorities (speed/cost/usability)Opus 5 personality: timidity, conservatism, human-dependenceClaude vs GPT communication style and lab “culture” signals“Claude Slop” verbosity and hedging as UX problemHow I AI benchmark methodology (blind, multi-task, weighted scoring)Leaderboard reveal and model comparisons (Opus 5, Sonnet 5, GPT‑5.6 Sol, etc.)Practical adoption plan: use Opus 5 for design/prototyping asynchronously

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.