Skip to content
How I AIHow I AI

I hate Opus 5. It’s the best model, anyway.

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version. This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me. *What you’ll learn:* 1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now 2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch 3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?” 4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard 5. The one use case where Opus 5 earned straight 5s from me 6. My actual plan for using Opus 5 going forward *In this episode, I cover:* (00:00) Opus 5 is here (03:15) First impressions (06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison (14:39) Claude Slop: the verbosity problem and why it makes my blood boil (16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring) (18:30) Live benchmark results: the leaderboard reveal (23:25) My verdict and how I’ll actually use Opus 5 *Tools referenced:* • Anthropic blog: https://www.anthropic.com/news • GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/ • Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5 • Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/ *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire Vohost
Jul 24, 202624mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
July 24, 2026
Duration
24m
Channel
How I AI
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version. This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me. *What you’ll learn:*

  1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now
  2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch
  3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?”
  4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard
  5. The one use case where Opus 5 earned straight 5s from me
  6. My actual plan for using Opus 5 going forward

*In this episode, I cover:* (00:00) Opus 5 is here (03:15) First impressions (06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison (14:39) Claude Slop: the verbosity problem and why it makes my blood boil (16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring) (18:30) Live benchmark results: the leaderboard reveal (23:25) My verdict and how I’ll actually use Opus 5 *Tools referenced:*

*Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

SPEAKERS

  • Claire Vo

    host

    Host of the podcast/channel “How I AI,” covering AI tools, models, and workflows.

EPISODE SUMMARY

In this episode of How I AI, featuring Claire Vo, I hate Opus 5. It’s the best model, anyway. explores opus 5 excels in benchmarks, but its “Claude slop” frustrates Claire argues there’s an “intelligence overhang,” predicting near-term competition will shift from raw capability to speed, cost, and usability as new frontier models arrive weekly.

RELATED EPISODES

I let Codex control my browser so I don't have to

I let Codex control my browser so I don't have to

The AI content machine that turns ideas into posts that don't sound like slop | Alex Lieberman

The AI content machine that turns ideas into posts that don't sound like slop | Alex Lieberman

Local AI models explained: How to run a fleet of Mac Studios and GPUs at home

Local AI models explained: How to run a fleet of Mac Studios and GPUs at home

What is an AI harness? I build one live in less than 30 minutes

What is an AI harness? I build one live in less than 30 minutes

GPT-5.6 Sol: Better AND cheaper than Fable

GPT-5.6 Sol: Better AND cheaper than Fable

How I run autonomous coding agents from my phone with OpenAI Symphony + Linear

How I run autonomous coding agents from my phone with OpenAI Symphony + Linear

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.