Skip to content
How I AIHow I AI

I hate Opus 5. It’s the best model, anyway.

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version. This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me. *What you’ll learn:* 1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now 2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch 3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?” 4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard 5. The one use case where Opus 5 earned straight 5s from me 6. My actual plan for using Opus 5 going forward *In this episode, I cover:* (00:00) Opus 5 is here (03:15) First impressions (06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison (14:39) Claude Slop: the verbosity problem and why it makes my blood boil (16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring) (18:30) Live benchmark results: the leaderboard reveal (23:25) My verdict and how I’ll actually use Opus 5 *Tools referenced:* • Anthropic blog: https://www.anthropic.com/news • GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/ • Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5 • Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/ *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire Vohost
Jul 24, 202624mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
July 24, 2026
Duration
24m
Channel
How I AI
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version. This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me. *What you’ll learn:*

  1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now
  2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch
  3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?”
  4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard
  5. The one use case where Opus 5 earned straight 5s from me
  6. My actual plan for using Opus 5 going forward

*In this episode, I cover:* (00:00) Opus 5 is here (03:15) First impressions (06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison (14:39) Claude Slop: the verbosity problem and why it makes my blood boil (16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring) (18:30) Live benchmark results: the leaderboard reveal (23:25) My verdict and how I’ll actually use Opus 5 *Tools referenced:*

*Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

SPEAKERS

  • Claire Vo

    host

    Host of the podcast/channel “How I AI,” covering AI tools, models, and workflows.

EPISODE SUMMARY

In this episode of How I AI, featuring Claire Vo, I hate Opus 5. It’s the best model, anyway. explores opus 5 excels in benchmarks, but its “Claude slop” frustrates Claire argues there’s an “intelligence overhang,” predicting near-term competition will shift from raw capability to speed, cost, and usability as new frontier models arrive weekly.

RELATED EPISODES

GPT-6 Astra blew away every one of my benchmarks

GPT-6 Astra blew away every one of my benchmarks

Grok Bot + Grok 4.6  + Cursor Origin - is Claude Code dead?

Grok Bot + Grok 4.6 + Cursor Origin - is Claude Code dead?

Claude runs my entire business

Claude runs my entire business

I built an AI code review bot in 30 minutes - here’s how

I built an AI code review bot in 30 minutes - here’s how

How this OpenAI engineer uses Codex + ChatGPT Work to automate everything

How this OpenAI engineer uses Codex + ChatGPT Work to automate everything

How this “non-coder” used Cursor to add AI to retro hardware

How this “non-coder” used Cursor to add AI to retro hardware

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.