At a glance
WHAT IT’S REALLY ABOUT
Opus 5 excels in benchmarks, but its “Claude slop” frustrates
- Claire argues there’s an “intelligence overhang,” predicting near-term competition will shift from raw capability to speed, cost, and usability as new frontier models arrive weekly.
- Her first-hand use of Opus 5 reveals a distinct “personality”: highly timid, apologetic, conservative, and unusually reliant on human confirmation before acting.
- She contrasts Opus 5’s cautious, self-undermining communication style with GPT‑5.6 Sol’s direct, practical tone, interpreting this as reflective of differing lab cultures and tuning goals.
- Claire criticizes “Claude Slop”—verbose, hedged prose that makes outputs irritating to read—while noting she’s happy when Opus runs asynchronously and she only sees results.
- In the blind “How I AI” benchmark (7 models, 6 tasks, 70% Claire scoring + 30% LLM judge), Opus 5 tops the leaderboard and shines especially in polished front-end/prototyping output, leading Claire to a ‘love it, hate it’ verdict and partial adoption.
IDEAS WORTH REMEMBERING
5 ideasOpus 5 can be top-tier while still being unpleasant to collaborate with.
Claire finds Opus 5’s work quality excellent, but its interaction style (apologies, hedges, reluctance to decide) creates friction—showing capability and UX can diverge sharply.
Model “personality” is a real differentiator at high capability levels.
When multiple models are “smart enough,” Claire shifts to evaluating autonomy, decisiveness, tone, and trust dynamics—factors that strongly affect day-to-day productivity.
Verbose, self-protective language increases human verification costs.
“Claude Slop” isn’t just annoying; it can hide the point, slow scanning, and make it harder to assess correctness—especially when outputs are cheap to generate but expensive to validate.
Asynchronous/agentic use can neutralize communication drawbacks.
Claire likes Opus 5 most when it runs in the background (benchmarking, generation, builds) and she doesn’t have to read or negotiate with it in chat.
Blind, multi-task evaluations can overturn strong prior biases.
Despite expecting to dislike Opus 5, the blind scoring plus artifact review pushes it to #1, illustrating why “taste tests” across real deliverables beat single benchmarks like SWE-bench for her workflow.
WORDS WORTH SAVING
5 quotesI think we have an intelligence overhang.
— Claire Vo
This model is neurotic AF. It is so timid. It is so apologetic. It is so scared.
— Claire Vo
I cannot read Claude Slop anymore. I am losing my mind with Claude Slop, and the Claude Slop is Claude slopping, baby.
— Claire Vo
I find the prose, the in-chat prose, like it makes my blood boil.
— Claire Vo
It is my most loathed, loathed colleague, and yet it, it does the best work.
— Claire Vo
High quality AI-generated summary created from speaker-labeled transcript.
