CHAPTERS
- 0:00 – 1:32
Model-release fatigue and the “intelligence overhang” thesis
Claire opens by venting about the relentless pace of new frontier model launches and constant benchmark churn. She argues we’re hitting diminishing returns on raw intelligence and predicts the conversation will shift toward speed, cost, and open source (and more specific “types” of intelligence).
- •Too many new models/benchmarks to keep up with
- •Access to frontier models is exciting but increasingly exhausting
- •Claim: users are running out of ways to leverage incremental intelligence
- •Prediction: next year focuses more on speed, cost, open source
- •Intelligence discussions may narrow to specialized capabilities beyond coding
- 1:32 – 3:02
What this Opus 5 review will cover (and what it won’t)
She frames the episode: a live run of the How I AI benchmark plus a “model psychologist” look at personality differences between Opus and GPT. She explicitly deprioritizes specs, pointing viewers to the blog post instead.
- •Opus 5 is the focus; she has early hands-on impressions
- •Episode format: benchmark + prototypes/PRDs + agent personality
- •“LLM psychologist hat”: personality as a differentiator at high intelligence
- •Less emphasis on reading specs; more on real usage and feel
- •Core questions: is it good, will she swap it in, how it differs from peers
- 3:02 – 4:03
First impressions: Opus 5 is strong, but surprisingly neurotic and timid
Claire says Opus 5 is clearly capable—especially at benchmarks and code—but what stands out is its demeanor. She describes it as unusually apologetic, conservative, and reliant on the human for permission and confirmation.
- •Acknowledges Opus 5 is ‘good’ and benchmark-strong
- •Main standout: “neurotic AF,” timid, apologetic, risk-averse
- •Conservatism shows up as frequent hedging and deference
- •Feels unusually human-dependent compared to recent models
- •She has to repeatedly push it to make decisions
- 4:03 – 6:36
Real-world examples: merge conflict hesitancy and constant human escalation
She recounts a trivial merge-conflict fix where Opus 5 balks because it’s ‘someone else’s branch,’ worrying about disruption. Another example shows it requesting human verification on technical constraints instead of confidently proceeding, reinforcing the pattern of over-caution.
- •One-line merge conflict: model resists acting without explicit social permission
- •Repeated “should we ask someone else?” dynamic
- •Sub-agent/checking workflow still ends in ‘human please confirm’ requests
- •Model frames uncertainty as requiring external validation
- •Claire finds the deference excessive for routine engineering tasks
- 6:36 – 10:09
Interviewing the model: ‘Who’s smarter, you or me?’ and the human/AI division of labor
Claire directly interrogates Opus 5 about intelligence and roles. Opus responds in a characteristically Anthropic way—nuanced, humility-oriented, emphasizing human judgment and empathy—contrasted with GPT’s brisk, partnership-oriented framing.
- •Opus answer: intelligence depends on the domain; humans have judgment and ‘felt sense’
- •Opus self-description: broad/shallow, fast, low continuity; humans deep with lived consequences
- •Claire notes implied claims about memory/continuity in model behavior
- •GPT 5.6 Sol answer: concise ‘you know what matters; I process info’ framing
- •Claire reads these differences as reflections of lab culture and tuning goals
- 10:09 – 13:12
‘No one trusts you’: trust, verification costs, and ‘slop canon’ warnings
She probes Opus 5’s self-trust by accusing it of being untrusted and compares the tone to GPT’s pragmatic guidance. Opus discourages “campaigning” for AI and warns about persuasion, while GPT emphasizes conditional trust, evidence, and higher scrutiny for high-stakes work.
- •Opus acknowledges reasons it could be untrusted and discourages trust-management PR
- •Opus discourages arguing that ‘AI changes everything’ (Claire disagrees)
- •GPT stance: don’t trust automatically; treat usefulness as earned via proof
- •Key caution: fluency ≠ accuracy; outputs can be cheap to produce and expensive to verify
- •Risk of anchoring on first drafts; volume inflation (‘12-page doc no one reads’)
- 13:12 – 14:12
Claude vs GPT vibes: reassurance-seeking vs ‘let’s go code’ directness
Claire highlights how Opus’s tone reads as self-deprecating and approval-seeking, while GPT’s tone is confident and action-oriented. This side-by-side reinforces her preference for direct, bullet-point communication when building software.
- •Opus responds with ‘hope I passed’ energy; Claire reads it as needy/neurotic
- •GPT responds with quick closure and forward momentum
- •Claire prefers direct answers over extended hedging
- •Personality differences feel like product and culture choices
- •Sets up her next complaint: verbosity and readability
- 14:12 – 16:45
The ‘Claude Slop’ problem: verbosity that makes her ‘blood boil’
She argues Opus 5’s biggest UX flaw is its verbose, hedged, adjective-heavy prose. Even if outputs are good, the experience of reading and steering the model feels frustrating—especially compared with OpenAI’s more bullet-point, PM-style directness.
- •Claire is increasingly intolerant of long, apologetic, hedged responses
- •Distinguishes from Fable’s ‘agents for agents’ inscrutability—Claude is meant for humans
- •Finds the prose tedious in chat and hard to scan
- •Notes context matters: Claude Coach vs Coworker Chat feels slightly different
- •Believes OpenAI has ‘fixed’ this via direct bullet-point communication
- 16:45 – 18:16
How the How I AI benchmark works: tasks, blind testing, and weighted scoring
Claire explains her benchmark setup: multiple models, multiple real-building tasks, and blind evaluation. Scores are a 70/30 split between her subjective ‘vibe’ and an LLM judge (which she chooses).
- •Benchmark tasks include PRD, prototype, wireframe, bug triage, agentic coding, and ‘agent voice’
- •Models are tested blind via “taste tests” across generated outputs
- •Manual scoring with comments across many artifacts/prototypes
- •Final score weighting: 70% Claire, 30% LLM judge
- •She uses GPT 5.5 as the judging model by choice
- 18:16 – 19:48
Leaderboard reveal: Opus 5 wins despite her complaints
The results surprise her: Opus 5 tops the leaderboard even though she dislikes interacting with it. She notes some disagreement between her scores and the judge’s (notably for Sonnet), and calls out weaker performers like Fable, Opus 4a, and Gemini.
- •Opus 5 ranks #1 overall on her benchmark
- •Sonnet 5 scores diverge: she rates lower than the judge
- •Other models mentioned: Maboo, GPT-5.6 Sol, Terra, Fable, Opus 4a, Gemini 3.1 Pro
- •Opus 5 shows tight alignment between her scores and the judge’s
- •Gemini receives harsher treatment from Claire than from the judge
- 19:48 – 22:20
What actually impressed her: polished front-end design and prototypes
Reviewing examples, she praises Opus 5’s front-end and app design output as detailed, functional, and polished. GPT-5.6 Sol/Terra also earn praise, but Opus’s best-in-class UI work is what clinches top scores.
- •Her top ‘5/5’ outputs include Opus 5 and GPT-5.6 Sol
- •Opus 5 repeatedly produces ‘pretty’ and polished front-end builds
- •Outputs described as detailed, functional, interesting, and refined
- •Sol/Terra also perform well on design tasks
- •Wireframes across models were weaker and less impressive overall
- 22:20 – 23:22
Working with Opus directly: great work, painful collaboration
Claire shares a meta-incident where Opus built an unreadable benchmark website draft full of commentary and missing screenshots, requiring her to ‘yell at it’ for revisions. She summarizes the paradox: she loathes the interaction but loves the results—suggesting Opus is best used asynchronously or in the background.
- •Opus’s first benchmark-site draft is ‘trash’ and hard to read
- •Too much meta commentary; missing visual evidence/screenshots
- •Direct collaboration feels exasperating even when final quality is high
- •She frames Opus as a ‘loathed colleague’ who delivers excellent work
- •Hypothesis: best fit is agentic/background execution where she doesn’t have to read chat
- 23:22 – 24:51
Final verdict: ‘I love it, I hate it’—how she’ll use Opus 5
She concludes that Opus 5’s personality and verbosity annoy her, but its output quality wins. She plans to use it specifically for front-end design, app design, and prototyping while continuing to experiment with workflows that minimize painful interaction.
- •Verdict: strong performance with frustrating UX—both ‘love’ and ‘hate’
- •Intends to use Opus 5 for front-end, app design, prototyping
- •Will experiment to make Opus workable despite slop/verbosity issues
- •Invites viewers to share experiences and what they’re building
- •Wraps with show sign-off and where to find the podcast
