CHAPTERS
- 0:00 – 1:00
Why Claude fell out of favor: “smart but annoying”
Claire explains why she stopped using Claude for months despite earlier excitement around Fable and Opus. The core issue wasn’t intelligence—it was the interaction style: rambling, frustrating, and hard to work with conversationally.
- •Loved early Fable’s intelligence, then it was removed
- •Opus/Sonnet felt frustrating to converse with
- •Too much rambling and “Claude slop”
- •Abandoned Claude for most interactive work
- 1:00 – 2:01
Opus 5.5 first impression: faster, cheaper, and finally less annoying
Anthropic’s pitch is Fable-like performance at lower cost and higher speed. Claire’s bar is simpler: is it still irritating to use? Her early access testing suggests Opus 5.5 is now “minimally annoying,” bringing Claude back into rotation.
- •Claimed ~40% cheaper than Opus 5 and ~30% faster
- •Claire prioritizes usability/voice over raw benchmarks
- •Early access across coding, knowledge work, and chat
- •Conclusion up front: interaction is much improved
- 2:01 – 2:31
What Anthropic says it is: product tour, pricing, and responsiveness
Claire runs through Anthropic’s positioning and basic specs, including pricing and speed. She notes Opus 5.5 feels noticeably “zippier” than Opus 5, with one UX exception she’ll describe later.
- •“Frontier performance at a fraction of the cost” framing
- •Pricing: $4 input / $20 output (higher in fast mode)
- •Cache discounts follow typical patterns
- •Perceived responsiveness is better than Opus 5
- 2:31 – 3:01
Benchmarks (briefly): comparisons that matter less than lived UX
She quickly acknowledges benchmark charts showing Opus 5.5 beating Opus 5 and matching/competing with other frontier models. But she downplays benchmark obsession in favor of practical, day-to-day usefulness and “will it blend” usability.
- •Anthropic’s charts show gains vs Opus 5
- •Competitive positioning vs other top models
- •Skepticism about benchmark storytelling and omissions
- •Main test remains: real workflows and annoyance factor
- 3:01 – 5:02
Safety, alignment, and cyber/bio guardrails (and the “Claude scold” vibe)
Claire covers Anthropic’s emphasis on safety: external evaluations, stronger alignment, and prompt-injection resistance. She also describes the practical downside—Claude’s conservative refusals and cybersecurity task bumping—framing it as a persistent personality trait.
- •First release since “pace the frontier” call; external evals
- •Stronger alignment and prompt-injection rejection claims
- •Cybersecurity requests may be bumped to a restricted model
- •Claude’s “square/scold” personality shows up in practice
- 5:02 – 5:32
How she tested it: real-work bench categories and what she’s measuring
Claire outlines her evolving “How I AI” bench and where Opus 5.5 was applied: PRDs, prototyping, code auditing, inbox triage, email writing, computer use, and creative tasks. She sets expectations: highlight strengths, show limitations, then give a verdict.
- •Bench includes PRD writing, prototyping, codebase auditing
- •Adds knowledge-work tasks like inbox triage and email-as-me
- •Also evaluates agentic voice and computer-use workflows
- •Will share selected examples rather than exhaustive scores
- 5:32 – 8:03
Voice test: clearer, more GPT-like brevity—plus a silence/latency tradeoff
She demonstrates improved conversational formatting: concise bullets and fewer annoying flourishes. The main downside is “quiet time” on long turns—less narration, but more uncertainty about whether the model is still working.
- •Outputs are straightforward and skimmable (bullets, clarity)
- •Feels tuned to be less verbose and less “Claude-y”
- •Perceived latency worsens when it goes silent for minutes
- •UX lesson: verbosity and responsiveness shape user perception
- 8:03 – 10:36
Long-running agentic tasks: reliability, step depth, and real catches
Claire reports that all four long-running agent tasks succeeded (inbox triage, backend feature build, research, computer use). She focuses on the meaningful signals: many-step endurance, catching outliers, and resisting prompt injection, with some remaining gaps like voice matching.
- •All four long-running tasks completed successfully
- •25–82 steps per run shows endurance on multi-turn work
- •Ignored a prompt injection during inbox triage
- •Found mislinked tickets/outliers and applied policy constraints
- •Backend work caught edge cases; still benefits from cross-checking
- 10:36 – 12:37
Frontend redesign win: ChatPRD homepage refresh and conversion-oriented layout
This is where Claude shines: Claire shows a redesigned homepage that brings more content above the fold and improves visual hierarchy. She notes remaining “slop” issues like misalignment and persistent border quirks, but plans to ship the update.
- •Stronger above-the-fold layout and more visual structure
- •Improved value prop highlighting, logos placement, and reviews
- •Interactive integration filtering component stands out
- •Minor slop remains: misalignment, icon failures, border oddities
- •Net: a meaningful upgrade worth shipping
- 12:37 – 14:40
Prototyping across app types: complex UIs vs “consumer app” weakness
Claire reviews multiple one-shot prototypes—some impressively complete but overly dense (“chaos rain”). Opus 5.5 performs best on SaaS/dev-tool workflows and complex enterprise-like UIs, but struggles with consumer aesthetics (e.g., plant app) and falls back into “Claude tan/orange” styling.
- •One-shot prototypes are interactive and detailed
- •Highly dense designs can be hard to reason about
- •SaaS dashboards look clean with good rhythm and color usage
- •Consumer-style apps come out sloppy and off-brand
- •Common style artifacts: Claude tan/orange, sidebars, rounding
- 14:40 – 17:12
SVG illustrations: a standout creative capability
Claire highlights a surprisingly strong area: generating cute, consistent SVG illustrations (plants and character sets). Compared to other models, Opus 5.5 produced cleaner designs with fewer structural bugs, making it a top pick for code-based illustration workflows.
- •Generates recognizable plant SVGs (fern/cactus/snake plant)
- •Creates consistent character sets suitable for animation
- •Outperforms other models in style coherence and correctness
- •Useful new bench category: SVG writing/illustration generation
- 17:12 – 20:43
Writing voice & email assistant behavior: good mimicry, but still scolds
In agent-assistant mode, Claire finds the writing quality fine and generally non-annoying, but she dislikes the model’s tendency to moralize or refuse (“told me no”). Email voice imitation is strong—short, direct, and close to her preferred style.
- •Agent-assistant writing is acceptable and familiar for Claude
- •Can be brown-nosey and overly corrective in tone
- •Refuses risky dev instructions (e.g., YOLO to prod)
- •Email-in-my-voice output is short, clean, and usable
- 20:43 – 21:44
Video editing benchmark: works, but taste and execution lag behind competitors
Using an ElevenLabs connector and MCP workflow to cut selfie videos into shorts, Opus 5.5 underperforms. Claire cites poor taste in edits, bad color grading, weak jump-cut pacing, and mediocre overlays—suggesting this task needs specialized skill/prompting and other models currently do better.
- •Can complete the workflow, but quality is low
- •Weak editing judgment: pacing/jump cuts and overlay design
- •Color grading choices are poor
- •Other models did better on short-form editing in her tests
- 21:44 – 24:52
Final verdict: where Opus 5.5 fits vs Codex and what she’ll keep testing
Claire summarizes Opus 5.5 as “Claude is back”: less annoying, strong on agentic tasks, and exceptional for frontend and SVGs—though still conservative and sometimes slower in perceived latency. She’ll still reach for Codex for harness/app UX and computer use, while using Opus 5.5 for PR review, architecture, and frontend work, with more benchmarking to come.
- •Pros: less annoying, cheaper/faster, reliable agentic runs
- •Best-in-class for frontend design and SVG generation
- •Cons: conservative/scold vibe; perceived latency on long turns
- •Claude vs Codex split: Codex for computer use and app UX
- •Availability + increased usage limits; full bench comparison upcoming
