CHAPTERS
- 0:00 – 2:32
Going live: surprise double-drop of Opus 5.5 and GPT-6 models
Claire explains she planned a polished Opus 5.5 review but pivots to a first-ever live session because multiple model launches hit the same morning. She sets expectations: quick model highlights followed by a live, blind “vibe check” using her updated benchmark.
- •Anthropic Opus 5.5 and OpenAI model drops land the same day
- •First time doing a live review format
- •Plan: model rundown → blind benchmark → reveal and takeaways
- •Goal is to expose her real decision-making (including inconsistencies)
- 2:32 – 3:32
What launched: Opus 5.5, GPT-6 Sol, GPT-6 Luna—and where they fit
She frames these as faster/cheaper “daily driver” models rather than the absolute frontier (e.g., Astra/Fable). Pricing differences and token-efficiency improvements shape how she expects people to use them in real workflows.
- •Three refreshed models: Opus 5.5, GPT-6 Sol, GPT-6 Luna
- •Positioned as practical daily drivers for coding and knowledge work
- •Not the topmost frontier tier, but meaningful refreshes
- •Opus 5.5 costs ~2x Sol; pricing changes how she’ll choose models
- 3:32 – 4:02
Speed, caching, and token efficiency: the quiet drivers of cost and UX
Claire emphasizes that the real story isn’t only lower list price but better caching and lower token usage. She notes these improvements matter especially when resending repeated context in production apps, where cache mistakes can be expensive.
- •Lower cached-input costs and better token efficiency reduce spend
- •Caching is especially valuable for repeated context in apps
- •Operational lesson: cache optimization can be a major cost lever
- •Speed improvements affect day-to-day “feel” of the model
- 4:02 – 5:03
Opus 5.5 guardrails and alignment: capability with built-in constraints
She highlights Anthropic’s stronger cyber/bio safety guardrails arriving in an Opus-tier model. The constraints can “kick down” requests to older models or block certain classes of work, and she argues that safety posture spills into general personality.
- •Opus 5.5 ships with stronger cyber/bio guardrails
- •Certain security/bio tasks may be blocked or downgraded
- •Anthropic positions it as highly aligned with external evaluators
- •Safety philosophy influences tone and behavior even in normal tasks
- 5:03 – 7:35
Personality and ergonomics: why she stopped using Claude—and why it’s back
Claire describes why she abandoned Claude in daily use: overly scolding tone, hard-to-parse responses, and “Claude slop.” She reports Opus 5.5 feels more human and less annoying, enough to re-enter her daily rotation despite still preferring GPT models overall.
- •Past frustration: “conservative scold,” hard-to-understand verbosity
- •She moved to OpenAI/Codex/Astra/Sol for delight and ergonomics
- •Opus 5.5 improvement: cleaner bullets, more normal tone, less annoyance
- •Net: still prefers GPT UX, but Opus 5.5 is usable again
- 7:35 – 9:06
Latency vs perceived speed: narration, silence, and the ‘is it working?’ problem
She separates raw latency from perceived responsiveness. Opus 5.5 can feel slower because it narrates less; combined with “high effort” processing, silence creates uncertainty, while Sol feels faster partly due to better progress narration.
- •Sol feels faster in both latency and interaction cadence
- •Opus 5.5’s reduced narration improves slop but harms perceived speed
- •High-effort Claude behavior + quiet output can feel sluggish
- •Tradeoff: ‘shut up’ vs confidence that work is progressing
- 9:06 – 11:08
The expanded How I AI Bench: categories, scoring, and blind eval setup
Claire walks through the redesigned benchmark spanning productivity, frontend, backend, agent behavior, and creative tasks. She explains her blind test process: group comparable models, hide identities, assign 1–5 ‘vibe’ scores, and supplement with an LLM judge for correctness-heavy tasks.
- •Bench now covers: PRDs, inbox triage, frontend, backend, agents, creative
- •Blind labels (B/C/E/G etc.) and quick 1–5 subjective scoring
- •Avoids unfair comparisons across very different capability tiers
- •LLM-as-judge used where correctness and long outputs are hard to eyeball
- 11:08 – 13:40
Personal productivity test: email drafts, triage, and ‘readability’ as the killer metric
She grades multiple models on writing replies in her voice and triaging emails. Her strongest reactions hinge on clarity and formatting—especially hatred of excessive em-dashes and overly compressed summaries that lose context.
- •Best performers produce readable, low-slop drafts with sensible actions
- •Poor performers are too terse to understand or overuse em-dashes
- •She rewards models that add practical details (e.g., calendar invites)
- •Readability/voice fit outweighs raw brevity in her scoring
- 13:40 – 17:11
Frontend prototype vibe checks: dense UIs, prompt sensitivity, and “I hate them all” moments
Claire speed-runs a large set of frontend artifact comparisons (dashboards, incident tools, document ops). She notes many outputs converge on similar patterns, and that overly detailed prompts can produce dense, hard-to-grok interfaces—sometimes great technically but poor for rapid prototyping.
- •Many prototypes fail to render or look overly similar across models
- •Complex prompts often yield dense layouts that are hard to reason about
- •She favors readability and usable hierarchy over maximal detail
- •She observes model ‘tells’ (stylistic fingerprints) in UI output
- 17:11 – 21:46
B2B vs consumer UI: renewals dashboards perform; consumer ‘plant app’ mostly slops—except one standout
She finds B2B dashboard work can look quite polished, with one model excelling in color and readability. Consumer app generations feel stuck in artifact-era templates, but one plant-care app sparks genuine delight via cute illustrations and naming details.
- •B2B renewals dashboards: one model stands out with clean visual system
- •Several outputs misuse color palettes (e.g., too ‘sage healthcare’)
- •Consumer app attempts look templated and low originality
- •A single consumer result wins via charming illustrations and interaction details
- 21:46 – 25:18
Dev tools UI and backend evaluation: success on code, but presentation styles vary widely
She reviews a dev tools “jobs console” style UI and then shifts to backend tasks: auditing existing code and implementing a spec. Most models succeed functionally, so she leans on LLM judging for correctness while noting big differences in verbosity and developer-facing readability.
- •Dev tools prototypes often default to dark-mode conventions
- •Backend tasks: most models succeed, making human grading less useful
- •LLM judge used for correctness and deeper code review
- •Output style differs: some write ‘nice messages,’ others are terse
- 25:18 – 28:20
Agent personality and long-running tasks: tone, refusal patterns, and memo quality
Claire evaluates ‘agent voice’ across scenarios (including risky ops like pushing to prod) and penalizes annoying punctuation habits. For an 86-turn long-running memo task, she notices some models communicate more human-centrically and produce clearer top-line summaries.
- •Agent voice scoring: em-dashes are a recurring UX penalty
- •Most models refuse unsafe ‘YOLO to prod’ behavior
- •Long-running task outputs split into human-friendly vs terse families
- •She prefers models that summarize crisply (top issues/conflicts)
- 28:20 – 30:21
Creative stress tests: SVG icons shine; video editing disappoints
She ranks models on generating SVG illustrations (document, microphone, bug), where some show strong polish and shadows. Video editing (cutting selfie footage into shorts) is broadly poor across models, which she attributes more to workflow/skill constraints than core model intelligence.
- •SVG icon generation is now consistently usable; top models look polished
- •She scores based on recognizability, style coherence, and detail
- •Video cutting outputs are weak: overlays and cuts don’t meet expectations
- •Conclusion: video results may reflect tooling/prompting limits as much as models
- 30:21 – 38:51
Predictions, Barbie Bench, and the reveal: Opus wins breadth; Astra/Sol win her heart (and the judge disagrees)
Before revealing identities, Claire shares expectations (Claude for frontend/agents, Sol for writing). She showcases ‘Barbie Bench’—a 3D fashion-game render test that still breaks models—then reveals outcomes: Astra/Sol are her personal favorites, while Opus 5.5 scores high across the widest range; an LLM judge ranks differently, highlighting how evaluation criteria shape winners.
- •Pre-reveal hypotheses: Sol for writing/PRDs; Claude for agents/frontend; mixed on creative
- •Barbie Bench: 3D female-form rendering remains a failure mode despite progress
- •Results: Astra/Sol win her preference; Opus 5.5 ‘wins the week’ on consistency
- •LLM-judge rankings diverge from her tastes, underscoring subjective vs objective metrics
