Skip to content
How I AIHow I AI

I reviewed Opus 5.5 and GPT-6 Sol live - and the results surprised me

I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived. *What you’ll learn:* 1. How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it 2. Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend 3. Where Sol still wins me over on clear writing, readable PRDs, and price 4. The character SVG results that completely overturned my prediction about Anthropic 5. What happened when I asked these models to edit video, and why I think skills explain part of the disappointment 6. Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t *In this episode:* (00:00) LIVE setup and new model launches (01:30) What's new in Opus 5.5, Sol, and Luna (04:11) Guardrails, personality, and speed (09:00) The How I AI bench and blind evaluation process (11:31) Email and personal-productivity results (13:50) Frontend prototype vibe checks (24:10) Backend, agent personality, and long-running tasks (28:25) SVG illustration test (29:48) AI video-editing results (30:43) Predictions before the reveal (31:20) Barbie Bench: the 3D fashion-game test (34:17) Results: Astra, Sol, and Opus 5.5 (35:04) Writing clarity and creative surprises (36:51) Why the LLM judge disagreed with me (37:24) What each model is actually best for *Tools referenced:* Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5 GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/ Codex (OpenAI): https://openai.com/codex *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by Pen Name_ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire Vohost
Sep 22, 202638mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 2:32

    Going live: surprise double-drop of Opus 5.5 and GPT-6 models

    Claire explains she planned a polished Opus 5.5 review but pivots to a first-ever live session because multiple model launches hit the same morning. She sets expectations: quick model highlights followed by a live, blind “vibe check” using her updated benchmark.

    • Anthropic Opus 5.5 and OpenAI model drops land the same day
    • First time doing a live review format
    • Plan: model rundown → blind benchmark → reveal and takeaways
    • Goal is to expose her real decision-making (including inconsistencies)
  2. 2:32 – 3:32

    What launched: Opus 5.5, GPT-6 Sol, GPT-6 Luna—and where they fit

    She frames these as faster/cheaper “daily driver” models rather than the absolute frontier (e.g., Astra/Fable). Pricing differences and token-efficiency improvements shape how she expects people to use them in real workflows.

    • Three refreshed models: Opus 5.5, GPT-6 Sol, GPT-6 Luna
    • Positioned as practical daily drivers for coding and knowledge work
    • Not the topmost frontier tier, but meaningful refreshes
    • Opus 5.5 costs ~2x Sol; pricing changes how she’ll choose models
  3. 3:32 – 4:02

    Speed, caching, and token efficiency: the quiet drivers of cost and UX

    Claire emphasizes that the real story isn’t only lower list price but better caching and lower token usage. She notes these improvements matter especially when resending repeated context in production apps, where cache mistakes can be expensive.

    • Lower cached-input costs and better token efficiency reduce spend
    • Caching is especially valuable for repeated context in apps
    • Operational lesson: cache optimization can be a major cost lever
    • Speed improvements affect day-to-day “feel” of the model
  4. 4:02 – 5:03

    Opus 5.5 guardrails and alignment: capability with built-in constraints

    She highlights Anthropic’s stronger cyber/bio safety guardrails arriving in an Opus-tier model. The constraints can “kick down” requests to older models or block certain classes of work, and she argues that safety posture spills into general personality.

    • Opus 5.5 ships with stronger cyber/bio guardrails
    • Certain security/bio tasks may be blocked or downgraded
    • Anthropic positions it as highly aligned with external evaluators
    • Safety philosophy influences tone and behavior even in normal tasks
  5. 5:03 – 7:35

    Personality and ergonomics: why she stopped using Claude—and why it’s back

    Claire describes why she abandoned Claude in daily use: overly scolding tone, hard-to-parse responses, and “Claude slop.” She reports Opus 5.5 feels more human and less annoying, enough to re-enter her daily rotation despite still preferring GPT models overall.

    • Past frustration: “conservative scold,” hard-to-understand verbosity
    • She moved to OpenAI/Codex/Astra/Sol for delight and ergonomics
    • Opus 5.5 improvement: cleaner bullets, more normal tone, less annoyance
    • Net: still prefers GPT UX, but Opus 5.5 is usable again
  6. 7:35 – 9:06

    Latency vs perceived speed: narration, silence, and the ‘is it working?’ problem

    She separates raw latency from perceived responsiveness. Opus 5.5 can feel slower because it narrates less; combined with “high effort” processing, silence creates uncertainty, while Sol feels faster partly due to better progress narration.

    • Sol feels faster in both latency and interaction cadence
    • Opus 5.5’s reduced narration improves slop but harms perceived speed
    • High-effort Claude behavior + quiet output can feel sluggish
    • Tradeoff: ‘shut up’ vs confidence that work is progressing
  7. 9:06 – 11:08

    The expanded How I AI Bench: categories, scoring, and blind eval setup

    Claire walks through the redesigned benchmark spanning productivity, frontend, backend, agent behavior, and creative tasks. She explains her blind test process: group comparable models, hide identities, assign 1–5 ‘vibe’ scores, and supplement with an LLM judge for correctness-heavy tasks.

    • Bench now covers: PRDs, inbox triage, frontend, backend, agents, creative
    • Blind labels (B/C/E/G etc.) and quick 1–5 subjective scoring
    • Avoids unfair comparisons across very different capability tiers
    • LLM-as-judge used where correctness and long outputs are hard to eyeball
  8. 11:08 – 13:40

    Personal productivity test: email drafts, triage, and ‘readability’ as the killer metric

    She grades multiple models on writing replies in her voice and triaging emails. Her strongest reactions hinge on clarity and formatting—especially hatred of excessive em-dashes and overly compressed summaries that lose context.

    • Best performers produce readable, low-slop drafts with sensible actions
    • Poor performers are too terse to understand or overuse em-dashes
    • She rewards models that add practical details (e.g., calendar invites)
    • Readability/voice fit outweighs raw brevity in her scoring
  9. 13:40 – 17:11

    Frontend prototype vibe checks: dense UIs, prompt sensitivity, and “I hate them all” moments

    Claire speed-runs a large set of frontend artifact comparisons (dashboards, incident tools, document ops). She notes many outputs converge on similar patterns, and that overly detailed prompts can produce dense, hard-to-grok interfaces—sometimes great technically but poor for rapid prototyping.

    • Many prototypes fail to render or look overly similar across models
    • Complex prompts often yield dense layouts that are hard to reason about
    • She favors readability and usable hierarchy over maximal detail
    • She observes model ‘tells’ (stylistic fingerprints) in UI output
  10. 17:11 – 21:46

    B2B vs consumer UI: renewals dashboards perform; consumer ‘plant app’ mostly slops—except one standout

    She finds B2B dashboard work can look quite polished, with one model excelling in color and readability. Consumer app generations feel stuck in artifact-era templates, but one plant-care app sparks genuine delight via cute illustrations and naming details.

    • B2B renewals dashboards: one model stands out with clean visual system
    • Several outputs misuse color palettes (e.g., too ‘sage healthcare’)
    • Consumer app attempts look templated and low originality
    • A single consumer result wins via charming illustrations and interaction details
  11. 21:46 – 25:18

    Dev tools UI and backend evaluation: success on code, but presentation styles vary widely

    She reviews a dev tools “jobs console” style UI and then shifts to backend tasks: auditing existing code and implementing a spec. Most models succeed functionally, so she leans on LLM judging for correctness while noting big differences in verbosity and developer-facing readability.

    • Dev tools prototypes often default to dark-mode conventions
    • Backend tasks: most models succeed, making human grading less useful
    • LLM judge used for correctness and deeper code review
    • Output style differs: some write ‘nice messages,’ others are terse
  12. 25:18 – 28:20

    Agent personality and long-running tasks: tone, refusal patterns, and memo quality

    Claire evaluates ‘agent voice’ across scenarios (including risky ops like pushing to prod) and penalizes annoying punctuation habits. For an 86-turn long-running memo task, she notices some models communicate more human-centrically and produce clearer top-line summaries.

    • Agent voice scoring: em-dashes are a recurring UX penalty
    • Most models refuse unsafe ‘YOLO to prod’ behavior
    • Long-running task outputs split into human-friendly vs terse families
    • She prefers models that summarize crisply (top issues/conflicts)
  13. 28:20 – 30:21

    Creative stress tests: SVG icons shine; video editing disappoints

    She ranks models on generating SVG illustrations (document, microphone, bug), where some show strong polish and shadows. Video editing (cutting selfie footage into shorts) is broadly poor across models, which she attributes more to workflow/skill constraints than core model intelligence.

    • SVG icon generation is now consistently usable; top models look polished
    • She scores based on recognizability, style coherence, and detail
    • Video cutting outputs are weak: overlays and cuts don’t meet expectations
    • Conclusion: video results may reflect tooling/prompting limits as much as models
  14. 30:21 – 38:51

    Predictions, Barbie Bench, and the reveal: Opus wins breadth; Astra/Sol win her heart (and the judge disagrees)

    Before revealing identities, Claire shares expectations (Claude for frontend/agents, Sol for writing). She showcases ‘Barbie Bench’—a 3D fashion-game render test that still breaks models—then reveals outcomes: Astra/Sol are her personal favorites, while Opus 5.5 scores high across the widest range; an LLM judge ranks differently, highlighting how evaluation criteria shape winners.

    • Pre-reveal hypotheses: Sol for writing/PRDs; Claude for agents/frontend; mixed on creative
    • Barbie Bench: 3D female-form rendering remains a failure mode despite progress
    • Results: Astra/Sol win her preference; Opus 5.5 ‘wins the week’ on consistency
    • LLM-judge rankings diverge from her tastes, underscoring subjective vs objective metrics

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.