Skip to content
How I AIHow I AI

I reviewed Opus 5.5 and GPT-6 Sol live - and the results surprised me

I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived. *What you’ll learn:* 1. How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it 2. Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend 3. Where Sol still wins me over on clear writing, readable PRDs, and price 4. The character SVG results that completely overturned my prediction about Anthropic 5. What happened when I asked these models to edit video, and why I think skills explain part of the disappointment 6. Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t *In this episode:* (00:00) LIVE setup and new model launches (01:30) What's new in Opus 5.5, Sol, and Luna (04:11) Guardrails, personality, and speed (09:00) The How I AI bench and blind evaluation process (11:31) Email and personal-productivity results (13:50) Frontend prototype vibe checks (24:10) Backend, agent personality, and long-running tasks (28:25) SVG illustration test (29:48) AI video-editing results (30:43) Predictions before the reveal (31:20) Barbie Bench: the 3D fashion-game test (34:17) Results: Astra, Sol, and Opus 5.5 (35:04) Writing clarity and creative surprises (36:51) Why the LLM judge disagreed with me (37:24) What each model is actually best for *Tools referenced:* Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5 GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/ Codex (OpenAI): https://openai.com/codex *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by Pen Name_ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire Vohost
Sep 22, 202638mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Live blind benchmark reveals Opus 5.5 consistency, Sol/Astra charm

  1. Claire Vo reviews three newly launched models—Anthropic Opus 5.5, GPT-6 Sol, and GPT-6 Luna—framing them as faster, cheaper daily drivers amid a broader price/speed efficiency race.
  2. She highlights Opus 5.5’s new Opus-tier cyber/bio guardrails and explains how safety philosophy can bleed into general model personality and willingness in normal workflows.
  3. She runs a live, blind “How I AI Bench” across productivity, PRDs, frontend/backend coding, agent personality, long-running tasks, SVG illustrations, and video editing, scoring outputs primarily on clarity and usability.
  4. Her revealed preferences: Astra and Sol are her favorite day-to-day models for clear, enjoyable outputs, while Opus 5.5 earns the most consistently high scores across categories and becomes viable again for her daily use.
  5. A major wrinkle is that an LLM judge ranks the models differently than she does (favoring Fable and penalizing Sol), underscoring that automated evaluation may not match human-centric criteria.

IDEAS WORTH REMEMBERING

5 ideas

Opus 5.5, GPT-6 Sol, and GPT-6 Luna are positioned as cheaper, faster “daily drivers,” not frontier models.

Claire argues these releases are aimed at everyday coding/knowledge-work workloads, with emphasis on lower prices, better speed/latency, cached-input savings, and token efficiency rather than pushing absolute frontier capability.

Opus 5.5 adds stronger guardrails that can shape usability beyond “unsafe” domains.

She highlights that Opus 5.5 brings stricter safety constraints (cyber/bio) to an Opus-tier model, which can change what it will help with and may subtly affect “personality” and willingness even in benign tasks.

Opus 5.5 is a meaningful UX/personality upgrade over Opus 5—enough for her to bring it back into rotation.

Claire’s lived experience is that older Claude outputs felt overly verbose, “scoldy,” and ergonomically frustrating; Opus 5.5 improves significantly by being clearer and less annoying, though she still perceives it as dense at times.

She expands her benchmark to reflect agentic, multi-skill workflows and evaluates models via blind vibe scoring plus an LLM judge.

Her “How I AI Bench” now spans messy notes→PRDs, inbox triage and replies-in-voice, frontend prototypes, backend tasks, agent personality, long-running agentic research, computer use, SVG illustration, and video editing—then she scores outputs blind with quick vibe notes.

Her results separate “personal preference/clarity” winners (Astra/Sol) from “most broadly strong” winner (Opus 5.5).

After the reveal, she summarizes: Astra and Sol “win her heart” (most enjoyable/clear for her), while Opus 5.5 “wins her week” (high scores across many categories). She repeatedly values outputs that are easy to read and practically usable over maximal detail.

WORDS WORTH SAVING

5 quotes

The Claude slop, as I say, was Claude sloppin'.

Claire Vo

The feedback I gave to the Anthropic team is, like, not annoying. I, I found myself annoyed zero times with Opus 5.5, and when I thought I was annoyed with Opus 5.5, I was actually accidentally on Opus 5.

Claire Vo

Find somebody that loves you like GPT Sol loves a forest green or a light green.

Claire Vo

Everybody else on YouTube, all the bros are making, um, video- or video games with, like, spaceships and all different stuff. Your girl wants a Barbie fashion video game.

Claire Vo

Drum roll, please. GPT-6 Astra and GPT-6 Sol win my heart. Uh, okay, A- Astra, um, and GPT-6 I love. And then it says, "Opus 5.5 wins my week."

Claire Vo

Model launches: Opus 5.5, GPT-6 Sol, GPT-6 LunaPricing, speed, token efficiency, cached inputsGuardrails, alignment, and personality ergonomicsHow I AI Bench methodology and blind vibe scoringPersonal productivity: inbox triage and replies-in-voiceFrontend prototyping and UI “vibe checks”Backend audits vs feature builds; agent personality and long-running tasks

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.