At a glance
WHAT IT’S REALLY ABOUT
Live blind benchmark reveals Opus 5.5 consistency, Sol/Astra charm
- Claire Vo reviews three newly launched models—Anthropic Opus 5.5, GPT-6 Sol, and GPT-6 Luna—framing them as faster, cheaper daily drivers amid a broader price/speed efficiency race.
- She highlights Opus 5.5’s new Opus-tier cyber/bio guardrails and explains how safety philosophy can bleed into general model personality and willingness in normal workflows.
- She runs a live, blind “How I AI Bench” across productivity, PRDs, frontend/backend coding, agent personality, long-running tasks, SVG illustrations, and video editing, scoring outputs primarily on clarity and usability.
- Her revealed preferences: Astra and Sol are her favorite day-to-day models for clear, enjoyable outputs, while Opus 5.5 earns the most consistently high scores across categories and becomes viable again for her daily use.
- A major wrinkle is that an LLM judge ranks the models differently than she does (favoring Fable and penalizing Sol), underscoring that automated evaluation may not match human-centric criteria.
IDEAS WORTH REMEMBERING
5 ideasOpus 5.5, GPT-6 Sol, and GPT-6 Luna are positioned as cheaper, faster “daily drivers,” not frontier models.
Claire argues these releases are aimed at everyday coding/knowledge-work workloads, with emphasis on lower prices, better speed/latency, cached-input savings, and token efficiency rather than pushing absolute frontier capability.
Opus 5.5 adds stronger guardrails that can shape usability beyond “unsafe” domains.
She highlights that Opus 5.5 brings stricter safety constraints (cyber/bio) to an Opus-tier model, which can change what it will help with and may subtly affect “personality” and willingness even in benign tasks.
Opus 5.5 is a meaningful UX/personality upgrade over Opus 5—enough for her to bring it back into rotation.
Claire’s lived experience is that older Claude outputs felt overly verbose, “scoldy,” and ergonomically frustrating; Opus 5.5 improves significantly by being clearer and less annoying, though she still perceives it as dense at times.
She expands her benchmark to reflect agentic, multi-skill workflows and evaluates models via blind vibe scoring plus an LLM judge.
Her “How I AI Bench” now spans messy notes→PRDs, inbox triage and replies-in-voice, frontend prototypes, backend tasks, agent personality, long-running agentic research, computer use, SVG illustration, and video editing—then she scores outputs blind with quick vibe notes.
Her results separate “personal preference/clarity” winners (Astra/Sol) from “most broadly strong” winner (Opus 5.5).
After the reveal, she summarizes: Astra and Sol “win her heart” (most enjoyable/clear for her), while Opus 5.5 “wins her week” (high scores across many categories). She repeatedly values outputs that are easy to read and practically usable over maximal detail.
WORDS WORTH SAVING
5 quotesThe Claude slop, as I say, was Claude sloppin'.
— Claire Vo
The feedback I gave to the Anthropic team is, like, not annoying. I, I found myself annoyed zero times with Opus 5.5, and when I thought I was annoyed with Opus 5.5, I was actually accidentally on Opus 5.
— Claire Vo
Find somebody that loves you like GPT Sol loves a forest green or a light green.
— Claire Vo
Everybody else on YouTube, all the bros are making, um, video- or video games with, like, spaceships and all different stuff. Your girl wants a Barbie fashion video game.
— Claire Vo
Drum roll, please. GPT-6 Astra and GPT-6 Sol win my heart. Uh, okay, A- Astra, um, and GPT-6 I love. And then it says, "Opus 5.5 wins my week."
— Claire Vo
High quality AI-generated summary created from speaker-labeled transcript.
