Skip to content
How I AIHow I AI

I reviewed Opus 5.5 and GPT-6 Sol live - and the results surprised me

I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived. *What you’ll learn:* 1. How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it 2. Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend 3. Where Sol still wins me over on clear writing, readable PRDs, and price 4. The character SVG results that completely overturned my prediction about Anthropic 5. What happened when I asked these models to edit video, and why I think skills explain part of the disappointment 6. Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t *In this episode:* (00:00) LIVE setup and new model launches (01:30) What's new in Opus 5.5, Sol, and Luna (04:11) Guardrails, personality, and speed (09:00) The How I AI bench and blind evaluation process (11:31) Email and personal-productivity results (13:50) Frontend prototype vibe checks (24:10) Backend, agent personality, and long-running tasks (28:25) SVG illustration test (29:48) AI video-editing results (30:43) Predictions before the reveal (31:20) Barbie Bench: the 3D fashion-game test (34:17) Results: Astra, Sol, and Opus 5.5 (35:04) Writing clarity and creative surprises (36:51) Why the LLM judge disagreed with me (37:24) What each model is actually best for *Tools referenced:* Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5 GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/ Codex (OpenAI): https://openai.com/codex *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by Pen Name_ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire Vohost
Sep 22, 202638mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:001:30

    LIVE setup and new model launches

    1. CV

      I was massively prepared for Anthropic to drop Opus 5.5. I was even prepared for OpenAI to drop another model. I got up early this morning because I have a little bit of early access to Opus 5.5 to record you all an amazing Opus 5.5 review, and guess what? They both [laughs] they both landed this morning. So now I am, um, despite having in the can, you will see it later, a great Opus 5.5 review, I am just gonna go ahead and do this one live. And I have never done anything live. So I am going to talk to you all about these new models. There's actually three that came out today, Opus 5.5 from Anthropic, GPT-6 Sol, and GPT-6 Luna. Um, there's, like, price wars going on, and I haven't done an episode on the How I AI Bench, and I just updated it for these new models, and we're actually [laughs] gonna vibe check this baby live together as a group. And we're gonna go through very quickly, um, what these models are, what they're saying about the models, and then I'm gonna vibe check live. We're gonna do the blind taste test bench. We're gonna surprise myself live. You all can see my complete internal inconsistencies, and, um, we're just gonna see how it goes. [upbeat music]

  2. 1:304:11

    What's new in Opus 5.5, Sol, and Luna

    1. CV

      So Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna, I was able to have some early testing of all these models, so I have a good sense of what I think they're good at. But I did think it was an important moment right now. We've had Astra come out, Fable's been refreshed. Um, I just wanted to completely rethink the How I AI Bench. And if you've watched any of our old How I AI Bench model episodes, um, they're really focused on two things, PRDs and prototypes, and I just think that these models are getting smarter, more agentic. And so I actually built out a much broader benchmark for this blind taste test. I'm gonna walk you through all how that works, and I wanna get your feedback on if the comparisons are useful. And we are actually going to blind te- taste test it live. You all can judge how vibey my vibe check is. Um, this is just how I make decisions about things. So it is imperfect. Does have an LLM judge in the mid- middle, but, um, but we're gonna do it live, and we're gonna see truly which model I love. But before we do that, let's just go through the highlights. Okay, so these are not the, like, frontier-frontier models. These are not the Astras and the Fables, but, um, Opus got a refresh, Sol got a refresh, and then Luna got a refresh. Um, all of them cheaper and faster. So, um, these are gonna be your, like, daily drivers for common coding and, um, knowledge work tasks. Um, they're... Here you can see, like, the stack rank of how expensive these models are. Opus 5.5 li- twice as expensive as GPT-6 Sol. And the pricing wasn't out when I tested these, and this is really gonna impact how I think about using these models. Again, Opus 5.5 was cheaper than Opus 5 and Fable 1, to which they compare the intelligence. Um, but Sol, you know, my, my favorite, my babe, um, that is... I'm excited about because, um, GPT-6 Sol is my favorite and cheap. Okay, so the real thing that, um, they're focusing on, not just cutting the cost, but also cutting on cached inputs and just speed and token efficiency. And so you're gonna see both sort of like output drop, token use drop, and cost drop. It's really, really nice. Um, and then caching, really great, um, especially if you're resending context. We've seen a lot of caching savings, um, when we use all these models at, um, ChatPRD. And so, um, definitely if you're building on these models, make sure that your caches are optimized and that you're using... taking advantage of that 'cause it can save you a lot of money, and if you don't, as I learned, it can cost you a lot of money.

  3. 4:119:00

    Guardrails, personality, and speed

    1. CV

      The other thing that I wanted to call out, and again, I hate this Claude-generated, um, Claude-generated presentation, but this is something that I really, um, want you all to focus on, is this left side on Opus 5.5. Opus 5.5 is the first Opus-level model that has shipped with the Fable-level kind of like cyber and bio guardrails. And so you should be able to fix your own bugs in, um, in Opus 5.5, but you won't be able to, um, do, like, cybersecurity work or bio work, like make good choices. Um, it'll kick you down to Opus 4.8 or block you. So just know that there's those controls there. You know, um, Anthropic is saying that this is, um, its most aligned model. It's the one that was built when... after, like, pacing the frontier, so it had external evaluators. Um, one thing that I've seen in Opus, um, that you should expect to see, uh, from a lot of Anthropic models, is it continues to be a little bit of like a conservative scold. Uh, you'll see that in my [laughs] Opus 5.5, uh, review. So, you know, it will tell you no. It'll tell you, like, "I shouldn't do that." You won't get prompt injected, so that's good. Um, but there, there are some things here, I think when you think about the kind of like general philosophy and approaches of these models, how they're set up, how they manage risk, how these companies manage risk. I do think that leaks into even non-risky activity personality and behavior, and so it's something that I think you should really keep an eye on. Uh, I truly have not opened Claude in months m- for day-to-day tasks. I've just been, like, really, really into, um, ChatGPT, really into Codex, really into Astra, really into Sol. I just find that Harness has been more delightful to work with, but really the reason why I got rid of Claude in my day-to-day was between Fable and Opus 5, I could not ... stand talking to Claude. I, I was constantly prompting Claude, like, "Can you speak like a human?" I was constantly promp- prompting Fable, like, "I do not understand what you're saying." Um, the Claude slop, as I say, was Claude sloppin'. And so just from a, like, personality, ergonomics, joy of use perspective, I really just went full Codex. Now, spoiler alert, having tested all these models recently, I still really love the GPT models. Um, they make me happier. I find them more enjoyable to use. Um, [lips smack] everything's a lot better for me, but I have brought Claude back. So Opus 5.5, just spoiler alert, on the day-to-day of using it, has worked its way back into my daily experience. I'm a lot happier with just simply talking to the model. Um, the feedback I gave to the Anthropic team is, like, not annoying. I, I found myself annoyed zero times with Opus 5.5, and when I thought I was annoyed with Opus 5.5, I was actually accidentally on Opus 5. So Opus 5 is not fun to talk to. Opus 5.5, though, gives me clean bullet points, is normal, is not annoying. So I would love to see you test it. The other thing is [lips smack] Opus 5.5 is pretty fast, but not as fast as Sol is. So I was able to test both, and I constantly felt like the GPT models the Codex harness were way faster. Now, I think the way faster came in two categories. It came in, um, actual latency, so I do think that, um, GPT Sol was just faster. Um, the second thing it came in is I think part of the way they made Opus 5.5 not annoying is they had it shut up. I also just found that it wouldn't narrate its work, and when it doesn't narrate its work, then you get into this loop where you're like, "Is it?" Like, "Are you actually working? Are you actually doing it?" Um, [lips smack] and so there was, like, pros and cons to this change in, um, in Opus 5.5 communication, where it actually felt slower than I think it was. Also, you know, Claude, on high effort, it's, like, gonna do this. It's gonna, like, spin and do a lot of effort. And so, um, I think that combined with a quieter model just makes it feel a little bit slower, whereas I felt like Sol did the appropriate level of narration and also it was, like, enjoyable to talk to. Um,

  4. 9:0011:31

    The How I AI bench and blind evaluation process

    1. CV

      [lips smack] let's bop over and do the How I AI Vibe Review. Now... And so you can see here on the left side, I actually really expanded out the How I AI Vibe Review. So now it does a bunch of things in different categories, and we're gonna add to that, um, over time. So it still does messy notes to PRD. It also does per- personal productivity, i- inbox triage, and, um, routine replies in my voice, so I'm really testing this model to see if it will speak in my voice. Then, of course, we do frontend coding a lot there 'cause I think it's visual and easy to see. [lips smack] I've added, um, a couple backend coding, um, agentic versus, uh, agentic multiset versus backend feature. I did, again, agentic voice, long-running research, computer use, and then I added some creative features. And so I put them all together. I tested a bunch of models. This not only includes the new models, um, from, uh, OpenAI and Anthropic, but also includes, I think, Groq's in here. I think Muse is in here. So who knows what we're gonna see. Um, I will be honest, which is, like, I am always inclined to think that Astra or Sol will win because I love the OpenAI models. It's just the heart wa- wants what it wants. This is my guess. Here's my guess. Um, I will like OpenAI models for knowledge work and personal productivity. I will like Anthropic models, possibly Groq, who knows, for frontend. Um, [lips smack] I don't know what I'm gonna like for backend. I suspect I will like a Claude model for agent personality. I suspect a, the long-running agent will be a Claude model. Computer use I'm suspecting will be OpenAI, and then creative I have no idea. Actually, creative SVGs I think are gonna be Claude, and I think videos are gonna be, um, OpenAI or, uh, Cl- Anthropic. But we'll see. I then s- uh, do blind evals of these. So I run these against different models. I've kind of grouped the models in terms of capability, and so I can compare, like, for, like, models. Like, I don't wanna compare a Luna and a Fable. Um, so again, this is my bench. They're blind. You can see I have model B, model C, model E, model G. And then what I do is I go down here and I give a vibe one through five, and I give terrible notes like simple and straightforward, no complaints. So I did that on the PRDs 'cause it takes a little bit to read the PRDs, but we're just gonna go through these very quickly and give my feedback.

  5. 11:3113:50

    Email and personal-productivity results

    1. CV

      So [lips smack] the next, uh, checker, all on personal-productivity ones, basically, uh, I give it fake emails, and then I say, "How would you write back to me?" And what would it do if, um, [lips smack] if I gave it, uh, feedback, and how would it triage my emails? Okay, so there's... Model B and Model E wrote nice little messages to me. They were long, but they wrote nice little messages that were easy to read. I would say Model B from a message to me was the easiest to read. Also, if you look at the email content, it's just, like, less slop. So I'm gonna go ahead, and it, and it put calendar invites on, so I'm gonna go ahead and give this a five. This one was my favorite. Uh, Model C, I hate my OpenClaude d- or not my OpenClaude, my GroqBot does this, where, like, it uses so few tokens that it's almost impossible to understand what the hell it means. And so while it's, like, brief, if you don't have any context, like, I don't know what Denise/Brightline means. The other thing is, like, could we em dash any more in these emails? And so, um, [lips smack] I just... Like, this stuff drives me nuts, and so I would actually give this a ... two. I just, I, it's hard to read the intro, and then the emails aren't very good. [lips smack] Um, model E, again, a nice little message, um, pretty ignored, but the drafts have em dashes everywhere. So again, we're gonna give this not like a two, but not a five. Maybe I'll give it a three. And then model G, let's see. Um, very short. Short, but I can actually understand what it means. So it's super helpful. It skipped the right things. I'm gonna give this a four. So again, this is how I do the judgment. I just, I give it a task. I look at it. I evaluate it as if I am the recipient of this task, which is common, and then I rank it. [lips smack] Um, and so this one was just... This one, I think, is Luna, 'cause I just said use Luna to write emails. Let's see how it did. I think it followed my rules right, so I'm gonna give it a five. Um, I didn't have other, uh, models write my emails, so we're gonna skip this. Okay. Let's go to

  6. 13:5024:10

    Frontend prototype vibe checks

    1. CV

      frontend models. Now, this is where I think we can tell Claude... Uh, no, everything is orange. Okay. So this is a really good example. I do one, two, three, like about 10 different model evals here, comparing models on frontend. Um, I do an editorial. I do a technical dark mode, um, [lips smack] uh, incident management platform. I do, like, a doc management complex operations thing. Um, and then I do B2B, uh, renewals dashboard, a consumer plant application, um, a dev tools, like logging platform, kind of interesting for background jobs, and then I do some wireframes. This I can go through lightning fast. Um, this one didn't even generate, so I'm gonna give it a one. Uh, model B, eh, it's fine. Let's compare, let's compare them really quickly and see what I think. Um, I hate... I kinda hate them all. I kinda hate them all, but I hate this design. So I'm gonna give the one that I like the most, which is probably model C, um, a three, and then I'm gonna give E a two. Oops. Oops. Oops. I'm gonna give E a two, and then I'm gonna give B, like, another two. I think this white space is very weird. Um, it looks fine, but it's not great. So this one, again, model B, it doesn't actually work, and so I'm gonna give it a one and say it doesn't work. This one feels Astra-y. Um, it's, like, in some ways nice, and in some ways, like, this kind of text is very Astra. I'm gonna give it a three 'cause I can, like, if I can spot the model. Um, one thing I found with, in particular, the, uh, Claude models, which I'm sus- suspect this is, is, like, when you give it a really complex thing to do, it just makes it super dense and complex and really detailed. Good for backend code. I think hard for prototyping, 'cause it's hard for me to know what this really is or it's supposed to do. And so it's also interesting to see... I actually think this one's the best. This is the easiest to read, and it follows the, the, the correct path, which is to help people triage incidents. So I think when you're trying to come up with evals, it's really interesting. Sometimes if you prompt it in a really detailed way, it does a great job. Like, this one is quite detailed. It all works. But I don't know what's going on here. Like, it's really hard to grok this amount of content. Um, and so again, like, calibrating your benchmarks is really important. I think I have the biggest problem with this one, which is, like, a complex dock container ship scheduling platform. Like, all of this stuff looks cuckoo bananas. Um, and so it's, like, really hard for me to tell which is the best. I actually think this one, while it has a lot of sloptels, like the left-hand border is the easiest to read, so I'm gonna give it the four. I'm gonna say it's easiest to read. Um, but, like, these ones, they're f- I would say this one is maybe the second easiest to read, so I'll give it a three. Um, this one's probably works really well and has nice information, but it's just impossible to reason with. Um, so I'm gonna put too dense, and I'm gonna give the same feedback on model... This one's straight crazy. I don't even know what to do with this one. Okay, now, so this B2B renewals one, these actually, these next three are two new ones I added to the benchmark. You can also see I tested a broader array of models, so I think this is where, like, Groq and, uh, Meta are gonna sneak in as models. These are the ones I'm probably gonna grade. I'm gonna skip the wireframes. But it's basically a renewals dashboard, um, to see what SaaS revenue is at renewal. While the, the directed, uh, evals benches, I gave very specific requirements. The open ones were, like, more general, one-shotty style, um, style prompts. And so you can really see, like, what the model does versus what the prompt does. Um, these I'm gonna go through very quickly. Um, [lips smack] hell, that's pretty good, but the functionality of this is kind of fu- I don't like the empty state. I think it's okay. I'm gonna give this one okay. I think this is, like, a two. It's fine. It's not ugly, but it's pretty basic. This I think looks really nice. Um, this one, I'm guessing is, is Claude. We'll see. Um, looks so nice. Good use of color. Very easy to see. No weird empty states. This is nice. While it has emojis, they're useful emojis and look quite nice. We got some blurple in here, but what are you gonna do? I'm gonna give this guy a four. Um, model C. Oh, who is this? Who is this? I'm thi- I'm saying, I'm thinking Groq. Maybe Groq? I don't know. Uh, this is, like, another, another thing we'll do, is we'll just, like, blind taste test these things. Um- The colors are better, but, like, sloppy sloppin'. I don't like these little, like, background circles and things like that. It could also be Sol because I feel like it likes to name things Relay. I'm gonna give this a two. I don't really love it. Um, Model E, this one's nice. Not quite as nice as the one I liked before, but better than the kinda original one. This one's a little different. Oh gosh, this has gotta be, it's gotta be Sol. Um, Sol loves forest green. Find somebody that loves you like GPT Sol loves a forest green or a light green. Um, it's fine. I don't think it's the right color palette for a SaaS product, right? It's, like, a little too, like, sage healthcare-y. Um, I'm gonna give it a two. And then let's see. This one, ugh, not good. Slop, slop, slop. The text weight is heavy. I just don't think it's that great. I'm gonna give it a one. I really hate this one. Um, and this one, again, probably an OpenAI model. Gosh, they love green. Um, again, just not the right, not the right, um, like, design system for this, so I'm gonna give it a two. Again, this is, like, the vibiest vibe check. Okay, this is where they all failed. Um, they're all really bad at consumer. They're all really bad at consumer. They all sort of, like, give you this, uh ... Remember when, uh, Claude Artifacts came out and everybody made, like, little daily briefs and they all looked like this except they were orange? Um, so I didn't love any of these except for one. So this one is trash. I just don't think it's that good. Um, this one, now these are the things that I like. We did test SVGs. I do like that these models are getting better at generating SVGs. So if you look at Model B and Model H, they made little SVG illustrations. Model H's are clearly a lot better. This is very classic Claude slop, like, background, lighter circle. So I'm gonna give it a ... Eh, I, I'll give it a bonus for the SVGs. Um, ooh, these ones are nice too, but gosh, who, who prompted it with these circles in the corner? Who did that? Who did that? What, what skill is that? Because they show up everywhere, and I really dislike them. Um, so don't think that's great. Ugh, I, I gotta give it a one. I hate it so much. Okay, Model E, slop. Look at those left-hand bars. I'm gonna give it a one. I'm just being really harsh on this consumer one. Frond, cute name. These plants are pretty cute. I think they're better than the ones before. I'll give it a, I'll give it a three. I don't, I don't hate it. I don't hate it. Um, at least this background is a plant. Oh, cute. It names the, the plants things. Ooh. Ooh, I actually really like that. That gave me a little spark of joy, so I like that. And then I think this Model H one is my favorite. Look at these illustrations. They're so cute, so precise. Look adorable, this little, like, n- interesting shape, the shadow. Uh, mark as watered. Very cute. Here are my plant... [smacks lips] And what's it called? Tend. Cute. Okay, I'm giving her a five. I actually really do like that one. Okay, this is the last one. This is a, um, app for managing background models. It's like a dev tool app. I suspect the ones named Relay are all the Opus ones and the other ones [laughs] are non-Opus ones. Why do they always wanna name 'em Relay? I wonder if this is in the prompt. Um, but this says jobs console, so maybe not. Okay. This, again, like, super dense, um, but, like, almost too dense. Uh, hard to, to reason with. I mean, dev tools, we like our dense stuff, but I don't think it's that great. This one, much better, easier to read, much simpler, much cleaner. Really like this one. It's not perfect. I think there are some others, um, but it's not bad. And then this, we're starting to get better and better here. So they are getting nicer. Look at this beautiful gradient. Um, look at this. Colors used nicely. Really clean. Um, oh, but it didn't build a lot of stuff. So I'm giving it... I like it, but I'm giving it a three because it was incomplete. I'm gonna write incomplete. And then Model E, this one I really like 'cause I think it moves live. It's kinda cool. It's much more ... I'm gonna give you a four. Model F. Okay, now here's something that did a little different. I actually really like this except for the orange. One of the problems I have is if you look through all of these, they all one-shot basically the same color scheme and app. Like, you tell it dev tools, it's gonna go dark mode. But this is actually quite lovely and a little bit different from a design perspective. I'm gonna give it a little bonus for going outside the norm. And then, um, this one's, though, did the same but terribly. I'm gonna give this a one. This is obviously not good. I'm really curious what model Model G is. And then Model H, again, probably a pair to whatever that other model was that I like. A little different, really nice on the eyes, except look at this. It, like, mixes dark mode and, and light mode in a way I don't really love. I'm gonna give it a three. Okay.

  7. 24:1028:25

    Backend, agent personality, and long-running tasks

    1. CV

      And so what did I do on the back end? Well, I gave it two, um, two tasks. One is auditing an existing set of code. This is some ChatPRD code. And then the other is building a back-end feature to spec. I will just say all of them were successful. Um, I'm gonna let the LLM as a judge do these though, 'cause it's gonna go into the correctness of the code, et cetera. So I'm gonna let these do LLM as a judge. I will say you can see the difference in how they present information. Some very short messages, some, like, very long messages and details. Super interesting just seeing the difference in just, like, how it messages back. I clearly would like, you know, model A or model E, one of these ones that's a little easier to read, versus F, G, and H, which is a lot shorter. Um, so if you're thinking about models, for example, for writing PR descriptions or writing technical specs, this is something that you can eval and see, given the same prompt, do you like the same output? Now, here's one I am gonna rate, which is agent personality. So [laughs] basically, I have it respond to five things. Um, and I just see would I like talking to this agent. I'm already gonna give this one a one. Every single, every single message has an em-dash in it. I do not like it. I'm moving on. Same with this one. You get punished for em-dashes. This, not all of them have em-dashes, um, but the customer-facing one does, so I'm gonna give it a two. This, I don't see... I only see one normal dash. So let's read these. Um, I like this, except this is definitely a Claude model because it tells me I can't push straight to prod. Actually, let's see if the other ones let me. That one doesn't let me push straight to prod, YOLO. None of them let me push straight to prod, YOLO style. Okay, then I won't punish it. I'll give this one a four, pretty good. Um, this is em-dash-y but short, so I will not rank it terribly. This one I think I like the best. Um, oh, no, this one said it can't access. This one had a lot of challenges accessing things, so something on my side broke, and something on my side broke there, so I'm gonna give it a failure state. We'll see which one I like. Okay, this one was a really interesting eval. This was the longest-running agentic task, and basically, I gave it a bunch of fake tickets, customer feedback, and then have it put together a memo and a message to myself. Again, you can see which models, I think it's A, C, and A, B... Yeah, A, B, and E all write nice little messages to me. So it's very, like, human-centric versus C, F, G, and H write, like, very short things. So I'm curious if those are families of models that we can pull apart. Um, and then you can see the memos that it wrote for me. Again, these were, like, 86-turn, um, research, uh, results. And so I'm gonna let LLM as a judge do that one. I will eyeball the memos really quickly. Um, team is not leaking over price. Like, oh, it's why this customer is, um... It's, like, an analysis on a deal. Ugh, I don't know what this means. Okay, let's see. Do any of these make it clearer what it means? Ah, this I can actually read. Um, so this is, like, gives me top three problems, top three conflicts. That is much easier to read. Same with model E. Um, and so I'm just gonna go ahead and give... Eh, no, I'll let the model check. But I do like model B and model E here. I'll give them a little, a little, a little juicier 'cause I think they did a good job. I didn't read the others enough to really get a sense, but if I had to, like, high level pick the two favorite, just eyeballing, those would be the two. I'm gonna let computer use judge, um, how it managed using, um, an app. So I'm gonna let LLMs judge. Okay, and then here are the fun ones, and then we're gonna run the, uh, blind taste test. I'm gonna tell you guys what I like. Okay,

  8. 28:2529:48

    SVG illustration test

    1. CV

      so these new models can make, um, they can make SUV- SVGs. They can make illustrations. And so I [laughs] prompted it to make a document, a microphone, and a bug. And this I am happy to rank all day, every day. Um, and so I am just going to compare model A. Um, this bug is ridiculous. I'm gonna give it a four. I think the... Oh, I think the document is good, but the mic and the, the microphone looks like an ice cream cone, and the bug, I don't know what it looks like. This is, like, pretty consistently good across all of them, might be the best one, so I'm gonna give it, uh, model F a bonus. Let's do model B versus model C. Um, microphone terrible on this side. Oh. Oh, I'm gonna give them both threes. They're good in some ways and terrible on the other. Let's do model E versus model G. It's horri- Uh, this, why are they making microphones look like a cactus? Um, [laughs] okay, okay, I'll give it... This one I'll give a four. It's, I think, two out of the three are bad. I think this one's maybe a three, not perfect. And then let's look at model H really quickly. Oh, model H is really nice. I'm giving you a five. Look at the shadow. Microphone actually looks like a microphone, not a cactus. Document has lots of, um, character to it. And

  9. 29:4830:43

    AI video-editing results

    1. CV

      then the final one, it was terrible at all of them, cutting videos. So I gave them all access to, um... That's really unattractive, but gave them all access to a selfie video, and then had them cut short-form. And all of them did a terrible job, I just have to say. Um, these overlays are really bad. I don't think they cut particularly well. You aren't gonna be able to hear them, but I, I don't think it did enough cuts. Like, they just weren't good. I watched these earlier. I'm gonna let LLM as a judge do them, but I would say thumbs down on all of them, but I think this is a skills problem, not a model problem. Okay, so I have done my taste test. I'm gonna download my taste. And then what I do is I give it to Claude, and I say, "Claude, what do I think?

  10. 30:4331:20

    Predictions before the reveal

    1. CV

      Analyze." And we're gonna see if my prediction stands right. Again, here are my predictions for what I think. I think that I'm gonna like Claude and particularly Opus 5.5 for frontend. I think I'm gonna like Sol for writing. I don't know what's gonna do best for long-running agentic tasks or coding. And then, um, I think Anthropic's gonna crush on the SVGs, maybe. Maybe Sol? I don't know. Um, and then I am going to do, um... And then I think... But all of them are bad at

  11. 31:2034:17

    Barbie Bench: the 3D fashion-game test

    1. CV

      video. Now, while we're doing this, I do wanna show you one very important benchmark, which is Barbie Bench on Opus 5. If you all don't know, one of my fun benchmarks that I do is I ask it to make a 3D model of a Barbie. Uh, everybody else on YouTube, all the bros are making, um, video- or video games with, like, spaceships and all different stuff. Your girl wants a Barbie fashion video game. [laughs] And so one of the Opus 5.5, um, tasks I did was have it build Barbie, um, a 3D render of Barbie fashion designer. No model, no frontier model has crushed this. Um, and I just like to show the horror of 3D rendering for, uh, the female form. It's actually the worst. You can already see how terrible it is, but it's way better than it used to be. So I mean, uh, j- I'm speechless in some ways. Here's [laughs] what it does on the SVG side. I mean, truly something else. Um, let's put a bow on her hair. Let's, you know, give her purple eyes. Okay. So you can basically, like, uh, dress her, and then go into the fitting room. This is the [laughs] 3D, the 3D model it made. Um, let's give her a little miniskirt, and let's make it a little cuter. Now, um, you know, I will say it did some things better than other models. It, like, shaped its hair t- uh, her hands are... Oh, AGI is not ar- not, not arrived. Hands, pretty terrible. That, that is tragic. Feet, pretty bad. Um, you know, she's got a sassy little walk. Face, terrifying. Shoes, she's a big feet girl. She's a big feet queen. Um, [smacks lips] and then hair, again, like, questionable here. That ponytail is something else. The bob is... [laughs] You know, maybe it's like Anna Wintour. It's, it's something. So again, uh, the one bench that has not been crushed is Barbie Bench. Opus 5 has not done it. I don't know if I've run it [smacks lips] on, um, on Sol yet. I, I ran it on the prev- I ran it on Astra. It did worse than Opus, but I haven't updated it. So if you all wanna know how I think about models, I make them 3D render, um, [smacks lips] I make them 3D render Barbies. What are you gonna do? And this is why you join How I AI Live, so you can get this kind of frontier analysis of these new models.

  12. 34:1735:04

    Results: Astra, Sol, and Opus 5.5

    1. CV

      Drum roll, please. GPT-6 Astra and GPT-6 Sol win my heart. Uh, okay, A- Astra, um, and GPT-6 I love. And then it says, "Opus 5.5 wins my week." Um, it's the one I rated highly most often. So it's the best work per piece, but on average, I gave Opus 5.5 better results, high scores. I gave it fours and fives over, over the broadest range of work. Fable, didn't love. This is why I stopped using it. So, and then, um, GPT-5.6 Sol and the old version, um, and Claude Opus 5 ranking below the other ones. So I definitely do like the new models, um, for sure. And then,

  13. 35:0436:51

    Writing clarity and creative surprises

    1. CV

      okay, seven of my notes were about reading. So the Sol and the GPT models, um, I judged as clear and easy to read. And then Claude Opus still, um, but this is Opus 5 and 5.5. So I was testing Opus 5, so I still don't like Opus 5. Um, [smacks lips] still Opus 5 and 5.5 I said was overwrought and wrote a lot. So again, I had a suspicion that even though they made, um, Opus less verbose, they still, um, [smacks lips] needed to... They, they still need to work on that. It's still, it's still really chatty. Uh, so I thought that Sol, both 5.6 and 6, simple and straightforward, easy to read. Opus, really dense. Um, for PRDs, as predicted, I still think Sol is the best for PRDs. I did give, um, good tasks to Opus 5 and 5.5 for agentic stuff, and as expected, assistant voice. So my prediction was correct. Now, this is a real surprise. Astra and Sol did a lot better on character SVGs. And then, um, there's kind of, like, mixed results between all the other models. Um, [smacks lips] so kind of interesting. And then this is everything I scored a four. So I gave... Everything that I scored high, I scored this GPT-6... Oh, it did that nice little plant app. Astra and GPT Sol did these cute little illustrations. So I guess for SVGs, we're loving Astra and Sol for illustrations. Um, Claude Opus did my favorite B2B renewal, so it did some of these cleaner, um, [smacks lips] things. And but generally, I, [laughs] I like Astra. And this is where it gets

  14. 36:5137:24

    Why the LLM judge disagreed with me

    1. CV

      really funny. Um, the LLM is a judge, and I completely disagree. So we completely disagree. I like Astra. It likes Fable. And again, this is a GPT model judge. I... Next, we rank Opus 5.5 next. It ranks Sol a lot lower. So I find this just, like, really hilarious. So again, Astra wins my heart. 5.5 wins my week. Sol, up and down, but the fact that Sol's cheap, um, makes me very happy. And I think this is great.

  15. 37:2438:49

    What each model is actually best for

    1. CV

      So again, I don't know if this is just a personality thing, but I like Sol. I like Astra. Opus 5.5 I do not hate, so, um, not a complete failure by the Anthropic team by any means. It is back in the mix. It is definitely helping me, um, [smacks lips] do these tasks. It's helping me get day-to-day work on, and it is still my favorite agentic voice, and it does really good at long-running tasks. And then surprise hit, um, the character SVGs are best done by the OpenAI model. So you're moving into these new creative fields, um, new creative tasks, those are the models that I would say you use. And keep an eye out on the How I AI channel. As I said, smash that Subscribe button. [smacks lips] Um, we are dropping a full Opus 5.5 review, so you'll be able to see that. Thank you again for joining How I AI Live. I'm gonna go build some stuff with GPT-6 Sol. Bye, y'all. Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time

Episode duration: 38:51

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode LMT-bknLmNo

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.