Skip to content
How I AIHow I AI

Gemini 3 vs. Claude Opus 4.5 vs. GPT-5.1 Codex: Which AI model is the best designer?

I put three cutting-edge AI models to the test in a head-to-head design competition. Using the exact same prompt, I challenged Google’s Gemini 3, Anthropic’s Opus 4.5, and OpenAI’s Codex 5.1 to redesign my blog page, evaluating them on visual design quality, user experience improvements, and SEO optimization capabilities. One model produced a beautiful, polished, production-ready redesign. One was fine. And one completely whiffed. If you’re trying to figure out where each model fits in your workflow—design, planning, back-end, or something else—this episode will save you a lot of trial and error. *What you’ll learn:* 1. How each AI model approaches the same design challenge differently 2. Why planning capabilities dramatically impact design quality 3. The specific visual and functional improvements each model made 4. Which model excels at front-end design versus back-end functionality 5. How to strategically choose the right AI model for different parts of your workflow 6. The importance of model-switching based on specific use cases *Blog design:* https://www.chatprd.ai/blog *Brought to you by:* Lovable—Build apps by simply chatting with AI: https://lovable.dev/ *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo *In this episode, we cover:* (00:00) Introduction to the AI design challenge (01:25) The question: Which model is the better designer? (03:08) The prompt used for all three models (04:10) Gemini 3 Pro’s approach and results (06:00) Opus 4.5’s approach and results (10:54) Codex 5.1’s approach and disappointing results (14:51) Comparing the three designs side by side (16:03) Analyzing the change logs and SEO improvements from each model (22:43) Final verdict (23:00) Conclusion and next steps *Tools referenced:* • Gemini 3 Pro: https://deepmind.google/models/gemini/pro/ • Anthropic Opus 4.5: https://www.anthropic.com/news/claude-opus-4-5 • OpenAI Codex 5.1: https://platform.openai.com/docs/models/gpt-5.1-codex • Cursor: https://cursor.com/ Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Claire Vohost
Dec 3, 202525mWatch on YouTube ↗

CHAPTERS

  1. 0:04 – 0:36

    AI models face off: redesigning a real blog page in one shot

    Claire sets up a mini design showdown: Gemini 3 Pro, Claude Opus 4.5, and GPT-5.1 Codex will each redesign an existing (and intentionally imperfect) blog page. The goal is to see which model behaves like a “trusted design engineer” when improving an already-live site, not just generating something pretty from scratch.

    • Framing the core question: which new coding model is actually the best designer?
    • Using a real target: the ChatPRD blog page Claire considers poorly designed
    • Constraint: each model gets a one-shot attempt
    • Focus on redesigning existing code—not greenfield page generation
  2. 0:36 – 1:06

    Sponsor: Lovable for building AI-generated apps and websites

    A brief sponsor segment explains Lovable’s pitch: build functional apps and websites by chatting with AI, then customize and deploy. Claire contrasts Lovable with static-page no-code tools by emphasizing real app functionality and speed to ship.

    • Lovable turns chat prompts into deployable apps/websites
    • Supports customization, automations, and live domains
    • Positioned for marketers, PMs, and founders
    • Differentiation: full functionality vs. static no-code pages
  3. 1:06 – 3:07

    Why this test: benchmarking ‘design’ on an existing site is harder

    Claire explains why side-by-side comparisons matter: many models can one-shot attractive UI when prompted well, but real work is improving messy existing pages. She frames this as a practical question of where each model fits in a workflow.

    • Recent wave of “coding models” claimed to be great at design
    • Social feeds show polished AI-generated landing pages and components
    • Redesigning an existing site tests taste + constraints + implementation
    • The objective: identify the most reliable design-focused model
  4. 3:07 – 4:08

    Setup in Cursor: same codebase, same prompt, three models

    Claire describes the experiment setup inside Cursor: identical input code and identical prompt across models. The prompt targets visual appeal, UX improvements, and SEO/navigation best practices.

    • Tooling: Cursor used for model-by-model comparison
    • Same directory/code provided for each run
    • Prompt goals: visual appeal + UX + SEO + navigation
    • Models compared: Gemini 3 Pro, Opus 4.5, GPT-5.1 Codex
  5. 4:08 – 6:10

    Gemini 3 Pro redesign: solid layout upgrades but imperfect polish

    Gemini 3 Pro produces a noticeably improved blog page with a featured hero post and a card grid below. Claire likes several enhancements (tags, dates, hover zoom), but notes spacing and navigation tightness issues and feels it’s not the best overall despite Gemini’s reputation.

    • Adds a featured “hero” post at the top plus card-based listing
    • Improves post metadata visibility (tags, dates)
    • Hover effects: image zoom and shadowing
    • Weak spots: tight spacing near navigation; pagination/featured-image handling not addressed
  6. 6:10 – 7:42

    Opus 4.5 planning-first workflow: to-do list and step-by-step execution

    Running the same prompt with Opus 4.5, Claire observes a different behavior: the model creates a structured to-do list and works through changes methodically. She highlights this planning/tool-use pattern as a key reason the final output is more cohesive.

    • Opus triggers an internal to-do list via tool calls in Cursor
    • To-dos include layout redesign, post display enhancements, and SEO specifics
    • More precise implementation plan than Gemini’s direct generation
    • Claire attributes stronger outcomes to planning + incremental execution
  7. 7:42 – 9:43

    Opus 4.5 results: best-looking page with thoughtful micro-interactions

    Opus 4.5 delivers Claire’s favorite redesign visually, pulling in existing brand assets and adding background imagery. It refines hover states with a subtle arrow call-to-action and enhances post cards with richer metadata and better empty-state handling.

    • Uses existing site assets (background imagery/design elements) for brand consistency
    • Maintains strong information hierarchy: featured article + grid layout
    • Polished interaction design: hover zoom plus arrow CTA micro-detail
    • Adds reading time and improved metadata presentation
  8. 9:43 – 10:43

    Opus 4.5 robustness: graceful placeholders when posts lack images

    Claire highlights a practical UX improvement: Opus detects missing featured images and inserts attractive placeholder cards with an icon, preserving layout consistency. This attention to edge cases makes the page feel more intentionally designed.

    • Detects missing images and renders a designed placeholder state
    • Prevents awkward collapsed cards seen in other outputs
    • Improves perceived quality and consistency of the grid
    • Reinforces Opus’s strength in detail-oriented front-end UX
  9. 10:43 – 14:46

    GPT-5.1 Codex attempt: generic planning, ‘AI purple’ aesthetics, broken UX

    Codex 5.1 also produces a plan, but it’s higher-level and less design-specific than Opus. The resulting page leans on an overused purple-blue gradient, has logo/contrast issues, and contains confusing or non-functional navigation and featured sections.

    • To-do list exists but is broad: investigate layout → redesign → SEO
    • Visual issues: default purple gradient and weak branding fit
    • Featured section lacks clear CTA and linking/interaction clarity
    • Category/jump links behave oddly and don’t reflect the real library content
  10. 14:46 – 15:46

    Model roles in a workflow: why ‘model switching’ matters

    Claire zooms out from the results to emphasize model specialization: some excel at design, others at planning, writing, image generation, or backend engineering. She argues that testing models on repeated, concrete tasks builds the skill of assigning the right model to the right job.

    • Codex may be strong for backend work but weak for front-end design
    • Gemini is serviceable for UI but benefits from better planning
    • Opus stands out for design execution quality
    • Key practice: evaluate models by use case and switch models across workflow stages
  11. 15:46 – 16:17

    Change-log comparison workflow: ask each model to summarize its edits

    Claire shares a practical tactic: request a categorized summary of changes (design, UX, SEO) after the agent modifies code. This makes it easier to compare outputs when you aren’t watching every step and helps validate improvements beyond surface visuals.

    • Workflow tip: have agents summarize what they changed
    • Compare changes across models in categories (design/SEO/navigation)
    • Useful when running multiple agents asynchronously
    • Helps verify improvements on listing pages and individual posts
  12. 16:17 – 22:22

    SEO and page-level details: Gemini vs. Opus vs. Codex summaries

    Reviewing the summaries, Claire notes Gemini made meaningful SEO and blog-post-page improvements (schema, breadcrumbs, semantic HTML, related articles). Opus’s redesign changes are extensive and polished, though some SEO elements (like explicit JSON-LD mention) are less clear; Codex provides the shortest summary despite adding schema.

    • Gemini: JSON-LD/schema, breadcrumbs, semantic HTML, and related-article enhancements
    • Opus: broad design system improvements (badges, pills, spacing, empty states) and enhanced post metadata
    • Opus also upgrades the newsletter CTA component (though color choice still skews purple)
    • Codex: minimal summary; includes schema.org embedding but weaker UX/detail overall
  13. 22:22 – 24:24

    Final verdict and takeaway: Opus 4.5 wins for design + usability

    Claire concludes Opus 4.5 is the clear winner for front-end design quality and overall usability, crediting strong planning and detailed implementation. She emphasizes how fast this workflow is—generating three viable redesign directions in under 20 minutes—and plans to ship the best version.

    • Winner: Anthropic Opus 4.5 for design and practical UX improvements
    • Reasoning: stronger planning + better detail execution
    • Gemini: decent and helpful, especially with SEO/post-page work
    • Codex: not the “designer,” better suited elsewhere in the stack
  14. 24:24 – 25:27

    Wrap-up: shipping the redesign and where to follow the show

    Claire celebrates the speed and leverage of AI-assisted redesigns, noting she’ll ship the chosen design and share it in the show notes. She closes with standard calls to like/subscribe, comment, and find the podcast on major platforms.

    • She plans to ship the redesign quickly and share links/show notes
    • Reinforces the episode’s purpose: practical model comparison on a real task
    • Call to action: like/subscribe/comment
    • Where to listen and learn more: howiaipod.com and podcast platforms

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.