Skip to content
How I AIHow I AI

An exclusive inside look at GPT-5

In this episode, I share my hands-on experience with OpenAI’s GPT-5, the company’s new frontier model. As one of the first users outside of OpenAI to test the model, I put GPT-5 head-to-head with GPT-4.1 across real-world product use cases—from writing PRDs to generating code to assisting with visual design work. This is my unfiltered look at what GPT-5 can (and can’t) do—and how it changes the game for builders. *What you’ll learn:* 1. How GPT-5 differs from previous models with its engineering-focused approach to problem-solving and tendency to prioritize technical details over business context 2. A comparative analysis of how GPT-5 and GPT-4.1 generate different types of product requirement documents and prototypes for the same prompt 3. Why GPT-5 excels at technical writing, functional requirements, and code generation while potentially skipping important business discovery questions 4. The model’s impressive spatial awareness capabilities when generating images for interior design and other visual tasks 5. Practical considerations for choosing the right model based on your specific use case and audience 6. How GPT-5’s extensive tool-calling behavior and bullet-point communication style reflect its engineering-oriented design *Brought to you by ChatPRD—an AI copilot for PMs and their teams:* https://www.chatprd.ai/howiai *25k giveaway:*  To celebrate 25,000 YouTube followers, we’re doing a giveaway. Win a free year of my favorite AI products, including v0, Replit, Lovable, Bolt, Cursor, and, of course, ChatPRD, by leaving a rating and review on your favorite podcast app and subscribing to the podcast on YouTube. To enter: https://www.howiaipod.com/giveaway *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo *In this episode, we cover:* (00:00) Introduction to GPT-5 (04:34) Testing GPT-5 in ChatPRD for document generation (07:10) Comparing GPT-5 and GPT-4.1 on business vs. technical orientation (11:22) Side-by-side comparison of PRDs generated by both models (15:23) Where GPT-5 excels: Technical considerations and documentation quality (17:35) Comparing prototypes generated from different model outputs (19:57) Testing homepage critique capabilities between models (23:14) OpenAI’s strengths in API design and developer support (25:37) GPT-5’s performance as a coding assistant (27:26) Examining GPT-5 in ChatGPT’s interface (28:50) Testing GPT-5’s front-end design capabilities (31:17) Personal use case: bathroom remodel planning (33:45) Comparing GPT-5 vs. GPT-4 for interior design visualization (38:10) Summary of key findings and recommendations *Tools referenced:* • OpenAI: https://openai.com/ • ChatGPT: https://chat.openai.com/ • Claude: https://claude.ai/ • Gemini: https://gemini.google.com/ • Cursor: https://cursor.sh/ • v0: https://v0.dev/ • Lovable: https://lovable.dev/ • Bolt: https://bolt.com/ • LaunchDarkly AI Configs: https://launchdarkly.com/docs/home/ai-configs *Other reference:* • Benjamin Moore paints: https://www.benjaminmoore.com/ _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire Vohost
Aug 7, 202540mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 2:31

    GPT-5’s “engineer-built” personality and who it’s for

    Claire shares her first impressions of GPT-5 as a highly technical model optimized for engineering workflows. She frames the episode around a practical question: where GPT-5 fits alongside other models as a specialized teammate rather than a universal upgrade.

    • GPT-5 feels “built by engineers for engineers” in both style and capabilities
    • Strengths: coding, refactoring, technical implementation thinking
    • Potential tradeoff: less business/stakeholder-friendly than GPT-4.1/o3
    • Episode roadmap: product/PRD workflows, coding, ChatGPT Canvas, personal tests
  2. 2:31 – 4:33

    Context: Claire’s model stack and why she tests models like a “team”

    She explains her day-to-day model usage across ChatPRD and coding tools, emphasizing deliberate model selection by task. This sets up her evaluation mindset: not “is it better,” but “where does it belong in the workflow.”

    • Uses multiple providers/models (OpenAI, Claude, Gemini) depending on task
    • ChatPRD has extensive prompt/model A/B testing and satisfaction targets
    • Evaluation criteria: output quality, user fit, cost later
    • Models as teammates with distinct personalities and strengths
  3. 4:33 – 7:03

    Switching ChatPRD to GPT-5: side-by-side document generation setup

    Claire moves directly into her primary benchmark: PRD/document generation inside ChatPRD. She describes using LaunchDarkly AI Configs to swap models and compare GPT-4.1 vs GPT-5 under identical system prompts and context.

    • ChatPRD’s core workflow: AI-assisted PRDs and feature brainstorming
    • LaunchDarkly AI Configs used for fast model switching in local/production
    • Comparison uses same prompts/context to isolate model behavior differences
    • Early GPT-5 quirk: defaults to markdown bullets and dev-speak
  4. 7:03 – 11:37

    Business lens vs execution lens: how GPT-4.1 and GPT-5 ask questions

    The first divergence appears in the questioning and framing. GPT-4.1 pushes on personas, goals, and metrics, while GPT-5 moves quickly toward concrete features and implementation detail—useful for building, riskier for product discovery.

    • GPT-4.1: business impact discovery (who/why, metrics, personas, goals)
    • GPT-5: solution-first behavior (what/how, user stories, numbers, build specs)
    • Implication for PMs: GPT-5 may skip “why” exploration and jump to execution
    • Model behavior reflects broader “coding tool wars” and engineering focus
  5. 11:37 – 13:39

    PRD artifacts and tone: GPT-5’s code-like fingerprints and verbosity

    Claire opens the generated PRDs and highlights stylistic artifacts that signal GPT-5’s training bias, like code-block comments and dense structure. She discusses the tradeoff between detail for engineers versus readability for stakeholders.

    • GPT-5 may insert code-adjacent artifacts even in prose docs
    • GPT-5 produces longer, denser, more detailed PRDs overall
    • More detail helps execution; too much can obscure the core message for alignment
    • Personas/use cases: GPT-5 becomes feature-centric vs GPT-4.1 goal-centric
  6. 13:39 – 16:11

    Where GPT-5 clearly wins: functional requirements, UX specificity, tech considerations

    The strongest difference shows up in the “how it works” sections. GPT-5 generates notably stronger functional requirements and technical considerations, making it compelling for specs, engineering handoff, and documentation-heavy teams.

    • Functional requirements: more complete, prioritized, and edge-case aware
    • UX descriptions: more concrete prose that can improve downstream prototyping
    • Technical considerations: detailed, engineer-native language and analysis
    • Possible workflow split: PMs may not need to author deep technical sections
  7. 16:11 – 19:43

    From PRDs to prototypes: how model detail changes v0 outputs

    Claire tests whether the different PRDs lead to better prototypes. She finds GPT-4.1 yields cleaner, more colorful visuals, while GPT-5’s verbosity produces richer component ideas—even if the aesthetic is less vibrant by default.

    • Prototype quality depends on the PRD as an input artifact
    • GPT-4.1 prototype: simpler, clearer, more colorful/pleasant visual style
    • GPT-5 prototype: more components/upsell mechanics to choose from (abundance)
    • Insight: GPT-5’s detail can improve ideation breadth for prototyping tools
  8. 19:43 – 23:14

    Homepage critique test: surprising differences in “mean vs nice” feedback

    Claire compares how each model critiques her homepage. GPT-4.1 is more blunt and harsh, while GPT-5 is more diplomatic—even when prompted to be more critical—highlighting differences in steerability and feedback style.

    • GPT-4.1 delivers sharper, more negative critique (“not up to standard”)
    • GPT-5 gives balanced feedback and a gentler tone by default
    • Re-prompting shows limits/differences in how easily each model is tuned
    • Takeaway: app builders must test promptability and tone control per model
  9. 23:14 – 24:45

    OpenAI ecosystem advantage: API design and developer tooling

    Before deeper coding demos, Claire credits OpenAI’s platform strengths beyond raw model quality. She emphasizes API ergonomics, primitives, and controls as decisive factors for building LLM-backed software.

    • OpenAI’s edge: API design, developer tools, ecosystem primitives
    • Model choice in products often hinges on integration simplicity and controls
    • Notes improvements around tool calling, reasoning, and parameters
    • Encourages developers to review the updated documentation
  10. 24:45 – 27:48

    GPT-5 as a coding assistant: fast, thoughtful, and extremely tool-hungry

    In Cursor and real shipping work, Claire finds GPT-5 strong at producing and refactoring production code. The standout behavioral trait is aggressive tool usage—often hitting tool-call limits—plus a strong preference for bullet-point communication.

    • Becomes her daily driver while shipping a major feature
    • Strengths: speed, code quality, refactoring, “thoughtful” engineering partner
    • Behavior: heavy tool calling (can hit Cursor tool-call limits)
    • Communication style: persistent bullet points; may feel inefficient at times
  11. 27:48 – 31:20

    GPT-5 inside ChatGPT: Canvas prototyping and front-end design taste

    Switching to ChatGPT, Claire tests GPT-5 Thinking with Canvas to generate a blog design matching her brand. She’s impressed by its polish and taste, though she flags recurring contrast/readability issues in generated UI styling.

    • GPT-5 vs GPT-5 Thinking options; uses Thinking for design/prototyping
    • Canvas workflow: reference image + prompt → code/UI output
    • Perceived improvement: more “classy” and polished design than past defaults
    • Issue: text/background contrast problems; likely CSS/model behavior to fix
  12. 31:20 – 34:53

    Personal benchmark: bathroom remodel planning and spatial reasoning

    Claire applies GPT-5 to a real consumer workflow: planning a bathroom remodel with layout constraints and visual outputs. She reports improved interpretation of spatial instructions compared with earlier experiences.

    • Uses AI for layout feasibility, visualization, and contractor communication
    • GPT-5 better at left/right/back-wall interpretation and room layout prompts
    • Requires some re-dos but produces more accurate spatial results
    • Positions spatial awareness as a differentiator across both code and images
  13. 34:53 – 37:57

    GPT-5 vs GPT-4o for interior design visualization: paint matching and mockups

    In a side-by-side test, GPT-5 generates paint recommendations with names and codes, then produces more faithful mockups to her tile/paint placement instructions. Compared to GPT-4 outputs, GPT-5’s renderings look more coherent and instruction-following.

    • Uploads tile photos; asks for Benjamin Moore paint matches
    • GPT-5 returns specific color names and paint codes; crisp, readable text
    • Mockups follow instructions more closely (half-wall tile, floor, wall materials)
    • GPT-4 comparison appears less sensical and less aligned to constraints
  14. 37:57 – 40:11

    Final takeaways: choose GPT-5 for engineering execution, use others for stakeholder clarity

    Claire concludes that GPT-5 is exceptional for engineering tasks: coding, technical writing, and detailed specs. For business storytelling and stakeholder alignment, GPT-4.1/4o/o3 may still be preferable, while GPT-5 meaningfully upgrades ChatGPT’s Canvas and image workflows.

    • Core identity: “for engineers, by engineers” (technical thinker/writer/coder)
    • PM impact: delivers more what/how than who/why—choose based on artifact purpose
    • Coding: few complaints beyond bullet-point style and tool-call intensity
    • ChatGPT upgrades: stronger Canvas/front-end design + improved image spatial awareness

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.