Skip to content
How I AIHow I AI

Local AI models explained: How to run a fleet of Mac Studios and GPUs at home

Alex Finn is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I’ve come across: three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a fleet dashboard he built himself. He’s spent five months figuring out which local models belong on which machines, how to wire them to Claude Code loops, and how to get a software factory running without babysitting it. *What you’ll learn:* 1. How Alex chose between a Mac Studio (512 GB unified memory), DGX Spark, and RTX 5090, and what each is actually good for 2. Why Tailscale is worth installing even on a single machine, and how it lets one agent manage your entire hardware fleet 3. How the build loop and review loop in Claude Code work 4. How to allocate tasks by machine and model 5. Why unlimited local inference changes the use-case math in a way a $20 cloud subscription never can 6. What OpenClaw and Hermes are each best suited for, and why Alex runs five agents total with failover baked in *Brought to you by:* Runway—The creative AI platform for images, video, and more: https://runwayml.com/howIAI Jira Product Discovery—Prioritize with insights, build with confidence: https://atlassian.com/howiai *In this episode, we cover:* (00:00) Intro (02:58) Alex's hardware stack (03:48) What "ambient AI" means (04:15) Alex's red-pill moment with OpenClaw (07:04) Mac Studio vs. DGX Spark vs. RTX 5090 (13:24) How to set up local models with no technical knowledge (Tailscale + OpenClaw/Hermes) (17:16) Fleet control dashboard: assigning 24/7 tasks across machines (20:42) Local models as security scanners feeding Claude Code (22:25) How Alex allocates GLM 5.2, Qwen 3.6, and Ornith 1.0 by task (24:28) OpenClaw vs. Hermes: the honest comparison (26:55) The software factory: build loop, review loop, rocket emoji (31:55) Lightning round: favorite hardware, favorite model, prompting style (34:46) Where to find Alex *Tools referenced:* • Claude Code: https://claude.ai/code • OpenClaw: https://openclaw.ai/ • Hermes: https://hermes-agent.nousresearch.com/ • Tailscale: https://tailscale.com/ • Codex (OpenAI): https://openai.com/codex • GLM 5.2 (z.ai): https://huggingface.co/zai-org/GLM-5.2 • Qwen 3.6 (Alibaba): https://huggingface.co/Qwen/Qwen3.6-35B-A3B • Ornith 1.0: https://github.com/deepreinforce-ai/Ornith-1 • Gemma 4: https://huggingface.co/collections/google/gemma-4 • Playwright (browser testing): https://playwright.dev/ • Vercel (preview deploys): https://vercel.com/ *Other references:* • DGX Spark (Nvidia): https://www.nvidia.com/en-us/products/workstations/dgx-spark/ • Mac Studio (Apple): https://www.apple.com/mac-studio/ • How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex: https://www.lennysnewsletter.com/p/how-to-design-ai-agent-loops-schedules *Where to find Alex Finn:* LinkedIn: https://www.linkedin.com/in/alex-finn-1848684a YouTube: https://www.youtube.com/@AlexFinnOfficial X: https://x.com/AlexFinn *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire VohostAlex Finnguest
Jul 13, 202635mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
July 13, 2026
Duration
35m
Channel
How I AI
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

Alex Finn is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I’ve come across: three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a fleet dashboard he built himself. He’s spent five months figuring out which local models belong on which machines, how to wire them to Claude Code loops, and how to get a software factory running without babysitting it. *What you’ll learn:*

  1. How Alex chose between a Mac Studio (512 GB unified memory), DGX Spark, and RTX 5090, and what each is actually good for
  2. Why Tailscale is worth installing even on a single machine, and how it lets one agent manage your entire hardware fleet
  3. How the build loop and review loop in Claude Code work
  4. How to allocate tasks by machine and model
  5. Why unlimited local inference changes the use-case math in a way a $20 cloud subscription never can
  6. What OpenClaw and Hermes are each best suited for, and why Alex runs five agents total with failover baked in

*Brought to you by:* Runway—The creative AI platform for images, video, and more: https://runwayml.com/howIAI Jira Product Discovery—Prioritize with insights, build with confidence: https://atlassian.com/howiai *In this episode, we cover:* (00:00) Intro (02:58) Alex's hardware stack (03:48) What "ambient AI" means (04:15) Alex's red-pill moment with OpenClaw (07:04) Mac Studio vs. DGX Spark vs. RTX 5090 (13:24) How to set up local models with no technical knowledge (Tailscale + OpenClaw/Hermes) (17:16) Fleet control dashboard: assigning 24/7 tasks across machines (20:42) Local models as security scanners feeding Claude Code (22:25) How Alex allocates GLM 5.2, Qwen 3.6, and Ornith 1.0 by task (24:28) OpenClaw vs. Hermes: the honest comparison (26:55) The software factory: build loop, review loop, rocket emoji (31:55) Lightning round: favorite hardware, favorite model, prompting style (34:46) Where to find Alex *Tools referenced:*

*Other references:*

*Where to find Alex Finn:* LinkedIn: https://www.linkedin.com/in/alex-finn-1848684a YouTube: https://www.youtube.com/@AlexFinnOfficial X: https://x.com/AlexFinn *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

SPEAKERS

  • Claire Vo

    host

    Host of “How I AI with Claire Vo” and a product leader focused on applied AI tools and workflows.

  • Alex Finn

    guest

    Creator who builds and runs local AI/hardware setups (e.g., Mac Studios and GPUs) and shares AI automation and coding workflows.

EPISODE SUMMARY

In this episode of How I AI, featuring Claire Vo and Alex Finn, Local AI models explained: How to run a fleet of Mac Studios and GPUs at home explores running local AI fleets at home: hardware, models, workflows, automation Alex Finn explains why expensive home hardware can beat cloud subscriptions when you want “unlimited” 24/7 token-burning for continuous background tasks.

RELATED EPISODES

Grok Bot + Grok 4.6  + Cursor Origin - is Claude Code dead?

Grok Bot + Grok 4.6 + Cursor Origin - is Claude Code dead?

Claude runs my entire business

Claude runs my entire business

I built an AI code review bot in 30 minutes - here’s how

I built an AI code review bot in 30 minutes - here’s how

How this OpenAI engineer uses Codex + ChatGPT Work to automate everything

How this OpenAI engineer uses Codex + ChatGPT Work to automate everything

How this “non-coder” used Cursor to add AI to retro hardware

How this “non-coder” used Cursor to add AI to retro hardware

I let Codex use my computer. Here’s what it can actually do.

I let Codex use my computer. Here’s what it can actually do.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.