Skip to content
How I AIHow I AI

Local AI models explained: How to run a fleet of Mac Studios and GPUs at home

Alex Finn is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I’ve come across: three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a fleet dashboard he built himself. He’s spent five months figuring out which local models belong on which machines, how to wire them to Claude Code loops, and how to get a software factory running without babysitting it. *What you’ll learn:* 1. How Alex chose between a Mac Studio (512 GB unified memory), DGX Spark, and RTX 5090, and what each is actually good for 2. Why Tailscale is worth installing even on a single machine, and how it lets one agent manage your entire hardware fleet 3. How the build loop and review loop in Claude Code work 4. How to allocate tasks by machine and model 5. Why unlimited local inference changes the use-case math in a way a $20 cloud subscription never can 6. What OpenClaw and Hermes are each best suited for, and why Alex runs five agents total with failover baked in *Brought to you by:* Runway—The creative AI platform for images, video, and more: https://runwayml.com/howIAI Jira Product Discovery—Prioritize with insights, build with confidence: https://atlassian.com/howiai *In this episode, we cover:* (00:00) Intro (02:58) Alex's hardware stack (03:48) What "ambient AI" means (04:15) Alex's red-pill moment with OpenClaw (07:04) Mac Studio vs. DGX Spark vs. RTX 5090 (13:24) How to set up local models with no technical knowledge (Tailscale + OpenClaw/Hermes) (17:16) Fleet control dashboard: assigning 24/7 tasks across machines (20:42) Local models as security scanners feeding Claude Code (22:25) How Alex allocates GLM 5.2, Qwen 3.6, and Ornith 1.0 by task (24:28) OpenClaw vs. Hermes: the honest comparison (26:55) The software factory: build loop, review loop, rocket emoji (31:55) Lightning round: favorite hardware, favorite model, prompting style (34:46) Where to find Alex *Tools referenced:* • Claude Code: https://claude.ai/code • OpenClaw: https://openclaw.ai/ • Hermes: https://hermes-agent.nousresearch.com/ • Tailscale: https://tailscale.com/ • Codex (OpenAI): https://openai.com/codex • GLM 5.2 (z.ai): https://huggingface.co/zai-org/GLM-5.2 • Qwen 3.6 (Alibaba): https://huggingface.co/Qwen/Qwen3.6-35B-A3B • Ornith 1.0: https://github.com/deepreinforce-ai/Ornith-1 • Gemma 4: https://huggingface.co/collections/google/gemma-4 • Playwright (browser testing): https://playwright.dev/ • Vercel (preview deploys): https://vercel.com/ *Other references:* • DGX Spark (Nvidia): https://www.nvidia.com/en-us/products/workstations/dgx-spark/ • Mac Studio (Apple): https://www.apple.com/mac-studio/ • How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex: https://www.lennysnewsletter.com/p/how-to-design-ai-agent-loops-schedules *Where to find Alex Finn:* LinkedIn: https://www.linkedin.com/in/alex-finn-1848684a YouTube: https://www.youtube.com/@AlexFinnOfficial X: https://x.com/AlexFinn *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

Claire VohostAlex Finnguest
Jul 13, 202635mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
July 13, 2026
Duration
35m
Channel
How I AI
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

Alex Finn is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I’ve come across: three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a fleet dashboard he built himself. He’s spent five months figuring out which local models belong on which machines, how to wire them to Claude Code loops, and how to get a software factory running without babysitting it. *What you’ll learn:*

  1. How Alex chose between a Mac Studio (512 GB unified memory), DGX Spark, and RTX 5090, and what each is actually good for
  2. Why Tailscale is worth installing even on a single machine, and how it lets one agent manage your entire hardware fleet
  3. How the build loop and review loop in Claude Code work
  4. How to allocate tasks by machine and model
  5. Why unlimited local inference changes the use-case math in a way a $20 cloud subscription never can
  6. What OpenClaw and Hermes are each best suited for, and why Alex runs five agents total with failover baked in

*Brought to you by:* Runway—The creative AI platform for images, video, and more: https://runwayml.com/howIAI Jira Product Discovery—Prioritize with insights, build with confidence: https://atlassian.com/howiai *In this episode, we cover:* (00:00) Intro (02:58) Alex's hardware stack (03:48) What "ambient AI" means (04:15) Alex's red-pill moment with OpenClaw (07:04) Mac Studio vs. DGX Spark vs. RTX 5090 (13:24) How to set up local models with no technical knowledge (Tailscale + OpenClaw/Hermes) (17:16) Fleet control dashboard: assigning 24/7 tasks across machines (20:42) Local models as security scanners feeding Claude Code (22:25) How Alex allocates GLM 5.2, Qwen 3.6, and Ornith 1.0 by task (24:28) OpenClaw vs. Hermes: the honest comparison (26:55) The software factory: build loop, review loop, rocket emoji (31:55) Lightning round: favorite hardware, favorite model, prompting style (34:46) Where to find Alex *Tools referenced:*

*Other references:*

*Where to find Alex Finn:* LinkedIn: https://www.linkedin.com/in/alex-finn-1848684a YouTube: https://www.youtube.com/@AlexFinnOfficial X: https://x.com/AlexFinn *Where to find Claire Vo:* ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email jordan@penname.co._

SPEAKERS

  • Claire Vo

    host

    Host of “How I AI with Claire Vo” and a product leader focused on applied AI tools and workflows.

  • Alex Finn

    guest

    Creator who builds and runs local AI/hardware setups (e.g., Mac Studios and GPUs) and shares AI automation and coding workflows.

EPISODE SUMMARY

In this episode of How I AI, featuring Claire Vo and Alex Finn, Local AI models explained: How to run a fleet of Mac Studios and GPUs at home explores running local AI fleets at home: hardware, models, workflows, automation Alex Finn explains why expensive home hardware can beat cloud subscriptions when you want “unlimited” 24/7 token-burning for continuous background tasks.

RELATED EPISODES

The AI content machine that turns ideas into posts that don't sound like slop | Alex Lieberman

The AI content machine that turns ideas into posts that don't sound like slop | Alex Lieberman

What is an AI harness? I build one live in less than 30 minutes

What is an AI harness? I build one live in less than 30 minutes

GPT-5.6 Sol: Better AND cheaper than Fable

GPT-5.6 Sol: Better AND cheaper than Fable

How I run autonomous coding agents from my phone with OpenAI Symphony + Linear

How I run autonomous coding agents from my phone with OpenAI Symphony + Linear

I benchmarked the NEW Sonnet 5. The results shocked me.

I benchmarked the NEW Sonnet 5. The results shocked me.

How to turn your company into AI builders | John Kim (Sendbird, co-founder and CEO)

How to turn your company into AI builders | John Kim (Sendbird, co-founder and CEO)

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.