Skip to content
The Twenty Minute VCThe Twenty Minute VC

How to Build Your Own Data Center & Why Every Startup Should Do It

Cliff Weitzman is the co-founder and CEO of Speechify, the world's leading AI voice and text-to-speech platform, used by more than 60 million people globally. Diagnosed with dyslexia as a child, Cliff first built Speechify at Brown University to help him consume written material through audio. ----------------------------------------------- Timestamps: 00:00 - Intro 00:55 - Why Speechify Is Spending Tens of Millions on Its Own GPUs 07:12 - Why Buying GPUs Can Be Better Than Renting Them 12:28 - Are NVIDIA’s Circular Economy Fears Overblown? 13:49 - What No One Understands About Buying AI Chips 18:26 - Why Owning Compute Can Become a Competitive Advantage 20:55 - How Big Can the AI Data Market Become? 22:59 - How ElevenLabs Leapfrogged Speechify 25:23 - Is Speechify Making a Mistake Going Into B2B? 30:17 - Why Every Winning AI Company Becomes a Compound Startup 31:10 - Is This the Hardest Time Ever for Startups to Hire Great Talent? 37:26 - How AI Is Completely Changing Software Engineering 39:18 - How Should Companies Think About AI Token Spend? 45:29 - How Your Hiring Process Needs to Change in the AI Era 47:03 - Is Voice AI Becoming Completely Commoditized? 50:18 - Is Customer Support the Wrong AI Market to Bet On? 53:34 - Sierra vs ElevenLabs: Who Becomes Bigger? 55:18 - Quick-Fire Round ----------------------------------------------- Subscribe on Spotify: https://open.spotify.com/show/3j2KMcZTtgTNBKwtZBMHvl?si=85bc9196860e4466 Subscribe on Apple Podcasts: https://podcasts.apple.com/us/podcast/the-twenty-minute-vc-20vc-venture-capital-startup/id958230465 Follow Harry Stebbings on X: https://twitter.com/HarryStebbings Follow Cliff Weitzman on X: https://twitter.com/cliffweitzman Follow 20VC on Instagram: https://www.instagram.com/20vchq Follow 20VC on TikTok: https://www.tiktok.com/@20vc_tok Visit our Website: https://www.20vc.com Subscribe to our Newsletter: https://www.thetwentyminutevc.com/contact ----------------------------------------------- Legal Disclaimer: The content of this podcast is for informational and entertainment purposes only and does not constitute financial or investment advice. Any discussion of stocks, public markets, or investment strategies reflects the personal opinions of the speakers and should not be relied upon when making investment decisions. Figures, valuations, and financial data referenced may be estimates or subject to error. Always consult a qualified financial adviser before making any investment decision. The views expressed are those of the individual speakers and do not represent the views of 20VC or its affiliates. ----------------------------------------------- #20vc #harrystebbings #cliffweitzman #ceo #speechify #ai #voiceai #startup #elevenlabs #sierra

Cliff WeitzmanguestHarry Stebbingshost
Sep 5, 20261h 4mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 2:05

    Speechify’s GPU bet: why spend tens of millions to own compute

    Cliff explains why Speechify began buying NVIDIA GPU racks instead of relying purely on rented cloud capacity. He frames it as removing friction for engineers and accelerating model iteration speed, especially for training. The segment sets up the core thesis: compute access is becoming strategically decisive for AI companies.

    • Bought first major GPU rack in 2022 primarily to accelerate training
    • Owning GPUs changes engineer behavior (less “parsimonious” experimentation)
    • Michael Jordan/“basketball hoop at home” analogy for unlimited practice
    • Speechify’s model performance/price positioning vs frontier labs and competitors
    • Early motivation: speed of iteration and model quality gains
  2. 2:05 – 4:28

    Rent vs buy economics: when GPU ownership beats the cloud

    They dig into the math of buying an H100 versus renting on AWS/GCP/Azure, arguing renting can cost 1.5× of purchase price per year. Cliff also introduces the practical training constraint: large-scale jobs require tightly co-located GPU clusters and memory access that’s hard or expensive to assemble ad hoc in the cloud.

    • Illustrative math: ~$30k to buy vs ~$35–50k/year to rent an H100
    • GPUs can remain useful beyond warranty; old cards still valuable
    • Large training runs require co-located GPUs + high-bandwidth memory access
    • Owning enables cheaper per-token inference when running open-source models
    • Cloud still used tactically (dedicated + spot) alongside owned baseline
  3. 4:28 – 7:07

    Depreciation, chip cycles, and why older GPUs still matter

    Harry challenges the lock-in and depreciation risk of buying hardware amid rapid chip cycles. Cliff argues GPUs don’t “wear out” like consumer devices, and older generations stay useful—especially for inference and lower-priority experiments. He also highlights load balancing via cloud bursts and the option to rent out spare capacity.

    • Training needs newest speed; inference can run well on older GPUs
    • Older cards (e.g., K80s) can still handle specific inference workloads
    • Experiment portfolio: not every job needs bleeding-edge hardware
    • Seasonal demand handled by cloud bursts on top of owned baseline
    • Excess capacity could be rented out (even if rarely expected)
  4. 7:07 – 13:44

    Operating a ‘startup-scale data center’: racks, colocation, and energy constraints

    Cliff describes what “buying GPUs” really entails operationally: ordering racks (Rubens/Blackwells), shipping logistics, and colocating them in third-party data centers. He notes energy as the biggest constraint and outlines how data centers provide power, networking, and on-site hands to install and service equipment.

    • Rubens described as 72 cards per rack; multiple racks planned
    • Mix of latest-gen (B300) plus early access strategy
    • Speechify rents colocation space; data center handles power/networking/install
    • Energy is the binding constraint more than rack space
    • ElevenLabs also builds clusters (early DIY-to-scale narrative)
  5. 13:44 – 13:49

    What founders miss about chip procurement: vendors, delays, insurance, and liquid cooling

    This chapter details the unglamorous realities of sourcing AI chips: buying through vendors like Dell, negotiating delivery priority, and dealing with delays. Cliff explains why teams pay large premiums to receive hardware early and why insurance and cooling infrastructure (especially liquid cooling for newer GPUs) become first-class concerns.

    • NVIDIA often sells via intermediaries; Dell as a major rack supplier
    • Global sourcing and delivery slippage can break plans and budgets
    • Paying big premiums to ‘skip the queue’ for earlier delivery
    • Late delivery is costly because colocation rent continues regardless
    • Liquid cooling adoption: sidecars, approvals, and data center retrofits
  6. 13:49 – 17:44

    NVIDIA ‘circular economy’ debate and the emerging secondary market floor

    Harry asks whether fears about a circular GPU economy harming NVIDIA are real. Cliff argues GPUs have strong intrinsic value (measured by FLOPs) and cites NVIDIA’s financing/underwriting efforts to create liquidity and a value floor in secondary markets. He likens it to asset-backed financing innovations like SolarCity.

    • GPUs framed as intrinsically valuable assets vs speculative instruments
    • Value anchored to compute throughput (FLOPs) and usability anywhere
    • NVIDIA + major financial firms underwriting GPU resale value (floor)
    • Better lending terms could follow with GPUs as collateral
    • Comparison to SolarCity-style amortization and asset-backed financing
  7. 17:44 – 19:39

    Compute, data, architecture: why ownership becomes competitive advantage

    Cliff ties compute strategy to competitive advantage: bigger training, faster iteration, and more capable teams. He emphasizes that modern AI progress requires the trifecta of compute, data, and strong architecture/product execution—especially as agents and RL-style iteration loops accelerate development cycles.

    • Owning compute expands feasible training scale and team leverage
    • Engineers run many agents/experiments concurrently; compute is the bottleneck
    • Compute also needed for data cleaning and synthetic dataset generation
    • Core resources: data + compute + architecture (plus user feedback loops)
    • Speed and iteration cadence become the primary moat in applied AI
  8. 19:39 – 22:25

    Buying data and the size of the data marketplace (and why it’s not ARR)

    The conversation shifts to data as an input: when to buy it, who sells it, and how to evaluate the market. Cliff points out that data sales are often one-time transactions rather than recurring revenue, making the business operationally intense and legally sensitive. Still, he sees major opportunity if providers can prove measurable model improvement.

    • Speechify buys some datasets, but mostly small/supplemental
    • Data vendors (e.g., Surge/Micro1/Mercor) accelerate time-to-revenue
    • Data deals are one-off, not ARR—riskier business model
    • Legal provenance/indemnification is a key part of the value proposition
    • Success metric: customer must see real model quality uplift from the data
  9. 22:25 – 25:18

    ElevenLabs leapfrogged Speechify: Cliff’s biggest strategic mistake (B2B wedge)

    Cliff admits Speechify’s biggest mistake was underestimating the power of starting with a B2B API wedge. He explains how he over-weighted the risk of commoditization and under-weighted the compounding effect of continuous product expansion from a simple initial entry point. The segment becomes a broader lesson on AI labs: ship a wedge, then relentlessly broaden capabilities.

    • Cliff believed TTS APIs would commoditize, so avoided B2B early
    • Mistake: forgot the ‘AI lab’ model—continuous innovation is the product
    • API as wedge → voice features → duplex/turn-taking → agents and outcomes
    • ElevenLabs’ launch excellence and product layering cited as key
    • Principle: get usage first (even free), then expand and monetize broader stack
  10. 25:18 – 30:12

    Should Speechify go B2B now? Competing with ElevenLabs and Sierra

    Harry argues B2B is brutally competitive and Speechify risks being ‘third place.’ Cliff counters that Speechify dominates B2C TTS distribution and can fund/learn B2B while the market remains large and non-monopolistic. They debate power laws, timing, and the necessity of being “in the race” despite incumbents.

    • Speechify claims ~98% of App Store TTS installs; massive usage scale
    • Harry’s concern: ElevenLabs momentum + Sierra/Brett Taylor execution
    • Cliff’s counter: markets shift; incumbents can fumble; oligopoly likely
    • Strategic stance: cannot opt out—must learn B2B and compound products
    • Value accrues to top players, but multiple huge outcomes still possible
  11. 30:12 – 31:14

    The ‘compound startup’ idea: why winners layer products and wedges

    Cliff describes how, at scale, startups can’t remain single-product—they must compound by layering new offerings and expanding along adjacent value pools. He connects this to market observations in speech-to-text productivity tools (WhisperFlow/Willow) and to Speechify’s own decision to expand into new surfaces, including Siri-like experiences.

    • At a certain point, it’s too expensive not to compound products
    • Wedge strategy: ship something simple, then expand into adjacent needs
    • WhisperFlow example: quality tradeoffs when switching models for cost
    • Lesson: even ‘commoditized’ categories can remain open if incumbents stall (Apple/Siri)
    • Speechify launching new products across STT, voice assistant, and B2B API/agents
  12. 31:14 – 37:21

    Hiring in the AI era: is it harder than ever, and what to hire for now

    They debate whether OpenAI/Anthropic make hiring impossible by concentrating talent and offering huge comp. Cliff argues it’s hardest for growth companies chasing elite leaders, while smaller startups can still win by hiring for raw aptitude and slope. He outlines a shift from handcrafted code obsession to technical intelligence and rapid learning ability with AI tools.

    • Anthropic/OpenAI attract CTO-tier talent with massive comp packages
    • Seed vs growth distinction: easier to hire for potential at early stage
    • Shift hiring signal: raw technical intelligence > polished coding style
    • Targets: math/physics talent, Olympiad/LeetCode/Kaggle profiles
    • Humans still balance risk vs certainty; liquidity and IPOs affect incentives
  13. 37:21 – 50:12

    AI-centric engineering in practice: agent tools, shipping discipline, and token spend

    Cliff explains how Speechify drives adoption of agentic development: showing internal demos, standardizing on tools (Claude Code/Cursor), and measuring performance by production output. They cover token budgeting philosophy, why leaderboards can be misleading, and how to structure goal-seeking loops with measurable targets and QA discipline.

    • Tool stack: Claude Code and Cursor dominate; some Codex usage
    • Change management: demo best workflows to teams to force adoption
    • Metric of success: shipped to production and used by real users
    • Token spend: budget within bounds; waste is a performance issue
    • Best practice: define target + measurement, iterate in a loop, QA hard
  14. 50:12 – 55:13

    Market bets: commoditization, customer support skepticism, and Sierra vs ElevenLabs

    Harry challenges the attractiveness of customer support as a crowded AI market; Cliff agrees it’s noisy but argues you still need frontline deployments to discover higher-value adjacent opportunities. Cliff positions Speechify’s B2B core as the API/model layer (quality, speed, cost), while agents and services help uncover the next wedge. They close with a pragmatic view: both Sierra and ElevenLabs can become enormous due to the size of the agents market.

    • Customer support crowded: many well-funded entrants + incumbents + in-house builds
    • Cliff: Speechify’s primary B2B offering is API/model performance, not support agents
    • Agents matter as a learning surface to uncover ‘next hill’ opportunities
    • ElevenLabs’ government adoption cited as evidence of compounding GTM
    • Sierra vs ElevenLabs: different games; both likely massive as agents expand
  15. 55:13 – 1:04:34

    Quick-fire + the bigger vision: voice-first interfaces and AI for biology

    In quick-fire, Cliff predicts voice becomes the dominant human-computer interface as latency and model quality improve. The conversation then shifts to Cliff’s personal passion: applying GPU-powered AI to biology and pharmacology, including sequencing, proteomics, and rare disease research. He ends by tying it back to his origin story—assistive technology and compute enabling breakthroughs in health and quality of life.

    • Prediction: voice becomes primary interface; screens used less
    • Current voice assistants criticized for latency and degraded model quality
    • Meta’s approach praised; expectation of constant conversational computing
    • Biology/pharma use case: multi-omics + longitudinal data analyzed on GPU clusters
    • Rare/orphan disease research: sequencing groups, pattern finding, molecule design via AlphaFold/CRISPR workflows

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.