Skip to content
The Twenty Minute VCThe Twenty Minute VC

Alex Wang: Why Data Not Compute is the Bottleneck to Foundation Model Performance | E1164

Alexandr Wang is the Founder and CEO @ Scale.ai, the company that allows you to make the best models with the best data. To date, Alex has raised $1.6BN for the company with a last reported valuation of $14BN earlier this year. Scale tripled their ARR in 2023 and is expected to hit $1.4BN in ARR by the end of 2024. Their investors include Accel, Index, Thrive, Founders Fund, Meta and Nvidia to name a few. ----------------------------------------------- Timestamps: (00:00) Intro (01:05) Diminishing Returns in AI Compute (09:08) Solving Reasoning to Overcome Limits (10:56) From Data Scarcity to Abundance (14:37) Challenges in Structuring Massive Enterprise Data (18:59) Fair Access to Proprietary Data for Models (22:02) Model Commoditization (26:51) Value Extraction Challenges in AI Commoditization (32:55) Navigating Data Regulatory Challenges for Innovation (36:53) A Military Asset in Global Conflict: China & Russia (42:49) The Future Landscape of Foundation Models (44:52) About Founder Brand & PR & Media (52:11) Hiring (01:00:41) Quick-Fire Round ----------------------------------------------- In Today’s Show with Alex Wang We Discuss: 1. Foundation Models: Diminishing Returns: What are the three core pillars that can meaningfully improve foundation models performance? Why is data the single largest bottleneck to the performance of models today? What data do we need to capture that we do not currently, that will have the biggest impact on model performance moving forward? Will we see the largest companies in the world revert back to on-prem with the increasing security challenges of migrating all customer data to foundation models? 2. AI: A Military Asset in Global Conflict: China + Russia Why does Alex believe that AI has the potential to be an even more powerful military asset than nuclear weapons? If this is the case, should we have open systems? Do we not have to have closed systems? Why does Alex believe that the CCP’s approach to industrial policy is better than anyone else’s? How does Alex evaluate the rise of Chinese EV car manufacturers in the last few years? Does Alex really believe that China is two years behind the US in the AI race? 3. “I Get Fairer Treatment in Congress than in the Press”: Why does Alex believe that the best PR is no PR? Why does Alex believe that he got fairer treatment in congress than he does in the media? Why does Alex believe that all founders should look to own their own distribution channels today? 4. Alex Wang: AMA: What are some of Alex’s biggest lessons from Patrick Collison on the impact that a hot company brand has on the ability for that company to hire the best? Does Alex think Trump is going to win? What would be the impact if he were to? Why does Alex believe that enterprise software will be changed forever in the next few years? What question is Alex never asked that he thinks he should be asked? ----------------------------------------------- Subscribe on Spotify: https://open.spotify.com/show/3j2KMcZTtgTNBKwtZBMHvl?si=85bc9196860e4466 Subscribe on Apple Podcasts: https://podcasts.apple.com/us/podcast/the-twenty-minute-vc-20vc-venture-capital-startup/id958230465 Follow Harry Stebbings on Twitter: https://twitter.com/HarryStebbings Follow Alexandr Wang on Twitter: https://twitter.com/alexandr_wang Follow 20VC on Instagram: https://www.instagram.com/20vchq Follow 20VC on TikTok: https://www.tiktok.com/@20vc_tok Visit our Website: https://www.20vc.com Subscribe to our Newsletter: https://www.thetwentyminutevc.com/contact ----------------------------------------------- #20vc #harrystebbings #alexandrwang #scaleai #openai #venturecapital #founder #chatgpt #ai #foundationmodels #china #military

Alexandr (Alex) WangguestHarry Stebbingshost
Jun 12, 20241h 6mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 2:58

    Compute spend is exploding, but model leaps have slowed since GPT-4

    Alex and Harry open on the apparent gap between surging GPU investment and the lack of a “jaw-droppingly better” base model than GPT-4. Alex frames the moment as a possible plateau while the industry waits for the next real breakthrough.

    • GPT-4 has been out since 2022 with no clear GPT-5-level jump yet
    • NVIDIA data center revenue and GPU spend have skyrocketed in the same period
    • The industry is investing heavily in compute but not seeing proportional capability gains
    • Sets up the question: is this a real asymptote or a temporary lull?
  2. 2:58 – 4:13

    The three pillars: compute, data, and algorithms—and the emerging data wall

    Alex explains that progress depends on compute, data, and algorithmic advances moving together. He argues today’s slowdown is largely due to hitting a “data wall” after consuming most easy-to-get internet-scale text.

    • AI progress historically requires compute + data + algorithms in tandem
    • Scaling compute alone is insufficient if data/algorithms lag
    • GPT-4-style training largely exhausts open internet data
    • The current plateau can be explained by data scarcity more than compute limits
  3. 4:13 – 7:19

    Why internet data can’t create true agents: missing real-world reasoning traces

    They dig into why pretraining on internet text tops out: the most valuable human reasoning that powers work isn’t written down online. Alex uses an enterprise example (fraud analysis) to show the missing step-by-step decision processes models need.

    • “Easy data” = crawlable/torrented/open web content
    • Models are great at “emulating the internet,” but that’s not enough
    • Complex workplace reasoning (e.g., fraud decisions) isn’t publicly documented
    • To do real tasks, models need data that captures decision-making processes
  4. 7:19 – 8:17

    Frontier data: reasoning chains, tool use, and agent behaviors as the next fuel

    Alex introduces the idea of “frontier data” as the essential input for moving beyond today’s capabilities. He describes it as complex reasoning chains, tool use, correction loops, and multi-step agentic workflows.

    • Need to move from a data-scarcity mindset to frontier-data abundance
    • Frontier data includes complex reasoning chains and multi-step agent behaviors
    • Tool use and iterative correction are core to agent performance
    • Capturing these traces is key to pushing models beyond GPT-4-era limits
  5. 8:17 – 11:18

    Where the data is: enterprise troves, plus mining vs forward data production

    Alex argues the largest untapped datasets live inside enterprises, dwarfing internet corpora (e.g., 150PB at a major bank vs ~<1PB internet training sets). He distinguishes between one-time mining of existing data and ongoing forward production of new data.

    • Enterprise data volumes are massive and largely off-internet for good reasons
    • Mining existing enterprise data can yield a one-time, meaningful benefit
    • Long-term progress requires continuous “forward” data production
    • If compute and data scaled in lockstep, model capability could jump dramatically
  6. 11:18 – 13:03

    Synthetic + human-in-the-loop production: ‘safety drivers’ for data quality

    To produce new frontier data, Alex outlines a hybrid approach: synthetic generation with expert humans guiding, correcting, and taking over when systems fail. He compares it to autonomous vehicle safety drivers who intervene during edge cases.

    • Frontier data production is a hybrid human + synthetic pipeline
    • Humans intervene when models get stuck or produce factual errors
    • Analogy: safety drivers disengage AVs during failure modes
    • Emerges new high-leverage roles: AI trainers/contributors
  7. 13:03 – 14:37

    Scaling ‘AI trainers’: why expert data contribution becomes a high-impact job

    Alex argues expert contribution to training data may be one of the most leveraged careers, since small improvements propagate to millions of downstream uses. He describes how domain experts can “transmit” their capabilities into widely used models.

    • Expert contributors can have society-wide impact through model improvements
    • Small quality gains compound across massive model usage
    • Roles like trainers/contributors formalize expert involvement
    • Particularly relevant for science, math, medicine, and other hard domains
  8. 14:37 – 19:29

    Structuring messy enterprise data, and why on-prem + open models matter

    They discuss the practical barriers of enterprise data governance: structure, cleanliness, and sensitivity. Alex predicts sophisticated enterprises will prefer on-prem or open-source models they can customize without leaking their data to external model vendors.

    • Enterprise data is rarely clean/structured for easy model ingestion
    • Most advanced firms will mine internal data; many won’t in 5 years
    • Enterprise leaders view proprietary data as a key future differentiator
    • On-prem/open-source models (LLaMA/Mistral) fit needs for strict data guarantees
  9. 19:29 – 25:17

    Data as the durable moat: competing on proprietary access as models commoditize

    Alex contends compute and algorithms are hard to defend long-term, but data can create durable competitive advantage. They discuss exclusive/strategic content deals and how future labs will compete on unique data strategies rather than just GPU counts.

    • Algorithms diffuse; compute can be purchased—data is the sustainable moat
    • Examples: publisher/library partnerships as early signs of data moats
    • Future competition shifts from “GPU bragging” to “data access bragging”
    • Labs must develop differentiated strategies to produce/secure unique datasets
  10. 25:17 – 33:26

    Where value accrues: models vs infrastructure vs services, and the future of pricing

    Alex discusses uncertainty about value capture in the AI stack, pointing to NVIDIA as “below the model” and apps/services above it. They explore the ‘End of Software’ thesis, rising customization, changing engineering work, and a shift from per-seat to consumption pricing as agents do more work.

    • Value capture is moving; models may commoditize faster than expected
    • Infrastructure (e.g., NVIDIA) and applications/services may capture outsized value
    • Enterprises will demand more customization as software creation costs drop
    • Per-seat pricing weakens when AI agents (not employees) do the work; consumption pricing rises
  11. 33:26 – 37:20

    Regulation and innovation: pro-data policy, pooling, and healthcare anonymization

    Alex warns restrictive data regimes can tie a country’s hands, similar to constraining chip production. He argues for sector-wide pooling where it benefits everyone (e.g., safety/fraud) and for workable frameworks (e.g., anonymization) to unlock medical data for better health outcomes.

    • EU approach is portrayed as restrictive; US/UK need pro-data lens
    • Countries should avoid handicapping future data production
    • Data pooling can advance whole sectors (aerospace safety, financial fraud/compliance)
    • Healthcare: revise/clarify HIPAA/PII constraints via anonymization to enable learning from patient data
  12. 37:20 – 41:31

    Geopolitics: China’s catch-up, industrial policy, and AI as a decisive military asset

    The conversation shifts to China’s rapid progress and the risk of centralized industrial policy “turning the crank” faster than the West. Alex frames advanced AI/AGI as a potentially dominant military asset and argues the West must prevent adversarial AGI advantage.

    • China has closed gaps quickly; cites strong Chinese models nearing top leaderboards
    • CCP industrial policy excels at scaling established industries (solar, EVs)
    • AI/AGI could surpass nukes as a strategic military asset via strategy, hacking, weapons production
    • Argues for Western focus on leadership to avoid catastrophic geopolitical imbalance
  13. 41:31 – 42:50

    Open vs closed AI: drawing the line for frontier systems

    Alex argues for a bifurcated approach: keep the most powerful frontier systems closed for security/geopolitical reasons, while allowing openness for less-advanced models that primarily generate economic value. The key challenge becomes defining where the line sits as capabilities grow.

    • Most cutting-edge systems should remain closed for national security reasons
    • Open models can still be fine below a certain capability threshold
    • Views current open models as not yet ‘military assets’
    • Critical future task: determine and monitor the openness threshold as models advance
  14. 42:50 – 44:48

    10-year foundation model landscape: ‘battle of giants’ and capital-driven consolidation

    Looking ahead, Alex predicts foundation model development will become so expensive that only nation-states and hyperscalers can sustain it. He anticipates consolidation and highlights long-term uncertainty in major lab–hyperscaler partnerships.

    • Training frontier models may rise to tens/hundreds of billions in cost
    • Only giants (nations/hyperscalers) can underwrite such programs
    • Smaller labs likely get acquired or deeply integrated via partnerships
    • Open questions: how Microsoft–OpenAI and Amazon–Anthropic partnerships resolve long-term
  15. 44:48 – 52:12

    Founder brand, PR, and media incentives: why ‘the best PR is no PR’

    Alex argues traditional media incentives (clicks) create a build-up/tear-down cycle that can harm companies. He advocates direct channels—podcasts and founder-led communications—where messages aren’t distorted, emphasizing founder brand as a key distribution asset.

    • Traditional press optimizes for clicks; tends to amplify extremes
    • Direct channels let founders explain work without message distortion
    • Founder brand often outperforms company brand in audience attention
    • Alex cites unfair media treatment (including around defense work) versus more balanced congressional engagement
  16. 52:12 – 1:00:41

    Hiring and leadership: talent density, founder approval, and the hypergrowth mistake

    Alex explains Scale’s emphasis on hiring people who deeply care, maintaining an elite bar as the company grows. He shares that he still approves every hire and reflects on a major mistake: equating company hypergrowth with headcount hypergrowth, which diluted excellence.

    • ‘People who give a shit’ sweat details and push through roadblocks
    • Alex claims he approves every hire; overrides team recommendation ~25–30%
    • Hyper-hiring (2020–2022) made it impossible to maintain the quality bar
    • Lesson: divorce team growth from company growth; prioritize talent density and excellence
  17. 1:00:41 – 1:06:04

    Quick-fire reflections: misconceptions about AGI, election take, and Scale’s long-term aim

    In rapid Q&A, Alex reiterates his belief that data—not just compute—is the major missing ingredient to AGI. He shares a cautious view on hype cycles (learning from autonomous vehicles), calls the US election a swing-state tossup, and positions Scale as a long-term “data foundry” for AI.

    • Changed mind: prioritize quality/excellence over headcount hypergrowth
    • Biggest AI misconception: compute alone leads to AGI—data is also required
    • Cautionary parallel to AV hype: promises can outpace technical reality and cause hangovers
    • Scale in 10 years: remain the ‘data pillar’/data foundry; interest in going public

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.