The Twenty Minute VCAlex Wang: Why Data Not Compute is the Bottleneck to Foundation Model Performance | E1164
CHAPTERS
- 0:00 – 2:58
Compute spend is exploding, but model leaps have slowed since GPT-4
Alex and Harry open on the apparent gap between surging GPU investment and the lack of a “jaw-droppingly better” base model than GPT-4. Alex frames the moment as a possible plateau while the industry waits for the next real breakthrough.
- •GPT-4 has been out since 2022 with no clear GPT-5-level jump yet
- •NVIDIA data center revenue and GPU spend have skyrocketed in the same period
- •The industry is investing heavily in compute but not seeing proportional capability gains
- •Sets up the question: is this a real asymptote or a temporary lull?
- 2:58 – 4:13
The three pillars: compute, data, and algorithms—and the emerging data wall
Alex explains that progress depends on compute, data, and algorithmic advances moving together. He argues today’s slowdown is largely due to hitting a “data wall” after consuming most easy-to-get internet-scale text.
- •AI progress historically requires compute + data + algorithms in tandem
- •Scaling compute alone is insufficient if data/algorithms lag
- •GPT-4-style training largely exhausts open internet data
- •The current plateau can be explained by data scarcity more than compute limits
- 4:13 – 7:19
Why internet data can’t create true agents: missing real-world reasoning traces
They dig into why pretraining on internet text tops out: the most valuable human reasoning that powers work isn’t written down online. Alex uses an enterprise example (fraud analysis) to show the missing step-by-step decision processes models need.
- •“Easy data” = crawlable/torrented/open web content
- •Models are great at “emulating the internet,” but that’s not enough
- •Complex workplace reasoning (e.g., fraud decisions) isn’t publicly documented
- •To do real tasks, models need data that captures decision-making processes
- 7:19 – 8:17
Frontier data: reasoning chains, tool use, and agent behaviors as the next fuel
Alex introduces the idea of “frontier data” as the essential input for moving beyond today’s capabilities. He describes it as complex reasoning chains, tool use, correction loops, and multi-step agentic workflows.
- •Need to move from a data-scarcity mindset to frontier-data abundance
- •Frontier data includes complex reasoning chains and multi-step agent behaviors
- •Tool use and iterative correction are core to agent performance
- •Capturing these traces is key to pushing models beyond GPT-4-era limits
- 8:17 – 11:18
Where the data is: enterprise troves, plus mining vs forward data production
Alex argues the largest untapped datasets live inside enterprises, dwarfing internet corpora (e.g., 150PB at a major bank vs ~<1PB internet training sets). He distinguishes between one-time mining of existing data and ongoing forward production of new data.
- •Enterprise data volumes are massive and largely off-internet for good reasons
- •Mining existing enterprise data can yield a one-time, meaningful benefit
- •Long-term progress requires continuous “forward” data production
- •If compute and data scaled in lockstep, model capability could jump dramatically
- 11:18 – 13:03
Synthetic + human-in-the-loop production: ‘safety drivers’ for data quality
To produce new frontier data, Alex outlines a hybrid approach: synthetic generation with expert humans guiding, correcting, and taking over when systems fail. He compares it to autonomous vehicle safety drivers who intervene during edge cases.
- •Frontier data production is a hybrid human + synthetic pipeline
- •Humans intervene when models get stuck or produce factual errors
- •Analogy: safety drivers disengage AVs during failure modes
- •Emerges new high-leverage roles: AI trainers/contributors
- 13:03 – 14:37
Scaling ‘AI trainers’: why expert data contribution becomes a high-impact job
Alex argues expert contribution to training data may be one of the most leveraged careers, since small improvements propagate to millions of downstream uses. He describes how domain experts can “transmit” their capabilities into widely used models.
- •Expert contributors can have society-wide impact through model improvements
- •Small quality gains compound across massive model usage
- •Roles like trainers/contributors formalize expert involvement
- •Particularly relevant for science, math, medicine, and other hard domains
- 14:37 – 19:29
Structuring messy enterprise data, and why on-prem + open models matter
They discuss the practical barriers of enterprise data governance: structure, cleanliness, and sensitivity. Alex predicts sophisticated enterprises will prefer on-prem or open-source models they can customize without leaking their data to external model vendors.
- •Enterprise data is rarely clean/structured for easy model ingestion
- •Most advanced firms will mine internal data; many won’t in 5 years
- •Enterprise leaders view proprietary data as a key future differentiator
- •On-prem/open-source models (LLaMA/Mistral) fit needs for strict data guarantees
- 19:29 – 25:17
Data as the durable moat: competing on proprietary access as models commoditize
Alex contends compute and algorithms are hard to defend long-term, but data can create durable competitive advantage. They discuss exclusive/strategic content deals and how future labs will compete on unique data strategies rather than just GPU counts.
- •Algorithms diffuse; compute can be purchased—data is the sustainable moat
- •Examples: publisher/library partnerships as early signs of data moats
- •Future competition shifts from “GPU bragging” to “data access bragging”
- •Labs must develop differentiated strategies to produce/secure unique datasets
- 25:17 – 33:26
Where value accrues: models vs infrastructure vs services, and the future of pricing
Alex discusses uncertainty about value capture in the AI stack, pointing to NVIDIA as “below the model” and apps/services above it. They explore the ‘End of Software’ thesis, rising customization, changing engineering work, and a shift from per-seat to consumption pricing as agents do more work.
- •Value capture is moving; models may commoditize faster than expected
- •Infrastructure (e.g., NVIDIA) and applications/services may capture outsized value
- •Enterprises will demand more customization as software creation costs drop
- •Per-seat pricing weakens when AI agents (not employees) do the work; consumption pricing rises
- 33:26 – 37:20
Regulation and innovation: pro-data policy, pooling, and healthcare anonymization
Alex warns restrictive data regimes can tie a country’s hands, similar to constraining chip production. He argues for sector-wide pooling where it benefits everyone (e.g., safety/fraud) and for workable frameworks (e.g., anonymization) to unlock medical data for better health outcomes.
- •EU approach is portrayed as restrictive; US/UK need pro-data lens
- •Countries should avoid handicapping future data production
- •Data pooling can advance whole sectors (aerospace safety, financial fraud/compliance)
- •Healthcare: revise/clarify HIPAA/PII constraints via anonymization to enable learning from patient data
- 37:20 – 41:31
Geopolitics: China’s catch-up, industrial policy, and AI as a decisive military asset
The conversation shifts to China’s rapid progress and the risk of centralized industrial policy “turning the crank” faster than the West. Alex frames advanced AI/AGI as a potentially dominant military asset and argues the West must prevent adversarial AGI advantage.
- •China has closed gaps quickly; cites strong Chinese models nearing top leaderboards
- •CCP industrial policy excels at scaling established industries (solar, EVs)
- •AI/AGI could surpass nukes as a strategic military asset via strategy, hacking, weapons production
- •Argues for Western focus on leadership to avoid catastrophic geopolitical imbalance
- 41:31 – 42:50
Open vs closed AI: drawing the line for frontier systems
Alex argues for a bifurcated approach: keep the most powerful frontier systems closed for security/geopolitical reasons, while allowing openness for less-advanced models that primarily generate economic value. The key challenge becomes defining where the line sits as capabilities grow.
- •Most cutting-edge systems should remain closed for national security reasons
- •Open models can still be fine below a certain capability threshold
- •Views current open models as not yet ‘military assets’
- •Critical future task: determine and monitor the openness threshold as models advance
- 42:50 – 44:48
10-year foundation model landscape: ‘battle of giants’ and capital-driven consolidation
Looking ahead, Alex predicts foundation model development will become so expensive that only nation-states and hyperscalers can sustain it. He anticipates consolidation and highlights long-term uncertainty in major lab–hyperscaler partnerships.
- •Training frontier models may rise to tens/hundreds of billions in cost
- •Only giants (nations/hyperscalers) can underwrite such programs
- •Smaller labs likely get acquired or deeply integrated via partnerships
- •Open questions: how Microsoft–OpenAI and Amazon–Anthropic partnerships resolve long-term
- 44:48 – 52:12
Founder brand, PR, and media incentives: why ‘the best PR is no PR’
Alex argues traditional media incentives (clicks) create a build-up/tear-down cycle that can harm companies. He advocates direct channels—podcasts and founder-led communications—where messages aren’t distorted, emphasizing founder brand as a key distribution asset.
- •Traditional press optimizes for clicks; tends to amplify extremes
- •Direct channels let founders explain work without message distortion
- •Founder brand often outperforms company brand in audience attention
- •Alex cites unfair media treatment (including around defense work) versus more balanced congressional engagement
- 52:12 – 1:00:41
Hiring and leadership: talent density, founder approval, and the hypergrowth mistake
Alex explains Scale’s emphasis on hiring people who deeply care, maintaining an elite bar as the company grows. He shares that he still approves every hire and reflects on a major mistake: equating company hypergrowth with headcount hypergrowth, which diluted excellence.
- •‘People who give a shit’ sweat details and push through roadblocks
- •Alex claims he approves every hire; overrides team recommendation ~25–30%
- •Hyper-hiring (2020–2022) made it impossible to maintain the quality bar
- •Lesson: divorce team growth from company growth; prioritize talent density and excellence
- 1:00:41 – 1:06:04
Quick-fire reflections: misconceptions about AGI, election take, and Scale’s long-term aim
In rapid Q&A, Alex reiterates his belief that data—not just compute—is the major missing ingredient to AGI. He shares a cautious view on hype cycles (learning from autonomous vehicles), calls the US election a swing-state tossup, and positions Scale as a long-term “data foundry” for AI.
- •Changed mind: prioritize quality/excellence over headcount hypergrowth
- •Biggest AI misconception: compute alone leads to AGI—data is also required
- •Cautionary parallel to AV hype: promises can outpace technical reality and cause hangovers
- •Scale in 10 years: remain the ‘data pillar’/data foundry; interest in going public