a16zInside the Race to Measure Frontier Intelligence
At a glance
WHAT IT’S REALLY ABOUT
Why independent, private benchmarks are now essential for frontier AI
- VALS was created because open, public benchmarks became saturated, narrow, and increasingly easy for labs to game, making self-reported capability claims unreliable.
- The episode describes how high-stakes evaluation now happens under tight operational constraints—sometimes a six-hour pre-release window—driving investment in distributed infrastructure and automation.
- As AI shifts toward agentic, long-horizon work, evaluation shifts from many simple questions to fewer complex tasks with richer rubrics, forcing organizations to formalize previously “fuzzy” definitions of good work.
- Beyond capability, enterprises must evaluate models on cost, latency, and token efficiency, especially as token spend threatens to become a major budget line item relative to salaries.
- On policy and geopolitics, the speakers argue government should set rules while competent third parties test enforceable, evidence-based standards—potentially enabling “trust but verify” regimes across nations.
IDEAS WORTH REMEMBERING
5 ideasPublic benchmarks stopped being trustworthy once they became gameable.
Open-source questions and rubrics quickly become targets for optimization, creating a gap between “leaderboard performance” and real-world robustness. VALS positions itself as an independent verifier using held-out benchmarks to reveal that gap.
The Llama 4 episode illustrates how public scores can diverge from private reality.
Krishnan cites Meta’s Llama 4 as a case where public benchmark scores looked strong but the model underperformed on VALS’s private, held-out evaluations. The episode uses this as evidence for why independent, non-public testing matters.
Frontier evals are operationally constrained—speed and infrastructure are part of the moat.
Model launches can provide only hours of access before release, forcing evaluators to maximize signal under strict rate limits and time constraints. VALS invested in distributed infrastructure and internal automation (“Steve”) to compress high-quality testing into that window.
The next bottleneck is turning messy human workflows into explicit, legible evals.
As evals move from simple Q&A to agentic, long-horizon tasks, the number of test items shrinks while rubrics and success criteria become more complex. This pushes evaluators to formalize “fuzzy” real-world expectations (e.g., what counts as partner-level legal work).
Recursive self-improvement can be benchmarked via proxies, not full self-training runs.
Instead of literally training a model to produce its successor (too slow/expensive), VALS builds proxy tests for pieces of the process: pre-training, post-training, engineering, and research behaviors. The goal is an apples-to-apples way to discuss recursive self-improvement potential across labs.
WORDS WORTH SAVING
5 quotesEvery time a new trillion-dollar industry emerges, there's a need for this independent testing group.
— Rayan Krishnan
When Meta released Llama 4, um, that was a bit of a disaster.
— Rayan Krishnan
It's forcing a lot of the, um, more fuzzy or distributed forms of evals to be made explicit.
— Rayan Krishnan
It's 10X more we were spending in tokens than employee salary for that month.
— Rayan Krishnan
The government kind of has an inclination of what it's afraid of, be it biohacking or cyber hacking or so forth, but th-then there becomes the question of, okay, can the model do it, and then can you get the model to do it?
— Ben Horowitz
High quality AI-generated summary created from speaker-labeled transcript.