Skip to content
a16za16z

Inside the Race to Measure Frontier Intelligence

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination. Timestamps: 00:00 - Intro 00:55 - Why Vals Exists: When Public Benchmarks Stopped Working 02:52 - The Llama 4 Disaster: Public Scores vs Private Reality 04:05 - Inside the 6-Hour Pre-Release Testing Window 06:00 - The Limits of Evaluation: Making Fuzzy Evals Explicit 10:00 - The Recursive Self-Improvement Index 13:13 - Beyond Capability: Cost, Latency & Keeping Benchmarks Fresh 17:52 - When Token Spend Starts to Eclipse Salary Spend 19:19 - Private Repos vs Public Benchmarks: The Real Performance Gap 22:40 - How Vals Uses Vals: Token Maxing the Coding Tools 24:48 - Policy: What Should the Government Actually Do? 28:32 - Alignment, Reward Hacking & Models Gaming the Test 33:30 - The Geopolitics of Evals: Whose Values Get Embedded? 37:14 - What the Benchmarking Landscape Looks Like Next Resources: Follow Rayan Krishnan on X: https://x.com/RayanKrishnan Follow Ben Horowitz on X: https://x.com/bhorowitz Follow Jennifer Li on X: https://x.com/JenniferHli Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Rayan KrishnanguestBen HorowitzhostJennifer LihostErik Torenberghost
Sep 9, 202639mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Why independent, private benchmarks are now essential for frontier AI

  1. VALS was created because open, public benchmarks became saturated, narrow, and increasingly easy for labs to game, making self-reported capability claims unreliable.
  2. The episode describes how high-stakes evaluation now happens under tight operational constraints—sometimes a six-hour pre-release window—driving investment in distributed infrastructure and automation.
  3. As AI shifts toward agentic, long-horizon work, evaluation shifts from many simple questions to fewer complex tasks with richer rubrics, forcing organizations to formalize previously “fuzzy” definitions of good work.
  4. Beyond capability, enterprises must evaluate models on cost, latency, and token efficiency, especially as token spend threatens to become a major budget line item relative to salaries.
  5. On policy and geopolitics, the speakers argue government should set rules while competent third parties test enforceable, evidence-based standards—potentially enabling “trust but verify” regimes across nations.

IDEAS WORTH REMEMBERING

5 ideas

Public benchmarks stopped being trustworthy once they became gameable.

Open-source questions and rubrics quickly become targets for optimization, creating a gap between “leaderboard performance” and real-world robustness. VALS positions itself as an independent verifier using held-out benchmarks to reveal that gap.

The Llama 4 episode illustrates how public scores can diverge from private reality.

Krishnan cites Meta’s Llama 4 as a case where public benchmark scores looked strong but the model underperformed on VALS’s private, held-out evaluations. The episode uses this as evidence for why independent, non-public testing matters.

Frontier evals are operationally constrained—speed and infrastructure are part of the moat.

Model launches can provide only hours of access before release, forcing evaluators to maximize signal under strict rate limits and time constraints. VALS invested in distributed infrastructure and internal automation (“Steve”) to compress high-quality testing into that window.

The next bottleneck is turning messy human workflows into explicit, legible evals.

As evals move from simple Q&A to agentic, long-horizon tasks, the number of test items shrinks while rubrics and success criteria become more complex. This pushes evaluators to formalize “fuzzy” real-world expectations (e.g., what counts as partner-level legal work).

Recursive self-improvement can be benchmarked via proxies, not full self-training runs.

Instead of literally training a model to produce its successor (too slow/expensive), VALS builds proxy tests for pieces of the process: pre-training, post-training, engineering, and research behaviors. The goal is an apples-to-apples way to discuss recursive self-improvement potential across labs.

WORDS WORTH SAVING

5 quotes

Every time a new trillion-dollar industry emerges, there's a need for this independent testing group.

Rayan Krishnan

When Meta released Llama 4, um, that was a bit of a disaster.

Rayan Krishnan

It's forcing a lot of the, um, more fuzzy or distributed forms of evals to be made explicit.

Rayan Krishnan

It's 10X more we were spending in tokens than employee salary for that month.

Rayan Krishnan

The government kind of has an inclination of what it's afraid of, be it biohacking or cyber hacking or so forth, but th-then there becomes the question of, okay, can the model do it, and then can you get the model to do it?

Ben Horowitz

Private vs public benchmarksBenchmark gaming and incentive conflictsPre-release evaluation windows and infrastructureAgentic and long-horizon evaluationsMaking “fuzzy” real-world rubrics explicitRecursive Self-Improvement Index (RSI)Enterprise ROI: cost, latency, token efficiency

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.