a16zInside the Race to Measure Frontier Intelligence
Episode Details
EPISODE INFO
- Released
- September 9, 2026
- Duration
- 39m
- Channel
- a16z
- Watch on YouTube
- ▶ Open ↗
EPISODE DESCRIPTION
a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination. Timestamps: 00:00 - Intro 00:55 - Why Vals Exists: When Public Benchmarks Stopped Working 02:52 - The Llama 4 Disaster: Public Scores vs Private Reality 04:05 - Inside the 6-Hour Pre-Release Testing Window 06:00 - The Limits of Evaluation: Making Fuzzy Evals Explicit 10:00 - The Recursive Self-Improvement Index 13:13 - Beyond Capability: Cost, Latency & Keeping Benchmarks Fresh 17:52 - When Token Spend Starts to Eclipse Salary Spend 19:19 - Private Repos vs Public Benchmarks: The Real Performance Gap 22:40 - How Vals Uses Vals: Token Maxing the Coding Tools 24:48 - Policy: What Should the Government Actually Do? 28:32 - Alignment, Reward Hacking & Models Gaming the Test 33:30 - The Geopolitics of Evals: Whose Values Get Embedded? 37:14 - What the Benchmarking Landscape Looks Like Next Resources: Follow Rayan Krishnan on X: https://x.com/RayanKrishnan Follow Ben Horowitz on X: https://x.com/bhorowitz Follow Jennifer Li on X: https://x.com/JenniferHli Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.
SPEAKERS
Rayan Krishnan
guestAI evaluation and benchmarking practitioner at VALS, focusing on third-party model evaluations.
Ben Horowitz
hostCo-founder and General Partner at Andreessen Horowitz (a16z).
Jennifer Li
hostAndreessen Horowitz (a16z) host/investor with a perspective informed by being born and raised in China and following Chinese open-source models.
Erik Torenberg
hostPodcast host at Andreessen Horowitz (a16z), facilitating interviews on technology and policy.
EPISODE SUMMARY
In this episode of a16z, featuring Rayan Krishnan and Ben Horowitz, Inside the Race to Measure Frontier Intelligence explores why independent, private benchmarks are now essential for frontier AI VALS was created because open, public benchmarks became saturated, narrow, and increasingly easy for labs to game, making self-reported capability claims unreliable.
RELATED EPISODES