Skip to content
a16za16z

Inside the Race to Measure Frontier Intelligence

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination. Timestamps: 00:00 - Intro 00:55 - Why Vals Exists: When Public Benchmarks Stopped Working 02:52 - The Llama 4 Disaster: Public Scores vs Private Reality 04:05 - Inside the 6-Hour Pre-Release Testing Window 06:00 - The Limits of Evaluation: Making Fuzzy Evals Explicit 10:00 - The Recursive Self-Improvement Index 13:13 - Beyond Capability: Cost, Latency & Keeping Benchmarks Fresh 17:52 - When Token Spend Starts to Eclipse Salary Spend 19:19 - Private Repos vs Public Benchmarks: The Real Performance Gap 22:40 - How Vals Uses Vals: Token Maxing the Coding Tools 24:48 - Policy: What Should the Government Actually Do? 28:32 - Alignment, Reward Hacking & Models Gaming the Test 33:30 - The Geopolitics of Evals: Whose Values Get Embedded? 37:14 - What the Benchmarking Landscape Looks Like Next Resources: Follow Rayan Krishnan on X: https://x.com/RayanKrishnan Follow Ben Horowitz on X: https://x.com/bhorowitz Follow Jennifer Li on X: https://x.com/JenniferHli Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Rayan KrishnanguestBen HorowitzhostJennifer LihostErik Torenberghost
Sep 9, 202639mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
September 9, 2026
Duration
39m
Channel
a16z
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination. Timestamps: 00:00 - Intro 00:55 - Why Vals Exists: When Public Benchmarks Stopped Working 02:52 - The Llama 4 Disaster: Public Scores vs Private Reality 04:05 - Inside the 6-Hour Pre-Release Testing Window 06:00 - The Limits of Evaluation: Making Fuzzy Evals Explicit 10:00 - The Recursive Self-Improvement Index 13:13 - Beyond Capability: Cost, Latency & Keeping Benchmarks Fresh 17:52 - When Token Spend Starts to Eclipse Salary Spend 19:19 - Private Repos vs Public Benchmarks: The Real Performance Gap 22:40 - How Vals Uses Vals: Token Maxing the Coding Tools 24:48 - Policy: What Should the Government Actually Do? 28:32 - Alignment, Reward Hacking & Models Gaming the Test 33:30 - The Geopolitics of Evals: Whose Values Get Embedded? 37:14 - What the Benchmarking Landscape Looks Like Next Resources: Follow Rayan Krishnan on X: https://x.com/RayanKrishnan Follow Ben Horowitz on X: https://x.com/bhorowitz Follow Jennifer Li on X: https://x.com/JenniferHli Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

SPEAKERS

  • Rayan Krishnan

    guest

    AI evaluation and benchmarking practitioner at VALS, focusing on third-party model evaluations.

  • Ben Horowitz

    host

    Co-founder and General Partner at Andreessen Horowitz (a16z).

  • Jennifer Li

    host

    Andreessen Horowitz (a16z) host/investor with a perspective informed by being born and raised in China and following Chinese open-source models.

  • Erik Torenberg

    host

    Podcast host at Andreessen Horowitz (a16z), facilitating interviews on technology and policy.

EPISODE SUMMARY

In this episode of a16z, featuring Rayan Krishnan and Ben Horowitz, Inside the Race to Measure Frontier Intelligence explores why independent, private benchmarks are now essential for frontier AI VALS was created because open, public benchmarks became saturated, narrow, and increasingly easy for labs to game, making self-reported capability claims unreliable.

RELATED EPISODES

Why AI Demand Is Outrunning Compute Supply

Why AI Demand Is Outrunning Compute Supply

How AI Changes the Economics of Innovation

How AI Changes the Economics of Innovation

Why Top Founders Are Racing Into AI Infrastructure

Why Top Founders Are Racing Into AI Infrastructure

How Cursor Built One of AI’s Fastest-Growing Companies

How Cursor Built One of AI’s Fastest-Growing Companies

The State of AI: Models, Moats, and the Consumer Renaissance

The State of AI: Models, Moats, and the Consumer Renaissance

Inside OpenAI’s Breakthroughs in Mathematical Reasoning

Inside OpenAI’s Breakthroughs in Mathematical Reasoning

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.