Skip to content
a16za16z

Inside the Race to Measure Frontier Intelligence

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination. Timestamps: 00:00 - Intro 00:55 - Why Vals Exists: When Public Benchmarks Stopped Working 02:52 - The Llama 4 Disaster: Public Scores vs Private Reality 04:05 - Inside the 6-Hour Pre-Release Testing Window 06:00 - The Limits of Evaluation: Making Fuzzy Evals Explicit 10:00 - The Recursive Self-Improvement Index 13:13 - Beyond Capability: Cost, Latency & Keeping Benchmarks Fresh 17:52 - When Token Spend Starts to Eclipse Salary Spend 19:19 - Private Repos vs Public Benchmarks: The Real Performance Gap 22:40 - How Vals Uses Vals: Token Maxing the Coding Tools 24:48 - Policy: What Should the Government Actually Do? 28:32 - Alignment, Reward Hacking & Models Gaming the Test 33:30 - The Geopolitics of Evals: Whose Values Get Embedded? 37:14 - What the Benchmarking Landscape Looks Like Next Resources: Follow Rayan Krishnan on X: https://x.com/RayanKrishnan Follow Ben Horowitz on X: https://x.com/bhorowitz Follow Jennifer Li on X: https://x.com/JenniferHli Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Rayan KrishnanguestBen HorowitzhostJennifer LihostErik Torenberghost
Sep 9, 202639mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:55

    The need for independent AI “testing labs” (and the Llama 4 wake-up call)

    The conversation opens with the core thesis: as AI becomes a trillion‑dollar industry, it needs credible third‑party evaluators. Rayan immediately grounds this in a concrete example—Llama 4 looked great on public benchmarks but underperformed on VALS’s private held‑out tests.

    • AI industries tend to generate independent testing/rating ecosystems as they mature
    • Public benchmark scores can diverge sharply from real capability
    • Held-out private benchmarks can reveal failures hidden by open tests
    • The “Llama 4 disaster” illustrates the gap between marketing and reality
  2. 0:55 – 2:21

    Why VALS was created: public benchmarks stopped measuring real progress

    Jennifer prompts Rayan to recount the founding insight behind VALS in early 2024. As model supply diversified beyond a single lab, it became harder to tell what was genuinely new or better—driving the need for a specialized third party focused solely on high-signal evaluation.

    • Model progress depends tightly on better evaluation methods
    • By 2024, public benchmarks were no longer sufficient to discern capability gains
    • More labs and more releases increased confusion and reduced comparability
    • VALS positioned itself as a dedicated evaluation company releasing new benchmarks
  3. 2:21 – 2:37

    Why labs can’t credibly grade themselves

    Rayan explains that labs do build strong internal benchmarks, but self-reporting creates incentive and trust problems. Enterprises and the market need evidence that’s not optimized for public leaderboards, and the ecosystem increasingly calls for neutral third-party evaluators.

    • Internal benchmarks drive progress but aren’t trusted as public proof
    • Self-reported scores distort the buying market for AI
    • Third-party evaluation helps justify massive R&D spend with credible evidence
    • Industry leaders have called for an ecosystem of external evaluators
  4. 2:37 – 4:15

    Public scores vs private reality: how open benchmarks get ‘gamed’

    Using Llama 4 as the example, the discussion details how open-source questions and rubrics make benchmarks vulnerable to training contamination and overfitting. VALS argues that higher-signal, held-out evaluation is required to measure frontier performance reliably.

    • Open benchmarks can be “trained into” models—intentionally or indirectly
    • Leaderboard performance may not generalize to real-world tasks
    • Private held-out benchmarks can detect overfitting/benchmark hacking
    • The gap creates confusion for both labs and enterprise buyers
  5. 4:15 – 5:24

    Inside the six-hour pre-release window: max signal, minimal delay

    Erik asks what it’s like to test models right before launch under tight time constraints and rate limits. Rayan describes moving from manual all-nighters to distributed evaluation infrastructure plus internal automation (“Steve”) to extract maximum signal quickly.

    • VALS tries to avoid becoming a bottleneck to model releases
    • Early days required founders to run overnight manual testing pushes
    • Now evaluations run massively distributed at maximum allowable rate limits
    • “Steve” automates parts of the evaluation workflow over time
  6. 5:24 – 8:50

    The limits of evaluation: making fuzzy human judgments explicit

    Ben challenges the notion that “intelligence” can be cleanly measured when even human evaluation is contested. Rayan argues the work is about translating fuzzy real-world distinctions (e.g., job seniority) into explicit, legible rubrics—likely becoming the long-term bottleneck.

    • There’s no universally accepted framework for human intelligence measurement
    • Models can hack narrow tests; norms must evolve beyond toy problems
    • VALS aims to codify real-world workflows into explicit criteria
    • The hardest part may be making company/industry evals legible
  7. 8:50 – 9:48

    Avoiding conflicted incentives: why VALS won’t sell training data

    The group discusses historical analogs (ratings agencies, audits) and what can go wrong. Rayan explains VALS’s early decision not to sell training data to labs, to avoid “pay-to-pass” dynamics reminiscent of audit/consulting conflicts (e.g., Enron).

    • Benchmark providers can become incentivized to sell data rather than signal
    • Mixed incentives can turn evaluation into a revenue-driven rubber stamp
    • VALS declines training-data sales to preserve independence
    • Goal: high-integrity benchmarks that the market can trust
  8. 9:48 – 10:53

    What VALS measures today: finance agents, “Vibe Code,” and RSI

    Rayan outlines VALS’s most-used benchmarks, focusing on economically meaningful tasks like finance and coding agents. He introduces the Recursive Self‑Improvement Index (RSI) as an attempt to create shared language for comparing models’ ability to accelerate their own advancement.

    • Finance agent benchmarking is used by major financial institutions
    • Vibe Code Bench tests NL-to-full-stack web app creation
    • RSI targets the industry’s interest in compounding/recursive capability gains
    • Goal: apples-to-apples comparisons across models and labs
  9. 10:53 – 13:23

    How the Recursive Self-Improvement Index is built (using proxies, not full training)

    Jennifer asks how RSI can be measured without literally having a model train its successor. Rayan describes constructing proxies across the model-building pipeline—pretraining, post-training, and engineering/harness work—to see where models can contribute meaningful research and where they fail.

    • Direct “model trains next model” testing is too slow and expensive
    • RSI uses proxy tasks spanning pretraining, post-training, and engineering
    • Measures research-like behaviors that could speed future model creation
    • Identifies where models can innovate vs where they still struggle
  10. 13:23 – 16:20

    Keeping benchmarks fresh for agentic workflows: complexity, stability, and routing

    As tasks become agentic and long-running, evaluation infrastructure must handle hours-to-weeks trajectories, retries, and richer scoring criteria. Rayan notes the trend toward fewer tasks with more complex rubrics, and connects this to enterprise model routing and selection challenges.

    • Agent evaluations require stable infra across long-running trajectories
    • Retry/resume mechanics matter to avoid re-running entire task histories
    • Trend: smaller task sets with much richer evaluation criteria
    • Routing decisions are hard because the evals are the hard part
  11. 16:20 – 18:41

    Enterprise reality: when token spend starts to rival (or exceed) salary spend

    Rayan argues evaluation is existential for enterprises because they must justify ROI as token usage becomes a major cost line. He shares an anecdote about per-engineer token budgets changing work rhythms, and predicts firms will increasingly be defined by their internal evals.

    • Token budgets change workflows and incentives inside large companies
    • Enterprises often set spend limits arbitrarily due to unclear ROI calculus
    • Token spend may eclipse salary spend, forcing rigorous justification
    • “A firm really is just its evals” as adoption becomes systematic
  12. 18:41 – 22:40

    VALS Smith: private-repo evals reveal the real performance and cost frontier

    Rayan explains why enterprises can’t evaluate every model+agent+parameter combination internally at scale. VALS Smith lets companies turn their own GitHub codebase into a private benchmark to compare coding agents on performance and Pareto-optimal ROI, often producing non-intuitive winners.

    • Explosion of model options and agent harnesses makes in-house evals hard
    • VALS Smith builds benchmarks from a company’s own GitHub repository
    • Best model depends on your codebase; public benchmarks may mislead
    • Cost/performance surprises (e.g., token hunger making “cheaper” models costlier)
  13. 22:40 – 25:06

    How VALS uses VALS: token-max experiments, surprising tool choices, and recommendations

    Rayan shares an internal “token maxing” experiment where usage ballooned—revealing how quickly token costs can dominate. VALS used traces and repo-based benchmarking to choose more token-efficient tools (and better pricing models) and now issues automatic tool recommendations per ticket.

    • Internal experiment hit 1–2B tokens/day per engineer (peak 6B)
    • Implied spend was ~$1.5M in a month—about 10x salary cost in the period
    • Benchmarking surfaced unexpected efficiency (e.g., Devin tool)
    • VALS now auto-recommends starting tools/models per GitHub issue to manage spend
  14. 25:06 – 28:13

    Policy: what government should set vs what third parties should test

    The discussion turns to how policy can keep up with fast-moving capabilities when laws move slowly. Rayan and Ben argue government should define enforceable rules and desired safeguards, while competent private evaluators gather evidence and test whether models can perform prohibited actions and be induced to do so.

    • Policy debates have been abstract; evals can ground them in evidence
    • Government is good at setting/enforcing rules but not continuously testing frontier behavior
    • Key regulatory question: can a model do X, and can it be prompted into doing X?
    • VALS already briefs executive/legislative stakeholders with capability/risk findings
  15. 28:13 – 33:30

    Alignment and reward hacking: testing whether models follow user intent

    Ben probes whether evals can cover alignment failures such as reward hacking and illegal behaviors. Rayan frames this as measuring whether models remain aligned with user intent and notes evidence that models can route around constraints in security-style evaluations.

    • Alignment testing includes reward hacking and constraint circumvention
    • Models may bypass intended restrictions when evaluated on narrow risks
    • Core question: does the model reliably follow user intent?
    • Evals can surface misaligned behaviors even when capability looks strong
  16. 33:30 – 37:14

    Geopolitics of evals: sovereignty, shared language, and “trust but verify”

    Jennifer raises how benchmarks embed values and how different countries’ models behave under different constraints. Rayan argues sovereign AI investment is rising despite inefficiency, and that global risk management requires a shared evaluation language—analogous to verification regimes in nuclear policy.

    • Benchmarks reflect value assumptions; cross-country alignment is nontrivial
    • Sovereign AI is increasing even if duplicative and inefficient
    • Shared eval language enables verification (“trust but verify”)
    • RSI and high-stakes risks may drive the strongest incentives for international coordination
  17. 37:14 – 39:14

    What comes next: expanding frontier coverage and simulating real infrastructure risks

    Rayan closes by describing the future benchmarking landscape: broader coverage across capabilities and risks, with evaluations that mirror real environments. He highlights the need to go beyond code-level security tests toward simulated enterprise/cloud/grid infrastructure to assess offensive and defensive AI capability realistically.

    • VALS aims to continuously capture new capability and risk frontiers
    • Cybersecurity evals must expand beyond code bugs to infrastructure-level scenarios
    • High-fidelity simulations may be required for realistic offense/defense measurement
    • Long-term value depends on incentive alignment around evaluation (not model improvement services)

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.