Skip to content
YC Root AccessYC Root Access

Improving Small Language Model Reasoning With A* Search

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Alexander Braverman about a test-time scaling method for improving reasoning in smaller language models. Instead of relying on a larger teacher model or an external reward model, the method uses the language model’s own self-critique as a heuristic in an A*-inspired search. It explores multiple reasoning paths, deprioritizes weaker branches, and searches for a stronger answer using the same underlying model. On mathematical reasoning benchmarks, the method improved accuracy more efficiently than other test-time approaches at comparable token and runtime budgets. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostAlexander Bravermanguest
Aug 6, 20266mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
August 6, 2026
Duration
6m
Channel
YC Root Access
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Alexander Braverman about a test-time scaling method for improving reasoning in smaller language models. Instead of relying on a larger teacher model or an external reward model, the method uses the language model’s own self-critique as a heuristic in an A*-inspired search. It explores multiple reasoning paths, deprioritizes weaker branches, and searches for a stronger answer using the same underlying model. On mathematical reasoning benchmarks, the method improved accuracy more efficiently than other test-time approaches at comparable token and runtime budgets. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

SPEAKERS

  • Ankit Gupta

    host

    Host/interviewer on YC Root Access (Y Combinator).

  • Alexander Braverman

    guest

    Researcher presenting work on improving small language model reasoning using A*-style test-time search and self-critique heuristics.

EPISODE SUMMARY

In this episode of YC Root Access, featuring Ankit Gupta and Alexander Braverman, Improving Small Language Model Reasoning With A* Search explores a* search uses self-critique to boost small-model reasoning accuracy The talk introduces a test-time scaling method that treats LLM reasoning as tree search and applies an A*-inspired algorithm to small language models.

RELATED EPISODES

Evaluating the Fine-Grained Planning Abilities of Web Agents

Evaluating the Fine-Grained Planning Abilities of Web Agents

ChartNet: Training Vision-Language Models to Understand Charts

ChartNet: Training Vision-Language Models to Understand Charts

Any-Horizon Reasoning for Video Agents

Any-Horizon Reasoning for Video Agents

LeanAgent: Lifelong Learning for Formal Theorem Proving

LeanAgent: Lifelong Learning for Formal Theorem Proving

Zero-Shot Predictive Models for Relational Databases

Zero-Shot Predictive Models for Relational Databases

Interpretability and Safety for Robot Foundation Models

Interpretability and Safety for Robot Foundation Models

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.