Skip to content
YC Root AccessYC Root Access

Any-Horizon Reasoning for Video Agents

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Jitesh Jain about building video agents that can adapt their reasoning to videos of different lengths. Current agents often struggle with long, open-ended video questions because temporal grounding is unreliable and training data is expensive. SAGE combines visual tools with transcripts and web search, then uses synthetic question-answer data, tool trajectories, and reinforcement learning to teach the model when each source of information is useful. As videos become longer, the agent takes more reasoning steps and improves more over the base model, suggesting that it is learning to spend effort according to the task. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostJitesh Jainguest
Aug 6, 20264mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
August 6, 2026
Duration
4m
Channel
YC Root Access
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Jitesh Jain about building video agents that can adapt their reasoning to videos of different lengths. Current agents often struggle with long, open-ended video questions because temporal grounding is unreliable and training data is expensive. SAGE combines visual tools with transcripts and web search, then uses synthetic question-answer data, tool trajectories, and reinforcement learning to teach the model when each source of information is useful. As videos become longer, the agent takes more reasoning steps and improves more over the base model, suggesting that it is learning to spend effort according to the task. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

SPEAKERS

  • Ankit Gupta

    host

    Host/interviewer on YC Root Access (Y Combinator).

  • Jitesh Jain

    guest

    Researcher presenting his CVPR work on any-horizon reasoning for long-video agents.

EPISODE SUMMARY

In this episode of YC Root Access, featuring Ankit Gupta and Jitesh Jain, Any-Horizon Reasoning for Video Agents explores video agents learn adaptive, any-horizon reasoning with tools and RL The work targets “any-horizon” behavior: agents should skim long videos but fully watch short clips, adapting effort to the question and video length.

RELATED EPISODES

Evaluating the Fine-Grained Planning Abilities of Web Agents

Evaluating the Fine-Grained Planning Abilities of Web Agents

ChartNet: Training Vision-Language Models to Understand Charts

ChartNet: Training Vision-Language Models to Understand Charts

LeanAgent: Lifelong Learning for Formal Theorem Proving

LeanAgent: Lifelong Learning for Formal Theorem Proving

Zero-Shot Predictive Models for Relational Databases

Zero-Shot Predictive Models for Relational Databases

Improving Small Language Model Reasoning With A* Search

Improving Small Language Model Reasoning With A* Search

Interpretability and Safety for Robot Foundation Models

Interpretability and Safety for Robot Foundation Models

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.