Skip to content
YC Root AccessYC Root Access

Any-Horizon Reasoning for Video Agents

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Jitesh Jain about building video agents that can adapt their reasoning to videos of different lengths. Current agents often struggle with long, open-ended video questions because temporal grounding is unreliable and training data is expensive. SAGE combines visual tools with transcripts and web search, then uses synthetic question-answer data, tool trajectories, and reinforcement learning to teach the model when each source of information is useful. As videos become longer, the agent takes more reasoning steps and improves more over the base model, suggesting that it is learning to spend effort according to the task. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostJitesh Jainguest
Aug 6, 20264mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Video agents learn adaptive, any-horizon reasoning with tools and RL

  1. The work targets “any-horizon” behavior: agents should skim long videos but fully watch short clips, adapting effort to the question and video length.
  2. Existing video agent systems underperform on open-ended queries and are too slow, largely because they depend on imperfect temporal grounding models not trained on long videos.
  3. The proposed system supplements visual reasoning with knowledge-grounded tools like web search and speech transcripts to reduce reliance on fragile temporal localization.
  4. To avoid expensive human annotation of long videos, the team generates synthetic question–answer data and tool-use trajectories (using Gemini 2.5 Flash) for supervised finetuning and as a precursor to RL.
  5. A multi-reward RL setup uses an LLM judge for open-ended evaluation, penalizes unnecessary tool calls on wrong answers, and selectively rewards correct use of visual grounding tools to encourage adaptive tool usage.

IDEAS WORTH REMEMBERING

5 ideas

Any-horizon reasoning is the core capability missing from many video agents.

Humans naturally change how much they watch based on the task (skip in long content, watch fully in short clips), while prior agents tend to apply a fixed, inefficient strategy.

Temporal grounding accuracy is a key bottleneck for long-video agents.

Many systems route decisions through temporal localization models that are not robust—partly because they haven’t been trained on long-video distributions—leading to wrong or slow retrieval of relevant moments.

Adding non-visual tools can compensate for weak visual grounding.

Using speech transcripts and web search provides external signals and world knowledge that help the agent narrow what to inspect visually, improving both effectiveness and runtime.

Synthetic data is a practical substitute for costly long-video annotation.

Since hour-long annotation can exceed ~$30/hour, generating QA pairs and tool-use trajectories with a strong model shifts cost to a scalable one-time generation pipeline.

Open-ended video QA needs judge-based rewards rather than string matching.

Because answers can be phrased many ways, an LLM-as-judge enables reward assignment for correctness and quality when exact-match metrics fail.

WORDS WORTH SAVING

5 quotes

This is what we term as the natural any horizon ability of humans. Humans are very good at adapting how much time do they spend on what based on the task at hand.

Jitesh Jain

We found that existing systems usually relied on temporal grounding, which is similar to grip for coding agents, and temporal grounding models for videos are not that accurate yet.

Jitesh Jain

They are not accurate because the models have not been trained on long videos-

Jitesh Jain

So we equip our system with tools like web search and transcribed speech, which usually have external signals that can help us alleviate the issue that arises from inaccuracy of the temporal grounding model, models, and that makes the system m-much more effective and efficient.

Jitesh Jain

Second with videos is videos are long, one hour, and if you use existing annotated, uh, platforms like Prolific, one hour of the annotated time can cost in ac- in excess of thirty dollars.

Jitesh Jain

Any-horizon human-like viewing behaviorFailures of temporal grounding on long videosKnowledge-grounded reasoning via web search and transcriptsSynthetic QA and tool-trajectory generationSFT-to-RL training pipelineMulti-reward reinforcement learning with LLM judgeRuntime and scaling with video duration

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.