Episode Details
EPISODE INFO
- Released
- August 6, 2026
- Duration
- 4m
- Channel
- YC Root Access
- Watch on YouTube
- ▶ Open ↗
EPISODE DESCRIPTION
At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Jitesh Jain about building video agents that can adapt their reasoning to videos of different lengths. Current agents often struggle with long, open-ended video questions because temporal grounding is unreliable and training data is expensive. SAGE combines visual tools with transcripts and web search, then uses synthetic question-answer data, tool trajectories, and reinforcement learning to teach the model when each source of information is useful. As videos become longer, the agent takes more reasoning steps and improves more over the base model, suggesting that it is learning to spend effort according to the task. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs
SPEAKERS
Ankit Gupta
hostHost/interviewer on YC Root Access (Y Combinator).
Jitesh Jain
guestResearcher presenting his CVPR work on any-horizon reasoning for long-video agents.
EPISODE SUMMARY
In this episode of YC Root Access, featuring Ankit Gupta and Jitesh Jain, Any-Horizon Reasoning for Video Agents explores video agents learn adaptive, any-horizon reasoning with tools and RL The work targets “any-horizon” behavior: agents should skim long videos but fully watch short clips, adapting effort to the question and video length.
RELATED EPISODES
