Skip to content
YC Root AccessYC Root Access

Interpretability and Safety for Robot Foundation Models

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Bear Häon about applying interpretability and AI safety techniques to vision-language-action models. The work examines activations inside a vision-language-action model, groups neurons associated with concepts such as speed or caution, and then steers the robot's behavior by amplifying those groups. This provides a way to better understand and control how robots translate language into physical actions. The broader goal is to develop safety methods for physical AI, where failures can have direct consequences in the real world. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostBear Häonguest
Aug 6, 20265mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
August 6, 2026
Duration
5m
Channel
YC Root Access
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Bear Häon about applying interpretability and AI safety techniques to vision-language-action models. The work examines activations inside a vision-language-action model, groups neurons associated with concepts such as speed or caution, and then steers the robot's behavior by amplifying those groups. This provides a way to better understand and control how robots translate language into physical actions. The broader goal is to develop safety methods for physical AI, where failures can have direct consequences in the real world. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

SPEAKERS

  • Ankit Gupta

    host

    Host at YC Root Access (Y Combinator) who interviews guests about machine learning and startups.

  • Bear Häon

    guest

    Founder of the Physical AI Safety Institute and researcher focused on interpretability and safety for robot foundation models.

EPISODE SUMMARY

In this episode of YC Root Access, featuring Ankit Gupta and Bear Häon, Interpretability and Safety for Robot Foundation Models explores interpreting robot foundation models to control behavior and improve safety The discussion presents one of the first interpretability papers for Robot Foundation Models (RFMs), adapting AI-safety interpretability techniques to embodied, physically acting systems.

RELATED EPISODES

Evaluating the Fine-Grained Planning Abilities of Web Agents

Evaluating the Fine-Grained Planning Abilities of Web Agents

ChartNet: Training Vision-Language Models to Understand Charts

ChartNet: Training Vision-Language Models to Understand Charts

Any-Horizon Reasoning for Video Agents

Any-Horizon Reasoning for Video Agents

LeanAgent: Lifelong Learning for Formal Theorem Proving

LeanAgent: Lifelong Learning for Formal Theorem Proving

Zero-Shot Predictive Models for Relational Databases

Zero-Shot Predictive Models for Relational Databases

Improving Small Language Model Reasoning With A* Search

Improving Small Language Model Reasoning With A* Search

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.