Skip to content
YC Root AccessYC Root Access

Interpretability and Safety for Robot Foundation Models

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Bear Häon about applying interpretability and AI safety techniques to vision-language-action models. The work examines activations inside a vision-language-action model, groups neurons associated with concepts such as speed or caution, and then steers the robot's behavior by amplifying those groups. This provides a way to better understand and control how robots translate language into physical actions. The broader goal is to develop safety methods for physical AI, where failures can have direct consequences in the real world. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostBear Häonguest
Aug 6, 20265mWatch on YouTube ↗

CHAPTERS

  1. 0:07 – 0:30

    Why robot foundation models need interpretability now

    Ankit introduces Bear Häon and his Berkeley work on robot AI safety and the Physical AI Safety Institute. Bear frames the talk around their effort to create the first interpretability paper specifically for Robot Foundation Models (RFMs).

    • Bear’s background: Berkeley research + founder of the Physical AI Safety Institute
    • Context: first interpretability work for RFMs appeared only last year
    • Goal: connect interpretability to concrete safety outcomes for physically acting systems
  2. 0:30 – 1:10

    From “AI safety for digital agents” to safety for embodied robots

    Bear explains the motivation: AI safety has developed interpretability, alignment, and control tools mostly for systems that reason and act digitally. The key question is whether these techniques can transfer to systems that also act in the physical world.

    • AI safety toolkits were built for mostly digital action spaces
    • Embodied robots add new failure modes due to real-world actuation
    • The paper is positioned as a first-pass feasibility demonstration
  3. 1:10 – 1:19

    What the model is: VLA as a fine-tuned vision-language model for action

    Bear describes the starting point for their method: a VLA model, which he characterizes as effectively a fine-tuned vision-language model (VLM). This grounds the interpretability approach in standard transformer internals.

    • VLA ≈ fine-tuned VLM adapted for robot control
    • Interpretability target: standard transformer architecture
    • Sets up later discussion of transformer blocks and FFNs
  4. 1:19 – 1:50

    Peering inside transformers: FFNs as concept “dictionaries” via token activations

    The technique focuses on transformer feedforward networks (FFNs), treated as MLPs within each block. By projecting activations onto the embedding/token space, they derive a representation akin to a dictionary showing which tokens strongly activate which FFNs.

    • Transformer blocks contain FFNs (MLPs) that can be analyzed per-layer
    • Projecting onto the embedding layer reveals token-level interpretability signals
    • Creates an FFN-to-token activation “dictionary” across the network
  5. 1:50 – 2:05

    Clustering internal features into concepts (fast/slow/cautious)

    With FFN activation representations in hand, Bear explains clustering them into higher-level concepts (e.g., quick, cautious). This turns low-level neuron activity into more human-legible behavioral dimensions.

    • FFN representations can be clustered by semantic similarity
    • Clusters correspond to behavioral concepts (e.g., fast vs cautious)
    • Enables systematic interpretation of what parts of the model drive behavior
  6. 2:05 – 2:21

    Steering robot behavior by hyper-activating neuron clusters during inference

    Bear describes a control mechanism: during the forward pass, selectively “hyper-activate” concept clusters to push the robot toward a desired behavioral regime. This is presented as a direct, mechanistic way to constrain or steer robot actions.

    • Intervention happens at inference-time during the forward pass
    • Hyper-activating clusters steers behavior along interpretable dimensions
    • Aims at concrete controllability rather than purely post-hoc analysis
  7. 2:21 – 2:35

    Demo outcome: same command, different action-space behavior (high vs low activation)

    Using their website demo, Bear highlights that identical natural-language commands can yield different robot trajectories depending on whether a low- or high-neuron cluster is activated. The result is a measurable shift in the robot’s action-space magnitude or style.

    • Demonstration compares low-cluster vs high-cluster activation
    • Same language instruction produces different execution behaviors
    • Shows causal leverage: internal features can be used for direct control
  8. 2:35 – 3:28

    How interpretability connects to safety: constraining unexpected harmful behavior

    Ankit reframes the work through an AI safety lens: better understanding internal components could enable constraints that reduce unexpected bad outcomes under generic commands. Bear agrees and situates the work across the spectrum of safety concerns, from reliability to catastrophic risk.

    • Interpretability as a path to constraint and reliability
    • Safety motivations span ‘robustness’ through ‘catastrophic risk’
    • Embodiment increases the risk surface compared to purely digital models
  9. 3:28 – 4:31

    Why embodiment escalates stakes: from knowledge to physical capability

    Bear gives an example of how an RFM embedded in a humanoid could turn dangerous knowledge into real-world action—acquiring parts and executing plans physically. He argues this creates a broader set of risks than language-only systems and motivates expanding technical safety work.

    • Embodied systems can operationalize harmful instructions end-to-end
    • Physical access + autonomy increases potential impact
    • Physical AI Safety Institute aims to mobilize more technical research
  10. 4:31 – 5:36

    Building a field: the ‘Science of Physical AI Safety’ workshop and its agenda

    To close, Bear describes an upcoming workshop designed to bring AI safety and robot learning together into a shared discipline. He outlines the submission timeline, the organizing lineup, and three central questions about transferring ideas from AI safety and classical robotics into RFM safety and evaluation.

    • Workshop: ‘The Science of Physical AI Safety’ (November); submissions open Aug 12
    • Goal: unify AI safety + robot learning communities
    • Three questions: (1) what AI safety methods apply to RFMs, (2) what classical robotics contributes (kinematics/dynamics/control), (3) how to evaluate RFMs differently
    • Invites paper or short demo submissions

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.