Skip to content
YC Root AccessYC Root Access

Interpretability and Safety for Robot Foundation Models

At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Bear Häon about applying interpretability and AI safety techniques to vision-language-action models. The work examines activations inside a vision-language-action model, groups neurons associated with concepts such as speed or caution, and then steers the robot's behavior by amplifying those groups. This provides a way to better understand and control how robots translate language into physical actions. The broader goal is to develop safety methods for physical AI, where failures can have direct consequences in the real world. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Ankit GuptahostBear Häonguest
Aug 6, 20265mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Interpreting robot foundation models to control behavior and improve safety

  1. The discussion presents one of the first interpretability papers for Robot Foundation Models (RFMs), adapting AI-safety interpretability techniques to embodied, physically acting systems.
  2. The core method inspects transformer feedforward networks (FFNs) in a vision-language-action model, mapping highly activating tokens to build concept-like “dictionaries” for internal components.
  3. By clustering FFN activations into human-interpretable concepts (e.g., cautious vs. fast) and “hyper-activating” clusters at inference time, the robot’s action style can be directly steered under the same language command.
  4. The conversation frames physical embodiment as an added risk surface: an RFM can translate knowledge into real-world action, amplifying concerns from reliability failures to catastrophic misuse.
  5. Bear introduces the Physical AI Safety Institute and a workshop aimed at unifying AI safety and robotics communities around methods, classical control insights, and evaluation standards for physical AI safety.

IDEAS WORTH REMEMBERING

5 ideas

FFN activations provide a practical foothold for interpreting RFMs.

By projecting FFN activity onto the model’s embedding/token space, you can identify which tokens/concepts each FFN responds to most strongly, creating a dictionary-like view of internal representations.

Interpretability can be used not just to understand robots, but to control them.

After clustering FFN representations into concepts (e.g., cautious/fast), selectively amplifying those clusters during the forward pass can steer the robot’s motion/trajectory style without changing the user’s natural-language command.

Behavior steering changes the robot’s action space under identical instructions.

The same command can yield “higher” or “lower” action intensity depending on which neuron clusters are activated, offering a lever for safety constraints like slower, more careful execution.

Physical embodiment raises the stakes compared to purely digital models.

A general reasoning model acting through a robot can acquire materials and execute real-world plans, turning “knowledge of how” into “ability to do,” which expands the threat model beyond typical LMs.

Physical AI safety spans reliability, alignment, and catastrophic-risk perspectives.

The talk explicitly acknowledges a spectrum—from making systems more dependable on tasks to addressing long-horizon misalignment and severe misuse scenarios—and argues embodiment can amplify all of them.

WORDS WORTH SAVING

5 quotes

We wrote the first ever interpretability paper for Robot Foundation Models.

Bear Häon

But could we take those techniques and apply them to create concrete safety outcomes for systems that both think and physically act?

Bear Häon

You actually get this representation that's almost like a dictionary of, for every FFN, what were the tokens across all the FFNs that were most activated by that FFN, uh, layer?

Bear Häon

An RFM embedded in a humanoid can, can, uh, know how to make it. It can go buy the parts by walking into Home Depot in a world that's, you know, very normalized that to have, you know, robot embodiments, and then actually go make it and place it in a high-density location.

Bear Häon

The workshop is called The Science of Physical AI Safety. The goal is to bring together the AI safety and robot learning communities to establish this shared field.

Bear Häon

Robot foundation model (RFM) interpretabilityVision-language-action (VLA) modelsTransformer FFNs as analyzable control pointsToken/embedding-based activation “dictionaries”Concept clustering and behavior steeringEmbodiment-driven risk amplificationPhysical AI Safety Institute and workshop agenda

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.