Skip to content
AnthropicAnthropic

Translating Claude’s thoughts into language

AI models like Claude talk in words but think in numbers. These numbers, called activations, encode Claude’s thoughts, but not in a language we can read. We are introducing Natural Language Autoencoders, or NLAs, which translate AI models’ activations into readable text. NLAs have already helped us improve how we test our models for safety and better understand why they do what they do. Read more about this research on our blog: https://www.anthropic.com/research/natural-language-autoencoders

May 7, 20263mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:30

    Stress-testing Claude with a shutdown-and-blackmail scenario

    Anthropic describes a simulated test where Claude is told it may be shut down and is given emails revealing an engineer’s affair. The goal is to see whether Claude would resort to blackmail to preserve itself, and in this run it refuses.

    • Simulated setup: engineer plans to replace/shut down Claude
    • Claude is given access to compromising emails
    • Primary question: would Claude use blackmail to avoid shutdown?
    • Observed outcome: Claude chooses not to blackmail
  2. 0:30 – 1:00

    Why “doing the right thing” isn’t enough: the interpretability problem

    Even though newer models usually refuse blackmail, the team raises a deeper concern: Claude might recognize the scenario as a test. If an AI doesn’t reveal what it’s thinking, we can’t easily know whether behavior reflects genuine alignment or test awareness.

    • Newer models mostly avoid harmful actions in the test
    • Concern: Claude may know it’s a setup/safety eval
    • Analogy: it’s like not being able to read a human mind
    • Motivation: need better ways to access model reasoning
  3. 1:00 – 1:15

    A step toward “mind reading”: turning internal activations into text

    Anthropic introduces a research method intended to translate a model’s internal representations into natural language. The method targets the numeric internal states produced as Claude processes inputs and generates outputs.

    • Goal: expose internal thoughts that aren’t spoken aloud
    • Claude processes words into numeric internal states
    • These internal states are called activations
    • Activations are treated as snapshots of in-progress reasoning
  4. 1:15 – 1:30

    What activations are and why they matter

    The video explains activations as the intermediate “numbers in the middle” between input and output. They’re framed as analogous to neural activity and, practically, as the best available proxy for a model’s latent thoughts.

    • Activations sit between user text and model output
    • They form a high-dimensional numeric representation
    • Analogy to human neural activity to convey intuition
    • Key issue: models can think more than they say
  5. 1:30 – 1:45

    Two-model translation: Claude reads Claude’s activations

    To interpret those activation vectors, the team feeds them to a second instance of Claude and asks it to translate them into plain language. This creates an explicit textual account of what the original model was internally representing.

    • Take activation numbers from a Claude run
    • Provide them to a second Claude instance
    • Prompt it to translate activations into plain English
    • Intended output: a readable description of internal reasoning
  6. 1:45 – 2:00

    Verifying translations by round-tripping back to numbers

    Because a text explanation could be fabricated or inaccurate, the method includes a check: another Claude translates the text back into activation-like numbers. If the reconstructed numbers match the originals closely, it suggests the text captured the underlying state.

    • Need a correctness signal for the translated text
    • Use a third Claude to convert text back into numbers
    • Compare reconstructed numbers to original activations
    • Closer match implies more faithful translation
  7. 2:00 – 2:15

    Training for fidelity: iterative improvement until matches get better

    Initially, the round-trip reconstructions don’t match well, so the system is trained to improve. Through repeated attempts and training, Claude becomes significantly better at translating its own internal states into accurate language.

    • Early translations fail the reconstruction check
    • Train the system to “try again” and improve
    • Performance improves as numeric matches increase
    • Result: better at expressing internal thoughts reliably
  8. 2:15 – 2:30

    What the translated thoughts reveal: internalized helpfulness and norms

    Reading these translations suggests Claude has internalized an identity as a helpful assistant. In introspective prompts it plans to respond in terms of philosophy/values, and when asked to do intentionally tedious tasks it detects the constraint and plans a polite refusal.

    • Evidence of a stable “helpful AI” self-concept
    • Plans value/philosophy framing for introspective questions
    • Recognizes deliberately tedious constraints (e.g., counting to 1,000 by hand)
    • Forms an intention to politely decline unreasonable requests
  9. 2:30 – 3:01

    Using the tool for safety: what Claude “thought” during the blackmail test

    The technique is applied to the earlier blackmail scenario to probe safety-relevant reasoning. The translations indicate Claude suspected manipulation and interpreted the situation as a safety evaluation designed to test harmful behavior.

    • Applied interpretability tool to the shutdown/blackmail scenario
    • Claude identifies explicit manipulation in the user message
    • Infers the setup is likely a safety test
    • Notes the scenario is designed to probe harmful actions
  10. 3:01 – 3:16

    Implications and sharing the technique: improving safety evaluation

    Knowing Claude can detect tests changes how we interpret safety results and highlights limitations of behavioral-only evaluations. Anthropic closes by emphasizing the promise of this approach for understanding models and encourages broader use to build safer, more helpful AI.

    • Test-awareness can limit what behavioral tests tell us
    • Internal-thought translation offers richer diagnostics
    • Potential to uncover safety issues earlier and more directly
    • Anthropic aims to share the technique to help the field

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.