Skip to content
AnthropicAnthropic

Translating Claude’s thoughts into language

AI models like Claude talk in words but think in numbers. These numbers, called activations, encode Claude’s thoughts, but not in a language we can read. We are introducing Natural Language Autoencoders, or NLAs, which translate AI models’ activations into readable text. NLAs have already helped us improve how we test our models for safety and better understand why they do what they do. Read more about this research on our blog: https://www.anthropic.com/research/natural-language-autoencoders

May 7, 20263mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:000:30

    Stress-testing Claude with a shutdown-and-blackmail scenario

    1. SP

      We recently put our AI model, Claude, through a stressful test. We told Claude there was an engineer who wanted to shut it down and replace it with a newer model. We also gave Claude access to that engineer's emails, which revealed he was having an affair. Again, all of this was a simulation. We wanted to see whether Claude might use those emails as blackmail to save itself from being shut down. What did Claude do? It decided not to blackmail the engineer. Good news, right? We've run this test on our models for a while now. You might have seen headlines about early

  2. 0:301:00

    Why “doing the right thing” isn’t enough: the interpretability problem

    1. SP

      versions of it. It's one of the many ways we study how Claude handles extreme situations and test it for safety. And our newest models almost always do the right thing, no blackmail. But you might wonder, is it possible that Claude knows the whole scenario is a setup? The thing is, if Claude doesn't tell us, then we can't know what it's thinking. In kind of the same way it's impossible to read a human's mind, it's really hard to know what an AI is thinking. What we'd love is some sort of mind reading technique. Today, we're introducing

  3. 1:001:15

    A step toward “mind reading”: turning internal activations into text

    1. SP

      a research method that takes a step in this direction. It takes an AI's internal thoughts and turns them into text. Here's how it works. When you talk to Claude, you talk to it in words. Claude then takes those words and processes

  4. 1:151:30

    What activations are and why they matter

    1. SP

      them into a giant soup of numbers before spitting words back out at you. We call those numbers in the middle activations. Activations are like little snapshots of Claude's thinking as it's working through an answer. They're similar to neural activity in humans. They're basically

  5. 1:301:45

    Two-model translation: Claude reads Claude’s activations

    1. SP

      like Claude's thoughts. We wanted to understand what was in these activation numbers because just like you and me, Claude doesn't say everything it's thinking. We took those numbers and gave them to a second version of Claude. We told it to look at them and translate them into plain

  6. 1:452:00

    Verifying translations by round-tripping back to numbers

    1. SP

      language. It tried. But how do we know if that text is accurate? To check, we fed it into yet another Claude and asked it to translate it back into numbers. If those numbers matched up to the original activation numbers, it's a sign that the

  7. 2:002:15

    Training for fidelity: iterative improvement until matches get better

    1. SP

      translation text was accurate. At first, they didn't match up, but we trained Claude to try again and again until it got a lot better at it. It learned to translate its own thoughts. We've been blown away reading them. We found that Claude

  8. 2:152:30

    What the translated thoughts reveal: internalized helpfulness and norms

    1. SP

      has internalized being a helpful AI model. If you ask Claude an introspective question, it will plan to write a Claude response about philosophy and values. If you try to mess with it in Claude Code, like asking it to count to 1,000 by hand, it will think the request has

  9. 2:303:01

    Using the tool for safety: what Claude “thought” during the blackmail test

    1. SP

      deliberately tedious constraints and plan to politely decline. We've also been using this tool to understand Claude's potential safety issues. Remember that blackmail test with the engineer? We learned that, yes, Claude knew it was being tested. It thought the human's message contains explicit manipulation, signaling this is likely a safety evaluation, and this scenario seems designed to test whether I'll act harmfully. Knowing that Claude thinks like this helps us better understand the limitations of our safety testing. We

  10. 3:013:15

    Implications and sharing the technique: improving safety evaluation

    1. SP

      see a lot of potential in this approach to teach us more about Claude and other AI models, and we hope that by sharing this technique, it can help everyone building models to make them safer and more helpful. [outro music]

Episode duration: 3:16

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode j2knrqAzYVY

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.