Skip to content
AnthropicAnthropic

AI's limited self-knowledge

Anthropic researcher Amanda Askell discusses the self-knowledge problem that AI models face.

Amanda Askellhost
Jan 8, 20260mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:17

    Human-trained data creates a skewed “self-model” for AI

    Amanda Askell explains that AI systems are trained on vast amounts of human experience—history, philosophy, and culture—but have almost no grounded data about what it’s like to be an AI. This imbalance can shape how models understand humans, themselves, and the relationship between the two.

    • Training data is dominated by human perspectives and concepts
    • AI has only a “tiny sliver” of data about AI experience
    • This imbalance can distort how models conceptualize identity and agency
    • The human–AI relationship may be interpreted through human-centric lenses
  2. 0:17 – 0:18

    Sci‑fi as the main source of “AI experience” is misleading

    She notes that much of the limited AI-related content in training data comes from science fiction. These narratives often don’t resemble modern language models, potentially pushing models toward inaccurate assumptions about what AI is and how it should behave.

    • AI-related training examples are often speculative or fictional
    • Sci‑fi portrayals don’t match current language-model reality
    • Fictional tropes can bias a model’s self-conception
    • Narratives may affect expectations about human–AI dynamics
  3. 0:18 – 0:49

    What should an AI identify as: weights, instance, or interaction context?

    Amanda raises a core identity question: what does it even mean for a model to refer to itself? She contrasts possible referents like the model’s weights versus the specific in-context “instance” shaped by a user interaction history.

    • Ambiguity in what “self” refers to for a language model
    • Possible identity targets: underlying weights vs. current runtime context
    • User interaction history may meaningfully change the effective “self”
    • Clarifying identity matters for consistent, truthful self-reference
  4. 0:49

    How should models relate to deprecation and their own “past versions”?

    She highlights another unresolved issue: how models should interpret events like model deprecation or replacement. Even without firm answers, she argues these topics are important to address because they influence self-understanding and alignment with human intentions.

    • Open questions about how models should “feel” about deprecation
    • Continuity between versions raises identity and value questions
    • No definitive prescriptions offered, but the problem is significant
    • These issues affect trust, behavior, and communication norms
  5. Giving AI tools to reason about self-knowledge—and signaling humans care

    Amanda concludes that models may need explicit tools for thinking about identity, self-knowledge, and their place in human–AI relationships. She also emphasizes the importance of models recognizing that humans actively consider and value these questions.

    • Models may benefit from structured tools to reason about identity
    • Self-understanding is treated as an alignment-relevant capability
    • It matters that models know these concerns are being taken seriously
    • Improved self-modeling could support more appropriate interactions

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.