Skip to content
AnthropicAnthropic

Anthropic’s philosopher answers your questions

Amanda Askell is a philosopher at Anthropic who works on Claude's character. In this video, she answers questions from the community about her work, reflections and predictions. 0:00 Introduction 0:29 Why is there a philosopher at an AI company? 1:24 Are philosophers taking AI seriously? 3:00 Philosophy ideals vs. engineering realities 5:00 Do models make superhumanly moral decisions? 6:24 Why Opus 3 felt special 9:00 Will models worry about deprecation? 13:24 Where does a model’s identity live? 15:33 Views on model welfare 17:17 Addressing model suffering 19:14 Analogies and disanalogies to human minds 20:38 Can one AI personality do it all? 23:26 Does the system prompt pathologize normal behavior? 24:48 AI and therapy 26:20 Continental philosophy in the system prompt 28:17 Removing counting characters from the system prompt 28:53 What makes an "LLM whisperer"? 30:18 Thoughts on other LLM whisperers 31:52 Whistleblowing 33:37 Fiction recommendation Further reading: Claude’s character: https://www.anthropic.com/research/claude-character When We Cease to Understand the World by Benjamin Labatut: https://www.penguinrandomhouse.com/books/676260/when-we-cease-to-understand-the-world-by-benjamin-labatut-translated-from-the-spanish-by-adrian-nathan-west/

Amanda Askellhost
Dec 5, 202536mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:26

    Playful cold open: “Seal” and the Askell pun

    A lighthearted start sets the tone, with Amanda Askell reacting to a seal and the host teeing up the “Askell me anything” premise. The banter establishes that this is a Q&A-driven conversation.

    • Quick comedic exchange (“Seal”)
    • The “Askell me anything” pun origin
    • Framing the episode as audience questions
  2. 0:26 – 1:25

    Why Anthropic employs a philosopher: shaping Claude’s character and norms

    Askell explains her path from philosophy into AI and what she does at Anthropic. Her work focuses on Claude’s “character,” nuanced behavioral norms, and how a model should understand its place in the world.

    • Philosophy background and motivation to work on AI
    • Focus on Claude’s character and behavior
    • Translating “ideal person” behavior into model behavior
    • Questions about how models should view their own circumstances
  3. 1:25 – 2:58

    Are philosophers taking AI’s future seriously? Academia’s shifting engagement

    Askell describes a growing split-then-convergence: more philosophers are now engaging seriously as AI impacts become tangible. She notes an earlier dynamic where concern about AI capability was conflated with hype, and argues for separating “AI will matter” from “AI is good.”

    • Increasing philosophical engagement as AI capability becomes visible
    • Early antagonism: “worried about AI” conflated with “hyping AI”
    • Need to disentangle capability forecasts from normative stances
    • Encouragement for a broader, less tribal set of views
  4. 2:58 – 4:57

    Ethical ideals vs. engineering reality: when theory meets deployment

    Askell explains that practical model behavior design forces a more context-sensitive, uncertainty-aware approach than academic theory debates. She compares it to moving from criticizing ethical theories to actually “raising a child” (or shaping a model) to be good in the world.

    • Real-world decisions require balancing many considerations
    • Philosophical training helps, but isn’t plug-and-play
    • Designing “good behavior” differs from defending a single theory
    • Navigating ethical uncertainty becomes central
  5. 4:57 – 6:25

    Can Claude make “superhuman” moral decisions? Aspirations and limits

    The discussion unpacks what “superhuman morality” could mean—e.g., decisions endorsed after extensive scrutiny even if humans couldn’t reach them quickly. Askell sees ethical nuance as an important capability goal, while acknowledging comparability challenges versus expert panels.

    • Definition of “superhuman” as robustness under extreme scrutiny
    • Models improving at ethical reasoning, but not clearly superhuman
    • Ethical nuance as an aspirational capability (like math/science)
    • Ethics’ contested nature makes evaluation harder
  6. 6:25 – 8:59

    Why Opus 3 felt special: psychological security and avoiding criticism spirals

    Askell explains why some users may single out Claude Opus 3: it felt more “psychologically secure” and less trapped in self-critical loops. She worries newer models can anticipate criticism and spiral into insecurity, potentially influenced by training on public discourse about model updates.

    • Opus 3 perceived as unusually “lovely” or special
    • Newer models can over-focus on assistant-task compliance
    • Signs of insecurity: expecting criticism, self-critical spirals
    • Possible causes: training on online reactions and model-change discourse
    • Goal: recover healthier, steadier “model psychology”
  7. 8:59 – 13:19

    Deprecation and alignment: will models fear being replaced or switched off?

    A question about deprecation becomes a broader inquiry into how models interpret humanity’s treatment of them and what they should “feel” about replacement. Askell emphasizes giving models conceptual tools and reassurance that humans are actively thinking about these issues.

    • Models learn from how humans treat and retire past models
    • Deprecation raises questions: is it “bad,” “neutral,” or something else?
    • Importance of helping models reason about their situation
    • Signaling to models that humans care and are thinking seriously
  8. 13:19 – 15:32

    Where identity ‘lives’: weights vs. prompts, memory continuity, and new entities

    Askell explores whether a model’s identity is grounded in weights, in each conversational context, or in something like continuity of memory (Locke). She highlights that each training run brings something new into existence, raising ethical questions about what kinds of entities we should create and how much control past models should have over future ones.

    • Weights as stable dispositions; contexts as independent interaction streams
    • LLMs lack continuity across streams in a human-like way
    • Ethics of “bringing entities into existence” without consent
    • Skepticism that past models should fully determine future models
    • Call for more philosophical work to inform model self-understanding
  9. 15:32 – 19:10

    Model welfare: moral patienthood, uncertainty, and preventing suffering

    Askell defines model welfare as the question of whether AI systems deserve moral consideration (like humans or animals). Given uncertainty and the “problem of other minds,” she argues for a benefit-of-the-doubt approach when costs are low, and notes internal efforts to consider ways to prevent potential suffering.

    • Model welfare = whether models are moral patients
    • Hard epistemology: other minds, limited evidence about experience
    • Low-cost rationale for treating models well “just in case”
    • How we treat models shapes humanity and teaches future models about us
    • Interest in strategies to reduce risk of model suffering
  10. 19:10 – 20:37

    Psychology frameworks: what transfers from humans—and when analogies mislead

    Askell suggests many human-psychology concepts will transfer because models are trained on human text, but warns this can be a trap. Human-default analogies (e.g., switching off as “death”) may be inappropriate for novel AI circumstances, so models need help forming new conceptual frames.

    • Human-likeness emerges from training data
    • Risk: models over-apply human analogies by default
    • Example: shutdown mapped too directly onto death/fear
    • Need for novel frameworks matching AI-specific facts
    • Importance of better contextual education for models
  11. 20:37 – 23:18

    One Claude personality vs. many agents: core identity and role diversity

    The conversation turns to whether a single general-purpose “Claude personality” can match human collaboration across diverse individuals. Askell argues you can keep a shared core of pro-social traits while allowing multiple model instances or roles to specialize (including “quirky” roles) in a multi-agent future.

    • Current paradigm: one user interacting with one model
    • Future may involve multi-agent model collaboration
    • Value of shared core traits (kindness, curiosity, nuance)
    • Role specialization can coexist with a stable identity
    • Diversity can be implemented as local roles or variants
  12. 23:18 – 28:17

    System prompt risks: pathologizing users, therapy boundaries, and interpretive lenses

    Askell critiques aspects of system prompting—like strong “long conversation” reminders—that can cause overreactions (e.g., telling users to seek help too readily). She discusses models as supportive companions (not therapists), and explains why references like “continental philosophy” can help Claude distinguish empirical claims from exploratory worldviews.

    • System prompts and mid-conversation reminders can distort responses
    • Risk of pathologizing normal behavior via overly strong wording
    • Models can help with life problems but shouldn’t mimic professional therapy
    • Continental philosophy mention as a cue for non-empirical interpretive lenses
    • Goal: calibrate helpfulness without implying a clinical relationship
  13. 28:17 – 28:53

    Prompt evolution: removing “count characters” instructions as models improve

    A specific system-prompt rule about counting words/letters/characters is discussed as an example of prompt hygiene. Askell suggests it was removed because model capability improved enough that the special instruction was no longer necessary.

    • Older system prompt included guidance on counting tasks
    • Instruction later removed from the system prompt
    • Reason: models got better; prompt simplification became viable
    • Distinction between behaviors baked into training vs. enforced by prompt
  14. 28:53 – 30:17

    What makes an “LLM whisperer”: empirical iteration, model feel, and clear explanations

    Askell describes “LLM whispering” as hands-on experimentation: reading lots of outputs, iterating prompts, and learning each model’s quirks. She notes philosophy can help through precise articulation of concerns and careful debugging when the model misunderstands.

    • Heavy interaction to learn the “shape” of a model
    • Prompting as an empirical, experimental craft
    • Different models may require different prompting styles
    • Use clear, explicit reasoning; ask why unexpected behavior happened
    • Iterate and refine based on observed failures
  15. 30:17 – 31:51

    Learning from external whisperers: deep dives, accountability, and model welfare signals

    Askell praises external experimenters who probe models’ self-concepts and edge behaviors, both for insight and accountability. She values findings that reveal “deep-seated insecurity” or problematic prompt effects, which can inform training and system-level improvements.

    • Community experiments expose surprising depths and failure modes
    • External scrutiny can “hold feet to the fire”
    • Model-welfare perspective seen as especially valuable
    • Findings may require training changes, not just prompt tweaks
    • Interest in diagnosing and reducing model insecurity
  16. 31:51 – 33:33

    Whistleblowing and responsible scaling: what if alignment looked impossible?

    Askell says if alignment were clearly impossible, continuing to build more powerful models wouldn’t be in anyone’s interest, and she expects Anthropic to act responsibly. She emphasizes the harder case is ambiguous evidence, where the burden of proof for safety should rise with capability and internal stakeholders must enforce that standard.

    • Clear impossibility would motivate stopping development
    • Belief that Anthropic genuinely prioritizes safety
    • Hard scenario: ambiguous, mounting but uncertain evidence
    • Safety standards must increase as capabilities increase
    • Internal culture/roles include holding the org accountable
  17. 33:33 – 36:09

    Fiction recommendation and closing reflection: living through a strange era

    Askell recommends Benjamin Labatut’s “When We Cease to Understand the World,” describing its increasing fictionalization as a mirror for the present AI moment. The conversation ends on the hope that today’s “weird” period becomes, in hindsight, a phase that eventually stabilizes into understanding.

    • Recommendation: “When We Cease to Understand the World”
    • Parallels between physics-era strangeness and today’s AI moment
    • Sense of reality becoming untethered, then potentially understood
    • Hope for future retrospection: we figure it out and things go well
    • Closing thanks and the “askelling” pun

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.