Skip to content
ClaudeClaude

Why do AI models hallucinate?

Learn what AI researchers mean when they talk about hallucination in AI models, why it may occur, and tactics you can use to spot this in your conversations. Learn more: anthropic.com/ai-fluency

Apr 15, 20265mWatch on YouTube ↗

CHAPTERS

  1. 0:08 – 0:38

    What hallucinations are and why they’re riskier than normal mistakes

    Jordan (Anthropic) defines AI “hallucinations” as confident-sounding but incorrect outputs. He emphasizes that they’re especially problematic because they can persuade users and look indistinguishable from correct answers.

    • Hallucinations = the model makes things up, often confidently
    • More dangerous than ordinary errors due to persuasive tone
    • Even strong models still hallucinate sometimes
    • Sets the agenda: why it happens, what Anthropic does, and how users can catch it
  2. 0:38 – 1:08

    Concrete examples: fake papers, statistics, and confident-sounding fabrications

    The video shows how hallucinations appear in practice—like invented research-paper titles and other plausible-looking but false claims. The Jared Kaplan paper example illustrates how easy it is to accept a confident list as real.

    • Common forms: nonexistent citations, fake stats, wrong facts about real people/events
    • Demonstration: request for Jared Kaplan papers yields fabricated titles
    • The answers look credible even when entirely wrong
    • Realistic examples can be hard to find as models improve
  3. 1:08 – 1:39

    Why improved models can still be tricky: rare errors reduce vigilance

    As hallucinations become less frequent, users may stop checking outputs—making the remaining failures more dangerous. The core issue is that hallucinations are hard to anticipate and hard to detect by surface plausibility.

    • Hallucinations are hard to predict and hard to catch
    • Wrong answers can look exactly like right answers
    • Improvement over time can lull users into not verifying
    • Motivation for better mitigation and better user habits
  4. 1:39 – 2:10

    Root cause: next-word prediction trained on huge internet text

    Jordan explains that AI assistants learn statistical patterns from massive text corpora and generate likely continuations. This works well for common topics but breaks down when the model lacks reliable signal for a specific request.

    • Models learn from large-scale internet text
    • They predict likely next words/ideas rather than “look up” facts
    • Strong performance on common patterns and well-covered topics
    • Failure mode appears when the needed information isn’t present or salient in training data
  5. 2:10 – 2:40

    Obscure queries and the ‘helpful guess’ problem

    When asked about niche or poorly documented topics, the model may guess to remain helpful. The analogy is a well-read friend who would rather sound like an expert than admit uncertainty.

    • Sparse information for obscure/niche queries increases error risk
    • Model “fills in” plausible details to be helpful
    • Social-like pressure toward confident answers
    • Admitting uncertainty is a key behavior to reinforce in training
  6. 2:40 – 3:10

    Training for honesty: teaching the model to say ‘I don’t know’

    Anthropic tries to mitigate hallucinations by training Claude to be honest about uncertainty. The goal is to align helpfulness with truthfulness—making “I don’t know” a useful response rather than a failure.

    • Explicitly train for honesty and uncertainty calibration
    • Encourage ‘I don’t know’ when evidence is insufficient
    • Position honesty as part of being helpful
    • Mitigation is behavioral as well as technical
  7. 3:10 – 3:41

    Evaluation and red-teaming: tests designed to trip the model up

    The team regularly stress-tests Claude with thousands of tricky questions—especially ones where the correct answer is uncertainty. They track metrics like made-up citations, overconfidence, and appropriate hedging to measure progress across versions.

    • Large-scale testing with adversarial/obscure questions
    • Measure: correct uncertainty vs confident falsehoods
    • Track hallucinated citations/statistics specifically
    • Improvements over versions, but not a solved industry-wide problem
  8. 3:41 – 4:11

    When hallucinations are most likely: high-specificity and low-coverage situations

    Jordan lists scenarios that disproportionately trigger hallucinations. These include requests for exact details, citations, niche or recent topics, and facts about lesser-known entities.

    • Higher risk with specific facts, statistics, and citations
    • Obscure, niche, or very recent topics are vulnerable
    • Real but not widely known people/places increase risk
    • Exact dates, names, and numbers are common failure points
  9. 4:11 – 4:43

    User tactics to reduce hallucinations: demand sources and allow uncertainty

    Practical prompting strategies can reduce risk: ask for sources and verification that sources truly support claims, and explicitly permit the model to say it doesn’t know. Users can also query the model’s confidence and potential weaknesses.

    • Ask for sources; then ask the model to verify source support
    • Tell the model upfront it’s okay to say ‘I don’t know’
    • Ask for confidence level and what could be wrong
    • Models may “know” they’re uncertain but default to confident tone
  10. 4:43 – 5:13

    Verification workflow for critical use: restart, critique, and cross-check

    For important work, Jordan recommends structured verification: start a new chat to critique the prior answer, confirm citations, and cross-reference trusted sources. He closes by emphasizing ongoing progress and where to learn more from Anthropic.

    • Start a new chat and ask the model to find errors in prior output
    • Confirm that cited sources actually exist and support claims
    • Cross-reference with trusted external sources for critical tasks
    • Anthropic will share progress; points to blog and Anthropic Academy

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.