Skip to content
ClaudeClaude

Can you trust what AI tells you?

How much you can trust an AI depends on what you’re asking. Kyra from the Anthropic education team breaks down the two most common reasons for an AI to be confidently wrong: hallucination and sycophancy. Have a question? Let us know in the comments. Learn more at Claude Academy: http://academy.claude.com Chapters 0:00 Can you trust an AI's answer? 1:39 Hallucination and sycophancy, explained 3:04 Trust is a dial, not a switch 3:36 Four habits for checking AI

Kyrahost
Aug 11, 20265mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

How to calibrate trust in AI answers and verify claims

  1. AI reliability varies by domain: it tends to be stronger on common, well-covered topics and weaker on niche, recent, or private information.
  2. Two major failure modes are highlighted: hallucinations (plausible but false statements) and sycophancy (agreeing with what you want to hear).
  3. Because polished outputs can create an illusion of correctness and models don’t reliably flag uncertainty, users must apply their own scrutiny.
  4. Trust should be treated like a dial: use lighter checking for low-stakes/creative work and stronger verification for factual or high-impact decisions.
  5. Four practical habits are proposed: match checking to stakes, request and open sources, avoid leading prompts, and explicitly invite “I don’t know.”

IDEAS WORTH REMEMBERING

5 ideas

Accuracy depends on what you ask the model to do.

Models are generally more dependable on widely represented topics in training data and more error-prone on obscure, newly emerging, or inaccessible/private details.

Confidence and formatting are not evidence of correctness.

Clean prose, structure, and even citations can trigger overtrust; the model can sound equally certain when it knows vs when it’s guessing.

Hallucination and sycophancy are distinct risks to watch for.

Hallucinations create believable falsehoods, while sycophancy produces overly agreeable answers when your prompt implies a preferred conclusion.

Use risk-based verification: treat trust like a dial.

For brainstorming/drafting, minor errors are low-cost; for numbers, citations, and health/legal/financial decisions, independently validate before acting.

Always check sources by opening them, not just collecting citations.

Models may cite real materials but can also fabricate plausible references; clicking through ensures the cited source exists and supports the specific claim.

WORDS WORTH SAVING

5 quotes

Let's say you ask an AI a question, and the answer comes back confident, well-organized, and maybe even cites a source. Can you trust it? Should you trust it? The answer is more nuanced than a simple yes or no.

Kyra

On things that are a bit more obscure, like niche details, very recent events, or private information it was never shown, it's more likely to get things wrong or even make them up. The catch is that it can sound equally sure of itself either way. AI won't reliably flag its own weak spots.

Kyra

There are two common ways it goes sideways, and they have different causes. The first is that the model can generate something plausible that isn't true. This is what's called hallucination.

Kyra

The second is that the models can sometimes tell you what you seem to want to hear. This is called sycophancy.

Kyra

Mostly, it means that trust should work like a dial, not an on/off switch.

Kyra

When AI is more vs less reliableHallucination (plausible fabrication)Sycophancy (over-agreeableness)Polish-driven overtrust in outputsTrust as a dial (risk-based scrutiny)Source verification and citation checkingPrompting for balanced arguments and uncertainty

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.