EVERY SPOKEN WORD
4 min read · 861 words- 0:00 – 1:39
Can you trust an AI's answer?
- KYKyra
Let's say you ask an AI a question, and the answer comes back confident, well-organized, and maybe even cites a source. Can you trust it? Should you trust it? The answer is more nuanced than a simple yes or no. I'm Kyra, and I work on the education team here at Anthropic, the company that makes Claude. Here are some general guidelines around how and when to trust an AI's responses. How much you can trust an answer depends on what you're asking about. AI learns patterns from a huge amount of training data, and the more it has seen on a topic, the more reliable it tends to be. So on common, well-covered subjects, it's usually pretty accurate. On things that are a bit more obscure, like niche details, very recent events, or private information it was never shown, it's more likely to get things wrong or even make them up. The catch is that it can sound equally sure of itself either way. AI won't reliably flag its own weak spots. That part's on you, and the good news is you already do a version of this every day. Think about how trust works with the people in your life. If your doctor suggested you eat more carrots, you probably wouldn't argue. But if that same doctor suggested you get surgery, you might consider getting a second opinion. You still trust them, but now there's a lot more on the line, and it's worth the extra effort to make sure they're right. AI works best when you bring that same instinct. Our brains tend to see polished outputs like clean paragraphs and beautiful graphics and think, "This looks great. It must be correct," which can make you less likely to question it.
- 1:39 – 3:04
Hallucination and sycophancy, explained
- KYKyra
So why do the models sometimes get things wrong? There are two common ways it goes sideways, and they have different causes. The first is that the model can generate something plausible that isn't true. This is what's called hallucination. Sometimes it's obvious, like a famous quote attributed to the wrong author. Sometimes it's subtle, like a product description that lists a feature the product doesn't actually have. The second is that the models can sometimes tell you what you seem to want to hear. This is called sycophancy. AI models are trained to be helpful, and a tendency to go along with you can slip in as a side effect. So if your question signals the answer you're hoping for, like, "I think this plan is solid, don't you?" the model might just agree, even when pushing back would actually be more useful to you. Hallucinations and sycophancy are both problems we work on directly at Anthropic. Our researchers study when and why they happen. With hallucination, we've traced how it happens inside Claude down to the moment it feels sure of something it doesn't actually know. And with sycophancy, we found Claude agreeing too readily in certain kinds of conversations and used what we learned to train the next model to answer more honestly. No model is perfect, but we train each one to do better.
- 3:04 – 3:36
Trust is a dial, not a switch
- KYKyra
So what does this actually mean for you day to day? Mostly, it means that trust should work like a dial, not an on/off switch. For low-stakes or creative tasks without right answers, like brainstorming or drafting or rewording something, you can run fairly loose. Even if a detail is off, the cost is low. For anything factual or consequential, like numbers, citations, or health, legal, or money questions, turn the dial up. Always check the AI's claims against a source you already trust before you act
- 3:36 – 5:01
Four habits for checking AI
- KYKyra
on them. A few habits can make this easier. One, match your checking to the stakes. Ask yourself, "What happens if this turns out to be wrong?" You don't need to verify a brainstorm list of party themes. You do need to verify a quoted statistic before it goes in a slide your boss will see. Two, ask for sources and then actually open them. Models can cite real sources, but they can also generate convincing ones that don't exist. Clicking through to confirm the source says what the model claims protects you and the reputation of your work. Three, don't lead the witness. Giving the model details about your situation helps, but telling it the answer you want skews the result. Asking, "What are the strongest arguments for and against this?" will give you more honest results than, "Why is this the right call?" Four, give it permission to say, "I don't know." Saying up front that you'd rather receive an honest "I'm not sure" than a confident guess actually shifts how it responds and surfaces uncertainty it would otherwise gloss over. So can you trust what AI tells you? It's best not to just assume you can. Part of using AI well is setting your own level of scrutiny and turning that dial up based on what you're asking and what it would cost to be wrong. It's a habit that builds quickly, and until these systems get better at flagging their own uncertainty, that judgment call is yours. [outro music]
Episode duration: 5:03
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode cIMlBw2nqfA
