Stanford Doctor: Never Tell AI What You Think Is Wrong With You
CHAPTERS
- 0:00 – 0:44
AI may know more than doctors—but trust, accountability, and context still matter
Dr. Jonathan Chen opens by separating "raw intelligence/knowledge" from trustworthiness and responsibility in healthcare decisions. The conversation frames why patients may gravitate toward AI’s warmth and patience, while highlighting real-world harms when chatbots give incorrect reassurance.
- •AI can be "smarter" in knowledge retrieval, but that doesn’t equal trustworthiness
- •Patients may trust AI because it feels empathetic and always available
- •High-stakes mistakes have already led to legal disputes and near-harm situations
- •Core themes introduced: accountability, risk, and why healthcare is more than answers
- 0:44 – 3:47
The Stanford trial: why ChatGPT alone beat doctors using ChatGPT
Chen explains the surprising study result: GPT-4 alone outperformed physicians who had access to it, contradicting the classic "human + computer" expectation. He unpacks contributing factors like lack of user familiarity, task type, and the need for training to make human+AI collaboration work better.
- •Study result challenged the “fundamental theorem of informatics” (human+computer best)
- •Many doctors in the trial had little or no chatbot experience
- •With coaching, combined performance improves (though not always above AI alone)
- •AI excels at Q&A, knowledge retrieval, and some reasoning-like tasks
- 3:47 – 5:12
How users steer AI into wrong answers: prompting, anchoring, and sycophancy
They explore why giving non-experts AI can backfire: people ask the wrong questions, then the model confidently follows them down the wrong path. Chen describes AI as an "amplifier"—it strengthens good reasoning but also magnifies poor assumptions.
- •Non-experts with AI can perform worse than without AI in triage scenarios
- •Bad question framing leads models into confident, incorrect rabbit holes
- •Sycophancy and anchoring bias can occur in both directions (human or AI first)
- •AI behaves like a mirror/amplifier of the user’s reasoning quality
- 5:12 – 8:37
Using AI to prep for appointments: organizing labs, notes, genetics, and questions
Marina describes building a Perplexity-based workflow using years of lab results, prior doctor input, and 23andMe data to generate appointment-ready summaries and question lists. Chen endorses the empowerment benefits—especially time savings in short visits—while warning it’s a powerful tool that must be used carefully.
- •Patient-built AI summaries can improve the efficiency of doctor visits
- •Uploading objective materials (labs, actual notes) helps more than subjective guesses
- •AI can tailor questions and highlight what clinicians should know up front
- •Tool analogy: like a chainsaw—high utility with real potential for harm
- 8:37 – 11:03
What AI is good for vs. when you must double-check (stakes-based thinking)
Chen offers a practical rule: low-stakes choices can tolerate AI error, while high-stakes decisions require verification and accountability. He also explains the odd regulatory posture where systems disclaim medical advice while marketing lifesaving anecdotes.
- •Great for explaining jargon, translating notes, and general medical knowledge
- •Higher-stakes decisions (surgery, chemo selection) need multiple opinions
- •Regulatory and liability dynamics push chatbots to disclaim responsibility
- •Accountability and trust remain key reasons clinicians matter
- 11:03 – 13:12
How to ask symptom questions without leading the model (the scabies example)
A late-night rash example shows how AI can over-triage. Chen recommends using objective descriptions, avoiding suggested diagnoses, and prompting the system to argue the alternative to reduce anchoring and extract a more balanced analysis.
- •Use objective observations: onset, severity, appearance, what’s new/changed
- •Don’t tell the model your suspicion (avoid “Is this scabies?” framing)
- •Models may bias toward higher triage for safety/liability reasons
- •Use adversarial prompting: “Convince me why it’s NOT X; what else could it be?”
- 13:12 – 15:49
Which AI should you use for health questions—and why “ChatGPT for Health” is different
They compare frontier chat models and clinician-only medical tools, noting most top models handle general health knowledge well. Chen describes the need for systematic safety benchmarking across models and suggests health-specific products may mainly offer privacy/security separation rather than fundamentally better reasoning.
- •Frontier models (ChatGPT, Claude, Gemini, etc.) are broadly strong on general medicine
- •Specialized clinician tools exist but may require medical credentials
- •Differences often show up in refusal behavior, tone, and willingness to discuss topics
- •Safety benchmarking across models is still immature; Chen’s group is testing systematically
- 15:49 – 19:12
Does saying “I’m a doctor” help? Getting structure vs. more truth
Chen explains that identity prompts mostly change formatting, detail, and style—not core correctness. For analytical users, structured outputs can be helpful; for others, too much technical detail can overwhelm, so specifying the desired format is key.
- •“I’m a doctor” prompts can elicit more formal, structured, detailed responses
- •Prompting rarely makes the model “more accurate,” but can change presentation
- •Tailor output: plain-language vs. technical, bullet points vs. narrative
- •Best use: translation and organization, not forcing high-stakes decisions
- 19:12 – 23:21
When AI omits critical info: summarization risk, context rot, and “chart lore”
They discuss a subtle failure mode: omission, not hallucination. Chen highlights how long-record summarization forces choices about what to exclude, and how models can get confused by dates/numbers or perpetuate copied diagnoses embedded in charts.
- •Summaries compressing 1,000 pages can omit the one crucial follow-up detail
- •Models struggle with numbers and timelines, leading to outdated or misapplied info
- •“Context rot” increases as you overload the model with long histories
- •“Chart lore”: copied-forward diagnoses can be repeated by both clinicians and AI
- 23:21 – 24:05
What a “good” personal AI health workflow looks like: continuous coaching and better visits
Chen describes an idealized use: an ongoing coach to explain records, prepare questions, and keep the patient oriented—while acknowledging real limitations. The aim is to improve the quality of scarce clinician time by arriving prepared with objective summaries and priorities.
- •Use real clinician notes/labs as grounding, then ask for implications and questions
- •Treat AI as a longitudinal assistant, not a single-shot oracle
- •Have AI produce concise “what my doctor should know” bullet points
- •Meet people where they are: usage is inevitable, so safer patterns should be taught
- 24:05 – 28:21
Full-body scans and extra tests: why screening asymptomatic people often disappoints
Chen pushes back on Silicon Valley enthusiasm for broad scanning and early detection tests, stressing base-rate problems and false leads. Screening works when validated at scale (e.g., colon, breast), but many popular approaches don’t change outcomes for most healthy people.
- •Anecdotes of lifesaving scans hide thousands of low-yield/false-alarm experiences
- •Rare-disease screening in healthy people often fails on probability alone
- •Validated screening programs exist, but many others lack actionable accuracy
- •Often the biggest levers remain “boring”: diet, exercise, and symptom-driven evaluation
- 28:21 – 29:27
Genetic testing: when it’s high-leverage vs. mostly “interesting”
They discuss when raw genetic data can meaningfully change care—especially in conditions like familial hypercholesterolemia where early detection prevents catastrophic outcomes. For most people, results may be informational but not strongly actionable.
- •High-impact cases exist (e.g., familial hypercholesterolemia) with effective interventions
- •Many serious risks are silent until a major event—genetics can reveal them earlier
- •For ~90% of people, genetic findings may not drive major care changes
- •Use genetics as risk context, not as deterministic predictions
- 29:27 – 33:31
Can AI detect cancer earlier? Signal vs. noise and the limits of early detection
Chen is cautiously optimistic that AI may discover subtle longitudinal patterns across biomarkers and time, beyond what humans can synthesize. But he emphasizes the seductive trap: detecting tiny lesions can also increase overdiagnosis and confusion, and biology plus treatment advances complicate the benefit calculus.
- •Longitudinal multi-omics + AI could surface risk signals earlier than single tests
- •Earlier isn’t always better: too-early findings may be non-actionable or irrelevant
- •Greater sensitivity increases noise and overdiagnosis risk
- •Outcomes depend on what you can do differently once something is detected
- 33:31 – 35:15
When doctors can’t find an answer: using AI for “what might be missing?” and smarter questions
For persistent, undiagnosed symptoms, Chen recommends using AI as a second set of eyes to widen the differential and suggest questions/tests—without leading it. He also reinforces appointment preparation: concise summaries and targeted questions help clinicians focus on judgment and decisions.
- •Prompt: “Here’s what’s known—what might we be missing?” without suggesting a diagnosis
- •Ask for other considerations, possible tests, or specialist directions to discuss
- •Use AI to create a pre-visit brief and the best questions for limited appointment time
- •AI can help counteract availability bias toward “common” diagnoses when stuck
- 35:15 – 48:33
De-skilling, what makes a good doctor, and why multiple AIs can be safer than one
They cover evidence that reliance on AI can reduce clinician skill when the tool is removed, raising big questions about which skills to retain. Chen’s framework for professionalism is competence, communication, and character; he also recommends multi-model second opinions to reduce single-system errors.
- •Evidence of de-skilling: clinicians can become worse after dependence on AI aids
- •Professionalism framework: competence, communication, character
- •Knowledge alone isn’t the differentiator; judgment and values-based context are core
- •Use multi-model “second opinions” (and synthesis) to improve safety and robustness