Skip to content
Silicon Valley GirlSilicon Valley Girl

Stanford Doctor: Never Tell AI What You Think Is Wrong With You

Bring your team’s research, context, and AI workflow together on one canvas: Try Miro – https://miro.pxf.io/2RLQv0 Dr. Jonathan Chen is a Stanford physician and computer scientist who teaches doctors how to use AI. His team found that GPT-4 alone outperformed doctors using it on diagnostic reasoning tasks. But being smarter doesn’t automatically make AI trustworthy. In this episode, we break down how to ask AI about your health, use your medical records to prepare for appointments, and avoid steering a chatbot toward the answer you want to hear. Jonathan also explains why AI can leave out critical information, why more health screening isn’t always better, and what doctors still need to do when a computer knows more than they do. Timestamps: 00:00 — Is AI smarter than your doctor? 01:38 — How AI outperformed doctors using AI 03:33 — How your questions can steer AI toward the wrong answer 05:12 — Using your medical records to prepare for appointments 08:27 — What AI can help with and when to double-check 11:28 — How to ask about symptoms without leading the AI 12:57 — Which AI should you use for health questions? 15:45 — Does telling AI “I’m a doctor” help? 20:48 — Can you give AI too much medical history? 24:05 — Are full-body scans and extra tests worth it? 28:21 — When genetic testing can make a difference 29:28 — Can AI help detect cancer earlier? 33:31 — What to ask AI when doctors can’t find an answer 35:15 — Can relying on AI make you lose your skills? 36:51 — What makes a good doctor when AI knows more? 39:39 — AI doctors and automated prescriptions 41:39 — Will AI cure all diseases in 10 years? 42:24 — What better access to care could change 44:52 — Who should make the final decision? 47:10 — Always use multiple AI models for a second opinion Links: 📩 Follow my Newsletter:https://siliconvalleygirl.beehiiv.com/subscribe?utm_source=youtube&utm_medium=video&utm_campaign=futureproof-sub&utm_content=JonathanChen 🔗 My Instagram: https://www.instagram.com/siliconvalleygirl/ 📌 My Companies & Products: https://partnerships.marinamogilko.co

Dr. Jonathan ChenguestMarina Mogilkohost
Sep 25, 202648mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:44

    AI may know more than doctors—but trust, accountability, and context still matter

    Dr. Jonathan Chen opens by separating "raw intelligence/knowledge" from trustworthiness and responsibility in healthcare decisions. The conversation frames why patients may gravitate toward AI’s warmth and patience, while highlighting real-world harms when chatbots give incorrect reassurance.

    • •AI can be "smarter" in knowledge retrieval, but that doesn’t equal trustworthiness
    • •Patients may trust AI because it feels empathetic and always available
    • •High-stakes mistakes have already led to legal disputes and near-harm situations
    • •Core themes introduced: accountability, risk, and why healthcare is more than answers
  2. 0:44 – 3:47

    The Stanford trial: why ChatGPT alone beat doctors using ChatGPT

    Chen explains the surprising study result: GPT-4 alone outperformed physicians who had access to it, contradicting the classic "human + computer" expectation. He unpacks contributing factors like lack of user familiarity, task type, and the need for training to make human+AI collaboration work better.

    • •Study result challenged the “fundamental theorem of informatics” (human+computer best)
    • •Many doctors in the trial had little or no chatbot experience
    • •With coaching, combined performance improves (though not always above AI alone)
    • •AI excels at Q&A, knowledge retrieval, and some reasoning-like tasks
  3. 3:47 – 5:12

    How users steer AI into wrong answers: prompting, anchoring, and sycophancy

    They explore why giving non-experts AI can backfire: people ask the wrong questions, then the model confidently follows them down the wrong path. Chen describes AI as an "amplifier"—it strengthens good reasoning but also magnifies poor assumptions.

    • •Non-experts with AI can perform worse than without AI in triage scenarios
    • •Bad question framing leads models into confident, incorrect rabbit holes
    • •Sycophancy and anchoring bias can occur in both directions (human or AI first)
    • •AI behaves like a mirror/amplifier of the user’s reasoning quality
  4. 5:12 – 8:37

    Using AI to prep for appointments: organizing labs, notes, genetics, and questions

    Marina describes building a Perplexity-based workflow using years of lab results, prior doctor input, and 23andMe data to generate appointment-ready summaries and question lists. Chen endorses the empowerment benefits—especially time savings in short visits—while warning it’s a powerful tool that must be used carefully.

    • •Patient-built AI summaries can improve the efficiency of doctor visits
    • •Uploading objective materials (labs, actual notes) helps more than subjective guesses
    • •AI can tailor questions and highlight what clinicians should know up front
    • •Tool analogy: like a chainsaw—high utility with real potential for harm
  5. 8:37 – 11:03

    What AI is good for vs. when you must double-check (stakes-based thinking)

    Chen offers a practical rule: low-stakes choices can tolerate AI error, while high-stakes decisions require verification and accountability. He also explains the odd regulatory posture where systems disclaim medical advice while marketing lifesaving anecdotes.

    • •Great for explaining jargon, translating notes, and general medical knowledge
    • •Higher-stakes decisions (surgery, chemo selection) need multiple opinions
    • •Regulatory and liability dynamics push chatbots to disclaim responsibility
    • •Accountability and trust remain key reasons clinicians matter
  6. 11:03 – 13:12

    How to ask symptom questions without leading the model (the scabies example)

    A late-night rash example shows how AI can over-triage. Chen recommends using objective descriptions, avoiding suggested diagnoses, and prompting the system to argue the alternative to reduce anchoring and extract a more balanced analysis.

    • •Use objective observations: onset, severity, appearance, what’s new/changed
    • •Don’t tell the model your suspicion (avoid “Is this scabies?” framing)
    • •Models may bias toward higher triage for safety/liability reasons
    • •Use adversarial prompting: “Convince me why it’s NOT X; what else could it be?”
  7. 13:12 – 15:49

    Which AI should you use for health questions—and why “ChatGPT for Health” is different

    They compare frontier chat models and clinician-only medical tools, noting most top models handle general health knowledge well. Chen describes the need for systematic safety benchmarking across models and suggests health-specific products may mainly offer privacy/security separation rather than fundamentally better reasoning.

    • •Frontier models (ChatGPT, Claude, Gemini, etc.) are broadly strong on general medicine
    • •Specialized clinician tools exist but may require medical credentials
    • •Differences often show up in refusal behavior, tone, and willingness to discuss topics
    • •Safety benchmarking across models is still immature; Chen’s group is testing systematically
  8. 15:49 – 19:12

    Does saying “I’m a doctor” help? Getting structure vs. more truth

    Chen explains that identity prompts mostly change formatting, detail, and style—not core correctness. For analytical users, structured outputs can be helpful; for others, too much technical detail can overwhelm, so specifying the desired format is key.

    • •“I’m a doctor” prompts can elicit more formal, structured, detailed responses
    • •Prompting rarely makes the model “more accurate,” but can change presentation
    • •Tailor output: plain-language vs. technical, bullet points vs. narrative
    • •Best use: translation and organization, not forcing high-stakes decisions
  9. 19:12 – 23:21

    When AI omits critical info: summarization risk, context rot, and “chart lore”

    They discuss a subtle failure mode: omission, not hallucination. Chen highlights how long-record summarization forces choices about what to exclude, and how models can get confused by dates/numbers or perpetuate copied diagnoses embedded in charts.

    • •Summaries compressing 1,000 pages can omit the one crucial follow-up detail
    • •Models struggle with numbers and timelines, leading to outdated or misapplied info
    • •“Context rot” increases as you overload the model with long histories
    • •“Chart lore”: copied-forward diagnoses can be repeated by both clinicians and AI
  10. 23:21 – 24:05

    What a “good” personal AI health workflow looks like: continuous coaching and better visits

    Chen describes an idealized use: an ongoing coach to explain records, prepare questions, and keep the patient oriented—while acknowledging real limitations. The aim is to improve the quality of scarce clinician time by arriving prepared with objective summaries and priorities.

    • •Use real clinician notes/labs as grounding, then ask for implications and questions
    • •Treat AI as a longitudinal assistant, not a single-shot oracle
    • •Have AI produce concise “what my doctor should know” bullet points
    • •Meet people where they are: usage is inevitable, so safer patterns should be taught
  11. 24:05 – 28:21

    Full-body scans and extra tests: why screening asymptomatic people often disappoints

    Chen pushes back on Silicon Valley enthusiasm for broad scanning and early detection tests, stressing base-rate problems and false leads. Screening works when validated at scale (e.g., colon, breast), but many popular approaches don’t change outcomes for most healthy people.

    • •Anecdotes of lifesaving scans hide thousands of low-yield/false-alarm experiences
    • •Rare-disease screening in healthy people often fails on probability alone
    • •Validated screening programs exist, but many others lack actionable accuracy
    • •Often the biggest levers remain “boring”: diet, exercise, and symptom-driven evaluation
  12. 28:21 – 29:27

    Genetic testing: when it’s high-leverage vs. mostly “interesting”

    They discuss when raw genetic data can meaningfully change care—especially in conditions like familial hypercholesterolemia where early detection prevents catastrophic outcomes. For most people, results may be informational but not strongly actionable.

    • •High-impact cases exist (e.g., familial hypercholesterolemia) with effective interventions
    • •Many serious risks are silent until a major event—genetics can reveal them earlier
    • •For ~90% of people, genetic findings may not drive major care changes
    • •Use genetics as risk context, not as deterministic predictions
  13. 29:27 – 33:31

    Can AI detect cancer earlier? Signal vs. noise and the limits of early detection

    Chen is cautiously optimistic that AI may discover subtle longitudinal patterns across biomarkers and time, beyond what humans can synthesize. But he emphasizes the seductive trap: detecting tiny lesions can also increase overdiagnosis and confusion, and biology plus treatment advances complicate the benefit calculus.

    • •Longitudinal multi-omics + AI could surface risk signals earlier than single tests
    • •Earlier isn’t always better: too-early findings may be non-actionable or irrelevant
    • •Greater sensitivity increases noise and overdiagnosis risk
    • •Outcomes depend on what you can do differently once something is detected
  14. 33:31 – 35:15

    When doctors can’t find an answer: using AI for “what might be missing?” and smarter questions

    For persistent, undiagnosed symptoms, Chen recommends using AI as a second set of eyes to widen the differential and suggest questions/tests—without leading it. He also reinforces appointment preparation: concise summaries and targeted questions help clinicians focus on judgment and decisions.

    • •Prompt: “Here’s what’s known—what might we be missing?” without suggesting a diagnosis
    • •Ask for other considerations, possible tests, or specialist directions to discuss
    • •Use AI to create a pre-visit brief and the best questions for limited appointment time
    • •AI can help counteract availability bias toward “common” diagnoses when stuck
  15. 35:15 – 48:33

    De-skilling, what makes a good doctor, and why multiple AIs can be safer than one

    They cover evidence that reliance on AI can reduce clinician skill when the tool is removed, raising big questions about which skills to retain. Chen’s framework for professionalism is competence, communication, and character; he also recommends multi-model second opinions to reduce single-system errors.

    • •Evidence of de-skilling: clinicians can become worse after dependence on AI aids
    • •Professionalism framework: competence, communication, character
    • •Knowledge alone isn’t the differentiator; judgment and values-based context are core
    • •Use multi-model “second opinions” (and synthesis) to improve safety and robustness

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.