OpenAIBuilding AI for better healthcare — the OpenAI Podcast Ep. 14
CHAPTERS
- 0:00 – 0:40
OpenAI’s health focus: safe, sensitive answers for patients and clinicians
Andrew Mayne introduces Dr. Nate Gross and Karan Singhal and frames the episode around building AI that can handle high-stakes, sensitive healthcare questions. The conversation sets expectations: model training, safety/privacy, and real-world impact for clinicians, patients, and health systems.
- •OpenAI’s health work spans patients, clinicians, and healthcare systems
- •Core emphasis on handling sensitive medical queries responsibly
- •Episode focus: how models are trained, evaluated, and deployed in healthcare
- 0:40 – 1:47
Nate Gross’s path into healthcare: policy, public hospitals, and tech gaps
Nate explains his entry into healthcare through health policy and value-based care, then medical school experience at a major public hospital. He contrasts consumer tech’s rapid progress with the outdated tools doctors used, motivating his drive to modernize healthcare workflows.
- •Early interest in health policy and access
- •Training in settings with high clinical volume and complexity
- •Frustration with legacy clinical tech (fax, clipboards, early EHRs) versus consumer apps
- 1:47 – 3:10
Karan Singhal’s motivation: AI safety roots and the opportunity in clinical AI
Karan describes a long-standing fascination with intelligence and a conviction that advanced AI would arrive within our lifetimes. He connects his safety/privacy background to healthcare, arguing the clinical world underestimated what LLMs could unlock and why it’s a responsibility to pursue.
- •Philosophy-of-mind interest leading into AI research
- •Focus on maximizing upside while avoiding downside (safety lens)
- •Healthcare as a major, underappreciated opportunity for LLM impact
- 3:10 – 4:59
Why healthcare is a prime target: fragmentation, missed care, and access at scale
Nate lays out systemic problems—fragmentation, reactive care, and limited clinician time—creating large gaps in outcomes and equity. He positions OpenAI’s technology as a scalable way to improve access to knowledge and support the entire health ecosystem.
- •Care gaps driven by fragmentation and limited patient engagement time
- •System tends to be reactive rather than proactive
- •OpenAI’s role: scale access for patients, clinicians, and builders/entrepreneurs
- 4:59 – 7:06
Strategy for ChatGPT Health: demand, privacy protections, and patient context
Nate explains why health is a major ChatGPT use case and how that creates responsibility to build dedicated experiences. The strategy centers on strong privacy guarantees and enabling users to bring their own context so advice is more personalized than one-size-fits-all search.
- •Scale of usage: large share of ChatGPT queries are health-related
- •Privacy/security: encryption and a commitment not to train on user health data
- •“Empowerment” via user-consented context to improve relevance over generic search
- 7:06 – 10:15
Training and evaluation approach: HealthBench and physician-built multi-turn rubrics
Karan details how OpenAI prioritized evaluation first, treating healthcare as a concrete proving ground for safety and alignment. He introduces HealthBench: a realistic, multi-turn conversational evaluation built with hundreds of physicians to measure nuanced performance and safety behaviors.
- •Healthcare used to ground safety/alignment work with real incentives
- •HealthBench evaluates multi-turn conversations, not just Q&A
- •~250 physicians involved across data generation and rubric design
- •Measures behaviors like asking for context before answering
- 10:15 – 14:22
Why OpenAI scores strongly on health evals: health integrated across the training stack
Karan argues performance comes from a cross-functional effort spanning pre-deployment evals, production monitoring, and physician collaboration. Nate adds that meaningful health performance requires more than exam-style accuracy—models must handle nuance, escalation, literacy adaptation, and uncertainty.
- •Health considerations integrated from pre-training through post-training
- •Continuous monitoring and safety work in production (privacy-preserving)
- •Healthcare isn’t multiple-choice: needs contextual reasoning and escalation
- •Adaptive literacy and better uncertainty calibration reduce harmful overconfidence
- 14:22 – 17:05
Deployment challenges and near-term frontiers: access, multimodal data, and workflow evaluation
Karan and Andrew discuss how lower AI costs improve access and broaden reach (including free users). Karan highlights future “zero-to-one” capabilities from integrating long-term, multimodal health data, and emphasizes evaluating real clinical workflows post-deployment.
- •Access improves as cost of intelligence drops; wider rollout goals
- •Multimodal integration (wearables, labs, longitudinal records) enables new predictions
- •Analogy: AI as a protective safety net (like self-driving cars)
- •Need for post-deployment workflow monitoring and measurement
- 17:05 – 18:11
Clinical co-pilot in Kenya: intervening only when needed to reduce errors
Karan describes a study with Penda Health clinics in Nairobi using an AI “clinical co-pilot” that monitors EHR entries and interrupts only when something looks concerning. The reported result was a statistically significant reduction in diagnostic and treatment errors, underscoring the value of workflow-embedded safety nets.
- •Deployed across ~20 clinics with workflow-sensitive interruptions
- •Monitors clinician documentation and flags potential issues
- •Study outcome: reduced diagnostic/treatment errors vs control
- •Represents a shift from benchmarks to real-world outcome evaluation
- 18:11 – 21:25
Trust and integration hurdles for professionals: grounding in guidelines and connecting silos
Nate outlines clinician needs: trust, up-to-date grounding (literature, guidelines, local institutional policies), and secure handling of sensitive data. He emphasizes healthcare’s siloed systems and the difficulty of connecting many tools and data formats into unified AI layers.
- •Grounding answers in latest evidence and local guidance builds trust
- •Need for secure/HIPAA-aligned environments to use sensitive context
- •Healthcare data/tooling is fragmented: structured/unstructured, cloud/non-cloud
- •AI value grows with connectors that prevent information from “falling through cracks”
- 21:25 – 22:26
Collaboration with health systems: standards-based record syncing and consented context
Nate explains that broad collaboration—government standards, EHR interoperability, and consumer health ecosystems—enables patients to bring records into ChatGPT in a few taps. The goal is consented, low-friction context sharing so AI can offer more relevant, actionable support.
- •Interoperability depends on national standards and ecosystem cooperation
- •Patients can sync EHR data with minimal steps under consent
- •Integration with wearables and biosensors expands context
- •Combining datasets can unlock insights not possible in isolated apps
- 22:26 – 25:33
Everyday health assistant use: translating data into decisions and follow-through
Nate and Andrew discuss practical uses: relating sleep, stress, and exercise to scheduling, menu planning, and daily choices. Nate frames ChatGPT as a unifying layer that extends partner tools into broader life contexts, helping patients adhere to care plans and reduce friction in navigation.
- •Turns wearable/behavior data into actionable daily recommendations
- •Supports lifestyle planning (food choices, activity, calendar priorities)
- •Acts as an organizing “spouse-with-a-clipboard” analog—only with consent
- •Reduces friction: parsing old info, tracking tasks, following care plans
- 25:33 – 26:42
OpenAI’s three-bucket framework: raise the floor, sweep the floor, raise the ceiling
Nate summarizes how OpenAI thinks about impact: broaden access, reduce administrative burden, and enable transformative advances. The framing ties patient empowerment and clinician time savings to longer-term breakthroughs in medical capability.
- •Raise the floor: make AI benefits accessible to everyone
- •Sweep the floor: cut admin/bureaucracy so clinicians have more patient time
- •Raise the ceiling: enable new levels of medical capability while clinicians stay in control
- •Clinician time scarcity is a core design constraint
- 26:42 – 30:54
Biggest “wow” moments and clinician feedback: from adoption to ‘indispensable’ safety nets
Karan’s highlight is the rapid growth of health/wellness usage, validating the need for responsible health AI. Both share field feedback: in Nairobi, clinicians hesitated to run another study because withholding AI felt unsafe; Nate also notes caregiver stories and rare but meaningful “miracle” cases where AI accelerates diagnosis or care access.
- •Rapid adoption of health use cases is a major milestone
- •Change management and workflow fit can make tools indispensable
- •Clinicians viewed non-AI control groups as potentially dangerous in one setting
- •User stories range from caregiver support to accelerated diagnosis in critical cases