Skip to content
OpenAIOpenAI

Episode 16: Building AI for Life Sciences

What does it take to build AI systems that can actually help scientists? Research lead Joy Jiao and product lead Yunyun Wang discuss how OpenAI is developing models for life sciences and what responsible deployment means in a field with real biosecurity stakes. They explore how AI is already improving research workflows and where it could lead in drug discovery and more autonomous labs — including why a future with less pipetting sounds pretty good to most scientists. Chapters 0:39 Introducing the Life Sciences model series 3:47 Joy’s path into life sciences 5:00 Autonomous lab with Ginkgo Bioworks 7:27 Yunyun’s path into life sciences 8:12 OpenAI’s life sciences work 9:48 Biorisk, access, and safeguards 15:43 What models can do in the lab 17:51 Building scientific infrastructure 20:14 Why compute matters for science 24:54 Where are we in 6-12 months? 29:51 Scientific adoption and skepticism 33:17 Advice for students and researchers 40:27 Where are we in 10 years?

Joy JiaoguestYunyun Wangguest
Apr 16, 202644mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 3:06

    Life Sciences model series: biochemistry-focused models, tools, and research workflows

    Andrew introduces OpenAI’s Life Sciences effort with Joy Jiao and Yunyun Wang, framing the goal: new models that can accelerate biology and medicine while being deployed responsibly. Yunyun outlines the Life Sciences model series, emphasizing mechanistic understanding (genomics/proteins), early discovery bottlenecks, and how product surfaces like ChatGPT and Codex support long-running, agentic workflows.

    • Launch of a biochemistry-focused “Life Sciences model series” anchored to real research workflows
    • Mechanistic understanding as a core direction: genomics and protein understanding first
    • Early discovery as a key bottleneck where more thinking time/compute can help
    • Workflow integration via ChatGPT (literature synthesis) and Codex (long-horizon agentic work)
    • Enterprise needs: reproducibility and repeatability in scientific pipelines
  2. 3:06 – 3:47

    From tool-using assistant to biochemistry expert: how models help computational biology today

    Joy describes how models already behave like computational biologists by calling established tools (e.g., protein structure prediction), inspecting outputs, and iterating inputs. She argues the next leap is endowing models with deeper biochemical intuition so they can choose and use tools more intelligently and reach correct answers faster.

    • Models can already run scientific tools, inspect outputs, and iterate like a computational biologist
    • Tool use is powerful, but expertise/intuition will unlock better decisions and speed
    • Protein structure prediction tools as an example of integrating open-source methods
    • Shift from ‘operator’ to ‘expert’ as a capability milestone
  3. 3:47 – 5:00

    Joy’s path: systems biology to software to accelerating science (without pipetting)

    Joy shares her trajectory from a Harvard PhD in systems biology to software and ultimately OpenAI, motivated by the slow pace and manual labor of lab work. She frames this as a “full circle” moment: using AI to accelerate the kind of biology she once did, ideally with robots handling wet-lab execution.

    • Background in systems biology; desire for faster iteration than academia typically allows
    • Manual wet-lab work (pipetting) as a key frustration and automation target
    • Software as a path to higher personal ‘velocity’ and scalable impact
    • Returning to biology via AI as a way to accelerate her past self’s workflow
  4. 5:00 – 6:20

    Autonomous lab collaboration with Ginkgo Bioworks: GPT-5 in the loop

    The conversation turns to OpenAI’s work with Ginkgo Bioworks, pairing GPT-5 with robotic lab automation to design and run experiments. Joy recounts early uncertainty about biological competence due to limited biology-focused training signals, and the surprise that the model produced non-zero protein output in initial runs—shifting expectations about AI’s potential to accelerate science.

    • Ginkgo collaboration began soon after GPT-5 training completed (mid-2025 timeline discussed)
    • Initial question: can the model do any biology and design workable experiments?
    • Surprising early success: experiments produced measurable protein (‘non-zero’ result)
    • Biology differs from math/CS because verification often requires real experiments
    • From skepticism to ‘it feels obvious’ models can accelerate science within months
  5. 6:20 – 7:27

    Scaling beyond human bottlenecks: parallel agents, compute-intensive data, and orchestration

    Yunyun interprets the Ginkgo results as proof of ‘the art of the possible’ and highlights ingestion of high-throughput experimental data as compute-heavy. She argues scientific progress is often constrained by human throughput, and envisions shifting the bottleneck from people to compute by running many coordinated sub-agents in parallel so researchers can focus on interpretation.

    • High-throughput experimental data is difficult and compute intensive to process
    • Human throughput is often the real bottleneck in experimental workflows
    • Future: many sub-agents running in parallel to divide-and-conquer research tasks
    • Researchers spend more time analyzing and interpreting the most meaningful outputs
  6. 7:27 – 9:46

    Yunyun’s path and OpenAI’s life sciences focus: from biorisk to beneficial bio R&D

    Yunyun explains she began with biorisk mitigation and biodefense initiatives, informed by infectious disease and virology work, then moved into life sciences product leadership. The team’s life sciences focus has been building for roughly two years via capability evaluations and early experiments, with more partners across protein/enzyme/chemical design and drug discovery.

    • Career arc: biorisk mitigations and biodefense → life sciences research/product
    • Wet-lab entry via infectious disease and virology; enduring biosecurity lens
    • Life sciences effort shaped by capability evals and model-in-the-loop experiments
    • Active/expected partnerships in chemical, protein, and enzyme design
    • AI’s potential across the pipeline: mechanism → target → drug design → approval acceleration
  7. 9:46 – 15:44

    Biorisk, information hazards, and differentiated access as a core safeguard strategy

    The discussion addresses dual-use risks: the same steps used for legitimate biology can resemble harmful workflows, making intent hard to infer from prompts. Yunyun and Joy describe a risk-averse approach—incremental deployment, refusals/high-level guidance for general users, layered mitigations—and argue that unlocking full capability requires differentiated access for verified institutions with strong controls.

    • Information hazards: benign-looking precursor steps can lead to dangerous outcomes
    • Dual-use similarity makes malicious intent hard to detect from prompts alone
    • Incremental deployment as capabilities rise; safeguards tied to capability gains
    • General access may trigger refusals or only high-level protocol explanations
    • Differentiated access: verified researchers/institutions with reagent tracking and enterprise controls
  8. 15:44 – 17:52

    What models can do in labs now: from pipetting optimization to idea scrutiny and target narrowing

    Joy and Yunyun describe current practical value: generating spreadsheets to reduce pipetting steps, writing analysis code, and supporting sophisticated design tasks when paired with tools. Yunyun highlights a key near-term role: models as rigorous critics/discriminators that test feasibility, synthesize literature, and narrow vast hypothesis/target spaces for experiments and drug discovery.

    • Low-level wins: spreadsheets/scripts to minimize pipetting steps and routine lab friction
    • Higher-end tasks: enzyme/protein design combined with biological design tools
    • Models as ‘skeptical’ reviewers to assess novelty/feasibility like human scientists
    • Literature synthesis and hypothesis triage to identify what’s worth testing
    • Disease target screening/selection: narrowing the aperture at scale
  9. 17:52 – 20:15

    Building scientific infrastructure on Codex: software, monitoring, visualization, and collaboration

    Joy outlines a vision where core scientific computing workflows live in Codex: running code across remote machines, monitoring logs, building fit-for-purpose analysis tools, and generating shareable UIs rather than raw data dumps. Yunyun extends this to organizational adoption—moving from individual assistants to institution-scale ‘agent workforces’ coordinating parallel research tasks.

    • Codex as a foundation for ‘everything possible on the computer’ in scientific workflows
    • Remote execution: run jobs across multiple dev boxes; monitor logs asynchronously
    • Rapid generation of analysis/visualization software tailored to experimental data
    • New collaboration norms: sharing interactive HTML/UIs instead of raw datasets
    • Adoption ladder: personal assistant → program/team agent workforce for institutions
  10. 20:15 – 22:33

    Why compute matters for science: model scale vs. test-time compute (reasoning time)

    Joy distinguishes two compute axes: bigger models yielding emergent capabilities, and test-time compute scaling where models ‘think’ longer for harder problems. She argues reasoning-time scaling can unlock new levels of difficulty and discovery—captured by the team’s internal motto about scaling compute to cure disease—connecting infrastructure investments to long-horizon scientific work.

    • Two compute axes: training larger models vs. scaling test-time compute
    • Emergent capabilities from parameter scaling (GPT-2 → GPT-3 analogy)
    • Reasoning models can allocate variable thinking time, potentially for days
    • Test-time compute as a pathway to harder scientific discovery, not just chat tasks
    • Team motto: ‘scale test time compute to cure all disease’
  11. 22:33 – 24:54

    Near-term impact areas: repurposing drugs, personalized medicine, and lab automation support

    Joy discusses where progress is already visible: drug repurposing suggestions based on mechanistic understanding and momentum in personalized medicine (e.g., RNA-based treatments like ASOs). She also emphasizes practical enablement—translating protocols to automation platforms and improving statistical rigor in analysis—positioning AI as an amplifier of scientists rather than a replacement.

    • Drug repurposing: mapping known FDA-cleared drugs to new indications/symptom relief
    • Personalized medicine acceleration, including RNA-based treatments (e.g., ASOs)
    • Protocol-to-automation translation is a major bottleneck; Codex helps bridge wet lab + coding
    • Models can guide rigorous statistics, suggest tests, and identify biases/QC issues
    • AI as an accelerator that keeps the scientist in the loop
  12. 24:54 – 27:05

    6–12 month outlook and the path to ‘virtual cell’ prediction: toxicity and complexity challenges

    Asked about the next 6–12 months, Joy and Yunyun focus on accelerating stages of the drug pipeline, with the long tail from discovery to market still taking years. Joy expects models to progress from strong chemical reaction outcome prediction toward harder biological predictions (like toxicity in specific people), while Yunyun highlights near-term wins such as enabling patentable discoveries on OpenAI’s platforms.

    • Hopeful milestones: AI-designed drugs/cures may take years, but earlier pipeline gains are near-term
    • Speed-ups may be most feasible in design, safety review, and clinical-trial-adjacent stages
    • Current strength: predicting chemical reaction outcomes
    • Hard frontier: predicting toxicity and complex biological outcomes at individual/system levels
    • Near-term success metric: researchers achieving patentable discoveries using the platform
  13. 27:05 – 29:51

    Evaluating biology models: experimental ground truth, synthetic ‘messiness,’ and wet-lab validation

    Joy explains biology evaluation strategies: predicting outcomes of already-run experiments, using massive datasets like single-cell RNA-seq for perturbation prediction, and generating synthetic datasets with planted biases to test computational biology competence. Both emphasize that wet-lab validation remains the ultimate test, and Yunyun notes evals should reflect real-world messiness and can progress from recreating baselines to pushing into novel antibody binding and design.

    • Evals via experimental ground truth: predict outcomes of known experiments
    • Virtual-cell style tasks using single-cell RNA-seq and unseen perturbations
    • Synthetic data with planted artifacts to test QC, bias detection, and statistical correction
    • Wet-lab experiments as the ‘real’ final evaluation in biology
    • Antibody binding prediction evals as a stepping stone to de novo antibody design and new treatments
  14. 29:51 – 33:17

    Adoption, skepticism, and making AI ‘just work’: demos, publications, and lowering cognitive load

    Joy notes a cultural split: West Coast enthusiasm vs. East Coast skepticism, and both argue the gap is bridged through hands-on utility and peer-reviewed publications with wet-lab proof. They also highlight a different barrier—stress about using AI ‘correctly’—and describe a product vision where scientists can state goals and the system handles prompting, multi-agent orchestration, and tool use automatically.

    • Regional/cultural variance in receptiveness to AI in life sciences
    • Winning trust via practical demos (e.g., dilution spreadsheets) and deep collaborations leading to publications
    • Skepticism is welcomed as a healthy force that demands rigor
    • Need for workflow-representative evals so labs can see clear adoption pathways
    • Product goal: reduce prompting burden so scientists can express intent and the system orchestrates the rest
  15. 33:17 – 44:25

    Advice for students and researchers + a 10-year vision: autonomous labs, democratized expertise, and defenses

    Joy advises students to shift from memorization toward exploration with AI—asking questions, reading papers, and connecting ideas—while Yunyun encourages early adoption via low-pressure projects and collaboration. Looking 10 years ahead, they envision autonomous robot labs directed by humans, accelerated work on rare diseases and personalized medicine, and improved biosecurity through environmental surveillance and faster medical countermeasures, with expert-level capabilities democratized to more people.

    • Students: focus less on memorization, more on exploration and connecting research with AI
    • Researchers: start with low-lift uses—paper understanding, data analysis, and hobby/passion projects to build fluency
    • Collaboration patterns: sharing scripts/conversations; ‘co-scientist’ agents working together
    • 10-year vision: autonomous robotic labs + AI exploring, designing experiments, and iterating with humans
    • Democratized expert knowledge; faster countermeasures and surveillance (wastewater/air) for emerging threats

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.