No Priors“Curing All Disease by next century is too conservative" - Mark Zuckerberg
CHAPTERS
- 0:00 – 1:05
Cold open: Open tools, individualized medicine, and protein world models
The episode opens with the core thesis: Biohub aims to build open tools that help the entire scientific community understand and engineer biology. The cold open previews personalized medicine, open-source distribution, and large-scale protein modeling as foundational capabilities.
- •Commitment to open tools for the scientific community
- •Vision for mechanistic, individualized understanding of disease risk and intervention
- •Open-source as a faster path to broad impact
- •Protein-language-model framing: models that “understand proteins” enable downstream design
- 1:05 – 1:41
Meet the guests and the Biohub “virtual biology” agenda
The hosts introduce Mark Zuckerberg, Priscilla Chan, and Alex Rives and set the stage: applying frontier AI to build world models across biological scales. The conversation frames Biohub as a long-term, platform-style effort rather than a single-disease program.
- •Biohub positioned as AI-at-scale for biological world models
- •Focus spans proteins, cells, and higher-order biological interactions
- •$500M virtual biology initiative as a major philanthropic commitment
- •Emphasis on tool-building rather than productizing one therapy
- 1:41 – 6:10
Origin story: Why “cure all disease” led to tool-building and shared infrastructure
Priscilla and Mark recount early conversations with scientists—some of whom laughed at the ambition—and the practical blockers they learned about. The key lesson: scientific progress is slowed by siloed work, poor tool-sharing, and lack of durable, community infrastructure.
- •Initial philanthropic goal: cure/prevent/manage all disease by end of century
- •Scientists’ skepticism prompted deeper inquiry into what’s missing
- •Structural issues: silos, slow sharing, and fragile one-off tools
- •Tooling and shared knowledge bases as the leverage point
- 6:10 – 8:27
From single-cell sequencing to data ecosystems: Human Cell Atlas and Cell by Gene
The discussion traces how early Biohub funding around single-cell sequencing evolved into large community datasets and annotation tooling. This progression becomes a template: build tools, catalyze communities, and let external contributors expand the corpus over time.
- •Early RFA focus on single-cell transcriptomics methods and sharing
- •Human Cell Atlas as a large-scale reference database
- •Cell by Gene as an annotation tool that seeded a broader community
- •Shift from “stamp collecting” critique to models that can extract insights
- 8:27 – 14:22
Integrating frontier AI and frontier biology: Building hierarchical models from proteins upward
The guests explain why biology requires coupling cutting-edge AI with wet-lab innovation to generate new data types. They outline a hierarchical modeling strategy—proteins to cells to systems—and the need for experiments that create “connective tissue” across levels.
- •Biology lacks an internet-scale dataset; new measurement methods are required
- •Hierarchical approach: proteins → cells → immune/system dynamics
- •Wet-lab + AI designed as one integrated feedback loop
- •Examples of bridging data: spatial transcriptomics, zebrafish imaging, molecular sensors
- 14:22 – 16:58
Mechanistic interpretability for protein language models: Opening the biological black box
Alex describes applying mechanistic interpretability to protein language models to uncover what the models represent about structure and function. The goal is not only prediction, but extracting new biological knowledge and connecting unknown proteins to known biology.
- •Protein language models learn biology emergently from sequence prediction
- •Interpretability tools can map representations to structure/function concepts
- •Potential to connect uncharacterized proteins to known mechanisms
- •Long-term ambition: interrogate models to propose new mechanisms of action
- 16:58 – 21:42
Why a nonprofit and why open source: Scale, time horizons, and inclusion of the long tail
Mark and Priscilla argue that open-source nonprofit structure maximizes impact by distributing tools broadly and attracting more collaborators. They highlight constraints unique to biology—data generation and method invention—and the importance of enabling work on rare and niche diseases.
- •Open source as the fastest path to widespread scientific adoption
- •Biological data requires new experimental methods, not just funding compute
- •10–15 year horizons fit nonprofit better than typical venture timelines
- •Decentralizing tools helps address rare diseases and under-resourced areas
- 21:42 – 26:26
From diseases to mechanisms: Personalized medicine and system-level targets like inflammation
Rather than predicting which disease gets cured first, the conversation centers on building mechanistic chains from gene variants to proteins to disease processes. Mark adds that Biohub often targets cross-cutting systems—like inflammation and immunity—that influence many conditions.
- •Goal: mechanistic links from genetics → proteins → disease processes
- •Personalized interventions enabled by generalizable, reusable models
- •Systems focus examples: inflammation (Chicago) and immune system modeling
- •Biohub builds tools; biotechs/academia can translate to specific therapies
- 26:26 – 28:05
The clinic translation gap: Rethinking clinical research and safe deployment
Priscilla emphasizes that faster basic science does not automatically translate into faster patient impact. She points to clinical research processes as a bottleneck and describes early efforts (e.g., CRISPR Cures partnerships) aimed at shortening bench-to-bedside timelines responsibly.
- •Scientific acceleration doesn’t guarantee clinical acceleration
- •Need to redesign clinical research and deployment practices
- •Focus on safely shortening the bench-to-patient distance
- •CRISPR Cures partnership as an early translation experiment
- 28:05 – 32:15
ESMFold2 launch: Folding 1.1B proteins and designing binders digitally
Alex details ESMFold2 as a fast, accurate protein-structure and interaction prediction system trained on billions of sequences. He describes using it as a “world model” to design proteins and even single-chain antibodies, then validating candidates experimentally to reach nanomolar binding.
- •ESMFold2: language-model-based protein world model
- •Blazing-fast structure prediction; strong interaction benchmarks
- •Folded/predicted structures for 1.1B proteins
- •Design loop: search digitally → synthesize small set → validate in assays/cryo-EM
- 32:15 – 38:41
Off-target effects, toxicity, and rare-disease edge cases: New economics for trials
The group explores how cell atlases and multi-scale models could predict off-target binding and toxicity earlier—before expensive trials. They also discuss rare disease communities organizing registries and trials, and how opt-in participation plus cheaper design could unlock the long tail.
- •Predicting off-target effects via receptor expression across cell types
- •Transcriptomic models as early warning for toxicity (e.g., kidney effects)
- •Example: targeted CRISPR success enabled by delivery practicality (liver)
- •Rare disease cohorts as organized, fast-moving trial enablers; edge cases teach system behavior
- 38:41 – 44:28
Putting technology in individuals’ hands: Open ecosystems, biosafety, and hiring mission-driven teams
Mark connects Biohub’s openness to a broader philosophy: progress comes from empowering many individuals rather than centralizing capability. They acknowledge biosafety tradeoffs, then discuss recruiting—arguing Biohub’s unique combo of frontier AI + frontier biology plus mission attracts top talent.
- •Decentralized innovation: tools for individuals, not a single central solver
- •Open source as one mechanism; biosafety must be balanced
- •Biohub’s differentiator: frontier AI tightly coupled to frontier wet labs
- •Talent thesis: small, elite teams can drive major progress when mission-aligned
- 44:28 – 49:55
Beyond ESMFold2: Agentic design workflows and the “virtual cell” as the next frontier
Alex and Mark outline what comes next: integrating ESMFold2 into agentic systems that automate iterative design, and laddering up to a virtual cell model. They describe desired inputs/outputs in broad terms—linking genetic, proteomic, and transcriptomic layers to phenotype with strong generalization.
- •Early community experimentation: connecting models to agentic systems
- •Research agenda choice driven by constraints: compute, data, and model scaling
- •Virtual cell goal: integrate multiple biological layers and predict phenotype
- •Key challenge: generate enough high-quality, generalizable data to close the gap
- 49:55 – 56:20
Defining success and strategy updates: Directed, closed-loop biology at scale
The final segment defines near-term success as producing uniquely world-class, open scientific contributions that others can build on. Mark and Priscilla describe strategy shifts—formalizing Biohub as the primary philanthropic focus, installing an AI-first leadership team, and tightening integration across teams to close the loop between models and experiments.
- •Five-year success metric: uniquely better, world-class hierarchical biology models
- •Expectation that broad external idea generation follows strong core releases
- •Strategy update: Biohub becomes central philanthropic focus; AI-first leadership shift
- •Teams “arms linked”: more directed execution; closing the loop between AI and wet lab