CHAPTERS
- 0:00 – 1:58
Moonshot vision: virtual cells to make science faster
Patrick Hsu lays out Arc Institute’s core ambition: use foundation models to build “virtual cells” that can simulate human biology. The goal is to turn slow, physical lab iteration into faster, more parallelizable computation that experimental biologists actually trust and use by default.
- •Science is constrained by real-world lab time (cells, tissues, animals), unlike GPU-only domains
- •Virtual cells as a foundation for simulating biology from the cellular unit upward
- •Emphasis on usefulness to wet-lab experimentalists, not just ML benchmarks
- •Framing: if we can’t model a cell well, modeling whole bodies is premature
- 1:58 – 2:36
Why scientific progress is slow: incentives, training, and complexity
The conversation unpacks the “Gordian knot” behind slow science, from misaligned incentives to the increasing multidisciplinary nature of modern biology. Patrick argues that many bottlenecks are structural and cultural, not just technical.
- •Science speed limited by incentives, funding structure, and academic career dynamics
- •Separation of basic science vs commercially viable work can slow translation
- •Modern problems require many disciplines; most groups can’t be excellent at five things at once
- •Complexity compounds across understanding, perturbation, and safety
- 2:36 – 5:11
Arc Institute as an organizational experiment: collision frequency and flagship projects
Patrick explains Arc’s design: co-locating multiple disciplines to increase collaboration and reduce incentive friction. Arc focuses on large “flagship” efforts that require coordinated, cross-domain work.
- •Physical proximity and shared goals increase “collision frequency” across disciplines
- •Universities can be too diffuse; labs are distributed and incentives are individualistic
- •Arc’s flagship projects: Alzheimer’s target discovery and virtual cells
- •Big projects demand infrastructure, people, and shared incentives beyond typical labs
- 5:11 – 6:53
Why AI is ahead in language/images vs biology: evaluation and lab-in-the-loop reality
Patrick argues biology is intrinsically harder to model than text or images, and harder to evaluate. Iteration cycles slow because validation requires experiments and interpretation of noisy, indirect readouts.
- •Humans can easily judge language/image outputs; biology outputs are less interpretable
- •We don’t “speak DNA” natively—evaluation is indirect and token-based
- •Biology requires lab-in-the-loop validation with experimental ground truth
- •Key need: faster measurement, better interpretability, and higher-dimensional feedback loops
- 6:53 – 10:23
What ‘virtual cells’ mean in practice: scaling measurement, mirrors, and data tiers
They discuss missing measurements in biology and how models can still learn useful representations from scalable data like RNA. Patrick introduces a practical view: bet on what scales now, and progressively add modalities like proteins, spatial context, and time.
- •We can’t measure everything (metabolites, spatial/temporal detail) at high throughput yet
- •RNA can act as a lower-resolution ‘mirror’ of protein/metabolic states at scale
- •Roadmap: single cells → pairs → tissues → intact physiological contexts
- •Framework: invention vs engineering vs scaling; Arc invests in novel tech, not just scaling existing assays
- 10:23 – 13:14
From AlphaFold to virtual cells: perturbation prediction as the wedge
Patrick frames the virtual cell “AlphaFold moment” as a tool that becomes the default for predicting perturbation outcomes. Arc operationalizes this by focusing on perturbation prediction across a manifold of cell states—directly aligned with how drugs work.
- •AlphaFold provides an analogy: not perfect physics, but useful end-state predictions
- •Virtual cells framed as predicting how perturbations move cells between states
- •Drug development seen as ‘click-and-drag’ from diseased to healthy cell states
- •Complex disease implies combinations and sequences of perturbations (purposeful polypharmacology)
- 13:14 – 15:27
Lab usefulness over benchmarks: co-pilot for biologists and in-silico target ID
Patrick stresses that virtual cells must guide real experimental decisions, not just improve ML metrics. The intended loop is prediction → targeted experiments → model improvement, ultimately enabling in-silico target identification and combination strategies.
- •Goal: actionable ranked experiments (e.g., ‘the 12 things to test’)
- •Lab-in-the-loop iteration to refine models with real measurements
- •Long-term: in-silico target ID and suggestions for drug compositions
- •Vision: AI-enabled vertically integrated pharma is compelling, but research capability must mature first
- 15:27 – 19:20
Where we are on the capability curve: GPT-1/2 stage and what GPT-3 would look like
Patrick estimates virtual cell modeling is still early—between “GPT-1 and GPT-2.” A ‘GPT-3 moment’ would be convincing biological rediscovery: models predicting canonical interventions and mechanisms familiar from textbooks and Nobel-level examples.
- •Current models yield ‘blurry pictures of life’; generated genomes likely not viable yet
- •Full-stack approach required: public data curation, private data generation, benchmarks, architectures
- •Proposed evals: rediscover Yamanaka factors (reprogramming) and differentiation factors (e.g., NGN2, ASCL1, MyoD)
- •Shift from MAE-style benchmarks to explainable, textbook-grade demonstrations
- 19:20 – 22:13
Textbooks as compressed ground truth; simulation vs mechanistic understanding
Erik probes whether textbooks are wrong; Patrick argues they’re compressed abstractions with exceptions. They distinguish between mechanistic explanations and forecasting-style simulation—where accurate predictions can be valuable even without full causal clarity.
- •Textbooks represent reliable knowledge but simplify multi-dimensional biology
- •Discovery often means finding exceptions and contextual dependencies
- •Virtual cells can be scoped and more rigorous than ‘digital twin’ talk
- •Analogy: weather prediction—useful without fully explaining underlying causality
- 22:13 – 30:09
Industry reality: why biotech growth lags—clinical trial bottlenecks and capital intensity
The discussion shifts to why scientific innovation hasn’t translated into stronger biotech business outcomes. Jorge highlights the hard constraint of clinical trials, the need to reduce cost/time, and why capital intensity plus weak valuation step-ups discourage early-stage investment.
- •Drug failures: wrong target vs wrong molecule (or both)
- •Clinical trials are necessary bottlenecks; some endpoints inherently take time
- •Levers: reduce capital intensity, compress timelines, increase effect sizes
- •Investment challenge: early investors bear capital burden without commensurate value inflection in valuations
- 30:09 – 35:49
GLP-1s as proof of value: ambition, patient populations, and risk management
Patrick argues GLP-1s created extraordinary market value by addressing huge patient populations, reshaping industry ambition. The high failure rate encourages ‘circling the wagons’ around validated mechanisms, but the GLP-1 example shows the upside of bold bets when impact is broad.
- •GLP-1s added enormous market cap (Lilly/Novo), dwarfing much of biotech’s aggregate value creation
- •High trial failure encourages focus on well-validated biology but smaller populations (lower EV)
- •Large-population, high-impact wins can reset incentives and ambition
- •Genetic medicines and new modalities expand what’s possible for hard diseases
- 35:49 – 38:15
Future breakthroughs and the next bottlenecks: making and testing, plus regulatory friction
Even if AI makes design abundant—‘a trillion binders in silico’—the physical world remains limiting. Patrick and Jorge emphasize that manufacturing and testing (animals → humans) will dominate, and Patrick notes how regulatory structures influence where trials happen and how fast progress can be.
- •Drug pipeline bottlenecks: designing vs making vs testing; making/testing may dominate in AI-success world
- •Testing sequence remains slow: mice → dogs → monkeys → humans
- •Regulatory regime shapes speed; discussion of running early trials overseas then returning for efficacy trials
- •Core need: increase success probability so the long journey is ‘worth it’ more often
- 38:15 – 42:04
AI drug discovery: where there’s hype, hope, and real ‘heft’ today
Patrick gives a grounded assessment of AI’s current strengths and weaknesses in biotech. He calls toxicity prediction overhyped, sees strong traction in protein-related modeling and some clinical imaging applications, and argues ‘AI-designed drugs’ will soon be a baseline capability rather than a marketing claim.
- •Hype: toxicity prediction from structure alone is still unreliable
- •Heft: protein binding and protein design; also pathology/radiology automation
- •Many ‘first AI-designed drug’ claims are marketing; AI will become a native part of the stack
- •Progress limited by feedback loops that require slow real-world testing
- 42:04 – 45:50
Parallelizing discovery: Dario Amodei’s thesis and what virtual cells could enable
Patrick reacts to predictions about rapid biomedical progress by emphasizing parallelization if discoveries are sufficiently independent. With predictive models spanning design, docking, and cellular response, discovery could increasingly resemble a compute problem—though validation remains critical.
- •Key premise: discovery tasks can be parallelized if largely independent
- •Layered model stack: design → docking → off-target analysis → cellular correction prediction
- •Potential timeline compression if steps become reliable and composable
- •Long-run plausibility paired with near-term focus on tangible, testable model capabilities
- 45:50 – 54:02
Patrick’s broader investing focus: synthetic biology, BCIs, robotics, and agents
The conversation widens to Patrick’s thesis on technologies that can meaningfully improve the human experience within a lifetime. He highlights synthetic biology, brain-computer interfaces, robotics, and AI agents—stressing the need for teams that combine technical innovation with product and commercial execution.
- •Priority areas: synthetic biology (health/longevity), BCIs, industrial & consumer robotics
- •Agents as ‘real work’ replacement—services economy opportunity beyond SaaS
- •Execution constraint: aligning technical, product, and business capabilities within teams
- •Expectation of new ML architectures beyond the 2017 transformer and renewed scaling of underexplored ideas
- 54:02 – 55:38
Closing: Arc’s Virtual Cell Challenge and an open benchmark for progress
Patrick closes by pointing to a concrete initiative: an open competition to evaluate perturbation prediction models, inspired by CASP’s role in protein folding progress. The aim is transparent measurement of capability over time and broader community participation.
- •Virtual Cell Challenge modeled after CASP for protein folding
- •Open competition with prizes and external sponsors to drive adoption and benchmarking
- •Goal: transparent tracking toward a ‘ChatGPT moment’ for virtual cells
- •Call for participation from bioML experts and engineers from other domains
