Skip to content
a16za16z

Faster Science, Better Drugs

Can we make science as fast as software? In this episode, Erik Torenberg talks with Patrick Hsu (cofounder of Arc Institute) and a16z general partner Jorge Conde about Arc’s “virtual cells” moonshot, which uses foundation models to simulate biology and guide experiments. They discuss why research is slow, what an AlphaFold-style moment for cell biology could look like, and how AI might improve drug discovery. The conversation also covers hype versus substance in AI for biology, clinical bottlenecks, capital intensity, and how breakthroughs like GLP-1s show the path from science to major business and health impact. Timecodes: 00:00 Introduction to Accelerating Science 00:35 Welcome to the Podcast 00:45 The Moonshot: Virtual Cells and Human Biology 01:57 Challenges in Scientific Progress 02:58 Interdisciplinary Collaboration at Arc Institute 05:11 The Role of AI in Biology 10:18 Understanding Virtual Cells 22:13 Biotech and Pharma Industry Insights 27:39 Challenges in Clinical Trials 28:13 Capital Intensity and Technological Advancements 29:02 Improving Drug Development 30:12 Market Impact and Industry Trends 33:19 Future Technological Breakthroughs 38:15 AI in Drug Discovery 45:50 Investment Focus and Future Prospects 54:02 Closing Remarks and Upcoming Initiatives Resources: Find Patrick on X: https://x.com/pdhsu Find Jorge on X: https://x.com/JorgeCondeBio Stay Updated: Find a16z on X: https://x.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX?si=3E8B3qT9TyiwAHJ7JnaKbg Listen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Patrick HsuguestJorge CondeguestErik Torenberghost
Sep 15, 202555mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:58

    Moonshot vision: virtual cells to make science faster

    Patrick Hsu lays out Arc Institute’s core ambition: use foundation models to build “virtual cells” that can simulate human biology. The goal is to turn slow, physical lab iteration into faster, more parallelizable computation that experimental biologists actually trust and use by default.

    • Science is constrained by real-world lab time (cells, tissues, animals), unlike GPU-only domains
    • Virtual cells as a foundation for simulating biology from the cellular unit upward
    • Emphasis on usefulness to wet-lab experimentalists, not just ML benchmarks
    • Framing: if we can’t model a cell well, modeling whole bodies is premature
  2. 1:58 – 2:36

    Why scientific progress is slow: incentives, training, and complexity

    The conversation unpacks the “Gordian knot” behind slow science, from misaligned incentives to the increasing multidisciplinary nature of modern biology. Patrick argues that many bottlenecks are structural and cultural, not just technical.

    • Science speed limited by incentives, funding structure, and academic career dynamics
    • Separation of basic science vs commercially viable work can slow translation
    • Modern problems require many disciplines; most groups can’t be excellent at five things at once
    • Complexity compounds across understanding, perturbation, and safety
  3. 2:36 – 5:11

    Arc Institute as an organizational experiment: collision frequency and flagship projects

    Patrick explains Arc’s design: co-locating multiple disciplines to increase collaboration and reduce incentive friction. Arc focuses on large “flagship” efforts that require coordinated, cross-domain work.

    • Physical proximity and shared goals increase “collision frequency” across disciplines
    • Universities can be too diffuse; labs are distributed and incentives are individualistic
    • Arc’s flagship projects: Alzheimer’s target discovery and virtual cells
    • Big projects demand infrastructure, people, and shared incentives beyond typical labs
  4. 5:11 – 6:53

    Why AI is ahead in language/images vs biology: evaluation and lab-in-the-loop reality

    Patrick argues biology is intrinsically harder to model than text or images, and harder to evaluate. Iteration cycles slow because validation requires experiments and interpretation of noisy, indirect readouts.

    • Humans can easily judge language/image outputs; biology outputs are less interpretable
    • We don’t “speak DNA” natively—evaluation is indirect and token-based
    • Biology requires lab-in-the-loop validation with experimental ground truth
    • Key need: faster measurement, better interpretability, and higher-dimensional feedback loops
  5. 6:53 – 10:23

    What ‘virtual cells’ mean in practice: scaling measurement, mirrors, and data tiers

    They discuss missing measurements in biology and how models can still learn useful representations from scalable data like RNA. Patrick introduces a practical view: bet on what scales now, and progressively add modalities like proteins, spatial context, and time.

    • We can’t measure everything (metabolites, spatial/temporal detail) at high throughput yet
    • RNA can act as a lower-resolution ‘mirror’ of protein/metabolic states at scale
    • Roadmap: single cells → pairs → tissues → intact physiological contexts
    • Framework: invention vs engineering vs scaling; Arc invests in novel tech, not just scaling existing assays
  6. 10:23 – 13:14

    From AlphaFold to virtual cells: perturbation prediction as the wedge

    Patrick frames the virtual cell “AlphaFold moment” as a tool that becomes the default for predicting perturbation outcomes. Arc operationalizes this by focusing on perturbation prediction across a manifold of cell states—directly aligned with how drugs work.

    • AlphaFold provides an analogy: not perfect physics, but useful end-state predictions
    • Virtual cells framed as predicting how perturbations move cells between states
    • Drug development seen as ‘click-and-drag’ from diseased to healthy cell states
    • Complex disease implies combinations and sequences of perturbations (purposeful polypharmacology)
  7. 13:14 – 15:27

    Lab usefulness over benchmarks: co-pilot for biologists and in-silico target ID

    Patrick stresses that virtual cells must guide real experimental decisions, not just improve ML metrics. The intended loop is prediction → targeted experiments → model improvement, ultimately enabling in-silico target identification and combination strategies.

    • Goal: actionable ranked experiments (e.g., ‘the 12 things to test’)
    • Lab-in-the-loop iteration to refine models with real measurements
    • Long-term: in-silico target ID and suggestions for drug compositions
    • Vision: AI-enabled vertically integrated pharma is compelling, but research capability must mature first
  8. 15:27 – 19:20

    Where we are on the capability curve: GPT-1/2 stage and what GPT-3 would look like

    Patrick estimates virtual cell modeling is still early—between “GPT-1 and GPT-2.” A ‘GPT-3 moment’ would be convincing biological rediscovery: models predicting canonical interventions and mechanisms familiar from textbooks and Nobel-level examples.

    • Current models yield ‘blurry pictures of life’; generated genomes likely not viable yet
    • Full-stack approach required: public data curation, private data generation, benchmarks, architectures
    • Proposed evals: rediscover Yamanaka factors (reprogramming) and differentiation factors (e.g., NGN2, ASCL1, MyoD)
    • Shift from MAE-style benchmarks to explainable, textbook-grade demonstrations
  9. 19:20 – 22:13

    Textbooks as compressed ground truth; simulation vs mechanistic understanding

    Erik probes whether textbooks are wrong; Patrick argues they’re compressed abstractions with exceptions. They distinguish between mechanistic explanations and forecasting-style simulation—where accurate predictions can be valuable even without full causal clarity.

    • Textbooks represent reliable knowledge but simplify multi-dimensional biology
    • Discovery often means finding exceptions and contextual dependencies
    • Virtual cells can be scoped and more rigorous than ‘digital twin’ talk
    • Analogy: weather prediction—useful without fully explaining underlying causality
  10. 22:13 – 30:09

    Industry reality: why biotech growth lags—clinical trial bottlenecks and capital intensity

    The discussion shifts to why scientific innovation hasn’t translated into stronger biotech business outcomes. Jorge highlights the hard constraint of clinical trials, the need to reduce cost/time, and why capital intensity plus weak valuation step-ups discourage early-stage investment.

    • Drug failures: wrong target vs wrong molecule (or both)
    • Clinical trials are necessary bottlenecks; some endpoints inherently take time
    • Levers: reduce capital intensity, compress timelines, increase effect sizes
    • Investment challenge: early investors bear capital burden without commensurate value inflection in valuations
  11. 30:09 – 35:49

    GLP-1s as proof of value: ambition, patient populations, and risk management

    Patrick argues GLP-1s created extraordinary market value by addressing huge patient populations, reshaping industry ambition. The high failure rate encourages ‘circling the wagons’ around validated mechanisms, but the GLP-1 example shows the upside of bold bets when impact is broad.

    • GLP-1s added enormous market cap (Lilly/Novo), dwarfing much of biotech’s aggregate value creation
    • High trial failure encourages focus on well-validated biology but smaller populations (lower EV)
    • Large-population, high-impact wins can reset incentives and ambition
    • Genetic medicines and new modalities expand what’s possible for hard diseases
  12. 35:49 – 38:15

    Future breakthroughs and the next bottlenecks: making and testing, plus regulatory friction

    Even if AI makes design abundant—‘a trillion binders in silico’—the physical world remains limiting. Patrick and Jorge emphasize that manufacturing and testing (animals → humans) will dominate, and Patrick notes how regulatory structures influence where trials happen and how fast progress can be.

    • Drug pipeline bottlenecks: designing vs making vs testing; making/testing may dominate in AI-success world
    • Testing sequence remains slow: mice → dogs → monkeys → humans
    • Regulatory regime shapes speed; discussion of running early trials overseas then returning for efficacy trials
    • Core need: increase success probability so the long journey is ‘worth it’ more often
  13. 38:15 – 42:04

    AI drug discovery: where there’s hype, hope, and real ‘heft’ today

    Patrick gives a grounded assessment of AI’s current strengths and weaknesses in biotech. He calls toxicity prediction overhyped, sees strong traction in protein-related modeling and some clinical imaging applications, and argues ‘AI-designed drugs’ will soon be a baseline capability rather than a marketing claim.

    • Hype: toxicity prediction from structure alone is still unreliable
    • Heft: protein binding and protein design; also pathology/radiology automation
    • Many ‘first AI-designed drug’ claims are marketing; AI will become a native part of the stack
    • Progress limited by feedback loops that require slow real-world testing
  14. 42:04 – 45:50

    Parallelizing discovery: Dario Amodei’s thesis and what virtual cells could enable

    Patrick reacts to predictions about rapid biomedical progress by emphasizing parallelization if discoveries are sufficiently independent. With predictive models spanning design, docking, and cellular response, discovery could increasingly resemble a compute problem—though validation remains critical.

    • Key premise: discovery tasks can be parallelized if largely independent
    • Layered model stack: design → docking → off-target analysis → cellular correction prediction
    • Potential timeline compression if steps become reliable and composable
    • Long-run plausibility paired with near-term focus on tangible, testable model capabilities
  15. 45:50 – 54:02

    Patrick’s broader investing focus: synthetic biology, BCIs, robotics, and agents

    The conversation widens to Patrick’s thesis on technologies that can meaningfully improve the human experience within a lifetime. He highlights synthetic biology, brain-computer interfaces, robotics, and AI agents—stressing the need for teams that combine technical innovation with product and commercial execution.

    • Priority areas: synthetic biology (health/longevity), BCIs, industrial & consumer robotics
    • Agents as ‘real work’ replacement—services economy opportunity beyond SaaS
    • Execution constraint: aligning technical, product, and business capabilities within teams
    • Expectation of new ML architectures beyond the 2017 transformer and renewed scaling of underexplored ideas
  16. 54:02 – 55:38

    Closing: Arc’s Virtual Cell Challenge and an open benchmark for progress

    Patrick closes by pointing to a concrete initiative: an open competition to evaluate perturbation prediction models, inspired by CASP’s role in protein folding progress. The aim is transparent measurement of capability over time and broader community participation.

    • Virtual Cell Challenge modeled after CASP for protein folding
    • Open competition with prizes and external sponsors to drive adoption and benchmarking
    • Goal: transparent tracking toward a ‘ChatGPT moment’ for virtual cells
    • Call for participation from bioML experts and engineers from other domains

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.