Skip to content
No PriorsNo Priors

No Priors Ep. 103 | With Vevo Therapeutics and the Arc Institute

On this week’s episode of No Priors, Sarah Guo is joined by leading members of the teams at Vevo Therapeutics and the Arc Institute – Nima Alidoust, CEO/Co-Founder at Vevo Therapeutics; Johnny Yu, CSO/Co-Founder at Vevo Therapeutics; Patrick Hsu, CEO/Co-Founder at Arc Institute; Dave Burke, CTO at Arc Institute; and Hani Goodarzi, Core Investigator at Arc Institute. Predicting protein structure (AlphaFold 3, Chai-1, Evo 2) was a big AI/biology breakthrough. The next big leap is modeling entire human cells—how they behave in disease, or how they respond to new therapeutics. The same way LLMs needed enormous text corpora to become truly powerful, Virtual Cell Models need massive, high-quality cellular datasets to train on. In this episode, the teams discuss the groundbreaking release of the Tahoe-100M single cell dataset, Arc Atlas, and how these advancements could transform drug discovery. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @Nalidoust | @IAmJohnnyYu | @PDHsh | @Davey_Burke | @Genophoria Show Notes: 0:00 Introduction 1:40 Significance of Tahoe-100M dataset 4:22 Where we are with virtual cell models and protein language models 10:26 Significance of perturbational data 17:39 Challenges and innovations in data collection 24:42 Open sourcing and community collaboration 33:51 Predictive ability and importance of virtual cell models 35:27 Drug discovery and virtual cell models 44:27 Platform vs. single hypothesis companies 46:05 Rise of Chinese biotechs 51:36 AI in drug discovery

Sarah GuohostJohnny YuguestNima AlidoustguestPatrick HsuguestDave BurkeguestHani GoodarziguestElad Gilhost
Feb 25, 202557mWatch on YouTube ↗

CHAPTERS

  1. 0:05 – 1:31

    Tahoe-100 launch: who’s involved and what the episode will cover

    Sarah Guo introduces leaders from Vevo Therapeutics and the Arc Institute and frames the conversation around the Tahoe-100 release. They set the agenda: why this dataset matters, what “virtual cell” models are, and when AI-driven biology might yield real treatments.

    • Introductions: Vevo (Johnny Yu, Nima Alidoust) and Arc (Patrick Hsu, Dave Burke, Hani Goodarzi)
    • Episode focus: Tahoe-100 dataset + virtual cell modeling
    • Motivation: move beyond protein-only AI to cellular/system-level models
    • Goal: accelerate drug discovery and disease understanding
  2. 1:31 – 4:18

    What Tahoe-100 is and why it’s an ImageNet-like milestone for cell biology

    The guests define Tahoe-100 as the largest single-cell drug-perturbation dataset to date and argue it can catalyze a step-change in AI-for-biology. They compare it to foundational datasets like ImageNet that enabled rapid progress in other AI domains.

    • Tahoe-100 described as the world’s biggest single-cell RNA-seq drug-perturbed dataset
    • Potential to enable ‘virtual cell’ modeling and new drug discovery workflows
    • Analogy to ImageNet: datasets as inflection points for model capability
    • Need to model biology at the cellular level, not just proteins
  3. 4:18 – 10:20

    Why virtual cell models complement protein language/structure models

    Arc and Vevo explain why protein structure prediction and protein language models are not sufficient to predict phenotypes and drug responses. They outline a systems view: cells are dynamic, context-dependent programs whose state is reflected in gene expression.

    • Protein models answer binding/structure questions; cells answer response/context questions
    • Cell-as-computer analogy: DNA as ROM, RNA as RAM, model as “CPU” mapping inputs to outputs
    • Virtual cell enables inverse design: choose perturbations to push diseased cells toward healthy states
    • Different constraints by domain: some areas compute-limited (DNA), others data-limited (cell state)
  4. 10:20 – 10:26

    Why perturbational single-cell data matters: from correlation to causation

    They argue observational atlases are mostly descriptive and often miss causal relationships and disease dynamics. Perturbations (chemical or genetic) expose how cell states move, helping models learn a generalizable latent manifold of responses.

    • Perturbations create clear before/after signals to infer causality
    • Models need diverse perturbations to learn the ‘manifold’ of cell states
    • Prior public single-cell data: mostly healthy, mostly observational
    • Perturbational public data was tiny (≈1–2M cells) compared to Tahoe-100 (100M)
  5. 10:26 – 13:33

    What’s wrong with prior data, and what Tahoe-100 adds (scale, diversity, low batch effects)

    The conversation zooms into practical data issues: fragmentation, poor labels, and batch effects make many public datasets hard to use for foundation modeling. Tahoe-100 is positioned as both a scale jump and a quality jump, spanning many patients and drug treatments with consistent processing.

    • Academic datasets are small, inconsistent, and batchy; batch effects can dominate signal
    • Tahoe-100 doubles cumulative public scale and is designed to minimize batch effects
    • Coverage: ~50 patient-derived cancer models and ~1,200 drug treatments
    • Information content depends on context diversity, not just raw cell counts
  6. 13:33 – 15:17

    Do scaling laws apply? Translating “tokens” to cells and gene expression

    They discuss whether 100M–330M cells is ‘enough’ and how to think about scaling. The team maps single-cell measurements to an LLM-style token budget by treating expressed genes as tokens, arguing this gets them into the hundreds-of-billions token regime—close to where scaling analyses become meaningful.

    • Hard to know sufficiency upfront; guidance comes from LLM/DNA model scaling experience
    • A ‘token’ analogy: genes (and expression levels) as tokens per cell (≈2k–5k)
    • 100M single-cell profiles can correspond to ~200–300B token-like units
    • Key caveat: not all tokens are equally informative; perturbations increase information density
  7. 15:17 – 24:59

    Choosing perturbations and the “mosaic” platform for scalable, hypothesis-light experiments

    Vevo explains how they select perturbations for cancer relevance while preserving broad biological utility. They describe the mosaic tumor approach—pooling cells from many patients—allowing thousands of drugs to be screened across diverse genetic backgrounds efficiently, shifting biology toward more unbiased exploration.

    • Perturbation selection: match tools (drugs/targets) to disease biology (e.g., cancer pathways)
    • Mosaic platform: pool cells from many patients into one screenable system
    • Scales from ‘one model at a time’ to tens/hundreds simultaneously
    • Shift toward hypothesis-free generation as costs fall and throughput rises
  8. 24:59 – 30:00

    Open-sourcing Tahoe-100 and launching the Arc Virtual Cell Atlas (plus SC Basecamp)

    They explain why a venture-backed company would release the dataset publicly: to set a new standard and enlist the broader community to build better models. Arc details the complementary Arc Virtual Cell Atlas initiative, including SC Basecamp—an agent-driven crawler/curation pipeline for assembling hundreds of millions of observational single-cell profiles.

    • Reasons to open-source: raise the bar, catalyze community progress, and remove data bottlenecks
    • Open sourcing supports a ‘small team of superstars’ model by leveraging external contributors
    • Arc Virtual Cell Atlas: curated resources to accelerate virtual cell modeling
    • SC Basecamp: automated crawling/standardization of public single-cell data (~230M cells)
  9. 30:00 – 32:41

    Data collection innovations: consistency, ‘few hands,’ and removing analytic batch effects

    They highlight how Tahoe-100’s experimental consistency reduces variability common in biology (‘in our hands’ effects). Arc discusses why reprocessing raw data matters: tools, genome builds, and pipelines change over time, so standardized reanalysis removes a major source of contamination in aggregated datasets.

    • Tahoe-100 produced with unusually few operators to reduce experimental variance
    • Scale note: ~60,000 drug–patient (or drug–model) interactions underlying 100M cells
    • SC Basecamp reprocesses raw data to remove analytic/pipeline-induced batch effects
    • Uniform processing aims to create a reliable foundation dataset for the field
  10. 32:41 – 34:10

    How to judge virtual cell models: predictive accuracy, benchmarks, and today’s limitations

    The group turns to evaluation: a virtual cell model should predict gene expression changes after perturbation. They note current performance is poor (around ~10% on differentially expressed genes) and argue the field needs shared benchmarks; the new data is framed as a prerequisite for meaningful improvements.

    • Core metric: ability to predict differentially expressed genes after perturbation
    • Current best models are weak; predictive performance around ~10% in practice
    • No widely accepted benchmark today; standardization would benefit the industry
    • Hypothesis: poor results are driven heavily by data quality and inconsistency
  11. 34:10 – 41:37

    Why virtual cells matter for drug discovery: speed, target choice, and traversing huge chemical space

    They argue virtual cell models are valuable because biology is slow and expensive; accurate in silico prediction could massively parallelize experimentation. The discussion connects model generalization across patient contexts and chemical libraries to reducing failure rates by improving target selection and candidate prioritization.

    • Biology runs on ‘biological time’; virtual cells enable fast, parallel exploration
    • Goal: predict how new chemical entities shift diseased cell state (or kill cancer selectively)
    • Two generalization axes: across cell contexts/patients and across chemical space
    • Drug failure rate (~90%) reflects wrong targets and/or inadequate drug matter; virtual cells can reduce search space
  12. 41:37 – 44:15

    Beyond single-cell: right abstraction level, context, organoids, and spatial extensions

    They address whether single-cell is enough and explain their view on abstraction: transcriptomics is a practical modeling layer that captures many pathway-level effects. They describe how organoids/spheroids and in vivo contexts can be incorporated, and note spatial data as an additive modality.

    • Choosing abstraction: transcriptome as a high-signal readout of cell state changes
    • Extending to organoids/spheroids: perturb and measure many cells to capture multicellular dynamics
    • Context is ‘filtered through the cell’; diverse conditions teach environmental effects
    • Future augmentation: spatial measurements and richer multi-modal data
  13. 44:15 – 52:21

    Company-building and industry shifts: platform vs single-hypothesis, China’s rise, and where AI helps most

    In closing hot takes, they contrast platform companies with single-hypothesis biotechs, arguing platforms enable more rigorous, less biased selection of what reaches the clinic. They discuss Chinese biotechs’ cost and speed advantages, the need for better operating models in the US ecosystem, and how AI can reduce costs both in discovery and downstream clinical workflows.

    • Platform vs single-hypothesis: avoid being ‘wedded’ to one idea; generate many hypotheses and select rigorously
    • China’s biotech efficiency pressures US cost bases; competition can benefit patients
    • Operational challenge: virtual biotech can be slow; full vertical integration is expensive—need a middle path
    • Near-term AI value: discovery (better targets) plus clinical development workflows (docs, cohorts, filings)
  14. 52:21 – 57:40

    Why now: inflection-point analogy, ‘GPT-1/2 for cells,’ and the long proof horizon in medicine

    They respond to skepticism about AI in biotech by drawing parallels to earlier AI winters and ImageNet-driven takeoff. The guests place biology at an early LLM-era stage (roughly GPT-1 to GPT-2 for virtual cells, further along for proteins) and emphasize that clinical timelines and small-number statistics mean validation will lag capability improvements.

    • AI progress historically needs compute + data + model advances to hit non-linear inflections
    • Evo DNA models show strong zero-shot learning signals; Tahoe aims to do similar for cells
    • Consensus: protein models are ahead (post-GPT-3-ish), cell state models are earlier (GPT-1/2)
    • Even big lifts in success rate take years to demonstrate due to 10+ year drug timelines

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.