No Priors Ep. 103 | With Vevo Therapeutics and the Arc Institute
Summary
Tahoe 100 shifts single-cell perturbation data toward machine-learning scale. Vivo reports 100 million single-cell measurements across 50 cancer models from different patients, 1,200 drug treatments, and 60,000 drug–cell or drug–cell-line interactions, versus roughly 1–2 million publicly available perturbational data points before it. Dave Burke compares the hoped-for impact to an “ImageNet moment” for cellular biology.
The differentiated asset is causal, diverse, consistently generated data—not another pile of healthy-cell observations. Perturbations provide a before-and-after signal, while disease models expose states absent from normal tissue. Four people performed consistent work across 60,000 experiments with effectively no batch effects. The team argues that scale and breadth can make the data itself generate hypotheses and surprises.
Virtual cells address the systems-level question that protein models leave open. Protein language models can learn folding, binding, and structural biology, but a drug acts on a protein embedded in a cell, tumor, and broader system. The proposed virtual cell is a “notional CPU” that predicts how a genetic edit or chemical changes the transcriptome—and eventually solves the inverse problem of moving a diseased state toward health.
Current virtual-cell models remain commercially immature, with predictive performance on differentially expressed genes on the order of 10% and no accepted benchmark. Tahoe 100 represents roughly 200–300 billion training tokens under current encodings. Dave describes roughly 1 trillion tokens as a comfortable comparison point from other domains, while stressing that biology’s scaling laws remain uncertain. Hani Goodarzi places protein models beyond GPT-3 maturity but cellular models “closer to GPT-1 than 2.”
Open sourcing is an operating strategy, not philanthropy detached from Vivo’s economics. Tahoe 100 joins Arc’s roughly 230 million-cell scBaseCamp collection to form a 330 million-cell Virtual Cell Atlas; the community can train and critique models while Vivo continues building its Mosaic data-generation platform. Nima’s framing is “a new stake in the ground”: a small internal team can recruit an external research ecosystem without building a large organization.
If virtual cells work, the potential prize is simultaneous compression of target risk, chemical search, and experimental time. Roughly 90% of drugs fail in clinical trials; the guests argue failures reflect both poor drug matter and selection of the wrong target. A sufficiently accurate model could search tens of millions of compounds, generalize across patient contexts, and help teams “measure twice and cut once”—but at 10% accuracy, Patrick Hsu cautions, “you’re just simulating noise.”
The technical inflection may be now, but therapeutic proof will remain slow. Chinese biotechs’ cost, pace, antibody manufacturing, and IND-enabling packages are pushing US companies toward leaner teams and stronger vendors; Nima calls it “morning in bio” and urges building now rather than promising results in three to five years. Sarah Guo stresses that treatments commonly require 11-plus years, and even moving success from 10% to 30% leaves a “law of small numbers” that may take a decade-long window to validate.
Deep dive
1. Tahoe 100 makes perturbational biology a machine-learning-scale resource
Johnny Yu defines Tahoe 100 as the world’s largest single-cell RNA-sequencing dataset: 100 million cells spanning 50 cancer models from different patients, 1,200 drug treatments, and 60,000 drug–cell or drug–cell-line interactions. His strongest claim is that it may be “the first data set that’s going to enable machine learning in this space.”
The relevant comparison is not merely total cells. Publicly available perturbational data previously amounted to roughly 1–2 million single-cell points, while most historical datasets were observational, fragmented, poorly labeled, and drawn from healthy tissue rather than disease states.
Dave Burke’s analogy is ImageNet in 2009: a sufficiently large, purposeful dataset can produce a nonlinear capability jump. The hope—explicitly still a hope—is that cellular perturbations do for systems biology what foundational protein-structure datasets, including PDB and CASP, helped enable for models such as AlphaFold.
2. A virtual cell models causality beyond protein structure
Nima Alidoust distinguishes two languages: protein models learn “the language of structural biology”—folding, molecular binding, and antibody–protein interaction—while virtual cells aim at “the language of systems biology,” where a target sits inside a cell, tumor, immune environment, and broader organism.
Dave’s computing analogy makes the abstraction concrete: DNA is ROM, encoding the cell; RNA is RAM, its changing working memory; and a virtual-cell model infers a “notional CPU” that maps an edited gene or applied drug into a transcriptomic response.
The inverse problem is the therapeutic one: given a diseased cell’s expression profile, which genetic or chemical perturbation might return it toward a healthy state—or, for cancer, kill it while sparing healthy cells? That capability remains an intended destination, not a demonstrated product.
Hani Goodarzi’s case for perturbational data is causality. Observing a liver sample yields associations; intervening genetically or chemically creates a defined before and after, allowing a model to learn changes that drive cell state rather than merely accompany it.
3. Information content matters more than another redundant million cells
Nima reports that early single-cell foundation models could discard 99% of a roughly 60-million-cell collection with little performance loss. His inference is that much of the available data repeated similar biological contexts, so raw scale overstated how much the model could learn.
Diversity supplies the missing information. A model intended to reason about heart, brain, liver, bone, or cancer needs normal and diseased states across cell and tissue types; perturbations help explore the high-dimensional “manifold” rather than repeatedly sampling one neighborhood.
Each cell contributes roughly 2,000–5,000 gene-and-expression tokens, making Tahoe 100 approximately 200–300 billion tokens under the team’s estimate. Dave cites about half a trillion tokens for GPT-3 and 700 billion for ESM-3, with roughly 1 trillion as a comfortable comparison point, but says the field will not know the relevant scaling law until it gets there.
The bottleneck also varies by domain. Decades of genome sequencing make DNA models increasingly compute- and context-length-limited; cellular models remain data-limited because scalable single-cell profiling emerged only recently.
4. Mosaic replaces one-patient-at-a-time screening with pooled experiments
Vivo’s Mosaic platform pools cells from many patient-derived models—including multiple cancer types and distinct genetics—into one “mosaic tumor,” then screens hundreds or thousands of drugs while resolving each model’s response. Sarah’s reaction captures the jump from serial experiments to multiplexed biology.
Initial perturbations target cancer-relevant genes, growth, DNA regulation, and pan-cancer pathways. Hani argues that these conserved pathways may broadly apply to neuroscience and immune-cell development, while future datasets could add rare diseases and use model feedback to fill gaps.
Scale changes experimental philosophy as well as throughput. Nima says Vivo can generate about 50 times more perturbational data in five weeks than was publicly available from the previous decade, reducing the need to preselect a narrow hypothesis: teams can go large on chemical and patient-sample space and become more unbiased.
5. Open data lets a tiny company borrow the field’s intelligence
Nima says Vivo chose to open-source the data within hours of discussing the opportunity for two reasons: to put “a new stake in the ground” for expected dataset scale, and to keep an internal team of roughly three or four people focused while recruiting outside researchers to expose strengths, defects, and model opportunities.
Arc’s Virtual Cell Atlas combines Tahoe 100 with scBaseCamp, an approximately 230 million-cell observational collection, for roughly 330 million cells. The proposed workflow is complementary: one could potentially pretrain on broad observations, then add perturbations to teach dynamics and improve prediction.
Arc built an agent likened to “a Google crawler and index” to find, annotate, and uniformly reprocess public sequencing data. That matters because changing tools, tool versions, genome builds, and reagent chemistries otherwise introduce analytical and experimental batch effects into merged datasets.
Operational consistency is itself part of the asset: Dave says four people performed exactly consistent work across 60,000 experiments. Fewer hands reduce the familiar biological caveat that a result works only “in my hands.”
6. Accurate simulation could compress biological time and chemical search
Patrick’s aging-project anecdote illustrates the constraint: one experimental round could require aging animals for two years. In-silico parallelism would be transformative, but he draws a hard line—if a virtual cell is only 10% accurate, “you’re just simulating noise.”
Vivo’s intended application is to predict how a new chemical entity interacts with cells from different patient models, asking whether it can move a diseased cell toward health—or, in cancer, kill it without killing healthy cells. Johnny’s future vision is a drug generated from a virtual-cell model, not a demonstrated capability today.
Nima separates two generalization problems. Across cells, models must transfer from observed patients to new individual variation; across chemistry, they must navigate tens of millions of candidate compounds and effectively unbounded biologics, identifying the small region worth synthesizing rather than screening familiar libraries incrementally.
With about 90% of drugs failing in clinical trials, Patrick argues the industry may have both bad drug matter—potency, toxicity, pH, kinetic profiles, and related properties—and the wrong targets. Virtual cells could narrow the target search before expensive chemistry and development begin.
7. The transcriptome is the abstraction, while cellular context remains visible
Dave describes RNA expression as a 1980s graphic equalizer with roughly 20,000 moving bars, responding to environment, stress, aging, health, and disease. Patrick says the group views the transcriptomic layer as rich enough to capture consequential state without simulating every molecular detail of the cell.
Multicellular modeling can then ladder upward through spheroids, organoids, in-vivo models, immune environments, and spatial data. Hani’s nuance is that environmental information is “filtered through the cell”; observing enough contexts may let a nominally single-cell model infer effects produced by surrounding tissue.
8. Platform companies are designed to abandon bad hypotheses
Nima contrasts Vivo with a single-hypothesis biotech, whose organization and incentives become committed to making one thesis work—sometimes advancing a drug after testing it on only three patient samples. A platform can generate competing hypotheses and remain “a lot more scientific” about what reaches the clinic.
Chinese biotech sharpens the operating challenge. Patrick points to lower cost bases, faster pipelines, strong safety and toxicology packages, IND-enabling work, and efficient antibody manufacturing; he sees that competition as potentially beneficial because patients, investors, and companies all want working molecules faster and cheaper.
Neither a vendor-chained “virtual biotech” model nor full vertical integration has solved the problem cleanly: the former proved slow, while owning everything proved expensive and bureaucratic. The proposed middle is capable vendors plus lean companies, with Nima arguing that the industry should incorporate these capabilities rather than rely solely on regulatory limits.
Nima’s “morning in bio” manifesto has three parts: build now rather than announce delivery in three to five years; organize small, concentrated teams of “superstars”; and challenge domain experts’ long lists of reasons cross-domain machine-learning techniques cannot work. His hedge remains intact: “Maybe it works, maybe it doesn’t work—but if you don’t try, you’ll never know.”
9. The technical inflection may be now, but therapeutic proof will lag
Dave reaches back to neural networks before ImageNet and AlexNet: ideas can appear unproductive until compute, data, and model architecture cross an inflection. Single-cell resolution, scalable perturbations, and modern training methods are the biological equivalents the group believes may now be converging.
Evo 2 is offered as evidence of emergent biological learning: Arc trained it on 9.3 trillion nucleotides without explicit DNA instruction, yet it learned ribosome-binding sites, codon degeneracy, and zero-shot BRCA1 variant pathogenicity with an area under the ROC curve of around 94%.
Maturity differs sharply by domain. Hani places protein language models beyond GPT-3, while Patrick and Dave describe cellular biology as roughly GPT-2-level or still developing toward it; Hani ultimately places virtual-cell models “closer to GPT-1 than GPT-2.” Existing best models predict differentially expressed genes with performance on the order of 10%, and no accepted benchmark yet settles comparisons.
Sarah notes her decade of skepticism about AI biotech and stresses that GPT-4-like cellular capability would not be instantly obvious: treatments generally require 11 years or more, and moving success from 10% to 30% would be extraordinary but still subject to the “law of small numbers.” Dave says better models could help “point the cannon in the right direction,” but proof would accumulate slowly over a roughly 10-year development window.