Pioneers Insight Method Research Author
🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Back to Episodes

🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"

· Source link · AI Summary archive

Summary

  • Genesis Molecular AI argues that the frontier of foundational AI research has shifted from familiar LLM architectures toward diffusion models for 3D molecular structure. GANs failed on proteins and protein–ligand systems, while diffusion supplied “the right primitive” for iteratively generating physical structures. The hosts’ sharper recruiting pitch is that LLM labs still largely rearrange transformer layers published in 2017, whereas “some of the most innovative diffusion research” now happens in structure prediction.

  • Genesis presents roughly 1 Å accuracy as the useful threshold for protein–ligand prediction, because the conventional 2 Å scale can conceal chemically fatal errors. At 2 Å, an aromatic ring may flip while the output still looks plausible; hydrogen bonds occupy only a 2.7–3.3 Å distance window. Evan Feinberg’s formulation is blunt: “Drug discovery really is a science of resolution,” and models at 1.8–1.9 RMSD risk producing agent-amplified “slop.”

  • Genesis has adapted parts of the LLM scaling stack to molecules: synthetic-data pre-training, iterative inference-time computation, and eventually reinforcement learning. The public structural corpus contains only roughly 200,000 entries, versus an estimated 10^60 drug-like small molecules, so physics simulations supply additional training data. During inference, the models “think in terms of crystal structures,” repeatedly refining an internal representation while physics-based guidance steers diffusion.

  • The highest-leverage AI opportunity, in Genesis’s view, is the missing design layer between known disease biology and clinical testing. Knowing the responsible target is orthogonal—and sometimes inversely related—to being able to drug it; Feinberg calls the blanket 10% clinical-success statistic “really a lowball” for candidates with strong genetics, pharmacokinetics, safety, and translational models. The opportunity spans zero-to-one binders and better successors to existing drugs, as later-generation ALK inhibitors demonstrate.

  • A viable drug requires simultaneous optimization across more than 30 ADMET-related endpoints, not merely an accurate binding pose. Potency, selectivity, solubility, membrane permeability, cytochrome P450 inhibition, hERG liability, tissue exposure, and other properties can invalidate a molecule independently. Worse, objectives anti-correlate: making a compound greasier may improve binding while damaging solubility, and adding polarity may then prevent cellular entry—“playing whack-a-mole” at molecular scale.

  • Genesis’s prospective-data advantage comes from coupling its models to real drug programs and laboratory feedback: Incyte provides disclosed partner programs, while Insitro supplies rapid compound production and measurements. Disclosed Incyte work ranges from advancing existing chemical matter toward a development candidate to finding the first known binders for a target with no patents, papers, or co-crystal structure. The Insitro collaboration creates repeated design–make–test–analyze cycles and training data spanning structure, potency, and ADMET, although synthesis and high-fidelity validation remain stubbornly difficult to automate.

  • Sapphire is Genesis’s attempt to turn specialist models into an always-on drug-discovery workforce without removing scientists from strategic control. An LLM orchestrates pose, potency, ADMET, and chemistry tools so medicinal chemists need not master every parameter; the envisioned result is “fleets of hundreds” of virtual scientists operating 24/7. Feinberg rejects full human replacement: experts set direction and evaluate outcomes while agents absorb execution and repetitive tool use.

  • OpenBind supplied the external generalization test Genesis says private partner data had previously prevented it from showing. On the unseen EV-A71 3C protease target, whose flexible loop must move around the ligand, Pearl reportedly produced a much wider performance gap than public in-distribution benchmarks and was “basically correct for every single pose.” The remaining constraint is compute: both guests named GPUs as their bottleneck, while Edunov argued that “the amount of alpha left in pure LLM space is just getting a little questionable” relative to life sciences.

Deep dive

1. Diffusion became the primitive molecular AI had been missing

  • Edunov’s route to Genesis ran from physics into software engineering, FAIR research, and leadership of Llama 2 and Llama 3 pre-training. Joining as CTO let him “recover my roots in physics” while applying large-scale AI methods to molecular systems.

  • Feinberg came from the complementary direction: a physics-and-computer-science background, a medical family, and graph-machine-learning work in VJ Pond’s Stanford lab. While Edunov studied “a lot of big graphs,” Feinberg studied molecules as “many small graphs” of atoms, bonds, and spatial interactions.

  • Feinberg remembers declaring GANs the future of image generation around 2017–2018, then watching mode collapse make them ineffective for proteins and protein–ligand complexes. The field had to “wait for the right primitive,” and diffusion proved much more useful for generating three-dimensional structures.

  • The surprising result is that foundational diffusion research is no longer concentrated in consumer image generation. Feinberg’s claim: “Some of the most innovative diffusion research is happening in our field,” making 3D structure prediction a pillar nobody would have forecast a decade earlier.

2. Drug-discovery AI is compounding, not awaiting an iPhone moment

  • When Genesis began roughly seven years earlier, its founders worried they were late: incumbents had raised orders of magnitude more capital, and investors questioned whether another AI drug-discovery company was needed. Feinberg now calls that hindsight almost absurd—“thought to be late, but turns out it was still early innings.”

  • His original thesis remains intact: roughly 20,000 protein-coding genes can contribute to disease, and no single advance can solve that universe. “There’s been no single iPhone moment”; even smartphones, ChatGPT, and autonomous vehicles became useful through repeated improvements rather than one clean zero-to-one event.

  • The expectation is continuing expansion of solvable targets, with “large leaps” embedded in cumulative iteration. Feinberg says current systems are vastly more useful than those of a decade ago and predicts another exponential improvement during the next ten years.

  • The host’s central challenge was generalization: molecular models historically became pattern matchers that told researchers what they already knew. Feinberg agreed this is the urgency wherever AI meets the physical world—models naturally interpolate near training data, while drug discovery demands reliable extrapolation.

3. Molecular design is the missing middle between biology and trials

  • Feinberg’s working analogy casts the protein as a lock and the drug as a key intended to change its function. Binding is “necessary but not sufficient”: the molecule must avoid anti-targets, reach the correct tissue, remain safe, and satisfy roughly 30 additional developability properties.

  • A decade ago, researchers hypothesized that accurate 3D protein–ligand complexes would improve affinity and potency prediction. They could barely test it: computational pose predictions were poor, while experimental crystallography or cryo-EM could cost tens of thousands of dollars and consume months, years, or an entire thesis.

  • Genesis says the recent breakthrough is not merely generating attractive structures but showing that systematic pose-accuracy improvements carry into potency prediction. Better complexes could also expose druggable configurations in proteins previously considered undruggable.

  • Target identification, molecular design, regulatory preparation, patient segmentation, and clinical trials require distinct models despite sharing some machinery. Feinberg places the highest leverage in design: patients are often told clinicians know what caused their condition but still lack a selective therapy capable of acting on it.

4. Pearl attacks a 10^60 search space with scarce structural data

  • Pearl accepts a protein sequence and ligand representation, then predicts their joint 3D complex—a co-folding task. Genesis deliberately concentrates on small and medium-sized molecules, including orally available drugs, macrocycles, and peptides, rather than treating every biological interaction as one problem.

  • “Small” does not mean computationally easy. Edunov estimates about 10^60 drug-like small molecules, each admitting rotations and alternative conformations; after the hosts proposed finding a needle in a haystack, Feinberg inverted it to “finding hay in a needle stack,” where most candidates either fail or are dangerous.

  • The nearest molecular equivalent to internet-scale pre-training is the RCSB Protein Data Bank, with only a couple hundred thousand structures, though each contains substantial latent information. New experimental structures arrive at a “glacial pace” because they are expensive and difficult to produce.

  • Genesis supplements that corpus by simulating small molecules with physics. Protein–protein systems are much larger and costlier to model, but molecular dynamics and related calculations can generate lower-cost structural data for small-molecule pre-training—provided their biases are handled carefully.

5. Pearl imports scaling laws without copying language models literally

  • Edunov maps the LLM recipe into three stages: pre-training scaling, post-training through fine-tuning or reinforcement learning, and inference-time scaling. Genesis creates synthetic pre-training data and then lets its models spend additional inference computation rather than immediately emitting one structure.

  • The analogy to reasoning tokens is functional, not linguistic. The model is “thinking in terms of crystal structures”—partially materialized internal representations that it revisits as the diffusion head iteratively refines predicted coordinates.

  • Because diffusion already unfolds across multiple denoising steps, Genesis can inject physics-based guidance during generation. The hosts framed this as balancing learned and physical force fields; Edunov accepted the steering intuition but declined to claim that researchers know exactly what the network internally represents.

  • Feinberg treats physical priors as disciplined representation choices, analogous to encoding images as pixel grids or language as token sequences. Genesis uses physics in inputs, architecture, and output validation while trying not to force models to inherit every assumption held by “we puny humans.”

6. One angstrom separates plausible pictures from useful chemistry

  • The discussion notes that at 2 Å, a structure is like a generative image whose detail is wrong—but unlike visible blur, the molecular error may look entirely credible. A whole aromatic or heterocyclic ring can flip while still passing a coarse structural metric.

  • Feinberg grounds the 1 Å objective in hydrogen bonds: donor-to-acceptor heavy-atom distances typically span 2.7–3.3 Å, only a 0.6 Å window. Too short is a clash; too long rapidly weakens the interaction. “Drug discovery really is a science of resolution.”

  • The hosts offered a serious counterpoint: a single pose is an abstraction over a probability distribution, and affinity also includes enthalpic, entropic, and dynamical contributions. Feinberg accepted the abstraction but defended its inspectability—a scalar potency output “might as well be completely hallucinated” if no structure lets scientists test whether it makes sense.

  • Feinberg distinguished a potent ligand’s tightly resolved binding core from solvent-exposed regions that may genuinely “flop around.” The objective is sub-angstrom correctness where protein interactions occur, while accepting dynamics elsewhere; that core then supports free-energy prediction and the practical question, “What molecule do I make next?”

7. The field’s two-angstrom benchmark created an eval crisis

  • Asked how Genesis crossed the threshold, Feinberg gave “an extremely boring answer”: “data, infrastructure and evals.” Because “you can only improve what you measure,” optimizing sub-angstrom accuracy changes filtering, curriculum, architecture, losses, and many small engineering decisions that compound.

  • Real partner and internal programs made 2 Å failures obvious in a way academic leaderboards did not. Feinberg traced RMSD below two to old docking studies built for publication and later inherited by AI, not to medicinal chemists establishing that it was sufficient for prospective design.

  • His SWE-bench analogy was intentionally provocative: a model can score well without becoming anyone’s preferred coding tool. Molecular evaluation is likewise moving beyond RMSD toward physical validity, PoseBusters, and lDDT; Feinberg describes “an eval crisis in our field that is now in transition.”

8. Structure prediction is only one pillar of a 30-property problem

  • Feinberg pushed back on popular claims that AlphaFold-era structure prediction solved drug discovery. A static, relatively low-resolution structure omits dynamics, selectivity, exposure, safety, and the other endpoints required to turn a binder into a medicine.

  • ADMET is represented operationally by more than 30 assays. Feinberg cited solubility, oral bioavailability, cytochrome P450 variants, and hERG inhibition, where an excessive effect can create cardiotoxicity.

  • These endpoints vary in learnability. Some correspond directly to interaction with a particular protein and might be addressable through 3D modeling; others aggregate many pathways, while public datasets can be “comically small.”

  • Genesis’s breadth predates Pearl: its Stanford lineage produced MoleculeNet and multitask graph networks for pharma-scale ADMET prediction. Feinberg’s emphasis is continuity—the company has focused on all models required for drug discovery, including molecular generation, without expanding into unrelated target-identification or clinical-trial problems.

9. Known biology leaves both first-in-class and best-in-class opportunity

  • Shawn Wang’s business challenge was whether targets with known biology and tractable structure had already been picked over. Feinberg separated the variables: biological validation is orthogonal to ease of drugging and may appear anti-correlated because the most compelling disease targets can be exceptionally hard to bind selectively.

  • Feinberg also disputed the undifferentiated 10% clinical-success statistic. Candidates with close genetic linkage, understood biology, translating animal models, adequate predicted pharmacokinetics, and strong safety profiles have “fairly high” approval rates from Phase 1 through Phase 3—though he supplied no universal figure.

  • Opportunity therefore spans true zero-to-one programs with no known binder and one-to-ten programs improving imperfect chemical matter. His public analogy was ALK inhibitors: later generations produced qualitatively better survival curves, showing why “we’ve drugged ALK” did not mean development should stop.

  • Genesis focuses on small and medium-sized molecules; small molecules alone remain about 65% of FDA-approved drugs, Feinberg said. The intended market is consequently both new target access and replacement of suboptimal clinical or approved agents.

10. Partner data turns models into prospective drug programs

  • Genesis disclosed work with companies including Gilead and an expanded Incyte collaboration. One Incyte program began with a challenging target and existing chemical matter; Genesis fine-tuned foundation models on partner data to move the program closer to the binary milestone of selecting a development candidate.

  • The opposite bookend began with strong disease linkage but no known chemical matter—no patent, paper, or ligand co-crystal. The teams found initial hits, then advanced them into inhibitors active in biochemical assays and living-cell assays.

  • The commercial model pairs Genesis’s AI specialization with pharma’s strengths in biology, development, and commercialization. Renaming Genesis Therapeutics as Genesis Molecular AI reflected that identity rather than abandoning medicines; internal programs still dogfood the models, generate candid medicinal-chemist feedback, and make the platform “battle tested.”

  • Feinberg described the organization as a “double helix”: deep AI research alongside experienced drug hunters, some of whom have helped produce multiple approved medicines. The goal is to place models directly with many drug developers while retaining enough in-house discovery to understand the work rather than behave like “keyboard jockeys.”

11. Wet-lab feedback is valuable precisely because automation remains messy

  • Feinberg’s near-term reinforcement-learning path begins with physics-based rewards and could extend into laboratory-in-the-loop rollouts: predict molecules, synthesize them, measure downstream properties, and feed those outcomes back into training. He said Genesis has already seen early signs that RL works with its models.

  • Insitro’s rapid compound production and measurement enable continuous design–make–test–analyze cycles across structure, potency, and ADMET. Feinberg called the collaboration unusually important because it combines historical and prospective data for joint foundation-model training.

  • The physical work resists simplistic robotic-lab narratives. Synthesis requires compatible reagents, catalysts, solvents, temperatures, and protocols; compounds then need purification and confirmation through NMR, mass spectrometry, or related methods to prove the vial contains what researchers intended.

  • High-throughput screens and DNA-encoded libraries can test millions or billions of compounds, yet their correlation with de novo resynthesis and high-fidelity assays may have “a shockingly low R squared.” Automation also favors constrained chemistry, while drug discovery searches for novel Pareto outliers among anti-correlated objectives—speed can exact a harsh cost in molecular quality.

12. Agents can scale drug hunters only after the underlying models work

  • Feinberg compared molecular agents with coding agents: both amplify positive and negative value, so “agents are only as useful as the underlying models that they’re orchestrating.” Pose systems around 1.8–1.9 RMSD would merely automate production of structures medicinal chemists reject as “slop.”

  • Sapphire, Genesis’s code name for an agentic platform, is envisioned as “fleets of hundreds” of medicinal chemists and CADD scientists working 24/7. Its LLM orchestrator can call pose, potency, ADME, and chemistry tools while handling parameter choices that no human can master across an entire software stack.

  • Crystal structure can become another model modality: an agent might inspect an image, invoke specialized geometric tools, or consume a natively tokenized 3D representation. That lets it reason from structural predictions rather than treating each predictor as an opaque scalar oracle.

  • Feinberg expects the Cursor trajectory—first autocomplete, then increasingly autonomous execution—but “I don’t believe in full automation or replacing humans.” Scientists remain strategic directors, while agents absorb routine execution; his version is that drug hunters become “grand strategists” with hundreds of computational workers.

13. OpenBind exposed generalization; GPUs now constrain the upside

  • Public benchmarks often compress model performance into a narrow band because every team has optimized against them. OpenBind’s unseen EV-A71 3C protease target offered a cleaner test: Pearl had not trained on or been developed around it, and the target’s flexible loop must move to accommodate the ligand.

  • Feinberg said Pearl’s numbers were “way higher” than published open models and that it was “basically correct for every single pose” in modeling the loop movement. Genesis presented this as public confirmation of the larger gaps it says it observes on confidential partner targets.

  • Both guests named GPUs as the bottleneck. Edunov argued LLM companies are consuming capacity needed for medicine discovery; Feinberg’s ideal intervention would be an enormous H100 cluster, while noting NVIDIA has invested twice in Genesis and collaborated on kernels and Pearl.

  • Edunov’s allocation thesis is that medicines remain valuable through economic cycles, whereas “the amount of alpha left in pure LLM space is just getting a little questionable.” The hosts contrasted mainstream LLM architectures, still closely related to the transformer published in 2017, with molecular diffusion as a more architecturally distinct field with direct clinical stakes.