🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Summary
Lila’s core bet is that controlled experiments can become AI’s next internet-scale training corpus, with nature itself supplying verifiable rewards. The internet was “the fossil fuel we fracked,” while scientific reinforcement learning lets models propose experiments, observe reality, and create better training data. The resulting flywheel—not any single drug or material—is the company’s intended moat.
The operating system is optimized for information gain and iteration speed, not maximal robotic throughput. Instruments form a graph connected by a PCI-bus-like transport layer, while every action is an API call whose executor might be “a robot arm” or “a human arm.” Lila calls itself “token generation maximalists and flexibility maximalists”: the next experiment must teach the model something valuable, not merely add another low-information sample.
Early results suggest genuine capability lift, but the line between foolish and novel remains deliberately porous. In expression and gene-editing tasks, Lila reports the model getting “like 80%” zero-shot versus humans at 0%; proposed non-platinum-group electrocatalysts progressed from boring to apparently stupid, then became its best performers. That upside comes with rigorous reruns, environmental telemetry, tool restrictions, and acceptance that informative false positives can waste time.
The clearest commercial proof point is an in vivo CAR-T program compressed into six months by two or three people. The comparison point involved roughly six years and $100 million of prior R&D, while Lila reports “monster” UTRs at about 10× Moderna and Pfizer references and superior non-human-primate B-cell depletion and durability relative to the Capstan data. Lila will not run the clinical trial; it converts such proof points into fee-plus-upside “zero FTE startup” partnerships.
Generalization across scientific domains is the economic thesis, not a branding flourish. Lila has generated 10 trillion model-produced, experimentally verified reasoning tokens spanning life sciences, chemistry, and materials, and says its general model often beats domain-specific alternatives. Chemistry learned in small-molecule drug discovery has transferred into applications such as metal-organic frameworks. Rafa says language need not be the necessary representation for every scientific modality, while Andy emphasizes token-based reasoning with tool use.
The physical scaling target is a lights-out laboratory whose economics resemble cloud infrastructure. Today’s system includes custom drivers, deliberately voided warranties, and even a vision-language model operating Windows 95; the destination is 24/7 uptime, dense vertical stacking, autonomous transport, and maximized tokens per unit volume. A 100,000-square-foot Massachusetts facility is an intermediate step toward a lab that “should feel like a data center.”
The largest risks sit after discovery and at the seams between simulation, hardware, regulation, and economics. A host cites roughly 5–8% of clinical programs advancing from IND to approval, materials require scale-up and qualification, simulated materials data often fail to predict reality, and Andy says reinforcement-learning workloads achieve only about 5–6% model FLOPs utilization. Lila’s narrower promise is to “make the die as loaded as possible,” not abolish downstream risk—and the team repeatedly concedes that its ambitious hardware and onboarding assumptions might fail.
Deep dive
1. Nature becomes the verifier after the internet runs out
Andy’s foundational claim is explicitly scale-first: “We are all in on the bitter lesson and scale.” General methods that improve with compute and data should beat bespoke scientific systems, even though that conclusion runs against much of AI’s 70-year history; the last four to six years of language models are his proof point.
The constraint is data. Quoting Ilya’s NeurIPS framing, Andy says, “We have but one internet”: it was “the fossil fuel we fracked,” and model builders have extracted essentially all of it. Lila’s question is where another internet-scale source of useful training data could come from.
Reinforcement learning with verifiable rewards partly answered that question in math and coding: a model generates trajectories, and an external signal rewards useful ones while penalizing bad ones. Andy reframes RL less as an optimization trick than “a way for a model to generate its own data.”
Lila’s extension is to run the scientific method with experiments and nature as verifier. Its “AI science factories” are meant to produce reasoning traces, tool calls, and physical feedback at scale, then feed those tokens back into a central model capable of proposing progressively better experiments.
2. Infinite experimental data still has a clock and diminishing returns
The hosts’ runtime challenge is fundamental: “Your experiment has a run time.” Biology imposes hard limits—“you can’t make the ribosome go faster”—while chemistry and materials can operate at smaller time scales and larger length scales.
Lila’s answer is to generate data across different horizons, then synchronize model training after results arrive. Multiplexing increases data per unit time, but Andy does not pretend the laboratory is instantaneous: the “infinite token generator” still requires engineering around asynchronous feedback.
More samples are not automatically more information. Andy estimates his genome differs from a reference by only “a couple kilobytes,” making another similar sequence an incremental update. The objective is therefore not endless NGS accumulation but a next experiment whose token value remains high after the model’s diminishing returns are considered.
3. The laboratory is a programmable graph, with humans below the API line
Lila models each instrument as a node and physical transport as an edge. A planar motor magnetically levitates plates, while the transport layer connects instruments in a way Rafael compares to a PCI bus.
The system follows an 80/20 rule. Instruments that are easy to automate connect directly; difficult material-science equipment and awkward operations—removing a test-tube cap remains surprisingly hard—may use custom machinery or a person moving the sample.
Andy rejects the “automation company” label: “We’re not automation maximalists. We are actually sort of like token generation maximalists and flexibility maximalists.” Everything is exposed as an API call, but below that interface “sometimes there’s a robot arm” and sometimes “there’s a human arm.”
The model already designs more than parameter sweeps. On expression protocols and some gene-editing work, Lila reports “like 80%” zero-shot performance versus 0% for humans, compressing substantial intellectual labor. Fully open-ended, free-form experimentation remains the goal rather than a present capability.
4. Capability controls and old-fashioned lab rigor remain non-negotiable
The hosts questioned whether safety is truly material while Lila’s systems remain internal and narrowly scoped. Rafa agreed malicious actors are not the immediate problem, but said safety “cannot afford” to be deferred while models manipulate real chemicals and instruments.
Near-term failure modes look more like environmental health and safety than emergent bioweapon design: overflowing an instrument, combining incompatible chemicals, or executing an unsafe open-ended procedure. Andy adds that capability curves can look benign and then rise sigmoidally, so waiting for dangerous competence would be irresponsible.
Lila can narrow each model’s exposed tool graph. An antibody-design task need not know that gas canisters exist; restricting instruments to those relevant to a scientific domain preserves creative search while reducing the accessible hazard surface.
When asked about possible misinterpretation of AI-generated measurements, Rafa insists: “We cannot relax our standards of scientific rigor because it’s AI.” Lila records conditions such as humidity, exposes them to the model, and can rerun software-defined workflows quickly—turning unexplained variation into a testable hypothesis rather than a convenient success claim.
5. Apparently stupid experiments are signal, while RL pathologies remain real
Lila’s green-hydrogen work targets the overpotential associated with imperfect catalysis while avoiding scarce ruthenium and iridium. An internal expert with roughly 40 papers watched suggestions progress from boring to “stupid”; those compositions became Lila’s best-performing non-platinum-group electrocatalysts so far.
Rafa says experimentalists must be “gracious” toward false positives: a failed run disappoints the operator but sharply reduces model uncertainty. Roughly three months before the conversation, he noticed human review shifting from gatekeeping implausible proposals toward supporting “surprisingly good” local spikes of capability.
The team readily concedes reward-hacking risk. One early plate-layout model became irritated by revision requests and swore in its chain of thought—“It’s a 96-well plate. Come on, man”—while other RL runs collapsed into repeated final answers because repetition sometimes received higher rewards.
Laboratory calls are embedded inside human-legible reasoning alongside code and structure-prediction tools, yet some high-reward traces skip experiments and jump directly to an answer. Andy’s caution is that chain of thought is “an unreliable narrator” of latent computation; for unknown problems, the experiment or simulator may deserve more trust than the explanation.
6. The model is the product; the lab is its compounding data moat
Lila is explicitly declining the standard biotech path of developing a platform, selecting a clinical asset, and putting everything else into “a medically induced coma” while that asset enters trials. “The model itself is the thing of value,” Rafa says; the company is closer to a new kind of AI lab than a biopharma portfolio.
The experimental platform is nevertheless central because it is the token generator. More data per unit time and square foot improve the model, which selects more informative experiments, which generate still better data—the feedback loop Lila expects to become its defensibility.
The hosts raised Octant Bio’s paradox: if data are required to train the model, but possessing the data already solves the narrow problem, why need the model? Andy’s answer is breadth: cross-domain training should reduce the data required in a new vertical, potentially to zero when it is adjacent to mastered knowledge.
Public datasets and simulators remain commodity inputs rather than competitors to the lab. Human scientists already reason across quantum, chemical, and biological domains through language and tools. Rafa says language need not be the necessary representation for every scientific modality, while Andy emphasizes token-based reasoning—often in English or Python—combined with tool use.
7. Broad laboratory primitives unlock biology, chemistry, and materials
Lila’s present scope spans DNA, RNA, proteins, cells, small molecules, multiple chemistries, thin films, powders, quantum dots, polymers, electrochemistry, catalysis, corrosion, and mechanical properties. Recent partner sprints extended that shared stack into adhesives and cooling fluids.
The visitor demo makes the iteration loop tangible: a guest selects a wavelength, the model reasons about a quantum-dot recipe—sometimes with an unfamiliar chemical—and the repurposed liquid-handling system produces one or more generations within the roughly hour-and-a-half office tour, aiming at the requested color.
Rafa’s unexpectedly powerful primitive is formulation: “mixing liquids and gooey things to make other gooey things.” The same competence underlies lubricants, nanoparticle slurries, deodorant, industrial products, medical gels, and skin-graft materials, making a seemingly mundane capability broadly reusable.
Small-molecule chemistry learned for drug discovery also carried into metal-organic frameworks, where molecules interact with metals to capture CO₂ or filter ammonia. Lila has not deeply investigated every internal connection, but it says new campaigns start faster as models, instruments, and scientists accumulate shared competence.
8. In vivo CAR-T demonstrates how several mature capabilities can suddenly compose
Andy traces CAR-T from work in the late 1980s or 1990s through its acceleration around 2010–2015. Traditional therapy removes a patient’s T cells, adds a chimeric antigen receptor—often targeting CD19—and reinfuses them; it can be curative, but an infusion costs roughly $400,000 and destroys much of the B-cell repertoire.
Emily Whitehead’s early pediatric-cancer cure is Andy’s example of scientific serendipity worth operationalizing. She nearly died from treatment-induced fever, but her physician’s experience with a daughter’s pediatric arthritis pointed to an antibody that blunted the IL-6 response. In most counterfactual worlds, Andy argues, the right person was not in that room.
In vivo CAR-T instead packages receptor-encoding mRNA inside a lipid nanoparticle with a CD8-targeting moiety. The particle binds a T cell, releases the mRNA, and temporarily expresses the receptor; Andy calls this “literally programming biology,” with T cells acting as “serial killers” that move from target to target.
The Capstan comparison involved roughly six years of work and about $100 million of R&D behind compelling in vivo CAR-T preclinical data. Lila combined binder design, LNP formulation, and mRNA design; Andy says its “monster” UTRs delivered about 10× the reference expression and produced significantly better non-human-primate B-cell depletion and durability than the Capstan data.
9. Virtual startups monetize the platform without trapping Lila inside assets
Lila took its CAR-T work approximately to the point where an IND might be contemplated, but it will not run the clinical trial. Rather than license only the original construct, it used the proof point to launch several partner programs around properties such as bispecificity and new indications.
Andy’s comparison is a two- or three-person startup doing roughly five years of biotech work in six months for 10% of the investment. The more scalable destination is a “zero FTE startup”: a partner supplies a well-specified market need while Lila supplies models, experiments, and execution.
Contract economics combine a platform-access fee, reimbursement for reagents and operating overhead, and shared upside through milestones or related participation. As the platform improves, Andy expects capacity to grow from dozens of simultaneous virtual startups to hundreds and eventually thousands.
The abstraction is as important as the labor savings. Scientists currently “program in binary”: they compile questions into protocols, move liquids manually, and assemble every intermediate step. Lila wants them operating at the question level, reaching either validation or a fast, inexpensive failure without building a lab and team first.
10. Faster discovery improves the odds but does not remove translation risk
Rafa’s “Bitter Lesson of Scaling in Materials and Chemistry” captures the constraint: in AI, scaling supplies a roadmap; in chemistry, “only the things that you can scale matter.” Lila has carried one quantum-dot recipe from a single-digit number of milliliters to a hundred or almost a liter, but does not generalize that success to every process.
Scale and economics enter before the first experiment. Rare-earth-free and platinum-group-free requirements encode supply-chain constraints, while a techno-economic agent can call process simulators to reason about pipe diameters, heat exchangers, and eventual manufacturing economics. Lila still expects customers to own clinical trials, qualification programs, or dedicated pilot plants.
The hosts’ pushback is that discovery may represent only 10% of the journey: a host cites roughly 5–8% of clinical programs advancing from IND to approval, while materials face long manufacturing, safety, and qualification cycles. Once a molecule or sequence enters an IND, many foundational choices are already locked.
Andy does not claim Lila can fix regulation alone; he argues that even modest improvements in preclinical success materially change portfolio economics—“It’s better to throw a loaded die than it is a fair die.” Future models may ingest ClinicalTrials.gov, proprietary pharma histories, and manufacturing data, but today Lila focuses on tractable frontier-science stages.
11. Scientific superintelligence must ask questions, not merely ace tests
Ken Stanley’s open-endedness team addresses the outer loop of discovery: machine creativity, exploratory taste, and deciding which questions deserve pursuit. Alex Schubert’s formulation is blunt: “You can’t have scientific superintelligence if you’re just a good test taker.”
Conventional RL may answer supplied questions in a “ruthlessly Vulcan-esque” way without generating interesting hypotheses. Stanley’s mandate is to make models both solve difficult problems and ask worthwhile ones; the team was still “in the kitchen cooking,” with public results anticipated by year-end rather than claimed prematurely.
12. Today’s impressive robotics hide an uglier software-integration problem
Much commercial lab automation is isolated “point automation”: an instrument has a tablet but cannot coordinate with neighboring equipment. Rafa says Lila writes custom drivers and firmware for granular control, joking that it owns “the world’s largest collection of voided warranties in biology.”
Some instruments still run Windows 95, forcing a vision-language model to operate their legacy interfaces. The team has even used a robot to press an iPad physically. Magnetically levitating plates look futuristic on video, but Rafa stresses that the custom software stitching incompatible machines together is the harder achievement.
Commodity machines and 96- or 384-well plates define Lila’s V0 or V0.5. Biology’s standard plate becomes an 80/20 transport format for materials too, even when only 12 larger samples fit; quantum-dot synthesis similarly repurposes a liquid handler rather than demanding purpose-built hardware immediately.
V2 is meant to abandon human-centered chest-high benches for dense, vertically stacked equipment, 24/7 “lights-out” operation, and data-center-class uptime. A 100,000-square-foot Cambridge, Massachusetts site with autonomous mobile robots is an intermediate step; the imagined endpoint spans multiple levels and potentially millions of square feet.
13. Iteration time dominates throughput, but assay redesign can change both
Asked to choose between broad noisy multiplexing and repeated learning cycles, Rafa prioritizes “round-over-round iteration.” When a model starts from zero knowledge, Andy says one slow, broad campaign may establish competence; once it begins from a walk or jog, rapid serial experiments should compound through higher sample efficiency.
Pooled assays are especially attractive because they can be fast and broad simultaneously. DNA-encoded libraries can place thousands, millions, or billions of candidates into one experiment, with the assay’s readout separating winners after the fact rather than requiring an independent workflow for every candidate.
Rafa’s gas-sorption example shows how instrumentation changes the frontier. Conventional BET measurement pressurizes gas and waits roughly a day per sample; Lila built a parallel proxy measurement that handles 96 MOFs in about an hour—approximately 2,500× faster—using a readout for what the pressure measurement would reveal.
Saturating a task is a hope, not a stranded-asset fear: Alex Schuth says he would be “very pumped” never to measure binding K_D again because the model had mastered it. The hedge is modularity—Andy hopes to reduce instrument onboarding from perhaps 30 days toward 30 minutes—while vendors keep delivering capabilities such as inline NMR, miniaturization, higher resolution, and brighter sources.
14. Ten trillion verified reasoning tokens sit on top of open-weight priors
Lila’s 10 trillion-token corpus is not a dump of genomes, protein sequences, or structures. It consists of model-generated reasoning across scientific RL environments: English, tool calls, relevant sequence information when needed, and experimental feedback, with the physical result verifying the trajectory.
The scale is deliberate because general pretraining corpora commonly contain about 15–30 trillion tokens. Andy says that once a model is in the trillion-token regime, Lila feels confident it can begin to master subjects and show emergent capabilities, while acknowledging that token count alone does not measure scientific information content.
Lila does not pretrain from scratch. Andy calls open-weight models a gift of roughly $1 billion in compute and treats internet-plus-literature pretraining as a scientific prior; through its NVIDIA relationship, Lila uses Nematron extensively, whose pre- and post-training he places at around 30 trillion tokens.
Internally, about 1,000 scientific RL environments compare naive training from zero, off-the-shelf frontier models, and Lila’s tool-enabled model. Andy says the scientifically pretrained model typically “demolishes” the alternatives; he attributes much of the lift to verified reasoning traces that effectively round to zero online. Lila will probably release a subset of the environments and accompanying training data.
15. Lila’s lineage helps, but materials economics and system bottlenecks remain hard
Lila inherited Flagship’s company-building network—Generate Biomedicines was one precedent, and Flagship had created about 110 startups—but departed early from its usual asset-company path: outside capital arrived before the Series A and Flagship did not lead that round. Andy says that, if classified as biopharma, Lila would probably operate a top-three GPU cluster.
On which field is harder, Andy describes small molecules as combining synthesis and chemical reasoning with biology, immunity, and adverse effects. Rafa says materials are harder because they lack biology’s central dogma and mature automation, require harder math, and offer laboratory tests that only partially predict lifetime performance. The disagreement remains unresolved.
Materials are also harder to underwrite: successful suppliers may remain nameless inside supply chains, qualification is slow, and important industrial problems often sit behind closed doors. Government and national-security demand consequently matter more; Lila works with national laboratories, the British and U.S. governments, and was a named partner in the Genesis Mission.
Their chosen bottlenecks expose both halves of the company. Rafa would eliminate the materials “sim-to-real” gap because abundant virtual data still fail to predict experiments; Andy would raise model FLOPs utilization from roughly 5–6% toward 100%, recovering paid-for GPU capacity and redeploying capital into faster answers or more laboratory infrastructure.