Pioneers Insight Method Research Author
A.I. Scientists Are Here. But Is Progress Accelerating? | EP 170
Back to Episodes

A.I. Scientists Are Here. But Is Progress Accelerating? | EP 170

Summary

  • Edison Scientific’s Kosmos compresses what collaborators described as three to six months of PhD-level data analysis into a roughly 12-hour run. The claim comes from giving the same objectives and datasets from unpublished research to Kosmos, which recovered researchers’ findings overnight; Sam Rodriques initially thought “there is no way that this is true.” He says its conclusions are right about 80% of the time, roughly comparable to asking a human to investigate independently.
  • The system’s economics look unusual because a $200 prompt can consume 1,500 papers and generate 42,000 lines of code. Kosmos coordinates hundreds of agents across OpenAI, Google, Anthropic, and internally trained models through a “structured world model” that keeps long tasks coherent. Rodriques calls $200 promotional pricing and contrasts it with the $5,000-$10,000 scientists may already have spent collecting the data.
  • Kosmos is already producing novel findings, but “AI discovery” still means accelerated analysis rather than autonomous scientific truth. Of seven conclusions in its paper, three replicated known findings and four were described as new contributions—including a proposed mechanism connecting a noncoding type 2 diabetes variant, a binding protein, gene expression, and SIRT1, which is involved in pancreatic insulin secretion. Scientists must still understand, cross-check, and experimentally validate the output.
  • The largest medical bottleneck remains clinical experimentation, not a shortage of plausible hypotheses. Rodriques says anyone assuming today’s hundreds-of-millions-of-dollars trials are already optimally designed is “off your rocker,” so AI can improve experiments and trials. But manufacturing clinical-grade material, recruiting scarce patients, dosing them, and waiting for outcomes remain stubbornly physical constraints.
  • Rodriques calls curing all disease within a decade “crazy,” while treating a 30-year leap as plausible. Even a drug that completely halted aging might require five or 10 years to prove effective, and “with infinite intelligence” some answers would remain unknowable without new experiments. Faster biomarkers could shorten that clock, which is why he allows that the aggressive forecasts might still surprise him.
  • Generative biology—not merely prediction—is the scientific capability with the largest current step change. Models can now propose antibodies, proteins, or organisms from scratch for desired characteristics; the envisioned workflow is to specify a disease protein and generate a matching antibody, though manufacturing, validation, and human trials still follow. Rodriques rates AlphaFold 3 underhyped, lab automation appropriately hyped, and “virtual cell” branding, quantum computing, and brain-computer interfaces overhyped.
  • Scientific agents are at the beginning of an adoption S-curve, but ordinary lab practice will change more slowly than the model frontier. Coding assistants and literature search offer immediate value to conservative biologists, whereas deeper agentic research requires trust built through demonstrated results. Rodriques once thought a majority of high-quality hypotheses being generated by agents in 2026 or 2027 was overhyped; now he calls 2026 ambitious but says 2027 may be real.

Deep dive

1. Kosmos compresses months of analysis into a 12-hour run

  • Rodriques’s original reaction to the six-month claim was “there is no way that this is true.” Edison Scientific tested it by giving Kosmos the same objectives and datasets from academic work that collaborators had not yet published. It rediscovered overnight findings that researchers said had taken three, five, or six months—not necessarily six months of continuous labor, but that amount of research effort.

  • Kosmos looks like a prompt box but is “not a chatbot”: users submit a research objective and wait roughly 12 hours. Each run uses models from OpenAI, Google, and Anthropic alongside Edison’s task-specific internal models, rather than relying on one frontier system.

  • The technical unlock is a “structured world model” recording the evolving state of knowledge about the task. That lets hundreds of agents work in parallel and sequence without forgetting the objective or drifting “off the rails,” producing one coherent investigation rather than disconnected model outputs.

  • The workload explains the price: an average run reads 1,500 papers and writes 42,000 lines of code, versus perhaps a few hundred lines from a Claude session. The current $200 charge is promotional and will rise; against the $5,000-$10,000 cost of gathering experimental data, users tell Rodriques they “can’t believe” it is only $200.

2. Novel findings arrive before proof—and still require scientists

  • The team’s Kosmos paper presented seven conclusions: three replications and four described as new contributions to the scientific literature. Rodriques says the system returns genuinely deep insights and is right about 80% of the time, “kind of similar” to asking a human researcher to investigate independently.

  • His strongest example involved millions of genetic variants associated with disease whose mechanisms remain unknown. Given raw data for type 2 diabetes, Kosmos connected a variant outside a gene to a protein that binds near it, identified the relevant gene expression, and connected the result to the mechanism of SIRT1, which is involved in pancreatic insulin secretion—a mechanism no human researcher had yet assembled from those inputs.

  • Casey Newton’s useful decomposition was that science chooses data, gathers it, then draws conclusions; Kosmos currently attacks the third step. Rodriques recalled spending six months analyzing his PhD dataset while earning roughly $40,000 a year. The agent can surface findings quickly, but scientists must first understand months’ worth of compressed work, rerun analyses, cross-reference it, and validate experimentally; this particular result is, in his view, unlikely to become a drug target.

3. Clinical reality, not idea generation, sets medicine’s clock

  • Casey Newton’s pushback—worth keeping—was that drug discovery may not be the binding constraint when trials, patient recruitment, and FDA approval take far longer. Rodriques largely agreed: the number of diseases scientists know how to cure in mice is “astronomical,” because experiments are easy to run there; human experimentation is intrinsically slower.

  • AI still matters upstream because assuming every pharmaceutical trial is optimally conceived from all available knowledge means “you are off your rocker.” Trials cost hundreds of millions of dollars, while existing datasets contain insights no one has had the capacity to find. Better analysis should therefore produce better experiments and trials, even if it cannot remove the trials.

  • Rodriques’s blunt timeline call: curing all or most diseases within a decade is “crazy”; a “humongous leap forward” over 30 years is plausible, though neither ending aging nor curing everything is guaranteed possible. If a drug halted aging between ages 25 and 65, researchers might still need five or 10 years merely to detect the effect.

  • Regulation is only part of the delay. Even with no regulator, teams must manufacture enough clinical-grade material for humans, locate scarce patients by forming relationships with doctors, dose them, and wait. Nor will “GPT-7” simply explain how to cure Alzheimer’s: “with infinite intelligence,” missing facts about the world would remain missing until experiments supplied them. Biomarkers might accelerate iteration, which Rodriques accepts as a reasonable route by which optimistic forecasts could beat his own.

4. Better science needs both world models and productive noise

  • Rodriques divides AI science into modeling the natural world and modeling how science is done. Kosmos belongs to the second category; protein-structure prediction, antibody generation, and organism design belong to the first. The biggest capability shift is generative design: producing proteins, antibodies, or organisms from scratch with requested properties, something he calls “a new capability that we have never had before.”

  • Kevin Roose challenged the reliability of scientific analysis with the example of Google answering that 2026 was not next year. Rodriques said researchers must spend substantial time checking agents, but publication already demands checking such work. Perfection is unavailable either way; the attainable standard is human-comparable reliability, while “checking the work is always going to be faster than producing it in the first place.”

  • The unresolved issue is whether optimization destroys serendipity. Rodriques expects mistakes to survive, invoking spores entering an open window and producing the observation behind penicillin. First-year graduate students likewise generate progress by doing “the most random kooky stuff” experts would reject. His response is that systems may need to preserve variation or “add noise,” just as biological evolution uses randomness to discover useful functions.

5. Adoption starts with coding while agents approach the S-curve

  • Most scientists have not radically changed their work. Rodriques describes biologists as conservative because inherited protocols work even when no one fully understands every step, and testing every alternative is impossible. Most labs will therefore continue familiar methods until they see peers producing demonstrably better results.

  • Coding and literature search are the immediate exceptions. Biology’s historical coding bottleneck is falling as researchers use Claude Code, OpenAI models, and Gemini without already knowing how to code, while agents can parse an otherwise unmanageable scientific literature. Full research agents sit further out on the frontier and may diffuse more slowly.

  • In the lightning round, Rodriques called vibe proving probably overhyped as an end use, though valuable for advancing AI; scientific lab robotics “appropriately hyped” but technically immature; and AlphaFold 3 probably underhyped despite intense attention. “Virtual cells” are overhyped as a name because current systems model narrow cellular functions, not complete cells; quantum computing and sci-fi-style brain-computer interfaces are also overhyped or further away than people imagine.

  • His three defining 2025 advances were scientific agents, de novo antibody design from groups including Chai and Nabla Bio, and the Arc Institute’s from-scratch bacteriophage design—the last “awesome” even if its utility is uncertain. He expects agents to “infiltrate everything” in 2026. A majority of high-quality hypotheses being agent-generated that year remains ambitious, but the prediction he once thought he was overhyping may be real for 2027.