Pioneers Insight Method Research Author
No Priors Ep. 138 | The Best of 2025 (So Far) with Sarah Guo and Elad Gil
Back to Episodes

No Priors Ep. 138 | The Best of 2025 (So Far) with Sarah Guo and Elad Gil

Summary

  • The strongest startup signal was blind workflow validation, not model spectacle. Winston Weinberg and Gabe tested GPT-3 on 100 landlord-tenant questions with early chain-of-thought prompts; three attorneys judged 86 answers suitable to send without edits, after being told nothing about AI. Even OpenAI’s general counsel replied, “I had no idea the models were this good at legal.”
  • AI can revive “graveyard” markets when adjacent infrastructure removes the original constraint. Arvind Jain argued enterprise search failed because pre-SaaS data was inaccessible; standardized, interoperable SaaS systems and APIs made a turnkey product possible just as internal content exploded. One Glean customer has more than 1 billion documents—the size of the entire internet Jain recalled from Google in 2004.
  • Reasoning models become materially more efficient when they recognize uncertainty and delegate work to tools. In visual tasks, models can admit they cannot see something, crop or manipulate the image, and achieve “very noticeably different” test-time scaling slopes. For valuation work, one researcher’s example was simpler: have the model write code to run the calculation and “know what the actual answer is.”
  • Digital labor displacement could arrive faster than physical automation or political adaptation. Brendan Foody expects rapid, painful displacement across roles such as customer support and recruiting, potentially producing “a big populist movement” and difficult wealth-allocation questions if superintelligence gains follow a power law. He expects physical automation to be slower, with potential work ranging from robotics-data creation to restaurants and therapy.
  • Superintelligence could turn shared vulnerability from a deterrent into a trigger for preemption. Dan Hendris compared advanced-AI strategy with nuclear retaliation, but argued automated AI research could become so destabilizing that China or the US might launch cyberattacks against the other’s data centers. He imagined Russia reassessing the balance, using espionage to monitor projects, and threatening retaliation.
  • Technical conviction does not eliminate model surprise or brittleness. Isa Fulford expected training on browsing tasks to work, yet found the first working model surprisingly strong. Sarah Guo likened the experience to a path “paved with strawberries,” and Fulford agreed. The same model can perform something exceptionally smart and then make a mistake that prompts, “Why are you doing that? Stop.”
  • Entrepreneurship and healthcare outcomes supplied different evidence about disciplined building. The Flagship speaker questioned treating entrepreneurship as random “shots on goal,” especially when deploying hard-earned capital in healthcare, climate, agriculture, and food security. At Abridge, product criticism is “oxygen,” while feedback from a doctor saying she would not retire—and a family saying the tool let a mother come home for dinner—provided the deeper purpose Rao described.

Deep dive

1. Workflow evidence revealed capability before the market noticed

  • Harvey’s origin was Winston Weinberg’s surprise that “no one was talking about GPT-3.” He and Gabe applied early chain-of-thought prompts to 100 r/legaladvice landlord-tenant questions, told three attorneys nothing about AI, and asked whether each answer was ethical and usable without edits. The attorneys answered yes to 86.
  • Dr. Fay Lee framed spatial intelligence as evolution’s difficult conversion of collected light into an internal 3D world for navigation, manipulation, and interaction. Humans still struggle to reconstruct their surroundings with their eyes closed; making rich 3D creation fluid and editable “at your fingertips” would create “a whole different world.”

2. Infrastructure changes can turn graveyards into markets

  • Arvind Jain’s diagnosis of enterprise search’s “graveyard” history was infrastructural: pre-SaaS vendors could not reliably locate and connect data across servers and storage systems. Shared software versions, interoperability, and APIs finally made unified, turnkey search feasible.
  • The need surfaced at Rubrik, where information sprawled across 300 SaaS systems and employees complained they could find nothing—yet Jain found “there was nothing to buy.” Scale deepened the opportunity: a large Glean customer holds over 1 billion documents, matching the entire 2004 internet in Jain’s comparison.

3. Tools steepen reasoning’s scaling curve

  • The OpenAI researchers described visual models recognizing their own uncertainty—“I can’t really see the thing”—then using tools to crop or manipulate an image. That makes inference tokens more productive and produces a noticeably steeper test-time scaling slope.
  • One researcher’s valuation example exposed the division of labor: a model could repeatedly fit coefficients and self-verify inside its context, or write a simple program to run the valuation model and determine the actual answer. Compute improves when work outside the model’s “comparative advantage” moves to a purpose-built tool.
  • Fulford expected browsing-task training for Deep Research to work, but was still surprised by how well it worked. Sarah Guo likened the experience to a path “paved with strawberries,” and Fulford agreed. Yet performance remains jagged: models can do “such smart things” before making an inexplicable error and prompting, “Why are you doing that? Stop.”

4. Digital acceleration raises labor and security risks

  • Foody expects displacement across digital roles such as customer support and recruiting to happen “very quickly” and painfully, creating a large political problem. Near superintelligence, the harder question becomes reallocating wealth if gains concentrate according to a power law.
  • The host’s pushback—“What does the physical world mean?”—forced specificity: robotics-data work, waiting tables, and therapy where people want human interaction. Foody thinks physical automation will be slower because it lacks the virtual world’s self-reinforcing improvement loops.
  • Hendris first drew a nuclear analogy: states could deter first strikes through shared vulnerability. He then argued that, once systems can automate most AI research, rapid bootstrapping to superintelligence or a “superweapon” could become so destabilizing that China or the US might preemptively cyberattack the other’s data center. He imagined Russia reassessing the balance, using espionage to monitor projects, and threatening retaliation.

5. Disciplined building earns legitimacy through outcomes

  • The Flagship speaker recalled starting a company in 1987 as a 24-year-old immigrant, when former Merck or IBM senior executives were the people entrusted with venture capital—roughly $23 million per round. That experience led him to ask why entrepreneurship could not become a profession instead of remaining random, improvisational, emotional, and “gamy.”
  • Challenged on “gamy,” he pointed to a culture where repeated failures and occasional wins become a scoreboard. When capital tackles “damn near impossible” problems in healthcare, climate, agriculture, and food security, “you can’t think of this as…shots on goal.”
  • Abridge routes positive feedback into a “love stories” channel accessible to anyone inside the company. Shiv Rao calls harsh product criticism “oxygen,” while praise includes a doctor saying she will not retire and a rural doctor’s family saying Abridge let her come home early for dinner. Rao contrasted hypergrowth’s short “dopamine hits” with purpose’s “oxytocin hits,” which he said the company is really after.