Pioneers Insight Method Research Author
The Decade of May 15-22, 2025: Google's 50X AI Growth & Transformation with Logan Kilpatrick
Back to Episodes

The Decade of May 15-22, 2025: Google's 50X AI Growth & Transformation with Logan Kilpatrick

Summary

  • Google’s AI usage rose 50× in a year—from roughly 10 trillion to 500 trillion tokens per month—amid an organizational rebuild, infrastructure scaling, and surging product demand. That is more than 50,000 monthly tokens per person on Earth, before counting other providers. Kilpatrick credits the mid-2023 Google Brain–DeepMind consolidation, repeated training-and-release cycles, TPU expansion, and DeepMind’s shift from foundational research into products; the key signal is “the slope of improvement” across all four.
  • Kilpatrick expects leading models to diverge as easy gains run out and structural advantages begin to dominate. Research initially converges because “people light the path,” while product teams race to avoid looking behind; from here, however, “the low-hanging fruit has been captured,” improvement gets harder, and Google’s compute infrastructure should matter more. Nathan’s pushback is the investor’s version: if DeepMind considers the next frontier difficult, tier-two foundation-model training looks brutal.
  • Application-layer economics remain unusually favorable even if frontier-model economics consolidate. Kilpatrick calls this “the best time in human history” to build a startup: billion-user incumbents solve general problems, while startups can attack narrow segments with speed, focus, and cheaper tooling. Nathan’s evidence is tangible—about $1,000 a month in AI subscriptions produces far more value, while one AI-assisted radio-ad project was priced roughly 75% below the incumbent workflow yet generated about $3,000 of project revenue per hour he personally spent.
  • The Windsurf episode exposes supplier risk, but Kilpatrick sees Google’s incentives favoring broad and early API distribution. He considers Anthropic’s decision to cut Windsurf off after its OpenAI partnership defensible as compute allocation to likely long-term partners. For Google, Cloud’s mandate, external evaluation signal, competitive momentum, and slow model migration inside billion-user products create “many levels of motivation and game theory” against withholding the best models—an expectation, not a Gemini 3 guarantee.
  • Gemini 2.5 Pro’s clearest differentiation is command of long context, not merely context-window size. Nathan’s 400,000–500,000-token codebase test felt like a step change; Kilpatrick cited OpenAI’s MRCR eight-needle evaluation, where the newest 2.5 Pro was “on the order of like 20% better.” His mechanism is that reasoning finally enables models to use the full window, pushing architectures toward more context loading even though RAG will remain necessary.
  • Reasoning models are turning agent architecture into a moving target: today’s necessary scaffolding can become tomorrow’s technical debt. Models are becoming “systems and agents out of the box,” while NotebookLM’s audio-overview pipeline has already fallen from 14 Gemini-heavy stages to four. The practical advice is to build structured workflows for present reliability while avoiding one-way-door designs that cannot be simplified as capabilities move into the model.
  • Diffusion language models could create another software discontinuity through sheer speed. Kilpatrick finds their coarse-to-fine generation more intuitive than strictly autoregressive writing and would not be surprised if the paradigm wins within two years. Nathan suspects diffusion could ultimately be better and notes that the transformer is not the end of history. Quality, cost, and performance trade-offs remain open, but Kilpatrick calls the demonstrated speed “unbelievably fast” and sees editing as a naturally strong fit.
  • Kilpatrick thinks AGI will be felt as a product assembly—and that authentic human perspective will remain scarce inside it. A model with perhaps 50% better reasoning, 50% better long context, and properly surfaced memory could produce the experience before any single release wins universal agreement as “AGI.” Yet he writes about 95% of his emails and posts without AI because agency and voice matter: “people want the Nathan experience,” even when NotebookLM can generate polished, topic-specific substitutes on demand.

Deep dive

1. Competition creates parity before it creates differentiation

  • Kilpatrick separates research convergence from simple copying. Reasoning is his specimen: once an innovation and the required “order of magnitude of investment” became legible, “people light the path”; competitors could incorporate that insight while combining it with independent bets already underway.

  • Product convergence follows a harsher incentive. AI is “the most competitive ecosystem in the entire world right now,” concentrating capital, talent, intellectual capital, and execution speed; teams must balance differentiated long-term bets against the immediate reputational cost of lacking whatever feature rivals just shipped.

  • That tension reaches Google’s developer platform directly: compatibility makes it easier for customers of other providers to adopt Gemini, but every unit of effort spent reaching feature parity competes with building next-generation APIs and model capabilities.

  • Nathan’s suspicion that clustered launches reflect coordinated readiness gets a practical rebuttal: some timing is intentional, but much is happenstance. Organizations larger than roughly 20 people cannot reliably pivot a major release with an hour’s notice; “companies are not that nimble,” including the frontier labs.

2. Google’s turnaround was an organizational rebuild, not a model launch

  • Google’s old fragmentation had a rational basis. Google Brain pursued broad research that produced work including the transformer; Google Research leaned more applied and upstreamed capabilities into products; DeepMind followed Demis Hassabis’s more specific view of the route to AGI.

  • Once one near-term approach was visibly working, dispersed bets became a liability. The mid-2023 combination of Brain, parts of Research, and DeepMind was, in Kilpatrick’s framing, the beginning of “putting itself in the position to be successful for the next 10 years.”

  • Reorganization did not instantly create frontier execution because “large human systems are extremely complex.” Leadership had to reconcile different cultures, establish team structures, build repeatable training and release cycles, and learn operational lessons that OpenAI had already practiced more frequently before that moment.

  • Compute had to scale alongside the institution: TPUs were required for both research and inference, and idle capacity was not waiting on demand. DeepMind also became a product organization responsible for the Gemini app and developers; Kilpatrick expects the next three to five years to reveal the payoff from those “good and hard decisions.”

3. Fiftyfold usage makes infrastructure a strategic constraint

  • Google went from roughly 10 trillion tokens per month to 500 trillion in about a year—a 50× increase and more than 50,000 tokens monthly for every human alive. Because usage is uneven and other providers add more volume, the active-user intensity is higher still.

  • Kilpatrick argues that incentives and corporate DNA reinforce that curve. Google had deployed transformer-based systems inside Search at multibillion-user scale long before the present generative incarnation, while better models now improve Docs, Sheets, YouTube, Waymo, Cloud, and other products rather than functioning as a detachable add-on.

  • His forecast is increasing divergence: “the low-hanging fruit has been captured,” each further advance becomes harder, and Google’s infrastructure advantage should become more visible. Nathan sharpens the implication—if DeepMind says the next level is difficult, being a tier-two foundation-model trainer looks especially unattractive.

  • Specialization remains one possible route: a lab could theoretically decide to build the world’s best coding model instead of the best general model. Kilpatrick uses Anthropic only as a hypothetical and immediately hedges that its broad mission makes such a code-only turn unlikely; smaller model companies are already pursuing areas in which they can build a differentiated perspective.

4. Startups retain speed and focus while frontier training consolidates

  • Nathan’s bearish case is that large technology companies can win any category they choose, including applications: all three major frontier developers had just launched coding agents, creating direct competition for products such as Cursor and Windsurf that also build on their models.

  • Kilpatrick concedes that Google’s Jules was “super early” and far behind the adoption of leading coding products; Nathan notes that one had just announced roughly $500 million in ARR. That gap is itself evidence that an incumbent model provider does not automatically inherit application distribution.

  • Training frontier models demands deep capitalization, but “there’s no better time in human history than right now to be building a startup” at the application layer. Software creation is faster, experimentation cheaper, monetization unusually rapid, and a million specific problems remain beneath the general solutions built for billion-user audiences.

  • The durable startup advantages are speed and focus. Large enterprises face security, privacy, and evaluation constraints that slow adoption of new tooling; a startup can solve one user’s problem without arbitrating among “a million and one innovative things.” Kilpatrick’s phrase: “the ability to focus on just a single thing is a blessing.”

5. Cloud economics argue against withholding Google’s best models

  • Windsurf illustrates dependency risk cleanly: after it agreed to a deal with OpenAI, Anthropic withdrew access to Claude, previously its main model. Kilpatrick has “empathy” for Anthropic’s stated desire to allocate scarce compute toward likely long-term partners; Nathan agrees that declining to support a competitor is conventional business strategy.

  • Nathan then poses the darker scenario: Google could deploy Gemini 3 internally across its coding agent, Gmail, and Docs, then delay the public API by months. Kilpatrick finds that hard to imagine because Google Cloud, which he describes as the world’s fifth-largest enterprise business, has a mandate to export Google-grade infrastructure so others can build without recreating it.

  • External developers can currently move faster than Google’s internal teams. A product with one billion—or even 150 million—users must prevent behavioral regressions, run extensive evaluations, and preserve user expectations; a small startup can switch models almost immediately. Forcing internal deployment first would therefore lengthen releases dramatically.

  • Against the AI 2027 scenario of widening internal-public gaps, Kilpatrick cites weak pre-release evaluation signal, the importance of projecting momentum, and customers’ switching costs. Per-million-token pricing might eventually change, but broad distribution still captures uses across a wide range of products: “there’s a great business to be built giving that to other people.”

6. Gemini 2.5 Pro’s clearest gap is usable long context

  • Default personality is harder to standardize than it appears. A baseline Gemini model serves products with radically different users, so Google has taken a relatively neutral position while leaving room for products such as the Gemini app to add character; changing models can otherwise feel as if “the person…is now gone.”

  • Nathan’s qualitative test was a 400,000–500,000-token research codebase. Gemini 2.5 Pro could retain command of the whole repository while repeatedly rewriting long files and debugging the resulting errors—an experience no other provider available to him supported at that context length.

  • Kilpatrick’s quantitative evidence is OpenAI’s MRCR benchmark. On the hardest eight-needle retrieval variant, the latest 2.5 Pro was “on the order of like 20% better”; single-needle retrieval was already near 100% with Gemini 1.5 Pro, but performance had historically decayed sharply as distinct targets accumulated.

  • The mechanism, developed in conversation with reasoning lead Jack Ray, is a fusion of reasoning and context: reasoning lets the model actually exploit the available window. Google is now seeing materially longer requests, and Kilpatrick expects more workloads to load information directly into context—“of course, they’ll still need RAG in cases.”

7. Early access is relationship-driven, while Veo 3 waits on scale

  • Kilpatrick makes the trusted-tester funnel unusually explicit: builders with an interesting project can email lkilpatrick@google.com . Google has a “super robust early access program,” and his criterion is useful developer feedback rather than membership in some mysterious inner circle; secrecy varies by project.

  • Veo 3 API access was not yet available for external onboarding. The constraint is infrastructure: API demand has a different order of magnitude from a high-priced consumer product, so Google must ensure that capacity can support the scale before opening it broadly.

  • Native audio changed Kilpatrick’s view of generated video. He had been skeptical because turning silent clips into practical media required substantial downstream work; synchronized speech and sound make the result feel alive. Nathan’s Waymark example is local-business advertising, where Veo 2 animated assets and Veo 3 could add clips with different voices beyond the traditional voice-over format.

8. AI spending is already producing disequilibrium economics

  • Nathan’s personal AI subscriptions had reached roughly $1,000 per month across OpenAI, Claude, Gemini, and about 20 accumulated tools. Much of that is duplicative testing, but he says the productivity gain remains “dramatically higher” than the cost even after allowing for experimentation.

  • The interaction model now feels like delegating to humans. While traveling, Nathan built two applications on Replit with the agent doing almost everything; instead of engineering completions token by token, he provides product notes and intervenes mainly when the result suggests that his instructions may have misled it.

  • Kilpatrick’s preferred future evaluation is economic productivity per dollar rather than another saturated academic benchmark. The useful question is how much positive value a system creates from a $20 subscription—and he expects the answer five years from now to be materially different from today’s.

  • Nathan’s sharpest example is a local-radio production job normally worth six figures at a few hundred dollars per location and version. AI enabled roughly a 75% client discount while project revenue per hour he personally spent approached $3,000—not all paid to him. Kilpatrick says such edges will continue to appear, and that “being on the frontier is likely to be disproportionately rewarded.”

9. Better reasoning collapses agent scaffolding into the model

  • Nathan’s “horseshoe theory of agents” puts both early chatbots and advanced coding agents into turn-based interaction: the new turns are simply larger. Reliable unattended automation still tends to live in the middle—structured, multi-prompt workflows with limited branches, rather than a model receiving 50 tools and permission to choose its own adventure.

  • Kilpatrick expects more scaffolding to move into the reasoning/model layer, with search, code execution, sandboxes, tools, and function calling increasingly available out of the box. Scaffolding will still be necessary, but builders should avoid one-way doors that require a fundamental rewrite once the model natively handles yesterday’s custom orchestration.

  • NotebookLM shows the transition numerically. Its audio-overview generation began as a 14-step, mostly Gemini-powered pipeline and is now four steps; eliminating sequential calls simplified the system and made the product faster, not merely easier to maintain.

  • A2A addresses production-agent requirements that MCP does not, including authentication and other parts of deploying agents at scale. Kilpatrick leaves the standards outcome open: MCP might expand into those functions, or it might preserve room for complementary frameworks such as A2A.

10. Diffusion language models could make software feel instantaneous

  • Kilpatrick finds diffusion’s coarse-to-fine process closer to his own cognition: thought begins fuzzy and structural, then breaks into parts, with sentence-level token sequencing only at the end. He would not be surprised if that paradigm won within two years, though he makes no categorical prediction. Nathan later says he suspects diffusion could ultimately be better and notes that the transformer is not “the end of history.”

  • Kilpatrick’s immediate reaction is speed—“unbelievably fast.” If comparable quality and cost hold, generation fast enough to render within “the blink of a human eye” could make personalized generative interfaces practical rather than forcing users to watch a screen stream tokens.

  • The trade-offs remain unknown, and autoregression may coexist with diffusion. The broader lesson is to keep exploring alternative paradigms, while diffusion’s ability to revise an output appears especially well matched to increasingly important editing workflows.

11. AGI may be assembled as a product before it is declared a model

  • Kilpatrick doubts a single release will produce universal agreement that “we’ve clearly built” AGI, partly because definitions already diverge. His expected AGI moment is experiential: somebody combines a very capable model with the right product-level components, including memory and context, and users conclude that the overall system feels general.

  • The underlying jump might be less cinematic than the product: perhaps long context becomes 50% better, reasoning becomes 50% better, and a team finally surfaces the right memories at the right moments. That last step may be as much engineering, neuroscience, and human psychology as model scaling.

  • Nathan points to Google’s Titans work as evidence that memory is progressing separately, citing demonstrations up to 10 million tokens with strong performance and the possibility of going further. Like the brain, the resulting intelligence may remain modular rather than residing wholly inside one monolithic model.

12. Human perspective becomes scarcer as synthetic content gets abundant

  • Kilpatrick’s worldview remains “fundamentally human-centric.” He personally writes roughly 95% of his emails and posts with zero AI assistance because he wants agency over his tone and public identity; even a convincing digital twin saying something plausible on his behalf feels foreign.

  • Synthetic abundance can increase the value of authored perspective. Kilpatrick cares less when someone sends content produced entirely by AI because it signals less craft, but he preserves an important exception: if software works beautifully, he generally does not care whether a human wrote it.

  • Nathan’s pushback is NotebookLM, which can create an on-demand program about a topic no human podcaster has covered and is already taking some of his listening time. Nathan then invokes Sundar’s comparison with Search: AI chat grew to hundreds of millions of users while Google queries still grew, suggesting related products can solve distinct intents rather than substitute one-for-one.

  • Attention remains the hard constraint—Nathan has raised default listening from 2× to 2.5× and wonders whether neural interfaces eventually become necessary as models outrun human reading. Kilpatrick’s counter-bet is relational scarcity: amid thousands of optimized substitutes, audiences will still want “the Nathan experience” and a differentiated human point of view.