Sharp Tech preview: Nvidia's answer to AI capital constraints
Sharp Tech preview: Nvidia's answer to AI capital constraints
Summary
- Ben frames the binding constraint on the AI buildout as money itself, after compute and power. He leans on the railroad analogy because the fundamental issue in 1873 “is the world ran out of money” — a year ago bubble talk was dismissed since hyperscalers paid out of free cash flow, but “we sort of blew through debt in like nine months,” with debt raised in the second half of last year and first half of this year “probably soon to be approaching, like, a trillion dollars” as issuers’ balance sheets get “sketchier and sketchier.”
- NVIDIA’s new $500B+ financing platform with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR is a pitch to reclassify AI as patient-capital infrastructure. Jensen Huang’s case is that “you’re all thinking about AI wrong” — GPUs run longer than you think and CUDA improves them over time, so the asset class fits pension-fund money that classically buys toll roads. Ben’s zoom-out: the case is being made now “because all the short-run capital’s been used up.”
- The A100 proof point — six-year-old, fully depreciated chips that CoreWeave says are contracting at higher rates than before — doesn’t prove what Jensen wants it to. Ben’s mechanism: the shift to water cooling means GB200s and the upcoming Vera Rubin can’t slot into old air-cooled data centers, so those facilities may be stranded and A100s may persist because “there’s no replacement for them.” In abundance, “the old compute’s gonna get retired very quickly.”
- Today’s scarcity reflects pre-2024 decisions, and correlated signals are how boom-bust cycles happen. Everyone sees demand exceeding supply at the same time: “It might be a 5X signal, but 10 companies invest, so you end up with double the capacity that you need” — so the current supply-demand environment is not representative of two, three, or thirty years out.
- Ben believes the infrastructure can pay off, but is bearish on certainty of timing. Unlike a railroad, AI can have digital-good scalability — “none of that stuff quite works now… it’s working pretty well, and it’s accelerating unbelievably rapidly” — but “it’s not enough to be right, it’s about timing,” and the risk is “an air pocket where we run out of money” before the spend cycles back as profit.
- The micro story: LLMs helped send NVIDIA’s stock to the moon while diminishing its CUDA moat. The developer platform shifted far above CUDA — “no one who’s writing an AI application today is using CUDA”; apps sit on OpenAI/Anthropic APIs or Bedrock-on-Trainium, fully abstracted from chips — so Ben agrees with Andrew Sharp’s read that the financing platform is partly defensive as cost-sensitive customers push toward Google TPUs.
Deep dive
1. The real question isn’t compute or power — it’s what happens when you run out of money
- This is a mailbag episode, opening with an emailer (Andrew, not the DC one) asking whether NVIDIA’s newly announced financing platform — Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR mobilizing “over $500 billion of third-party capital” — echoes the re-securitization of mortgages into CDOs that set up the subprime crisis.
- Ben’s macro framing runs through the railroad analogy he “reluctantly” linked (even Satya Nadella cited it): the fundamental issue in 1873 “is the world ran out of money.” A year ago you couldn’t call this a bubble when companies paid from free cash flow — “once we start getting into debt, then we need to have a conversation.” Then: “we sort of blew through debt in like nine months,” a sum “probably soon to be approaching, like, a trillion dollars,” raised by great businesses whose balance sheets are “getting sketchier and sketchier.”
- The hosts’ victory lap — they’d predicted in Madison a week earlier that lending would tighten. “Good job by us.”
2. Jensen’s pitch: AI is long-run infrastructure that deserves long-run capital
- The untapped pool is patient capital — pension funds whose classic investment is a toll road (Ben detours through the “doctor plan,” the catch-up pension structure for late-starting high earners, to explain the mechanics). Pensions in theory would have been a good match for railroads: money that must exist long-term but pays out slowly.
- Huang’s post argues “you’re all thinking about AI wrong” — it’s a long-term investment: GPUs run longer than you think, CUDA makes them better over time, and data shells are 30-year assets. Ben’s zoom-out: “it’s like, yeah, because all the short-run capital’s been used up.”
- Ben sees a “beautiful symmetry” with Google’s equity issuance, which he’d compared to Berkshire using high-margin See’s Candies cash flow to buy BNSF — a lower-margin business throwing off high absolute, predictable cash.
3. The A100 evidence is real but not representative
- Ben flags the choreography as no accident: Huang makes the case, then CoreWeave’s earnings tout A100s — a six-year-old, previous-generation chip — “contracting out at a higher rate than before,” fully depreciated, pure profit. On the surface, “a pretty good argument. There’s just a couple problems.”
- Problem one: the shift to water cooling means GB200s and the upcoming Vera Rubin can’t slot into old passively-cooled data centers (the H generation may have been air-cooled or half-and-half). Those facilities may be stranded, so A100s may stay in place because “there’s no replacement for them” — fine in a compute-scarce world, but “not representative of what your expectations should be for GPUs going forward.”
- Problem two: today’s supply reflects 2024-and-earlier decisions (two-year lead times), when markets were “freaking out about CapEx” — and the spenders were wrong only in not spending more. But when everyone gets the same signal: “It might be a 5X signal, but 10 companies invest, so you end up with double the capacity that you need.” If GB200s become abundant, “the old compute’s gonna get retired very quickly.”
4. The bull case isn’t insane — but being right isn’t enough
- Ben’s distinction from the railroad: there’s no way to accelerate a railroad’s revenue — brutal terrain, land development, finite trains — whereas AI can have digital-good scalability, especially in the possibility of AI writing its own programs or being “set loose on a company” to create agents. “None of that stuff quite works now, but… it’s working pretty well, and it’s accelerating unbelievably rapidly.” He mentions pushback from an emailer calling him a Luddite for taking six months to vibe code, and says the criticism is kind of valid.
- The unavoidable math: more supply depresses prices; the bet is demand accelerates even faster. And debt can’t fund things forever — “at some point you need to actually make money.” Ben’s bottom line: “I believe this stuff will pay for itself. The question is will it pay for itself in time to avoid, like, an air pocket where we run out of money?” Andrew’s translation: “a whole bunch of bag holders.”
5. The micro story: LLMs shifted the platform above CUDA, and this deal is partly defense
- Andrew’s read — which Ben endorses as “the NVIDIA-specific question” — is that as everyone gets cost-sensitive and Google brings TPU infrastructure online, NVIDIA wants to encourage buildouts using NVIDIA hardware and software.
- Ben’s history: pre-ChatGPT GTCs threw every parallel-computing library at the wall; he recalls GTC 2024 as oddly boring because LLMs, while sending the stock to the moon, “were bad for NVIDIA” — the developer platform moved far above where NVIDIA sits. “No one who’s writing an AI application today is using CUDA”; apps run on OpenAI or Anthropic APIs, or Bedrock on Trainium with a Chinese open-source model, “totally abstracted away.” The moat “has been tremendously diminished.”
- The earned-it caveat Ben insists on: NVIDIA almost went under building CUDA when nobody understood why, bottoming out as recently as October 2022 — three weeks before ChatGPT, when Ben wrote “NVIDIA in the Valley.” “They have earned every dollar they’ve gotten through 25 years of taking massive risks.”