Dylan Patel on the infrastructure powering the AI revolution | The Next Big Thing
Dylan Patel on the infrastructure powering the AI revolution | The Next Big Thing
Summary
- The memory call is the episode’s spine: SemiAnalysis flipped from “memory is the biggest loser from AI” (early ‘23) to biggest winner when o1 led Patel to predict KV cache would explode, and Patel now says “it’s a shortage that’s going to last years” — capacity growing only 20-30%/yr for the next three years while demand doubles. Pricing is up ~4x with “another 2x, 3x” coming; the inelastic buyers get forced out (mid/low-end Chinese smartphone shipments already down 40%), iPhone and MacBook prices “have to go up” next year — “a few hundred bucks,” not $100 — until “AI gets its fill.” Gross margins head to 85-90% before halving back to the 70s or lower.
- CPUs: real inflection, but a mini-cycle, not a supercycle. Reinforcement-learning environments and agentic tool calls genuinely inflected demand — SemiAnalysis called it in November institutional research and ARM, Intel, and AMD have “ripped” since — but the sell-side “who doesn’t really understand technology at all is just making up stuff.” The math: at ~$5k per CPU versus “$50 something thousand” per Blackwell, $300-500B of Blackwell implies only $30-50B of CPU sales; today’s frenzy is a one-time catch-up on ~10M GPUs or other AI chips shipped over three years with no CPUs attached. “This market was underpriced and it’s more fairly priced now.”
- CPO is pushed out — favor copper instead. “People are a little bit too excited on CPO”: not 2027, “tail end of 28, but 29 is the real ramp” for scale-up co-packaged optics. Reuben is all copper and Fineman on the GPU is still copper, so likely Amphenol and non-CPO optics are favored medium term while CPO manufacturing and yields catch up.
- Power is the binding constraint on a buildout going 20 GW this year → 30 GW → 50 GW, and the answer is increasingly on-site: within a couple years half of incremental datacenter power will be generated behind the meter — from GE Vernova combined-cycle turbines down to truck engines converted to gas and serviced by car mechanics (“it’s a pain in the ass, but it will work”). Solar-plus-battery undercuts gas in ~2 years; Nvidia dropping 800V from likely Reuben Ultra Kyber pushes that conversion supply chain out.
- The ROI debate, answered with receipts: Anthropic was free-cash-flow positive and profitable in April and May with June trending the same, revenue past $50B ARR at 70%+ gross margins — “Anthropic is printing.” SemiAnalysis’s own AI spend went from under $100k annualized in November to ~$11M today (peak week annualizing $14M) for a 90-person firm once Claude Code inflected; companies clamping down on AI spend “are going to get left in the dust.”
- Cost optimization means the newest model, not the cheapest: a task that took Claude 4.6 Opus 100,000 tokens over several turns takes 4.8 Opus 25,000 in one shot. Token efficiency, not benchmark edge, is why “Anthropic has been beating OpenAI” — OpenAI’s models win frontier edge cases but burn 3-4x the tokens with a slower human feedback loop. For AI baked into fixed processes, the opposite logic: freeze quality and ride the ~60x/yr cost decline (DeepSeek was 600x cheaper than likely GPT-4 roughly two years later, versus 3,600x implied by two 60x steps).
- Blackwell measured 30x faster than Hopper somewhere on the continuum on DeepSeek V3 in SemiAnalysis’s open-source InferenceX benchmarks — above even Jensen’s ridiculed 25x launch claim (“Jensen, I was wrong. You were sandbagging”), and Jensen spent five minutes at GTC on Patel’s charts and “Inference King” belt as proof he doesn’t sandbag numbers.
- The reusable framework for every “next shortage” (MLCCs, PCB drill bits, copper foil): flow-through × elasticity × market structure — how many cents of each AI dollar reach the product, whether pricing is spot-commodity (memory) or partner-stable (TSMC takes “5, 10%”), and whether the market is an oligopoly or hundreds of names traded across Taiwan, Japan, and Korea.
Deep dive
1. From tween shitposter to 90-person research firm
- Patel’s own origin story: “the origin of SemiAnalysis really comes from shitposting” — posting about smartphone SoCs and display specs “before I ever even had a smartphone,” moderating Android/Intel/Nvidia forums by age 12, then two years as a quant: “yes, you make money, but it’s not as amazing as it seems.” He quit in 2020 for a WordPress blog under his real name, shaped by growing up living in his parents’ motel in rural Georgia — “I kind of just know business.”
- The first post set the template: when Huawei lost access to TSMC, the U.S. market thought Qualcomm would win; Patel called MediaTek the biggest winner because “geopolitically China would rather buy from a Taiwanese firm than from a US firm.” Technology, supply chain, finance, and geopolitics melded — fed by 40 conferences a year up and down the stack, some 300-person and Japanese-only: “I didn’t live anywhere.”
- The pivot to an institutional firm came via hire #3, Myin — a hire with a hedge-fund background who answered a hiring note buried in the paid section of the early-‘23 post arguing memory was the biggest loser from AI. Models and data services followed, and “the ball started tumbling down the hill”: headcount 2→7 (2023-24), 7→20, 20→60, now 90, with 30 added this year.
- The moat, as Patel tells it, is talent density no one else has: ex-ASML, Applied Materials, and Lam Research engineers upstream; ex-Intel, TSMC, Nvidia, OpenAI, Tesla FSD, and someone likely from Cohere downstream; “someone at my company who built a power plant in Kazakhstan.” The other half: ex-hedge-fund, or “random people from the internet who are super passionate.”
2. Jensen on stage: “Dylan said I was sandbagging — but I wasn’t”
- InferenceX is SemiAnalysis’s open-source benchmarking suite running every single night — because any night a CUDA, PyTorch, driver, or inference-engine version can drop — across eight GPU types plus Google TPUs and Amazon Trainium, on $50M+ of donated hardware from OpenAI, Microsoft, Amazon, Google, CoreWeave, Nebius, Crusoe, and Oracle.
- The call it produced: Jensen claimed 25x Blackwell-over-Hopper at launch and “a lot of people were like, no, no, no, it’s like 3x”; SemiAnalysis’s simulator said 15-20x. InferenceX then measured 30x faster than Hopper somewhere on the continuum on DeepSeek V3. Patel emailed: “Jensen, I was wrong. You were sandbagging.”
- At GTC — 20,000 people in the stadium — Jensen put Patel’s charts and the WWE-style “Inference King” belt on stage for five minutes as proof he doesn’t sandbag numbers: “He talked about us longer than anyone else in the entire presentation.” The only thing that got comparable airtime was OpenClaw.
3. Anthropic is printing — and SemiAnalysis’s own AI bill went over 100x
- Patel’s first answer to the ROI skeptics: Anthropic was free-cash-flow positive and profitable in April and in May, with June trending the same; recurring revenue “soared past $50 billion ARR” at gross margins above 70%. “Anthropic is printing.” OpenAI’s revenue is inflecting too as Codex adoption grows.
- The demand side, from his own books — he calls it ARS, “annual reoccurring spend”: under $100k in November (a $200 ChatGPT seat per employee), $4M by end of January once Claude Code hit its inflection with Claude Opus 4.5 and 4.6, ~$11M today with a peak week annualizing at $14M. AI is already more than a third of employee spend, “probably half by the end of the year depending on how Methos and other models” land.
- For good developers on ~$300k salaries, AI spend “is starting to approach one to one” — and at SemiAnalysis, “a lot of our biggest spenders are people who don’t know how to code,” who just “iterate, iterate, iterate.”
- The corporate fork in the road: companies blew through full-year AI budgets by Q2. Some cut legacy SaaS, some cut employees instead of AI, some clamp down on AI — “but those companies are going to get left in the dust in terms of productivity gains.”
4. Cost optimization means the newest model, not the cheapest
- Patel splits AI workloads in two. Process-integrated AI (check every inbound document for XYZ): hit a quality bar, freeze it, then ride the cost curve down — models get ~60x cheaper per year at fixed quality. “People freaked out about DeepSeek because it was 600 times cheaper than likely GPT-4.” About 2 years after that comparison, the 60x annual curve would imply 3,600x; it actually ended up 600x.
- AI-as-assistant inverts the logic: “cost optimization is often times taking the newest model.” A task that took Claude 4.6 Opus 100,000 tokens and several back-and-forths takes 4.8 Opus 25,000 tokens in one shot — fewer tokens, less human time, lower true cost.
- Token efficiency is “the main reason why Anthropic has been beating OpenAI”: OpenAI’s models can crack edge cases in leading science, math, and code that Anthropic’s cannot, “but they take 3x as long and 4x as many tokens,” and the human-in-the-loop feedback cycle degrades. So SemiAnalysis remains “a majority Anthropic shop,” reserving overnight tasks for OpenAI Codex.
- After both the 4.6→4.7 and 4.7→4.8 Opus releases, his cost fell for about a week — then soared past prior highs as people adjusted: “okay, the work I was doing is done, let me do more.”
5. Memory: a shortage measured in years, paid for by your next iPhone
- The intellectual arc matters: SemiAnalysis called memory the biggest loser in early ‘23 (AI servers carried far less memory content than regular servers), then flipped in December 2024 when o1 launched — reasoning led Patel to predict KV-cache usage would explode, and while weights are read identically at 1,000 or 100,000 tokens of context, KV cache memory reads scale with context while compute barely moves. Conclusion then: “memory was going to be the biggest winner.”
- The January 2026 note doubled down when people asked if +50% was the top: “no, no, no — I don’t think you guys get it.” Capacity grows 20-30% a year for the next three years; demand is doubling. “This is not a short-term shortage. It’s a shortage that’s going to last years.”
- The clearing mechanism is brutal: prices soar until inelastic buyers drop out. Chinese mid/low-end smartphone makers like Xiaomi report shipments down 40%; the high end is untouched so far, which is why “next year iPhone prices have to go up, next year MacBook prices have to go up” — and since $100 won’t move that market, “they’re going to have to go up a few hundred bucks” until “AI gets its fill.”
- He is explicit this stays cyclical: pricing already ~4x with 2-3x more coming, memory margins headed toward 85-90% gross — “memory necessarily doesn’t deserve a margin of 85%” — before halving back to the 70s or lower. Cycles survive; the trough just keeps rising.
6. The framework: flow-through × elasticity × market structure
- The episode’s reusable tool, laid out before every sector answer: for each dollar of AI spend, how many cents flow to this product — “it might be one cent on this product, but it might be 5 cents on this product”? Is the end market up 50%, doubling, or quadrupling? Then, who can reprice: TSMC is non-elastic — “pretty fair with their customers… we’ll take up price 5, 10%” — while memory lets spot and contract markets clear, so it 4x’s. ASML barely oscillates at all.
- Apply it to the tail: MLCCs, PCB drill bits, copper foil — “you’ll go online and see this is the next shortage, this is the next shortage.” The local bumps are real but “very small and many,” and the names trade in Taiwan, Japan, Korea — “not just easily accessible to investors.”
7. CPUs: agents and RL made the forgotten chip more in demand
- SemiAnalysis flagged it in November institutional research: OpenAI and Anthropic were striking deals to rent essentially all the CPUs in Amazon’s, Google’s, and Microsoft’s fleets. The host’s observation: three years of AI without hearing “CPU,” and now it’s everywhere.
- The mechanism: pre-training barely touched CPUs, but reinforcement learning checks every generation against an environment — unit tests, compilers, sandboxed websites and shopping flows — and agentic inference makes constant tool calls into the regular world (searches, databases, Python interpreters). Both are CPU-hungry in ways chat never was.
- A third demand leg: deployed output. GitHub commits are up “multiple X” versus last year — “a lot of the code is sloppy, but a lot of code is being deployed” — and web scrapers and business-process automations land on standard, cost-effective CPU cores.
- The winners map: Intel and AMD both raised prices; ARM entered and its stock “has gone gangbusters”; Amazon extracts “incredible margins” renting Graviton rather than selling it; Nvidia’s standalone Vera carries $20B of CPU revenue guidance.
8. But CPUs are right-sizing, not the next supercycle
- Patel’s pushback on his own call’s momentum: “the sell side, who doesn’t really understand technology at all, is just making up stuff” — ratios now imply more CPU than AI compute, which is “false.” A full Blackwell runs “$50 something thousand per chip” versus ~$5,000 per CPU; even a 1:1 ratio on $300-500B of Blackwell yields only $30-50B of CPU sales.
- What’s actually happening is a backlog catch-up: roughly 10 million GPUs or other AI chips shipped over three years “that don’t have any CPU attached really.” Once that fleet is matched, demand drops to the incremental attach rate — “we’re in sort of a mini cycle of CPU.” His verdict: “this market was underpriced and it’s more fairly priced now.”
- The design space splits by workload, per the core-count law — double a core’s size and per-core performance rises only ~50%. Nvidia’s Vera bets on fewer than 100 fast cores for workloads where AI compute stalls waiting on a CPU response; AMD’s 256 cores and Graviton offer the more-core option when workloads are highly batched and no one is waiting on any single core. “For some workloads you do want Vera and for some workloads you want the Graviton or the AMD CPU.”
9. CPO slips to 2029 — favor copper in the meantime
- Networking content is growing faster than any other category — from sub-10% to above 10% of AI-chip-associated spend, and 20-30% once CPO arrives; telecom optics—likely including Ciena—have been ripping.
- But on co-packaged optics itself: “people are a little bit too excited.” “It’s not coming in 27 in my view — really the tail end of 28, but 29 is the real ramp for scale-up co-packaged optics.” Manufacturing volumes, yields, and chip designs simply aren’t there; switch-level CPO comes earlier than GPU-level.
- The Monday institutional note: medium-term bullish copper and non-CPO optics, “kind of bearish on CPO” — Reuben is all copper, Fineman on the GPU is still copper, and Reuben is likely only just starting to ship. So backplane makers—likely including Amphenol—“are actually going to do way better over the next few years than previously expected.”
- Why copper keeps winning locally: “at the end of the day, integrating optics is so much more expensive than sending something electrically” — until distance forces repeaters or optics. Close your eyes for five years and optics is “way bigger”; some of that is priced in, some isn’t.
10. Power: from truck engines to space, the constraint that yields to money
- The scale: 20 GW of datacenters deployed this year, 30 GW next year, 50 GW the year after — gated by energy first, politics second, construction third. Of the three power legs, transmission is “the hardest to be bullish on” (utility monopolies, cost amortization rules); generation and conversion are where the action is.
- Behind-the-meter is the release valve: SemiAnalysis predicts half of incremental new datacenter power will be generated on-site within a couple of years. The supply chain runs from GE Vernova/Mitsubishi/Siemens combined-cycle gas down to train, boat, and truck engines converted into power generation; truck engines can be converted to gas, back-driving electric motors, buffered by batteries, serviced by “a bunch of people from car mechanic shops” — 10 GW+ of datacenters planned on such tech. “It’s a pain in the ass, but it will work” — the continuum runs from “going full dirty” to “fully into space,” where solar panels need no battery at all.
- The next crossover: “in about 2 years, solar plus battery will be cheaper than gas” thanks to China’s manufacturing scale — with the caveat of reliability nines (enough battery for one night is cheap; three rainy days is not).
- Conversion is its own supply chain — IGBTs, silicon carbide, GaN MOSFETs, the 12V→54V→800V DC transition, solid-state transformers, supercapacitors — and it just took a hit: likely Reuben Ultra Kyber no longer has 800 volt, pushing that chain out. Fittingly, SemiAnalysis’s biggest research vertical isn’t semiconductors — it’s the “DEI team” (datacenters, energy, industrials, “it’s a pun internally”), tracking every datacenter and power plant, because “Google’s interested in what Meta is able to deploy.”
Verification Notes
- Raw captions contain an internal contradiction: “memory isn’t a shortage” is followed by “It’s a shortage that’s going to last years”; the digest follows the latter.