Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis
Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis
Summary
- The core call: co-design is the 100x. Optimize hardware, systems software, and model separately and three 2x’s multiply to 8x; co-optimize across all layers and “it’s actually 100x.” The consequence is model-silicon lock-in: DeepSeek V3’s expert shapes were built for Hopper (V4’s for Blackwell and Huawei’s chip), so “TPUs suck at running DeepSeek” — and “the way OpenAI’s models are headed, it would be a terrible decision for them to use TPUs,” with the mirror image true for Anthropic and Google training on GPUs.
- The CUDA moat isn’t CUDA anymore. Models write kernels now — “all software gets commoditized” — so the moat has migrated to the ecosystem: Chinese open models (DeepSeek, Kimi, Alibaba) are co-designed for Nvidia GPUs, dragging every downstream inference API and RL company onto them. If Google open-sourced really good models, the same gravity would pull toward TPUs — hence Gemma.
- The compute crunch is demand-led and durable. 20 GW comes online this year, 30+ GW next even counting delays, but Fable/Mythos 5’s TAM is “not just 2x that of Opus” while world compute didn’t double in the seven-odd months since Opus 4.5. Anthropic was net-income profitable in Q2 ex-stock comp, with Opus 4.8 API token margins “north of 80%” — so it can rent GPUs above market and still print: “whatever price I want to pay, I can pay.”
- Gigawatts aren’t fungible. Trainium rents at sub-$10B per GW, GPUs at $12–13B, the SpaceX–Google deal at ~$25B per GW; colocation has gone from $60/kW/month to $120–160, up to $200. Google stuffs ~1.5 GW of hardware into a 1 GW shell by sloshing power, and “a gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI.”
- Everyone builds ASICs, but general-purpose survives — labs “literally don’t know what architecture they’re going to be doing in a year,” so specialized silicon risks racing to a local minimum. Google runs three different-architecture TPU programs yet still pays xAI $11 per GPU-hour. Cerebras’s risk: dollars concentrate in the best model, and if frontier models hit 10T+ parameters with million-token context, SRAM-based chips can’t fit them.
- Jensen is engineering a multipolar world: a world where hyperscalers own all compute or three closed labs own all models “is one in which he’s screwed,” so he backstops neoclouds and neolabs and loves Chinese labs — “you throw a bunch of bait into the water and the best fish will figure out and survive.”
- Space compute is real but not yet: sub-1% of compute by 2030, but “more than half of incremental compute” by 2040. Scale forecast: OpenAI plus Anthropic alone exceed 100 GW combined by 2030, terawatts by 2040, with inference “many percentage points of GDP.”
- The bear case he respects vs the one that infuriates him: Sonya’s leverage worry (a Crusoe customer just asked to halt a buildout) is live if models’ useful work stops expanding faster than compute — “that tide turns.” But “AI has no ROI” infuriates Dylan, while Shaun’s model-progress denial is captured by: “the line has been up and to the right in terms of capabilities this entire time.”
Deep dive
1. From motel step-stool to Starcraft grandmaster — the making of an obsessive
- The origin story as told: his parents ran a motel and a gas station, and he was too short to reach the cigarettes — so he learned to pre-position the step stool by who walked in. “The first neural network I trained” was racially and visually profiling people based on when they enter the gas station, which cigarette to pick. The small-business economics tinge never left.
- His Xbox 360 hit the red ring of death; he opened it up and shorted the temperature sensor — “opened Pandora’s box.” By 12 he was moderating Reddit hardware forums, and even as a teen argued Nvidia over the crowd-favorite AMD on margins: “they use a smaller chip to get better performance at better power efficiency.”
- The pattern is obsession: at one point grandmaster on the North American Starcraft 2 ladder. “Obsession is good.” Grades? “Fine enough for Asian parents.”
2. The 2020 crash-out behind SemiAnalysis
- The founding was a pile-up: he got “screwed out of a bonus” after making his quant firm well over $10M of risk-free revenue, his grandmother developed dementia and died after a fall, COVID hit. Then an internet argument got him doxxed — he stopped posting for three weeks, asked “what am I doing? Why do I care?”, and launched the SemiAnalysis blog on his 24th birthday.
- He was then effectively homeless from mid-2020 through 2024: a truck-and-tent tour of national parks reading semiconductor textbooks, then Latam, then 40+ conferences a year anywhere in the supply chain. SPIE lithography humbled him — understood ~10% the first time, half the second, 75% the third: “some parts of the supply chain are so arcane and so deep… even now I don’t understand everything.”
- The lore that stuck, from a Japanese guy in broken English: in the 1980s the world’s only factory for one chemical burned down and memory prices doubled or tripled. “Not too different from today.”
- Today SemiAnalysis is ~90 people — a big chunk are supply-chain technologists and engineers, and another big chunk are people formerly at hedge funds — and the internal fights (“that doesn’t matter” / “but cost” / “this technology is the coolest”) happen organically. On the $100M revenue rumor: “It’s as accurate as the information is.”
3. InferenceX tracks 60x model-cost declines
- Point-in-time benchmarks die on arrival: new models weekly, vLLM/SGLang/PyTorch updating twice a week — “relentless breakthrough after breakthrough” that has driven model cost down ~60x a year for equivalent quality. So InferenceX runs daily, automated, on ~15 chip types across the best Chinese and US open models, on over $50M of donated hardware (over $100M once TPUs and Trainium land) from CoreWeave, Crusoe, Nebius, Oracle, Microsoft, Amazon, Google, and OpenAI.
- The point is the Pareto-optimal frontier: vendors compare their optimal point against a rival’s suboptimal one — “if I drove a Porsche versus some race car driver, obviously I’d drive it slower.” InferenceX open-sources containers for the optimal configuration at every point on the curve.
- His master framework: the throughput-vs-interactivity curve — “most things in hardware, infrastructure, model, application layer — everything is downstream of that curve.” One-size-fits-all inference ends: batch workloads take the 4x cost saving, latency-sensitive users pay 4x up — Claude Code fast mode and OpenAI’s priority queue are examples of price-segmented points.
4. Space compute and the intelligence-per-watt ledger
- On what share of inference happens in space: sub-1% by 2030, but “in 20 years I think the vast majority of compute will be going in space” — by 2040, “probably more than half of the incremental compute.” He’d buy the SpaceX IPO if he could buy stocks — “not investment advice,” stated twice.
- The scale behind that: by 2030 “just OpenAI and Anthropic will have over 100 gigawatts combined,” terawatts by 2040, with inference bigger than oil — “many percentage points of GDP.”
- Intelligence per watt has improved roughly 40x (versus 60x on cost). Versus the human brain “we’re many orders of magnitude away — thankfully, doesn’t really matter”: it’s much easier to power computers than humans, who come with “sickness, disease, food preferences… sleep.”
5. The core thesis: co-design turns 8x into 100x
- Shaun’s framing — three layers (hardware, kernel-level systems, model), with recent gains mostly from hardware — gets flatly rejected: “Shaun, I completely disagree with you.” Hopper to Blackwell gave ~30x on DeepSeek at the most optimized deployment, but the model layer gave more: three years ago the frontier was GPT-4; today a small Qwen — “27B parameters total and like 2 billion active” — is way better.
- The mechanism: DeepSeek V3’s expert shapes were optimized for Hopper; V4’s for Blackwell and Huawei’s chip. Result: “TPUs are objectively an amazing chip… [but] TPUs suck at running DeepSeek.” Shapes, network IO, collectives, and attention arithmetic intensity are co-optimized through the stack — “it’s hard to say you can disentangle the gains.” Same at Google: each Gemini generation is built for that generation’s TPU, and pulled onto old hardware “it’s really not that great.”
- On the idea that China does this better: no — “the West doesn’t tell people what they do.” GPT-4o was roughly the same size as DeepSeek V3, slightly smaller, and came out earlier.
- The money line: layer-by-layer optimization gives 2x·2x·2x = 8x; co-design across all three and “instead of being multiplicative to 8x, it’s actually 100x.” A company likely associated with Naveen Rao (a Sequoia investment) is the long-horizon version — silicon, software abstraction, and model simultaneously, potentially analog compute with energy-based models: “Probably won’t work, but that’s exciting… definitely won’t work quickly.”
6. The CUDA moat isn’t CUDA anymore
- Shaun’s observation: model companies no longer fear other chips — “Claude and Codex are actually quite good at doing a lot of that optimization work,” and there are only tens of model companies, not the thousands the CUDA-moat thesis assumed. Dylan concedes: “the CUDA moat and software moat is at least partially disentangled… models are just great at coding and all software gets commoditized.”
- What remains is ecosystem gravity: DeepSeek, Kimi, Zhipu (likely), Alibaba, Tencent (likely), Xiaomi co-design their open models for GPUs, so everyone downstream — inference API providers, RL customization shops — concludes “I guess I need to use Nvidia because the ecosystem uses Nvidia.” If Google open-sourced really good models, the pull would reverse toward TPUs — that’s the point of Gemma.
- Big labs don’t need the open stack at all: OpenAI forked PyTorch long ago. The frontier game is “I’ll choose the best hardware and I’ll co-design my model and infrastructure software through and through… and I’ll have AI help me write all that software.”
7. Nvidia vs TPU: both win — the real risk is local minima
- He refuses the binary: two years out, Google makes 10M+ TPUs (~$100B+ of TPU a year) while Nvidia ships tens of millions of GPUs ($500B+, “not a specific estimate”). “I could with a straight face argue GPUs are way better than TPUs or TPUs are way better than GPUs” — the resolution is co-design: “the way OpenAI’s models are headed, it would be a terrible decision for them to use TPUs,” and potentially terrible for Anthropic and Google to train on GPUs.
- The concrete divergences: OpenAI’s models are much sparser, Anthropic’s sparser-but-more-dense; matrix-multiply unit sizes differ; NVLink switches connect 72 GPUs while Google’s switchless ICI connects 8,000 chips at high bandwidth (passing through other chips). Google runs three separate TPU design programs — Broadcom, MediaTek, and a third undisclosed one — with genuinely different architectures.
- End state: everyone deploys billions to tens of billions of their own ASICs (Google: hundreds of billions a year), but general-purpose compute survives because labs “literally don’t know what architecture they’re going to be doing in a year” — a specialized chip can race to a local minimum and be stranded by one attention-mechanism breakthrough. Even Google pays xAI $11/hour per GPU, and some non-Gemini Google bets (drug discovery or Waymo — “I won’t say which one”) run primarily on GPUs.
- On Cerebras: “really innovative” — SemiAnalysis uses fast mode “almost exclusively” — but the risk is that fast mode matters most on the best models, revenue concentrates there (large numbers of users switched to Fable and Mythos on launch day despite the price), and a 10T+ parameter model at million-token context won’t fit on SRAM chips like Cerebras and Groq. On tokens versus dollars: “who cares about volume by tokens? It’s about the dollars” — his F-150s-vs-Camrys point.
8. The physical bottlenecks: ancient memory cells, the 1W/mm² wall, diesel engines
- Memory, from the technology angle rather than supply chain: the NAND cell is ~25 years old, DRAM ~40, and HBM progress has just been more, faster stacks. Coming in the next few years: stacking memory directly on the compute die, which “makes your bandwidth explode.”
- Power density has been pinned at ~1 watt per mm² for two decades — chips hit 1,400W today, 2,000W with Rubin, ~4,000W with Rubin Ultra, all by adding silicon. Breaking past that wall (in development now) means less silicon per chip, at the cost of hard thermal and electrical-interference engineering. Co-packaged optics is a when-not-if: “the debate is like ‘27, ‘28, ‘29, 2030.”
- Energy has unglamorous fixes: convert the millions of diesel truck engines the US can build to gas, back-drive electric motors as generators, and staff the sites by pulling mechanics out of car shops. The deeper Western problem, in Shaun’s framing that Dylan seconds: “why would you want to go work in hardware when you can make ads to ads?”
9. The crunch is real — and Anthropic can pay above market
- Supply is compounding — 20 GW this year, more than 30 GW next, delays included — but demand compounds faster: Mythos 5/Fable 5’s TAM is “not just 2x that of Opus,” yet world compute didn’t double in the seven-to-eight months since Opus 4.5.
- The economics beneath it: Anthropic was net-income profitable in Q2 excluding stock-based comp, maybe including it by Q3, with Opus 4.8 API token margins “north of 80%.” At 75% gross margin, doubling compute cost still leaves 50% — so “every GPU I rent… I can immediately turn around and sell tokens on it at a positive margin… whatever price I want to pay, I can pay.” They even bought GPUs above market from SpaceX (still below what Google later paid, having signed earlier).
- Sonya’s pushback — worth keeping: everyone is levered to “we got to build,” and Crusoe publicly said a customer asked to halt construction on a buildout; “high leverage high growth makes me very nervous as an investor.” Dylan teases (“small amount of equity has huge upside — you’re not a debt investor… go to the school of private equity”) but concedes the crux: if models’ economically valuable work stops expanding faster than compute capacity, “that tide turns.”
- His base case is it doesn’t turn: models are improving faster than six months ago via a “pseudo recursive self-improvement loop” — models writing the infrastructure that launches the next model sooner. Capital is the binding constraint: Google raised despite ~$100B of SpaceX stock unlockable in nine months (Larry Page’s $1B at a $10B valuation — “one of the greatest investments of all time. Good job, Larry.”); Meta announced a raise and the stock tanked. And the take that “infuriates me”: “AI has no ROI” and model-progress denial — “bro, the line has been up and to the right in terms of capabilities this entire time… look at the new benchmark, they’re skyrocketing.”
10. Gigawatts aren’t fungible — Jensen wants a multipolar world
- Pricing already discriminates: Trainium rents at sub-$10B per gigawatt, GPUs historically $12–13B, and the SpaceX–Google deal “was like 25 or something crazy billion dollars per gigawatt — $25 million per megawatt a year.” Colocation went from $60/kW/month to $120–160, as high as $200 (weak-credit tenant, good facility), as low as $80 in India. And on the compute layer: “a gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI” — both could sell every gigawatt they have, especially since Codex 5.5.
- Operational skill compounds it: Google puts ~1.5 GW of hardware into a 1 GW datacenter by sloshing power around from workload knowledge, and cuts utility deals for 2 GW “except three days of the year.” Shaun’s addition: SpaceX’s Starlink networking and Tesla power-management experience “might be missing from the analysis a lot of people are doing” — and SpaceX sells compute that’s “running now, buy it,” versus Google selling six months forward on paper to finance its own purchase orders.
- Why neoclouds exist at all — his 2023 “Amazon Cloud Crisis” report: hyperscaler advantages (Nitro NICs, tenant isolation, custom SSDs, custom networks) were built for time-sliced CPU clouds and turned neutral-to-detrimental when AI customers rent whole racks on long-term contracts; Microsoft’s datacenter team “fell on their face” when forecasts doubled. Plus incentives: nobody at a hyperscaler gets rich building faster, while Crusoe’s “hyperlevered equity owners” do.
- Jensen’s chess, in Dylan’s telling: “Jensen absolutely hates a world where all the hyperscalers have all the power” — a world of only OpenAI/Anthropic/Google models, or only hyperscaler compute, “is one in which he’s screwed.” So he backstops neoclouds, funds neolabs, and champions Chinese labs — because Crusoe and CoreWeave existing in five years makes TPU and Trainium weaker. “You throw a bunch of bait into the water and the best fish will figure out and survive.” Early returns: Thinking Machines’ Tinker at a few hundred million dollars of ARR under about six months from launch.