Pioneers Insight Method Research Author
Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China
Back to Episodes

Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China

Summary

  • The Nvidia–Intel tie-up is a strategic endorsement that could lower Intel’s cost of capital while redrawing the PC and data-center map. Nvidia committed $5 billion after SoftBank’s $2 billion and the U.S. government’s $10 billion, but Dylan Patel still thinks Intel needs roughly $50 billion; Jensen Huang supplies a semiconductor “Buffett effect” before a larger capital raise. Guido Appenzeller called the integrated x86-plus-Nvidia product compelling and warned that when “two arch nemeses suddenly team up,” AMD and Arm face the worst possible news.

  • Huawei is technologically credible, but manufacturing volume—especially high-bandwidth memory—remains the load-bearing constraint. Huawei reached market with a 7 nm Ascend AI chip in 2020, later obtained roughly 2.9 million TSMC-made chips through intermediaries, and now proposes separate prefill and decode products with custom HBM. China can probably produce substantial 7 nm logic and perhaps reach 5 nm with existing equipment, yet Patel stressed that design is not production: HBM3 yields, etch capacity and the transition from stockpiles to domestic scale remain unresolved.

  • China’s rejection of Nvidia chips is simultaneously industrial policy, a risky capacity bet and potentially negotiating leverage. ByteDance and other model builders still prefer Nvidia because it is “way better,” but Beijing can compel domestic adoption while Huawei publicizes an ambitious roadmap—a maneuver Patel called “10,000 IQ,” with Washington “playing checkers while they’re playing chess.” China may temporarily backtrack if domestic supply cannot ramp fast enough, forcing a choice between sovereignty and deploying “super powerful AI” at U.S.-competitive volume.

  • The near-term Nvidia bull case rests on AI capex running materially above Wall Street’s model, not on further market-share gains. Bank consensus puts six hyperscalers at roughly $360 billion of capex next year; Patel’s data-center and supply-chain work points to $450–500 billion, with Nvidia largely growing alongside the market while defending share. OpenAI alone signed more than $300 billion with Oracle, and the industry bull case becomes “multiple trillions a year on AI infrastructure”—though Patel refuses to forecast beyond five years because “the fifth year is sort of YOLO.”

  • Nvidia’s moat is a repeated willingness to risk inventory, redesign late and ship working silicon before competitors finish revising theirs. Jensen allegedly ordered Xbox volume before Microsoft formally awarded the business, placed non-cancellable capacity bets above customers’ own plans, and added Volta’s tensor cores only months before fabrication. Nvidia commonly ships A0 silicon while one Intel data-center processor reached E2—roughly 15 revisions—capturing the cultural difference between “I hate spreadsheets. I just know” and quarter-by-quarter caution.

  • Amazon and Oracle can both gain AI-cloud share without having the best accelerator, because powered capacity and balance-sheet willingness are scarce. Patel expects AWS revenue growth to trough and reaccelerate above 20% as Anthropic, Trainium and GPU deployments fill Amazon’s spare capacity; Trainium remains “very hard to use,” but a lab serving a few high-volume models can hand-optimize it. Oracle’s hardware-neutral engineering and willingness to underwrite OpenAI’s demand make it the bolder counterparty, although whether OpenAI can pay more than $80 billion annually in 2028–29 remains the central risk.

  • GB200 economics are workload-dependent, and reliability can erase headline performance for teams lacking sophisticated infrastructure. Against roughly 1.6x H100 total cost, GB200 may offer only about 2x performance in some training cases but north of 6–7x per GPU for DeepSeek inference; the problem is that one failure now sits inside a 72-GPU NVLink domain. Some clouds consequently promise about 99% availability for 64 GPUs but only 95% for all 72, making “the blast radius of a failure” as important as benchmark speed.

  • The GPU market is tightening again while Nvidia’s next strategic problem becomes what to do with potentially $250 billion of annual free cash flow. Patel also used an ambiguous “$200 million” figure in the same sentence. Hopper capacity at several major neoclouds sold out as reasoning-model inference surged, Blackwell deployment took longer than Hopper, and prices bottomed months ago before creeping upward—small allocations remain easy, large immediate clusters do not. Patel’s preferred outlet is data centers and power, the bottlenecks to GPU growth, but even that may not absorb the cash without turning Nvidia into a culturally different company.

Deep dive

1. Nvidia’s Intel investment converts rivalry into strategic dependency

  • Patel’s opening reaction was financial theater: Nvidia announced a $5 billion Intel investment, Intel’s stock jumped roughly 30%, and the stake was already “a billion-dollar profit” by his telling. More importantly, a major prospective customer is committing capital and product roadmaps, giving Intel validation before it returns to public debt or equity markets.

  • The product logic is unusually strong. Intel would package one of its chiplets beside an Nvidia chiplet, reversing a history in which Intel faced antitrust litigation over chipsets and paid Nvidia a settlement. Patel called the turn “poetic”: Intel is “sort of crawling to Nvidia,” yet an x86 laptop with fully integrated Nvidia graphics might be the market’s best device.

  • The announced checks remain small against Intel’s needs: $5 billion from Nvidia, $2 billion from SoftBank and $10 billion from the U.S. government, versus Patel’s prior estimate that Intel needed roughly $50 billion immediately. More strategic investors—he floated Apple as one possibility—could create a “Warren Buffett coming into a stock” effect before Intel raises the balance.

  • Appenzeller’s customer view was enthusiastic but brutal for competitors. Intel could reset its uncompetitive internal graphics and AI programs, including the “Gaudi F4” mentioned in the discussion; AMD now faces two historic enemies acting together, while Arm loses the pitch that it is the natural partner for anyone avoiding Intel. “It remixes the cards.”

2. Huawei entered sanctions as a near-frontier chip company

  • Patel asked listeners to begin in 2020, not with the current roadmap. Huawei submitted an Ascend accelerator to impartial public benchmarks, became the first to bring a 7 nm AI chip to market and had only a narrow technological gap with Nvidia before the ban cut off full foreign-supply-chain access. It had also surpassed Apple in TSMC orders.

  • The Trump administration’s restrictions severed that access just as Huawei had the ingredients to challenge the market. Nvidia accelerated while Huawei rebuilt around SMIC, pursued Korean memory and simultaneously used shell companies to place TSMC orders. Huawei had already trained significant models on its limited original Ascend inventory.

  • By late 2024, Patel said that intermediary channel had produced roughly 2.9 million TSMC-made chips from about $500 million of orders before authorities stopped it. He referenced reporting about a possible $1 billion U.S. fine for TSMC but explicitly was unsure whether it had been issued; importantly, visible Ascend deployment had not yet consumed all that inventory.

  • The 2025 H20 ban then removed what SemiAnalysis estimated as more than $20 billion of Nvidia China revenue. Nvidia wrote off inventory, later received permission to resell it and faced a harder decision: restart a supply chain that might again be interrupted, or let Huawei and Cambricon absorb a market whose present “domestic” capacity still uses foreign wafers and memory.

3. China can design frontier architecture sooner than it can manufacture it

  • On logic, Patel’s reading of the equipment rules was more permissive than their public label. Although Washington describes restrictions reaching 14 nm, he argued that the equipment actually blocked is primarily needed below 7 nm. China should therefore make substantial volumes of 7 nm AI chips and might stretch existing tools to 5 nm, albeit with harder economics.

  • Huawei’s roadmap splits inference into specialized products: one chip for recommendation systems and prefill, another for decode. Nvidia and multiple startups are making the same architectural separation, so Huawei’s surprising claim was not the split itself but custom HBM for decode—an approach Nvidia and AMD were also only preparing to adopt the following year.

  • The import trail still signals a bottleneck. China previously devoted 30–40% of equipment imports to stockpiling lithography, versus a historical fab mix around 17–18% and roughly 25% in the EUV era; now etch imports are surging. HBM needs through-silicon vias etched through each layer before stacks are assembled 12-high or 16-high.

  • Huawei had sampled HBM2 but, by Patel’s account, had not begun volume production of HBM3, a technology introduced years earlier. Equipment availability and yield learning both matter: China can catch up faster than it took to invent the technology because the process already exists, but a few months of imports cannot reproduce years of Korean capacity. Torenberg framed the eventual outcome as “a matter of when, not if”; Patel’s own emphasis was that manufacturing capacity and yields remain the bottlenecks.

4. Beijing’s Nvidia ban creates a dangerous transition gap

  • China can initially reject Nvidia because it is converting the 2024 stockpile into accelerators. The difficult interval comes when those parts run down before domestic logic and HBM reach volume. Patel expects China eventually to ramp, but “it’ll take a little bit longer,” potentially creating a period when Beijing backtracks and relaxes policy.

  • Commercial incentives point the other way from industrial policy. ByteDance was described as “begging for Nvidia chips”: it uses some Huawei and Cambricon hardware but wants Nvidia to build the best models and deploy inference efficiently. The government can mandate domestic purchases, yet that does not mean Nvidia has ceased to be competitive.

  • Smuggling and re-exportation continue at low to lower-medium volume, but cannot close a national deployment gap. China ultimately has to choose how heavily it weights an internal supply chain against keeping pace in powerful AI; otherwise it deploys “so many fewer AI chips” than the United States.

5. Huawei may be hyping strength precisely because it still wants Nvidia

  • Patel’s negotiating interpretation was delightfully adversarial: if China wants better U.S. chips, it should advertise domestic self-sufficiency, unveil “the most crazy” multi-year Huawei roadmap possible and announce that Nvidia is banned. American suppliers then warn officials that an irreplaceable market is disappearing and lobby for looser export limits.

  • He called that play “10,000 IQ”—“we’re here playing checkers while they’re playing chess”—while conceding that much of Huawei’s architecture is real. The exaggeration lies principally in manufacturing certainty: the roadmap treats hoped-for capacity and yields as if they already exist.

  • Export policy therefore cannot be reduced to selling everything or nothing. Patel proposed comparing China’s achievable volume at each performance tier, then deciding what U.S. products near or somewhat above that level can be sold. AI’s potential end market is far larger than semiconductor equipment, so maximizing current chip revenue alone is an inadequate objective.

6. Huawei is the competitor Jensen fears beyond China

  • Patel took Jensen’s description of Huawei as “formidable” literally. Before sanctions, Huawei surpassed Apple in TSMC orders and phone share across multiple markets, then began recovering without Western supply chains. Against that history, fearing Huawei more than AMD is not lobbying theater alone.

  • Jensen’s best argument is to make Huawei’s aspirational roadmap become geopolitical reality: persuade policymakers that manufacturing capacity is no constraint and that Huawei will capture China, the Middle East, Southeast and South Asia, Europe and Latin America. Patel’s objection is narrow but crucial—capacity and yield are genuine constraints, even if temporary.

  • The alternative strategy is to “Japanize” China: isolate it until hardware and software become hyper-optimized for its domestic market, like the unusual Japanese PCs Patel invoked, leaving the global platform to U.S. firms. But isolation can also force China onto a superior branch while Western hardware–software co-design settles into a local optimum. “I don’t know if it’s accurate, but it’s an interesting one.”

7. Nvidia’s measurable bull case is $450–500 billion of hyperscaler capex

  • Bank consensus had Microsoft, CoreWeave, Amazon, Google, Oracle and Meta spending about $360 billion next year. Patel’s site-by-site data-center, component and supply-chain model produced roughly $450–500 billion. Nvidia cannot meaningfully add share from its dominant base; the earnings lever is defending share while total infrastructure expands faster than consensus.

  • Oracle’s contract illustrates the magnitude. OpenAI committed more than $300 billion over several years, rising beyond $80–90 billion annually, despite lacking the present cash to pay it. OpenAI had reached roughly $20 billion ARR; estimates for the following year’s exit ranged from $35 billion to $45 billion, while projected annual cash burn ran around $15–25 billion before profitability targeted for 2029.

  • Stack similar revenue and fundraising across OpenAI, Anthropic and other labs, and $500 billion of hyperscaler capex becomes plausible. Nvidia’s maximal thesis is “multiple trillions a year on AI infrastructure,” with GPUs mediating business agents, coding and consumer companionship alike.

  • When pressed for Nvidia’s ultimate ceiling, Patel refused false precision. If AI repeatedly improves AI, value creation could reach hundreds of trillions, but supply chains are only visible three or four years out: “the fifth year is sort of YOLO.” Whether white-collar workers become twice as productive, replaced, or dependent on a constant token stream is beyond an investable five-year forecast.

8. Nvidia built its moat by repeatedly risking the company

  • Nvidia failed early and repeatedly made existential commitments. An industry veteran told Patel that Jensen ordered Xbox production before Microsoft formally awarded the contract—likely with more nuance than the legend preserves—but the order came first. Nvidia’s first successful chip similarly had to work from its only affordable mask set or the company would run out of money.

  • During crypto booms, Nvidia persuaded suppliers that demand was durable gaming, visualization and data-center growth, prompting them to build capacity. When crypto collapsed, Nvidia absorbed inventory write-downs; suppliers retained empty lines. AMD had more efficient mining silicon but declined to raise production aggressively, a sensible risk policy that surrendered the upside.

  • The behavior continued at hyperscaler scale. Nvidia booked non-cancellable, non-returnable capacity above Microsoft’s internal plan, then Microsoft raised its plan toward Nvidia’s number. Patel summarized the founder’s decision system with Jensen’s own line to his CFO: “I hate spreadsheets. I don’t look at them. I just know.”

  • Jensen’s pinball metaphor explains the time horizon: “The reason you win is so you can play again.” Winning funds the next generation, not a fixed 15-year plan. Founder memory matters here—he remembers nearly going bankrupt, yet concludes that Nvidia must keep taking comparable risks rather than optimize for predictable Wall Street quarters.

9. First-pass silicon and late redesigns turn vision into market share

  • Nvidia’s long-tenured engineering organization contains both visionary architects and operators willing to say, “We need to get this silicon out now.” Patel described one private engineering leader as almost mythical and another fellow as famous for cutting features that technologists love, preserving them for the next chip rather than delaying shipment.

  • The measurable edge is stepping. Nvidia often ships A0 silicon and occasionally A1; one Intel data-center processor reached E2, roughly its fifteenth revision. Each new stepping can cost about a quarter, so verification quality becomes a go-to-market weapon rather than engineering hygiene. Patel recalled that Intel was openly jealous that Nvidia “consistently delivered on the first revision.”

  • Nvidia also begins transistor-layer production, pauses before final metal wiring if necessary, then releases volume once validation arrives. That combination of simulation, verification and calculated work-in-process lets it respond before competitors finish revising their masks.

  • Volta is the canonical bet: after observing AI workloads on P100 Pascal, Nvidia added tensor cores only months before sending the design to fabrication. Without that late change, somebody else might have captured the AI accelerator market. Equally important, its software organization delivered drivers and infrastructure quickly enough for first-pass hardware to be useful immediately.

10. Nvidia’s balance sheet is becoming a strategy problem

  • Patel cited roughly $250 billion of annual free cash flow; the transcript also contains an ambiguous “$200 million” figure in the same sentence. Regulators did not allow an Arm acquisition, even the $5 billion Intel investment requires review, and no obvious acquisition can absorb hundreds of billions without creating antitrust or integration problems.

  • Small investments in CoreWeave, model labs and other neoclouds help diversify buyers, but remain “small fries.” Nvidia could finance an entire Anthropic, xAI or OpenAI round, yet selecting winners would alarm every unchosen customer and strengthen their motivation to adopt AMD, TPUs, startups or internal silicon.

  • The cleaner use is to expand the complement: data centers and power. Nvidia could finance data-center and energy infrastructure without operating the cloud layer itself, removing the physical bottleneck to GPU growth while letting increasingly capable cloud competitors handle rentals.

  • The unresolved risk is cultural. Pouring concrete and building power infrastructure require different people from designing accelerated computing, while endless buybacks invite comparison with Apple under Tim Cook—excellent supply-chain execution, but in Patel’s view little transformative investment for nearly a decade. “Nothing requires $300 billion of capital” while also offering an obvious return.

11. Amazon’s spare power can reverse its AI-cloud slowdown

  • Patel’s Q1 2023 “Amazon’s Cloud Crisis” argued that AWS was optimized for the previous era: elastic scale-out networking, custom CPUs and cost reduction, not tightly coupled AI systems maximizing performance per dollar even when absolute cost rises. Neoclouds would commoditize that advantage, and AWS subsequently became the weakest-performing hyperscaler.

  • His new call is a turn, not a retraction of those structural criticisms. AWS year-over-year growth should trough in the current quarter and reaccelerate above 20% as Anthropic, Trainium and Nvidia capacity starts generating revenue. In today’s shortage, having powered space to fill can matter more than offering the industry’s cleanest software or accelerator.

  • Amazon historically ran unusually dense facilities—about 40 kW racks when peers used 12 kW—and its obsessive efficiency made data halls feel “like a swamp,” hot and humid. Retrofitting networking and cooling is less elegant than purpose-built infrastructure, but inexpensive relative to GPUs; Patel cited 2 GW sites with power, transformers, wet chillers and dry chillers already secured.

  • That retrofit creates component winners. SemiAnalysis highlighted Astera Labs around $90 based partly on Amazon connectivity orders; Patel noted the stock reached about $250 the following month. His point was not that Amazon has superior architecture, but that additional networking and cooling hardware is immaterial if the resulting GPU racks can be rented.

12. Anthropic can tolerate Trainium because it serves a narrow workload

  • Trainium remains “very hard to use” and was described by one host’s portfolio company as nearly impossible in summer 2023. Patel did not claim the developer experience had become good. His narrower claim was that production inference already requires hand-written kernels, low-level optimization and sometimes assembly-like work even on Nvidia.

  • TPUs and Trainium use larger, simpler cores with less general functionality, which some Anthropic engineers reportedly prefer once operating at that low level. A general customer still struggles, but Anthropic can optimize a handful of models rather than support the world’s architectures.

  • Patel’s intentionally rough example was serving mostly Sonnet—“Sonnet 3.5, or sorry, 4.5, whatever it is”—while leaving other models on GPUs or TPUs. If Anthropic reaches tens of billions of ARR and one model carries most traffic, spending heavily to optimize perhaps $15 billion of Trainium capacity becomes economically rational despite poor usability.

13. Oracle won by underwriting demand Microsoft would not

  • Oracle combines a large balance sheet with little hardware dogma. It will deploy Arista Ethernet, white-box networking, Nvidia InfiniBand or Spectrum-X, backed by strong network and software teams. That flexibility makes it a natural counterparty for OpenAI’s extreme demand.

  • Microsoft’s exclusivity became a right of first refusal: OpenAI can present an $80 billion annual or $300 billion multi-year requirement, and Microsoft can decline. Oracle accepted the bet because OpenAI lacks a balance sheet and Oracle has one; Microsoft’s caution is defensible, but it created the opening.

  • SemiAnalysis mapped Oracle’s signed and prospective sites through permits, satellite imagery, power equipment, chillers, transformers and generators. Using a simplified GB200-era assumption—a GPU at roughly 1,000 watts, a whole system at roughly 2,000 watts, and about $50,000 of all-in accelerator capex per GPU—it estimated approximately $12 million of annual rental revenue per megawatt, then rolled each site’s quarterly energization into Oracle forecasts.

  • That method closely matched Oracle’s announced 2025–27 revenue path and most of 2028; unseen 2028–29 capacity created the miss. OpenAI’s ability to pay more than $80 billion annually remains uncertain, but Oracle buys GPUs only one or two quarters before rental. Its earlier commitment is mainly data-center capacity, limiting stranded-asset risk, with debt available for later GPU purchases.

14. xAI’s advantage is treating regulation as another engineering constraint

  • AI infrastructure has moved from percentage growth to order-of-magnitude growth. A 100,000-GPU cluster exceeding 100 MW was once extraordinary; Patel’s team now tracks roughly ten of them and reacts to another 200 MW site with a yawning emoji. “It’s only exciting if you do gigawatt scale.”

  • Elon Musk’s first Memphis build still stands out: xAI bought a factory around February 2024 and trained models within six months on roughly 100,000 GPUs. It deployed large-scale liquid cooling, mobile substations, generators and CAT turbines, and tapped a nearby natural-gas line—galvanizing resources while conventional operators would search for another site.

  • Colossus 2 repeats the feat near gigawatt scale. After political and environmental resistance limited expansion in Memphis, xAI bought another facility roughly ten miles away, placed it near the Mississippi border and acquired a power plant in Mississippi, where regulation differed. Patel’s first-principles summary: others say power cannot be built there; Musk says, “Just go across the border.”

15. GB200 rewards the right workload and punishes weak operations

  • SemiAnalysis estimated GB200 total cost of ownership at roughly 1.6x H100. If a workload sees only 2x performance, the upgrade is worthwhile but marginal; for DeepSeek inference, Patel cited more than 6–7x performance per GPU with continuing optimization, turning a 60% cost premium into roughly 3–4x performance per dollar.

  • B200 is operationally simpler: eight GPUs in a conventional server, with less upside but familiar reliability. GB200 NVL72 creates a coherent 72-GPU domain and much larger inference gains, yet heat, novelty and system complexity make it finicky. A single GPU failure has a far larger “blast radius” than one failure in an eight-GPU box.

  • Sophisticated labs run high-priority work on 64 GPUs and low-priority jobs on eight, borrowing a healthy GPU when another fails and postponing physical service. Clouds cannot casually rent those spares to another customer because the NVLink domain must remain coherent.

  • SLAs now encode the compromise: Patel’s stylized example was roughly 99% uptime for 64 GPUs but 95% for all 72, varying by provider. Large labs can schedule around that; small companies may lose the theoretical gain through idle GPUs, downtime and an inability to mix high- and low-priority workloads.

16. Rubin CPX turns prefill into a distinct silicon market

  • Modern training increasingly consists of inference-like reinforcement-learning generation, while inference itself divides into prefill and decode. Prefill processes the prompt and constructs the KV cache; decode autoregressively emits tokens. They stress hardware differently enough that leading labs already place them on separate GPU pools.

  • Chunked prefill can fill unused batch capacity, but it slows concurrent decode. Disaggregation lets operators autoscale long-input and long-output traffic separately while guaranteeing time to first token—the latency users notice most. “I can’t read that fast anyways,” one host observed, but users still abandon products that hesitate before streaming.

  • Decode repeatedly moves model parameters and user-specific KV caches, making memory bandwidth central. A 64,000-token prefill request instead contains enough computation to occupy an accelerator by itself, making raw FLOPS more valuable than loading parameters rapidly.

  • Nvidia’s Rubin CPX specializes for that compute-heavy prefill phase and strips out expensive HBM, which Patel said represents more than half a GPU’s cost. If Nvidia passes through comparable margins, CPX can make long-context inference materially cheaper without forcing data-center operators to redesign the surrounding facility.

17. Large GPU buyers are back in a tightening spot market

  • Patel’s procurement rule was deliberately memorable: “How you buy GPUs, it’s like buying cocaine.” Buyers call or message several suppliers—“Yo, how much you got? What’s the price?”—rather than conduct a pristine enterprise RFP. His team maintains Slack connections with roughly 30 neoclouds and circulates live customer requirements.

  • Several major neoclouds had sold out of Hopper while their Blackwell capacity was still months away. Reasoning-model inference and revenue accelerated faster than supply, while Blackwell’s reliability and deployment learning curve delayed usable capacity relative to Hopper, which could be installed and producing within a month or two.

  • Hopper pricing bottomed roughly three to six months earlier and had begun creeping upward. Patel stopped short of declaring a return to the 2023–24 shortage: a small number of GPUs is easy to obtain, but a large cluster available immediately is difficult. Capacity, once again, matters before the formal price sheet.