The End of DeepSeek Week: Moneyball for AI, The Future of Compute Demand, Geopolitical Reality Checks, and More
Summary
DeepSeek is better understood as Moneyball for AI than as proof that compute no longer matters. Its $6 million figure covered one final V3 training run—not all the experiments, data collection, or R&D—and the disclosed optimizations are “a tremendous gift to the US.” Thompson’s analogy: the Oakland A’s found the tactics, but American labs can become the Red Sox by pairing them with vastly larger budgets.
DeepSeek achieved genuine efficiency breakthroughs without overtaking the US model frontier. Distillation is nearly impossible to prevent—rate limits merely turn it into a piracy-style cat-and-mouse game—and OpenAI itself used distillation to make ChatGPT economically viable. Producing GPT-4o-, Sonnet-, or o1-level performance months later means “by definition you’re not ahead”: these were “efficiency breakthroughs,” not capability breakthroughs.
Constraints made DeepSeek exercise an optimization muscle that abundant compute had allowed US software culture to neglect. Its mixture-of-experts training, expert-balancing method, and hand-programmed communication work showed that “the nerfing worked.” Thompson conceded his own prior bias against optimization and argued for dual tracks: one team pursuing capability without constraint, another thinking about efficiency exclusively.
The apparent shortage of transformative AI products looks more like a normal adoption lag than proof the technology has failed. ChatGPT launched only in November 2022; the web needed roughly 13 years before Facebook’s 2006 feed found a native commercial format. Chat clients already deliver “massively untapped efficiency gains,” but adoption remains gated by user imagination—and Apple Intelligence, which Sharp expected to make AI “idiot-proof,” has been “an absolute disaster.”
Cheaper models could increase rather than destroy aggregate compute demand because reasoning models can manufacture training data. A reasoning model generates chains of thought, those outputs train a cheaper base LLM, the improved base produces a better reasoning model, and the cycle repeats: “the AI’s making the AI better.” Thompson’s conclusion was categorical about the direction, though not the timing—this era could demand “a basically infinite amount of compute,” including inference performed specifically for training.
Microsoft now holds attractive AI optionality while OpenAI and SoftBank pursue a venture-scale frontier bet. Microsoft retains OpenAI API access and first refusal on compute while stepping back from funding models that may become commodities within months. Thompson put the alternative at perhaps a 5% chance with astronomical upside—reasonable for equity, not debt—and said the evidence currently tilts toward commoditization.
DeepSeek exposes the geopolitical wager embedded in US chip controls: a durable AI lead must compensate for America’s industrial weakness. Thompson remains “in a pile of mud” because EUV is an enforceable chokepoint, yet cutting China off may accelerate domestic substitution, weaken TSMC’s deterrent value, and encourage America to defend past innovation instead of advancing. The policy ultimately shares OpenAI’s takeoff premise: if models diffuse and commoditize, the supposedly compounding technological advantage may not endure.
Deep dive
1. DeepSeek’s Moneyball play is a gift to richer rivals
Thompson endorsed the Oakland A’s analogy: DeepSeek found tactics that richer US labs can adopt with “a huge payroll,” just as the Red Sox absorbed Moneyball methods. More compute remains useful; the claim that it is suddenly worthless is “ridiculous.”
His geopolitical tell was disclosure itself. If DeepSeek’s sole objective were weakening US technology, the “smart thing” would have been to keep the optimizations secret; publishing them buttresses the paper’s broad credibility and lets the entire industry copy the work.
Sharp traced the distorted narrative to a seductive contrast: Sam Altman had announced, a week earlier, that Stargate needed $500 billion, then DeepSeek appeared with a $6 million figure. But that number described one final training run, not all the experiments, data collection, or R&D behind it.
The investor distinction is timing. Compute may ultimately “fill all available space,” yet efficiency can defer absorption, change inference economics, and benefit the companies paying for capacity even while suppliers sell into enormous long-term demand.
2. Distillation compresses costs, not the frontier gap
Asked whether labs could prevent distillation, Thompson’s answer was “not really.” Providers can restrict APIs and impose rate limits, but digital inputs and outputs make enforcement resemble piracy; determined operators can even script ordinary chat clients to harvest answers.
Distillation is also how frontier labs serve users economically. OpenAI initially faced overwhelming ChatGPT demand while serving an expensive full model, so necessity drove breakthroughs that let it support hundreds of millions of users at lower cost: “It’s not a good thing or a bad thing.”
Thompson preserved the capability distinction: matching GPT-4o, Sonnet, or o1 months after release does not put DeepSeek ahead. “The efficiency breakthroughs are real, but they’re efficiency breakthroughs. They’re not capability breakthroughs.”
Sharp rejected both American denial and declinist panic. His middle account: DeepSeek was well-funded, used substantial NVIDIA hardware, likely distilled OpenAI outputs for V3 and R1, and contributed impressive innovations—without turning a $6 million “side quest” into the end of US leadership.
3. Chip constraints forced genuine “coding to the metal”
A listener compared DeepSeek with console developers extracting exceptional graphics from aging Xbox and PlayStation hardware. Thompson accepted the core mechanism—static constraints reward deep optimization—while noting that modern game economics pushed engines toward cross-platform abstraction once asset creation became more expensive than engine work.
The PS3 supplied the counterexample: difficult hardware drove developers to build for Xbox first and port later. Sony responded by buying studios and securing exclusives, illustrating how optimization choices eventually reshape platform strategy rather than remaining a purely technical concern.
DeepSeek’s mixture-of-experts work was substantive. It found a more efficient way to train every expert, moving balancing logic into a bias factor within the experts instead of calculating everything centrally—challenging the assumption that mixture-of-experts gains came mainly during inference.
Engineers also hand-programmed compute units to improve communication under restricted bandwidth. Thompson’s verdict retained both sides of the story: DeepSeek remained months behind the frontier, but “the nerfing worked” and the engineering was “really amazing.”
4. Optimization debt makes Google’s infrastructure matter again
Thompson accepted the criticism that he had previously discounted optimization: “I never said this before this week,” and often argued the opposite. Rationally following comparative advantage can abandon a learning curve whose loss only becomes visible years later—much as he thinks the US eventually lost the ability to build things.
Google’s original infrastructure strategy offered the model: replace bespoke Sun servers with cheap x86 hardware, assume components will fail, and build resilient software around that failure. The result was cheaper, more scalable infrastructure foundational to Google’s success.
Thompson now favors dual tracks—one team unconstrained by optimization, another thinking about nothing else. Gemini 1.5’s million-token context window may reflect Google’s integrated infrastructure, chips, and software; “when it comes to serving the world, Google is still the best.”
That strength has a disruptive catch. If models commoditize, Google can serve them economically and undercut rivals, but the entity most exposed to a successful low-cost AI product may be Google Search itself.
5. AI’s product overhang is only two years old
Thompson’s unsatisfying but historically grounded answer was that foundational technologies “take longer than you think.” Transformers arrived in 2017, GPT-3 sat underused in plain sight, and ChatGPT’s wake-up call came only in November 2022.
Early products usually bolt the new technology onto an old format. Putting newspaper-style display ads beside web text did not unlock internet economics; Facebook’s feed arrived in 2006, roughly 13 years after the early web, and created a genuinely native commercial structure.
Existing chat clients are already powerful, but their usefulness is “gated by your willingness and capability of thinking of things to do.” Thompson’s mundane specimen was asking ChatGPT to convert every H4 Markdown heading in an article to bold, eliminating perhaps 20 minutes of irritation.
Sharp’s pushback was distribution, not capability: he expected Apple Intelligence to make LLMs effortless and universal, but called it “an absolute disaster.” The eventual breakthrough may instead follow the classic founder story—someone solves a personal problem, productizes it, and removes the need for user volition.
6. OpenAI’s safety rhetoric collides with its closed strategy
A listener corrected Thompson’s history: OpenAI was genuinely open relative to Google’s DeepMind, while model cards and CBRN risk categories created useful language for responsible release. Thompson immediately conceded, “Great point,” and granted nearly the entire business argument.
Closed weights are rational for the leader; open weights are a strategy credit for Facebook and DeepSeek. DeepSeek also benefits directly by distributing techniques across a Chinese ecosystem facing chip constraints, while recipients can use the weights instead of distilling outputs.
Thompson’s remaining objection was rhetorical scope. When “safety” expands from catastrophic risk to misinformation or bias, critics of the latter can be accused of ignoring human extinction—a “motte-and-bailey” move that, in his experience, shuts down legitimate disagreement.
Governance sharpened the concern: if one lab believes AI will control the world, its nonprofit board structure implicitly says, “we will control the world.” Thompson prefers “more AI, not less AI”; the Rubicon has been crossed, diffusion is inevitable, and a safety-first nonprofit should not pair closed models with “self-righteous posturing.”
7. Microsoft owns optionality while OpenAI buys the moonshot
Thompson saw Microsoft in an enviable position after the Sam Altman–Satya Nadella reconciliation selfie. It retains full OpenAI API access and first refusal on compute while reducing its obligation to finance frontier models that could be commoditized within months.
Sharp challenged whether OpenAI’s strategy remained reasonable when distillation has no durable defense. Thompson answered probabilistically: perhaps the chance of one dominant model is only 5%, but its upside is astronomical enough to justify a venture-capital-style equity wager.
That payoff shape makes debt financing hard to understand; the equity case is a venture-style bet whose upside can justify its risk. Microsoft can decline the same risk without severing access to OpenAI’s models.
Thompson nevertheless thinks evidence is moving toward commodity models. His character assessment supplied the residual governance risk: Altman has an “Icarus sort of fetish” and may always fly too close to the sun.
8. Synthetic data turns efficiency into more compute demand
DeepSeek’s 2,048-GPU V3 run did not describe its entire operation. Thompson cited a report of 50,000 Hopper GPUs—H800s count as Hopper—and argued restricted memory bandwidth likely capped the synchronized cluster, not the firm’s desire for more hardware.
Nor was the final run the main consumption pool: researchers require many experimental runs, while DeepSeek’s startlingly cheap inference had already “completely destroyed the pricing model” in China. Large data centers increasingly exist to serve model usage, not merely one-off training.
R1’s deeper unlock is a feedback loop: reasoning models generate chains of thought and effectively limitless synthetic data; that data improves cheap vanilla LLMs; stronger base models then enable stronger reasoners. “We have entered this virtuous cycle.”
Thompson called this “the bitter lesson” at work—the brute-force creation and acquisition of possible answers. Efficiency may delay some near-term inference purchases, but training, inference, and inference-for-training collectively create an appetite “larger than ever.”
9. Compute growth pulls power and networking along with it
More capable clusters still require NVIDIA’s systems advantage. Thompson noted that an individual AMD chip may be faster, but NVIDIA remains “a gazillion times better” at programming through CUDA and connecting accelerators without communication overhead collapsing the cluster.
The result preserves demand for networking, power distribution, and liquid cooling even if individual runs become cheaper. Reasoning workloads need continual generation, while synthetic outputs become an input to the next training cycle.
Thompson said the Trump administration was preparing an executive order for dedicated, off-grid generation tied one-to-one with data centers, alongside a dramatic loosening of regulations. Predictable loads could bypass grid bottlenecks and send results elsewhere through fiber.
Nuclear is the natural fit because data centers require continuous power; solar could work in remote locations but would need substantial batteries for nighttime operation. The through-line: DeepSeek did not close the infrastructure opportunity—it may have “actually unlocked” more appetite.
10. Export controls wager America’s industrial future on AI takeoff
Thompson argued that NVIDIA “does not operate like an American company”: it allocated scarce chips to startups such as CoreWeave, invested in some recipients, limited hyperscaler leverage, and sold into China while Amazon or Microsoft might have bought more. “NVIDIA’s hands are not clean.”
Yet his export-control position remained deliberately unresolved. EUV is a clear, enforceable chokepoint whose commercialization took roughly 10–13 years and repeated ASML rescues; broader DUV and chip restrictions may instead handicap US suppliers, accelerate Chinese substitution, and spend “a card you can only play once.”
TSMC dependence also deters conflict: if China relies on Taiwanese fabrication, invasion carries an extra cost; cutting that access changes the calculus. Thompson therefore worried that perfectly effective controls could be more destabilizing than leaky ones—the leakage may function as a “pressure valve.”
His honest conclusion was, “I land in a pile of mud.” The unresolved trade is between slowing Chinese capability today and teaching America to defend yesterday’s lead rather than rebuild its own capacity to innovate and manufacture.
11. DeepSeek punctures America’s geopolitical self-assurance
Sharp’s Mistral counterfactual captured the emotional distortion: had a French lab produced DeepSeek’s work, Western observers might have celebrated. Because it came from China, the same result triggered both dismissal and panic—and perhaps a useful “infusion of humility.”
Thompson’s military reality check was physical: China can build ships and dominates components for drones, robotics, motors, actuators, and batteries, while the US struggled even to expand artillery production. “Fighting a war is not aggregating users and capturing demand.”
Taiwan exposes the contradiction. Its leading-edge fabs support US security but staying on-island is also Taiwan’s leverage; Thompson argued America still needs domestic capacity whether China takes Taiwan or war destroys TSMC. He preferred direct Intel purchasing guarantees over merely subsidizing supply, with tariffs a less efficient demand tool.
The chip ban ultimately shares Altman and Dario Amodei’s premise: AI may take off, improve itself, and preserve a compounding US lead large enough to overcome manufacturing weakness. DeepSeek matters because if models instead diffuse and commoditize, many strategic decisions were built on an assumption “that might not be true.”