Your Guide to the DeepSeek Freakout: an Emergency Pod
Summary
- DeepSeek’s claimed economics challenged the assumption that frontier AI requires the richest labs, best chips, and tens of billions in infrastructure. The year-old Chinese company said it trained V3 on export-limited chips for $5.5 million—roughly 100 times less than a comparable OpenAI model—though Kevin Roose cautions that the figure deserves “a grain of salt.” If it checks out, barriers to entry, model pricing, and margins all fall.
- The technical story became a market event when DeepSeek released its R1 reasoning model and then brought its models to consumers. NVIDIA fell about 18%, erasing hundreds of billions of dollars in market value, while DeepSeek’s free, ad-free app hit No. 1 in the US App Store. R1 was positioned as a counterpart to OpenAI’s o1 and o3, and its visible thought process made prompting feel less like “throwing a penny in a fountain.”
- Casey Newton’s countercase is that rapidly declining AI costs were already the trend, not a sudden invalidation of the industry. An Ethan Mollick chart showed that, in some cases, GPT-4-level inference costs had fallen 1,000 times over a couple of years, and prior American investment supplied techniques DeepSeek could build upon: “One of the reasons that it was cheap for them is because it was expensive for everyone else.”
- Pricing pressure is immediate even if Big Tech’s data centers do not become stranded assets. One person at a large AI company told Kevin that customers were already asking whether switching from OpenAI APIs to DeepSeek could save 80%; meanwhile, the same training servers and chips can be redirected toward inference, where scaled incumbents may benefit as usage grows.
- DeepSeek shows that open models can close proprietary labs’ lead faster, while raising geopolitical and safety risks. R1’s weights can be downloaded and modified, it refuses some politically sensitive queries such as Tiananmen Square, and DeepSeek has disclosed no clear safety program. Casey nevertheless warns against using China reflexively to accelerate an AI race, especially when advocates may profit from conflict.
- Jevons paradox is the incumbents’ bullish escape hatch, but the hosts do not treat it as settled. Satya Nadella argued that greater efficiency will make AI “a commodity we just can’t get enough of,” preserving demand despite falling unit costs. Kevin notes that this is also exactly what Microsoft’s CEO would say during a selloff; Casey’s verdict remains that DeepSeek is significant, but some reactions are “over their skis.”
Deep dive
1. DeepSeek turned an efficiency paper into a consumer shock
Kevin traces DeepSeek to High-Flyer, the hedge fund from which the roughly year-old Chinese AI company emerged. Its first major attention came around Christmas with V3, a model it said could compete with leading American systems.
The startling claim was not merely performance: DeepSeek said V3 used export-limited, “second-rate” chips rather than top H100s, with a raw training-run cost of $5.5 million. Kevin keeps the necessary hedge—take that figure “with a grain of salt”—but says it implies roughly 100-fold cheaper training.
R1 then supplied DeepSeek’s attempt at a reasoning model, in the mold of OpenAI’s o1 and o3, followed days later by an accessible app. Millions downloaded a free, ad-free chatbot that seemed “as good or better than ChatGPT,” propelling a rare Chinese consumer app to No. 1 in the US App Store.
Casey sees a product-design breakthrough in DeepSeek displaying how it understood a query while answering. Instead of prompting being like “throwing a penny in a fountain,” users can understand the model’s approach and formulate better follow-ups.
2. The selloff repriced the AI moat
NVIDIA’s roughly 18% fall—hundreds of billions of dollars in lost market value—made DeepSeek broader than another model launch. Kevin calls the investor fear “declining margins and commoditization”: a little-known entrant may have matched leaders with weaker chips and a fraction of their spending.
The threatened premise was that “bigger was better”—frontier models required billions or hundreds of billions of dollars, vast data centers, and the best GPUs. If DeepSeek’s claims hold, leading models might instead cost single- or double-digit millions, inviting more competitors and limiting what providers can charge.
Casey’s pushback is worth keeping: DeepSeek did not emerge independently of expensive American innovation. “One of the reasons that it was cheap for them is because it was expensive for everyone else”; earlier labs absorbed the cost of figuring out the techniques DeepSeek could build upon.
3. Cheaper models do not necessarily strand compute
Casey is “freaking out a bit less” because cost declines were already extreme and widely expected. Citing a chart shared by Ethan Mollick, he notes that GPT-4-level inference costs had fallen 1,000-fold over a couple of years; DeepSeek underscores a curve already in motion.
Data-center spending is also reusable. The same chips and servers bought for giant training runs can serve inference, and as AI adoption expands, the incumbents with capacity can answer the queries and build businesses around them. DeepSeek’s own demand, Casey argues, means it would still “love to have all of these chips,” so its success does not prove export controls failed.
Kevin preserves the near-term bear case: a person he knows at a large AI company said customers were already asking whether replacing OpenAI APIs with DeepSeek could save 80%. Even if compute demand survives, model suppliers face pressure because buyers want inference “as cheap as possible.”
Satya Nadella’s answer was Jevons paradox: efficiency can increase consumption enough to offset savings. His claim that AI use will “skyrocket” could preserve Microsoft’s economics—but Kevin notes it is precisely the message investors would expect from its CEO during a panic.
4. Open weights are shortening every proprietary lead
Casey describes part of DeepSeek’s achievement as “a fancy ripping-off of techniques” pioneered in the United States. The real change is speed: open models once remained somewhat behind systems from OpenAI or Google, whereas DeepSeek suggests they can now catch up much faster.
Kevin says DeepSeek has shown how models such as Llama 3 or Llama 4 can be distilled—made smaller and cheaper without sacrificing much performance. Once techniques become legible, Casey says, it becomes harder to maintain a defensible moat.
Meta, which has spent billions on AI while freely releasing its models, may feel that pressure most acutely. The Information reportedly described four internal war rooms responding to DeepSeek; Casey says he has no original reporting but trusts the reports. His tart framing is that Meta was “supposed to be the best company at ripping other people off,” and now faces a faster imitator.
5. Open proliferation expands both geopolitical and safety risk
DeepSeek’s models remain recognizably shaped by China: R1 reportedly refuses prompts about Tiananmen Square. Kevin worries that if Chinese firms lead, their censorship laws and values could become embedded in widely used AI and prove difficult to remove.
Casey accepts the concern but distrusts calls to accelerate simply because China exists. His eyebrows “arch up” when people with investments that could benefit from a prolonged conflict invoke geopolitical competition to justify a faster AI race.
AI safety experts worry that growing capabilities could soon reach general intelligence or superintelligence and end badly for humanity. R1’s open weights let anyone download, run, or modify it, sharpening the fear that dangerous future systems will be “very hard to control”—and difficult to contain within one country or company.
Casey sees no disclosed DeepSeek safety ethos or even a clearly identified safety researcher: the apparent approach is “build AGI, give it to as many people as possible” and see what happens.