Jonathan Ross: DeepSeek Special - How Should OpenAI and the US Government Respond | E1253
Jonathan Ross: DeepSeek Special - How Should OpenAI and the US Government Respond | E1253
Summary
- Ross’s headline verdict: DeepSeek is “Sputnik 2.0” — the NASA-space-pen-versus-Russian-pencil story “just happened again” — but the $6M figure is marketing. “It is true that they spent about $6 million or whatever on the training — they spent a lot more distilling or scraping the OpenAI model,” and since OpenAI reportedly loses money on every API token, it was “effectively subsidizing, accidentally, the training of this model.” The genuine innovation was fully automated verifiable-reward RL, no humans in the loop.
- Models are now nakedly commoditized — “if there was any doubt before, that doubt’s over” — and LLMs have “no switching cost whatsoever,” which kills the cloud analogy. Ross’s move if he were Sam Altman: gear up to open-source OpenAI’s models in response — “open always wins, always,” Linux proved it — because “it’s pretty clear you’re going to lose that, so you might as well try and win all the users and the love,” then fall back on OpenAI’s real Seven-Powers moat: brand.
- Stargate’s $500B is “not enough spending,” not too much: Google spent 10–20x more on inference than training in Ross’s TPU days, Harry thinks Jenson said half of Nvidia’s revenue is already inference, and Ross thinks inference could reach 95% — “you don’t train to become a cardiovascular surgeon and then perform for 5% of your life.” Test-time compute compounds it: one DeepSeek answer burned 18,000 intermediate tokens.
- The tradeable call: Harry “just bought a shitload of Nvidia” on the 16% dump — “the most screaming buy of the century” — and Ross agrees on the weighing-machine view: Nvidia is “actually more valuable thanks to DeepSeek, not less.” Jevons paradox: compute cost drops ~1,000x a decade and consumption rises ~100,000x, so spend rises 100x. Training is the high-margin “mainframe” niche; inference is the larger market, and Groq taking low-margin volume is “probably the best thing that’s ever happened for Nvidia stock.”
- The CCP risk is the data, and it’s about to get worse: “right now the CCP is probably going to be taking the safeties off the weapons… now we want the data” — Harry puts it at 100% that Beijing treats DeepSeek as another TikTok. Meanwhile export controls are theater: “you can literally log in, swipe a credit card, and rent GPUs” — “it’s like the Maginot line, you just go around it.”
- What everyone copies next: DeepSeek’s very sparse mixture-of-experts (~671B parameters, ~250 experts of ~2B each, only a handful active) plus synthetic-data retraining — Llama 3.3 70B already beat 3.1 405B via fine-tuning on better data. And DeepSeek restricting signups to Chinese phone numbers means they ran out of inference compute: “training scales with the number of ML researchers you have; inference scales with the number of end users you have.”
- For the foundation-model complex: “pivot, get over it — just pivot.” The model is the engine, not the car; Perplexity is “perfectly positioned” for the moment hallucination rates drop (building on today’s models is like “trying to create Uber before we had smartphones”); and Europe’s prescription is 100 Station Fs by year-end, a thousand by next year.
Deep dive
1. Sputnik 2.0 — but the $6M number is marketing
- Ross doesn’t hedge the significance: “Yes, it is Sputnik 2.0.” But the cost story is spin — “they spent a lot more distilling or scraping the OpenAI model” than the ~$6M of GPU time, which he believes was roughly the same GPU time—4,000 GPUs for 30 days—as the original Llama 70B (Llama’s first model cost ~$5M of GPU time “and it set the world on fire — in a good way”). His summary: “they’re really good at marketing.”
- The mechanism, spelled out: scaling laws assume uniform data quality, but better data beats more tokens. AlphaGo Zero showed the ladder — train, generate better games, retrain, level up. DeepSeek’s shortcut: “if there’s a really good model already right there, just have it generate the data and you go whoop — right up to where it is. And that’s what they did.”
- He refuses the “China just copies” frame — the RL was genuinely innovative in its simplicity: instead of humans grading outputs, “here’s the box, output the answer here, and then check it… no need to involve a human, completely automated.” On DeepSeek’s reward-modeling innovation, an honest non-answer worth keeping: “that area I’m not as familiar with… why don’t you tell me what you saw and I could tell you if it tracks.”
- The irony he flags: OpenAI, likely unprofitable per API token, “was effectively subsidizing, accidentally, the training of this model” — losing a little money on every token while DeepSeek harvested training data. OpenAI probably still has that data and could train on it; distilling DeepSeek back is unnecessary because “they’re actually better still.”
2. Export controls are a Maginot line
- The “biggest gaping hole”: nobody needs to smuggle chips when “you can literally log in, swipe a credit card, and rent GPUs” from any cloud provider. “It’s like the Maginot line — you just go around it. You need to seal it up a little more.”
- Groq blocks Chinese IP addresses — “I believe we might be unique in doing that” — but Ross admits it’s “a little bit fruitless” since anyone can rent a server elsewhere and log in from there. His verdict on the whole regime: “it’s a big Swiss-cheese wall,” and IP blocking probably isn’t the right tool anyway.
3. The CCP will treat DeepSeek as another TikTok — 100%
- The data concern is “probably the most significant”: even well-meaning companies don’t delete — “they write delete right next to your data… it’s still there. Do you really think the CCP doesn’t have all your data?” And the exposure is collective: a neighbor’s complaint or a spouse’s health data can make you vulnerable.
- Groq’s 2016 no-China decision was commercial, not geopolitical, and yielded a formula: “you must send more money to China than you take out” — plus hand over all data and shape answers. Today’s tell: low-temperature DeepSeek won’t discuss Tiananmen. The scarier version Ross sketches: “What about TikTok, should it be banned? Absolutely not, here’s why — and it gives you a cogent reason. That’s kind of scary.”
- His forecast — DeepSeek is a hedge fund acting on its own, but “right now the CCP is probably going to be taking the safeties off the weapons… they’re going to be like, why are you making this model open source? Now we want the data.” Asked directly if Beijing will see it as another TikTok, Harry says: “100%.” Harry’s counter: TikTok you can ban tomorrow; “here it’s open source” — there’s no off switch.
- Why Groq broke its own rule and hosts R1: once DeepSeek hit #1 on the App Store, people were putting data in regardless — so offer an alternative where “we store nothing — we don’t even have hard drives… when the power goes off, everything goes away.”
4. Commoditization is now naked — open-source it, Sam
- “This has just made it absolutely nakedly clear that the models are commoditized… if there was any doubt before, that doubt’s over.” Through Hamilton Helmer’s Seven Powers (which Ross says he fills out for every single investment and requires every person at Groq to fill out): OpenAI’s actual power is brand — “no one else in this space” — and Stargate is Sam trying to bridge from brand to scale economies.
- The prescription: “if I was in that position, I would be gearing up to open-source my models in response — it’s pretty clear you’re going to lose that, so you might as well try and win all the users and the love.” Cannibalization worry? People still pay for Dell over Super Micro because of trust; brand survives the giveaway. The only wrinkle is timing — done now “it looks like a response as opposed to an intentional thing… and it is a response.”
- Why open wins: Linux won when everyone thought open source was less secure and buggier; “now people expect open to be more secure, less buggy, and have more features — how is proprietary ever going to win?” LLMs have no switching cost whatsoever, which is why the cloud analogy “doesn’t hold up at all,” even though Linux has switching costs. Meta, powered by network effects, could give everything away free — “the more it goes open source, the more of an advantage they have. I am completely jealous of that.”
- Inside OpenAI today, he imagines: foot soldiers asking “is my equity going to be worth anything,” seniors managing morale. The line worth keeping: “the number one driver of bad decisions is fear. They have to pick something, commit to it hard, and be brave about it.”
5. $500B Stargate isn’t enough — inference goes to 95%
- Harry: doesn’t DeepSeek ridicule the $500B announcement? Ross: “actually, I don’t think it’s enough spending.” The precedent is Jeff Dean’s two-slide presentation to Google leadership circa 2011–12: “Slide one: good news, machine learning finally works. Slide two: bad news, we can’t afford it” — doubling or tripling Google’s global data-center footprint at $20–40B, for speech recognition alone.
- Google spent 10–20x as much on inference as training in Ross’s era; Harry thinks Jenson said inference is already half of Nvidia’s revenue; Ross thinks the future number could be 95%. “You don’t train to become a cardiovascular surgeon and then perform for 5% of your life — it’s the opposite.” Test-time compute accelerates this: one DeepSeek query took 18,000 intermediate tokens before answering.
- On whether the $500B is real: Gavin Baker tweeted some math and Ross independently “came up with spookily similar math,” though people in the know say they’ve got it — “but then you keep pressing and it’s like, well, maybe there’s some cutesiness to it.” His read: Stargate is an admission that models are commoditized and infrastructure is the moat — but it’s slow capex, and “the real win here is brand. I’d be hiring the best brand firms I could.” OpenAI’s brand in three years: “much stronger.”
6. Jevons paradox — Nvidia is more valuable thanks to DeepSeek
- Harry’s live position: “I just bought a shitload of Nvidia” on the 16% dump — “the most screaming buy of the century.” Ross answers via Buffett/Munger — short term a popularity contest, long term a weighing machine — and on the weighing: “it’s actually more valuable thanks to DeepSeek, not less.” The sell-off assumed compute is mostly training and cheap models mean fewer chips; both premises are wrong.
- Jevons paradox (Ross says he tweeted it before Sacha: “just as Sacha likes to say he made Google dance, I made Sacha dance”): more efficient steam engines meant more coal bought, because “when the opex comes down, more activities come into the money.” For five to six decades, compute cost fell ~1,000x per decade while consumption rose ~100,000x — spend up 100x every decade. Groq sees developer counts “skyrocket” every time token prices drop.
- On “your margin is my opportunity” versus margin-as-defensibility: “training is a niche market with very high margins” — the mainframe business, still worth hundreds of billions a year — while inference is the larger market. Groq taking “low-margin, high-volume inference so Nvidia can keep its margins nice and high” is “probably the best thing that’s ever happened for Nvidia stock.” Even raising money in late 2024, he still had to explain why inference beats training — Groq’s thesis since 2016.
7. What gets copied next: sparse MoE, synthetic data — and DeepSeek is out of compute
- The efficiency trick everyone now adopts: Ross believes R1 is ~671B parameters against Llama’s 70B, structured as roughly 250 experts of roughly 2B parameters each with only a handful active per query — “not every neuron in your brain fires.” More parameters retain more from less data; sparsity skips the compute. Ross recalls GPT-4 was reportedly around 16 experts before being reduced to 8; DeepSeek went the opposite direction, and “part of the cleverness was figuring out how they could have so many experts.”
- The precedent for the coming wave: Meta’s Llama 3.3 70B outperformed its own 3.1 405B — and it wasn’t retrained from scratch, just fine-tuned on a small amount of higher-quality data. Now everyone with hundreds of thousands of GPUs will “create a lot of synthetic data and train” hard against this architecture — bigger models where users are few, cheaper ones where users are many.
- DeepSeek restricting new signups to Chinese phone numbers has one reading: “they ran out of compute” — inference compute. The structural line: “training scales with the number of ML researchers you have; inference scales with the number of end users you have” — which is why chip startups “are going to do just fine.”
8. Steroids, R&D theft, and Europe’s thousand Station Fs
- China practices “RDT — research, development, theft… it’s just part of the culture, and it’s not just against Western companies, it’s against each other too” — the famous Huawei switches booting with Cisco’s logo “and all the bugs.” Should the West steal back? “That’s just viscerally disgusting to me — I’m literally repulsed by the idea.” Harry’s pushback: racing someone on steroids means taking steroids. Ross’s concession: maybe governments have to get involved — he’d love a fair fight with DeepSeek’s “really smart people,” but “the government keeps putting its thumb on the scale.”
- Harry floats the CCP underwriting free access for data capture, as with BYD subsidies “destroying the European car market”; Ross wants automated retaliation modeled on Cold War deterrence — “if you subsidize this industry, we will automatically subsidize the equivalent industry… so don’t do it.” Ross, bluntly: Xi cares only about power retention, so “rational discourse about rules of play is bluntly unrealistic.”
- China’s own anxiety, per Ross: its chief advantage is people — but “what if a GPU becomes the equivalent of a contributor to the workforce… does China’s advantage erode?” Hence his push for Europe’s potential 500 million people to enter the fight, with a concrete prescription for the EU: “by the end of this year you should have 100 Station Fs, and by the end of next year, a thousand” — 3,000 entrepreneurs surrounded by other risk-takers.
- What most worries him: AI-automated cyberwar. Google just announced the first zero-day found by an LLM; nation-states can now automate vulnerability scanning, attacks are deniable (“is it really China? Russia? North Korea? A friendly making it seem like one of them?”) — and unlike nuclear MAD, “I’m just hacking you — and that could spiral out of control.” Even reputation attacks: sullying a public figure “could be worse in some ways” than shooting them, “but you can get away with it.”
9. For foundation models: pivot; for apps: craftsmanship
- To his VC friends mourning “hundreds of millions” in foundation-model losses: “How many companies became incredibly successful without pivoting? Few. Pivot, get over it — just pivot.” The Suno founder (likely) is his exemplar: “he saw it from the beginning — models are going to be commoditized… the model is an engine. What is the car?” On Mistral surviving: “each company has to find their own thing… possible to pivot” — a hedge, not a yes. Who loses: “anyone who just wants to keep going in a straight line.”
- Perplexity is “perfectly positioned for the moment the hallucination — really confabulation — rate comes down”: then medical diagnosis and legal work open up. Until then it’s “like trying to create Uber before we had smartphones” — but people pay for Perplexity anyway, so it rides the wave while waiting for the tsunami. On wrapper apps versus models: “everyone was like, there’s no value in these wrapper apps; everyone was like, there’s no value in these foundation models — where the [__] is the value? That’s part of the exciting part, discovering it.” His answer: craftsmanship — “the details aren’t the details, the details are the thing.”
- On a possible plateau, Ross says self-driving had a much higher threshold because machines get zero tolerance for fatalities; poetry and code are different. And the generative age gets speedrun because “we are the smartphones” — we know where this technology goes, so everyone pre-positions capital. If he were Elon/xAI: better about the hardware bet, worse about building a model — “why is Elon doing that? Just pick one up off the ground.”