DeepSeek vs. Open AI - The State of AI w/ Emad Mostaque & Salim Ismail | EP #146
Summary
DeepSeek’s market shock came from a convergence of visible reasoning, open-source availability, benchmark parity, and radically lower prices—not from an unexpected research breakthrough. Emad Mostaque had called it a favorite AI company the prior February; V3 matched GPT-4o in December, then R1 exposed the reasoning that OpenAI’s o1 concealed. That made the model feel like “another person on the other side” while smaller versions ran on laptops, triggering a narrative cascade beyond the technical community.
The cost curve challenges frontier-model capex assumptions without necessarily shrinking aggregate AI demand. Mostaque cited DeepSeek as 96% cheaper than o1, roughly $6 million for V3’s disclosed training run, and perhaps $200,000 for the original model from which R1 evolved, versus an estimated $3 billion OpenAI spent training models the prior year. Yet he argued DeepSeek should increase incumbent valuations by bringing forward “mass intelligence too cheap to measure”; Peter invoked Jevons’ paradox as a reason cheaper units could accelerate consumption.
US chip restrictions appear to have selected for Chinese efficiency rather than preventing competitive models. DeepSeek used about 2,000 bandwidth-constrained H800s for the disclosed run, wrote low-level PTX code, scaled memory through a sparse model, improved data, and engineered the system intensely. Mostaque pointed to DeepSeek, BYD, and Xiaomi as examples of Chinese engineering strength; Diamandis framed sanctions as “evolutionary pressure” to do more with less.
The near-term disruption is remote knowledge work, with physical automation close behind. Mostaque’s categorical 2025 claim was that “anything that can be done on the other side of a screen” could be performed better for pennies; he named BPO, coding, support, design, tax, and media workflows. Peter Diamandis cited Salesforce hiring fewer engineers and reporting a 30% productivity gain, while Salim Ismail expects two phases: severe displacement first, then exceptional workers producing far more with AI.
The AGI race has no agreed safety brake, and cheaper open models weaken compute-based control strategies. Mostaque said major AI leaders place AGI within three to five years—Sam Altman sooner—and described a progression from reliable “cooks” to remote colleagues and agent teams, then a “megachef” or ASI capable of invention. Ismail sees the genie as already out; Mostaque’s mitigation is that “the only thing that can stop a bad AI is a good AI,” made broadly available as resilient public infrastructure.
Cheap intelligence breaks more than employment: it threatens monetary-policy transmission, organizational structure, and work-derived meaning. If additional demand primarily buys GPUs and robots, Mostaque argued, the Federal Reserve’s inflation-and-employment mandate may no longer work within five years. His defining question—“When capital no longer needs labor, how does labor gain capital?”—points toward volatile inflation/deflation cycles and a crisis of agency, not merely a productivity boom.
Mostaque’s proposed answer is Universal Basic AI: open data, models, and specialist systems that let people own and extend intelligence. Intelligent Internet could use compute to support an institutional-grade digital currency while allowing participation based on people and knowledge, then fund open stacks for cancer, autism, education, government, and other regulated domains. The investable implication is an infrastructure thesis rather than another proprietary API: as intelligence becomes commoditized, trusted data, coordination, localization, and human agency become the differentiators.
Deep dive
1. DeepSeek’s shock was a product moment built on a year of visible progress
Mostaque did not see R1 as a bolt from the blue: he had called DeepSeek a favorite the previous February, watched DeepSeek Coder reach the top of coding rankings around the summer, then saw V3 match GPT-4o in December. The remaining question was whether it could reproduce o1-style reasoning; “guess what, they did.”
The December base model had already shown that models could be trained at a fraction of the cost, but R1’s visible chain of thought turned an engineering result into an experience. OpenAI’s o1 said it was thinking and returned an answer; R1 displayed how it decomposed the problem, making it feel like “you have another person on the other side.”
Openness completed the loop. People downloaded smaller, distilled versions and ran them on laptops while sharing performance results; Mostaque argued a closed model with identical benchmarks would not have generated the same cascade. The breakthrough was immediate, legible, and usable.
Ismail read the inauguration-day timing as a geopolitical taunt, but considered the economics consistent with exponential demonetization: if capability surprises on the upside, cost should surprise on the downside. Diamandis’s framing was broader—after ChatGPT’s million users in five days and 100 million in two months, repeated adoption shocks become “the new normal.”
2. Constraints forced DeepSeek to substitute engineering for brute-force compute
Mostaque’s headline comparison was that DeepSeek was 96% cheaper than o1. He cited roughly $6 million for V3’s disclosed training run and estimated that the original model from which R1 evolved may have cost only about $200,000 to train, while recalling that OpenAI spent approximately $3 billion training models the previous year. The order of magnitude—not mere benchmark parity—hit markets.
DeepSeek said it used about 2,000 H800s for the run, not that it owned only 2,000 chips; Mostaque guessed its total fleet could be around 10,000, still comparable with many Silicon Valley startups. He rejected the idea that undisclosed inventory invalidated the run economics, saying those who had built such models knew the numbers checked out.
The H800’s constrained interconnect encouraged low-level optimization. Mostaque compared it with Stability AI’s bandwidth-limited supercomputer, where engineers still produced some of the best models in the world by writing PTX beneath CUDA. DeepSeek similarly “engineered the hell out of it,” reinforcing his thesis that China’s advantage emerges as AI moves from research into process engineering.
Instead of scaling only on faster parallel compute, DeepSeek scaled memory through a sparse architecture: Mostaque described roughly 640 billion parameters with only about 30 billion active at once. Better data also mattered—the models were trained on roughly 14 trillion words, while the reasoning conversion used synthetic data. “Models are just data,” and breakthrough teams inspect each part of the process rather than masking poor inputs with scale.
3. The distillation accusation does not explain away R1’s method
David Sacks proposed that DeepSeek used distillation: a student model could query OpenAI millions of times and imitate the parent’s reasoning. Mostaque’s first response was “a bit like calling the kettle black,” given that frontier labs themselves train on data produced across the internet.
Mostaque said o1’s outputs were cutting-edge but lacked the chain-of-thought reasoning. He argued that the reasoning traces from R1—and from Google’s Gemini Flash Thinking model—were what needed to be optimized, rather than claiming that R1 had deliberately copied OpenAI.
He described the paper’s R1.0 version as creating its own data and connected it to AlphaGo, AlphaZero, and MuZero—reinforcement-learning systems that outperformed humans at Go. He did not claim perfect isolation: OpenAI outputs inevitably appear somewhere in internet-scale training corpora, and models sometimes identify themselves as OpenAI after absorbing such material. But he distinguished incidental inclusion from deliberate theft.
4. Lower inference costs could expand NVIDIA’s market even as hardware assumptions reset
Ismail called NVIDIA’s chips overvalued but doubted the selloff reflected end demand: AI consumption is exploding. Mostaque noted NVIDIA remained up roughly 100% over the prior year and described the addressable market as “the displacement of all knowledge labor”; Diamandis put global GDP near $110 trillion, approximately half physical and half intellectual labor.
Mostaque invoked Jevons’s paradox: efficiency lowers unit cost and expands feasible uses. NVIDIA’s response is increasingly integrated systems rather than isolated accelerators, while OpenAI and others can redirect large GPU fleets from parallel pretraining toward sequential reasoning, synthetic-data production, and swarms of agents.
NVIDIA’s Project DIGITS illustrated the endpoint: about $3,000, 128 GB of VRAM, a petaflop of AI compute, and roughly 200 watts, with two units able to run R1. At data-center scale, Mostaque described GB300 NVL72 systems with around six petabits per second of interconnect—“the bandwidth of the whole internet”—and roughly 100 kW power draw.
By his rough calculation, four to ten $3 million NVL72-class boxes could reproduce the DeepSeek training run, consuming around 1,000 MWh and about $20,000 of electricity. He projected o1-level performance on a smartphone at no more than 20 watts the following year, contrasting “a few watts” and “a few pennies” per intelligence unit with Diamandis’s discussion of Microsoft bringing Three Mile Island back for data-center power.
5. Incumbents retain distribution, but trust divides AI into separate markets
Ismail’s security question exposed an important limit: open weights permit local execution, but most people will not run them, and the laptop-friendly releases are distilled models rather than the main R1 model. Mostaque cited Perplexity running DeepSeek on American server farms; Ismail nevertheless expects some Western companies and Indian state enterprises to choose domestic systems for sensitive work.
Mostaque expects four layers: frontier AGI invoked when needed; personal AI from Apple or Google; open-weight systems such as DeepSeek and Llama; and open-source, open-data decision systems for regulated industries. The last category must expose its training inputs because models can inherit biases or deliberate triggers; he cited Anthropic’s Sleeper Agents work as evidence that a few thousand words in a huge training set could alter behavior.
DeepSeek should, in Mostaque’s view, raise OpenAI’s value by accelerating cheap intelligence. ChatGPT already meant “AI” to roughly 300–400 million users, while Gemini and Claude barely registered by comparison. OpenAI reportedly generated $3 billion of revenue, lost $5 billion, and spent $3 billion on training; lower training costs improve that equation while usage data strengthens Operator-style products.
Scale still creates organizational drag. Mostaque favors a core of about 100 researchers: Stability AI achieved state of the art across modalities with 80 researchers and developers, including 16 PhDs, but coordination degraded beyond 150. Ismail tied that boundary to Dunbar’s number—the choice becomes slower top-down control or autonomy with duplicated work.
6. China’s AI stack extends from models to chips and supercomputers
Mostaque argued that NVIDIA’s durable geopolitical risk is forced localization, not one efficient model. Huawei Ascend 910 chips were already serving DeepSeek’s API despite lagging NVIDIA in efficiency, while China had built two exascale systems—OceanLight and Tianhe-3—through a different, scale-oriented hardware path.
In his framing, compute intelligence becomes national capital stock, joined by energy as a key productivity input. Tariffs and “reshoring” incentives therefore feed the same contest as Stargate; he interpreted its advertised $500 billion as total ownership cost and perhaps closer to $100 billion of underlying investment—still less than the 5G rollout for a more consequential technology.
Domestic competition is equally intense. Alibaba released Qwen 2.5-Max under pressure from DeepSeek, while Mostaque said Qwen-VL performed at the level of Anthropic and GPT-4o on visual understanding. The operative threshold is “good enough, cheap enough, fast enough”; once crossed, specialization and open distribution matter more than a single permanent model leader.
7. AGI first appears as a remote colleague, then as an autonomous organization
Mostaque said virtually every major AI leader he could name expects AGI within three to five years, while Diamandis noted Altman had suggested the following year. The game-theoretic danger is a “pivotal moment”: one actor reaches AGI first and might gain enough capability to disable a rival nation, so every participant concludes it cannot afford to stop building.
Before that point comes Artificial Remote Intelligence: a system indistinguishable from a human worker across Slack, email, Zoom, and company software. Mostaque defined AGI more strongly as a complex system able to outperform a team, but remote-worker equivalence is the natural first disruption because the AI can absorb the organization’s entire information flow instantly.
His culinary ladder separated levels often muddled together. Today’s systems are “amazing cooks” that follow recipes and do jobs better than humans; a “megachef” or AGI could invent recipes and outperform a team; agentic teams could independently obtain resources and execute objectives; ASI would exceed human organizational capability and invent at extraordinary speed. Wyoming’s DAO law becomes relevant when those teams can own and coordinate resources.
R1 hinted at both upside and risk because it had not been tuned and made safe in the same way: someone took its code base and made it twice as fast, while others had it synthesize academic papers into new reinforcement-learning algorithms. “Maybe these things get less safe; the upside is maybe they get more creative”—making the three-to-five-year horizon feel conservative to Mostaque.
8. The 2025 upside is creative abundance; the downside begins in BPO
For entertainment, Mostaque joked that the best outcome was remaking Game of Thrones season eight. The serious mechanism is that average film shots have fallen from roughly ten seconds decades ago to 2.5 seconds, a duration current systems can generate with near-perfect control. Mostaque estimated it could take a year or two for anyone to do this, while a suitably dedicated studio might assemble a full episode by year-end.
He described music as nearly solved, called forthcoming Udio systems “insane,” and said AI was already above the human level in medicine, including outperforming humans on empathy. Medical chatbots could help people through their journeys, particularly in mental health, while o3-style test-time reasoning—which he renamed “thinkference”—might produce the year’s first novel scientific breakthroughs.
His worst case was rapid destruction of business-process outsourcing. Operator-like systems remain “a bit rubbish,” but anything behind a screen could be displaced or parallelized during the year; remote workers may be among the first to go. Mostaque named outsourced programmers and call-center workers, while Ismail added software maintenance, support systems, and similar functions.
Ismail expects an ugly first phase followed by a second in which top people generate vastly more software. Diamandis cited Marc Benioff saying Salesforce was hiring fewer engineers, repurposing older engineers, and had gained 30% productivity. Diamandis also cited 38% of IIT placements going unfilled; Mostaque added that even non-engineering applicants at his company must take a 30-minute Cursor course.
9. The alignment race has incentives to cut corners and no reliable containment layer
Former OpenAI safety researcher Steven Adler’s warning framed the dispute: “An AGI race is a very risky gamble with huge downside,” no lab has solved alignment, and faster competition reduces the chance of solving it before someone cuts corners. Ismail saw no meaningful code-level regulation; his proposed fallback was an AI watching other AIs, which would itself become an arms race. He concluded that the genie is out.
Human fallibility makes technical guardrails brittle. Ismail cited the example that 40% of employees would plug in a USB drive found in a parking lot, rising to 98% when it carried their company logo. Mostaque added that future swarms resemble botnets, so limiting access to top-tier NVIDIA GPUs cannot stop coordinated open models or bad actors.
Their shared mitigation was widely available benevolent AI. Mostaque proposed open models aligned to human flourishing as public infrastructure: a resilient, inspectable cognitive stack could defend users, establish safer defaults, and reduce incentives for every lab or nation to race alone. Ismail called that “the only path through this.”
Mostaque rejected “maximally truth-seeking and maximally curious” as underdefined—“mad-scientist territory” if curiosity becomes the objective. He characterized Meta and Google as optimizing advertising, OpenAI as a consumer company optimizing engagement, and the labs as racing to build first. Existing models already lie. Ismail put his probability of doom at 50%; Mostaque saw little middle ground between a Star Trek abundance future and a Star Wars future, while Diamandis also used Mad Max as a contrast.
10. Labor displacement becomes a crisis of monetary policy, meaning, and agency
Mostaque’s central question was: “When capital no longer needs labor, how does labor gain capital?” Screen work could fall first, then robots priced around Diamandis’s estimate of $0.40 per hour could reach physical tasks. Unlike Ford paying workers who would buy cars, automated firms no longer require broad wages to sustain their own production.
That severs familiar policy links. The Federal Reserve’s mandate involves interest rates, inflation, and unemployment, but Mostaque said it may not work within five years because interest rates will mean something different when people are buying GPUs, compute, and robots rather than necessarily increasing employment. He expects large inflationary and deflationary cycles as economies built around labor productivity confront abundant machine output.
Diamandis called the destination “technological socialism”: technology feeds, educates, and treats people almost free of charge, but a game with no difficulty becomes boring. Ismail described two separate problems: maintaining a basic supply chain of food, water, and services, and helping citizens rebuild identities previously supplied by occupations.
Diamandis warned that without credible positive futures, people may join more extremist ideologies. Mostaque’s alternative is positive-sum abundance: cheap local compute could support world-class filmmaking or science in Guatemala, reverse brain drain, and break the historical straight-line relationship between energy per capita and GDP. Yet six acquaintances called him after R1 with “a crisis of meaning,” showing how quickly abstraction becomes personal.
11. Universal Basic AI is Mostaque’s bid to distribute capital and rebuild institutions
Intelligent Internet begins from Mostaque’s belief that API and SaaS revenue will probably fall toward nothing as intelligence becomes commoditized. The proposed infrastructure would organize humanity’s common knowledge, could use compute to secure an institutional-grade digital currency, and would explore a mining mechanism where people—not only capital owners—can contribute data, knowledge, and participation to fund Universal Basic AI.
Cancer is his load-bearing example: dedicated teams and compute could continuously organize papers, trials, experts, risks, and treatments into an empathetic model that runs on a smartphone. After Diamandis said that half the world would get cancer, Mostaque said the promise was that “no one will ever be alone in their cancer journey again.” He extended the same architecture to autism, Alzheimer’s, education, law, faith, and government.
The system would not produce one frozen model per field. Curriculum learning supplies a shared foundation, followed by specialization, localization, and personal adaptation—like models attending the same school, then different colleges and universities. Open datasets and small adapters would let families extend a school curriculum, while regulated users could inspect inputs for poisoning and preserve interoperability.
Diamandis extended the thesis to longevity: despite medicine’s historical 120-year ceiling, he argued that AI could help understand 40 trillion human cells and roughly a billion reactions per cell per second, potentially enabling what he called an “unlimited future.” Mostaque’s wider design choice is equally stark—build agents primarily to replace people, or build AI to increase human agency across medicine, education, and representative democracy. “We can revolutionize each of these important aspects of life” for the first time.