Pioneers Insight Method Research Author
DeepSeek Panic, US vs China, OpenAI $40B?, and Doge Delivers with Travis Kalanick and David Sacks
Back to Episodes

DeepSeek Panic, US vs China, OpenAI $40B?, and Doge Delivers with Travis Kalanick and David Sacks

Summary

  • DeepSeek’s R1 release cut the market’s estimate of China’s AI lag from six-to-12 months to roughly three-to-six months, challenging the scarcity premium around frontier models. David Sacks called R1 comparable to OpenAI’s o1 and its API roughly one-twelfth the cost, but rejected the “$6 million versus $1 billion” framing: $6 million reportedly covered only the final training run, while one analyst estimates DeepSeek and its affiliated hedge fund control 50,000 Hopper GPUs worth more than $1 billion.

  • NVIDIA’s 17.7% plunge and roughly $600 billion market-cap loss exposed the core semiconductor debate: does cheaper model training destroy compute demand or multiply it? Chamath Palihapitiya argued DeepSeek’s use of GRPO and bare-metal PTX showed that constrained teams can route around memory-heavy orthodoxy and CUDA lock-in. Travis Kalanick offered the counterweight: “When AI gets cheap…there’s going to be a lot more AI,” while Sacks invoked what he called Jin’s Paradox.

  • The DeepSeek story combines genuine engineering innovation with unresolved evidence of distillation from OpenAI. Sacks said V3 identified itself as ChatGPT-4 in five of eight tests and noted that DeepSeek disclosed roughly 800,000 reasoning samples without clearly explaining their source; he nevertheless called the team technically strong. The innocent explanation is training on publicly posted ChatGPT output, while the disputed one is mass use of OpenAI’s API—an uncertainty the panel refused to smooth over.

  • AI equity value may migrate from interchangeable models toward routing layers, applications, proprietary data and physical execution. Chamath would first build a “shim” that can hot-swap OpenAI, Claude, Llama, DeepSeek or a future R2; Jason Calacanis argued models are becoming infrastructure like storage or GPS; and Kalanick separated the opportunity into “wrapper,” tools and vertical-specialist businesses. Sacks pushed back that OpenAI’s o3 frontier remains ahead of R1 and that declaring closed-model returns dead is “a little premature.”

  • OpenAI’s reported attempt to raise $40 billion at a $340 billion pre-money valuation is the cleanest test yet of whether more capital creates a moat or institutional softness. Kalanick called access to capital a strategic weapon but warned that overcapitalization can make companies “too bureaucratic, too loose, too weak, too soft” while “a thousand flowers” bloom in open source. His experience with Masa: refuse the money and it may subsidize competitors; accept it and assume the investor’s intelligence will inform other bets.

  • Export controls may slow China while simultaneously forcing the engineering workarounds that make it more self-sufficient. Chamath highlighted unverified claims that up to a quarter of NVIDIA revenue flows through Singapore, whose roughly 100 data centers consume about 876 MW, and questioned where the chips ultimately land. The panel’s harder policy problem: restrictions could push China toward domestic fabs, simpler process nodes and AI-designed chips—“constraint as a feature, not a bug.”

  • DOGE’s claimed $1 billion of daily savings matters most through the Treasury market, not the headline arithmetic. The panel framed the fiscal objective as cutting roughly $1 trillion to $1.1 trillion from a $2 trillion annual deficit to get below 3% of GDP; faster cuts could lower inflation and long-term yields, reducing future interest expense. But courts must determine how much congressionally mandated spending the executive can stop, while Social Security, Medicare and Medicaid remain far harder than leases, headcount or discretionary procurement.

  • Cheap AI could accelerate autonomous transport and robotic food production, but electricity and real estate may become the binding assets. Kalanick’s Bowl Builder was rolling out with five customers in April, while his Waymo thesis is that “cheap good AI makes cheap good autonomy”; yet converting all California ride-share miles to EVs could, by his rough calculation, require doubling the state’s energy capacity. If autonomous fleets need roughly one-tenth as many parked cars, 20%-30% of urban land could be repriced—making charging depots, grid access and repurposable parking more consequential than another model benchmark.

Deep dive

1. Food automation needs purpose-built infrastructure, not a robot wedged into an old restaurant

  • Kalanick’s hundred-year endpoint is food that is high-quality, inexpensive, convenient and precisely matched to each person’s dietary preferences, with machines making and delivering it at a cost approaching—or surpassing—the grocery store. People will still cook, but as a soulful hobby: “I love horses, but I don’t ride a horse to work.”

  • CloudKitchens is his real-estate, software and robotics stack for that transition—“either the AWS or the Nvidia…for food.” Restaurants remain customer-facing; CloudKitchens serves “those who serve others,” handling the infrastructure that could eventually connect an order to health preferences, ingredient quantities, provenance and a photograph of the finished bowl.

  • The Bowl Builder lets a restaurant prep ingredients in the morning and leave. DoorDash or Uber Eats orders then trigger hot and cold dispensing, sauces, bagging, utensils, sealing, conveyor movement and an automated locker; a courier waves a phone at a camera and only the correct compartment opens. Five customers were scheduled to begin using the machine in April.

  • Friedberg had pursued a similar canister-based system years earlier, with working demos, but it never reached production. Kalanick’s distinction is deployment: retrofitting a human-optimized QSR can require major capex and two-to-three months closed, while delivery-only kitchens were designed around the machine. One 800-square-foot restaurant did about $3 million annually and reportedly handled roughly 800 orders an hour at lunch.

2. DeepSeek changed the perceived AI race before anyone resolved its cost claims

  • The immediate market verdict was brutal: NVIDIA fell 17.7%, erasing roughly $600 billion, while TSMC, Arm and Broadcom also sold off. The catalyst was DeepSeek’s claim that R1 reached performance comparable to leading Western reasoning models after a final training run costing about $6 million on only 2,000 GPUs.

  • Sacks argued the story became unusually explosive because two conflicts landed together: United States versus China and closed models versus open source. Remove either dimension and it probably does not become a global story capable of erasing close to a trillion dollars in market value.

  • His substantive update was narrower but important. Before R1, industry estimates placed China six-to-12 months behind; because R1 followed o1 by roughly four months, he would now call the gap three-to-six months. A Chinese company becoming the second major public release of a reasoning model was “legitimately surprising.”

  • The $6 million comparison, however, mixed a final run with fully loaded Western R&D, hardware and operating costs. Sacks cited an estimate of 10,000 H100s, 10,000 H800s and 30,000 H20s across DeepSeek and its founder’s hedge fund—a 50,000-plus-Hopper cluster costing more than $1 billion—while stressing that outsiders cannot empirically validate either total spending or undisclosed hardware.

3. Constraint produced engineering paths the compute-rich West had little reason to explore

  • Chamath’s first specimen was reinforcement learning. Rather than follow the prevailing PPO approach, DeepSeek used GRPO, which he said requires less memory while remaining highly performant. His inference was not that Western engineers lacked the ability, but that surplus compute gave them little reason to abandon orthodoxy.

  • His second specimen was DeepSeek’s route around CUDA. The team worked through PTX, which Chamath compared to writing assembly and controlling the NVIDIA hardware closer to bare metal. That weakens the assumption that CUDA’s high-level ecosystem is an unassailable moat and points toward a more heterogeneous hardware and software environment.

  • The disagreement remained useful: semiconductor bulls have incentives to discredit the $6 million claim, challengers have incentives to celebrate it, and neither side can fully audit DeepSeek. Jason called the team “badass,” while Sacks agreed that some of the work was technically brilliant and made constraint look like “a feature, not a bug.”

4. The distillation evidence is substantial, but it does not erase DeepSeek’s own work

  • Friedberg described distillation as using a large model’s answers to train a smaller, cheaper one. A demonstration in which DeepSeek began discussing China before deleting the answer illustrated post-generation censorship, but was not by itself proof of which model supplied the underlying reasoning.

  • Sacks’s stronger evidence came from V3: when asked its identity, it reportedly answered ChatGPT-4 in five of eight cases. That implies substantial exposure to ChatGPT output, but leaves two explanations—DeepSeek could have crawled publicly posted conversations, or it could have made large-scale API calls that violated OpenAI’s terms.

  • DeepSeek disclosed roughly 800,000 reasoning samples used in moving from V3 to R1 but remained hazy about where they originated. Sacks called the small sample count itself remarkable; Jason emphasized the white paper’s science and thoroughness, while Sacks agreed that some of the work was technically brilliant.

  • Jason’s pushback—worth keeping—was that OpenAI itself abandoned its original open-source posture while training on publishers, artists and other internet content, then objected when a Chinese rival allegedly reused its outputs. Sacks was sympathetic to OpenAI’s operational shock, while Chamath warned against answering it with universal KYC or cloud restrictions that would slow legitimate innovation.

5. Open source is simultaneously a community benefit and a Chinese catch-up strategy

  • Chamath put the immediate burden on Meta: the next Llama must “embrace and extend” DeepSeek’s improvements, attract developers and exceed Gemini and R1. In his framing, dependence on one chip, one high-level framework and one model family has become an unacceptable single point of failure.

  • Sacks resisted the leap from R1 to inevitable model commoditization. R1 was comparable to o1, released four months earlier and trained internally perhaps nine or ten months earlier; OpenAI had already moved to o3, while Google, Anthropic and Meta had their own reasoning work. “It’s a little premature to conclude that there’s no reward for being at the frontier.”

  • Nor did Sacks accept the image of a purely altruistic “plucky upstart.” A Chinese company that is behind has a strategic reason to open-source capable models and undercut leading American firms. Jason’s reconciliation was cleaner: DeepSeek can be advancing open access and pursuing geopolitical self-interest at the same time.

6. Model churn moves the investable moat upward, downward and sideways

  • Friedberg asked where equity value belongs if capable models become fast, cheap and broadly available. His electricity analogy: the utilities did not capture all the gains from electrification; much of the value accrued to the factories, products and wider economy that cheap power made possible.

  • Chamath’s first startup requirement would be a model-neutral “shim.” Engineers should not become locked to Claude, OpenAI or Llama; a company must be able to substitute R1, a future R2, an Alibaba model or another leader without rebuilding its application whenever benchmark leadership changes.

  • Jason pushed hardest toward applications, comparing models with storage beneath YouTube or GPS beneath Uber. He revived Gavin Baker’s phrase that “the fastest depreciating asset in the world was a large language model,” arguing that application distribution and real-world hardware will retain value longer.

  • Kalanick divided the opportunity into wrapper, tools and specialized vertical AI, then challenged the simple NVIDIA bear case: “When AI gets cheap…there’s going to be a lot more AI.” Sacks connected that to what he called Jin’s Paradox—lower unit cost can unlock so many economically feasible uses that aggregate spending rises rather than falls.

7. China’s copy cycle has already flipped into operating-model innovation

  • Kalanick remembered Uber China as an “all-out war” in which a painstaking product launch could be copied within two weeks, then one. Uber staffed an entire San Francisco floor with roughly 400 Chinese nationals and advertised in Chinese along Highway 101 to recruit people “to serve the homeland.”

  • His conclusion changed over time: extreme copying compresses the gap until there is nothing left to copy, after which the practiced organization turns toward creativity. For the future of online food delivery, he would now study Shanghai rather than New York; features reaching Uber Eats or DoorDash may have existed in China three or four years earlier.

  • The best specimen was the office-building handoff. Hundreds of perimeter lockers receive food and packages from couriers, while a second class of runners completes delivery inside the building—an “epically efficient” system adapted to local labor economics.

8. Export controls may be creating the very workarounds they were meant to delay

  • Chamath highlighted claims—explicitly still speculation—that up to a quarter of NVIDIA revenue was associated with Singapore and that chips could be moving onward to China. Singapore is only about 250-260 square miles; its roughly 100 data centers consume approximately 876 MW, and the entire local data-center industry was described as only a $1.5 billion-to-$2 billion revenue business.

  • The inference was not proof that every chip leaves Singapore. It was a policy question: if export controls have an easily used shell-company route, the administration must determine where the hardware actually ends up and whether closing one channel changes outcomes or merely redirects trade.

  • Friedberg’s second-order objection was that sanctions encourage China to reproduce fabs, supply chains and eventually ASML-dependent capabilities. If any modern industrial system can coordinate the required effort, he argued, China is a plausible candidate.

  • Chamath made the workaround more concrete: AI can help design chips for older, simpler manufacturing processes. Groq’s choice of 14 nanometers—“VHS and Beta” technology in his analogy—showed that useful architectures need not depend on cutting-edge two-nanometer yields. The alternative Chinese strategy is centralized frontier-model capex, followed by subsidized distillation for domestic firms.

9. OpenAI’s proposed mega-round tests capital as both weapon and anesthetic

  • The reported transaction was $40 billion at a $340 billion pre-money valuation, potentially led by Masa, alongside the Stargate narrative of Sam Altman, Larry Ellison and large-scale infrastructure. Jason framed the consumer contest as OpenAI versus Meta: ChatGPT has more than 1 billion monthly active users and hundreds of millions of daily active users, while Meta has billions of eyeballs and daily active users.

  • Kalanick corrected the record that Uber did not take Masa’s money during his tenure, despite years of pursuit. He regarded Masa as a “promiscuous investor”: information gained from backing one company could inform investments in every competitor. Uber instead secured a $3.5 billion Saudi investment before the Vision Fund existed.

  • The catch was that refusing the capital did not neutralize it; Masa’s money flowed into competitors including DoorDash and subsidized their markets. Kalanick’s rule: when access to capital is a competitive weapon, “you must play ball,” while understanding that the resulting intelligence and influence may be deployed elsewhere.

10. More data centers do not guarantee the strongest long-run moat

  • Kalanick’s warning for OpenAI was that excess capital can produce an organization that is “too big, too bureaucratic, too loose, too weak, too soft.” Against a decentralized open-source ecosystem where “a thousand flowers” bloom, sheer infrastructure may be an overwhelming advantage in some full-stack sectors and irrelevant in others.

  • Friedberg questioned the premise of one giant do-everything model. His alternative resembles a mixture of experts: repeatedly shrink copies into many smaller models, some specializing in mathematics, reading or writing, until a coordinated network uses less energy and time than a monolith.

  • Chamath nominated proprietary content as the durable leverage—Reddit, Quora, newspapers or Disney—while Friedberg argued text is only a fraction of the prize. He claimed YouTube’s video library could be 100-200 times larger than the rest of the internet’s and cited Tesla’s years of camera data as the kind of domain-specific moat that improves a product and generates still more proprietary data.

  • Friedberg’s caution was conditional: eventually data or algorithm quality becomes “the long pole in the tent,” and additional compute no longer repairs the constraint. Jason’s own licensing example was a blanket $2,500 offer to put his book into Microsoft’s training pool, which he favored to help establish a properly licensed market.

11. DOGE begins with people and buildings, but the large savings sit inside systems

  • Jason said DOGE was claiming approximately $1 billion of savings per day—about $3 per American daily or $1,000 annually—with an ambition to triple the pace. The early program combined return-to-office mandates, an eight-month resignation offer expected to attract perhaps 5%-10% of federal workers, hiring restraint and lease cancellations.

  • Friedberg saw execution of a previously advertised playbook, not a surprise; Bill Clinton had also used buyouts during his deficit-reduction program. The unsettled issue is authority: courts will decide how far the executive can suspend spending that Congress required by statute and what must return to the legislature.

  • Jason focused on controllable friction. Even where spending is authorized, the executive can make hiring, procurement and replacement harder, impose competency reviews and slow administrative processes. Return-to-office itself may produce substantial attrition if buildings are as empty as the panel had heard.

  • Chamath’s “three-layer onion” was people, physical infrastructure, then IT and services. The first two save money, but engineers embedded at the nucleus of DOGE teams can enter systems of record and forensically trace payments; he suspected that identified waste could ultimately exceed $2 trillion.

12. The Treasury market makes speed more valuable than the nominal size of each cut

  • The panel’s benchmark was a roughly $2 trillion annual deficit and Ray Dalio’s recommendation to bring it below 3% of GDP, implying approximately $1 trillion to $1.1 trillion of cuts. One comparison suggested that applying 2019 spending to 2024 revenue would yield a $500 billion surplus instead of a $1.5 trillion deficit—a $2 trillion swing.

  • Friedberg’s feedback loop was central: credible spending cuts reduce inflation pressure and increase confidence in 30-year repayment, lowering Treasury yields and future interest expense. “The faster you make the cuts, the less you have to cut.”

  • The market was not yet convinced. The 30-year yield had touched 5% on January 13 and eased to about 4.77%, while nearly 30% of federal debt was expected to refinance that year around current rates; Chamath noted one auction had barely two-times coverage, and a senior capital-markets contact expected 5.5% before relief.

  • Political constraints remained the bearish case. Congress members still wanted benefits for their own districts, while Social Security, Medicare and Medicaid cannot be slow-rolled like discretionary procurement. Chamath’s hope was that Elon Musk could popularize “use a pencil” solutions and name-and-shame waste; Jason argued that dividing each saving by 330 million Americans makes the benefit legible.

13. Waymo has crossed the psychological threshold from experiment to transportation

  • Kalanick contrasted his first Uber autonomous ride in Pittsburgh—leaving him unable to stand straight from adrenaline, with only a large red stop button—with Waymo now: “You’re not even thinking twice.” Normalization is itself part of deployment because riders increasingly experience the system as ordinary rather than theatrical.

  • He still viewed today’s retrofitted vehicles as transitional. Tesla’s Cybercab represents the destination: no steering wheel and an interior designed around passengers rather than a backup human driver. He said Tesla’s newer FSD models had improved perhaps tenfold in miles per intervention over one three-month period.

  • His causal chain was concise: cheap AI makes cheap autonomy, which could in turn commoditize driving intelligence. Manufacturing then becomes a harder bottleneck and a Tesla advantage, while Waymo’s manufacturing partners and Uber’s collection of autonomous-vehicle partners determine who can deploy physical fleets at scale.

  • On safety, Kalanick argued autonomous vehicles are becoming “provably safer” despite individual failures and also remove interpersonal risk between riders and drivers. Regulation should ease as cities accumulate ordinary experience, though Chinese systems such as BYD may face a separate barrier: whether the United States permits the technology to enter at all.

14. Electricity, depots and obsolete parking may matter more than the autonomous-driving stack

  • Kalanick’s back-of-the-envelope calculation was that electrifying all California ride-share miles could require doubling the state’s energy capacity. Even 10%-20% additional capacity would be a five-to-ten-year challenge in many places, leading to his contrarian possibility: combustion-engine autonomous vehicles may scale faster than EV autonomy.

  • Electric fleets therefore need urban land for charging, cleaning, maintenance and robotized servicing; the panel compared these depots with data centers that require their own substations. Permanent magnets and China’s rare-earth position add another physical dependency.

  • Autonomy also attacks parking economics. Kalanick estimated cars on the road could be utilized about 15 times more than before, allowing perhaps ten times fewer vehicles after preserving rush-hour capacity. Because parking occupies roughly 20%-30% of urban land, even a partial transition could release an extraordinary amount of property.

  • Jason flagged the downside before the redevelopment opportunity: falling land values could affect household retirement accounts and pension-fund balance sheets. Kalanick imagined housing, charging infrastructure or even hydroponic “farm-to-table” production, but repeated that grid upgrades are already the long pole in many CloudKitchens developments.

15. The DCA tragedy exposed a safety system still dependent on voice radio and human reaction

  • Jason relayed a commercial pilot’s warning that DCA was “the sketchiest airport we fly into,” with controllers playing “fast and loose” and helicopter traffic nearly impossible to see while moving at 150 mph. The pilot said both crew members remain on “red alert” during every approach.

  • Wisk CEO Brian Yutko’s proposed inversion was that collision-avoidance software should be permitted to take control even in piloted commercial aircraft. Some fighters already use automatic ground-collision avoidance when pilots lose consciousness; comparable authority could prevent a commercial crew from continuing into a collision the aircraft already predicts.

  • The second failure was infrastructure: critical instructions still travel through VHF voice communications, while air-traffic control runs on technology the panel described as dating to the 1960s. Better data links, tower automation and advanced pilot training were presented as investable upgrades after a tragedy that “should never have happened.”