DeepSeek DeepDive + Hands-On With Operator + Hot Mess Express!
Summary
- DeepSeek narrowed the perceived US-China model gap without proving that China has won the broader AI race. Its app logged 1.9 million recent iOS downloads and roughly 1.2 million on Google Play, while a US Navy ban and an Italian ban after a data-protection inquiry, plus a Microsoft-OpenAI distillation investigation, underscored security and data-use risks. Jordan Schneider’s base case remains that Chinese labs can “fast follow” frontier models.
- DeepSeek’s organizational design helped produce technical breakthroughs. Born from a successful quant hedge fund, it hired young researchers and operated without an immediate direct-profit motive while Alibaba, Tencent, ByteDance, and Huawei struggled to create the institutional structure for breakthrough collaboration. Schneider called its team “dreamers” resembling OpenAI from 2017 to 2022, though success may now push it toward a hyperscaler partnership and potential national-champion duties.
- Export controls retain strategic value because efficient models do not eliminate the need for compute. China can still buy Nvidia’s H20, which Schneider described as basically world-class for deploying AI and letting everyone use it; he said the hardware gap was “actually not that big,” but DeepSeek’s founder has identified controls as a constraint. Kevin Roose’s consumer-hardware scenario became perhaps 10 percentage points likelier; Schneider’s counter-signal was that Nvidia fell 15%, not 95%.
- A six-month model lead may matter less than deployment capacity and institutional adoption. Casey Newton characterized DeepSeek as efficiently reproducing capabilities US companies had trained nine to 12 months earlier; he would reconsider if China first delivered a phenomenal agent or “virtual coworker.” Schneider argued the ultimate contest also runs through data centers, regulation, labor resistance, and society’s ability to absorb Industrial Revolution-scale reorganization.
- Operator makes OpenAI’s $200 monthly price more intelligible as an early labor product than as a premium chatbot. Its hosted “browser within a browser” clicks human websites without requiring APIs, aiming eventually at virtual engineers, consultants, lawyers, doctors, and research assistants. As Roose put it, “$200 a month is a lot to pay for a version of ChatGPT, but it’s not a lot to pay for a remote worker.”
- The hands-on results showed genuine autonomy surrounded by costly unreliability. Operator completed roughly 90% of a domain-purchase, hosting, and DNS project, yet an Instacart attempt searched for milk in Des Moines and took at least 10 times longer than doing it manually; another run earned only $1.20 from 45 minutes of surveys. The experience resembled training “a very new, very insecure intern.”
- Agent capability is improving fast, but the product still sits across a privacy-and-security chasm. Anthropic’s Computer Use scored 14.9% on OSWorld, while OpenAI’s CUA reached 38.1% three months later—a failing grade with a striking slope. Seamless use would require access to the user’s logged-in browser and payments, while autonomous action could eventually enable anything from manipulated purchases to cyberattacks.
- The week’s smaller failures illustrated the liabilities accompanying rushed automation and physical technology. Fable removed every AI feature after racist summaries; Amazon paused drone deliveries after two rainy-weather crashes; and Fitbit paid $12 million after 174 overheating reports and 118 injuries. Waymo vandalism was only a “lukewarm mess” to Roose, but one with escalation potential as autonomous vehicles become a physical target for anti-technology anger.
Deep dive
1. DeepSeek’s breakthrough became a distribution and policy event
Casey’s opening snapshot captured the speed of the shock: 1.9 million recent iOS downloads and about 1.2 million Google Play downloads, followed by a US Navy ban and an Italian ban after a data-protection inquiry.
OpenAI said it had evidence DeepSeek distilled its models; Microsoft and OpenAI were investigating possible API abuse. The hosts noted the irony of OpenAI objecting to data use “without payment or consent” while facing The New York Times’ copyright lawsuit.
Schneider called DeepSeek an “odd duck”: its CEO came from a successful quant hedge fund and, after ChatGPT, committed money, compute, and “fresh young graduates” to building language models without an obvious near-term business model.
2. DeepSeek’s freedom produced breakthroughs that incumbents struggled to organize
Schneider contrasted DeepSeek with Alibaba, Tencent, ByteDance, and Huawei: the large companies had resources, but struggled to create the “organizational and institutional structure” that rewarded the collaboration and research needed for genuine breakthroughs.
Freedom from a direct profit motive helped DeepSeek produce notable innovations beginning in late December and then the R1 chatbot. Schneider’s working model, based on two extended CEO interviews and employee posts, was simple: “They’re dreamers.”
The closest comparison was OpenAI from 2017 to 2022, when Sam Altman could say he had no idea how the company would make money. DeepSeek’s ambition likewise appeared to be: “We’ll build it, and we’ll make it cheaper for everyone. You know, we’ll figure it out later.”
Its trading strategies could fund that mission for now, but Schneider thought the attention might start a new phase requiring DeepSeek to “shack up with a hyperscaler”—perhaps ByteDance, Alibaba, Tencent, or Huawei—as government scrutiny begins.
3. Becoming a national champion could compromise DeepSeek’s original mission
Schneider rejected the idea that Chinese technology CEOs primarily dream of spreading Xi Jinping Thought: they want to compete with Mark Zuckerberg and Sam Altman and prove they are “really awesome and great technologists.”
His warning came through ByteDance’s history: Zhang Yiming’s earlier Weibo posts supported freer expression, but by 2018 he publicly apologized and pledged closer adherence to Chinese socialist values after the political environment closed around the company.
Didi supplied the harder precedent. It listed on a Western exchange after the government told it not to, was taken off app stores, and entered a “rectification process”—an example of an apolitical company getting “zapped” after crossing the state.
DeepSeek had flown under the radar, but Schneider said that period was over. National-champion status could bring government contracts, distraction, and potentially greater state access, even if those responsibilities conflict with its mission to develop and deploy AI broadly.
4. Open-source success creates a censorship dilemma, not a capability ceiling
The hosted DeepSeek product refused questions about Tiananmen Square, Xi Jinping, and the Great Leap Forward, yet Casey noted that the open-source V3 model apparently had not been trained with the same avoidance behavior and could be downloaded, hosted elsewhere, and stripped of some guardrails.
Schneider called the censorship-capability argument “a bit of a red herring.” Claude will not produce racist material on demand and remains capable; similarly, political restrictions do not necessarily make Chinese models structurally less intelligent than Western ones.
The real dilemma is political: an open Chinese model can be induced abroad to produce statements that could bring punishment inside China. Yet DeepSeek has brought Chinese AI “the most positive shine” it has ever received globally, forcing regulators to balance control against prestige.
5. Compute controls still matter, even if efficiency changes the odds
Whether DeepSeek possessed smuggled banned chips was “a question more for the US intelligence community than Jordan Schneider on Twitter.” China’s legal options are not trivial: Nvidia’s H20 remains basically world-class for deploying AI and letting everyone use it.
Schneider nevertheless called compute a core input regardless of future distillation, noting that DeepSeek’s founder had described export controls as the company’s main constraint. If the US holds the line on chips and semiconductor-manufacturing equipment, China should struggle to leap ahead, much less fast-follow beyond developing comparable models.
His nightmare was a trade exchanging “soybeans for ASML EUV machines.” The controls work only conditionally; a political bargain that weakens them could undermine the advantage he believes the US and its allies still possess.
6. Model parity is only one frontier in the AI competition
Kevin’s pushback was that algorithmic efficiency might eventually place the “biggest, baddest, smartest” model on a MacBook, making chip controls and proliferation limits ineffective. Schneider allowed that DeepSeek may have raised this outcome’s probability by 10 percentage points.
His market-based rejoinder: Nvidia fell 15%, not 95%. If consumer hardware had truly made advanced chips irrelevant, he expected a much more dramatic repricing; a democratized local-model future remained possible but “pretty unlikely” to his “history major brain.”
Casey saw V3 and R1 primarily as optimizations of models US firms trained nine to 12 months earlier. His threshold for declaring the race transformed would be China launching the first exceptional virtual coworker or agent, not merely catching up efficiently.
Kevin doubted that a six-month lead meaningfully prevents a bipolar AI world. Schneider’s answer widened the frame: frontier experiments still require giant data centers, while implementation brings teachers’ unions, regulatory resistance, and societal reorganization comparable to the Industrial Revolution.
7. Operator turns the agent thesis into a visible, expensive prototype
OpenAI followed Anthropic’s Computer Use and Google’s Project Mariner with Operator, available through ChatGPT’s $200-a-month Pro tier. It runs on a dedicated site and opens a browser hosted on OpenAI’s servers rather than taking control of the user’s computer.
This “browser within a browser” has no existing bookmarks or sessions, though users can take over when needed. Operator visually navigates ordinary human websites, trading the speed and structure of APIs for generality wherever no API has been built.
The commercial North Star is virtual labor: engineers first, then consultants, lawyers, doctors, and perhaps research assistants. Casey would gladly pay $200 for useful research help; Kevin observed that the same price looks modest beside a customer-service or billing employee.
8. Operator’s best run looked like an intern; its worst needed constant rescue
On TripAdvisor, Operator found London walking tours and summarized them within minutes. Casey could have searched faster himself, but watching the “self-driving browser” navigate, read, and report still felt like a legitimate technical feat.
Instacart went worse: Casey’s best explanation was that Operator inherited the server’s Iowa location, searched for milk in Des Moines, and then treated his preferred San Francisco grocery store as the delivery address. He had to take over, log in, and correct it after spending at least 10 times the normal effort.
Kevin’s harder test was materially better. Asked to find the best available domain under $50, buy it, arrange hosting, and configure DNS, Operator completed about 90% while he got coffee; he intervened for payment, registration details, and a hosting-plan choice.
The winning instruction resembled management of “a very new, very insecure intern”: “Only ask for my intervention if you can’t progress any farther.” Elsewhere it ordered two small tacos as food for two people and needed prompting—“Does that sound like enough food?”—to reconsider.
9. Rapid benchmark gains collide with a hostile internet and real safety risks
Operator earned $1.20 from online surveys in 45 minutes, and Kevin said he was sure it consumed hundreds of dollars’ worth of GPU computing power, while refusing online poker. It was blocked by The New York Times, Reddit, and YouTube, while GoDaddy would not let it buy a domain name, demonstrating how quickly websites can constrain general-purpose agents.
Casey’s optimistic datapoint was OSWorld: Anthropic’s Computer Use scored 14.9%, while OpenAI said CUA reached 38.1% three months later. “38.1% is a failing grade,” but comparable improvement over three to six months could produce a capable computer user.
His product objection was the gap between an agent using a browser and one using your browser. Existing logins and payment details would remove friction, but also create much greater privacy and security exposure; OpenAI says it stops taking screenshots while the human controls Operator.
Kevin warned that “the internet is not going to sit still.” If agents become 10%, 20%, or 30% of traffic, advertising assumptions could break, merchants might optimize messages to influence bots’ purchases, and greater autonomy could enable an agent to launch cyberattacks or steal from crypto wallets.
10. The Hot Messes ranged from offensive software to physical injury
Fable’s AI told one reader to “surface for the occasional white author” and questioned another’s interest in a “straight cis white man’s perspective.” After its mitigations apparently failed, Fable removed every AI feature and submitted a replacement app version—an unambiguous hot mess.
Amazon paused commercial drone deliveries in Texas and Arizona after two latest-model aircraft crashed in rainy December testing. No one was injured, so the hosts rated it moderate or warm, while emphasizing the quality-control stakes of heavy machines flying over people.
Fitbit agreed to pay $12 million after allegedly failing to report burn risk quickly. From 2018 through March 2022, it received at least 174 overheating reports and 118 reported injuries, including two third-degree and four second-degree burns: “If technology physically burns you, it is a hot mess.”
11. Maps complied with politics while Waymo became a physical target
Google said Maps would follow government updates renaming the Gulf of Mexico the Gulf of America and Denali Mount McKinley. Kevin called routine place-name compliance a mild “tempest in a teapot”; Casey rated the confusion a hot mess, preserving their rare split verdict.
A crowd dismantled a Waymo during an illegal Los Angeles street takeover and used pieces to smash its windows. Motive remained unknown. Casey noted that Waymo had only become officially available in Los Angeles in November, so novelty and curiosity might have contributed; Kevin rated the incident a “lukewarm mess” with escalation potential as some people see Waymos as the physical embodiment of technology invading every corner of life.