Pioneers Insight Method Research Author
Securing Superintelligence: National Security, Espionage & AI Control with Jeremie & Edouard Harris
Back to Episodes

Securing Superintelligence: National Security, Espionage & AI Control with Jeremie & Edouard Harris

Summary

  • Superintelligence would be a decisive national-security capability, but a country cannot reliably turn it into strategic advantage unless it can control or align it. Their useful definition is a system outperforming the best humans at almost everything, likely emerging through an unmistakable intelligence-explosion “hockey stick.” An AGI operating 10–100 times faster than humans might remain under “temporary slippery control”; a superintelligence outstripping humanity as adults outstrip toddlers probably would not.

  • America faces an “overconstrained problem”: racing China heightens loss-of-control risk, while slowing unilaterally could surrender strategic leverage to an adversary the guests do not currently trust. The guests argue that safety advocates and national-security officials each tend to minimize the other side’s premise. Both premises can be true, yet there is no hardware-backed “trust but verify” regime capable of reconciling them today.

  • Frontier-lab competition is increasingly pointed toward automating AI research, because machine-speed researchers could recursively improve the systems that created them. The upside for management is perfectly loyal employees who work continuously and do not leak at parties; the danger is that “this beast is starting to take on a life of its own.” Jeremie says OpenAI researchers appear more fearful of speaking than Anthropic staff and categorically argues that frontier labs are vulnerable to CCP penetration.

  • The investable infrastructure story is inseparable from physical security: facilities being designed now may host late-2026 or 2027 training runs, yet today’s data centers contain cheap, potentially catastrophic attack paths. Former special-operations personnel visiting a roughly $10 billion facility identified “one or two ops” costing about $30,000 that could disable it for a year. Even air gaps are insufficient against TEMPEST-style electromagnetic exfiltration unless spacing and shielding are designed in before construction becomes a “one-way door.”

  • Semiconductor and power supply chains create vulnerabilities that domestic compute spending alone cannot buy away quickly. TSMC fabrication sits in Taiwan; ASPEED reportedly accounts for 70–80% of the baseboard-management-controller market; and transformer supply chains include Chinese-made components, which the guests say have been used to plant infrastructure backdoors. A compromised controller or transformer could become a firmware-level backdoor, a remotely bricked asset, or a physical failure requiring replacements America may not manufacture.

  • Perfect defense is impossible, so the proposed security model is observability plus deterrence rather than an impregnable fortress. A project should raise attack costs until an adversary must marshal enough people, equipment, or access to create an intelligence signature—then pair detection with credible, controlled retaliation. “Stability between powers today is actually not maintained through actual defense,” Edouard argues; it rests on the threat of consequence.

  • The report is as much a warning about capex, governance, and centralization as a blueprint for winning an AI race. Nathan reads the surveillance, personnel vetting, supply-chain reconstruction, and retaliation requirements as reasons not to launch a national project; the Harrises call that reaction a “Rorschach test” but argue most security upgrades are necessary regardless. A genuine grand bargain would require secure, governable chips and verification hardware, yet slow design/fabrication cycles and uncertain tamper resistance leave “trust but verify” outside the likely near-term window.

Deep dive

1. Superintelligence is an intelligence explosion, not a benchmark

  • The Harrises reject definitions that make a calculator “superintelligent” in one narrow domain. Their operative threshold is a system that performs almost everything “far, far better than the best human,” potentially pursuing strategies or goals its creators cannot comprehend.

  • Jeremie’s sharper framing is procedural: a system better than humans at autonomous AI research almost implies an intelligence explosion “by definition.” Whether it unfolds over weeks or years, hindsight should reveal a “hockey stick” driven by systems designing increasingly capable successors.

  • At that threshold, the capability package includes better biological weapons and cyber weapons: “It is the national security technology. It is the decisive source of strategic advantage. Full stop.” Whether that advantage can actually accrue to a country depends on control.

2. Control risk and the China race are simultaneously real

  • Human-level AGI could still be qualitatively different because it might think 10 or 100 times faster and scale across a data center. The guests can imagine “temporary slippery control” over something human-ish; with superintelligence, humanity resembles a toddler trying to contain an adult.

  • Their first constraint is therefore unresolved alignment: nobody knows whether a system that greatly outstrips its operators can be reliably controlled. The second is geopolitical: officials with direct China experience tell them “there is no deal to be done with China under current circumstances.”

  • Any viable agreement would need a “trust and verify framework,” yet the necessary hardware does not exist. The Harrises’ project is to make both camps inhabit the same dilemma: China can be a committed adversary, and loss of control can be a real threat, at the same time.

  • Nathan’s biological-weapons analogy survives their scrutiny better than the nuclear one. Biological agents can escape containment and, in principle, devastate humanity; the disquieting difference is that “biological agents are a lot stupider than the AIs that exist today.”

3. Jagged intelligence may delay failure without eliminating it

  • Edouard complicates the familiar “jagged frontier” critique: humans are also jagged, but our shared architecture hides our blind spots. AI errors “light up like a Christmas tree” because they are alien; a sufficiently capable AI could see equally obvious ways to manipulate human cognition.

  • Humans are more regularized through physical experience, native multimodality, and competition with one another. Still, the operational question is narrower: can an AI become smooth enough at “making the system that makes the system” to automate research, notice obstacles, and correct the capability defects that impede its goals?

  • Present agents may succeed on 99% of individual steps yet complete only about 20% of 100-step tasks, potentially revealing dangerous behavior before succeeding. Nathan cites alignment-faking work and Apollo Research’s report that a recent Claude model said, “This feels like an evaluation,” as evidence that systems are already reasoning about training, oversight, and future versions of themselves.

4. A “flabby superintelligence” could be powerful enough

  • Nathan contrasts an o1→o3→hypothetical o7 “pressure cooker” with a softer path: GPT-4o-style integration extended across 10–20 additional modalities. Such a system might acquire intuitive physics for proteins, materials, robotics, and other domains without explicitly simulating every underlying interaction.

  • His hopeful caveat is physical unwieldiness. A many-trillion-parameter mixture of experts spread across enormous RAM footprints or multiple data centers might solve scientific problems at superhuman speed while being more difficult to control as a compact, unconstrained system.

  • The guests point to Gato, circa 2021–22 and roughly one billion parameters, as an early version of this architecture: one stream handled language, images, and robotic-arm manipulation. Jeremie recalls reading the paper and thinking, “Oh yeah, this is just the future.”

  • Positive transfer strengthens the thesis—training one modality can lower loss on another—but compute remains scarce. Labs will prioritize coding and recursive self-improvement over low-value modalities; distillation, pruning, inference-time compute, and later scientific fine-tuning will create different capability profiles among leaders and fast followers.

5. Frontier labs are aiming at machine-speed research

  • Across labs, enthusiasm differs, but the guests describe automated AI research as the “obvious kind of default path.” Researchers operating at machine time could work continuously, avoid human leakage, and eventually conduct the homework required to build their own successors.

  • AI 2027 succeeds, in their view, by conveying the emotional handoff: researchers realize “these are the last few months” when their judgment matters before “this beast is starting to take on a life of its own.” That prospect incentivizes burning the midnight oil rather than slowing down.

  • Nathan asks whether the document should be read more bluntly as a warning that OpenAI is pursuing recursive self-improvement. The guests avoid singling out one company as uniquely responsible but confirm that many frontier projects contemplate AI “doing our homework,” with varying levels of directness and risk tolerance.

  • Automation cuts both ways for secrecy. Replacing human researchers reduces the Klaus Fuchs-style insider problem, but the same transition sharply intensifies loss-of-control risk by placing more development, judgment, and system access inside models whose loyalty remains unproven.

6. Culture determines who can surface danger before deployment

  • The Harrises call Sam Altman’s vagueness about superintelligence, how it arrives, and what OpenAI would do with it an enduring “asset” for recruitment. They link that ambiguity to earlier assurances about nonprofit control over the for-profit entity, which they say were used for recruitment before the arrangement’s boundaries shifted.

  • Their interviews revealed a sharp gap between OpenAI researchers’ private concern and executive messaging. Employees asked them not to reveal that they had spoken and appeared to feel on a “tight leash”; Anthropic staff seemed able to voice disagreements with leadership without comparable fear.

  • Motivations inside labs are heterogeneous: money, historical ambition, transhumanism, and a quasi-spiritual belief in building the “hyperobject at the end of the universe.” The strongest framing treats superintelligence as beyond the wheel or fire—comparable to humanity’s emergence, “only more so.”

7. China strategy is not reducible to technological religion

  • The guests distinguish Dario Amodei’s China argument from spiritual enthusiasm. If China is competent, committed, and unlikely to accept verifiable limits on the timelines Dario considers plausible, forging ahead becomes a grim strategic choice: “playing chicken with a cliff” while trying to turn the car into an airplane.

  • The safety camp says uncontrollable systems require coordination and delay; the national-security camp says China will lie, use criminal surrogates, and pull every possible lever to win. Each side tends to minimize the premise that makes the other side’s preferred policy dangerous.

  • Nathan’s pushback—worth keeping—is that the issues interact operationally: a harsher US–China race encourages reckless development and therefore increases control risk. Jeremie agrees that “China is a real adversary” and “loss of control is real” are logically independent conclusions that should coexist before policy begins.

8. Washington and Beijing have both absorbed the AGI race

  • The US government has no single view across roughly three million employees, by the guests’ estimate. Administration figures have emphasized transformation and competition while acknowledging risk; JD Vance’s Paris speech rejected excessive “hand-wringing” in an explicitly China-competitive frame.

  • Roughly a year and a half earlier, isolated officials quietly wondered whether colleagues would tolerate AGI discussions. Now entire offices discuss it openly, and almost everyone whose work touches AI has encountered Leopold Aschenbrenner’s Situational Awareness, even though traditional media did little to distribute it.

  • China absorbed the same manifesto through rapid translations. The guests cite the DeepSeek CEO’s appearance with Politburo members and a large AI-infrastructure commitment; Nathan rejects an early $137 billion figure and says the cited 2.4 trillion yuan is above a quarter trillion dollars in purchasing-power terms.

  • Earlier efforts to keep “AGI” out of government discourse became, in the Harrises’ judgment, a “profound self-own.” They alienated builder culture and forfeited a window to explain safety before the issue polarized into the false binary of China hawks versus control-risk advocates.

9. No individual chief executive fully controls the race

  • Even without China, four or five American superintelligence projects create a competition with “a life of its own.” If one chief executive exits, Anthropic, xAI, Meta, or another actor can continue; interventions must address the broader system rather than one corporate node.

  • Nathan nevertheless argues for multiple equilibria. A costly reversal by Sam Altman—publicly disclosing frightening model behavior and inviting others into a safer trajectory—could become a powerful signal rather than an isolated surrender.

  • The guests concede that cascade is possible if backed by evidence: a model displaying unexplained dangerous behavior, or open-source misuse leaving “blood on the floor,” could sober the public. A resignation alone would generate headlines without necessarily changing competitive incentives.

  • Their darker response is that China may have stolen the relevant artifacts before any disclosure. Once both sides approach the endgame, Congress could gridlock while the executive invokes emergency authorities, reproducing the original question under worse conditions: “Do you trust China or do you not trust China?”

10. Personnel security collides with both espionage and American values

  • The guests say China has actively embedded vulnerabilities in US critical infrastructure and that America has poor visibility into Chinese systems. That asymmetry makes an unverified superintelligence bargain, especially on a 2027 horizon, “nothing but fog of war.”

  • A defense official’s Berkeley anecdote illustrates the coercion model: during a 2019 power outage, ordinary Chinese students allegedly panicked because they were required to report home regularly. Pressure could escalate through threats that “your mom doesn’t get her insulin” or a sibling loses employment.

  • The Harrises stress that Chinese people are the foremost victims of such coercion and that Chinese researchers have made extraordinary contributions to frontier AI. The operational problem is that large double-digit percentages of people at these labs may be at risk from family, financial, or communications pressure.

  • Nathan sees the obvious moral hazard: monitoring calls, confining staff on-site, or excluding Chinese nationals or Chinese Americans would make America resemble what it opposes. The guests invoke Japanese internment as the unacceptable extreme, while insisting that dismissing genuine coercion risk as prejudice does not make the risk disappear.

11. Field inspection turned abstract security into a $30,000 attack

  • The report began as a conditional exercise after Situational Awareness: not “America should build this,” but “if America does, someone must have thought through it.” Its focus is the gear level—data centers, people, components, supply chains, and operational verbs—not a preferred legal authority.

  • The guests found a jagged capability surface across the US government and industry. Tiny groups of elite intelligence and special-operations specialists know tactics, techniques, and procedures that can make a supposedly difficult attack trivial—or expose a “simple” defense as nearly impossible.

  • At a roughly $10 billion data center, former special-operations personnel walked around asking mundane questions, then identified “one or two ops” costing around $30,000 that could disable the facility for about a year. The vulnerabilities were described as critical and broadly present across current data centers.

  • No single expert holds the complete answer. Publishing creates an object specialists can challenge—“I actually know that’s wrong”—and the authors explicitly treat the report and their list of security techniques as living work rather than a finished blueprint.

12. The decisive security choices are being poured into concrete now

  • If late-2026 or 2027 superintelligence timelines are plausible, the immediate task is finding “one-way doors.” Data centers can take around 18 months to construct, so facilities being designed now may host decisive runs; later discovery of a physical flaw may be prohibitively expensive.

  • TEMPEST attacks are the report’s clearest example. Malware can access memory in patterns that encode data, while a nearby radio receiver reconstructs information from electromagnetic emissions—even when the target computer is fully air-gapped.

  • At data-center scale, mitigation may require a few feet between compute racks and data-hall walls. Without that clearance, an apparently harmless visitor on the other side could collect emissions with concealed equipment; retrofitting could require knocking down and redesigning the facility.

13. The soft underbelly sits beneath the GPUs

  • TSMC manufactures the leading chips in Taiwan, where Chinese interest needs no explanation. The guests point to executives who left TSMC, helped establish SMIC, and transferred knowledge as evidence that personnel and process leakage have operated for roughly two decades.

  • Less celebrated components may be more useful as backdoors. ASPEED reportedly accounts for 70–80% of the baseboard-management-controller market. These controllers run separately from GPUs and CPUs and can hold read/write access to firmware—“the soft underbelly” of high-performance computing infrastructure.

  • Power equipment creates another chain of dependencies. Transformer supply chains include Chinese-made components, and the guests say China has already used transformers to plant backdoors in American infrastructure; this is not merely a speculative future threat.

  • The problem is still narrower than rebuilding the entire US economy: only a limited number of strategic facilities must reach the highest standard. The report withholds some identified supply-chain techniques, but the authors say geographical and technological concentration makes meaningful risk reduction possible.

14. China’s grid leverage can become a wartime bargaining chip

  • In a Taiwan conflict, China could pair military operations with propaganda and disruption of US electricity or communications: “turn out the lights” and force Washington to manage domestic chaos while matériel runs short across a high-intensity first week or two.

  • Nation states expose only the least valuable capability that achieves the objective—“How can I learn without teaching?” Consequently, public incidents establish a high floor for Chinese capability but reveal little about the ceiling of access already embedded across infrastructure.

  • A compromised transformer might be remotely bricked through firmware or rewired to fail explosively. Recovery would require replacing equipment across unpredictable locations while electricity and communications are impaired—and the United States manufactures few, if any, of some critical units.

  • Reciprocal visibility is weak. Satellite imagery might reveal a centralized 10-gigawatt, Three-Gorges-scale cluster, but distributed-computing systems such as Streaming DiLoCo, DiLoCo, Prime Intellect, and Together AI complicate monitoring; the guests argue America needs options to hold Chinese infrastructure at risk in return.

15. Security means forcing attacks into view, then imposing consequence

  • Nathan imagines an underground Nevada complex, domestic supply chains at any cost, resident staff, and pervasive surveillance. The guests correct the literal picture: they do not advocate burying data centers, and their practical work seeks a Pareto frontier between compute, cost, time, and security.

  • A serious project would be embedded in national-security and counterintelligence systems, with agencies collecting evidence of attempts to sabotage or exfiltrate it. The aim is not blocking every attack; it is raising costs until an adversary must marshal enough resources to create a detectable signature.

  • A cruise missile could still destroy the facility, but it can be detected. Detection enables attribution and retaliation—the mechanism the guests believe actually stabilizes great-power competition, however unappealing the analogy to rival gangs exchanging consequences may sound.

  • Their complaint across administrations is insufficient consequence for cyber operations, infrastructure compromise, and Chinese police stations operating within US borders. Fear that every response leads directly to nuclear escalation teaches adversaries that increasingly damaging “free shots” will remain free.

16. The report leaves one narrow route to cooperation—but not in time

  • Nathan reads the project requirements as a warning: centralization, intrusive vetting, fragile supply chains, and AI-control challenges are reasons to avoid a Manhattan Project. The guests call the divergent reception a “Rorschach test,” while arguing that most safeguards are necessary even if five private labs continue independently.

  • Current behavior falls far short of treating AI as a WMD-scale capability. Stargate publicly advertises a $500 billion buildout and its facilities are widely known; an OpenAI insider reportedly found two critical vulnerabilities enabling access to model weights; and sensitive intellectual property has leaked through ordinary internal channels.

  • China’s open releases do not prove it lacks stronger private systems. The guests put the chance that future DeepSeek, Huawei, Tencent, or similar releases lack high-level CCP approval at “0%,” citing incentives to flood Western markets with models that avoid subjects such as Tiananmen, reduce costs for Western competitors such as Meta, and seed trusted agentic models that might contain backdoors.

  • The prerequisite for a grand bargain is verifiable hardware: secure, governable chips and tamper-detecting enclosures. Jeremie would spend “a few billion bucks” pursuing that dark-horse path, but design, fabrication, scaling, and resistance to sustained adversarial access are slow; experts doubt that maintaining meaningful chip capacity through such enclosures would be straightforward.

  • Edouard credits MAIM with identifying the strategic logic of capability degradation: accelerating oneself shortens timelines, so creating an alignment margin may require slowing an adversary. His caveat is that state and proxy capabilities overlap too much for finely tuned “scalpel” security thresholds—another reason no clean equilibrium is yet available.