Pioneers Insight Method Research Author
National Security Strategy and AI Evals on the Eve of Superintelligence with Dan Hendrycks
Back to Episodes

National Security Strategy and AI Evals on the Eve of Superintelligence with Dan Hendrycks

Summary

  • Hendrycks’s core call is that AI safety is primarily a statecraft problem, not solely a lab-level alignment problem. Labs are “predetermined to race”; they can cheaply gate expert bio access and address cyber risks, but obedient US and Chinese systems still leave strategic competition, rapid military integration, labor automation, and rising risk tolerance intact.
  • The immediate national-security picture is jagged: reasoning models are approaching expert virology capabilities, while Hendrycks does not think AI is currently relevant to a malicious actor’s devastating grid attack. That could “well change within a year’s time.” The discussion touches cyber defense, drones, electronic warfare, command-and-control reliability, and compute security as near-term application and risk areas.
  • Hendrycks rejects both a voluntary pause and a clean US sprint to superintelligence because neither survives adversarial response. Treaties need verification or force, model weights could be stolen, and a “trillion-dollar compute cluster in the desert totally visible from space” would look like a dominance bid that China would try to deter.
  • His proposed regime, mutually assured AI malfunction (MAIM), relies on shared vulnerability: states refrain from destabilizing superweapon projects because rivals can spy on or disable their data centers. Competition in chips and drones continues; coordination centers on preventing rogue-actor access, much as nuclear, chemical, and biological regimes separated rivalry from shared nonproliferation interests.
  • DeepSeek-style efficiency weakens capability denial against China, so Hendrycks would use export controls chiefly to track chips and constrain rogue actors while deterrence constrains great-power intent. His 80/20 package is concrete: a CIA AI-espionage cell, Cyber Command sabotage options, licensing and shipment notifications, and end-use checks—including investigating where the cited 10% of NVIDIA chips were going.
  • The economic phase change arrives when models move from impressive “oracle-like skills” to reliable agency. Humanity’s Last Exam is meant to expire closed-ended academic benchmarks near its ceiling—signaling superhuman performance in parts of math and STEM—but agents remain “near the floor”; once they can carry out digital work, Hendrycks expects “the vibes really shift.”

Deep dive

1. AI safety extends far beyond making models obedient

  • Hendrycks says he chose AI safety early because following AI’s trajectory made it look like “the most important thing during this century,” while its weirdness and unpleasant implications kept others away. The mandate was broader than failure prevention: think clearly, channel the technology productively, and cover tail risks that are “systematically under-addressed.”

  • His institutional split is blunt: “There aren’t that many safety efforts in the labs even now,” and companies are “kind of predetermined to race” or lose relevance. They can refuse “help me make a virus,” reduce terrorism or some accidents, and research controllability—but geopolitical competition determines far more of the outcome.

  • Alignment, in his taxonomy, is only a subset of safety. An AI can be reliably obedient to the US or China and still sit inside a competition that forces rapid military integration and higher risk tolerance; “whether they do what you want” does not remove concentration of power, strategic pressure, or the risks of never obtaining highly intelligent systems.

  • Guo’s pushback is that lab leaders would not say they can do nothing—and everyone has “equity in this equation.” Hendrycks’s reply is structural: design tweaks and refusal data do not stop mass disruption to labor or automation of digital work, while policymakers still think AI is “just selling hype” and do not believe company employees actually expect AGI in the next few years.

2. National-security capability is arriving unevenly

  • Hendrycks separates present capability from trajectory: he does not think AI is currently relevant to pulling off a devastating grid attack by a malicious actor, but that “could well change within a year’s time.” Virology has moved faster; reasoning models are approaching expert-level literature knowledge and may even assist with practical wet-lab situations, giving them national-security relevance already.

  • Guo points to Culminate and Sibyl in defensive cybersecurity, plus Chai and Somite in biotech, as examples of benefits arriving too. Hendrycks sees little trade-off: keep expert virology capabilities behind an enterprise relationship—“just speak to sales”—rather than exposing them to a new account asking how to culture a virus.

  • Beyond bio, Hendrycks’s stack includes state and non-state cyberattacks, drone development, exotic EMP research, and better situational awareness. Merely locating hardened launch sites or nuclear submarines could threaten second-strike capability and mutual assured destruction: an informational improvement, not itself a weapon, that is nonetheless “extremely disruptive” and destabilizing.

3. Voluntary pauses and a clean superintelligence sprint both fail

  • Hendrycks’s objection to voluntary slowing is enforcement: without verification or “some sort of threat of force,” restraint only makes the compliant actor weaker while worse actors advance. Guo’s comparison is cyberattacks and corporate espionage, where treaties and norms have produced little reliable compliance.

  • A sealed US sprint has the opposite flaw. Hendrycks notes that “30%-plus” of employees at some top AI companies are Chinese nationals; excluding them could send indispensable scientists to China and lower America’s odds of winning, while retaining them necessarily leaves information-security exposure. His immigration preference is still to make it easier for very talented AI researchers to stay.

  • The deeper failure is second-order reasoning: China would not watch passively while the US built a visible trillion-dollar cluster intended to produce a system that could “boss us around… till the end of time.” Hendrycks compares the idea to the brief nuclear-era argument for preemptively destroying the USSR, while saying AI’s multinational talent and knowledge flows mean the opportunity window does not really exist.

  • MAIM—mutually assured AI malfunction—turns that vulnerability into deterrence. If rapid automated R&D looked like a bid for decisive dominance, rivals could threaten cyberattacks against data centers or monitor labs through something as simple as “a zero day on Slack”; shared exposure discourages superweapon projects without ending competition in drones, chips, or conventional systems.

4. Deterrence constrains intent while compute controls police proliferation

  • Hendrycks’s immediate statecraft is deliberately unglamorous: create a CIA cell focused on foreign AI programs, then have Cyber Command prepare options to disable overseas data centers running destabilizing projects. Espionage supplies warning; building those capabilities and “buying those options” puts teeth behind deterrence.

  • For chips, he proposes a licensing regime modeled on tracking fissile material: allies receive exemptions if they notify the US when hardware changes location, while enforcement officers perform prioritized end-use checks. His complaint is implementation, not impossibility—he asks whether anyone went to Singapore to see where the cited 10% of NVIDIA chips were going, saying a basic end-use check would have revealed the destination.

  • DeepSeek’s efficient training does not overturn this view; it undermines the “Manhattan Project” fantasy that one superpower can robustly deny another the ability to build models. China will probably keep getting some fraction of the chips, and if controls were tightened completely it could steal model weights anyway. Controls should chiefly prevent proliferation to rogue actors such as Iran, potentially with Chinese cooperation, while deterrence constrains destabilizing intent.

  • Guo’s challenge is economic: if AI is powerful, no great power will renounce the capability. Hendrycks agrees and has no settled view on whether a magical three-times slowdown would help; turning the export-control “pain dial” fully against China could instead raise its incentive to invade Taiwan if AI chips become the currency of future economic power.

5. Evals expose a frontier that is brilliant but still unable to act

  • Hendrycks places Humanity’s Last Exam in the progression from MMLU and the MATH dataset to pre-ChatGPT work such as ImageNet-C. Professors and researchers worldwide contributed exceptionally difficult, closed-ended questions with objective answers—the kinds of problems they encounter in research—creating an attempt at the end of the road for exam-style academic benchmarks.

  • Performance near the ceiling would mean the closed-answer genre is “roughly expired” and could indicate a superhuman mathematician or STEM scientist in domains where such questions matter. It would not demonstrate open-ended competence, so saturation should not be confused with a system that can autonomously complete real work.

  • The next evaluations therefore give models digital tasks, let them work for hours, and check completion. Current systems remain “extremely defective as agents” and near the floor, although Hendrycks allows that this “could possibly change overnight”; without unsaturated tests, the public cannot see either the level or rate of improvement.

  • The frontier is radically jagged: systems can solve hard physics while failing to fold clothes, and may surpass humans in mathematics while remaining unable to book a flight. Proof checkers can validate superhuman mathematics, while qualities such as better “taste” may resist verification; for now, Hendrycks expects strong “oracle-like skills” before dependable agency.