No Priors Ep. 105 | With Director of the Center of AI Safety Dan Hendrycks
Summary
Hendrycks argues that AI safety is primarily a geopolitical and economic problem, not something laboratories can solve through model alignment alone. Labs are “predetermined to race” and can add basic safeguards, but they cannot design away labor disruption, concentrated power, or US–China strategic competition. Even perfectly obedient national systems could be integrated rapidly into competing militaries and force both sides toward higher risk tolerance.
Near-term capability is uneven: AI is not yet decisive across national security, but reasoning models have recently begun to create biological national-security implications and are rounding the corner toward expert-level literature knowledge and wet-lab assistance. Hendrycks doubts today’s systems could independently execute a devastating grid attack, while warning that biological capabilities have recently become more consequential. His proposed control is simple identity-gating: an unidentified user asking how to culture a virus gets refused; a legitimate biotech company can “just speak to sales.”
The military discussion extends well beyond chatbots into drones, electronic warfare, command-and-control, and situational awareness. Hendrycks points to drones, weapons research, and better detection of nuclear submarines or hardened launch sites, which could disrupt second-strike capabilities without itself being a weapon. Guo adds electronic warfare and a Wall Street analogy in which human judgment became rows of people clicking “accept, accept, accept.” Reliability therefore becomes both a product constraint and a strategic-security variable.
Neither unilateral restraint nor a clean race to superintelligence survives Hendrycks’s second-order test. A voluntary pause without verification or force merely weakens the participant, but a visible trillion-dollar desert cluster intended to secure dominance would invite espionage, cyberattack, or preemption. He says that in some places, more than 30% of employees at top AI companies are Chinese nationals; excluding that talent would damage the US effort and potentially strengthen China.
Mutually assured AI malfunction, or MIM, is proposed with Eric Schmidt and Alexandr Wang as deterrence against destabilizing superweapon projects rather than a halt to ordinary AI competition. States would monitor rival programs and retain cyber options to disable data centers attempting a decisive strategic breakout: “We will not build a superweapon, and we’re going to be watching for other people building them too.” Competition in chips, drones, and conventional systems would continue.
Compute controls should prioritize chip visibility and nonproliferation to rogue actors, not assume China can be denied advanced capability indefinitely. Guo raises DeepSeek and recent releases as a challenge to a simple 100,000-chip compute-security premise; Hendrycks agrees that efficient training undermines a Manhattan Project strategy and says China could ultimately steal model weights. He nevertheless argues that basic end-use checks might have revealed where the “10% of NVIDIA’s chips” discussed in relation to Singapore were going.
The key capability inflection is reliable agency, not simply higher scores on closed-ended academic benchmarks. Humanity’s Last Exam measures the remaining frontier of closed-ended expert knowledge; near-ceiling performance would imply something like a superhuman mathematician or STEM scientist. Guo also mentions Enigma, though Hendrycks’s answer focuses on HLE and the broader closed-ended/agent distinction. Agents remain “near the floor”: models can solve difficult physics while failing to book a flight, and Hendrycks expects the economic “vibes really shift” once they can reliably complete multi-hour digital work.
Deep dive
1. Safety is constrained by geopolitics before model design
Hendrycks entered AI safety because the technology’s trajectory looked consequential while its unpleasant tail risks were “systematically under-addressed.” His objective is broader than preventing catastrophe: understand the trajectory, channel it productively, and manage disruptions that technical model work cannot settle.
His institutional diagnosis is blunt: labs can refuse requests such as “help me make a virus,” but they are “kind of predetermined to race.” A company that meaningfully opts out risks becoming irrelevant, while different refusal data cannot change the prospect of mass labor disruption or automated digital work.
Alignment is therefore a subset of safety, not its synonym. An AI reliably obedient to the US and another reliably obedient to China could still be embedded rapidly into competing militaries; strategic pressure would raise risk tolerance even if both systems performed their principals’ bidding perfectly.
Guo’s pushback — worth keeping: lab leaders would not say they can do nothing, and every participant has economic equity in the outcome. Hendrycks allows that companies can pursue controllability research and policy advocacy, but maintains that the broader problem is geopolitically determined.
2. Near-term danger is jagged, but access controls buy a lot
Hendrycks distinguishes trajectory from present capability: in many national-security domains AI is not yet powerful, though “this could well change within a year’s time.” Today’s systems probably cannot execute a devastating grid attack by a malicious actor, while reasoning models are rounding the corner toward expert-level biological literature knowledge and practical wet-lab assistance.
Guo points to defensive cybersecurity and biology-discovery companies as evidence that competition produces near-term benefits. Hendrycks rejects a large safety-versus-benefit tradeoff here: unidentified users requesting step-by-step virus-culturing help can be blocked, while verified biotech customers receive the capability through an enterprise account — “just speak to sales.”
His threat map separates actors and applications: bioweapons make more sense as a non-state risk; cyber operations matter for both states and non-state actors. Hendrycks points to drones, weapons research including exotic EMPs, and situational awareness, including locating submarines or hardened launchers. Guo extends the discussion to electronic warfare, radio, radar, targeting, and command-and-control, especially in Ukraine.
Guo offers the Wall Street anecdote, where human oversight degraded into rows of people clicking “accept.” Hendrycks says automating more decision-making would not surprise him and turns the issue into one of reliability research.
3. Pauses and monopoly races both fail the second-order test
A voluntary pause without “teeth” simply lets worse actors advance. Treaties require verification, credible enforcement, or a threat of force; cyberattacks and corporate espionage offer little evidence that norms alone would hold.
A straight race is defensible for chips or drones, but Hendrycks rejects racing to turn superintelligence into a weapon. Espionage makes durable monopoly unlikely, yet removing Chinese researchers would also be self-defeating: he says that in some places more than 30% of employees at top AI companies are Chinese nationals, and many could return to China. He favors easier immigration for very talented AI researchers while saying the issue should remain separate from Southern-border policy.
China would not watch passively as the US built a trillion-dollar compute cluster “totally visible from space” as a bid for dominance. Hendrycks compares the logic to early nuclear proposals for preempting the USSR: multinational talent, interdependence, and limited time may mean the supposed window for a secure monopoly never exists.
4. MIM turns shared cyber vulnerability into deterrence
Mutually assured AI malfunction, or MIM, proposed with Eric Schmidt and Alexandr Wang, applies nuclear-style shared vulnerability once AI becomes pivotal and can automate AI research. A rival approaching a decisive superweapon would face espionage, sabotage, or a cyberattack on its data center; that expectation is intended to deter the destabilizing project without requiring every form of AI development to stop.
The analogy has limits. Hendrycks expects coordination where interests overlap — keeping dangerous capabilities from terrorists, as with chemical and biological weapons, and avoiding projects that could let one country “crush” another — while conventional competition in drones and other systems continues.
Calibration matters because, if AI chips become the “currency of economic power,” turning China’s export-control pain dial fully upward could strengthen its incentive to invade Taiwan. Hendrycks notes that China already wants to; this would give it additional reason.
The practical package is narrower: a CIA cell tracking foreign AI programs, Cyber Command purchasing disablement options, chip-location licensing, allied notification exemptions, and prioritized end-use inspections.
Against Guo’s DeepSeek challenge, Hendrycks agrees efficient training undermines the desert-cluster strategy: controls cannot robustly eliminate a great power’s capability or make it abandon the economic value of AI. China may obtain a fraction of the chips or steal weights anyway; controls remain useful for visibility and rogue-actor nonproliferation, particularly if officials actually investigate routes such as Singapore. He says AI-chip export controls were not a priority for Bureau of Industry and Security leadership and that basic checks could have exposed diversion to China.
5. Evals show oracle-like intelligence before reliable agency
Humanity’s Last Exam extends Hendrycks’s earlier benchmark work, including MMLU and the MATH dataset. Guo also mentions Enigma, though Hendrycks’s response here focuses on HLE and the broader closed-ended/agent split. Professors and researchers submitted unusually difficult research-grade questions with definitive, closed-ended answers, creating what he describes as a possible “end of the road” for exam-style academic evaluation.
Near-ceiling performance would roughly exhaust that benchmark genre and signal something like a superhuman mathematician or STEM scientist on closed-ended problems. It would not establish competence at open-ended work, where defining, pursuing, and completing a task matters as much as knowing an answer.
Agent evaluations should instead assign real digital tasks, allow a few hours of work, and check completion. Current systems remain “extremely defective as agents” and near the floor, although Hendrycks hedges that this “could possibly change overnight.”
The frontier remains jagged: systems answer hard physics questions but cannot fold clothes, and may surpass humans at mathematics while failing to book a flight. For verifiable reasoning, humans might still run an AI proof through a proof checker; by contrast, if systems develop better “taste,” that would be harder to confirm.
Hendrycks thinks AI is on track for “really good oracle-like skills” before it can reliably act on people’s behalf. Once it acquires agent skills, few barriers remain to major economic impact, and AI becomes “in a category of its own.”