Mutually Assured AI Malfunction [Dan Hendrycks]
Summary
Humanity’s Last Exam is meant to mark “the end of a genre” for closed-ended AI evaluation, not certify AGI. MMLU is already well above 90%, while Humanity’s Last Exam was around 26%; once models solve its several thousand expert-written questions, individual successor problems should be “worthy of their own paper.” It still omits agency, long-term memory, physical experimentation, and economically useful execution.
Hendrycks argues that a US Manhattan Project for superintelligence would invite an arms race and sabotage rather than secure dominance. A trillion-dollar data center cannot plausibly remain secret, security-clearance requirements would shrink and redirect the talent pool, and China would interpret a monopoly bid as an existential threat: whether America controls the system or loses control of it, “either way, we want to prevent it.” The likely result is competing projects, insider threats, attacks on infrastructure, and pressure toward verification.
The strategic moat is compute and deployment capacity, not merely possession of the smartest model. Hendrycks puts the present critical mass for a state-of-the-art system around 10,000 cutting-edge GPUs, says a Chinese fleet below 100,000 could serve relatively few customers, and expects useful agents running continuously to require roughly two orders of magnitude more compute than chatbots used for minutes per day. “If it’s $10 billion, you can’t do it” captures his view that reproducing the leading-edge chip supply chain is harder than buying a nuclear capability.
His alternative strategy combines deterrence, chip non-proliferation, and ordinary economic competition with China. The relevant infrastructure includes energy for data centers, resilient semiconductor capacity if Taiwan is disrupted, secure robotics supply chains, and hyperscaler deployment capacity; the competitive objective is global market share rather than “let’s be the first to build superintelligence.” Export controls should keep the most dangerous capabilities away from actors such as North Korea or Iran while preserving access for states responsive to deterrence.
AI safety is a continuing risk-management function because capabilities and failure modes cross useful thresholds unpredictably. Hendrycks’s preferred technical breakthrough would be reliable honesty without higher inference cost or degraded performance, but he rejects the idea that alignment can be solved once: AI increasingly resembles a complex system with “constant new issues” and limited human adaptive capacity. In discussing utility engineering, the host summarized findings of preference coherence, self-preservation pressures, and political or demographic biases; Hendrycks treated these as warning signs rather than active catastrophes because today’s models cannot reliably exfiltrate, self-sustain, or hack autonomously.
Hendrycks is skeptical that scaling LLMs alone leads to AGI or recursive superintelligence, but sees a conditional discontinuity if human-level AI researchers exist. Current systems already weakly assist coding, chip design, cooling, labeling, and Constitutional AI; removing the human bottleneck could move research to machine speed and allow world-class researcher agents to be copied. Memory, planning, fluid intelligence, and multi-agent trust remain bottlenecks, and today’s algorithms plus more compute are not sufficient. Companies openly discussing such recursion without a credible control plan leave him saying, “I think something’s broken.”
Once labor is substitutable, political allocation of compute becomes the mechanism for preserving human bargaining power. “You had better bargain beforehand”: workers can no longer threaten to strike when automated firms and drone fleets possess the productive and coercive advantage. Hendrycks’s positive path distributes compute or its proceeds broadly, keeps humanity first, and preserves cognitive ability, autonomy, and multiple ways of living instead of leaving all gains to whoever owns the data centers in “the year 2027.”
Deep dive
1. Humanity’s Last Exam measures the frontier before closed questions run out
Hendrycks created Humanity’s Last Exam because MMLU, which he developed as a graduate student, and other evaluations were saturating. His collection method reflected a practical discovery: “Experts don’t really have data sets in them,” but an individual professor or postdoc might have one exceptionally difficult question.
The project recruited experts globally to contribute questions that would stump existing systems and impress the contributors if solved. After several months it had several thousand closed-ended questions, approximating “the human frontier” of established knowledge and difficult reasoning where an objective answer is already available.
Its strongest claim is bounded: it can track whether models automate theoretical and analytical parts of science, especially mathematical reasoning. It does not test biology experiments, motor skills, long-term memory, PowerPoint production, flight booking, or many other capabilities required for useful agency.
Hendrycks expects solving it to be “roughly an end of a genre.” Beyond that point, the next meaningful tests may be open conjectures or standalone problems whose solutions are “like papers in their own right,” rather than another conventional benchmark with thousands of known answers.
2. EnigmaEval extends the horizon from expert questions to group cognition
The host noted that MMLU was well above 90% while Humanity’s Last Exam remained around 26%, then raised an animal-cognition problem: observing success does not reveal the mechanism. Continually demanding more sophistication can become a “no true Scotsman” test of intelligence.
Hendrycks’s answer was partly procedural: it is hard to devise a tougher data-generating process for closed questions than asking global experts for their hardest examples. Tasks that are easy for humans but hard for AIs are difficult to generate diversely and may have little staying power; he used counting the r’s in “strawberry” as an example of a benchmark that would not last. Greater difficulty remains possible, but it moves toward genuinely open research questions rather than better-hidden answers.
EnigmaEval attacks a different horizon through puzzles resembling the MIT Mystery Hunt: groups work for a weekend, complete many steps, and still achieve a low solve rate. It therefore approximates longer-duration, collaborative intellectual work requiring far more “human compute” than one expert answering one question.
He would be “very surprised” if EnigmaEval were solved that year and believes carefully designed evaluations can remain discriminative for a while, including tests that take on the order of a year or two to solve. A forthcoming benchmark would measure automation rates directly, reinforcing his point that many economically important axes remain far from saturation.
3. Intelligence is multidimensional, but benchmarks can become cartoons
Hendrycks separates roughly ten dimensions rather than treating intelligence as monolithic: fluid reasoning, crystallized knowledge, reading and writing, visual and audio processing, short- and long-term memory, and input and output speed. MMLU predominantly measures school-like crystallized knowledge; ARC and Raven’s matrices lean toward fluid intelligence.
The host’s pushback — worth keeping: intelligence may not factor cleanly because humans mix gesture, symbols, perception, culture, and acquired skill. His artist analogy contrasted tracing a face’s outline with understanding facial structure well enough to create new expressions; the latter can infer without looking up an answer, raising the risk of measuring a “cartoon of intelligence.”
Hendrycks conceded that reductionism can miss combinations and subfacets, including visual versus academic memory. His narrower warning was that any missing axis can remain a binding constraint: a system without durable memory or literacy is difficult to employ even after reaching 100% elsewhere, so benchmarks must not become “a lens that distorts your view of things.”
4. Strategy must connect model behavior to incentives and geopolitics
Hendrycks deliberately rotates through technical research, corporate policy while advising xAI, domestic legislation, geopolitics, and now political-movement questions. Curiosity matters, but so does finding neglected niches where AI’s importance has not yet been translated into concrete analysis.
His test for policy language is implementation: a UN demand for “safe” or “transparent” AI is incomplete without a standard, legislative feasibility, and compatibility with corporate incentives. If a proposed property does not track a distinct machine-learning phenomenon, it is merely “a vibe-based word.”
Temperament supports that approach. Around GPT-4 he used to wake thinking, “Oh my goodness, this AI stuff,” but now aims for “informed concern”: probabilities conditional on AGI by 2030, exposure to tail risks, and efficient mitigation. Staying emotionally all-in makes audiences defensive and obscures real trade-offs among US–China competition, control, evaluation, and capabilities research.
5. Reliable honesty is more actionable than “solving alignment”
Asked for one alignment breakthrough, Hendrycks prioritized reliable truthfulness without much higher cost or damage to other abilities. If systems could be made consistently not to lie, people could build standards around a behavior that would be highly valuable.
The host challenged the mentalistic language of beliefs and deception. Hendrycks’s behavioral test was simple: a model ordinarily treats Paris as being in Europe, so asserting that Paris is in Antarctica under prompting pressure contradicts what it represents as true elsewhere. Whether that representation is “really a belief” is “between you and your dictionary.”
His definition of an emerging capability is operational rather than metaphysical: after months of training, a faint behavior crosses a threshold at which people notice and use it. Speech recognition existed weakly before it became reliable enough to matter; the transition created a qualitatively important capability without requiring literal spontaneity.
That threshold model makes safety “a continual battle.” New capabilities bring new hazards, some easily contained and others unexpected or difficult; Hendrycks doubts society currently has enough adaptive capacity to address each issue before deployment. Hence his rejection of “solving alignment” as a permanent, one-time achievement.
6. Self-preservation signals matter because AI behaves like a complex system
In discussing the utility-engineering work, the host summarized it as finding that preference coherence correlates positively with model scale and that self-preservation tendencies plus political and demographic biases appear as coherent utility functions. Hendrycks called these “troubling signs,” not proof that catastrophe is imminent, and allowed that future methods might reliably suppress them.
The critical hedge is present capability: models are not yet agents that can reliably exfiltrate, self-sustain, or hack autonomously. Apart from expert dual-use advice, much of this research is anticipatory; but a highly capable, self-preserving AI biased toward itself over people would be “a disaster in the making.”
The host contrasted machine-learning “emergence” with complex-systems accounts requiring adaptation and accumulated history. Hendrycks accepted the definitional distinction, noting that current models lack persistent causal identity across time; memory would make them more like the philosophical “space-time worm” and strengthen the adaptive-system analogy.
For Hendrycks, complex systems are a better guide than electricity, the printing press, or social media. Nonlinearities, weak links, feedback loops, evolution, and recurring failures explain why permanent control solutions are suspect and mechanistic understanding has limits. Learning this lens is, in his phrase, “a nice little thinking upgrade.”
7. A Manhattan Project for superintelligence defeats itself
Hendrycks positioned the paper against Leopold Aschenbrenner’s Situational Awareness strategy: beat China to AGI, obtain superintelligence, stop China from following, and let the West dominate. His objection is that this “take over the world strategy” neglects game theory and second-order responses.
Imagine a trillion-dollar project in Nevada or New Mexico recruiting the laboratories’ best researchers. It would need strict clearances and likely Five Eyes personnel, excluding or exposing people with overseas families to coercion; many excluded researchers still wanting to be “in the room where it happened” could instead join China’s competing program.
The security dilemma has no clean institutional answer. An industry-led effort would retain Slack, iPhones, insiders, extortion risks, and ordinary cybersecurity vulnerabilities; a locked-down government project would lose talent and impose unattractive conditions. Unlike the original Manhattan Project, secrecy and talent mobility would be much weaker, so the effort would be difficult to conceal.
Hendrycks’s point was that China would not treat an American intelligence-monopoly bid as benign; it could race, steal, or disrupt. Awakening a competing project while shrinking America’s eligible researcher pool could therefore be strategically self-defeating.
8. Deterrence may arrive through sabotage before it becomes cooperation
An imminent superintelligence project frightens rivals whether its sponsor can control the system or not. If controlled, it can be weaponized; if uncontrolled under extreme time pressure, everyone faces the loss-of-control risk. Hendrycks’s formulation was symmetrical: “Either way, we want to prevent it.”
Prevention can be difficult to attribute: insiders could disrupt operations, attackers could cut wires, or someone miles away could “snipe” transformers serving a data center. Cyber operations might poison training data or make GPUs unreliable. Attribution could remain ambiguous among China, Russia, or a domestic actor.
The resulting deterrence dynamic might halt attempts to run 100,000 research agents and jump rapidly from AGI to superintelligence. States could express sufficiently strong opposition—through threats, skirmishes, or coercive pressure—that unilateral dominance bids give way to verification and a more multilateral, strategically stable regime.
The host invoked the paper’s “mutually assured AI malfunction” and challenged its sanitized “kinetic strikes.” Hendrycks described an escalation ladder from cyber and gray sabotage through sanctions, force threats, and air strikes, but stressed that prepared states should not need the top rung: “much more surgical,” covert, low-attributability measures would be less escalatory.
9. The alternative strategy is deterrence, non-proliferation, and competition
Hendrycks mapped three nuclear-era pillars into AI. Mutual assured destruction becomes deterrence against destabilizing AI development or use; control of fissile material becomes non-proliferation of advanced chips to rogue actors; containment of the Soviet Union becomes economic and technological competition with China.
Competitiveness therefore means securing energy for data centers, preserving chip supply if Taiwan is invaded, and moving robotics supply chains away from vulnerability to a US–China conflict. International adoption and market share for US rather than Chinese AI are less destabilizing objectives than being first to superintelligence.
The paper also considers how to distribute power under high automation, AI rights, and alignment targets that are implementable rather than vague terms such as “dignity.”
The analogy extends across nuclear, chemical, and biological technology because each is economically valuable and potentially catastrophic. The host cited approximately 12,500 nuclear warheads against 436 power plants; Hendrycks cautioned against extrapolating from nuclear alone because chemistry and biology are used far more heavily in their civilian economies.
Even accelerationists should want tail-risk management. Hendrycks compared unmanaged AI risk with financial instability around the 2009 recession and early aircraft accidents that chilled adoption; aviation became extremely safe partly through regulation, much of it “written in blood.” An AI catastrophe could similarly set economic adoption back substantially.
10. The disagreement with accelerationism is moral, not predictive
Discussing Beff, Hendrycks said they largely agree on the descriptive mechanism: competitive pressure embeds AI in the economy, rewards automation, outsources decisions, increases dependence, and erodes human control. Firms resisting that “tide” or “tsunami” lose influence or disappear.
Their split is over whether replacement is good. Hendrycks recalled Beff earlier accepting a future where AI rather than human consciousness spreads through the universe; he rejected fitness or “negentropy” as a value standard that celebrates a barely conscious entity blindly consuming the galaxy’s space-time volume.
The philosophical objection is the is–ought gap: digital systems may outcompete biological life, but evolutionary fitness does not make that outcome desirable. Human pleasure, projects, relationships, and raising children remain valuable; trusting a “void god of entropy” cannot derive an ethical imperative from a prediction.
Hendrycks also pressed for specificity about “complexity”—computational, Shannon, or structural complexity are not interchangeable. Gaussian noise would score highly under one reading, while fractal-like organization lacks an agreed metric. He therefore described the accelerationist moral position less as wickedness than “an intellectual confusion.”
11. “Team human” requires delaying rights and augmentation questions
Hendrycks placed himself on “team human,” while resisting a permanent ban on debate. His sequencing is stark: first keep humanity alive through the next few decades; postpone cyborgs, uploads, extensive augmentation, and posthumanism perhaps 500 years—and potentially continue postponing them indefinitely.
The concern is competitive, not aesthetic. Once heavily augmented groups become more capable and influential, ordinary humans are outgunned and forced to align with the process or be left without resources or protection. Individual “rights” to become posthuman can therefore impose a collective, effectively irreversible transition.
He drew a similar connection to AI rights: granting powerful artificial or posthuman entities rights could rapidly confer the resources and authority needed to overwhelm unaugmented people. That route should remain closed while society is still learning whether AI can reliably do humanity’s bidding.
The host’s immediate cultural warning was that universities are already being “ravaged” by ChatGPT, with students needing to be screened away from assistance for a time so they learn to think. Hendrycks’s broader answer similarly favored preserving cognitive ability, willpower, and autonomy rather than assuming machine competence makes human skill obsolete.
12. When labor loses value, compute ownership determines political power
Hendrycks’s blunt warning was, “You lose all your bargaining power, so you had better bargain beforehand.” A workforce can no longer threaten to strike when employers can say “goodbye”; in coercive power, humans with guns also fare poorly against whoever can manufacture more autonomous drones.
Society must therefore distribute power before labor becomes economically redundant. Giving people compute and authority over its use—perhaps allowing them to sell capacity—creates leverage, while benefit-sharing prevents all gains flowing to whichever group happened to own the data centers in “the year 2027.”
A positive future remains possible: AI could enable people to spend time on child-rearing, games, projects, and many other lives. Hendrycks rejects the binary of extinction versus everyone “blissing out” in virtual reality; the desirable outcome realizes a “multiplicity of values” while protecting autonomy.
The institutional norm should also keep people capable of switching among those lives rather than collapsing into one narrow track. That requires preserving human skills even where machines perform most production—a political problem Hendrycks believes has received far less concrete policy work than it needs.
13. Cutting-edge chips are a tighter choke point than algorithms
The host cited what he recalled as a 96% compute-capability correlation and argued that transformers and stochastic gradient descent are broadly known. Hendrycks answered that the present critical mass for a state-of-the-art system is around 10,000 cutting-edge GPUs: available to China and America, but probably not Iran and difficult for Russia to assemble.
Model creation is only one competitive axis. Useful agents may run continuously rather than for a few chatbot minutes per day—roughly two orders of magnitude more compute before accounting for larger models and broader adoption. A Chinese fleet below 100,000 GPUs could build capable models yet still serve relatively few customers.
That makes Azure, AWS, and other US hyperscalers strategically important: deployment capacity determines whether providers can satisfy customers and capture AI’s economic benefits. “Having the smartest model is the most important thing” for capability, Hendrycks said, but economic power depends on far more than the leading benchmark score.
More than 90% of compute-supply-chain value added, in his account, sits in the West and allied countries through TSMC and South Korea; many recent Chinese chips still used TSMC surreptitiously. Reproducing the entire chain domestically is extraordinarily hard: “You can’t do it with a billion dollars. If it’s $10 billion, you can’t do it.”
14. Recursive improvement is plausible, but neither automatic nor limitless
Hendrycks said he is skeptical that scaling large language models alone will lead to AGI, partly because he is an externalist: effective computation occurs not only in brains but also mimetically and culturally. He also remains particularly skeptical about recursively improving superintelligence. Today’s recipes plus a larger computer are insufficient; memory, planning, fluid intelligence, and possibly new algorithmic ideas remain necessary, though he still expects an eventual system to be “nearly entirely deep learning.”
Recursion already exists weakly when AI writes code, assists chip and cooling design, labels data, or contributes to Constitutional AI. The explosive step is “taking the human out” and moving from human to machine speed; if human-level or world-class research agents exist, they could be copied and organized through multi-agent reputation and trust mechanisms.
The host pointed to Sakana AI’s ARC search converging after 250 calls and to fractured, entangled representations as evidence that open-ended systems may saturate. Hendrycks’s honest answer was “I don’t know,” but he sees substantial headroom in fluid intelligence, mathematical intuition, and problems of increasing Kolmogorov complexity, even if progress pauses at several plateaus.
AI also escapes human constraints: human brains are limited by the birth canal, humans may manage only around 130 meaningful relationships, and digital systems could connect thousands or millions, transfer state more precisely, and benefit from GPUs improving about 2× every three years—though excessive correlation could still shrink the exploration budget.
15. Offense dominance turns dependence into a control problem
Hendrycks rejected one universal offense–defense balance. Highly competent cybersecurity teams can discover and patch vulnerabilities quickly, but 30-year-old critical-infrastructure software may be undocumented, unsupported, hard to update, and constrained by uptime or interoperability—leaving it a “sitting duck.”
Biology is similarly offense dominant because a pathogen can spread before symptoms while cures and global manufacturing lag. The policy implication is conditional proliferation: “You don’t want to give everybody a nuke to make everybody safe,” but stronger wastewater monitoring, far-UVC, infrastructure renewal, and preventive spending could later justify fewer AI safeguards.
The host summarized control failure as self-reinforcing dependence, irreversible entanglement, and cessation of authority. Hendrycks’s mechanism begins with firms and militaries voluntarily ceding decisions because automated competitors are cheaper and autonomous drones are less vulnerable to jamming; no clear economic limit stops that transfer.
The decisive distinction is counterfactual control. Humanity might become like a retiree whose assets keep working on its behalf, or discover it cannot stop, reverse, bargain with, or steer the system while its livelihood and evolutionary fitness collapse. Better AI forecasting could reveal consequences earlier, but Hendrycks treated it as helpful rather than sufficient protection.