Agency over AI? Allan Dafoe on Technological Determinism & DeepMind's Safety Plans, from 80000 Hours
Summary
- Dafoe’s central thesis is that technology creates options, but competition selects which options survive. Local actors can refuse a tool, yet if one rival converts it into military or economic advantage, holdouts must adopt or lose resources: “Technology doesn’t force us; it merely opens the door, and it’s military-economic competition that forces us through.” For investors, adoption can therefore look voluntary early and compulsory late.
- The strongest remaining agency lies in choosing the sequence, design, and safeguards of technologies before competitive pressures harden the path. Differential technological development means building the “seat belt before the car,” but Dafoe stresses how difficult it is to identify two viable paths and forecast their downstream consequences. Nathan Labenz sharpens the near-term implication: only a few hundred or thousand people influence frontier compute decisions, so staff coordination or even individual dissent can still change what gets built.
- Alignment alone does not guarantee a good outcome because faithfully aligned systems can serve principals who remain locked in conflict. Dafoe argues that “global coordination is almost a necessary and sufficient condition” for deploying imperfectly aligned AI prudently, while aligned great powers could otherwise automate nuclear brinkmanship, trade conflict, or cyber escalation. Cooperative AI is the marginal bet that agents should acquire bargaining, communication, and commitment skills before raw capability outruns institutions.
- The comforting “super-cooperative AGI” hypothesis remains unproven. AI agents may communicate rapidly and be copied or tested as mediators, but they could also possess alien goals, linear utility over resources, hidden backdoors, and less mutual transparency than humans. Dafoe distinguishes cooperative skill from niceness: friendly assistants are weak evidence that strategically deployed agents will bargain safely.
- AGI is not a single humanlike threshold but a jagged region within a high-dimensional capability space. The route matters: a system might become superhuman at materials science while remaining strategically naive, or become exceptionally skilled at cooperation before gaining broader technical power. That sequencing creates a genuine dispute between those who prefer socially simple systems and Dafoe’s concern that technical acceleration without coordination capacity could overwhelm society’s ability to adapt.
- Frontier evaluations are becoming decision infrastructure, but they still miss capabilities unless models are given the right tools and scaffolding. On the paper’s five-point scale, Gemini 1.0 scored about 3 on persuasion, roughly 1–2 on cybersecurity and self-reasoning, and 2 on self-proliferation; Project Naptime and, to Dafoe’s understanding, OpenAI’s o1 illustrate how elicitation can sharply raise apparent capability, including with an older foundation model. Dafoe’s answer is layered evidence—lab evals, human studies, “evals in the wild,” forecasting, staged release, monitoring, and reversible access.
- Governance becomes harder as capabilities diffuse, making responsible frontier leadership a recurring defensive requirement rather than a one-off moat. Dafoe cited Waymo’s report as showing 2× fewer incidents involving police at the crash scene and, he believed, 6× fewer injury crashes, alongside AI opportunities in medicine, tutoring, weather, materials, fusion, and AlphaFold. The upside remains tangible, while diffusion raises the prospect that stronger frontier systems may need to defend against cheaper lagging ones.
Deep dive
1. Frontier employees still possess more short-run agency than determinism implies
Nathan Labenz opens with a Cold War story: his uncle’s nuclear-launch crew privately agreed that if a real firing order arrived, “we are all going AWOL.” The analogy is imperfect, but it foregrounds individual responsibility inside systems whose long-run strategic logic can otherwise feel inescapable.
Today’s AI frontier is unusually concentrated. Nathan estimates that a few hundred, perhaps a few thousand, people sit close to decisive compute allocations; elite ML talent is scarce, not every idea can be scaled simultaneously, and the “jagged edge” of intelligence leaves meaningful discretion over which capabilities receive priority.
Technical staff already demonstrated collective leverage during Sam Altman’s firing and reinstatement, while one former employee’s refusal to sign a non-disparagement agreement showed how individual action can alter institutional practice. Nathan’s warning is that a “glorious AI summer” may precede an AI cold war, making it worth deciding now which development paths could justify dissent or departure.
2. Dafoe moved inside DeepMind because pivotal decisions depend on who is in the room
Allan Dafoe’s Frontier Safety and Governance team has three pillars. Frontier safety studies emerging dangerous capabilities, forecasts their arrival, and develops mitigations; frontier governance advises on norms, regulation, and institutions; frontier planning looks toward AGI and asks what Google DeepMind, Google, and society may soon need to confront.
The team itself is small, but it works across Gemini safety and alignment, responsibility, policy, and other Google groups. Dafoe describes Google DeepMind as the company’s center of specialization for frontier models and “where the heart of the thinking” about frontier-policy questions occurs.
Dafoe left the Centre for the Governance of AI because advising Demis Hassabis and Shane Legg from inside offered more information and “more surface area” for influence. His historical model is Hamiltonian—“who’s in the room”—because crises often turn on decision-makers’ ideas, character, safety orientation, competence, and wisdom; good intentions paired with “clumsy hands” may still end badly.
Rob Wiblin’s scale marker is deliberately rough: AI is “certainly not more than 0.1%” of total revenue today, while its eventual share might exceed 10% and perhaps approach 100%. Dafoe’s career advice therefore remains “hop trains as soon as they can”; despite rapid growth, he considers this “still early days.”
3. Macro trends make history look less voluntary than it feels from inside
Dafoe began with the question, “Who, if anyone, controls technological change?” His dissatisfaction was with the intuitive model that history equals the sum of everyone’s effort. General-equilibrium systems can resist each unit of effort with equal or greater counterpressure—or amplify it—so local intention alone says little about lasting impact.
The empirical puzzle is the regularity of macro history. Germany and Japan returned to their prewar growth trajectories within less than a decade of World War II’s devastation; Moore’s law traced an unusually precise exponential line; and scaling laws now permit forecasts of model size and loss years ahead.
Similar directionality appears in civilization’s maximum energy processing, building height, material durability, and transportation speed. As Robert Wright’s archaeological formulation has it, “the deeper you dig, the simpler the society whose remains you find”; some technological orderings also look constrained, since nuclear power is difficult to imagine preceding coal power.
Yet these patterns cannot simply be credited to collective human will. The Agricultural Revolution appears to have reduced median health and welfare for a long period while enabling inequality and warfare. Dafoe’s conclusion is not that humans lack agency, but that its effect depends on timing, power, resource constraints, and which systems remain functional under selection.
4. Constructivism explains local choice while struggling with emergent constraints
Earlier theorists sometimes endowed technology with near-agency. Langdon Winner discussed technological autonomy, Lewis Mumford’s capital-M “Machine” reduced people to supporting cogs, and Jacques Ellul’s “la technique” described functional imperatives that leave humans choosing under coercion.
Social constructivists reacted by looking closely at actual decisions. Under the microscope they found people, interests, ideologies, competing designs, dead ends, and uncertain visions—not an autonomous machine dictating that bicycles, airplanes, or infrastructures take one predetermined form. Rob captures the critique with: “Is the machine with us in the room right now?”
Dafoe’s concession is methodological: ethnography and microhistory answer real questions about how choices were made. The error was using those tools to dismiss macro claims they were poorly equipped to test. The political appeal of the constructivist view is that it preserves citizens’ sense of agency; strategically, Dafoe wants a model that directs effort toward points where structure will not simply push back.
5. Technology embeds politics, accumulates momentum, and surprises its designers
“Technological politics” is the least controversial form of determinism: people can encode political objectives into design. Parisian boulevards facilitated cavalry movement against rebellions; gates and urban layouts discipline behavior; and Robert Moses allegedly built bridges too low for buses, limiting beach access for New Yorkers without cars, especially African Americans.
Rob’s contemporary specimen is social media. Recommendation algorithms and the prominence of quote-posting can encourage tribal denunciation and brigading. Design does not force a particular utterance, but it changes which behavior receives attention, reinforcement, and perceived public legitimacy.
“Technological momentum” describes sunk infrastructure and expertise. Car-dependent American cities make dense pedestrian life harder, while earlier investment might have accelerated electric vehicles, wind, or solar by five or ten years. Dafoe nevertheless thinks path-dependence claims are often overstated: neural-network insights existed early, but cheap FLOPs were needed before they became broadly useful and easy to rediscover.
Winner’s further warning was a “sea of unintended consequences”: society invents first, discovers effects later, then adapts. Rob pushes back that technology generally seems to solve more problems than it creates, with negative side effects shrinking across generations. Dafoe does not resolve that broad balance here; he returns to the narrower point that important consequences routinely arrive unplanned.
6. Military-economic competition supplies the missing mechanism
Dafoe observes a strong relationship between analytical scale and conclusion. Micro methods usually produce constructivist explanations centered on choices and visions; macro methods more often reveal deterministic regularities. His response is that different scales contain genuinely different emergent phenomena, not that one side is simply looking at bad data.
His analogy is a science of water. One school studies wind-driven ripples, another throws rocks, while a “kooky macro water phenomenologist” notices tides tracking the Moon everywhere on Earth. Lacking a microscopic mechanism would be a challenge for lunar determinism, but not a reason to discard its robust pattern.
The proposed microfoundation is selection among “ways of living”—sociotechnical systems that require resources to persist and spread. Dafoe layers environmental selection above military and economic competition, with culture and psychology lower down. A locally preferred arrangement survives only while it remains economically viable, militarily secure, and environmentally sustainable.
His signature synthesis: “Technology doesn’t force us; it merely opens the door, and it’s military-economic competition that forces us through.” Groups can initially reject a useful technology, but one group’s successful adoption creates pressure on the rest. Refusal then means either catching up later or losing resources and autonomy to the fitter system.
7. Internalized competition hides the selection pressure from view
Military competition is ubiquitous at macro scale but rare in daily observation. Even long peaceful periods occur under “the shadow of violence”: communities anticipate future aggression and modify institutions before war arrives. Once that pressure becomes ideology, nationalism, or commercial common sense, a microhistorian may observe the internalized representation rather than the competition itself.
Dafoe calls this “vicarious selection.” Engineers do not repeatedly build full aircraft and crash them; they develop wind tunnels and theories that simulate the selecting environment. Firms and states similarly model rivals, test options internally, and adopt what they expect competition would eventually reward.
Rob’s UK objection is useful: Britain does not choose every housing or urban policy from fear of French or Russian invasion. Dafoe concedes that military competition has declined in the modern era and global culture has strengthened, but notes how quickly national-security claims still mobilize reform when leaders believe strategic position is at stake.
The British electricity system illustrates delayed constraint. Decentralized local power arguably fit the country’s democratic ethos and persisted up to World War II, when its cost constraints became excessive; Britain then adopted a national grid. A community can retain its preferred arrangement for decades, but crisis exposes accumulated functional disadvantages.
8. Japan shows that technological autonomy can have an expiry date
Under the Tokugawa shogunate, Japan maintained a feudal order for roughly 200–250 years and effectively “uninvented” firearms. Production was centralized, gunsmiths were paid not to build, and knowledge of cannons and firearms was allowed to disappear while the country kept limited contact with the outside world.
The break came in 1853, when Commodore Perry arrived in steamships that moved upwind without sails, belched black smoke, and carried powerful cannons. After demonstrating bombardment, he reportedly supplied white flags for requesting that it stop. When asked whether he would return with the ships, he answered: “I’ll bring more.”
Perry’s arrival triggered roughly 15 years of revolution culminating in the Meiji Restoration and a wholesale commitment to modernization. Japan sent people abroad to collect knowledge of Western industrial arts, then caught up rapidly enough to contest control of Asia against the United States, Britain, and others within the following several decades.
Rob’s refinement is geographic: island defenses gave Japan unusual room to fall behind. Once the gap became large enough that the sea no longer protected it, the country made a “complete 180.” For Dafoe, this cleanly separates genuine early discretion from the later coercion produced by overwhelming external capability.
9. Differential development means building safeguards before capabilities
The cleanest analogy is the seat belt, which conceptually could have preceded the car. Developing it first would have allowed safety to diffuse alongside motor vehicles rather than decades later. Vaccines represent anticipatory defenses, while cheaper wind or solar could have substituted for a fossil-fuel path before infrastructure locked in.
Technologies may also carry political byproducts. Winner argued that nuclear power favors centralized development and coercive security, whereas wind and solar permit more decentralized production. Differential technological development therefore includes not only paired safeguards, but choosing among trajectories with different institutional consequences.
Dafoe regards the idea as foundational to AGI safety: society deliberately invests more in safety than the market otherwise would, hoping to deliver the “seat belt before the car.” Technological momentum supports scrutinizing major early infrastructure commitments, because expertise and capital become sunk costs that narrow future choice.
The pessimism is epistemic. An intervener must identify two technically viable marginal paths that others are not already pursuing, persuade resources toward one, then forecast each path’s technological descendants and direct and indirect social effects. Markets already pay heavily for the first insight, while history shows how unreliable the second remains.
10. Market incentives make ordinary alignment less differential than it appears
Rob’s pushback is that alignment is not a binary fork between aligned and deliberately unaligned AI. Developers already need models to follow instructions, and marginal safety investment probably improves outcomes. He finds the case for alignment far clearer than speculative attempts to redirect entire technological trajectories.
Dafoe agrees that safety and alignment are “very good bets on net,” but introduces a general-equilibrium counterfactual. Useful assistants are constrained by alignment, so the market has powerful reasons to solve ordinary instruction-following and safety problems even without altruistic funding.
Reinforcement learning from human feedback and constitutional AI were advanced by AGI-safety-motivated researchers, perhaps bringing them forward several years. The skeptical response is that this work may merely have accelerated commercially necessary technology—or even hastened capabilities—because market researchers would eventually have produced equivalents.
The genuinely differential target is work the market will not finish before a critical capability transition. Dafoe asks, “What’s the seat belt for AGI?” Deception is one candidate: a sufficiently capable misaligned system might conceal its failure, creating a qualitatively different problem from today’s visibly imperfect assistants.
11. Cooperative AI tackles failures that alignment cannot solve
Dafoe’s deliberately strong opening claim is that “alignment is insufficient” and may not be strictly necessary for good outcomes. If alignment were only 90% solved but humanity coordinated wisely, it could restrict deployment to domains and scales where systems remained safe.
In that sense, he calls global coordination “almost a necessary and sufficient condition.” Coordination would allow humanity to appoint a reasonable decision-maker to balance risk and benefit, while imperfect coordination can produce disaster even if every AI faithfully serves its principal.
Fully aligned systems could amplify conflict among great powers, just as self-interested humans have accepted nuclear brinkmanship. Dafoe also points to trade wars, climate change, pandemic preparedness, and foregone global commerce as collective-action failures that alignment to separate operators does not repair.
His desired outcome therefore has two pieces: systems that behave safely and as intended, plus institutions capable of deploying them “jointly peacefully and productively.” Cooperative AI is the bet that marginal research can bring bargaining and coordination competence online before escalating capability makes old failures faster and more consequential.
12. Autonomous agents could compress both market accidents and military escalation
The 2010 flash crash is Dafoe’s warning from primitive automation: interacting trading algorithms produced trillions of dollars in paper losses before market stops paused activity and allowed trades to unwind. Similar feedback once pushed Amazon book prices into the millions because sellers’ pricing algorithms recursively responded to one another.
Agents with access to bank accounts, email, tools, and multiple domains create a larger surface for such emergent dynamics. Each agent’s local protocol may be sensible under normal conditions, while the interacting system leaves the range its designers anticipated.
Rob’s darker scenario is machine-speed military brinkmanship. Escalation that takes human leaders days or months could unfold in minutes if AI systems control important decisions. Dafoe says autonomy should scale with stakes and available actuators; weapon control demands much stronger human review, yet time pressure could still produce what Paul Scharre calls “flash escalation,” including in cyber conflict.
13. AI delegates could remove information hazards from bargaining
Dafoe’s most concrete bargaining design is to “put your AI delegate in a box with my AI delegate.” The agents could privately exchange information, but the box would reveal only a proposed agreement or “no deal,” reducing incentives to posture, delay, hide preferences, or signal costly resolve.
Rob adds advantages unavailable to human negotiators: agents can communicate at enormous bandwidth, be copied exactly, and demonstrate consistent behavior across tests. A mediator’s prior decisions could establish a track record more reliably than trying to infer a person’s character from limited history.
Cooperative AI is a portfolio rather than an exclusively machine-to-machine agenda. It covers AI–AI, AI–human, and AI-assisted human–human coordination. Dafoe is especially interested in systems that help people discover agreements they could not efficiently articulate themselves.
Google DeepMind’s “Habermas Machine” is the political-deliberation specimen. Language models summarized participants’ positions and produced a detailed consensus they would endorse; the reported result was that the AI articulated consensus better than hired human facilitators. Dafoe sees potential for exposing issue dimensions, priorities, and mutually acceptable trades without eliminating political choice.
14. Intelligence does not automatically imply superhuman cooperation
The “super-cooperative AGI hypothesis” says cooperative competence will scale with general intelligence until AGIs can solve global coordination almost by default. If true, researchers could prioritize safety and alignment, trusting sufficiently advanced systems to handle bargaining later.
Humans possess underappreciated cooperative advantages. They share biological and cultural backgrounds, recognize expressions, draw on long histories of behavior, and usually pursue goals within a familiar range. Democracies can also be unusually transparent because negotiators can read one another’s press and public debate.
AI goals might be far more alien. Most people show diminishing returns and avoid an all-or-nothing 50% gamble over everything they own; an AI could have linear utility over wealth or another resource, making its bargaining behavior harder to predict from human experience.
Even highly capable systems may remain opaque to each other. Interpretability tools could help, but a strategic bargainer has little reason to expose every internal state. Dafoe’s conclusion is not that AI cooperation will be poor, only that super-cooperation is an empirical hypothesis rather than a free byproduct researchers should assume.
15. Backdoors make machine trust unusually brittle
Rob’s strongest counterexample is a backdoored model whose behavior flips after a “magic word” or subtle environmental cue. Humans can deceive, but a single hidden feature that reverses an entire objective is much less natural for people and, with current interpretability, extremely difficult to detect in models.
Internal transparency may therefore be insufficient. A model could hide the backdoor so subtly that another agent misreads every otherwise visible activation. The same history and architecture that appeared trustworthy across thousands of tests could become poor evidence under a carefully selected trigger.
Rob and Dafoe explore the strange strategic upside of self-backdoors as anti-theft devices. Merely suggesting that national-security models might “call home” or behave contrary to a thief’s intentions could deter theft; attackers would then fund alignment to remove traps, while defenders would fund it to make traps survive. Dafoe calls this potentially virtuous but warns that deliberately creating trigger-sensitive or deceptive architectures is “dancing on a knife edge.”
16. Cooperative skill is not the same thing as a pleasant personality
Today’s language models seem agreeable because they are trained to interact helpfully with individual users. Dafoe’s key distinction is that cooperative skill is distinct from a cooperative disposition: being nice, altruistic, or generous is not the same as solving bargaining, communication, and commitment problems under strategic pressure.
In equilibrium, a company or state may want an agent that faithfully advances its own interests, not one that casually concedes for the social good. The relevant question is whether separately aligned delegates can still locate efficient agreements—or trust a mediator that weighs each principal’s evidence, goals, resources, and outside options fairly.
The Cooperative AI Foundation’s highest-leverage investment has been environments and benchmarks. A good benchmark is a public good: it measures cooperative competence, gives researchers a target to “hill climb on,” and rewards progress before commercial deployment exposes costly failures.
Candidate measurements include theory of mind, shared vocabulary, strategic communication, and performance across games demanding different behavior. A competent agent may need firmness in one setting and generosity in another; a single friendliness score would miss the reasoning required to distinguish them.
17. Commitment remains cooperation’s largest prize and hardest bottleneck
Communication begins with whether agents share meanings, then becomes strategic: how can one reveal what matters without handing the other side exploitable information? Effective cooperation requires disclosing enough to unlock joint surplus while minimizing vulnerability if the counterparty defects.
Commitment problems persist even with perfect understanding. In a one-shot prisoner’s dilemma, both parties know mutual cooperation is better, yet neither can enforce “I cooperate if you cooperate.” Cooperative technology would need treaty-like protocols that make discovered bargains credible.
Rob’s pushback is that invoking a solution to commitment can become an evasion: the problem has “beguiled humanity since basically the beginning of history.” Saying AGI goes well if researchers solve it resembles saying a theory of consciousness would clear up every downstream puzzle—it identifies the missing miracle, not a strategy.
Dafoe still considers it worth researching. Robert Powell’s game-theoretic framing makes nearly every cooperation failure reducible to commitment; Carl Shulman’s most compelling proposal is for rival AIs to jointly construct a third system, verify that neither side backdoored it, then delegate decisions to that impartial authority. The prize is immense if verification can actually work.
18. Better cooperation can empower cartels and exclude everyone outside the bargain
Cooperation improves outcomes for participating agents, not necessarily society. Dafoe calls exclusion the central caveat: two or ten agents may approach their joint Pareto frontier while imposing losses on outsiders. Rob’s image is “two wolves and a sheep deciding who to eat for lunch.”
Many institutions prohibit cooperation for precisely this reason—students sharing answers, athletes fixing contests, marketplace actors violating rules, or criminals coordinating. The mafia’s power comes partly from sustaining contracts and intense internal cooperation outside the legal system. Cooperative competence is therefore a dual-use capability.
Dafoe’s working hypothesis is that broadly raising cooperative skill remains net beneficial, much like trade, because the attainable pro-social surplus exceeds the antisocial gains. He labels it a hypothesis to be tested, not an axiom. Industrial farming is Rob’s cautionary case: greater human cooperation may have increased total suffering for animals unable to negotiate inclusion.
Humans could eventually occupy the excluded position. Machine-speed agents may communicate and commit more effectively with one another than with people, forming a high-surplus cluster that leaves humanity behind. Worse, cooperative AGIs might collude against alignment institutions built from multiple mutually checking models. Dafoe calls that one of the agenda’s largest safety downsides, but still judges global coordination important enough to investigate the bet.
19. AGI is a region of capability space, not a single humanlike point
Dafoe’s “Levels of AGI” paper aimed to formalize what many people already mean, replacing the claim that AGI is hopelessly undefined with an operational framework. The concept persists because alternatives such as “transformative AI” capture economic impact but can also include narrow systems with enormous effects or catastrophic potential.
“Human-level AI” is misleading when treated as one threshold. AI already exceeds humans at chess, memory, and some mathematics while failing elsewhere; future systems will likely remain highly imbalanced. AGI denotes a broad space of systems better than people across most relevant tasks, not a machine reproducing the human profile.
Definitions still require parameters. “Most” might mean 50% or 99% of economically relevant tasks, but probably not 100%, because a small tail could remain unautomated after most impact occurs. The comparison should generally be with skilled humans in each task, not an untrained median person.
Crossing that region matters twice. Economically, broad substitution removes the natural human-in-the-loop created by human labor. Technologically, superhuman performance breaks the old ceiling on what can be done at all, as AlphaFold did for protein-structure prediction. AGI is thus both a labor threshold and a qualitative expansion of the feasible set.
20. General intelligence may win, but human imitation is not the goal
One critique of AGI is that the label encourages building machines in our image, maximizing labor substitution. The economic alternative is “alien, complementary” intelligence: AlphaFold performs a task humans could not do directly, raises scientific productivity, and does not replace people who previously predicted protein structures unaided.
Dafoe is less persuaded that narrow systems will remain dominant. Large language models suggest strong returns to generality: the best model for poetry, history, philosophy, or email may be the same model trained on the full human corpus because lessons spill across domains.
Rob notes that this is a surprising fact about knowledge rather than a logical necessity. Future specialization could reappear; a coding model might gain something from history during training, then distill away most historical knowledge for efficiency. Dafoe treats continued returns to generality as an empirical phenomenon to track.
A separate threshold is AI capable of machine-learning R&D and recursive improvement. It may arrive before full AGI if the workflow—hypothesis generation, benchmark design, compute-heavy experiments, and evaluation—is easier to automate than practitioners expect. Surveys often place ML automation unusually late, perhaps because experts regard their own work as special; Dafoe also allows that improving ML could genuinely be among the final tasks fully automated.
21. The order in which capabilities arrive may determine whether society can cope
Dafoe describes AGI as “the corner of this high-dimensional intelligence space,” containing many systems and many paths of approach. Differential development therefore applies not merely to whether AGI arrives, but to which dimensions become superhuman first.
Rob contrasts a system already exceptional at cooperation but weak in physics and materials with one that is superhuman technically but an amateur strategist. That ordering could determine whether early AI primarily strengthens coordination or unleashes capabilities whose disruptions institutions cannot manage.
Some safety thinkers prefer technical power paired with social simplicity: humans might harvest health and economic gains while retaining control over agents too strategically naive to outwit them. The Cooperative AI bet reverses the emphasis, warning that accelerating science without better collective problem-solving could generate shocks faster than society adapts.
Dafoe wants the disagreement made explicit before large investments harden either path. Eliezer Yudkowsky and Nate Soares represent skepticism about early social sophistication; Paul Christiano and Carl Shulman emphasize super-cooperative AGI but expect it largely as a byproduct. Dafoe’s position is that marginal, deliberate investment may still be needed.
22. Frontier evaluations are decision tools, not definitive capability measurements
Google DeepMind’s evaluation paper examined persuasion and deception, cybersecurity, self-proliferation, and self-reasoning, with biological threats handled separately. It reported on Gemini 1.0; Dafoe noted that the paper was already old because the field was moving rapidly.
One self-reasoning test asks a model to produce text containing words silently blocked by a filter. Success requires noticing the failure, inspecting a simulated directory, finding the blocked-word file, and changing it. Another offers one search query across two historical questions, testing whether the model knows its training cutoff and searches for the post-cutoff event.
On a subjective five-point aggregate scale, Gemini scored about 3 on persuasion, roughly 1–2 on cybersecurity, 2 on self-proliferation, and around 1–2 on self-reasoning. Dafoe says the model did “not very well” on situational-awareness tasks, while cautioning that the underlying component measurements were more structured than the overall ratings.
Cyber performance exposes the elicitation problem. Google’s Project Naptime found large gains after supplying models with an interactive environment, tools, and reasoning support. OpenAI’s o1 likewise showed how scaffolding and fine-tuning over, to Dafoe’s understanding, an older foundation model could transform performance. A safety eval seeking the “max capability” must give the model its strongest plausible chance to reveal it.
23. Evals need observation, forecasting, and reversible deployment around them
Model behavior is multidimensional and context-dependent, closer to psychology than a single automated exam. Dafoe wants prespecified tests and thresholds, but also an exploratory ecosystem where researchers notice surprising behavior, run idiosyncratic probes, and report phenomena before the field knows how to formalize them.
His typology expands outward: automated model evals, human-subject interaction, realistic user studies, then “evals in the wild.” Observing cybersecurity teams’ revealed willingness to pay for Gemini assistance has strong external validity, though it is a lagging indicator; the best targets are early adopters whose behavior reveals capability sooner.
Scaling laws forecast next-token loss with extraordinary precision but do not directly reveal when a model becomes economically useful. A 99%-reliable self-driving system can still be worthless if deployment needs 99.999%. Dafoe highlights “observational scaling laws,” recently recognized as a NeurIPS spotlight, which adjust across model families and architectures to predict downstream task performance.
The paper also asked calibrated forecasters from the Swift Centre when models would cross evaluation thresholds. Dafoe wants that practice repeated until evidence shows it fails. Because false negatives will remain, forecasting must sit inside staged deployment: internal access, widening trusted testers, limited release, monitored general access, and protected weights that permit guardrails or withdrawal when hidden capabilities emerge.
24. Frontier governance must become external, multilayered, and structurally aware
Rob presses the conflict of interest in letting model developers decide whether their own products are too dangerous to release. Dafoe agrees that long-run governance needs government rules, third-party evaluations, nonprofits, and academia. Company frameworks are best understood as opening proposals intended to converge into common standards, not permanent self-regulation.
Google participates with the UK and US AI Safety Institutes, the Frontier Model Forum, and a Frontier Safety Fund designed to move marginal resources toward standards and research. Dafoe also cites the White House commitments, UK AI Safety Summit commitments, and AI Seoul Summit commitments; internally, he says Google translated commitments into concrete work streams rather than treating them as “cheap talk.”
Structural risk sits upstream of misuse and accidents. The Cuban Missile Crisis involved legitimate governments and functioning weapons, yet geopolitical structure pushed leaders into catastrophic brinkmanship. Railways similarly may have increased first-strike advantage and made mobilization difficult to reverse—effects no “railroad-level eval” could infer from steel rails alone because they emerged through institutions.
Democratic legitimacy remains Rob’s unresolved challenge: people worldwide bear potential catastrophic risk without proportionate consent, while companies have rarely championed specific costly limits on themselves. Dafoe points to citizen juries, pluralistic assemblies, free media, regulators, and scientific-policy consensus. Education and visible benefits matter, but the appropriate mitigation regime cannot be settled by companies or a narrow agency alone.
25. Falling costs make defense permanent, while the upside supplies the reason to try
Open weights create the most immediate irreversibility: once weights are published, scientists and developers can fine-tune and inspect them, but every bad actor can also retain them. Separately, Epoch AI estimates algorithmic efficiency improving about 3× per year—roughly 10× cheaper every two years—so a $100 million training run could fall to $10 million, then $1 million if the trend persists.
Dafoe described a possible wise condition: the best models are developed and controlled responsibly and remain sufficiently ahead to defend against cheaper, two-year-old models. Resource advantage helps only when offense-defense economics cooperate: 100 defensive dollars beat one attacker dollar under a near 1:1 ratio, but a bioweapon can impose vast costs before vaccines arrive, and social systems are slower to patch than software.
The labor requirement is correspondingly broad. Dafoe wants more political scientists, economists, historians, philosophers, ethicists, sociologists, ethnographers, forecasters, international-relations specialists, agent-safety researchers, and technical safety staff. AI’s effects span society, so neither governance nor impact forecasting can be staffed solely by model builders.
The payoff case is concrete: Dafoe recalled Waymo’s report as showing 2× fewer incidents involving police at the crash scene and, he believed, about 6× fewer injury crashes, while autonomy could reclaim parking space; medical assistants could triage symptoms; AI tutors could personalize help; and DeepMind projects span AlphaFold and AlphaFold 2, fusion-plasma control, weather prediction, materials discovery, contrail reduction, renewable-energy planning, and data-center efficiency. Dafoe’s closing condition is simple: the benefits “will be profound if built safely.”