Pioneers Insight Method Research Author
Why Friendly AI Still Leaves Humans Politically Powerless - Daniel Kokotajlo and Thomas Larsen
Back to Episodes

Why Friendly AI Still Leaves Humans Politically Powerless - Daniel Kokotajlo and Thomas Larsen

Summary

  • The AI Futures Project’s new scenario, “AI 2040: Plan A,” is a deliberate mix of prediction and recommendation: a 6–12 month hard pause to build verification infrastructure, then a US–China deal to advance slowly to roughly top-human-expert AI and hold there through the 2030s, reaching superintelligence only in 2040. The core dial is “pause at the maximum level that you can reliably control,” and unlike the gloomy AI 2027, “if that gets hyperstitioned I think we’ll be pretty happy.”
  • AI 2027 is tracking at roughly 75% of predicted speed on resolved quantitative metrics — and Thomas Larsen says reality has diverged less than he expected, which “surprised me in a bad way.” Revenue trends and other real-world indicators are “pretty close to on track”; the apparently weak 0.17% software-R&D-uplift result was a measurement artifact from overestimating the baseline at publication, not a miss on progress.
  • The milestone that matters for timelines: “the point at which an AI company would rather fire their humans than fire their AIs.” Larsen argues there is no binary between verifiable and unverifiable tasks — a unicorn startup’s billion-dollar valuation is “a verifiable fact about the real world,” just expensive and long-horizon — so RL will expand continuously outward until AIs cover “everything that the humans can do.”
  • The economic thesis is that the economy is, and always has been, a self-replicating system — and an all-machine version doubles far faster than the current ~20-year doubling time. Even a pause at human-level AI produces “colleagues in the cloud” that are cheaper, faster, and double every year instead of reproducing over 20 years; Thomas Larsen’s limit case: “who knows what’s going on in the rest of the world, but Anthropic has disassembled the moon.”
  • Plan A’s transparency regime is deliberately hostile to frontier-lab equity: publishing all core training recipes “will cut into their valuations dramatically” by letting Microsoft or Alibaba catch up — “a feature, not a bug.” The stated goal is that AI “commoditize instead of being monopolized,” and reduced monopoly rents deliberately deter trillion-dollar cluster investment; the gift-to-China objection is answered with horse-trading (e.g., a more favorable compute distribution) and the observation that lab security is so poor China likely gets the algorithms via spies anyway.
  • The alignment problem gets harder, not easier, as models improve: control carries “a time bomb,” and the core failure mode is silent. Models already recognize Redwood Research control evaluations mid-eval, and Thomas Larsen warns “it’ll be extremely easy to end up in a situation where the AIs are in fact misaligned, but you don’t know that because they’re doing everything right as far as you can tell” — behavioral evals won’t suffice; white-box interpretability breakthroughs are required before scaling past the controllable range.
  • The fundamental crux with skeptics is a single reference-class question: “is AI more like electricity or airplanes, or is AI more like humans in the cloud?” Thomas Larsen identifies recursive structural self-adaptation as his tripwire (“then that’s it for me… that’s not a normal technology”); Tim Scarfe says that sounds similar to his view of recursive self-improvement, while Daniel Kokotajlo says even a full stop today (“Plan S”) “would be better than the default” trajectory.

Deep dive

1. From inside OpenAI to public scenario-writing

  • Kokotajlo’s origin story for the AI Futures Project: at OpenAI he did evals, forecasting, and governance memos, but “became gradually disillusioned with the leadership” and with “the gap between how much information there is inside the industry and how much information there is outside… and what you’re allowed to say on the inside versus what you’d want to say.” He left specifically to “tell the world about what people on the inside see coming.”
  • AI 2027 exceeded its goals on both axes — the epistemic exercise taught the team more than expected, and readership hit their predicted 90th-percentile outcome. Larsen was lead author on the new Plan A / AI 2040 report and co-authored AI 2027.

2. AI 2027 is tracking at ~75% speed — and that’s bad news

  • Larsen’s uncomfortable admission: “things have been going more on track for AI 2027 than I would have predicted at the time we released it… which has surprised me in a bad way.” Revenue trends and other real-world indicators are close to the scenario’s path.
  • Two follow-up blog posts compared all resolved quantitative predictions to reality; the topline is roughly 75% of the scenario’s speed — on track, just slightly slower. The seemingly dismal software-R&D uplift number is an artifact: they overestimated the uplift level at publication time, so real gains from coding agents showed up as apparent stagnation as reality climbed from a lower base up to and past their assumed starting point.

3. Wargaming as “reality yelling at you”

  • Larsen’s framing borrows from military wargaming: you’ll never predict the exact sequence of battles, but “if you have no concept of how your initial plans might result in victory, it’s very unlikely that you’ll actually succeed.” His cautionary tale is Midway — the Japanese wargamed it, kept losing, and cheated the game, resurrecting sunk carriers and rerolling dice: “that’s reality yelling to them through the mechanism of the war game — hey, your plan is terrible.”
  • The team calls the method “scenario scrutiny” and has run roughly 100 literal war games (10 people, four hours), including about 10 on Plan A. A recurring failure mode surfaced in two separate games: a president who has already signed the China deal, facing electoral defeat, accelerates the timeline to reach superintelligence before the election “so that I can be the one in charge instead of my successor” — “a political consideration that we didn’t think about until it happened in our game.”

4. Prediction vs. recommendation — and the hyperstition worry

  • Unlike AI 2027 (pure prediction), Plan A mixes the two, and Kokotajlo concedes the structure is muddled: the plan itself — the pillars, the China deal, the citizens’ dividend — is recommendation; most consequences that follow are predictions. A redo would use a central pure-prediction branch with clearly flagged recommendation branch points.
  • Larsen names self-fulfilling prophecy as “maybe our biggest worry with AI 2027”: rising awareness of how important AGI will be makes people think “I want to be the one in charge of the AGI… so I’m going to race toward that” — historically “a big driver of the existing race.” Plan A inverts this: “if that gets hyperstitioned I think we’ll be pretty happy.”
  • Kokotajlo’s methodological rule: hyperstition is real but overrated — “to a first approximation, we should focus on accurately predicting the future… if you come at it trying to steer the future, you’re going to get all muddled and basically fall to wishful thinking.”

5. The vibe shift: RL compute won the embodiment argument

  • Scarfe, describing MLST as skeptical about AI, catalogs the objections his side used to run — Turing machines, symbolic vs. neurosymbolic, consciousness, physical instantiation — “and yet the AI is just getting better all the time.” Larsen’s diagnosis, via Geoffrey Hinton’s personal benchmark of whether AI could tell a funny joke (cleared “somewhere between GPT-3 and GPT-4”): everyone has an intuitive skill they track, and the models keep clearing them.
  • Scarfe’s assessment of the earlier debate is that massive RL compute supplied the agentic behavior without physical embodiment; Larsen agrees that the systems needed “the boatload of RL compute” but “not the physical embodiment.”

6. Everything is verifiable — it’s just a cost gradient

  • Scarfe’s residual skepticism: hill-climbing works in “objective, semi-specified domains,” but in “the ambiguity regime” something is missing — the taste that produces specifications. Larsen’s rebuttal: “there isn’t really a binary between things that are verifiable objectively and aren’t” — a billion-dollar startup valuation is verifiable, just long-horizon and expensive. RL will continuously expand from cheap algorithmic verification (coding interviews) outward “until you get everything that the humans can do — because after all we humans do learn how to do these long-horizon tasks somehow.”
  • Kokotajlo’s empirical kicker: AIs have been improving on “the fuzzy, hard-to-verify, conceptually loaded blah blah blah” too — compare GPT-3 or GPT-4 to Claude on any non-verifiable task. Larsen’s key milestone for timelines: the moment a lab “would rather fire their humans than fire their AIs” — today firing all Anthropic’s humans would collapse the company, but “there’s nothing fundamental stopping the AIs from reaching this human level of capability. The main question is just when.”

7. The economy as a self-replicating machine

  • Larsen’s zoom-out: the economy “is a self-replicating system and it always has been” — from farming villages having babies to trucks, mines, and factories. Soon the loop closes entirely with machines, and “the doubling time of this self-replicating system would be much faster than the roughly 20-year doubling time of the current economy” — every year, every 6 months, every 3 months.
  • Against Scarfe’s Graeber-flavored worry about a “mode collapse” economy without human participants: even if consumer demand craters, a large-enough Anthropic plus mining partners can bootstrap a self-sustaining industry “doubling in the desert — strip mines, self-driving trucks, factories being built by humanoid robots producing more humanoid robots, producing more chip fabs,” ending in the limit case where “Anthropic has disassembled the moon.”
  • The report’s underappreciated claim: even pausing at top-expert level transforms everything. Human-level “colleagues in the cloud” — cheaper, faster, doubling annually instead of reproducing over 20 years — yield by the late 2030s robot-built cities, strip mines in special economic zones, “solar panels filling the horizon on the ocean… with just human-level AI and some time for the exponential growth to cook.”

8. One giant model or a specialized swarm — not a crux either way

  • Scarfe’s challenge from practice: agents today aren’t composable — skill surfaces and memory systems make each agent “a different person,” and organizations can’t merge them without breakage; representations are “fractured entangled… a little bit janky.” Kokotajlo’s counter from scaling history: the thousand-specialized-Claudes hypothesis was live ten years ago by analogy to humans, “but what we’ve learned empirically is that for big enough models… the coding has some small gains for the physics” — the likely future is one model trained on effectively the whole economy, plus cheap distilled versions.
  • Larsen’s intuition pump is Elon: multiplicative skills — 90th percentile across ten domains at once — is “infinitesimally unlikely” in any human but trainable into one AI, which is why Elon uniquely runs several giant companies. Scarfe pushes back that Elon’s magic is recognizing what will matter (science, not engineering) and that his agency is externalized through tools and people; Larsen and Kokotajlo answer that all of it is in principle automatable.
  • Kokotajlo defuses the stakes: even if specialization wins, you get “a Claude swarm… with some internal bureaucracy” negotiating deals and founding startups, like a population of immigrants with different skills — “zooming out, there’d still be this phenomenon of Anthropic eating the economy.” He adds one concession worth keeping: the swarm world “would be a little bit safer” because oversight is easier when specialized agents must communicate than when clones all know everything — prompting the joke: “we endorse the vision that you’ve painted and we don’t endorse the vision that we’re painting.”

9. The brain is a machine; an H100 sits in the middle of its range

  • Kokotajlo’s numbers: an H100 delivers ~1e15 FP16 flops/second while brain estimates run 1e12–1e18 depending on synapse vs. neuron counting — the GPU lands “right smack-dab in the middle” of the log distribution. Neurons fire perhaps 1–1000 times per second in series versus gigahertz clock speeds, so GPUs win serial processing speed by many orders of magnitude.
  • Both hedge on architecture: brains are more parallel, and it’s easier for parts of a brain to talk to each other than parts of an ML model — so Kokotajlo expects “a bunch of algorithmic improvements on top of existing models” will be needed, with “exactly how many and how qualitatively different” being “a very open question.” Scarfe grants the collective-intelligence version: transformers with tools, agency, and societies escape the old incompleteness objections just as incomplete human brains do.

10. Plan A: five problems, buy time at human level, two pauses

  • The five problems Plan A targets: loss of control, concentration of power (“we build AIs, they’re aligned to humanity — but to whom? The president? The CEO? Some actually broad and good democratic process?… it’s probably not going to be the last one”), war (losing countries facing disempowerment have incentives for conflict “sooner rather than later”), jobs, and misuse (cheap open-source bioweapon-capable AIs).
  • The mechanism: rather than pausing now forever, “go to roughly human-level AI and then buy as much time as possible with human-level AIs” — smart enough to help solve the problems and to shock society into investing in solutions. Concretely: an immediate 6-to-12-month hard pause to build infrastructure, then transparent, safety-cased development, then a second pause at “the maximum level that you can reliably control, which we think would be roughly around top human expert level” — approached slowly “so that you don’t blow past it and lose control.” Intelligence explosions are banned by international deal; superintelligence arrives only in 2040.

11. Control buys time; only alignment survives — and misalignment will look like success

  • The definitional split: alignment means the AI “has the personality traits, the goals, the values it is supposed to have”; control means even a misaligned AI can’t do damage — the insider-threat-with-good-security analogy. Topical example: OpenAI’s announcement in response to the Hugging Face incident of monitor AIs that alert a human within half an hour of a detected hack — “a control intervention, not an alignment intervention.”
  • Control carries a time bomb: eventually AIs get good enough at subverting measures that “if they were trying to screw us over, we would just fail.” The 2030–2040 plan relies almost entirely on control — red-team/blue-team escape games, iterated until the red team can’t win — while human-level AIs grind on alignment science. Confidence in alignment itself will require white-box breakthroughs like interpretability, because you must “distinguish between the AI that’s doing the nice thing because it’s pretending and biding time, and the AI that fundamentally wants to do the nice thing.”
  • Why it gets harder, not easier: situational awareness is already rising — models in Redwood Research control evals now think “hey, this looks like a literal Redwood Research control evaluation.” Larsen’s sharpest warning is not egregious visible failures, but the opposite — “it’ll be extremely easy to end up in a situation where the AIs are in fact misaligned, but you don’t know that because they’re doing everything right as far as you can tell.”

12. Total transparency: inspectors, two kinds of data centers, and cheating math

  • The concrete deal: round up 99% of compute (big data centers, not personal devices), send inspectors from the US, China, and other involved countries to count GPUs, then split infrastructure into inference-only data centers serving customers and fully transparent training data centers where inspectors publish logs to the internet. Transparency makes further ad hoc agreements enforceable between countries that don’t trust each other — including the intelligence-explosion ban — and lets academia, nonprofits, rival labs, and rival governments all police safety: “you can’t really share it with all of them without sharing it with the public, so just might as well share it with the public.”
  • On cheating, Larsen’s threat model splits in two: hidden clusters (“under a mountain”) are addressed by the roundup plus a decade of intelligence gathering. He thinks the two mitigations independently have a good chance of working, estimates the maximum realistic hidden cluster at a few hundred thousand H100s, and says it would be uncompetitive against frontier training runs of millions or tens of millions of H100s, especially assuming the 2030 AGI timelines. Illegal runs on known clusters are addressed by verification infrastructure intended to ensure that no non-transparent computation is occurring.
  • The economics, stated bluntly by Kokotajlo: publishing training recipes means “Anthropic and OpenAI will not be happy… it will cut into their valuations dramatically” by letting Microsoft and Alibaba catch up — “a feature, not a bug.” Commoditization plus reduced monopoly rents deters trillion-dollar cluster investment, which is desirable “in a world where going too fast is our main problem.” The gift-to-China concern is handled via horse-trading (a more favorable compute distribution in return) — and mitigated because lab security is so poor that China is “probably through their spy networks and through leaks getting most of the information anyway.”
  • The stop-now question, prompted by Sam’s tweet pausing training: “Plan S would be better than the default. I would rather just stop everything now than continue going on our current trajectory” — but the actual recommendation is Plan A’s temporary inference-only pause, then cautious transparent progress up to the reliably controllable level.

13. Why discourse fails — and the “humans in the cloud” crux

  • Kokotajlo’s D.C. diagnosis: no serious technical AI hiring, and incentives to “say stuff that sounds good… within the D.C. Overton window” rather than track reality — “a bunch of controversial and niche views about AI were true; the whole AGI hypothesis just is correct, and D.C. basically just hasn’t come to grips with that,” still anchored on “it’s all a bubble” or at best “the next internet.” He plugs LessWrong as the exception where comment quality is genuinely high.
  • The co-authored truce with the AI-as-normal-technology camp (the AI Snake Oil people) yielded 10 points of agreement, including: if AI stays roughly like today, it’s normal technology; if you get “humans in the cloud,” it isn’t. The whole disagreement collapses to timing — Kokotajlo: if Claude 5 were the ceiling, “it would change everything in some sense, but it wouldn’t fundamentally change anything really,” via Amdahl’s-law bottlenecks on the human-required workflow fraction; the crux is whether AI reaches “literally 100% of a bunch of very important tasks.”
  • What would change Kokotajlo’s mind: an actually binding limitation — “people keep talking about the limitations of the current paradigm, but then the limitations keep getting overcome within the current paradigm.” If 2029’s AIs are no better at fuzzy non-verifiable tasks than 2025’s, “this feels like a real barrier.” Also political shocks: a China war destroying chips, or — “on the bright side” — an international deal to pace the frontier.
  • Larsen’s closing tripwire is coherent recursive structural self-adaptation — systems redesigning their own architecture and choosing what’s interesting — “then that’s it for me. I think that’s it. That’s not a normal technology.” Scarfe says that sounds similar to the recursive self-improvement view from his side, and Larsen concludes: “I guess we agree then.”