AGI is still 30 years away — Ege Erdil & Tamay Besiroglu
AGI is still 30 years away — Ege Erdil & Tamay Besiroglu
Summary
- Headline call: full automation of remote work lands around 2045, not 2027. Tamay Besiroglu gives 2045; Ege Erdil says he is more bullish, while Dwarkesh suggests shaving Ege’s prediction by five years or 20%. The logic: AI has unlocked roughly one major capability (gameplay, language, reasoning) per ~3 orders of magnitude of training compute, we’ve burned through 9-10 OOMs since AlexNet, and energy/fab constraints leave “maybe three or four orders of magnitude of scaling left” before AI infrastructure becomes a non-trivial share of world output.
- The intelligence explosion is a category error — “like calling the Industrial Revolution a horsepower explosion.” Their thesis: transformative growth comes from broad, simultaneous upgrading of supply chains, capital, data, and deployment, not from geniuses in a data center; “it isn’t the case that the world today is totally bottlenecked by not having enough good reasoning.”
- Software-only singularity is unlikely because compute and cognitive effort are complements: software progress has historically tracked hardware (~30%/yr, matching Moore’s law), the big innovations (transformer, flash attention, chinchilla) were all about harnessing compute, and GPU-rich labs — not academia — produce the algorithms. Dwarkesh’s counter-evidence (Kokotajlo’s survey showing 1/30th compute only costs 2/3 of progress) gets the reply that the imbalanced-scaling experiment “was never run.”
- Falsifiers they’d accept: an AI that downloads a post-training-cutoff Steam game and beats it cold; OpenAI hitting $500B revenue (a mere $100B gets “maybe a 40 percent chance” and is not a huge update — “people pay trillions of dollars for oil”). Near-term marker: agents book flights by end of this year, which proves little since “nobody gets a job where they’re paid to book flights.”
- They still forecast ~30% explosive growth, quantitatively: an H100 does ~1e15 flop/s (roughly a human brain), costs ~$30K, and running “the software of the human brain” at $50-100K wages pays itself back in about a year — an economy doubling time near one year. Growth differentials will be “determined by regulatory jurisdiction boundaries more than anything else”; Dwarkesh floats a 10-20% chance that global coordination throttles it, which Ege says is not unreasonable.
- Takeover risk is overrated, lock-in doubly so: a dominant AI economy taking over humans is like asking “why doesn’t the US just invade Guatemala?” — war is inefficient when you already earn most of the income. Dwarkesh’s East India Company counter (“they just took over, right?”) challenges this, but the deeper claim is that value lock-in “looks very unlike anything that has happened in the past”: slavery ended from economic incentives, not abolitionist virtue.
- Mechanize’s case for accelerating: each year of delay might cost tens of trillions in consumption and perhaps 100-200M lives, valued at up to ~$10M per statistical life — and pausing may not buy safety, because “imagine you were trying to make progress on alignment in 2016 with the compute budgets of 2016. You would have gotten nowhere, basically.”
- The underrated unlock is the AI firm, not the AI genius: copyable workers give firms the missing third leg of evolution — high-fidelity replication — plus alignable preferences and a “hyper inference scale mega-Jensen” reviewing every pull request. Central planning arguments get stronger, though not decisive: Apple could run the economy of ancient Uruk, “but Apple as it exists today cannot manage the world economy as it exists today.”
Deep dive
1. “Intelligence explosion” is a horsepower explosion — wrong lens for the transition
- Tamay’s opening frame: the concept “is not a very useful” one — “it’s kind of like calling the Industrial Revolution a horsepower explosion.” Yes, raw physical power exploded, but the Revolution was complementary change across agriculture, transportation, law, finance, urbanization. Similarly, “we’ll get a lot of very smart AI systems, but that will be one part among very many different moving parts.”
- The load-bearing disagreement with short-timeline SF: intelligence isn’t the binding constraint on growth. “It isn’t the case that the world today is totally bottlenecked by not having enough good reasoning.” Acceleration requires complementary innovations, upgraded supply chains, demand for new products, and the whole economy scaling — not “very, very, very good reasoning tokens.”
- Both treat ASI as an unhelpful operative concept: coherent, but “I prefer just thinking about what actually happens in the world” — AI capability profiles are jagged, and you could get drastic acceleration without a system better than humans at everything, or an ASI that’s “very expensive or very slow” and changes nothing.
2. The timeline math: one capability unlock per ~3 OOMs, and only 3-4 OOMs left
- The headline: a drop-in remote worker that can do “literally everything” remote — Tamay: “maybe for me, that would be around 2045.” Ege is more bullish; Dwarkesh suggests shaving Ege’s prediction by 5 years or by 20%.
- The reasoning chain: since AlexNet, AI went through 9-10 orders of magnitude of compute and unlocked a handful of core competencies — gameplay (2015-20), sophisticated language, then abstract reasoning/coding/math — roughly one unlock per three OOMs. Remaining gaps (long-horizon coherence, agency, full multimodal understanding) imply more unlocks needed, but extrapolations of energy, GPU production, and fab constraints suggest “maybe three or four orders of magnitude of scaling left” before you’re spending a non-trivial fraction of world output on data centers.
- Against naive trend-reading in either direction: bulls extrapolate trend lines (“obviously it’s going to happen in 2027”), while Robin Hanson’s approach — extrapolate the tiny automated share of the economy — says “it’s going to take centuries.” They reject both.
- A worldview split Ege makes explicit: bulls think jobs are simpler than they are. “A lot of people look at jobs in the economy and they’re like, ‘oh, that person, their job is to just do X.’ But then that’s not true” — booking flights is a task fragment, and “automating that actually wouldn’t automate their job.”
3. The unhobbling debate: reasoning only looks easy in hindsight
- Dwarkesh channels Leopold’s view — these are “baby AGIs” held back artificially, and ChatGPT proved 1% of extra compute in post-training unlocks a whole capability; why not agency next? Ege’s counter: “you could have made similar points five years ago” about AlphaZero — people tried to extend it to math “and it didn’t work very well.”
- One guest’s deeper rebuttal — worth keeping whole: reasoning looks cheap “standing on this huge stack of technology” — internet-scale data, distillation, inference efficiency — but “if you ask someone to build a reasoning model in 2015, it would have looked insurmountable.” The prediction is that agency will likewise look simple in three-to-five years’ retrospect, after years of complementary innovation nobody’s counting today.
- On the METR doubling curve (task length doubling every seven months), the objection is construct validity: eval tasks must be “compact and closed” with clear metrics — “those are not problems that you actually work on in AI R&D. They’re very artificial problems.” A human good at them is probably a good researcher; an AI good at them “lacks so many other competences that a human would have — not just the researcher, just an ordinary human.”
4. Moravec’s paradox and the missing creativity: knowledge without recombination
- The framework: AI races ahead on what impresses humans — chess, Go, competitive math, multiplying 100-digit numbers (“the one that got solved first”) — precisely because these are evolutionarily recent skills evolution barely optimized. In humans these correlate with general competence; in AIs the correlation breaks: o3 mini high tops competitive programming “but it isn’t the best at actually helping you write code” — enterprise coding revenue “is just Claude.”
- The best specimen of the agency gap: Claude Plays Pokemon. The model could tell you exactly what to do if you’re stuck in Mount Moon — “but that doesn’t stop it from getting stuck in Mount Moon for 48 hours.” Explicit knowledge, unconnected to action.
- One guest’s sharpest empirical jab: reasoning models know “in literal terms more than any human does,” yet — “has a reasoning model ever come up with a math concept that even seems slightly interesting to a human mathematician? And I’ve never seen that.” Not even Sunday-magazine-scale novelty, despite enormous scope for cross-field recombination. Dwarkesh’s pushback — they’ve existed six months and aren’t trained to find connections — draws the reply that the sheer knowledge base makes zero output “actually quite remarkable.”
- Update conditions, stated precisely: an agent that downloads an arbitrary post-cutoff Steam game with no tutorials and finishes it; or OpenAI at $500B revenue — “$100 billion, that seems pretty likely to me… maybe a 40 percent chance,” and not a huge update, because “people pay trillions of dollars for oil. The fact that people pay a lot of money for something doesn’t mean it’s going to transform the world economy.”
5. Against the software-only singularity: compute and cognition are complements
- The two-part thesis: “Research is harder than people think, and depends a lot on compute scale.” Evidence for the second half: traditional software improves ~30%/yr, “basically matching Moore’s law”; AI’s algorithmic acceleration coincided with the deep-learning compute ramp; innovation concentrates in GPU-rich labs, not GPU-poor academia; and the era’s big wins — transformer, flash attention, chinchilla — “were just about how to harness your compute more effectively.”
- Dwarkesh’s counter-battery: OpenAI beat a compute-richer DeepMind; Daniel Kokotajlo’s survey found researchers with 1/30th the compute would still make 1/3 the progress; and a “cracked” researcher told him current models save “four to eight hours a week” on autocomplete-adjacent work but “24 to 36 hours a week” in unfamiliar domains — “just draw that forward.”
- The reply on the survey and anecdotes: complementarity doesn’t mean the most-compute lab wins (either factor can bottleneck; dysfunctional culture wastes compute), and the decisive experiment has never been run — “you don’t get to see at an industry level what would have happened if the entire industry had 30 times less compute,” and individuals benefit from spillovers. Dwarkesh’s call to labs: publish the multi-team pre-training resource experiments. The response is that even those won’t settle it — you need “very imbalanced scale-ups,” which are inefficient and rare.
- An honest conditional from Ege: if AGI arrives by 2027, a software-only singularity is “quite high” probability — “you’re conditioning on compute not being very large, so it must be that you get a bunch of software progress.”
6. Explosive growth is broad, not “Shenzhen in the desert”
- Against the enclave model: replicating the semiconductor supply chain — “tons of inputs and materials from probably tons of random places in the world” — is brutally hard, and frontier training free-rides on 30 years of internet data the whole economy produced. Broad deployment is itself the data engine: ChatGPT’s “which response was better?” prompts are “getting user data through this extremely broad deployment… just imagine that thing to continue.”
- On regulation intuitions: people extrapolate from housing, nuclear, supersonic flight — but those technologies wouldn’t double output. AI’s value “is just extremely large, even just for workers,” and 401k- and home-owning households “would do enormously better,” so political support may run deeper than the unemployment-backlash default assumes.
- The geographic prediction: expect heterogeneity across countries, with growth “more delineated by regulatory jurisdiction boundaries than anything else” — and, provocatively, the winning norms may not be classical-liberal: “we might adopt the kind of values and norms that get developed in, say, the UAE,” optimized for AI deployment. Hedged explicitly as illustrative, not a strong prediction.
- On China: its boom is half the picture — massive capital accumulation without labor-force scale-up, whereas AI scales both at once. And against “Shenzhen-in-your-backyard” alarm: the relevant scale is dollar output, where the US leads, and “people already think China is a big deal” — no strong update warranted.
7. Technology is capital deepening, not eureka — the beaver-meme view of history
- The through-line: invention-focused history underplays deployment. Solar panels: “nobody 20 years ago had the sketch for the 2025 solar panel” — efficiency came from the interplay of ideas, building, learning curves, and better materials. Edison’s light bulb: the idea is trivial; the filament experiments, then building power plants and lines to homes that had no electricity, were the work. The meme: two beavers at the Hoover Dam — “well, I didn’t build that, but it’s based on an idea of mine.”
- One guest’s favorite causal chain: “we discovered the Big Bang as a result of World War II” — war → radio communications → radio telescopes → cosmic microwave background. Serendipity scales with economic surface area: “more economic activity means we have more exposed surface area to have more discoveries.”
- Dwarkesh’s meta-observation, endorsed: Conquest’s law — “the more you understand about a topic, the more conservative you become about that topic.” He knows AI best of any industry and sees how much went into it; journalists asking “should we get in touch with Geoffrey Hinton?” are “kind of missing the picture.” The Gell-Mann-amnesia implication: extend that respect to every other industry.
- On credit attribution, the O-ring-adjacent point: R&D is necessary but so is producing food — “the point is not that you can get to a Dyson sphere by just scaling labor and capital… you also can’t get to it by just scaling TFP. You need to scale everything at once.” (Output elasticities: labor ~0.6, capital ~0.3, neither sufficient alone.)
8. Takeover: possible, but war is inefficient — and lock-in is a weak argument
- The core analogy: yes, a coordinated AI majority-economy could take over — “but that’s also probably true in our world… Why doesn’t the US just invade Guatemala? Seems like they could easily do it.” When AIs already earn “almost all of the income in the economy,” the marginal value of expropriating humans is small, and undermining the norms they grew up inside has costs. Misalignment alone doesn’t cause conflict — the literature requires misjudged relative strength or sacred non-negotiables on top.
- Dwarkesh’s strongest counters, preserved: the East India Company “could have just kept trading with the Mughals — they just took over, right?”; and peaceful negotiation doesn’t dispose of takeover — the Qing settlement after the Opium Wars “wasn’t better than never having interacted with the British empire in the first place.” Also: AIs are copies with low transaction costs, better coordinators than “why don’t the young people coordinate?” implies. Ege concedes possibility but notes today’s models: “I agree there’s some evidence that they’re good boys. No, there’s more than some evidence” — against Dwarkesh citing the OpenAI chain-of-thought reward-hacking paper.
- On value lock-in: the one-utility-function-forever picture “looks very unlike anything that has happened in the past” — even digital information suffers link rot, and digitization has coincided with faster cultural change. The deeper claim: future values are determined by which values are functional in future techno-economic environments, not by actions taken a thousand years earlier.
- The slavery case, contested at length: Dwarkesh credits British abolitionists (and Christianity) as proof individuals bend the arc; the counter — Russia abolished serfdom in the 1860s under no British pressure, brutal colonial labor grew less prevalent, and Industrial-Revolution wealth let suppressed egalitarian preferences express. “The forces that led to the end of slavery seemed like they were not contingent forces.” Dwarkesh’s riposte: if wealth lets values win, “this makes value alignment all the more important.”
9. Epistemic humility as policy: the WWII bombing miss, and why Mechanize accelerates anyway
- The best specimen of failed foresight: pre-WWII Britain expected hundreds of thousands of casualties from aerial bombardment in the first weeks of war. Reality: “they had less casualties in six years than they expected in three weeks” — a ~2-OOM miss, from boring practicalities (daytime bombing meant getting shot down, night bombing missed, firefighters worked, allied bombing cost $4-5 per $1 of German capital destroyed). Moral: detailed AI wargaming today will “just look fairly silly”; build adaptive capacity and better institutions instead.
- The underlying latent variable of the whole disagreement: “what is the power of reason?” The answer offered is limited. “There’s just this enormous amount of richness and detail in the real world that you just can’t reason about it. You need to see it.” Dwarkesh’s own takeaway after weeks of flip-flopping across guests: “classical liberalism is the way we deal with being this epistemically uncertain” — decentralize, avoid brittle high-volatility moves like nationalizing labs.
- Why acceleration is defensible: they don’t accept that speed clearly trades against safety — “imagine you were trying to make progress on alignment in 2016 with the compute budgets of 2016. You would have gotten nowhere, basically” — scaling drives alignment progress too. And the cost of delay is conditional but concrete: tens of trillions per year in forgone consumption, plus perhaps 100-200M deaths per year of delay at up to $10M per statistical life. To oppose that you need long-termist leverage they think is very low: “be more humble in what you think you can achieve, and just focus on the nearer term.”
- For the suffering-digital-minds worry Dwarkesh presses (“I don’t want galaxies worth of suffering”), the practical counsel is: don’t give up — discount the future for epistemic, not moral, reasons; align present systems to value happiness; and build institutional capacity to intervene if factory-farming-equivalents emerge.
10. Objections priced, and the AI firm as the real unlock
- Tyler Cowen’s Sub-Saharan-Africa point (clean water needs no scarce intelligence, yet we can’t deliver it) is agreed with — it defeats “geniuses in a data center,” not their broad-buildout model. Baumol and O-ring get the quantitative treatment: qualitative bottleneck talk “is not sufficient” — name the sector, its output share, and the elasticity of substitution. Even maximal bottlenecks leave big gains from reallocation: “software engineers become physical workers” at much higher wages. Their affirmative math: an H100 ≈ 1e15 flop/s ≈ a human brain, costs ~$30K, earns $50-100K wages → economy doubling time on the order of a year.
- The objection they respect: regulation — Dwarkesh floats a 10-20% probability that global coordination blocks explosive growth, and the response is that this is not unreasonable — “the world does have a surprising ability to coordinate on just not pursuing certain technologies,” viz. human cloning, though AI is more valuable and more national-security-relevant. On arms races: even a year’s lead may not beat ICBM-era deterrence — swap in 1990-China and war still looks unattractive.
- The AI-firm thesis (their joint blog post): firms today have selection and variation but not high-fidelity replication — copyable AIs complete evolution’s triad. Copy Jeff Dean into any domain; run a “hyper inference scale mega-Jensen” on $100B/yr of inference writing every press release and reviewing every pull request, merged back into himself. Plus alignable preferences (the principal-agent problem “might go away”), economies of scale where 2x GPUs buys smarter-and-more workers, and learning once instead of every human relearning from scratch.
- Central planning, their long-running offline debate, gets stronger but not vindicated: centralized sensing/processing, orders-of-magnitude-bigger planner brains (CEOs today don’t have bigger brains than workers), and Tesla FSD-style HQ-pushed updates all cut for it — but complexity scales too. Ege’s closer: “imagine if Apple was tasked with managing the economy of Uruk. I think it actually could… But Apple as it exists today cannot manage the world economy as it exists today.”