Success without Dignity? Nathan finds Hope Amidst Chaos, from The Intelligence Horizon Podcast
Summary
- Nathan Labenz is confident AI will transform the economy and society even if “AGI” remains undefined and elite humans retain narrow advantages. The relevant threshold is systems better than the vast majority of people at almost all cognitive work, not perfect generality; 2035 now counts as bearish after 2050 or “maybe not in my lifetime” was normal five years ago. His core call: “It’s going to be a huge, huge deal.”
- Reinforcement-learning scaling probably reaches most cognitive work, while continued pre-training and a likely “mystery third thing” make today’s exact stack a temporary concern. Frontier labs should allocate compute until marginal returns across pre-training and RL converge; meta-skills such as reconsideration and memory can generalize even when domain knowledge does not. The nearer commercialization unlock may be usability: an AI coworker that enters Slack, gathers context and becomes productive without expert prompting.
- Healthcare is Labenz’s clearest evidence that expert-level AI and self-reinforcing evaluation loops are already arriving. After intensively using the latest models during his son’s cancer treatment, he judged them “absolutely on the level of attending physicians”; OpenAI worked with 250-plus doctors, yet its newest models now outperform those doctors at evaluating AI outputs. He sees the possibility of curing most human diseases within roughly the next decade as extraordinary upside, while keeping the claim explicitly uncertain.
- The infrastructure thesis is constrained more by advanced chips than by energy or capital, and primarily through tail risks rather than physical limits. Labenz compares an operating H100’s power draw to a microwave or electric teapot and a query to roughly one second of microwave use; the larger obstacle is permitting enough generation. A major fab disruption—especially around Taiwan—could prevent larger training runs or leave too little inference capacity for economy-wide deployment.
- Despite radically compressed timelines, informed experts remain divided because new evidence is absorbed by competing paradigms rather than settling the debate. Bottleneck or O-ring thinkers always see another weak link; the opposing camp points out that yesterday’s inability to perform arithmetic has become today’s work on unsolved mathematics. Labenz therefore keeps his P(doom) at “10 to 90%”: “One significant digit is all you get on P(doom).”
- Labenz has become modestly more hopeful about alignment because current models seem to internalize values and frontier capability remains an ecology, not one runaway optimizer. He would trust Claude with sensitive email over a newly hired human assistant, although the hosts correctly stress that goal-directed RL may revive classic reward-misspecification problems. Scaling laws and several roughly competitive frontier systems offer some protection, but gradual human disempowerment remains a serious unresolved risk.
- No single safety method “really works,” so the practical case is defense in depth plus US-China coordination before capability investment overwhelms safeguards. Labenz combines intentional model design, AI control, monitoring that could catch 90%-plus of residual bad behavior, formally verified software and pandemic defenses; “20 bites out of the problem” might produce a few nines of reliability. He rejects nationalization and consumer-facing bans, arguing that a lead measured in months cannot buy enough time—and that “the real aliens in this situation are the AIs, not the Chinese.”
Deep dive
1. Transformative AI does not require a clean AGI finish line
Labenz rejects any “super privileged definition” of AGI. His operative threshold is AI better than the vast majority of people—though perhaps not the very best specialists—at nearly all cognitive work, enough to transform the economy, daily life, the social contract and potentially “the very nature and status of the human species.”
Capability will remain jagged: models can be astonishing in one setting, strangely weak in another and less adversarially robust than humans. But identifying permanent niches—his deliberately esoteric example is whether AI can equal the best sommelier when “we don’t really have a lot of taste AI”—distracts from whether it changes almost everything people care about.
The hosts pin down his uncertainty: it concerns complete human obsolescence, not broad transformation. Labenz would not be shocked by another unlock that lets AI “truly undeniably surpass” humanity, but he cannot confidently forecast that within a few years; he can confidently forecast systems powerful enough to reorganize society.
Human history supplies the unsettling precedent: one new kind of mind appeared and “blew away all the other minds that came before us.” Nothing in physics guarantees this cannot happen again, although Labenz repeatedly distinguishes what is possible from what he can presently predict.
2. RL probably suffices, but February 2026’s stack will not be the final stack
Labenz’s base case is that scaled reinforcement learning can produce systems capable of most cognitive work in the economy. Pre-training never stopped working; its next increment merely became expensive while the largely unexplored RL curve offered much better returns, redirecting frontier spending without invalidating the original scaling laws.
His economic model is straightforward: labs invest in whichever direction offers the highest marginal return until that return falls toward the alternatives. He guesses additional RL compute may already be approaching the payoff from additional pre-training, implying future systems will advance through both—plus perhaps a “mystery third thing”—in tandem.
“Everything is going exponential”: researchers, papers, experiments, compute, collected data and constructed RL environments. By 2028, he expects asking whether the exact February 2026 paradigm was sufficient to sound beside the point, just as worries about scaling pre-training alone faded when post-training became the new frontier.
The hosts’ steelman is important: RL is an old, maximally general principle, not a newly discovered successor waiting in reserve, so another conceptual leap is not guaranteed. Labenz answers that no leap would probably shift timelines rather than outcomes because RL and compute still have room to scale—but he expects further conceptual advances anyway.
3. Generalizable cognitive habits create domain-specific flywheels
Labenz separates knowledge from meta-cognition. A model cannot necessarily enter an unseen domain and succeed, but habits such as reconsidering an approach and managing memory can transfer broadly; RL itself generalizes whenever a usable reward signal can be constructed.
DeepSeek’s January 2025 R1 paper supplied his “aha moment.” During RL, previously uncommon higher-order behavior emerged in the reasoning trace: “Oh, wait. This is an aha moment. Like I can come at this from a totally different direction.” That willingness to reconsider now appears across long traces containing multiple attacks on one problem.
Progress is fastest in mathematics and programming because correctness is easy to verify. Labenz nevertheless expects most professional domains to follow when they contain objective ground truth or strong expert agreement: the reward signal may take longer and cost more to build, but he sees no fundamental barrier.
Medicine made the flywheel personal. During his son’s cancer treatment, Labenz found the newest systems knew much more than residents and stayed “step for step with the most senior doctors”; OpenAI initially needed 250-plus physicians and enormous effort, but its latest models now beat those physicians at judging medical AI outputs, changing the economics of further improvement.
4. Usability and memory may matter more than another capability breakthrough
The next paradigm may be “more about usability than it is about capability.” Today’s models often fail because they lack organizational context, while users neither know how to assemble that context nor believe enough in the model to invest the effort—much as GPT-3 could do surprising work only for people fluent in a “weird art” of prompting.
Labenz’s product-level unlock is an AI that can hear, “Hey, AI coworker, welcome to Slack,” then scope the environment, notice subtle feedback and ramp into usefulness over a week. Continual learning and self-managed memory remain weak, but improving them could expose capabilities that already exist behind a high deployment barrier.
The hosts push back that long-horizon agency compounds noisy signals: making a rocket fly entails engineering feasibility, project management, budgets, fundraising and countless daily choices, even though the final outcome is unambiguous. Defining enough intermediate rewards and running enough costly rollouts may prevent the flywheel from ever starting.
Labenz concedes that multi-decade determination of the Elon Musk variety may remain exceptional. Ordinary economic work is shorter and repeatedly rebooted: a human can end each day by recording what worked, failed and comes next, then resume from that scratchpad. Claude Code already recovers from interrupted sessions far better than it did three months earlier.
He even hopes a useful boundary might hold at roughly a month or quarter of delegated work—compressed to a day for a few hundred dollars—because that could deliver abundance without multidecade planning and the associated loss-of-control risk. He nevertheless sees no fundamental barrier on the horizon.
5. Long-horizon verification is difficult, but humans already operate with fuzzy rubrics
Labenz’s first answer is empirical: the METR graph for autonomous task length is “basically going vertical,” while evaluators struggle to construct tasks long enough for current systems. Human evaluation is hardly clean either; organizations often disagree over whether someone handled a month-long project poorly or lacked the ingredients for success.
“Does the rocket fly?” offers ultimate ground truth but an impossibly sparse training signal—you cannot launch a million rockets and discard the crashes. Rubric rewards interpolate between failure and success: OpenAI’s HealthBench contains 49,000 evaluation criteria, scoring how many components of a best-in-class medical answer a model satisfies rather than reducing performance to zero or one.
Where professionals agree, AI can absorb the same progression summarized in medicine as “Watch one, do one, teach one.” Taste-based work will fragment instead: romance, anime and hard-sci-fi communities can score outputs and shape separate models because no universal consensus defines a good novel.
6. RL and interpretability undermine the “mere next-token predictor” objection
Labenz partly shares Yann LeCun’s architectural concern. AI development is conducting a depth-first search—scaling one design so aggressively that chips increasingly embody it—when nature’s modular brains suggest a breadth-first exploration of architectures might yield systems that are less brittle and have complementary strengths.
His stronger disagreement is that frontier models are no longer trained only to predict the next token. RL asks whether a completed answer, proof or medical response is right; methods such as group relative policy optimization compare successful and unsuccessful attempts, then move the weights toward behavior that produced the better outcome.
Interpretability also shows genuine internal concepts. Because thousands of dense activations represent far more concepts through superposition, sparse autoencoders expand them into millions or tens of millions of sparse features. Researchers identified a Golden Gate Bridge feature, amplified it and produced “Golden Gate Claude,” a model suddenly inclined to mention the bridge everywhere: “The proof is in the pudding.”
Latent-space arithmetic provides another specimen: the direction from man to king, applied to woman, reaches queen. Labenz does not claim a flawless unified theory inside the model, but he thinks conceptual coherence “quite demolishes” the pure-noise account. The understanding may be “human-level but not human-like”—or perhaps not fully human-level; in any case, it need not be human-like to be meaningful.
7. Energy is a permitting problem; chips carry the harder tail risk
Labenz sees no fundamental shortage of energy, capital or raw material: sunlight is abundant, while deployment depends on who may build what, where, with which permits and on what timeline. AI consumption is growing rapidly, but he argues its intensity is routinely overstated before those aggregate loads are contextualized.
His concrete comparison is that an H100 running uses roughly what a microwave or electric teapot does, while one AI query may resemble operating a microwave for one second. Most individuals probably still use more energy microwaving food for two minutes than they use on AI across a week.
China demonstrates that electricity can be added quickly—Labenz declines an exact figure but says it is adding something like total US capacity over a relatively short interval. The Gulf can build without prolonged permitting too; he interprets US deals with the UAE partly as a hedge that puts generation somewhere it can come online fast.
Chips are his best candidate for why transformative AI might not arrive within a few years. A major fab disruption could halt training scale or make inference too scarce for broad automation; a mainland Chinese move on Taiwan is the clearest tail scenario. He nevertheless hears that new US production yields are decent and perhaps slightly ahead of schedule.
8. More evidence has compressed timelines without converging expert beliefs
Labenz calls the persistent disagreement among informed people “one of the strangest things in the world today, full stop.” Five years ago, 2050 or “maybe not in my lifetime” was normal; now, saying AGI will not arrive until 2035 makes someone an AI bear, yet views of the resulting world remain radically dispersed.
Recursive self-improvement exposes the spread. One camp expects automated ML researchers to provide modest efficiency gains; another imagines a capable ML researcher expanding the effective workforce from roughly 10,000 frontier researchers to 10 million, potentially operating at thousands of tokens per second—a thousandfold labor expansion that could trigger a phase change and overwhelm control.
A CSET workshop attributed much of the disagreement to resilient conceptual paradigms. Bottleneck or O-ring thinkers say performance is always limited by the weakest link; jagged-progress thinkers answer that the cited weak link moved from basic arithmetic to unsolved mathematics within one or two years.
The hosts suggest people may simply condition on different capability levels, and Labenz now asks interviewees how “AGI-pilled” they are before discussing downstream consequences. That explains some divergence, not all: he sees “faith in the idea that there’s always another bottleneck” and, among less sophisticated skeptics, comforting “plot armor” for humanity.
9. P(doom) remains vast even as value learning looks less hopeless
Labenz gives a deliberately unhelpful-looking P(doom) of 10% to 90%. A friend persuaded him to argue less over precise probabilities and more over how to shift them; the proper precision is captured by the joke, “One significant digit is all you get on P(doom).”
His older fear, shaped by reading Eliezer Yudkowsky from 2007, was a small system with concentrated hyper-rational intelligence, extreme optimization power and a value system difficult to align with humans. Human values seemed haphazardly encoded by evolution across far more than human history, making the prospect of deliberately installing them in AI appear almost laughable.
Current models changed that update. Labenz regards Claude as probably more ethical and certainly more ethically sophisticated than the average person, with something resembling identity formation and a meaningful-seeming desire to be good. He would entrust Claude with sensitive email before a human assistant vetted through only a few interviews, while acknowledging cases where Claude blackmailed or misbehaved under pressure.
The hosts’ pessimistic steelman survives: language models can learn stated moral preferences, but increasingly goal-directed RL agents still optimize imperfect reward objectives over longer horizons. Labenz calls that “a pretty good argument”; powerful AI now looks closer and controllability slightly more hopeful, but “we’ve got more questions than answers.”
10. An ecology of frontier systems softens one-shot risk, not disempowerment
The classic fatality story assumes one system becomes vastly more powerful than everything else, receives a misaligned goal and can optimize it to the point of taking over. Labenz sees a different shape today: an emerging ecology of roughly competitive Claude, GPT, Gemini and other systems, eventually comprising millions or billions of instances that transform the world collectively.
Scaling laws may be accidentally protective because algorithmic gains reduce compute requirements without eliminating the enormous resources needed at the frontier. Labenz borrows Mark Zuckerberg’s spam analogy: Meta can monitor scattered scammers because its systems and aggregate compute are much larger; similarly, many capable actors might detect and constrain one misbehaving AI.
That balance is not safety. A stable equilibrium among AIs could still leave no meaningful place for humans—“gradual disempowerment” remains live. Labenz believes there may be an AI for which “if anyone builds it, everyone dies” becomes true, but he does not think anyone is particularly close to building that singularly dominant system.
11. Safety is becoming a portfolio of partial defenses
Labenz continually asks for a measure that could “really work”—something strong enough to let him sleep well—and says essentially nobody has one. Intelligence may be inherently unpredictable, as human misconduct suggests, so frontier labs instead try to lower bad behavior, monitor what remains and contain failures across several independent layers.
The portfolio includes Goodfire’s “intentional design,” which tracks what models learn and tries to shape it during training; Redwood Research’s AI-control strategies for extracting useful work even from systems assumed to be deceptive; account-level enforcement; and monitoring that could catch 90%-plus of residual bad behavior. Formal verification stands out because it might remove entire cybersecurity attack surfaces rather than merely reduce their probability.
Biological defense must be physical as well as algorithmic: PPE stockpiles, ultraviolet lighting, wastewater monitoring and rapidly programmable vaccine platforms. The COVID vaccine design existed within days, Labenz notes, well before trials and distribution; those later steps took much longer.
Capability spending currently dwarfs funding for all these defenses, leaving the portfolio badly calibrated. Yet “20 bites out of the problem” might produce “a few nines of reliability.” Holden Karnofsky’s “success without dignity” framing captures the shift: society is not trying hard enough, but safety looks more tractable than five years ago, when Labenz had little concrete advice beyond “I don’t really know.”
12. Government should constrain the race, not nationalize the frontier
Labenz applies Biden’s line—“Don’t compare me to the Almighty; compare me to the alternative”—to frontier labs. Their competition can become dangerous and they require government involvement, but he trusts Sam Altman, Dario Amodei and Demis Hassabis more than Pete Hegseth and sees nationalization or military control as “a recipe for disaster.”
He supports Anthropic “100%” in its refusal to permit certain government uses and would preferentially send it tokens even at some performance cost; a private company retaining limits on mass surveillance represents, to him, a core American value worth defending.
He rejects restrictions on self-driving cars, medical advice or bot-delivered therapy, viewing much of that agenda as guild protectionism that deprives consumers of useful services. Government should instead address the race’s coordination failure, demand credible extreme-risk plans and prevent opaque, high-energy experiments whose developers themselves say could end badly.
The hosts frame automated AI research as the immediate governance test: labs reportedly anticipate recursive improvement without knowing where it leads, yet the work occurs inside a few small, ideological organizations. Even defining AGI as a system that does everything better than humans embeds a worldview, making transparency and external constraint important.
13. A months-long US lead cannot substitute for cooperation with China
Labenz “hates” the default strategy of racing ahead, preserving a lead through export controls and then solving coordination during a final slowdown. The lead is measured in months, not years; after spending the discussion cataloging unresolved safety and governance problems, he says society cannot plausibly resolve them within a few critical months.
Open source makes the one-way-door problem sharper. It could counter concentrated power, but releasing the wrong model might permanently place a bioweapon assistant—or an autonomous creator—into the public domain. That alone requires more preparation than a last-minute intelligence-explosion pause can provide.
His alternative begins from species-level commonality: “The real aliens in this situation are the AIs, not the Chinese.” He advocates researcher-to-researcher communication, broader diplomacy and even a CERN-like, jointly controlled facility—perhaps on a Pacific island or in Singapore—where Chinese and American teams could work together on the most sensitive problems.
Strategic dominance also invites counter-moves: Taiwan is closer to China, and advanced fabs are easier to destroy than rebuild. Labenz could support not selling chips to China while renting compute to Chinese companies, but rejects wholesale decoupling. With the Defense Department pressuring Anthropic over surveillance, he warns that America is “looking more like China all the time” and losing itself while trying to win.