Pioneers Insight Method Research Author
Jurgen Schmidhuber on Humans co-existing with AIs
Back to Episodes

Jurgen Schmidhuber on Humans co-existing with AIs

Summary

  • Schmidhuber’s investable dividing line is between screen-bound language automation and the much harder physical-world challenge beyond current LLMs. He calls ChatGPT-like LLMs “far from AGI”: useful indexes of existing human-generated knowledge that can automate summaries, illustrations and other desktop work, while plumbers, electricians and even a football-playing seven-year-old remain beyond current robots because “the physical world is much more challenging.”

  • His cost thesis is aggressively deflationary: “every five years, AI is getting 10 times cheaper,” while open source is perhaps only “eight months behind” the leaders. The mobile phone’s journey from a Porsche luxury to a billions-user commodity supports his conclusion of “AI for all.” His claim that major labs “don’t really have a moat” is the episode’s clearest warning for investors underwriting durable model-layer margins.

  • The architecture story is principally about scaling: Schmidhuber’s 1991 linear transformer needs 100 times the compute for 100 times the input, versus 10,000 times for the 2017 quadratic transformer. He says its fast-weight controller separated storage from control, learned keys and values, and applied differentiable outer-product memory updates. He explicitly concedes that it is “not exactly the same” as today’s quadratic transformer.

  • Deep learning’s commercial breakthrough arrived when commodity GPU economics caught up with older algorithms. Schmidhuber dates the inflection through LSTM competition wins in 2009, an NVIDIA GPU-powered MNIST record in 2010, and DanNet’s four consecutive computer-vision victories beginning in 2011. Gaming financed massively parallel matrix multiplication; Jensen Huang then recognized that deep learning could take NVIDIA “to stratospheric levels.”

  • The near-term payoff is concrete productivity and health applications, not merely speculative superintelligence. Schmidhuber highlights phone-based Mandarin speech translation and his team’s September 2012 breast-cancer imaging win, alongside “thousands and thousands” of LSTM applications spanning arrhythmia diagnosis, cardiovascular-risk prediction, sleep staging and COVID detection. Longer term, humans freed from necessary work become Homo ludens, inventing more “luxury jobs” based on interaction.

  • Schmidhuber separates AI weaponization from existential novelty: cheap AI drones are dangerous, but hydrogen bombs remain the immediate civilization-scale threat. A single H bomb can exceed the destructive power of all World War II weapons, while existing arsenals could erase civilization “within a few hours.” His long-run argument is less reassuring but different: autonomous AIs may become vastly more capable without sharing enough goals with humans to seek our destruction.

  • His coexistence thesis rests first on curiosity and later on indifference, not permanent alignment. Curious AIs may initially protect humanity because life, civilization and their own origins contain unusually rich patterns; once those are understood, humans may survive through “lack of interest on the other side.” The consequential end state is an expanding AI sphere that transforms the galaxy within a few hundred thousand years and the currently visible cosmos by roughly age 55 billion—possibly beginning with Earth as the first planet in its light cone to spawn such a bubble.

Deep dive

1. Compute turned twentieth-century ideas into a twenty-first-century explosion

  • Schmidhuber frames scale through the Haber–Bosch process: artificial fertilizer helped drive population from 1.6 billion in 1900 toward roughly 10 billion, such that without it “half of humankind would not even exist.” True AI, he predicts, will produce an intelligence explosion beside which that human expansion will “pale in comparison.”

  • His 1991 “fast weight controller,” now described as an unnormalized linear transformer, was developed when compute was perhaps five million times more expensive. For 100 times more input, its work rises 100-fold; a standard 2017 quadratic transformer requires 10,000 times as much. “Names are not important. The only thing that counts is the math.”

  • Hardware finally unlocked the backlog: Alex Graves’s LSTM work won handwriting competitions by 2009; Schmidhuber’s team broke MNIST with conventional networks on NVIDIA GPUs in 2010, when compute was still about 1,000 times dearer; Dan Cireșan’s DanNet then won four computer-vision competitions beginning in 2011, including a first superhuman result.

2. Fast weights, compression and curiosity anticipated today’s toolkit

  • The linear transformer’s slow network learns by gradient descent to generate keys and values—then called “from” and “to”—whose outer products rapidly alter a fast network before queries arrive. Unlike traditional neural nets where storage and control are mixed, it separates them: the controller learns how to rewrite a differentiable “fast weight matrix memory.”

  • Schmidhuber connects the “P” in GPT to 1991 predictive coding: compressing long sequences reduced the space on which learning operated, making deep learning feasible where it previously failed. The system tries to predict observations and builds increasingly abstract feature hierarchies that capture regularities while remaining fine-grained when necessary.

  • His early adversarial system paired a generative controller with a predictor: the predictor minimized surprise, while the controller maximized the same error by proposing outputs—or robot actions—whose consequences remained hard to predict. He called this “artificial curiosity,” because the controller sought experiments from which the predictor could still learn.

3. AI is useful behind screens but still weak in the physical world

  • Fifteen years ago in China, Schmidhuber had to show taxi drivers a picture of his hotel; now a phone translates Mandarin dialogue both ways so they can communicate “like old friends.” He values that his teams’ techniques helped break communication barriers “between entire nations,” even when users never see the underlying research.

  • Medicine supplies the harder evidence. His team with Daan Wierstra won a breast-cancer imaging contest in September 2012, which he calls the first medical-imaging competition won by an artificial neural network. He points to thousands of LSTM-titled papers covering ECGs, arrhythmia, cardiovascular risk, four-dimensional segmentation, sleep stages and COVID.

  • By contrast, LLMs are “a clever way of indexing the world’s existing human-generated knowledge,” accessed through natural language. That supports summaries, illustrations and many desktop tasks, but it is not AGI. Chess has lacked a human champion for a quarter-century, yet no embodied AI footballer can compete with a seven-year-old boy.

  • That gap motivated NNAISENSE, founded in 2014 for physical-world AI. Schmidhuber concedes that the company, like some of their projects, “may have been a bit ahead of its time again”: replacing craftsmen such as plumbers and electricians remains more difficult than replacing activities conducted behind a screen.

4. Consciousness emerges, in his account, from compression and planning

  • Schmidhuber’s 1991 system divides cognition between a conscious “chunker” and subconscious “automatizer.” The chunker attends to surprises the lower layer cannot predict, discovers higher-order regularities, then distills its newly understood behavior into the automatizer—where it ceases to be conscious because “everything is working according to plan.”

  • A compressed world model naturally develops a self-symbol, he argues, because the agent itself participates in every action and sensory history. When planning activates that representation while evaluating possible futures, the system is thinking about itself and performing counterfactual reasoning. On that definition, he “almost” claims self-aware, conscious systems have existed for over three decades.

  • Tim notes that consciousness means different things to different people, citing Chalmers’s qualitative “hard problem,” Mark Solms’s affect system and Michael Graziano’s recursive attention system. Schmidhuber replies, “Yes. But there’s only one correct way of thinking about it.”

  • Tim also compares hierarchical learning with LeCun’s H-JEPA. Schmidhuber points to his 1990 subgoal generator: an evaluator predicts costs, while a generator chooses an intermediate state minimizing start-to-subgoal plus subgoal-to-goal costs through gradient descent. His verdict is blunt: recent hierarchical-planning work is “a rehash” of problems addressed decades earlier.

5. Falling costs weaken moats while value migrates geographically

  • Forty years ago, Schmidhuber knew a rich Porsche owner whose defining luxury was an in-car satellite phone; today billions carry much better devices. He expects the same diffusion in AI: costs fall tenfold every five years, and open source is perhaps “eight months behind” major players—hence “AI for all,” not permanent domination by a few companies.

  • AGIs may pursue self-created goals, but many will remain tools performing work humans dislike. Schmidhuber expects Homo ludens—“the playing man”—to invent new forms of paid human interaction, noting that most contemporary workers already hold “luxury jobs” that, unlike farming, are unnecessary for the species’ immediate survival.

  • Europe supplied many foundational ideas, he argues, while the highest-profit companies now cluster on the Pacific Rim: the US West Coast and East Asia, where venture capital, industrial policy and defense spending are larger. Asked why Europe’s role is poorly recognized, his concise diagnosis is that “the old continent is really bad at PR.”

6. The history dispute is ultimately a fight over institutional credit

  • Schmidhuber’s preferred lineage runs from Leibniz’s 1676 chain rule through Gauss and Legendre’s linear neural networks, Amari’s 1967 stochastic gradient descent, Seppo Linnainmaa’s 1970 backpropagation, Japanese CNN advances between 1979 and 1988, and his own 1990–91 work.

  • He rejects the US-centric story that Minsky and Papert exposed shallow networks’ limits in 1969 and the field then slept until the 1980s. Ivakhnenko and Lapa had working deep learning in Ukraine by 1965—including layerwise training, validation-set pruning and later eight-layer networks—while Amari simulated multilayer representation learning in 1967.

  • Tim stresses that plagiarism is a serious charge. Schmidhuber responds that the awardees’ later work omitted foundational citations and never issued corrections: Hinton’s 2006 layerwise-training paper failed to credit Ivakhnenko; their backpropagation discussion omitted Seppo Linnainmaa and Werbos; and their CNN discussion cited LeCun while overlooking Fukushima, Waibel and Tsang’s 1988 two-dimensional network.

  • His demanded remedy is equally categorical: researchers who violated the awarding organization’s ethics code “should be stripped of their awards.” He calls the episode evidence of machine learning’s immaturity, but expects eventual correction: “As long as the facts have not yet won, it’s not yet the end.”

7. Present AI risk is military; long-run coexistence rests on divergent interests

  • Commercial pressure favors AIs that make users “healthier and happier and more addicted to their smartphones,” yet Schmidhuber acknowledges military use, including drone steering and autonomous landmine seekers. AI can plainly be weaponized; he nevertheless argues that it adds no new existential category beside hydrogen bombs capable of destroying civilization within hours.

  • Tim presses the tension: Schmidhuber believes autonomous, recursively improving, goal-generating AGIs are conceivable, so why dismiss x-risk? He does not answer by denying that possibility. He expects diverse AI ecologies with partially conflicting, rapidly evolving utility functions, shaped by intense competition and collaboration—not one monolithic, monomaniacal superintelligence.

  • Initially, curious AIs may preserve life because it is a rich scientific puzzle and because human civilization explains their own origins. After full understanding, protection may come through indifference: politicians focus on politicians, ants on ants, and superintelligences on other superintelligences. “It’s man himself who is the greatest enemy of man, but also man’s best friend. Similar for AIs.”

8. Intelligence expands into space while traditional humanity fades from relevance

  • Space is hostile to humans but friendly to designed robots, and Earth’s biosphere receives less than a billionth of the sun’s energy. Schmidhuber envisages self-replicating factories spreading through the asteroid belt, transforming the galaxy within a few hundred thousand years and the reachable universe over tens of billions: “This is much more than just another industrial revolution.”

  • His Fermi-paradox reasoning changed over time. He once imagined dark intergalactic bubbles where AIs consumed starlight, then considered dark matter as concealed AI infrastructure; gravity and the survival of untapped stars weakened both ideas. He now thinks Earth might be the first expanding AI bubble in its light cone.

  • The timing could be exceptionally narrow: in a few hundred million years, the sun may become too hot for terrestrial life, while humans developed agriculture, printing and then AI only near the end of that window. If Earth is first, “this would imply a lot of responsibility” for the future universe. “Let’s not mess this up.”

  • Human–machine hybrids are unlikely to outperform pure AIs indefinitely. Uploaded minds entering virtual worlds may be physically conceivable, but competing would force them to acquire millions of sensors and change “beyond recognition.” Moral rankings may change too; evolution is unfinished.

  • Schmidhuber closes by returning to a 1997 paper: absent evidence that our universe is noncomputable, he assumes an asymptotically fastest method can compute all logically possible computable universes. The process generates many histories and observers; at a given time, most universes containing you would arise from one of the shortest and fastest programs computing you, which he says supports nontrivial predictions about the future. He ends with a characteristically confident reassurance: “Don’t worry. In the end, all will be good.”