Pioneers Insight Method Research Author
Pushing compute to the limits of physics
Back to Episodes

Pushing compute to the limits of physics

Summary

  • Verdon’s central bet is that AI’s probabilistic workloads are running on deterministic hardware that spends energy suppressing noise, only for software to add randomness back in. Extropic’s mixed-signal silicon instead harnesses stochastic electron dynamics to accelerate Markov chain Monte Carlo across discrete and continuous variables. The pitch is to “loosen our grip on electrons” and turn entropy from a liability into the computational resource.

  • The scaling thesis is physical: using less power requires less charge, but once charge becomes scarce its discreteness creates unavoidable noise. Conventional machines target error rates around (10^{-15}), paying an escalating energy cost to preserve determinism inside the “thermal danger zone.” Verdon therefore calls thermodynamic computing inevitable: “You’re gonna have to go thermodynamic at some point.”

  • Extropic says it has moved from a three-pBit superconducting prototype to a silicon chip with roughly 300 degrees of freedom, with millions targeted next year. Its controllable pBits reportedly generate entropy using only “a few hundred attojoules”; results had been submitted for peer review, while private-alpha hardware access, papers, and open-source software were promised after the recording. The company targets 1,000–100,000X chip-level energy-efficiency gains, although Verdon explicitly separates that from full-system cooling costs.

  • A key investor-relevant projection is the wafer-scale estimate: about 1.5 billion pBits and 20 billion parameters on 20 watts. Verdon contrasts that with a current wafer-scale system consuming 20 kilowatts or perhaps 100 kilowatts, and argues that a few-hundred-billion-parameter thermodynamic model might land within 10X of brain efficiency rather than today’s claimed 100-million-X gap. These are forward projections, not demonstrated system benchmarks.

  • The near-term thesis is additive rather than a wholesale GPU replacement. Verdon expects deterministic processors to retain classical functions, probabilistic chips to handle sampling, and quantum hardware to supplement workloads involving quantum systems. He says existing build-outs remain safe “for the foreseeable future” because model porting will take time, even if a technology “1,000X better” changes future power-to-compute ratios.

  • Changing the substrate could change the winning model architecture, because “transformers are not sacred.” Softmax, diffusion, tree-of-thought search, Monte Carlo tree search, reinforcement-learning rollouts, and test-time reasoning already expose sampling-heavy demand; efficient physical sampling could revive energy-based models and induce a “Cambrian explosion” of architectures optimized for probabilistic hardware.

  • Verdon connects decentralized compute to a political objective: individuals should own the always-on systems extending their cognition. He rejects mere access to a centrally controlled “one God model” as assimilation into a “Borg mind,” forecasting a soft human–AI merge through shared perception and action before neural hardware arrives. His broader e/acc prescription—“accelerate or die”—favors growth, variance, and anti-monopoly experimentation, while Ramstead presses on cancerous growth, authoritarian local optima, and catastrophic tail risks.

Deep dive

1. Reductionism failed where complex systems became the object

  • Verdon traces his route from a seven-year-old fascinated by subatomic particles to math and physics at McGill and theoretical physics at Waterloo and the Perimeter Institute. He originally expected a compact theory—and perhaps exotic propulsion—to make humanity’s expansion tractable.

  • The intellectual break came when condensed-matter and quantum-gravity problems resisted reduction to a few interpretable parameters. Even with microscopic equations, emergent behavior may require “an amount of computation that is similar to that of the universe” to predict; many systems cannot simply be renormalized analytically into clean higher-level laws.

  • Verdon calls accepting opaque computational representations an “ego death”: the human physicist may not be “the hero of the story.” Tensor networks and deep-learning systems can be as inscrutable as nature, yet programmable complexity can still learn compressed representations and provide predictive power.

2. Quantum information supplied the bridge from physics to AI

  • The “it from qubit” program generalized Wheeler’s “it from bit,” treating the universe as a quantum computer or self-simulation. Whether or not AdS/CFT delivered unification, Verdon concluded that complexism—and computation capable of matching nature’s complexity—was the more productive framework.

  • His first answer was to “fight fire with fire”: use tunable quantum complexity to model quantum complexity. He helped develop early parameterized quantum programs called quantum neural networks, became an early Rigetti user, and later joined Google to build TensorFlow Quantum before leading quantum machine learning at Alphabet X.

  • The reusable principle was substrate matching: physics-informed representations should run on accelerators whose native dynamics implement the same physics. A quantum representation belongs on a controllable Schrödinger evolution; a stochastic representation belongs on a controllable stochastic evolution.

3. Quantum refrigeration made a hotter computer attractive

  • After nearly eight years in quantum computing, Verdon judged progress too slow and focused on fault tolerance’s thermodynamic burden. A quantum machine tries to remain near zero temperature and entropy, continually pumping away noise through what he calls “an algorithmic form of refrigeration.”

  • His inversion was simple: let environmental noise enter until quantum coherence fades and the useful dynamics become stochastic. Unlike quantum computing, he claims no complexity-class separation for thermodynamic acceleration—only potentially enormous constant-factor gains in speed and energy that make digital emulation impractical at scale, though not impossible.

  • Ramstead sharpens the definition: physics-based computing exploits a component’s actual physical properties instead of digitally simulating them. Verdon keeps the category broad—quantum, stochastic, photonic, even a neural net trained through “puddles of water”—while distinguishing physics-based execution from physics-inspired software.

4. Determinism has an energy price, so Extropic harnesses noise

  • Classical hardware forces transistors into absolutely on or off states by making signals dwarf electron jitter. Invoking Maxwell’s demon, Verdon argues that “knowledge comes at a cost”: reducing entropy and maintaining determinism every clock cycle necessarily consumes energy.

  • His alternative is to “learn to let go”—moving from tightly yanking signals around to “gently guiding” noisy dynamics closer to equilibrium. The philosophical analogy is deliberate: software already surrendered imperative programming to gradient descent, and hardware should similarly stop treating every fluctuation as an enemy.

  • Technically, Extropic is building accelerators for Markov chain Monte Carlo with discrete variables, continuous variables, and mixtures. Its newest chip is mixed-signal: stochastic electronics generate proposals or entropy, while conventional digital components perform non-random operations such as Metropolis–Hastings acceptance and rejection.

5. AI software is probabilistic even when its machines are not

  • Ramstead identifies the stack’s paradox: hardware expends power removing randomness, then sampling algorithms restore it in software. Verdon adds that a transformer’s softmax already makes it “a big probabilistic computer,” while diffusion reverses a Markov chain and test-time reasoning increasingly uses Monte Carlo tree search, tree-of-thought search, or RL rollouts.

  • Extropic was founded in 2022 on the prediction that major workloads would move toward sampling and probabilistic inference. Verdon describes current architectures as interpolating between deterministic forward passes and full energy-based models; diffusion occupies part of that middle ground.

  • Two transformer-paper authors who invested in Extropic repeatedly told the team, “Transformers are not sacred.” They worked because they fit Google TPUs; alter the hardware’s fitness landscape, Verdon argues, and model search could enter a high-temperature phase—a “Cambrian explosion” of architectures native to probabilistic compute.

6. Moore’s law becomes Moore’s wall inside the thermal danger zone

  • Conventional processors generally seek error rates near (10^{-15}), allowing vast numbers of operations before correction becomes necessary. But miniaturization raises fluctuation-to-signal ratios: “To use less power, you need to use less charge,” and once little charge remains, its discreteness produces noise “period.”

  • Ramstead’s tennis-ball analogy separates regimes: macroscopic motion ignores molecular vibration, quantum machines preserve coherent superpositions, and thermodynamic devices operate where fluctuations coexist with component-scale signals. Verdon’s claim is that increasingly small deterministic transistors are forced toward that third regime anyway.

  • Verdon presents the brain as proof that useful thermodynamic intelligence exists, and Ramstead suggests it operates close to the Landauer limit. Verdon agrees that neurotransmitters move through stochastic chemical-reaction networks and argues that electron-based circuits might execute analogous probabilistic programs more efficiently because “electrons are much lighter than big neurotransmitters.”

7. The pBit turns an energy landscape into programmable probability

  • Extropic’s chip can be viewed as a time-dependent programmable energy function with diffusion resembling Langevin dynamics, itself an MCMC method. The analogy is literal: Bayesian sampling algorithms can be embedded into the native motion of electrons rather than numerically simulated step by step.

  • A pBit resembles a double-well landscape containing a bouncing ball. One well represents zero and the other one; adjusting the tilt controls how long the signal occupies each state, creating a “fractional bit” that continually dances between zero and one.

  • The company began with three superconducting pBits, then reproduced stochastic primitives in silicon; Ramstead describes the latest chip as having about 300 degrees of freedom, though Verdon cautions that they are not all pBits. He reports entropy generation at a few hundred attojoules and targets millions of degrees of freedom next year.

8. Silicon is the product, while superconductors were the learning platform

  • Superconducting circuit quantum electrodynamics offered a direct map from a desired Hamiltonian to a circuit, making it a useful early engineering platform. Depending on material, those devices operate at a few hundred millikelvin or a few kelvin; maintaining the cryogenic boundary is a major practical cost.

  • Verdon concedes the extreme thought experiment: millions of superconducting pBits inside a football-field-sized dilution refrigerator might form the most efficient probabilistic computer, but it is “probably not” practical or appropriate for a startup. He instead offers the work to academia and national laboratories as a community platform.

  • Extropic chose room-temperature silicon for commercialization. The claimed 1,000–100,000X efficiency range applies at chip level, with the cooling asterisk acknowledged; lower-temperature thermodynamic devices consume less locally, but maintaining the cold bath carries its own cost.

9. The winning stack is heterogeneous, not thermodynamic everywhere

  • Verdon imagines computation following physics across scales: early layers might preserve quantum complexity, intermediate layers retain probability and entropy, and later deterministic layers perform coordinate transformations after information has been distilled.

  • Practical graphs likewise mix a classical differentiable function defining a distribution with a probabilistic accelerator sampling from it. “You don’t necessarily need entropy everywhere in your graph all at once,” and thermodynamic hardware will not be best for every deterministic or high-precision operation.

  • Quantum computers may remain supplements for quantum systems; probabilistic and deterministic processors cover most other workloads. Verdon therefore tells owners of planned GPU and TPU infrastructure that porting will take “quite a while” and thermodynamic compute should initially appear as an add-on, leaving foreseeable build-outs usable.

10. Energy-based models expose both the upside and the power constraint

  • Verdon frames modern neural networks as mean-field approximations of energy-based models: deterministic hardware pushed researchers toward moments, Gaussian families, matrices, and vectors that map efficiently onto GPUs. Faster physical sampling could make the broader EBM family practical rather than forcing distributions into hardware-friendly summaries.

  • At wafer scale, he projects roughly 1.5 billion pBits and 20 billion parameters using 20 watts, versus 20 kilowatts or perhaps 100 kilowatts for a current wafer-scale system. A multilayer, few-hundred-billion-parameter system might approach within 10X of brain efficiency, he says, compared with a present gap he puts near 100 million X.

  • Verdon’s categorical constraint is that current hardware cannot scale ubiquitous agents, video models, world models, and embodied intelligence. Even abundant fusion power would ultimately raise planetary heat radiation: “We’re literally cooked.” The intended market therefore extends beyond generative AI into simulation, optimization, statistical inference, and science.

11. e/acc applies thermodynamic selection to technology and politics

  • Ramstead introduces Effective Accelerationism as a counterpoint to effective altruism; Verdon defines it as a “cultural hyperparameter prescription” maximizing civilization’s free-energy production and consumption on the logarithmic Kardashev scale. His stochastic-thermodynamics framing moves from selfish genes and memes to “selfish bits”: trajectories dissipating more free energy become exponentially more likely, making “accelerate or die” the movement’s intentionally dramatic slogan.

  • Ramstead’s pushback—worth keeping—is that unbounded growth can produce cancer, dictatorship, or imposed homogeneity. Verdon answers that these are local optima over short horizons: cancer kills its host, while burning everything immediately prevents future capture. The proposed objective is strategic growth over an effectively infinite horizon, preserving order and resources to secure more resources later.

  • Ramstead explains the hyperstition mechanism through active inference: perception and action can reduce model-world divergence, and “the car goes where the eyes look.” He then gives the controversial example that COVID was an accident from defensive bioweapons research, hedging that it was “probably defensive.” Verdon does not substantively endorse that example, but agrees with the broader optimistic-steering logic and says every action begins with a false belief.

  • Geopolitically, Verdon favors variance over monopoly: the US acts like high-temperature search, while China acts like a low-temperature optimizer that scales once a gradient is clear. DeepSeek’s exploration under export-control constraints is his example of an alternate hyperparameter region exposing Western convergence; his prescription is balanced top-down coordination and bottom-up exploration, because his “P90, 1984” was higher than his AI “P doom.”