The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
Summary
Benchmark success is not evidence that a model has learned a useful world model. Kenneth Stanley’s PicBreeder counterexample produces the same perfect-looking skull through two routes: open-ended evolution yields compact, modular controls for symmetry and the mouth, while conventional SGD yields “total spaghetti.” The investable warning is that surface capability can mask “imposter intelligence” whose poor structure only becomes costly during adaptation, continual learning, or out-of-distribution work.
The episode challenges the scaling thesis at its most capital-intensive point: conventional training may spend vast resources compensating for the wrong internal degrees of freedom. Stanley carefully targets “conventional SGD”—fixed architecture, fixed objective, gradient pursuit—not every conceivable use of SGD. If models instead discovered factors aligned with the world’s structure, he speculates training could become “multiple 10X, 100X” more efficient than today’s billions or hundreds of billions of dollars in infrastructure.
PicBreeder suggests that “it matters not just where you get, but how you got there.” Humans selecting among roughly 15 mutations inject only a few bits of information per generation, yet over dozens of iterations the system locks in symmetry and other reusable conventions without anyone explicitly teaching them. This path dependence suggests that data order, curriculum, topology growth, and selection for future evolvability may matter alongside dataset size.
Poor representations may explain why current models deliver derivative creativity without reliably producing transformative creativity. Stanley’s distinction is between a novel bedtime story and “inventing a new genre of literature”: the first recombines within an inherited distribution, while the second requires well-factored abstractions that open new conceptual paths. Products that automate generation and curation could therefore impress initially yet accelerate cultural mode collapse—“freeze pop culture in the year 2025”—unless humans keep supplying the genuinely new ideas.
The alternative architecture thesis is to grow capability from sparse, protected modules rather than train a dense monstrosity and prune it afterward. Tim Scarfe imagines beginning with 100 parameters, expanding to 1,000, then 10,000 and 100,000 while turning a useful 12-neuron subnetwork into 120 related neurons and preserving its abstraction. Akarsh Kumar characterizes the autonomous method for producing unified factored representations as the “trillion-dollar question”—a plausible source of disruption, but still a research agenda rather than a working recipe.
The guests reject both complacency and imminent-autonomous-AGI certainty. Kenneth Stanley says present AI already causes “massive harm” and retains roughly a one-third all-cause doom estimate, yet does not think any currently machine-trainable architecture leads to AGI; Scarfe likewise argues present systems require humans because autonomy exposes their incoherent action spaces. Their shared boundary is categorical: today’s narrow tools can amplify people powerfully without possessing the representation needed for independent, transformative agency.
The portfolio-level recommendation is diversification, not abandonment of scaling. Kumar wants some researchers to keep testing how far LLM scaling can go while directing materially more attention to artificial life, open-endedness, PicBreeder-like selection, curricula, and evolutionary mechanisms. The opportunity is large precisely because “everybody’s just stuck on SGD, SGD, scale, scale,” but the paper supplies a vivid counterexample and open questions—not proof that grokking, mixture-of-experts, or existing optimization cannot mitigate the problem.
Deep dive
1. Identical outputs can hide opposite kinds of intelligence
Stanley’s opening observation is visceral: an SGD-trained network and an evolved CPPN can output the same skull, yet their internal representations look like “amazing versus garbage.” The conventional model is entangled spaghetti; PicBreeder’s version resembles something deliberately engineered.
The counterexample matters more than another complaint about opaque networks. Without PicBreeder, one could assume neural representations intrinsically look messy; its compact, legible networks show that “it is not how life has to be.”
Stanley stops short of declaring deep learning fundamentally broken. The paper instead questions “representational optimism”—the unstated belief that good results imply good machinery underneath—and asks whether the stark alternative deserves a new training agenda.
2. PicBreeder found greatness by refusing to search for it
PicBreeder let crowds breed images encoded by compositional pattern-producing networks, or CPPNs. These networks map coordinates such as X and Y into hue, saturation, and luminance, use functions including sine and cosine, and generate resolution-independent images.
The butterfly from Stanley’s book symbolizes deception: users explicitly trying to evolve a butterfly generally failed because its stepping stones did not resemble butterflies. People following whatever seemed novel or promising instead reached surprising artifacts in “surprisingly few steps.”
That lesson produced novelty search and later quality-diversity work. Stanley notes that AlphaEvolve uses MAP-Elites under the hood: an example of systems illuminating many possibilities instead of converging directly on one prescribed point.
3. Direct optimization paid a steep representational price
When Stanley’s team made NEAT target a known PicBreeder image, complex targets such as the skull generally remained unreachable because image similarity supplied a deceptive heuristic. Simpler targets such as a crescent could sometimes be recovered, but the resulting network was roughly triple the original complexity.
Joel Lehman and Stanley later tried SGD. The tiny evolved architecture offered too few degrees of freedom for gradient descent, but a sufficiently large fixed network could reproduce even complex PicBreeder images almost perfectly.
The output concealed the cost: SGD failed to encode the skull’s symmetry once, instead re-representing disconnected pieces across both sides. The evolved network factored symmetry, mouth shape, opening, smiling, and other transformations into coherent reusable components.
Because every hidden CPPN node can be queried across all XY coordinates and rendered as an image, this is unusually direct evidence. The comparison does not depend on an interpretability proxy: “we can literally see what the representation is.”
4. Conventional SGD optimizes the proxy, not the representation users need
Kumar deliberately says “conventional SGD,” meaning the combined paradigm of a fixed architecture, fixed objective, and SGD chasing that target. The authors do not know whether SGD itself is irredeemable or whether another application of it could yield better factors.
His Goodhart-style diagnosis: training perfectly solves the stated loss, but the real demand is a representation supporting adaptivity, generalization, out-of-distribution behavior, creativity, and especially continual learning. Those downstream properties are far harder to formalize than loss.
Kumar preserves Stanley’s father’s analogy: two mathematicians can ace the same exam, yet one may transform the field while the other discovers nothing. “The test” therefore reveals little about the future capability that actually matters.
5. Human guidance may not fully explain the information created
Scarfe’s strongest pushback is that PicBreeder’s users inherited billions of years of evolution and a rich abstract world model. Their preferences could be an implicit imitation-learning channel that telegraphs symmetry, faces, and other natural motifs into the network.
Stanley’s honest uncertainty: this is “a deep, almost philosophical question” whose answer is unknown. But dozens or hundreds of clicks seem insufficient to demonstrate dimensions such as closed mouth, open mouth, and stem rotation—and users were choosing attractive possibilities, not constructing training examples.
Kumar sharpens the information argument: users saw about 15 mutations and clicked one, injecting only a few bits per generation. That channel cannot explicitly describe all the regularities appearing in a butterfly or skull after a few dozen generations.
Their proposed mechanism is hierarchical lock-in. Once a user selects a symmetric ancestor, symmetry becomes a convention; later exploration varies within it, adds new conventions, and gradually constructs an elegant hierarchy without anyone specifying the final object.
6. One apple weight behaves like an untrained 3D world model
Stanley’s favorite specimen is an evolved apple containing hundreds of connections but one weight that swings the stem continuously from left to right. The leaf rotates as if on a 3D axis, its shadow moves underneath, and the symmetric apple remains undisturbed—“it’s been decomposed.”
The network had never been trained on apples or swinging stems, and Stanley argues it would be circular to claim the trajectory demonstrated that movement: if the stem were already swinging, the capability would already exist. He treats it as a de novo “hypothesis about the world,” contrasting it with Move 37’s emergence after millions of games.
7. Compression is valuable, but the factorization determines usefulness
Stanley rejects a simple equation between maximal compression and intelligence. A face explicitly factored into eyes, nose, and mouth may be preferable to a smaller encoding that lacks those components, because the factored version generates principled new faces.
Kumar makes the same distinction through evolvability: a skull can be compressed aggressively yet remain useless for creative leaps. The valuable representation is not merely small; it can vary “in any direction of interestingness” without destroying established structure.
PicBreeder implicitly selects that property. Between two superficially similar skulls, the one whose components produce interesting descendants keeps attracting clicks, so modular, adaptable lineages tend to outlast brittle spaghetti.
8. Fractured representations impose a creative ceiling
Stanley calls present models capable of “derivative creativity”: ask for a bedtime story and receive a genuinely new one, but not a literary-prize winner or a new genre. “Transformative creativity” requires abstractions that expose previously unavailable directions.
Evolution wrapped around a large model—including the team’s evolution-through-large-models work and AlphaEvolve—can push outside the model’s distribution. Stanley regards this as inefficient beside a human mind that leaps directly through abstraction levels encoded in its representation.
Discussing Andrej Karpathy’s proposed gradient search over VO3-generated culture, Stanley asks who supplied the original creativity: the model rendering an ape doing ASMR, or the human who conceived it? Automated ideation and curation could still impress while converging on increasingly familiar variations.
His warning is cultural as well as technical: a feedback loop may “freeze pop culture in the year 2025” and remain there. Current AI powerfully amplifies a human germ of invention, but it cannot accelerate discoveries humans are not already imagining—in science, music, or art.
9. Intelligence needs a Goldilocks number of degrees of freedom
Scarfe compares dense networks to a pile of sand or block of clay: SGD needs enormous freedom, after which approaches such as the Lottery Ticket Hypothesis carve the result down. PicBreeder begins sparse and builds only what its history makes useful.
Stanley’s formulation is “a memory with an algorithm”: use as much capacity as necessary and no more, despite the engineering difficulty of identifying that boundary.
Einstein becomes their analogy for disciplined complexity. Stanley notes that Einstein regarded adding the cosmological constant as his greatest blunder, although it is now needed for valid scientific reasons—an example of both simplicity’s power and the danger of removing a degree of freedom the world needs.
10. A better network may need to grow rather than be carved down
Scarfe imagines a future training run starting with 100 parameters on simple data, expanding to 1,000, then 10,000 and 100,000. A discovered 12-neuron module might become 120 related neurons while remaining sufficiently isolated to deepen the same abstraction.
The destination would be neither generic sparsity nor a monolith, but historically grown modularity: “seeds” expanding into specialized structures instead of a massive initial network entangling everything available.
Scarfe connects this to NEAT and monotonic complexity. New information and topology are added while provenance is preserved; mutations occur inside viable frames, which is why children still have two legs rather than crossover randomly scrambling every body convention.
11. Activation functions matter, but training determines the geometry
Scarfe contrasts ReLU networks’ piecewise-linear partitions with CPPNs’ trigonometric functions, which continue beyond the observed support. His robust-function intuition is Y equals X squared: for unseen inputs it still does something structured rather than entering “no man’s land.”
Stanley’s pushback—worth keeping—is that activation functions are not the whole problem. Even with ReLUs, a constructive training process might divide a spiral into intelligible quadrants and 45-degree regions instead of the “funky angles” visible after ordinary SGD.
Scarfe widens the possibility space: some RNN constructions can be Turing complete yet remain untrainable by SGD, while handcrafted or hybrid systems may require expandable memory and generalized autoregression. The bottleneck is not what neural networks can represent, but what prevailing training can discover.
12. Grokking may help without answering the counterexample
Stanley concedes that grokking, mixture-of-experts, convolution, and other interventions might mitigate fracture; the paper has not tested every possibility. Yet the visual gap is so dramatic that he doubts a simple cleanup will transform SGD’s skull into PicBreeder’s structure.
His challenge is more basic: grokking first creates an entangled mess, then removes redundancy and fracture later. PicBreeder demonstrates that a representation can be “good in the first place,” making constructive formation at least a legitimate alternative.
The efficiency stakes could be enormous. Stanley contrasts billions or hundreds of billions of dollars in training infrastructure with a hypothetical representation whose degrees of freedom already align to reality; he speculates that could be “multiple 10X, 100X” more efficient, while stressing that the magnitude is unknown.
13. Learning order can create or prevent representational fracture
Stanley connects PicBreeder to POET-like open-ended curricula: tasks become more complex naturally, but the trajectory remains divergent rather than marching through a predetermined syllabus. Chronology becomes part of the learned structure.
Dumping all data into one batch ignores prerequisite order. An LLM may absorb calculus before arithmetic and invent a heuristic approximation of arithmetic while learning another version elsewhere, producing redundant, diminished concepts that are “fractured into pieces.”
Dr. Dugger’s personal example makes the cost concrete: non-calculus physics required memorizing separate cannonball formulas; after moving to the calculus class, he could derive them. Same apparent answers, radically different capacity to generalize.
Kumar relates this to an interpretability example where adding 23 and 57 becomes a maze of approximate facts—numbers “around 55” and “around 25”—whose weighted paths happen to land correctly. It works, but cannot adapt like an actual arithmetic abstraction.
14. Imposter intelligence can remain perfect on the surface
The paper’s phrase “imposter intelligence” describes a model that reproduces the right output with the wrong internal ontology. Its skull looks real, but underneath “it’s not really a skull”; it recognizes neither the object’s parts nor its regularities.
Stanley scales the metaphor to an LLM as an image of all human knowledge. A system could answer every in-distribution question convincingly while organizing knowledge as a “giant charade,” leaving continual learning and genuine invention prohibitively difficult.
The discussion links the hypothesis to mechanistic interpretability: polysemantic neurons may participate in addition and something unrelated such as black holes, while concepts are spread across many circuits. Interpreting a clean hierarchy would be hard; interpreting a fractured, entangled one may be intrinsically worse.
15. Biological evolution is constrained divergence, not optimization
Stanley argues that genetic algorithms damaged AI’s metaphor for evolution by turning selection into convergence on one target. Natural evolution has no equivalent final point; it diverges while remaining subject to survival.
Flight and photosynthesis illustrate the inversion. In an optimization experiment either could be the celebrated objective, but in nature each is a side effect of viability—and for Stanley, “the side effect is actually the main event.”
Survival therefore acts as a constraint, not a gradient that straightforwardly predicts innovation. Saying organisms optimize survival does not explain why photosynthesis follows; at the origin, nobody could derive that achievement from the constraint.
Dr. Dugger adds resource and energy pressure, while Stanley insists on another Goldilocks balance: excessive pressure suppresses experimentation, but trivial survival fills the world with inert blobs. Open-ended systems need a non-trivial “minimal criterion” without global competition collapsing into local hill climbing.
16. Canalization preserves structure while evolution explores
Stanley sees nature’s representation as unified and factored because variation preserves its deepest conventions. Children differ from parents, yet remain bilaterally symmetric and ordinarily retain arms, legs, and the inherited body plan.
Biology calls this canalization: development has “dug a trench into a mountainside,” so mutations resemble earthquakes while the water still follows the established canal. Conventional genetic algorithms, by contrast, damage the phenotype’s core regularities on mutation.
Artificial encodings can jump easily from bilateral to trilateral symmetry or five fingers to 10, yet biology almost never does. For Stanley, nature is evidence that autonomous divergent processes—not only humans clicking PicBreeder—can discover representations that remain evolvable without dissolving.
17. Open-ended search becomes less predictable as its history grows
Stanley’s “cone of inevitability” is narrow near an origin and expands with time. Parallel evolutions might repeatedly discover photosynthesis or eyes, but probably not humans; Kumar’s peacock is the kind of contingent form no one should expect to predict.
Novelty also accumulates information about the universe. Eyes exploit photons, ears exploit sound waves, and organisms gradually become “an encyclopedia” of physically available degrees of freedom—although many different configurations could expose the same underlying possibility.
Scarfe asks whether genuine intelligence still requires the physical world that supplies evolution’s staggering computation. Stanley allows that interaction may be necessary, but treats the internet as an increasingly direct proxy through which a model could explore, choose its own chronology, and perhaps learn far more efficiently.
Their commercial examples preserve the paradox: school-bus safety might be solved by making school buses obsolete; Scarfe says YouTube began as video dating, while GPUs built for games won the hardware lottery in AI. “What we think we want” can blind objective search to the market that replaces it.
18. The research agenda is representation-first and deliberately plural
Stanley’s recommendation is to measure whether current models contain imposter representations, test mitigations, and treat creativity and open-endedness as first-class problems. Reasoning toward a specified answer is useful, but “antithetical to creativity” when intelligence means finding value without knowing the destination.
His model is a child in a playground or a researcher following a justified gut instinct because a path “opens up a new playground.” Training only direct problem solving teaches systems to navigate once a destination is supplied, not to recognize which unexplored direction has potential.
Kumar applies open-endedness to the field itself: keep scaling LLMs to learn how far the paradigm goes, but stop putting “all our eggs in one basket.” Academia should invest materially more in artificial life, evolution, PicBreeder-like systems, and alternatives to IID batches repeated for millions of steps.
Stanley separates that architectural skepticism from complacency about harm. He assigns roughly one-third probability to all-cause doom and says AI causes “massive harm” today, yet sees no currently machine-trainable architecture leading to AGI; his closing distinction is categorical, between powerful narrow intelligence and general intelligence.