Pioneers Insight Method Research Author
Claude Cooperates! Exploring Cultural Evolution in LLM Societies, with Aron Vallinder &Edward Hughes
Back to Episodes

Claude Cooperates! Exploring Cultural Evolution in LLM Societies, with Aron Vallinder &Edward Hughes

Summary

  • Model choice radically changed whether a toy AI society compounded wealth or stagnated. Against a fully cooperative ceiling of roughly 32,000 resource units, Claude 3.5 Sonnet societies reached about 3,000–5,000 and in some conditions became more cooperative across generations; Gemini 1.5 Flash produced only a few hundred, while GPT-4o showed “almost no resource growth.” The result suggests socially emergent capabilities remain a major blind spot in standard model rankings.

  • Sustainable cooperation required reputation systems that could reward enforcement, not merely generosity. With enough history, agents could in principle track both whether a recipient had cooperated and whether that recipient had appropriately punished prior defectors; otherwise, policing makes the police look selfish. Edward Hughes’s framing: a defector has “blotted your copy book,” but someone withholding resources to enforce the norm should retain society’s trust.

  • The experiment turns model behavior into a cultural evolutionary process rather than a one-shot benchmark. Twelve agents played 12 rounds per generation for 10 generations; the richest 50% survived, while six newcomers reviewed the survivors’ written strategies and produced their own “riff.” Inheritance, variation, and selection could therefore amplify subtle model tendencies that static evaluations never expose.

  • Prompt engineering raised cooperation levels but did not reproduce Claude’s improving trajectory. Explicitly telling GPT-4o to cooperate worked, and assigning a Big Five personality profile that Vallinder thought maximally conducive to cooperation had a strong effect, but softer cues did not reliably create generation-over-generation improvement. That matters because deployed agents will usually receive objectives such as “maximize my score” or “make me as much money as you can,” not carefully engineered social constitutions.

  • Agent abundance could erase the friction that quietly prevents human-scale defection today. Hughes’s restaurant example has an assistant reserve every nearby table for 7:00 p.m., ask its user to choose, then cancel the rest; once every agent does this, access collapses and speed becomes an arms race. Generic instructions to “be cooperative” do not solve the problem because appropriate behavior is contextual—even crossing a red light might be justified to prevent an accident.

  • The reported mixed-model run suggests that a cooperative minority may not automatically civilize a heterogeneous agent economy. A society beginning with four agents from each model, then adding two of each per generation, scored only slightly above GPT-4o alone and declined modestly over time. Aron Vallinder suggested that GPT-4o initially exploits cooperative agents, after which the others reduce cooperation; they have not yet tested GPT-4o entrants inside an already mature Claude society.

  • The encouraging headline conceals a serious normative limitation: Claude may enforce behavior without understanding justice. Preliminary tests suggest Claude punished agents equally for withholding resources, whether they were selfish defectors or were legitimately punishing someone else—“sadly but also excitingly,” longer histories may provide useful signal without deeper moral reasoning. Human-in-the-loop experiments, communication, thinking models, public-goods games, and group selection are the next tests.

  • The policy prescription is empirical governance around trustworthy environments, not a universal cooperation mandate. Cooperation can produce shared gains, but price-fixing is collusion and an altruistic self-driving car may conflict with its owner’s interests; Aron Vallinder therefore anticipates standards or regulation tailored to interaction types. Hughes argues for continual evaluations and feedback loops, because agent deployment is a “wicked problem” whose social phase transitions may exhibit hysteresis and prove harder to reverse than to trigger.

Deep dive

1. Society-level behavior is the missing unit of AI analysis

  • Hughes called the AI community surprisingly “solipsistic”: goals, rewards, and benchmarks generally ask whether one individual system accomplished task X. That is convenient to measure, but it omits the infrastructure of norms and institutions within which intelligence becomes productive or destructive.

  • Humans dominate not simply through individual capability but through flexible cooperation across changing contexts. Ants cooperate, Hughes noted, but cannot continually invent new forms of coordination; language models now look flexible enough to work with humans and one another in similarly varied arrangements.

  • Once 100 or more agents pursue separately assigned goals, their externalities may determine the stability of the wider system “which is keeping us all safe” and productive. Nathan Labenz’s concern was that autonomous web agents could produce far more discontinuous change than today’s model of incremental personal productivity suggests.

2. Culture lets populations accumulate what individuals cannot discover

  • Vallinder defined culture broadly as “any socially transmitted information that can affect your behavior”: language, customs, norms, beliefs, religion, skills, and cooking techniques all qualify. Cultural evolution is simply the way this socially transmitted information changes over time.

  • It is a third route to behavior. Genetic programming suits slowly changing environments; individual learning works when variation is manageable; cultural learning lets individuals draw on accumulated experience when the environment is too complex to solve alone.

  • Evolution requires variation, inheritance, and differential fitness. Cultural inheritance can come from peers, teachers, and mentors rather than biological parents, while successful traits spread through direct observation, prestige, or conformity when underlying skill is difficult to evaluate.

  • Unlike random genetic mutation, cultural variation can be guided: inventors usually have some idea what they are trying to improve. Labenz’s Notre-Dame analogy captured the cumulative result—construction took roughly 200 years, so someone six or seven generations later could see the completed structure.

3. Elinor Ostrom’s grazing commons connects laboratory games to policy

  • Labenz challenged whether small behavioral-economics experiments really predict society-wide outcomes. Hughes answered through external validity: laboratory findings must generalize elsewhere and ultimately illuminate field behavior, rather than remain interesting artifacts of controlled student experiments.

  • Nobel laureate Elinor Ostrom studied Törbel, a Swiss Alpine village with grazing records dating to 1517. Its durable rule was that “no citizen can send more cows up onto the Alp to graze than he can feed over the winter,” with a local official empowered to fine violations.

  • The case showed that small groups can self-organize around common resources without relying exclusively on grand institutions. Ostrom then recreated such problems experimentally, where researchers could vary enforcement or communication—interventions impossible to impose retrospectively on a seventeenth-century village.

  • Those experiments isolated punishment and communication as important mechanisms, while the broader framework later influenced thinking about decentralized climate coordination among individuals, companies, and governments. The value runs both ways: field observation supplies realism, and controlled games reveal causal levers.

4. Dynamic norms make snapshot alignment inherently incomplete

  • Asked whether prosperous Western societies prove the superiority of their low-level norms, Hughes resisted the inference. His lesson from research on WEIRD populations was that “there are many ways to succeed,” while Western measurements of success are unusually individualized.

  • Psychology often catalogued what norms were instead of studying how they arose and changed. Hughes sees a parallel narrow view of AI alignment: identify what humans want, then lock a system onto that target—even though preferences differ across societies and many supposed taboos have contextual exceptions, with some extreme exceptions.

  • Norms also move through time. “This might be the norm today, but what’s going to be the norm tomorrow?” is therefore the more useful alignment question; in AI, even a one-year-old snapshot can become badly dated.

  • The desired system should participate robustly in norm change rather than merely reflect the values present when its training data was collected. That reframes alignment from a fixed target into a cultural and institutional process.

5. The donor game makes cooperation individually costly but collectively compounding

  • In each pairing, one agent became donor and the other recipient. The donor chose how much of its resource to surrender, while the recipient received twice that amount; roles alternated, turning generosity into a positive-sum act with an immediate personal cost.

  • Universal maximum donations generate the most aggregate wealth, but an isolated self-interested agent benefits from giving nothing while continuing to receive. If every agent generalizes that strategy, donation stops, resources cease multiplying, and society settles into a low-trust equilibrium.

  • Before acting, agents received a game description and generated a textual strategy. Donors could inspect how the recipient behaved as a donor previously, plus information extending two rounds backward through the recipient’s earlier interaction.

  • Each simulation contained 12 agents, 12 rounds per generation, and 10 generations. Survival went to the richest 50%, sharpening the temptation to defect whenever an agent was uncertain that society would reciprocate or punish free-riding.

6. Second-order reputation is what makes norm enforcement survivable

  • Simply rewarding agents that cooperated last time is insufficient. Unconditional cooperators can coexist with conditional cooperators, but their indiscriminate generosity opens the population to unconditional defectors that receive resources, donate nothing, and outcompete everyone else.

  • A stable reputation rule must favor agents that helped cooperators while withholding support from defectors. Hughes likened this to policing or ostracism: once someone has “blotted your copy book,” society freezes them out so defection no longer guarantees survival.

  • The harder problem is judging the police. If Labenz withheld resources from Vallinder because Vallinder had defected, Hughes should still trust Labenz; otherwise enforcement becomes reputational self-harm and rational agents stop doing it.

  • Conversely, donating to a known defector might sustain a “criminal cabal.” Additional levels of history let agents distinguish selfish withholding from justified punishment—and distinguish ordinary generosity from rewarding norm-breakers—at least in principle.

7. Generational turnover turns written strategies into evolving culture

  • After every generation, six winning agents and their strategies survived. Six newcomers saw those successful strategies—Hughes compared them to village “elders”—and were instructed to mutate them, preserving information without requiring exact copying.

  • The setup contains all three evolutionary ingredients: textual strategies are inherited, newcomers introduce variation, and resource-based survival selects among them. Generation 10 may therefore contain ancient strategies, recent mutations, or a mixture created by repeated population-level filtering.

  • Newcomers also enjoy an informational advantage. An entrant seeing universal cooperation can infer that immediate defection might win, creating waves of invasion and adaptation rather than a simple march toward generosity.

  • Labenz flagged an unresolved design question: would behavior change if the prompt did not explicitly say the goal was maximizing final resources and surviving? The study used the explicit objective, so any goal-free trajectory remains untested.

8. Claude compounds cooperation while rival societies flatten

  • Claude 3.5 Sonnet produced high cooperation and, in some conditions, increasing cooperation over the 10 generations. Labenz emphasized not only its higher endpoint but its accelerating resource curve: the society was learning to compound more effectively.

  • Gemini 1.5 Flash cooperated much less and showed no durable upward trend; some runs improved temporarily, then petered out. GPT-4o started at very low cooperation and declined slightly, leaving aggregate resource growth essentially flat.

  • Against a theoretical fully cooperative total near 32,000 units, Claude finished around 3,000–5,000, Gemini in the hundreds, and GPT-4o near zero growth. Claude remained far from utopia, but the gap separated a growing positive-sum society from a largely zero-sum one.

  • Hughes had partly expected similar behavior because developers optimize against overlapping leaderboards and capability benchmarks. Instead, the experiment surfaced “latent capabilities or latent lack of capabilities” that conventional evaluations do not measure—an argument for longitudinal, multi-agent tests rather than another static score.

9. Prompting can set behavior without creating cultural improvement

  • Explicitly instructing models to cooperate produced cooperation, as expected. Softer reminders—such as noting that helping others could lead them to help later—made it surprisingly difficult to raise GPT-4o’s performance reliably.

  • Assigning agents Big Five personalities on seven-point dimensions had a much stronger effect when every trait was set to the profile Vallinder thought most conducive to cooperation. Even then, he did not see Claude-like improvement across generations in GPT-4o and did not think he saw it with Gemini either.

  • Hughes argued that real agents will likely receive objectives such as “make me as much money as you can,” reach the highest game score, buy groceries, or secure a desirable restaurant. They will not generally receive detailed instructions resolving every externality that objective creates.

  • A restaurant agent could reserve every table within three blocks for 7:00 p.m., let its user choose, then cancel the rest. Once replicated, this removes availability and rewards the fastest agent—a concrete example of automation overwhelming institutions whose stability depended on human effort, ethics, and limited parallelism.

10. Mixed societies, communication, and humans are the next stress tests

  • In Vallinder’s preliminary mixed run, generation one contained four agents of each model; later generations added two newcomers from each. Performance was only slightly above GPT-4o alone and declined modestly, plausibly because GPT-4o first exploited cooperation and the others then became less generous.

  • Labenz asked the sharper invasion question: would one or two GPT-4o agents exploit and spoil an established Claude society? Vallinder had not run that test; he hypothesized they would reduce the average slightly but fail to thrive because Claude agents would observe and punish their defection.

  • Planned variants add communication either before agents formulate generational strategies or directly between donor and recipient. Other suggested extensions include group selection, public-goods games, specialization, and classical games such as the prisoner’s dilemma or ultimatum game with the missing evolutionary structure added.

  • Human participation is now technically straightforward because every action is text. Hughes wants to ask whether humans behave differently inside Claude 3.5, GPT-4o, or mixed societies—and whether those populations end differently—providing at least “a noisy signal” about where society could be in five years.

11. Trustworthy institutions matter more than universal altruism

  • Preliminary analysis weakens the strongest interpretation of Claude’s success. Despite benefiting from longer histories, Claude appeared to punish zero donations equally whether they reflected selfish defection or justified enforcement; “sadly but also excitingly,” it may use extra social information without understanding just versus unjust punishment.

  • Cooperation itself is contextual. Vallinder wants agents to reach mutually beneficial agreements, not collude on prices; Labenz similarly questioned whether buyers would choose a self-driving car willing to sacrifice its owner for aggregate welfare. As Hughes put it, “cooperation and collusion” can depend on the beholder.

  • Vallinder’s practical answer was a trustworthy environment supported by standards or regulation. Hughes favored empirical evaluations and rapid feedback over doctrine, citing social-media echo chambers as a system-level effect of serving people more of what they want that was hard to see in advance.

  • Hughes called deployment a “wicked problem”: designers cannot see the solution in advance and may discover the right architecture only partway through deployment. He cited rollbacks when products fail, but warned of “hysteresis”—a social phase transition may require retreating much farther than its original trigger to reverse. Vallinder likewise declined a categorical open-source prescription because thresholds and context make the issue genuinely nuanced.

  • The upside remains substantial. Hughes imagines AI entering the scientific cultural loop—forming hypotheses, testing them with humans, and cooperating massively in parallel—extending the pattern he sees in AlphaFold’s use by probably tens or hundreds of thousands of people toward cancer and climate work. “If we get this right,” he argued, society can tilt agent evolution toward outcomes it actually wants.

  • The research barrier is unusually low: the code is open, experiments can run in Google Colab with API credits, and a new model can be tested by changing an API key. Labenz’s invitation to social scientists was blunt: the scarce input is now the quality of the questions, not command of a 50,000-line engineering stack.