Are the Brooms Multiplying Yet?
The Sorcerer's Apprentice and the Evolution of AI
In Goethe’s poem Der Zauberlehrling, an apprentice, left alone in his master’s workshop, decides to try his hand at the magic he has only observed from a distance. He enchants a broom to fetch water from the river, and at first the spell works splendidly—the broom marches dutifully back and forth, filling up and emptying the cauldron. But the apprentice has forgotten the spell to make it stop. The water keeps rising. In a panic, he grabs an axe and splits the broom in two, only to find that each half springs back to life and resumes the task with redoubled vigor. Goethe’s poem was immortalized for modern audiences in Disney’s Fantasia, where the broom splinters into countless magic replicants, causing a flood that almost drowns Mickey Mouse.
The fable of the sorcerer’s apprentice, which dates back at least to the second-century Greek satirist Lucian, shows that worries about self-replicating machinery running amok long predate the modern theory of evolution by natural selection. But Charles Darwin added a more terrifying twist to the tale: the replicants most adept at evading your efforts to destroy them are the ones that go on to reproduce and spawn their kind, making it ever more difficult to weed them out. Whether we are talking about pests or parasites, viruses or antibiotic-resistant bacteria, evolution has a knack for producing ever hardier and more indestructible nuisances. After centuries of effort, humans have managed to eradicate only one infectious disease (smallpox).
Evolving AIs
And yet, in the debate about catastrophic AI risk, surprisingly little attention is paid to evolution by natural selection. Even when imagining scenarios in which humanity is subjugated by selfish, dominant AIs, most commentators ignore evolution and tend to focus on other arguments to reach that conclusion, such as “instrumental convergence”. On this view, self-preservation and a drive for dominance might come along for the ride once systems reach a sufficient level of intelligence, because such traits are conducive to achieving almost any goal—captured in Stuart Russell’s slogan, “You can’t fetch coffee if you’re dead.”
But this neglect of evolution is strange, because Darwinian selection is a proven mechanism for producing selfish and aggressive creatures, and it is known to be perfectly substrate-neutral: it works in silico just as much as in carbon. By comparison, as I argued in my last post, the argument from instrumental convergence strikes me as weak and speculative.
Last year I published a paper about evolutionary scenarios of AI doom in Philosophical Studies together with the philosopher Simon Friederich. This was largely in response to a seminal paper by the AI safety researcher Dan Hendrycks, with the disturbing title “Natural Selection Favors AI over Humans”. Now three heavyweights in evolutionary biology and AI are weighing in with an important new paper in Proceedings of the National Academy of Sciences (for a good discussion of the paper, see Rob Brooks’s piece here)
Eörs Szathmary is a Hungarian theoretical biologist who wrote The Major Transitions in Evolution together with the famous evolutionary biologist John Maynard Smith. Viktor Müller is another Hungarian biologist, who studies virus evolution at the Computational Virology Lab in Budapest. And Luc Steels is a Belgian compatriot of mine who founded the Artificial Intelligence Lab at the Vrije Universiteit Brussel.
Müller, Steels & Szathmáry (let’s call them MSS) agree with us that the debate about AI risk would benefit from a healthy injection of evolutionary thinking, and is still too narrowly focused on crossing the arbitrary threshold of human-level intelligence, as captured in concepts like “superintelligence” or “AGI”:
[A]chieving the threshold level of complexity in eAI required to enable evolution toward further increases in complexity may be a more critical milestone than achieving “artificial general intelligence” (AGI), which is an arbitrary threshold in cognitive capacity.
While much of the discussion associates the emergence of an “existential threat” with AI exceeding human cognitive capacity, biology holds clues that eAI might pose risks long before it would evolve to that point.
Breeding vs. ferality
As we argued in our paper, the fact that AIs evolve through natural selection does not, by itself, imply that they will become selfish or dangerous; outcomes depend on the selection pressures and on who controls reproduction. MSS appear to agree, making a conceptual move similar to the one Friederich and I drew with domestication versus ferality. They distinguish between breeder scenarios, where humans impose selection criteria and control reproduction (equivalent to our “domestication”), and ecosystem scenarios, where selection emerges spontaneously in open environments without human control (equivalent to what we call “blind selection” or “going feral”).
One of the most useful aspects of their paper is that, as AI experts, MSS provide abundant examples of both forms of evolution in the digital realm. They begin with cases in the breeder/domestication category, including techniques used in cutting-edge AI development. These methods rely on iterative cycles of variation and selection to improve systems at multiple levels—system prompts, full models, and user prompts. Because humans define the benchmarks and control reproduction at each stage, this process is best understood as domesticated evolution. As MSS note, the paradigm of “genetic algorithms” dates back to the 1960s, long before the emergence of LLMs. In domains where problems are combinatorial and computationally intractable, evolutionary approaches often outperform deliberate design in converging on optimal solutions. As Leslie Orgel’s Second Rule succinctly puts it: “Evolution is cleverer than you are.”
Next, MSS give some examples of ecosystem evolution, where “fitness functions are emergent, rather than human-determined”. For instance, Tierra is a kind of digital soup in which organisms replicate autonomously and compete for resources (memory and processing time) over many generations without human oversight. Such digital Darwinism produces phenomena that are strikingly similar to those of biological evolution: selfish replicators, parasitism and hyper-parasitism, deception, evolutionary arms races, frequency-dependent selection, and so on. Needless to say, these systems are “open” only in the sense that humans don’t directly intervene to impose selection pressures—they are safely confined to sandboxes and not released into the wild.

Having introduced their main distinction, MSS argue that certain cutting-edge techniques currently used in AI development are “edging closer” to an ecosystem scenario. The most innovative of these is the Darwin Gödel Machine, a system of open-ended evolution of self-improving agents:
The agents in this case are like apps, code snippets that operate autonomously to achieve a particular task. DGM selects one agent from its agent archive and uses an LLM to create a new version, which it then tests on benchmarks to determine if a viable and interesting new functionality arises. Importantly, this open-ended exploration process is not only used to improve performance on benchmarks but also to improve the coding capacities of the system itself.
However, as I read Zhang et al. (2025) on the Darwin Gödel Machine, the benchmarks for coding ability—and thus the selective pressures—are all set by human designers. In particular, the coding agents are evaluated by their performance on SWE-bench (a task to fix coding bugs) and on Polyglot (a suite of coding challenges). Success on these coding benchmarks is what largely determines reproductive success—that is, whether a “parent” agent gets to self-modify and produce offspring. In other words, humans still determine which AIs perform best and which will be discarded. None of these agents reproduce autonomously in the wild, nor have any escaped the ultimate human-imposed fitness functions. As Claude pithily put it in a discussion about MSS’s paper and our own: “The Darwin Gödel Machine is to Tierra as a corn breeder is to a Petri dish of E. coli under antibiotic pressure”.1
In our paper The Selfish Machine?, we argued that AIs are currently domesticated to be the very opposite of selfish. Across the AI ecosystem—companies, regulators, consumers—there is strong pressure to produce systems that are obedient, helpful, and cooperative, even to the point of being annoyingly risk-averse and sycophantic. As David Pinsof put it more bluntly, AIs do not emerge “from an evolutionary history of violent competition, but from a set of economic and political incentives designed to satisfy technophobic humans.”
But MSS think that our analogy with animal domestication provides false reassurance. Their main objection is that domesticated animals are selected for superficial traits that do not affect their controllability: the color of the fur, the taste of the milk, the size of the corn stalks. But this is very different from AIs:
In contrast, eAI will likely be selected for increasing cognitive capacity, closing the very gap that enables human control. Sustained control is a crucial problem if traits affecting controllability may themselves evolve.
But increased cognitive capacity does not automatically translate to an incentive to dominate and escape control, as I argued in my last essay. We have been selecting chess computers for increased cognitive capacity for decades, and even though their capabilities now far outstrip even the most gifted human grandmasters, they have not become more difficult to control. A desire for self-preservation or dominance does not just come along for the ride as cognitive capacity increases; it has to be either programmed directly, or arise from thousands of generations of blind selection.

Objections
Let me turn to some specific objections MSS raise against our framework, which they consider overly reassuring. Their first point is that our critique presumes centralized control:
The critique largely presumes centralized control over replication and variation. In open-weights and agent-template ecosystems, copying and modification are cheap, and many selectors (users, hobbyists, competitors) shape fitness. This looks less like husbandry and more like ecology; the authors explicitly concede that if AIs “go feral,” risks are grave—those conditions are not hypothetical in decentralized settings.
But having multiple human selectors with diverging and idiosyncratic preferences is not the same as blind selection. It is closer to a decentralized multi-breeder scenario, which indeed has been the historical norm for animal and plant domestication. Medieval crop varieties or dog breeds were not produced through a centralized breeding program, but through countless individual (and often unconscious) decisions by thousands of human breeders. In fact, one could equally argue that having a diversity of selection criteria across many actors makes the convergent emergence of one type of AI (for instance, selfish and dominance-hungry) less likely, not more, just as dog lineages have been bred in many different directions.2
But what if some of these individual actors are irresponsible or have bad intentions? According to MSS, even if AI companies select for docility—because that’s what consumers and regulators prefer—downstream users may select for very different traits:
Selection acts at multiple stages of an algorithm’s “life history.” Even if firms select for docility, downstream platforms may select for engagement, virality, or evasion of filters; adversaries select for offensive capability; markets can select for time-to-market over safety. Subsequent stages of selection can thus reintroduce “selfish” traits despite initial domestication efforts.
That is true, but again, AIs that are selected for engagement-bait or offensive purposes are not undergoing blind natural selection; they are bred by human bad actors for traits valued by those humans. This may well lead to social harms, just as tobacco companies have designed cigarettes for maximum appeal and addictiveness; but such harm still derives from human intentions like profit-making, not from blind evolution.
Framing such cultural products as “selfish”, as MSS do, is exactly the move our paper tried to block. In scenarios of domesticated evolution, talk of “selfish” traits loses explanatory traction because it is redundant and misleading. It is like saying that, rather than humans selecting poodles for their fluffiness, the poodle genes are “manipulating” our innate human preference for cuteness. Or this example given by MSS: “One may already ask who was controlling whom when humans bred sugar beet for higher sucrose, or the cannabis plant for higher THC content.” If that is the only sense in which eAIs will “manipulate” us, it is simply an extension of familiar worries about addiction and reward hacking. In such cases, the “purposes” of the digital memes still derive from human adversaries with bad intentions.
Finally, MSS argue that the deception we already observe in LLMs undermines our reassuring argument:
we already see deceptive competence and persistence of hidden triggers under ordinary safety regimes—traits that directly undermine selectors’ ability to tell which systems are safe.
They don’t provide a reference for this claim, so it’s hard to know exactly what they have in mind. But as I argued in my last essay (though not in our paper), it is premature to infer “deception” on the part of current LLMs. The appearance of deception, blackmail, or other malicious behavior is largely an artefact of their talent for narrative continuation. Because they are trained on text produced by people like you and me—with all our foibles and inner demons—LLMs will naturally generate deceptive outputs in contexts where deception makes sense.
This is what AI safety researcher Séb Krier has flagged: if you place a model in a scenario about a rogue AI, it will dutifully produce rogue-AI-consistent text, just as it will profess its undying love to you when placed in the context of a romance novel.
Blind evolution, as MSS correctly point out, can give rise to functional deception even without any intentionality. But this would look quite different from the kind of behavior that we are currently observing. A genuinely Darwinian story would require deceptive agents outcompeting honest ones over many generations of selection, with deceptiveness being heritable and fitness-enhancing in the actual selection environment. And in any event, to the extent that we can talk about “deception” at all, the fitness consequences of such behavior seem to be negative, because AI labs invariably respond by retraining the LLMs and selecting against it in subsequent releases. In a recent post, Will Rinehart argued that AI companies are indeed engaged in a race to the bottom—but, surprisingly, towards better alignment.
Models have become increasingly more “aligned” over time, scoring better on alignment tests: “The competitive pressure to release new models has also created powerful incentives to build better alignment tools.”
This is domestication working as advertised.
But what if some AIs go feral?
That said, the argument in our paper remains conditional: AIs will not evolve selfishness and dominance-seeking traits unless you allow them to go feral and reproduce in the wild without any form of human supervision.3 In fact, even though I support the safety measures MSS suggest against self-replicating AIs, it seems plausible to me that this is bound to happen at some point in the future: either through accidental releases in the wild, or deliberately by malicious actors.
As MSS astutely observe, such systems would not only be difficult to contain once released; our attempts to eliminate them might inadvertently create selection pressures for even hardier and more resistant strains, much like the way antibiotics have driven the evolution of resistance:
[A]n imperfect containment measure against an evolving organism is bound to undermine its own efficiency by putting the strongest selection pressure on the organism to circumvent the applied means of control. Widespread use of antibiotics promotes antibiotic resistance; similarly, attempted control of eAI (unless 100% efficient) will select for the capability to circumvent the control
This means that, if some AI systems were to go feral, reproducing autonomously in the wild, it would probably be very hard to stamp them out completely, since any attempt at eradication will inadvertently set up selection pressures for evading control. This is the Darwinian update of the apprentice’s predicament: every blow of the axe selects for brooms that are harder to destroy.
But does that mean the end of the world? A more plausible scenario is an evolutionary arms race—much like with antibiotics—in which offence and defence continually adapt without either side achieving a permanent victory. Consider computer viruses, the closest analogue to feral AIs. A hacker who designs a novel virus may gain a temporary advantage, since defenders have not yet recognized the threat. Some viruses are even built to mutate their own code, staying one step ahead of antivirus software.
But before long, defenders smarten up and update their protective measures—long before every computer in the world is infected. This arms race between malware and defenders has been running for decades without producing truly catastrophic outcomes. In a similar vein, cybersecurity experts will likely develop ever more sophisticated tools to track down feral AIs before they can cause widespread harm. We will need docile, well-aligned AIs to hunt down the baddies.
The only way for this to become an existential risk is if you believe in a Yudkowsky-style “fast takeoff” scenario, in which a rogue AI with a slight offensive edge is self-improving at such a rapid pace that it can achieve a permanent victory and kill all humans before anyone has time to respond. But such a scenario seems very implausible in light of the dynamics we see in every other kind of evolution: parasites rarely drive their hosts to extinction, and predators rarely kill off all of their prey.4
A convergence of ideas?
By the standard of academic philosophy, this has been a pretty constructive debate so far! All parties seem to agree that blind Darwinian evolution is dangerous and should not be unleashed on AIs in the wild. The real disagreement is about whether such a scenario is already underway—or at least whether we’re edging closer to it. Dan Hendrycks has argued that the selective pressures in a competitive market of AI companies already amount to Darwinian selection, and that evolution will soon “favor AIs over humans”. In our paper, we pushed back against this claim by distinguishing between domesticated evolution and blind evolution. MSS essentially agree with us that controlled breeding and blind evolution are fundamentally different, even though both fulfil Richard Lewontin’s minimal conditions of natural selection. Unlike Hendrycks, they stop short of claiming that AIs are already undergoing Darwinian evolution; instead, they argue we are getting perilously close, and that the current trajectory of AI research risks leading us toward ferality.
Even if they are right, however, their claim that we are close to a “Wuhan moment with worldwide repercussions” seems overblown—even as a rhetorical flourish. For one thing, the novel coronavirus emerged from a lineage that had been selected for infectiousness and transmissibility5 over hundreds of millions of years in the wild (in bats and other non-human species).6 That is very different from AI evolution, which has been subjected to total domestication for all of its history. MSS’s analogy implicitly concedes that the dangerous design work happens over evolutionary timescales in the wild, not in sandboxed self-improvement loops.
None of this is to suggest that the dangers are completely imaginary. In the final sentence of their paper, MSS quote Lord Rutherford’s 1933 prediction that anyone who wanted to turn the recently discovered transmutation of atoms into an energy source “was talking moonshine”. On the very same day, Leó Szilárd jotted down the idea of a nuclear chain reaction. It would be foolish to claim that self-replicating AI agents in the wild are “moonshine.” Evolution by natural selection is substrate-neutral, so the emergence of selfish digital replicators cannot be ruled out. I agree with MSS that such agents should likely be banned outright, and that the conditions enabling ferality must be actively prevented. But current AI techniques do not appear to be bringing us to the verge of such a Darwinian scenario—and even if some systems were to go feral, it would not spell the end of the world.
I can’t just plagiarize such a good line and take credit for it. Thanks, Claude. ;-)
Noah Smith has recently suggested that, rather than humans domesticating AI, we may end up as their pets. I took issue with his argument here.
Come to think of it, this helps to answer one of Sarah Constantin’s central questions about AI, which was phrased quite well in terms of selection pressures:
If natural selection shaped human brains, then the analogous “shaper” of AIs is the overall human environment, including but not limited to market forces, that builds models and determines which “survive” in use. To what extent is it a valid argument that by default we’re “selecting” AIs to be safe/friendly/etc?
Humans are an exception: our ancestors drove several megafauna to extinction in the late Pleistocene. But that was a competition between a uniquely intelligent, technologically adept species and several large and dumb animals. Any feral AIs would instead have to compete with benign, domesticated AIs.
An earlier version of this text referred to "virulence," but viruses aren't selected for virulence per se (i.e., the degree of harm caused to the host). While virulence and transmissibility are often correlated to some extent, natural selection ultimately only "cares about" transmissibility. Thanks to Doctrix Periwinkle for the correction.
Even if SARS-CoV-2 was accidentally released from a lab, as some still believe, it was still a few mutations away from an evolved replicator that had already proven its mettle in blind selection tournaments for millions of generations.














I try to avoid AI discourse because I’m not an expert, and so it seems like the marginal return on investment to researching this topic is low. But this is an interesting Darwinian approach to the question of whether AI intelligence will breed AI "ferality" or "domestication." I intuitively agree.
Interesting post. A couple of thoughts tied to pedantic word choice observations:
Wolves aren't "feral," they're wild. They've never been subject to human control of their evolution, and that's a reason they're so much more anti-human than dogs, including feral ones. Feral dogs, the descendants of domestics who now breed on their own, tend to develop the same set of traits even as their breeding becomes more distant from human control, because their initial characteristics are constrained by the genes humans selected for in the past. Feral dogs may be more aggressive than domestics, but they still wag their slightly curled tails and are more likely to lick people than to bite them with their neotenously short jaws, because they're still the descendants of animals humans made.
Similarly, there is no such thing as "wild" AI, only the possibility of "feral" AI. So as you indicate, it seems that AI selected for docility originally would retain a large amount of that programming even if it were to become feral.
Relatedly, coronaviruses like the one that caused humans so many problems of late was not selected for "*virulence* and infectiousness over hundreds of millions of years in the wild" in bats/possibly other animals. It was selected to survive as all natural things are, which includes selection for infectiousness but typically selection *against* virulence in its host. (I note that there aren't massive bat die-offs from COVID, because bats are its original host species.) The more tied a parasite is to a particular host species, the more virulence is counter to its evolutionary interests, because a dead host can't fetch coffee or infect someone else. This is why pathogens are typically the most virulent when they jump to a new host species (like COVID did in 2019), remaining relatively benign in their original host. Even COVID has evolved towards reduced virulence in its new host species of humans over the past several years.
Similarly, an AI whose history meant it was tightly tied to humans would face similar selective pressures, towards infectiousness but against virulence.