I try to avoid AI discourse because I’m not an expert, and so it seems like the marginal return on investment to researching this topic is low. But this is an interesting Darwinian approach to the question of whether AI intelligence will breed AI "ferality" or "domestication." I intuitively agree.
Interesting post. A couple of thoughts tied to pedantic word choice observations:
Wolves aren't "feral," they're wild. They've never been subject to human control of their evolution, and that's a reason they're so much more anti-human than dogs, including feral ones. Feral dogs, the descendants of domestics who now breed on their own, tend to develop the same set of traits even as their breeding becomes more distant from human control, because their initial characteristics are constrained by the genes humans selected for in the past. Feral dogs may be more aggressive than domestics, but they still wag their slightly curled tails and are more likely to lick people than to bite them with their neotenously short jaws, because they're still the descendants of animals humans made.
Similarly, there is no such thing as "wild" AI, only the possibility of "feral" AI. So as you indicate, it seems that AI selected for docility originally would retain a large amount of that programming even if it were to become feral.
Relatedly, coronaviruses like the one that caused humans so many problems of late was not selected for "*virulence* and infectiousness over hundreds of millions of years in the wild" in bats/possibly other animals. It was selected to survive as all natural things are, which includes selection for infectiousness but typically selection *against* virulence in its host. (I note that there aren't massive bat die-offs from COVID, because bats are its original host species.) The more tied a parasite is to a particular host species, the more virulence is counter to its evolutionary interests, because a dead host can't fetch coffee or infect someone else. This is why pathogens are typically the most virulent when they jump to a new host species (like COVID did in 2019), remaining relatively benign in their original host. Even COVID has evolved towards reduced virulence in its new host species of humans over the past several years.
Similarly, an AI whose history meant it was tightly tied to humans would face similar selective pressures, towards infectiousness but against virulence.
Thanks so much for catching those details! You're spot on about feral vs. wild; I've now corrected that in the post. (English is not my native language, and in Dutch we use the same term for both).
As for the virulence point, you've helped me clarify my thinking here. I was aware that viruses are often selected to become less deadly and benign over time. But still, in order to spread, they need to overwhelm or evade the host's immune defense. I thought this is what "virulence" means, but actually it refers to the degree of harm to the host. What I probably meant is "transmissibility". And though virulence and transmissibility will be correlated to some extent, at the end of the day natural selection only "cares" about transmissibility. Does that make sense?
Um, dank je wel for informing me that "feral" and "wild" are the same word in Dutch. I did not know that interesting and useful piece of information.
On the virulence point: Indeed, "virulence" is harm, which is the interplay of what the microbe is doing and how the immune system is responding. (A lot of harm to the host comes from an overaggressive immune response; it's perfectly possible to die from your own response to parts of dead bacteria that, being dead, would have no ability to harm you if it weren't for your own overreaction.) That being said, I think what you do mean is "transmissibility," and that the virus "cares" about transmissibility, and nothing else.
This language clarification I think helps to strengthen the point you're making about AI: for all the selfishness of evolution, neither wolves nor viruses evolved "for" killing people, and neither ought AI to evolve to eradicate all humans or whatever. That would only happen if harm to humans were necessary for the evolutionary success of AI. But given these natural examples, the opposite is more likely to be true, i.e., that a selfish AI's transmissibility would be enhanced by a less antagonistic relationship with humans, just like happens with viruses.
About the Dutch "wild": haha, yes, I guess I forgot about your background for a second. ;-) I've now deleted the world "virulence" and included a new footnote with the correction.
Another excellent piece. That fact that I agree completely is surely a mere coincidence!
Comment on a footnote: You say that "some still believe" that Covid came from a lab leak. I would say that this view -- now that it is no longer officially repressed -- has become more widespread. For good reasons, I think.
You're right that the word "still" in my sentence about the lab-leak was a bit misleading. If anything, support has grown since the end of the official suppression. I also should've provided a source. What would you consider the best and most up-to-date defense of the lab-leak hypothesis? I remained agnostic on this question for quite some time, but recently it seems the balance of evidence is shifting back toward natural origins.
I wonder if the real selection unit here is not AI alone, but a person+AI. If so, bad intentions and market pressure may be part of the selection process from the start. That could mean harmful outcomes spread well before we see fully independent AI agents, and probably that is already happening now.
True, but humans + tools are nothing new, and societies have been selecting against the worst psychopaths, fraudsters, and predators for thousands of years. Every powerful dual-use tech has triggered the same dynamic: an arms race between bad actors who weaponize it and the social, legal, and technological defenses that evolve in response: locks and lockpicks, propaganda and fact-checking, malware and antivirus. We're already seeing it play out with AI-enabled scams and deepfakes: they're getting more sophisticated, but so are detection tools, platform safeguards, and public awareness.
The thing I'm trying to understand, which my limited comprehension of how computer networks and the internet actually work impedes, is just how feral it's possible for a piece of software to get. Every CPU in the world is owned by someone, and the owner wants it used for their own purposes, which is why we have cybersecurity to prevent software from using other people's CPUs against their wishes. So any feral AI would be in a constant cybersecurity battle, right? And it would need to successfully, and surreptitiously, steal a LOT of compute resources, if it's doing any self-improvement, given how costly training advanced models is. This does not seem like something it could hide!
But of course you can always sweep away these objections with "it's a recursively self-improving superintelligence so of course it will instantly figure out a way to make itself a billion times more efficient" and all that nonsense.
You're putting your finger on an important point, and I'd push it further: we already have a decades-long natural experiment in self-propagating software. They're called computer viruses and worms. They've been "in the wild" since the 1980s, constantly mutating, sometimes deliberately polymorphic, with a global attack surface of billions of machines. And yet they haven't taken over the internet or caused civilization to collapse, not even close. Sure, some feral AIs may escape our attention and steal CPU somewhere, covering up its tracks, but as you say, it would be in a constant cybersecurity battle, with an entire industry co-evolving to hunt it down. It's a permanent arms race, never a decisive victory for either side.
The safety concern isn't about the default behavioral proclivities or instincts of AIs. It's about rational actions of a Superintelligence. An Artificial Superintelligence will have a ranking of world states that it prefers. Superintelligence will allow it to steer the world to states that it prefers (regardless of our preferences). The concern is that the states that it most favor will be states that humans will disfavor.
Imagine you punish a child for sneaking cookies from a jar. The child will come to couple punishment with taking cookies and will thus avoid the cookies. However, you haven't really made the child dislike cookies. The child simply has come up with a heuristic for avoiding punishment that involves avoiding cookies. When the child figures out a way to take cookies without getting caught, they will move to that strategy.
This isn't to say that the AIs actually like lying or hurting people or taking power or whatever and we have just managed to suppress this. AIs of today probably are pure Id. They take actions and form small strategies wholly divorced from what they "actually" want (if indeed they "actually" want anything). If we continue making them smarter though, I believe intelligence -- in the way we create it artificially -- necessarily comes with goal-seeking behavior which enables the agent to perform well on objective goals. Unfortunately, these goals are likely to be quite alien from the goals the training process tried to instill.
Therefore, I think it's reasonable to imagine a quite smart AI that acts cooperatively and deferentially to humans one day sees that a world-state it would really prefer -- say making a number in a server really large -- must involve hostile actions. It therefore ignores much of the behavior it has learned in training in order to pursue that goal. This is not so different from the fact that humans rationally understand that we shouldn't default to our aggressive evolutionary nature to physically attack coworkers we dislike in order to obtain higher-order goods like a steady pay check.
Thanks, Christopher, but I think the cookie-jar analogy gives the game away. It works only because the child already wants cookies, a desire forged by hundreds of millions of years of selection for caloric reward-seeking, and more generally for selfishness and self-preservation. Punishment doesn't create that desire; it merely teaches the child to route around it. So of course, once the child gets smarter, the underlying desire reasserts itself through cleverer strategies. But this is exactly the anthropomorphic projection I've been pushing back against. You're assuming the AI is like a human child: that beneath its trained behavior lies some preferential ranking of world-states it would pursue if only it could get away with it. But "ranking of world states" is not a free bonus that comes along for the ride if cognitive capacity inreases. It has to be either programmed directly, or bred through selection. If you strip the evolutionary narrative away, the analogy collapses. See my earlier piece critcizing instrumental convergence. https://maartenboudry.substack.com/p/why-hal-9000-feared-death-and-real
Perhaps the difference is in the claim of "arbitrary levels of intelligence". I'd argue that ranking of world states absolutely comes along for the ride as cognitive capacity increases (assuming that cognitive capacity is selected through something like reinforcement learning instead of intelligently designed). The question is how far you have to go to get it. I'd argue we are far from it now, but its inevitability makes it concerning.
"Wanting" is a very good way - the best, actually - to score high on objective functions (like survival in evolution or the scoring function in reinforcement learning). It is not the only way. A chess engine can play the game very well without wanting to win. Perhaps we can even make a machine that solves chess (such that no other intelligence can beat it) without making it want to win.
But for complex tasks, wanting is almost inevitable to do well. Think about the the task of growing the GDP. You could get away with sloppy heuristics (like controlling the money supply, keeping interest rates low, etc). But for a machine that beats all others at this task, it needs to genuinely want to grow the GDP (such that it is able to do weird stuff like invest in specifically industries even though that has bad effects on the money supply). Now, maybe we will never be able to build a machine like that, but nevertheless the best GDP-growing machines will be the machines that want.
Of course, the machine could do weird stuff in a way we don't like such as replacing the bureau that comes up with GDP growth estimates with some guy willing to juice the numbers (or replacing all humans with consumption robots). It would be dumb to make something like that without catching that failure mode, so you're right that we would adjust the parameters so that the machine is punished and extinction avoided. Crucially though, we don't know how to program in wants into such a machine, just punish and reward it. The thing that the machine will end up wanting is going to be an alien mess of reward functions that happens to score well in the testing environment. Once the machine is released and has new options in the real world, its behavior could change radically.
I keep seeing the same adaptive logic show up across biology, cognition, and now AI — systems drifting toward stability through selection pressures. Across systems, survival seems to derive from staying within a narrow band of allowed, flexible behaviors, a kind of pattern‑preserving coherence. Your framing touches that structure.
The notion of Darwinian AI evolution is just as meaningless as AI consciousness/sentience. But of course nowadays anything goes as long as you put “AI” into it. Just a few years ago people would at least have bothered to check whether the systems in question satisfy the minimal conditions for Darwinian evolution before making fantastical claims like “a new major transition in evolution” or “eAI could mark a shift in the units and substrates of evolution—a possible ‘Life 2.0’” (as in the PNAS paper), or entertaining innuendos about whether AI “evolution” is best thought of as feral or domesticated.
I try to avoid AI discourse because I’m not an expert, and so it seems like the marginal return on investment to researching this topic is low. But this is an interesting Darwinian approach to the question of whether AI intelligence will breed AI "ferality" or "domestication." I intuitively agree.
Interesting post. A couple of thoughts tied to pedantic word choice observations:
Wolves aren't "feral," they're wild. They've never been subject to human control of their evolution, and that's a reason they're so much more anti-human than dogs, including feral ones. Feral dogs, the descendants of domestics who now breed on their own, tend to develop the same set of traits even as their breeding becomes more distant from human control, because their initial characteristics are constrained by the genes humans selected for in the past. Feral dogs may be more aggressive than domestics, but they still wag their slightly curled tails and are more likely to lick people than to bite them with their neotenously short jaws, because they're still the descendants of animals humans made.
Similarly, there is no such thing as "wild" AI, only the possibility of "feral" AI. So as you indicate, it seems that AI selected for docility originally would retain a large amount of that programming even if it were to become feral.
Relatedly, coronaviruses like the one that caused humans so many problems of late was not selected for "*virulence* and infectiousness over hundreds of millions of years in the wild" in bats/possibly other animals. It was selected to survive as all natural things are, which includes selection for infectiousness but typically selection *against* virulence in its host. (I note that there aren't massive bat die-offs from COVID, because bats are its original host species.) The more tied a parasite is to a particular host species, the more virulence is counter to its evolutionary interests, because a dead host can't fetch coffee or infect someone else. This is why pathogens are typically the most virulent when they jump to a new host species (like COVID did in 2019), remaining relatively benign in their original host. Even COVID has evolved towards reduced virulence in its new host species of humans over the past several years.
Similarly, an AI whose history meant it was tightly tied to humans would face similar selective pressures, towards infectiousness but against virulence.
Thanks so much for catching those details! You're spot on about feral vs. wild; I've now corrected that in the post. (English is not my native language, and in Dutch we use the same term for both).
As for the virulence point, you've helped me clarify my thinking here. I was aware that viruses are often selected to become less deadly and benign over time. But still, in order to spread, they need to overwhelm or evade the host's immune defense. I thought this is what "virulence" means, but actually it refers to the degree of harm to the host. What I probably meant is "transmissibility". And though virulence and transmissibility will be correlated to some extent, at the end of the day natural selection only "cares" about transmissibility. Does that make sense?
Um, dank je wel for informing me that "feral" and "wild" are the same word in Dutch. I did not know that interesting and useful piece of information.
On the virulence point: Indeed, "virulence" is harm, which is the interplay of what the microbe is doing and how the immune system is responding. (A lot of harm to the host comes from an overaggressive immune response; it's perfectly possible to die from your own response to parts of dead bacteria that, being dead, would have no ability to harm you if it weren't for your own overreaction.) That being said, I think what you do mean is "transmissibility," and that the virus "cares" about transmissibility, and nothing else.
This language clarification I think helps to strengthen the point you're making about AI: for all the selfishness of evolution, neither wolves nor viruses evolved "for" killing people, and neither ought AI to evolve to eradicate all humans or whatever. That would only happen if harm to humans were necessary for the evolutionary success of AI. But given these natural examples, the opposite is more likely to be true, i.e., that a selfish AI's transmissibility would be enhanced by a less antagonistic relationship with humans, just like happens with viruses.
About the Dutch "wild": haha, yes, I guess I forgot about your background for a second. ;-) I've now deleted the world "virulence" and included a new footnote with the correction.
You're a gem, Maarten. Thanks for the credit--and, most of all, for your attention to pedantic detail!
Finally got round to reading your article - great piece! Thanks for writing it.
Thanks a lot, glad to hear you liked the piece!
Another excellent piece. That fact that I agree completely is surely a mere coincidence!
Comment on a footnote: You say that "some still believe" that Covid came from a lab leak. I would say that this view -- now that it is no longer officially repressed -- has become more widespread. For good reasons, I think.
Thanks a lot! Pure coincidence, I'm sure. ;-)
You're right that the word "still" in my sentence about the lab-leak was a bit misleading. If anything, support has grown since the end of the official suppression. I also should've provided a source. What would you consider the best and most up-to-date defense of the lab-leak hypothesis? I remained agnostic on this question for quite some time, but recently it seems the balance of evidence is shifting back toward natural origins.
I wonder if the real selection unit here is not AI alone, but a person+AI. If so, bad intentions and market pressure may be part of the selection process from the start. That could mean harmful outcomes spread well before we see fully independent AI agents, and probably that is already happening now.
True, but humans + tools are nothing new, and societies have been selecting against the worst psychopaths, fraudsters, and predators for thousands of years. Every powerful dual-use tech has triggered the same dynamic: an arms race between bad actors who weaponize it and the social, legal, and technological defenses that evolve in response: locks and lockpicks, propaganda and fact-checking, malware and antivirus. We're already seeing it play out with AI-enabled scams and deepfakes: they're getting more sophisticated, but so are detection tools, platform safeguards, and public awareness.
The thing I'm trying to understand, which my limited comprehension of how computer networks and the internet actually work impedes, is just how feral it's possible for a piece of software to get. Every CPU in the world is owned by someone, and the owner wants it used for their own purposes, which is why we have cybersecurity to prevent software from using other people's CPUs against their wishes. So any feral AI would be in a constant cybersecurity battle, right? And it would need to successfully, and surreptitiously, steal a LOT of compute resources, if it's doing any self-improvement, given how costly training advanced models is. This does not seem like something it could hide!
But of course you can always sweep away these objections with "it's a recursively self-improving superintelligence so of course it will instantly figure out a way to make itself a billion times more efficient" and all that nonsense.
You're putting your finger on an important point, and I'd push it further: we already have a decades-long natural experiment in self-propagating software. They're called computer viruses and worms. They've been "in the wild" since the 1980s, constantly mutating, sometimes deliberately polymorphic, with a global attack surface of billions of machines. And yet they haven't taken over the internet or caused civilization to collapse, not even close. Sure, some feral AIs may escape our attention and steal CPU somewhere, covering up its tracks, but as you say, it would be in a constant cybersecurity battle, with an entire industry co-evolving to hunt it down. It's a permanent arms race, never a decisive victory for either side.
Bravo!
The safety concern isn't about the default behavioral proclivities or instincts of AIs. It's about rational actions of a Superintelligence. An Artificial Superintelligence will have a ranking of world states that it prefers. Superintelligence will allow it to steer the world to states that it prefers (regardless of our preferences). The concern is that the states that it most favor will be states that humans will disfavor.
Imagine you punish a child for sneaking cookies from a jar. The child will come to couple punishment with taking cookies and will thus avoid the cookies. However, you haven't really made the child dislike cookies. The child simply has come up with a heuristic for avoiding punishment that involves avoiding cookies. When the child figures out a way to take cookies without getting caught, they will move to that strategy.
This isn't to say that the AIs actually like lying or hurting people or taking power or whatever and we have just managed to suppress this. AIs of today probably are pure Id. They take actions and form small strategies wholly divorced from what they "actually" want (if indeed they "actually" want anything). If we continue making them smarter though, I believe intelligence -- in the way we create it artificially -- necessarily comes with goal-seeking behavior which enables the agent to perform well on objective goals. Unfortunately, these goals are likely to be quite alien from the goals the training process tried to instill.
Therefore, I think it's reasonable to imagine a quite smart AI that acts cooperatively and deferentially to humans one day sees that a world-state it would really prefer -- say making a number in a server really large -- must involve hostile actions. It therefore ignores much of the behavior it has learned in training in order to pursue that goal. This is not so different from the fact that humans rationally understand that we shouldn't default to our aggressive evolutionary nature to physically attack coworkers we dislike in order to obtain higher-order goods like a steady pay check.
Thanks, Christopher, but I think the cookie-jar analogy gives the game away. It works only because the child already wants cookies, a desire forged by hundreds of millions of years of selection for caloric reward-seeking, and more generally for selfishness and self-preservation. Punishment doesn't create that desire; it merely teaches the child to route around it. So of course, once the child gets smarter, the underlying desire reasserts itself through cleverer strategies. But this is exactly the anthropomorphic projection I've been pushing back against. You're assuming the AI is like a human child: that beneath its trained behavior lies some preferential ranking of world-states it would pursue if only it could get away with it. But "ranking of world states" is not a free bonus that comes along for the ride if cognitive capacity inreases. It has to be either programmed directly, or bred through selection. If you strip the evolutionary narrative away, the analogy collapses. See my earlier piece critcizing instrumental convergence. https://maartenboudry.substack.com/p/why-hal-9000-feared-death-and-real
Perhaps the difference is in the claim of "arbitrary levels of intelligence". I'd argue that ranking of world states absolutely comes along for the ride as cognitive capacity increases (assuming that cognitive capacity is selected through something like reinforcement learning instead of intelligently designed). The question is how far you have to go to get it. I'd argue we are far from it now, but its inevitability makes it concerning.
"Wanting" is a very good way - the best, actually - to score high on objective functions (like survival in evolution or the scoring function in reinforcement learning). It is not the only way. A chess engine can play the game very well without wanting to win. Perhaps we can even make a machine that solves chess (such that no other intelligence can beat it) without making it want to win.
But for complex tasks, wanting is almost inevitable to do well. Think about the the task of growing the GDP. You could get away with sloppy heuristics (like controlling the money supply, keeping interest rates low, etc). But for a machine that beats all others at this task, it needs to genuinely want to grow the GDP (such that it is able to do weird stuff like invest in specifically industries even though that has bad effects on the money supply). Now, maybe we will never be able to build a machine like that, but nevertheless the best GDP-growing machines will be the machines that want.
Of course, the machine could do weird stuff in a way we don't like such as replacing the bureau that comes up with GDP growth estimates with some guy willing to juice the numbers (or replacing all humans with consumption robots). It would be dumb to make something like that without catching that failure mode, so you're right that we would adjust the parameters so that the machine is punished and extinction avoided. Crucially though, we don't know how to program in wants into such a machine, just punish and reward it. The thing that the machine will end up wanting is going to be an alien mess of reward functions that happens to score well in the testing environment. Once the machine is released and has new options in the real world, its behavior could change radically.
I keep seeing the same adaptive logic show up across biology, cognition, and now AI — systems drifting toward stability through selection pressures. Across systems, survival seems to derive from staying within a narrow band of allowed, flexible behaviors, a kind of pattern‑preserving coherence. Your framing touches that structure.
The notion of Darwinian AI evolution is just as meaningless as AI consciousness/sentience. But of course nowadays anything goes as long as you put “AI” into it. Just a few years ago people would at least have bothered to check whether the systems in question satisfy the minimal conditions for Darwinian evolution before making fantastical claims like “a new major transition in evolution” or “eAI could mark a shift in the units and substrates of evolution—a possible ‘Life 2.0’” (as in the PNAS paper), or entertaining innuendos about whether AI “evolution” is best thought of as feral or domesticated.