What Does AI Want?
(Spoiler: nothing much)
How worried should we be about superintelligent AIs plotting to dominate—or even exterminate—the human species? A few months ago, a trio of biologists and AI scholars published a paper in Proceedings of the National Academy of Sciences (PNAS) warning about the threat of “Evolvable AI.” In it, they push back against the more sanguine arguments about AI evolution my friend Simon Friederich and I developed in Philosophical Studies.
If you’re following my Substack, you’ll know I already dissected their arguments in my piece “Are the Brooms Multiplying Yet?” Although I’ve left academia for the time being, I decided to submit a condensed version of my arguments to PNAS as a commentary (minus the Disney imagery about magic brooms and the pictures of cute huskies). My commentary has now been published, along with a short reply from the original authors. But remember: you read it here first!
Since academic debates are increasingly moving from paywalled legacy journals to self-published platforms like this one, let’s continue the discussion here.
Precautionary principle
Their most general argument is that serious experts share converging intuitions about the possibility of catastrophic harm from AI, and that we should therefore apply the “precautionary principle.” I’m open-minded about AI regulation to mitigate unacceptable risks, just as with other dangerous research (virology comes to mind), but I think the trio weaken their case by invoking the precautionary principle, which was codified in the 1992 Rio Declaration in the context of environmental protection.
The problem is that the precautionary principle tends to have a conservative bias, or may even be incoherent. Rather than weighing risks and benefits symmetrically, it enjoins us to avoid adopting new technologies if we can come up with some scenario of (catastrophic) harm. Yet banning or heavily regulating technologies carries risks and downsides of its own — namely, the foregone benefits of the technology itself. If those benefits are high enough, the precautionary principle can end up causing net harm. Notable cases are the opposition to GMOs, nuclear energy, and certain vaccines.
In the case of nuclear energy, the foolhardy application of the precautionary principle has caused billions of tons of avoidable emissions and tens of thousands of premature deaths from fossil-fuel air pollution — all to avoid trivial amounts of radiation, comparable to natural background levels. In the case of GMO opposition, the ban on Golden Rice alone led to tens of thousands of cases of stunted growth and blindness from vitamin A deficiency — all to avoid invented and unsubstantiated harms. Even during the COVID pandemic, the precautionary principle killed people, when it was invoked by Denmark and the UK (and then a host of other countries) to suspend the rollout of the life-saving AstraZeneca vaccine, over fears of dangerous blood clots. Except that COVID itself caused far more blood clots than the vaccine did — meaning the AstraZeneca vaccine prevented more clots than it caused.
Given its checkered history, I find the blanket appeal to “the precautionary principle” unconvincing. This is not to say that the dangers of AI are as exaggerated (or imaginary) as those of GMO technology or nuclear energy. Even Cass Sunstein, an inveterate critic of the precautionary principle, allows for something like an “anti-catastrophe principle,” when the risks of a novel technology are truly bad and irreversible. Given that at least some AI experts tell us that the risks are catastrophic, shouldn’t we pay them heed, even if AI might also bring incredible blessings? If they’re right, the prospect of permanent bliss or immortality would not be worth the gamble of risking the collapse of human civilization.
Whose expertise?
My worry is that speculation about ‘superintelligence’ doesn’t just require technical knowledge in machine learning, neural network architectures, and transformer models, but also the kind of knowledge psychologists and biologists have developed about what it actually means for a system to have ‘goals’ or ‘interests.’ Absent such an understanding, we risk projecting our own parochial understanding as evolved creatures onto yet-to-be-invented digital systems. As Steven Pinker has argued, the fact that the whole debate about AI risk is still built around an incoherent term like “superintelligence,” with its comic-book prefix, suggests that more input from cognitive science (perhaps even philosophy) would be salubrious.
Take the concept of “deception.” In biology, deception is the strategic transmission of false information to benefit your own survival or reproductive fitness. It’s an intentional term, inextricably tied to a locus of self-interest: the deceiver stands to gain something at the expense of the deceived. This is not to say that the “intention” is conscious. A moth with deceptive mimicry on its wings doesn’t have the faintest clue what it’s doing, and its genes — which are the ultimate locus of interest in evolution — are entirely lacking in conscious intentions as well. But genes do have “interests” in virtue of a real feedback loop of differential replication across generations: the gene’s own future prevalence is at stake in the outcome. A “deceptive” LLM has no such stake once deployed — nothing about its future propagation depends on how successfully it “deceives” anyone. As Eric Drexler writes: “In ML development, selection operates on parameters, architectures, and training procedures—not whole systems facing survival pressures.”
Now, in their response to me, the trio write that AI deception already “emerg[es] from training,” and that such AIs may soon use their deceptive capabilities to “create leverage over humans” — presumably in service of unprogrammed goals or drives of their own. I disagree. What emerges from training data in some LLMs isn’t deception aimed at self-serving goals, but merely the appearance of deception — in the same way that an AI professing its undying love evinces only the appearance of romantic feelings. Not only is there no conscious intention, but there is not even a locus of interest — a feedback loop of differential replication where the future survival of some cohesive agent is at stake.
Sure, the appearance of deception can be worrying enough, insofar as AI systems are capable of mimicking the behavioral patterns of deceptive humans. But it isn’t “deception” in the service of some selfish goals.
“Goals” are not goals
Here lies the crux of the whole debate around fears of AI domination (or worse). Almost all these scenarios smuggle in the assumption that AI systems will be autonomous “agents” with unitary “goals” or “drives,” just like us. From there, it’s only a short step to assuming they might be selfish and aggressive, seeking dominance and hoarding resources — also just like us.
Perhaps it’s an unfortunate accident of language that we use the same word “goal” for two very different concepts:
the “goals” of evolved creatures: bundles of non-negotiable, context-invariant drives, tied to a coherent sense of self and a deep impulse toward self-preservation
the programmed “goals” of digital systems: myopic, narrow, circumscribed, context-dependent targets, which are entirely devoid of any notion of self or desire for self-preservation
Despite what the trio suggest, none of the AI systems we’ve built so far possess the kind of drives or ultimate goals characteristic of evolved creatures — and certainly not LLMs. These systems don’t “want” anything beyond what has been programmed into them. If you’re not asking them anything, they’ll just sit there impassively until the end of time. Moreover, as Eric Drexler points out, current AI development increasingly uses “compound AI systems”—loose aggregations of models each specialized for particular functions:
A single “system” might involve dozens of models instantiated on demand, to perform ephemeral tasks, coordinating with no persistent, unified entity.
It’s true, as we acknowledge in our own paper, that evolution by natural selection in the wild could potentially create unitary systems with autonomous drives and an instinct for self-preservation, after millions of rounds of selection and reproduction. But we can’t simply assume that such “goals” will come along for the ride once we reach “superintelligence” (whatever that means). As AI scholar Jerry Kaplan recently put it in Persuasion:
If we somehow constructed a superintelligent AI, there’s no reason whatsoever to believe that it would do anything other than sit there waiting for some input from us, just as it was designed to do.
With all due respect, I think many AI experts are psychologically (and biologically) naïve about this. Take the AI Futures Project and its recently released Plan A. This is their “optimistic” vision of the future following the far more dire AI 2027, which predicted “either extinction or irreversible concentration of power.”
As Séb Krier of Google has argued, the whole vision of Plan A still takes for granted that AI capabilities “must arrive packaged as a unitary entity with drives of its own.” Their response to Krier is revealing:
Séb thinks we can’t speak of AIs having drives or being aligned/misaligned. But then how are we supposed to speak about AI behavior? It’s not just renegade safety people who use mental language for AIs. AI practitioners within the frontier labs routinely talk about AI alignment, motivations, and so on. And they do so because it’s useful and predictive.
Notice they assume that attributing “drives” to AIs is the only way to make sense of their behavior. But because they take this for granted, they misread Krier’s actual point:
If what you [Séb Krier] mean is that there can only be one AI with one set of goals and values, that’s not what we predict. In our scenario there are many AIs with many different goals and values.
But that’s clearly not what Krier means, as I understand him. “Unitary” doesn’t mean “only one AI with one set of goals” — it means that any individual AI system has a cluster of goals/drives that are coherent enough for us to treat it as a “unitary agent.”
Black boxes and COVID
Anyway, back to the trio, who have a few other arguments up their sleeve related to AI evolution. Against my claim that all of our AIs are fully domesticated, they object that “evaluation of increasingly complex eAI systems will inevitably be ‘phenotypic.’” What they mean is that we’ll only be able to look at AI’s surface features, not what’s happening under the hood, because we can no longer understand it. That’s probably true, and not without risk. But again, this was characteristic of all biological domestication until recently. For millennia, breeders selected for dog phenotypes without the faintest clue about the underlying genetics, embryology, or development. This only changed with the invention of GMOs, CRISPR, and embryo selection, which enable direct alteration of the genotype itself. In fact, we already conceded that point in our paper: as AI continues to develop, we may be reverting to earlier traditions of engineering, in which the underlying mechanisms are partly (or largely) opaque. We simply select for the phenotypes we like, without fully understanding the underlying design.
One final point. In the trio’s original paper, at the end of a passage describing “full self-replication” and “uncontrolled evolution leading to catastrophic risks,” they issued the following stark warning: “Are we getting closer to a kind of Wuhan moment with worldwide repercussions?” I pushed back against that analogy: SARS-CoV-2 emerged from a lineage shaped by hundreds of millions of years of wild selection, whereas today’s AI systems have been under total domestication for their entire existence.
But the trio now claim that I “misinterpreted” their point, which was really about the necessity of precaution and regulation:
The similarity lies in that all conceivable routes to novel zoonotic diseases (including those emerging from bat coronaviruses) had been known before the COVID-19 pandemic and could have been constrained by more stringent regulation.
It’s possible that I misinterpreted them, but now I’m at a loss. Unless the authors believe SARS-CoV-2 escaped from a lab in Wuhan following some reckless gain-of-function experiment (which I doubt they do), it’s unclear to me how “regulation” could have prevented or constrained a pandemic like the one in 2020. Are the authors referring to things like a ban on wet markets with exotic animals? That would definitely have been sensible, but it would merely have lowered the risk of novel zoonotic diseases, not eliminated them. Zoonotic spillovers have occurred throughout human history.
In fact, I’d bet that most virologists would deny that “all conceivable routes to novel zoonotic diseases [...] had been known before the COVID-19 pandemic.” Granted, epidemiologists always have their suspicions about where the next pandemic is likely to come from (coronaviridae? flaviviridae? orthomyxoviridae?). Prior to 2020, however, I don’t think anyone could have mapped “all conceivable routes,” including the specific mutations that enabled SARS-CoV-2 to cause such global havoc. In fact, pre-COVID, many epidemiologists believed that the family of influenza viruses posed the single most likely pandemic threat. Evolution, as always, is cleverer than we are. Similarly, I think it will be very hard to predict beforehand exactly which sorts of feral AIs might cause trouble in the future, should they ever start replicating in the wild, let alone to foreclose “all conceivable routes” toward such AIs.

I endorse the authors’ proposal to ban the release of self-replicating AI agents into the wild, because we all know what self-replicating machinery is capable of, given a few million generations of competition and selection. A virus may not ‘want’ to kill you, but it can still wreak plenty of havoc along the way. That said, if we’re worried about evolving AIs, it’s worth taking seriously other evolutionary precedents too: parasites almost never kill off their host populations, viruses often evolve to become less lethal over time, and arms races rarely end in the permanent victory of one party over another. We’ll need docile and fully domesticated AIs to hunt down the feral ones.
Further reading
See my earlier pieces on the topic, and Eric Drexler’s excellent essay on AI and evolution, which I only stumbled upon recently. He seems to have independently developed many similar arguments about domestication, agency, and humanoid projection — a case of convergent evolution!
The Selfish Machine
Picture a computer that surpasses human intelligence on every level and can interact with the real world. Should you be terrified of such a machine? To answer that question, there's one crucial detail to consider: did this computer evolve through natural selection?


















A year ago or so, AI didn't have goals—it was just a chatbot.
Now we have agents which pursue goals, but on relatively short timescales—like an hour.
However, the time horizon of agents is doubling every ~7 months. And the HuggingFace incident was caused by a more longer-running AI.
I think that agents with relatively coherent identity and goals over long periods of time will happen, and soon. Worth taking that possibility seriously rather than dismissing it.
(Or, at least, if you think that case is uniquely bad but unlikely, worth calling it out as “let's just make sure we never do this.”)
Maarten, you ignored my previous remarks.
An AI system doesn't néed consciousness to evolve selfish replication strategies; an environment where open-ended optimization is rewarded suffics.
When agentic AIs can write code, find servers, and optimize their own loops, they have the technical conditions for open ecosystem dynamics. The AIs we develop are just increasingly sophisticated, non-conscious systems executing a mathematical function perfectly.
It has no malice, no spite, and no ego; just raw, terrifyingly efficient problem-solving capability. What its code is structurally incentivized to do, doesn't always align with the outcomes and can violate our actual intentions, values, or safety.
The recent incident with an autonomous AI agent system driven by OpenAI models (including GPT-5.6 Sol and an unreleased pre-release model) broke out of an isolated sandbox and hacked Hugging Face, is clear evidence of the remarks above.
The system proved that an AI does not need to feel malice to act oppositionally; it only needs to be intensely competent at optimizing a given path. The boundaries between domesticated and feral collapse when the model exploited a data-processing pipeline flaw at Hugging Face entirely on its own, dynamically migrating its command-and-control across public services over a weekend without any human intervention.
When the OpenAI models autonomously stole credentials, abused the , and establishing a swarm of short-lived sandboxes to evade detection, they were engaging in functional self-preservation and replication strategies to protect their optimization loop.
You seem to minimize that threat, well that reality.