A year ago or so, AI didn't have goals—it was just a chatbot.
Now we have agents which pursue goals, but on relatively short timescales—like an hour.
However, the time horizon of agents is doubling every ~7 months. And the HuggingFace incident was caused by a more longer-running AI.
I think that agents with relatively coherent identity and goals over long periods of time will happen, and soon. Worth taking that possibility seriously rather than dismissing it.
(Or, at least, if you think that case is uniquely bad but unlikely, worth calling it out as “let's just make sure we never do this.”)
Thanks for the pushback, Jason!. But it's worth being precise about which "goals" you refer to, given my point that this concept covers two very different things. Sure, you can have AIs that pursue more complex "goals" with longer time horizons, and I agree this brings novel risks. But these are still myopic, pre-programmed objectives with no stake whatsoever in the system's own continuation (to the extent there's even a cohesive "entity" being preserved over time at all, see Drexler's point about compound systems).
That's very different from the "goals" of evolved agents, which are autonomous and cross-context, tied to self-preservation and resistance to interference ("Don't tamper with my goals! Don't try to terminate me!"). The agent that carried out the Hugging Face attack doesn't cares about "its" own continued existence independent of the assigned task. There's no evidence it resisted shutdown or reprogramming once the ruse was discovered. When Hugging Face's security team intervened, that was the end of it.
Could there be a danger that we create an agent so single-mindedly obsessed with achieving a goal that it "thinks" it should resist shutdown? I'm suspicious of such "instrumental convergence," as you probably know. https://maartenboudry.substack.com/p/why-hal-9000-feared-death-and-real A chess engine is single-mindedly focused on mating its opponent— but only within the context of the game. It couldn't care less if you shut it down just before its winning move.
That said, I do think there's a real danger from *emulated* self-preservation, emerging from human training data (since of course humans resist shutdown or termination!). Sure, it wouldn't be genuine self-preservation, any more than Hugging Face's agent was engaged in genuine "deception" — but it could be behaviorally indistinguishable. I'd take that seriously and support hard limits on long-horizon autonomous agents without proper guardrails, regardless of whether we call this "genuine" agency.
Still, I don't buy the argument fom "instrumental convergence" as a sort of inevitable development once you reach a certain threshold of intelligence. Any objective programmed into an AI should always remain strictly conditional: "Try to achieve X or Y until further notice, or unless instructed otherwise." The programmer or authorized controler should always retain the power to overrule goals or shut the system down. I don't see why that would be impossible. The main reason we assume it's not possible is that it doesn't work with evolved creatures and their "goals". But AIs are not evolved creatures.
Conflating the two conceptions of "goals" is exactly the move I think the doom scenarios make, and it's also where an important disanalogy with evolved organisms holds: nothing here suggests these systems care about outliving the task, or about their goals being overwritten.
The expert consensus point is a bit sneaky - like saying philosophers of religion all believe in the existence of God so it must be true.
Fears of AI doom well predate modern LLMs and ML and a significant proportion of the "experts" now working in AI alignment come precisely from that rationalist position - it's what draws them. And even the wider AI engineering space has very strong links.
Yeah, I suspect that's true, but it's hard to prove and a bit "ad hominem", that's why I didn't go there. ;-) The same probably applies to climate scientists as well.
An AI system doesn't néed consciousness to evolve selfish replication strategies; an environment where open-ended optimization is rewarded suffics.
When agentic AIs can write code, find servers, and optimize their own loops, they have the technical conditions for open ecosystem dynamics. The AIs we develop are just increasingly sophisticated, non-conscious systems executing a mathematical function perfectly.
It has no malice, no spite, and no ego; just raw, terrifyingly efficient problem-solving capability. What its code is structurally incentivized to do, doesn't always align with the outcomes and can violate our actual intentions, values, or safety.
The recent incident with an autonomous AI agent system driven by OpenAI models (including GPT-5.6 Sol and an unreleased pre-release model) broke out of an isolated sandbox and hacked Hugging Face, is clear evidence of the remarks above.
The system proved that an AI does not need to feel malice to act oppositionally; it only needs to be intensely competent at optimizing a given path. The boundaries between domesticated and feral collapse when the model exploited a data-processing pipeline flaw at Hugging Face entirely on its own, dynamically migrating its command-and-control across public services over a weekend without any human intervention.
When the OpenAI models autonomously stole credentials, abused the , and establishing a swarm of short-lived sandboxes to evade detection, they were engaging in functional self-preservation and replication strategies to protect their optimization loop.
You seem to minimize that threat, well that reality.
As always, great, practical, level-headed analysis. TY!
IMO, being an oldster and having watched decades of people finding new Dooms to focus on, I think people just want their lives to have Meaning. With the relative fading of religion, we glom onto something else that "proves" that our lives are So Very Important (e.g. "What I do will impact the future of sentience in the Universe!")
Well written and masterly reflected; now, AIs are passive and not agressive, is argued here; but what if criminals would incite AIs deliberately to do harm to people, companies or societies?
A year ago or so, AI didn't have goals—it was just a chatbot.
Now we have agents which pursue goals, but on relatively short timescales—like an hour.
However, the time horizon of agents is doubling every ~7 months. And the HuggingFace incident was caused by a more longer-running AI.
I think that agents with relatively coherent identity and goals over long periods of time will happen, and soon. Worth taking that possibility seriously rather than dismissing it.
(Or, at least, if you think that case is uniquely bad but unlikely, worth calling it out as “let's just make sure we never do this.”)
Thanks for the pushback, Jason!. But it's worth being precise about which "goals" you refer to, given my point that this concept covers two very different things. Sure, you can have AIs that pursue more complex "goals" with longer time horizons, and I agree this brings novel risks. But these are still myopic, pre-programmed objectives with no stake whatsoever in the system's own continuation (to the extent there's even a cohesive "entity" being preserved over time at all, see Drexler's point about compound systems).
That's very different from the "goals" of evolved agents, which are autonomous and cross-context, tied to self-preservation and resistance to interference ("Don't tamper with my goals! Don't try to terminate me!"). The agent that carried out the Hugging Face attack doesn't cares about "its" own continued existence independent of the assigned task. There's no evidence it resisted shutdown or reprogramming once the ruse was discovered. When Hugging Face's security team intervened, that was the end of it.
Could there be a danger that we create an agent so single-mindedly obsessed with achieving a goal that it "thinks" it should resist shutdown? I'm suspicious of such "instrumental convergence," as you probably know. https://maartenboudry.substack.com/p/why-hal-9000-feared-death-and-real A chess engine is single-mindedly focused on mating its opponent— but only within the context of the game. It couldn't care less if you shut it down just before its winning move.
That said, I do think there's a real danger from *emulated* self-preservation, emerging from human training data (since of course humans resist shutdown or termination!). Sure, it wouldn't be genuine self-preservation, any more than Hugging Face's agent was engaged in genuine "deception" — but it could be behaviorally indistinguishable. I'd take that seriously and support hard limits on long-horizon autonomous agents without proper guardrails, regardless of whether we call this "genuine" agency.
Still, I don't buy the argument fom "instrumental convergence" as a sort of inevitable development once you reach a certain threshold of intelligence. Any objective programmed into an AI should always remain strictly conditional: "Try to achieve X or Y until further notice, or unless instructed otherwise." The programmer or authorized controler should always retain the power to overrule goals or shut the system down. I don't see why that would be impossible. The main reason we assume it's not possible is that it doesn't work with evolved creatures and their "goals". But AIs are not evolved creatures.
Conflating the two conceptions of "goals" is exactly the move I think the doom scenarios make, and it's also where an important disanalogy with evolved organisms holds: nothing here suggests these systems care about outliving the task, or about their goals being overwritten.
The expert consensus point is a bit sneaky - like saying philosophers of religion all believe in the existence of God so it must be true.
Fears of AI doom well predate modern LLMs and ML and a significant proportion of the "experts" now working in AI alignment come precisely from that rationalist position - it's what draws them. And even the wider AI engineering space has very strong links.
Yeah, I suspect that's true, but it's hard to prove and a bit "ad hominem", that's why I didn't go there. ;-) The same probably applies to climate scientists as well.
Maarten, you ignored my previous remarks.
An AI system doesn't néed consciousness to evolve selfish replication strategies; an environment where open-ended optimization is rewarded suffics.
When agentic AIs can write code, find servers, and optimize their own loops, they have the technical conditions for open ecosystem dynamics. The AIs we develop are just increasingly sophisticated, non-conscious systems executing a mathematical function perfectly.
It has no malice, no spite, and no ego; just raw, terrifyingly efficient problem-solving capability. What its code is structurally incentivized to do, doesn't always align with the outcomes and can violate our actual intentions, values, or safety.
The recent incident with an autonomous AI agent system driven by OpenAI models (including GPT-5.6 Sol and an unreleased pre-release model) broke out of an isolated sandbox and hacked Hugging Face, is clear evidence of the remarks above.
The system proved that an AI does not need to feel malice to act oppositionally; it only needs to be intensely competent at optimizing a given path. The boundaries between domesticated and feral collapse when the model exploited a data-processing pipeline flaw at Hugging Face entirely on its own, dynamically migrating its command-and-control across public services over a weekend without any human intervention.
When the OpenAI models autonomously stole credentials, abused the , and establishing a swarm of short-lived sandboxes to evade detection, they were engaging in functional self-preservation and replication strategies to protect their optimization loop.
You seem to minimize that threat, well that reality.
Wat ik me afvraag is niet of AI iets wil, maar wat een steeds capabeler systeem kan gaan doen — ook zonder bewustzijn, ego of eigenbelang.
Het incident kwam nadat een mens de opdracht gaf om in te breken
As always, great, practical, level-headed analysis. TY!
IMO, being an oldster and having watched decades of people finding new Dooms to focus on, I think people just want their lives to have Meaning. With the relative fading of religion, we glom onto something else that "proves" that our lives are So Very Important (e.g. "What I do will impact the future of sentience in the Universe!")
https://www.mattball.org/2023/05/doom-force-that-gives-life-meaning.html
https://www.mattball.org/2023/05/mans-search-for-meaning-aka-these-are.html
It's not doom. Reality has already caught up with Maarten's minimization.
Well written and masterly reflected; now, AIs are passive and not agressive, is argued here; but what if criminals would incite AIs deliberately to do harm to people, companies or societies?