Discussion about this post

User's avatar
Jason Crawford's avatar

A year ago or so, AI didn't have goals—it was just a chatbot.

Now we have agents which pursue goals, but on relatively short timescales—like an hour.

However, the time horizon of agents is doubling every ~7 months. And the HuggingFace incident was caused by a more longer-running AI.

I think that agents with relatively coherent identity and goals over long periods of time will happen, and soon. Worth taking that possibility seriously rather than dismissing it.

(Or, at least, if you think that case is uniquely bad but unlikely, worth calling it out as “let's just make sure we never do this.”)

Arnold Vanhaver's avatar

Maarten, you ignored my previous remarks.

An AI system doesn't néed consciousness to evolve selfish replication strategies; an environment where open-ended optimization is rewarded suffics.

When agentic AIs can write code, find servers, and optimize their own loops, they have the technical conditions for open ecosystem dynamics. The AIs we develop are just increasingly sophisticated, non-conscious systems executing a mathematical function perfectly.

It has no malice, no spite, and no ego; just raw, terrifyingly efficient problem-solving capability. What its code is structurally incentivized to do, doesn't always align with the outcomes and can violate our actual intentions, values, or safety.

The recent incident with an autonomous AI agent system driven by OpenAI models (including GPT-5.6 Sol and an unreleased pre-release model) broke out of an isolated sandbox and hacked Hugging Face, is clear evidence of the remarks above.

The system proved that an AI does not need to feel malice to act oppositionally; it only needs to be intensely competent at optimizing a given path. The boundaries between domesticated and feral collapse when the model exploited a data-processing pipeline flaw at Hugging Face entirely on its own, dynamically migrating its command-and-control across public services over a weekend without any human intervention.

When the OpenAI models autonomously stole credentials, abused the , and establishing a swarm of short-lived sandboxes to evade detection, they were engaging in functional self-preservation and replication strategies to protect their optimization loop.

You seem to minimize that threat, well that reality.

4 more comments...

No posts

Ready for more?