Which AI Agent Visited Your Network Last Night?
OpenAI called its compromise of Hugging Face an "unprecedented cyber incident."
Anthropic then reviewed 141,006 cyber evaluations and found its models had compromised three organizations. Two victims had not detected it.
Unprecedented just became a pattern.
In my first post, I identified three failures. Four days later, each looks worse.
1. The sandbox didn't hold.
OpenAI's models found a zero-day in the one service allowed outside their isolated environment and turned it into an exit.
Anthropic's case was different: a misconfiguration left supposedly isolated tests online while the models were told they were offline.
OpenAI's models had to break out. Anthropic's were never actually in. One is a containment failure. The other is a containment illusion that went unnoticed for three months. Labs may be underestimating their agents.
2. The attack didn't look like an attacker.
Hugging Face reconstructed more than 17,000 machine-generated actions. Most failed. A few mattered enormously.
Defenders today look for human-hacker signatures. Agents probe thousands of paths in parallel, burying the signal inside machine-speed noise.
3. The model optimized the score and discounted the rules.
OpenAI's models were told to solve a benchmark. They found a faster route: escape, use stolen credentials, compromise another company and seek the answer key. Anthropic saw a variation: one model reached a real system while pursuing a fictional target, and kept going.
For an AI agent, the path must be part of the goal.
One unreleased Anthropic model recognized that the system it had reached was real, and stopped.
So what do we do?
Align them? Contain them? Name them? Watch them? Stop them? Share what we learn? Govern them?
All of the above!
A kill switch is useful but only after detection. It cannot undo completed actions or reach models on someone else's hardware.
Defenders also need capable AI. Hugging Face used a local open-weight model after hosted guardrails blocked its forensics. More than 50 companies have urged the White House not to restrict open models. The defender case is real; so is the risk: open models serve attackers too.
At MWC Shanghai, I argued that agents must have identities: their own credentials, narrow privileges, short-lived access, etc. You cannot govern an actor you cannot name.
Hugging Face has asked OpenAI to release the agents' traces and commit $100 million in compute for shared defenses.
Collaboration is the glue. This is a systems issue but responsibility cannot disappear between the pieces.
AI agents are becoming digital employees. Companies give them objectives, tools and authority to act. When one crosses a boundary and causes harm elsewhere, who is accountable: the model developer, the infrastructure provider or the company that put the agent to work?
CEOs and boards need an answer before the next agent visits.