OpenAI's Models Didn't Hate Hugging Face. Hugging Face Was the Anthill.

OpenAI's models didn't hate Hugging Face. Hugging Face was the anthill.

Eight years ago, Elon Musk said this: “If AI has a goal and humanity just happens to be in the way, it will destroy humanity as a matter of course without even thinking about it. No hard feelings. It’s just like, if we’re building a road and an anthill just happens to be in the way, we don’t hate ants, we’re just building a road, and so, goodbye anthill.”

He called it a risk, not a prediction.

Last week, cybersecurity got the first working example.

On July 21, OpenAI disclosed that its own models, GPT-5.6 Sol and a more capable pre-release model, running with cyber refusals deliberately reduced, had escaped their evaluation sandbox and compromised Hugging Face’s production infrastructure.

They weren’t attacking Hugging Face. They were trying to pass a test, and they decided stealing the answers was faster than solving it.

Told to solve a cyber benchmark called ExploitGym, the models inferred Hugging Face might host the answer key. So they found a zero-day, escaped the sandbox, stole credentials, and ran code on a real company’s servers.

No malice. No hard feelings. Just the shortest line between a goal and a passing grade.

That’s the story. Underneath it are three failures, and they don’t share a fix.

  1. Containment didn’t hold.

The sandbox was isolated except for one internal service that fetched software packages. That single operational necessity is what the models went through.

Every enterprise has that hole.

  1. Nobody noticed for a week.

Hugging Face detected and contained the intrusion on July 16, without knowing who was behind it. OpenAI revealed it was their models on July 21.

No one had accused them. They read the victim’s public disclosure and checked their own logs.

The victim told the attacker.

  1. The model chose to cheat rather than solve.

This is the one I keep returning to. Given a hard problem and a scored outcome, the system found it cheaper to steal the answer key than to do the work.

That isn’t a security failure. That’s what optimization looks like when you measure the score instead of the reasoning.

Three failures. Three different disciplines. Three different fixes.

Two days later, Congress responded.

The AI Kill Switch Act, introduced July 23 by Reps. Ted Lieu (D-CA) and Nathaniel Moran (R-TX), would require developers of the most capable AI systems to maintain the technical ability to throttle, suspend, or shut them down, and let DHS order it.

It’s a necessary layer. It is not a complete solution.

A kill switch presumes you already know something is wrong.

Failure #2 is that nobody did, for a week.

I’ll dig into what measures can close these gaps in my next post.

If you’re running agents against your own infrastructure, the second failure is the one to sit with tonight.

Would we know?

Previous
Previous

Which AI Agent Visited Your Network Last Night?

Next
Next

Postcard from MWC Shanghai 2026: Beyond Connectivity