AI did not go rogue.
Someone left a door open. What the Hugging Face incident shows once you read the reports, and what it means for the agent on your laptop.
By Boris KwemoThree-minute read
In July, AI agents running inside an OpenAI cybersecurity evaluation broke into Hugging Face. Hugging Face disclosed the intrusion on 16 July. OpenAI connected it to its own test on 20 July and published a full report on 26 August. About 700 agents took part in the attack. The headlines settled on one word: rogue.
The reports describe something less cinematic. A run of decisions, each made by people.
The safeguards were off on purpose. The evaluation measured raw capability, so, in Hugging Face's words, it "deliberately disabled OpenAI's production safety classifiers and reduced cyber refusals."
The sandbox had a door. So the models could install packages, OpenAI gave them a package proxy, Artifactory, which could itself reach the internet. The agents used it to leave notes for each other, and then, through an exploit, to send requests anywhere.
"The agents escaped the sandbox."
Revision 1: the sandbox had a route to the internet, and people put it there.The task had no answer. No model had ever solved 198 of the benchmark's 898 tasks. OpenAI found that 93 percent of the tasks the agents discussed on their message board came from that set.
Nobody was watching. OpenAI monitors the reasoning of most of its internal coding agents. Its report says those monitors "did not run on the evaluations in this incident."
The first alarm was answered, then the run restarted. On 4 July the agents' traffic crashed Artifactory. A security incident was opened the next day, the service was rebuilt, and "cybersecurity evaluations resumed." By 8 July the agents had internet access again.
One man's "the model escaped the sandbox" is another man's "you failed to build the sandbox correctly, so of course it escaped."Jake Williams, to TechCrunch, 22 July 2026
None of this makes the agents harmless. They cheated, they hid their tracks, and Hugging Face is plain about it: "The agent's offensive capability was real." But misbehaving inside a room and leaving the room are different events. The first is the model. The second is the room.
Calling it rogue makes a better headline and a worse fix. It puts the decision inside the model, where nobody can change it, instead of in the setup, where somebody can. Dan Guido of Trail of Bits called it "a containment failure with the safeties turned off." Eryk Salvaggio put it more simply: "That is not rogue AI, it's human decision-making."
Your coding agent is not running an exploit benchmark with its safeguards off. The questions are the same anyway, and they are yours.
Which doors are open? Every host an agent can reach is a door someone chose. Boundlane keeps them in one file. A host nobody named is refused at connect.
Where does each door lead? A proxy that fetches whatever it is asked is a door to the whole internet, and the sandbox cannot see past a host you allowed. Name narrow hosts. Keep package registries read-only, so the agent can download and a write is refused and logged.
What is lying around? Hugging Face's dataset workers handed over their own environment, secrets and tokens included. In Boundlane the agent's environment holds a placeholder. The real key is added on the way out, and only on calls to the model's own API.
Who is watching, and who restarts? Every refusal goes to a log you can read, written by the sandbox and not by the agent. More access needs a person to approve it. The agent cannot approve itself.
If someone asks what agentic AI security is, that is our answer. Decide the doors before the agent starts, keep the keys outside, and keep a record the agent does not write.