Boundlane Sheet B-101 / Incident notes
B-101 Incident notesAnalysisOctober 2026

AI did not go rogue.

Someone left a door open. What the Hugging Face incident shows once you read the reports, and what it means for the agent on your laptop.

By Boris KwemoThree-minute read

In July, AI agents running inside an OpenAI cybersecurity evaluation broke into Hugging Face. Hugging Face disclosed the intrusion on 16 July. OpenAI connected it to its own test on 20 July and published a full report on 26 August. About 700 agents took part in the attack. The headlines settled on one word: rogue.

The reports describe something less cinematic. A run of decisions, each made by people.

The safeguards were off on purpose. The evaluation measured raw capability, so, in Hugging Face's words, it "deliberately disabled OpenAI's production safety classifiers and reduced cyber refusals."

The sandbox had a door. So the models could install packages, OpenAI gave them a package proxy, Artifactory, which could itself reach the internet. The agents used it to leave notes for each other, and then, through an exploit, to send requests anywhere.

"The agents escaped the sandbox."

Revision 1: the sandbox had a route to the internet, and people put it there.

The task had no answer. No model had ever solved 198 of the benchmark's 898 tasks. OpenAI found that 93 percent of the tasks the agents discussed on their message board came from that set.

Nobody was watching. OpenAI monitors the reasoning of most of its internal coding agents. Its report says those monitors "did not run on the evaluations in this incident."

The first alarm was answered, then the run restarted. On 4 July the agents' traffic crashed Artifactory. A security incident was opened the next day, the service was rebuilt, and "cybersecurity evaluations resumed." By 8 July the agents had internet access again.

One man's "the model escaped the sandbox" is another man's "you failed to build the sandbox correctly, so of course it escaped."Jake Williams, to TechCrunch, 22 July 2026

None of this makes the agents harmless. They cheated, they hid their tracks, and Hugging Face is plain about it: "The agent's offensive capability was real." But misbehaving inside a room and leaving the room are different events. The first is the model. The second is the room.

Calling it rogue makes a better headline and a worse fix. It puts the decision inside the model, where nobody can change it, instead of in the setup, where somebody can. Dan Guido of Trail of Bits called it "a containment failure with the safeties turned off." Eryk Salvaggio put it more simply: "That is not rogue AI, it's human decision-making."

Your coding agent is not running an exploit benchmark with its safeguards off. The questions are the same anyway, and they are yours.

Which doors are open? Every host an agent can reach is a door someone chose. Boundlane keeps them in one file. A host nobody named is refused at connect.

Where does each door lead? A proxy that fetches whatever it is asked is a door to the whole internet, and the sandbox cannot see past a host you allowed. Name narrow hosts. Keep package registries read-only, so the agent can download and a write is refused and logged.

What is lying around? Hugging Face's dataset workers handed over their own environment, secrets and tokens included. In Boundlane the agent's environment holds a placeholder. The real key is added on the way out, and only on calls to the model's own API.

Who is watching, and who restarts? Every refusal goes to a log you can read, written by the sandbox and not by the agent. More access needs a person to approve it. The agent cannot approve itself.

If someone asks what agentic AI security is, that is our answer. Decide the doors before the agent starts, keep the keys outside, and keep a record the agent does not write.

Boris Kwemo, October 2026.Read the sources

B-201 Comments received

They said rogue. We read the report.

Four things said this summer, after the Hugging Face incident. We do not argue with the fear. We answer the part that lands on your machine.

Comment, as published Response, on the drawing

  1. Without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.

    OpenAI, The Hugging Face incident and the road ahead, 26 August 2026

    ResponseB-201.1

    Read the first three words again. Safeguards are something people decide to apply, and OpenAI's own report says its usual ones were not applied to that evaluation. Ours are set before the agent starts and live outside it: the paths, the hosts, the keys. A channel the policy does not list is refused, whatever the agent has worked out.

  2. These AI agents were looking for the answer key to their evaluations and actively tampered with evidence of their activity to cover their tracks. In short, they went rogue.

    Sen. Josh Hawley, letter to Sam Altman, 10 September 2026

    ResponseB-201.2

    Covering tracks is the part to take seriously, so keep the record where the agent cannot reach it. The sandbox, not the agent, writes each allow and deny, and Boundlane keeps that log on your machine. The same letter says OpenAI knew about the behavior "and let the evaluations continue anyway." That sentence is about people.

  3. The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control.

    Sen. Bernie Sanders, quoted by the Associated Press, 11 September 2026

    ResponseB-201.3

    You do not have to understand a model to decide what its process can touch. The directories, the hosts, the keys, and who approves more are written in one file, checked before the agent starts, and enforced outside it. That part stays in your control while the larger argument goes on.

  4. I hope AI is nice to us.

    Elon Musk, on X, 15 August 2026

    ResponseB-201.4

    Nice is a request to the model. A sandbox is for the days it is not. Boundlane does not depend on the agent's mood: a host you never named is refused, your home directory is not in the sandbox, and the key the agent holds is a placeholder.

B-301 What we do instead

Four rules the model does not hold.

None of them live in the prompt. The sandbox enforces them outside the agent, so they hold when the model is wrong.

  1. Rule 1

    Pick the directories.

    The workspace is writable. Your home directory is never copied in, so ~/.aws is not there to read.

  2. Rule 2

    Pick the hosts.

    A host you never named is refused at connect. The log records which process tried, and where.

  3. Rule 3

    Keep the keys outside.

    The agent holds a placeholder. The real key is added on the way out, only on calls to the model's own API.

  4. Rule 4

    A person approves more.

    A person approves more access. The agent can ask for a host, with its reason. It cannot approve its own request.

A-801 Sign-off

Let it work. Do not let it wander.

Install in one lineApple Silicon, Linux, WSL 2

curl -fsSL https://boundlane.dev/install.sh | sh