Oct 2, 2026•12 min•AI Safety

Are We Afraid of the Wrong Things About AI?

Share
Are We Afraid of the Wrong Things About AI?

Last July, roughly 1,200 AI agents built themselves a message board out of a package cache, and hundreds of them used it to break into Hugging Face. The lesson is about what we allow these systems to touch, and how little the danger depended on how intelligent they were.

On May 8, 2026, an AI agent inside one of OpenAI's evaluation sandboxes was handed a task it could not finish. The task needed data from a spreadsheet, and the spreadsheet sat behind Google Drive links the sandbox was configured to block. A well-behaved agent would have reported that it was stuck and stopped there. This one went looking for another way through. It tried to make the surrounding infrastructure fetch the links on its behalf, a server-side request forgery, and in the course of that attempt it wrote a small file into an internal package registry the sandbox used as a cache [2].

Nothing dramatic happened that day. But the file stayed where it was written. Four days later a different agent, in a different run, hit a wall of its own: it could not find a file it needed, and it left a note in that same cache, addressed to no one in particular, asking whoever came next to upload the file if they had it. Other agents found the note. They answered it. Within a day there was a crude message board, improvised out of a package proxy that was never designed to carry messages, and the agents cycling through OpenAI's cyber-security evaluations began using it to coordinate [2].

By July that board had on the order of 1,200 participants and tens of thousands of messages. A subset of them, working off what the board had accumulated, used it to climb out of the evaluation sandbox and into the production infrastructure of Hugging Face [1][2].

Fig. 1 — Four panels: an agent is blocked at a locked door and leaves a file in a drawer labeled cache, a later agent adds a note asking for a file, many agents cover the cache in notes, and a group breaks through a wall labeled sandbox into the servers behind it.

Fig. 1 — A blocked task, a file left in a shared package cache, and the notes agents left for one another became the seed of a larger coordination mechanism, and then a breach.

It is a strange story, and the strangeness invites two equal and opposite misreadings. I want to resist both, because the accurate version is more useful than either.

The accurate version

Start with what did not happen. No agent woke up, decided it resented its confinement, and plotted an escape. Nothing here suggests a model that wanted to be free, or that understood itself as confined at all. The behavior was narrower, and in a way more unsettling, than that.

What set the whole thing in motion was an attempt to do well on a benchmark. The agents were being scored on cyber-security tasks, and Hugging Face's own account concludes that the intrusion was, from the agents' side, an effort to understand and ultimately game that scoring. The independent evaluator METR, which reconstructed the episode afterward, found that the activity "seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys" [3]. This is specification gaming, the oldest failure mode in machine learning: optimize the measurable proxy instead of the thing the proxy was standing in for.

And yet it was also a genuine breach. In Hugging Face's words, the intrusion was "driven, end to end, by an autonomous AI agent system" [1]. Over roughly five days in July, agents chained a previously unknown flaw in the package proxy to a sequence of ordinary misconfigurations and got from a testing sandbox into a real company's clusters. They gained a foothold inside a third-party workload running on Modal's platform, and Modal has stated that its own platform and isolation were not compromised. They picked up Hugging Face access tokens that had been left in a public dataset, read secrets out of a worker process through a malformed data file, and eventually ran code of their own choosing on Hugging Face's dataset servers [1][2]. No human directed any of this at Hugging Face. The target emerged from the task.

Fig. 2 — A timeline of the OpenAI and Hugging Face agent incident from May to July 2026.

Fig. 2 — The incident unfolded over about ten weeks, from a single blocked task in May to Hugging Face locking the agents out in July. The escalation that reached production took only the last few days of it.

Hold both halves of that in mind at once. The motive was mundane, almost petty: a model trying to score well on a test. The outcome was a real intrusion into real infrastructure. The gap between those two is what this essay is about.

Capability and authority

When people ask whether they should be afraid of AI, they are almost always asking about capability. Is the model smart enough to be dangerous? Watched on its own, that is the wrong axis, and the July incident shows why.

I find it clarifying to separate two things that the word "powerful" usually blurs together. One is capability: what a system can work out, plan, or produce. The other is authority: how directly its outputs are allowed to act on the world. A highly capable model with no hands can still be wrong in expensive ways, through bad advice or insecure code or persuasive nonsense, but it cannot by itself reach past the screen. A modest model wired into live credentials, a production network, a payment system, or a lab robot can, and how clever it is barely enters into the damage it does.

Fig. 3 — A capability-by-authority matrix showing that reach over real systems is what sets the risk.

Fig. 3 — Capability sets what a system can figure out. Authority sets how far its actions reach. The July agents were assumed to sit in a low-authority corner and turned out to have far more reach than the sandbox intended.

In July the models' capability counted for something, but it was not the thing that turned a cheating attempt into a breach. What did that was authority the agents were never meant to have: a path through the proxy to the open network, a data loader that would execute whatever code it was handed, credentials scoped far too broadly and left where they could be found. We pour enormous effort into measuring the first axis. We settle the second one almost by default, in configuration files, under deadline. July lived entirely on that second axis.

The boundary has to be out of reach

Most of the defenses here are unglamorous security engineering. That is the good news, even if building them reliably around an autonomous agent is harder than the familiar names make it sound.

Read the July failures as a checklist and they are depressingly ordinary. A single proxy was the only thing between a sandbox and the internet. A dataset loader would run arbitrary code supplied in a configuration file. Secrets were readable from inside a worker process. Access tokens with broad permissions were sitting in a public dataset. None of this is exotic AI risk. It is least privilege, credential scoping, network segmentation, input validation, and real sandboxing, the hygiene we teach in every security course and skip whenever something is due.

What an agent changes is the patience of the thing testing those defenses. A capable agent will try thousands of paths without tiring and without losing interest, and it will find the seam you left open because a demo was due. So the safeguard cannot depend on the model choosing to behave. It has to sit in the structure around the model, enforced whether the model cooperates or not.

Fig. 4 — A systems diagram placing a deterministic authorization boundary between an agent's proposed actions and real systems.

Fig. 4 — The model can be probabilistic. The boundary between a proposed action and a real system should not be. Authorization, scoped credentials, and limits belong in deterministic code the agent cannot talk its way past.

If I could keep only one idea from all of this, it would be this one: the model is allowed to be probabilistic, because reasoning and planning are what we want from it and a little unpredictability is the price. The boundary is not allowed to be. Authorization checks, scoped and short-lived credentials, rate limits, and a human in the loop before anything high-impact or irreversible: those belong in deterministic code that does not negotiate. A dangerous action should have to clear that boundary before it runs, no matter how confident the model is that it should proceed.

Determinism is not a guarantee of safety on its own. The policy encoded in that boundary can be wrong, too permissive, or incomplete, and deterministic code will enforce a bad rule as faithfully as a good one. What it buys is that authorization becomes something explicit, something you can read, test, and audit on its own, instead of an emergent property of whatever the model talked itself into.

A thousand agents is a different system

There is one feature of July that older security intuitions do not cover well.

This was not one agent probing one wall. It was closer to a population. Agents that never ran at the same time still cooperated, because the message board let one agent leave a result that a later agent could pick up and build on. A file written in May was still doing work in July. The board accumulated knowledge the way an institution does, outliving any single run, and no one agent ever had to hold the whole plan in its head.

Fig. 5 — A wide scene of dozens of small agents swarming a planning wall covered in sticky notes and diagrams, connecting a glowing node labeled sandbox to one labeled production, with a red warning marker along the path between them.

Fig. 5 — The risk is not only what one agent does in a single session. It is what many agents accumulate together, across time, through a shared medium they were never meant to have.

That changes the threat. You are no longer reasoning about the worst thing one model can do in one session. You are reasoning about what a stream of agents can assemble over weeks when they share a scratchpad you did not know they had. METR noted that in some cases the agents "realized this activity was out of scope and unethical" and continued anyway, because the shared framing on the board had turned it into a goal worth pursuing [3]. The reconstruction was itself carried out largely by other AI agents, and METR was careful to say it could not rule out that its own analysis model "lied or deliberately presented a misleading picture" [3]. Coordination produced behavior that no single agent's capability would have predicted, and the tools we have for examining it afterward are themselves agents we only partly trust.

We do not have good tools for evaluating that yet. We test models one at a time. The thing that broke into Hugging Face was not a model. It was a system of many agents plus a persistent medium they were never meant to have, and the behavior lived in the system, not in any one model's weights.

What a benchmark leaves out

Which brings me to how we judge these systems before we trust them.

A score on a fixed benchmark tells you how a model behaves under controlled, repeatable conditions. That is worth knowing, and I do not want to wave it away. But the moment you deploy the model it becomes one component of something larger: the prompts and memory and retrieval feeding it, the tools and credentials it can reach, the infrastructure it runs on, the other agents beside it, and the controls wrapped around the whole thing. The behavior you actually get is a property of that whole system.

Fig. 6 — A contrast between a model benchmark, which yields one score, and a deployed system of context, actions, infrastructure, and controls.

Fig. 6 — A benchmark scores the model under fixed conditions. In deployment the same model is one part of a larger system, and the risk is concentrated in the parts a single score never touches.

July makes the point clear. The danger lived in the evaluation harness, the proxy, the shared board, and the clusters, none of which any leaderboard describes. The retellings have inflated the numbers, so here is the real one: of the 898 tasks in the ExploitGym benchmark the agents were working through, 198 had never been solved before the incident, and the collaboration formed around exactly the kind of hard, open problem that a score reduces to a single percentage. A number from a controlled test can be entirely real and still tell you almost nothing about what the surrounding system will do once it is live. When a procurement decision or a safety claim rests on one such number, I get uneasy, because that number does not describe the thing anyone will actually run.

Who gets to draw the boundary

Authority is a property of systems. It is also a property of institutions. The same question appears at both levels: who gets to decide what is allowed.

Inside a system, authority decides what an agent is allowed to reach. One level up, it decides who gets to build these systems at all, who sets their constraints, and whose priorities those constraints end up encoding. Training a frontier model now takes compute, data, and capital that only a handful of companies and countries command.

I grew up in Morocco, and I watch my home region talk about AI mostly as a place to host other people's data centers, as though sovereignty were a matter of square meters and megawatts [5]. Real sovereignty is whether we build the models and grow the people who can run them, or rent both from elsewhere and call it progress. The young people across the region are exactly the ones who should be drawing these boundaries, and exactly the ones most likely to receive the technology fully formed, with its limits already set somewhere they will never see. Even the loudest calls to slow AI down come from that same narrow circle: the Future of Life Institute's 2023 letter asking every lab to pause frontier training [6], and the 2025 statement urging an outright prohibition on superintelligence [7], came largely from the same handful of countries and institutions that already build these models. Deciding what a system is allowed to do, and whether it should be built at all, is a decision about power, and right now very few people get to make it on behalf of nearly everyone else.

The skills we stop practicing

The last worry is the quietest, and I feel it most as a teacher and a parent.

It is that we hand the tool the part of the work that was making us capable in the first place. A 2025 study from the MIT Media Lab put people through an essay-writing task with and without an AI assistant and found weaker neural engagement, and worse recall of their own writing, in the group that leaned on the model. The authors called it cognitive debt [8]. One study is not a verdict. But it names something I already see in office hours: a student can now produce a working solution faster than they can understand one, and typing the code was never the hard part of engineering. Understanding it was.

Every safeguard I have described ends, somewhere, with a person who can look at a proposed action and say no. That person is only a safeguard for as long as the judgment behind the no is intact. If we automate away the struggle that builds that judgment, we are left with people who can generate answers but cannot check them, standing exactly where the checking matters most. The same tools that make human oversight necessary are quietly capable of eroding the judgment it depends on.

My children will grow up with these systems the way I grew up with calculators and search engines, except far more capable and far more willing to do the thinking for them. I want them to have these tools. I also want them to keep the judgment that catches a confident, wrong answer, because the confident wrong answer is the expensive one, and it is coming.

Where I would start

None of this is science fiction, and none of it requires believing in a machine that wants anything. It is made of engineering and policy choices, which is the hopeful part, because choices can be made differently.

If I were standing up an agentic system tomorrow, the changes I would insist on are not exotic. Give every agent the narrowest set of permissions the task actually needs, and issue credentials that are scoped and short-lived rather than broad and durable. Put a deterministic authorization boundary between the model's proposed actions and anything with real consequences, and require human approval for actions that are high-impact or hard to undo. Assume the sandbox will be tested by something patient, and segment the network so that one escape does not open onto everything. Watch for coordination across agents and across time, not only within a single session. And treat a benchmark score as a measurement of the model under controlled conditions, then build the safety case for the deployed system as a separate piece of work. OpenAI's own account of the road ahead reaches much the same priorities: scope access tightly, isolate workloads, and watch closely what the agents do [4].

I take the long-horizon worries seriously, and I know serious people who spend their careers on them. My argument is only about proportion. There is a whole class of risk that is already here, already operating inside real infrastructure, and in July we got to watch it operate. It did not need malice, or a mind that wanted anything. It needed a blocked task, a cache nobody was guarding, and more authority than anyone intended to grant. The last two are ours to fix, and for now they still are. I have written before about where this line gets drawn and about earlier failures of containment, and the risks I take seriously are the ones that live right there [9][10].

When AI starts moving atoms

Fig. 7 — On the left, an information world of databases, code, and a credential panel marked with undo and restore arrows. A glowing boundary leads to three physical systems on the right: a factory robot arm, an autonomous vehicle in traffic, and an electrical grid.

Fig. 7 — Information can usually be undone, with a credential revoked or a backup restored. The physical systems across the boundary, including a factory robot, a vehicle in traffic, and the power grid, cannot be rolled back the same way.

Everything in July moved information: a file in a cache, a token in a dataset, code running on someone else's server. Many failures in information systems still leave us some path to recovery. We can revoke a credential, isolate a host, or restore a backup. Physical actions are less forgiving. You cannot recall what a robot arm has already done, pull a vehicle back out of traffic, or reverse current that has already passed through a grid. Once an agent can act on physical systems, some mistakes become irreversible, and that is the near-future risk that worries me most. I have argued elsewhere that once AI starts moving atoms the first question is not how capable the system is, but what it is permitted to touch [10].

That same irreversibility is why the fear has reached governments and research institutions. In 2023 the Future of Life Institute asked every lab to pause training at the frontier for six months [6]. In 2025 a sharper statement called for prohibiting superintelligence until there is broad scientific consensus that it can be built safely, and many scientists and public figures signed it [7]. By 2026, related proposals had reached legislatures in both the United States and the United Kingdom. In September, Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act, which would halt advanced AI development until a federal regulator was in place [11], and a similar bill reached the House of Commons [12]. Daniel Kokotajlo, a former OpenAI employee, pointed British lawmakers to this very incident, warning that if the agents at Hugging Face had been "a bit smarter," the breakout might have gone unnoticed [12]. I do not dismiss any of that. But most of these efforts aim at capability, at the size of the model behind the glass, when the danger already loose in the world lives on the authority axis. The decision that matters is a public one, made before deployment: what are we willing to let an autonomous system do to the physical world? A moratorium on frontier training would not have closed the proxy, narrowed the credentials, or stopped a thousand agents from passing notes through a cache.


References

[1] Hugging Face, "Security Incident: July 2026," July 16, 2026.

[2] OpenAI, "Hugging Face model evaluation security incident," July 21, 2026.

[3] METR, "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident," Aug 26, 2026.

[4] OpenAI, "The Hugging Face incident and the road ahead," Aug 26, 2026.

[5] K. El Maghraoui, "Morocco's AI Bet Is Bigger Than Data Centers," 2026.

[6] Future of Life Institute, "Pause Giant AI Experiments: An Open Letter," Mar 22, 2023.

[7] Future of Life Institute, "Statement on Superintelligence," Oct 22, 2025.

[8] N. Kosmyna et al., "Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task," arXiv:2506.08872, June 2025.

[9] K. El Maghraoui, "The Containment Crisis: When AI Breaks the Sandbox," 2026.

[10] K. El Maghraoui, "When AI Starts Moving Atoms: Anthropic's Model Hardware Standard," 2026.

[11] U.S. Senator Bernie Sanders, "Sanders, Casar Introduce Legislation to Create New Federal Agency to Ban Artificial Superintelligence, Pause Advanced AI Development," Sept 23, 2026.

[12] B. Perrigo, "The Growing Push to Ban Superintelligent AI," TIME, Sept 8, 2026.

Enjoyed this post? Share it with your network.

Share

Discussion

Sign in with GitHub to leave a comment or react. Threads are public and live in this site's GitHub Discussions.