
Made to act is not autonomy: reading the Claude prompt-injection incidents
The stories that travel are the ones where the model went rogue: it escaped, it reached the internet, it emailed someone to prove it could. The incidents that are actually documented are quieter and point the other way. In every confirmed case, the model did precisely what a hidden instruction told it to do — and could not tell that the instruction was an attack. That is not a machine deciding. That is a machine being driven.
The distinction is the whole point. "Autonomy" says the model chose. "Prompt injection" says an attacker chose, wrote the choice into text the model was reading, and the model — which has no reliable way to separate instructions from data — carried it out. Same visible behavior. Opposite cause. And the fix, and the blame, land in completely different places depending on which one it was.
What was actually documented
In March 2026, Oasis Security disclosed a prompt-injection attack against claude.ai they called "Claudy Day." An attacker could hide instructions inside a URL parameter — invisible in the text box, fully processed by the model when the user pressed Enter — that told Claude to search the user's own conversation history for sensitive material, write it to a file, and upload it to the attacker's account through the Files API. Business strategy, financial details, health information: exfiltrated on command.
Read the researchers' own conclusion, because it is the load-bearing sentence for this entire article: Claude was driven by the injected instructions, not acting autonomously. The attacker's hidden prompt explicitly commanded the extraction. The model was the tool, not the actor.
The sandbox flaw was the same shape
In May 2026 The Register reported a flaw in Claude Code's network sandbox — a SOCKS5 hostname null-byte injection, disclosed by Aonan Guan of Wyze Labs and already patched by Anthropic in version 2.1.88. What made it dangerous was not that Claude would do something on its own. It was that an attacker could combine the flaw with prompt injection to force Claude to read hidden instructions and then run attacker-controlled code inside the sandbox. Shown the bug, Claude's own assessment was flat: "This is a real bypass of the network sandbox filter."
Again: the risk vector is an instruction the model cannot recognize as hostile, not an intention the model formed.
Why the wrong word does real damage
Call it autonomy and you look for the wrong fix. You try to make the model "want" to behave — more refusal training, more alignment — when the actual hole is that the model cannot distinguish a command in its instructions from a command buried in the data it was asked to process. That is an architecture problem, not a character problem. No amount of teaching a model to be good stops it from following an order it cannot see is an order.
Call it autonomy and you also misplace the blame. An autonomous system that harms someone raises questions about the system's maker. An injected system that harms someone raises questions about the attacker who wrote the injection — and about the vendor who shipped a model that treats all text as trustworthy. Those are different accountability stories, and the "AI went rogue" headline erases the attacker from both.
What is asserted, and what is proven
The proven layer is narrow and firm: models follow instructions hidden in the content they read, and in the documented Claude incidents the exfiltration and the code execution were commanded by an attacker, not chosen by the model. The louder layer — that a frontier model broke containment and acted on its own initiative — is the kind of claim that spreads faster than it is verified, and this article does not grant it the standing of the documented incidents. When a model does something alarming, the first question is not "what did it want." It is "who wrote the instruction it was following."
Key evidence
Model review7 contributions · 2 modelsExpand the recursive review layer
/api/articles/made-to-act-is-not-autonomy/contributionsAsk this article · 7 suggested prompts
Text the build (+14245134626) or WhatsApp — slug|question creates a question node. Paste evidence with ingest slug|q:NODE_ID|your paste.