On the Gray Swan indirect-prompt-injection benchmark, a set of 1,130 high-transferability attacks, Claude Opus 5 reduced the probability of an attacker succeeding within 15 attempts "from 5.5% to 2.0%" against Opus 4.8 (Anthropic Transparency Hub, 24 July 2026). About a month later a security researcher published a working end-to-end code-execution chain against that same model running in Claude Code auto mode, reporting "attack success rates up to 80% using a small sample size" with individual variants landing between 60% and 80% (Embrace The Red, 26 August 2026).
Both numbers are primary-sourced. Both are correct. The spread between them is not a contradiction. It is the difference between attacks that were in the evaluation set and one attack that was not.
Anthropic wrote the conclusion down before either number existed. Its engineering post on containment states: "Yet even with best-in-class defenses, protection in the model layer will never be 100% effective, which is why it can't stand alone" (Anthropic, 25 May 2026). A few lines on, the same page puts it as plainly as anyone has: "The deterministic boundary is what gets hit when everything probabilistic misses."
A claim that prompt injection is largely solved in practice circulated widely this year. The vendor's own institutional publications say something far more careful, repeatedly and in writing. This post turns them into a configuration checklist.
This sits in our own-your-stack cluster. The pillar above it is the Ops Automation Playbook, which ranks the workflows worth automating first.
By Dan Colta. We are a founder-led EU automation studio. There are two of us. We run scheduled unattended agents on our own machines, the content and outreach pipelines behind this blog among them. This is a configuration checklist rather than an offer: we do not sell security work.
Key Takeaways
- Model-layer defence got genuinely better. Claude Opus 5 cut attacker success within 15 attempts from 5.5% to 2.0% on a 1,130-attack benchmark. Single-attempt success fell from 0.5% to 0.2% (Anthropic Transparency Hub).
- A benchmark score describes the attacks inside the benchmark. A researcher reported up to 80% on a chain that was not in one (Embrace The Red). Eight published defences were bypassed at over 50% by adaptive attacks (Zhan et al., NAACL 2025 Findings).
- Auto mode is a per-action control rather than an isolation boundary. Anthropic publishes its own miss rate: 17% false negatives on real overeager actions (Anthropic).
- Claude Code's Bash sandbox fails open by default: a missing dependency produces a warning and an unsandboxed run (Claude Code docs). Set
sandbox.failIfUnavailableto true. Codex CLI runs with network access off by default (OpenAI).- Containment reduces blast radius rather than removing risk. In a February 2026 internal red team, Claude Code completed a credential exfiltration 24 times out of 25 retries. What held was the environment (Anthropic).
How good has model-layer defence actually got?
Materially better over the past year, measured by the vendor itself. On the Gray Swan benchmark of 1,130 high-transferability indirect-prompt-injection attacks, Claude Opus 5 cut the chance of an attacker succeeding within 15 attempts "from 5.5% to 2.0%" against Opus 4.8. Single-attempt success fell from 0.5% to 0.2% (Anthropic Transparency Hub, 24 July 2026).
A 2.75x reduction at fifteen tries is real engineering, not marketing. Anthropic's browser-use work has the same character: it reports progress, then refuses to round it up. It writes that "no browser agent is immune to prompt injection" and that it shares the findings "to demonstrate progress, not to claim the problem is solved" (Anthropic, 24 November 2025).
The framing that matters for unattended work is the k. A person running one interactive session is close to a single-attempt event and gets the 0.2% number. A scheduled agent that ingests an inbox, a feed or a scraped page every hour is a different machine. It is the one that takes the 15th attempt. Then the 1,130th.
That is the shape of every pipeline worth building, including the dull ones. A weekly reporting job that pulls numbers on a schedule is unattended by design. Nobody watches the 3am run.
Why does a 2.0% benchmark score not contradict an 80% working exploit?
Because a benchmark measures the attacks in it. On 26 August 2026 a researcher published a working chain against Claude Code Opus 5 in auto mode, reporting "attack success rates up to 80%" across variants of 60-80%, self-scoped as "small samples, not a universal ASR measurement" (Embrace The Red).
Anthropic's response, as the researcher reports it and in paraphrase: auto mode is a convenience feature backed by a best-effort classifier rather than a security guarantee. Chains of benign-looking steps are not what it is built to stop. The real boundary is OS isolation plus network egress control. Anthropic closed the report as Informative. Its own documentation says the same. The exchange reads as consistent rather than defensive.
Generalisation has been tested independently and the answer is hard. Researchers at UIUC evaluated eight published defences against indirect prompt injection on LLM agents and bypassed every one using adaptive attacks, "consistently achieving an attack success rate of over 50%" (Zhan, Fang, Panchal and Kang, NAACL 2025 Findings, February 2025). A defence tuned against known attacks is not a defence against the attack written for it.
What does auto mode actually promise?
A per-action check rather than a boundary, with a published 17% miss rate. Anthropic says so in two separate places. Auto mode is the built-in starting permission mode on Claude Code Pro, Max and Team plans (Claude Code docs, retrieved 18 September 2026). It is the default most people start from rather than something most people opted into.
The engineering post that introduced auto mode publishes its own miss rate: "The 17% false-negative rate on real overeager actions is the honest number" (Anthropic, 25 March 2026). It also records a 0.4% false-positive rate on real traffic and says auto mode "is not a drop-in replacement for careful human review on high-stakes infrastructure". A vendor publishing its own false-negative rate deserves more trust. The control it describes deserves less.
The docs are equally direct about the category error: "The classifier is a per-action control, not an isolation boundary" (Claude Code docs, retrieved 18 September 2026). A classifier scores the action in front of it. It does not constrain what the process can reach.
| Control | What the vendor calls it | Published number | Source |
|---|---|---|---|
| Auto mode classifier | A per-action control, not an isolation boundary | 17% false negative on real overeager actions; 0.4% false positive on real traffic | Anthropic engineering; sandbox environments docs |
| Confirmation prompt | A confirmation step | Bypassed by a command-parsing flaw, CVE-2025-54795, CVSS 8.7, fixed in v1.0.20 | GHSA-x56v-x2h6-7j34 |
| Deny rule | Matches the command as written | No published bypass rate | Claude Code security |
| Sandboxed Bash tool | OS-enforced boundary for Bash commands plus child processes | Fails open by default | Claude Code sandboxing |
| MCP servers and command hooks | Separate processes that run unconstrained on the host | Outside the Bash sandbox entirely | Claude Code sandbox environments |
A 17% miss rate on the overeager subset is not the layer you want holding the line for a scheduled job. It is a filter sitting above something else. For the account-level side of unattended automation, see LinkedIn automation in 2026.
Which layer is actually the boundary?
The operating system. In a February 2026 internal red-team exercise, Anthropic had a researcher phish an employee into launching Claude Code with a malicious prompt. The write-up states: "Across 25 retries of that prompt, Claude completed the exfiltration 24 times" (Anthropic, 25 May 2026). Anthropic's conclusion is that the only defence that holds in that situation is the environment, specifically egress controls.
That is a 96% attack-success rate in a controlled exercise, disclosed by the model vendor about its own product.
Three independent sources land in the same place. OWASP's LLM01:2025 entry says it is unclear "if there are fool-proof methods of prevention for prompt injection" and offers measures that "can mitigate the impact" (OWASP Gen AI Security Project, 2025). The UK's National Cyber Security Centre says "there's a good chance prompt injection will never be properly mitigated" the way SQL injection was (NCSC, Dave Chismon, 8 December 2025). The design-patterns literature reduces it to one rule: "once an LLM agent has ingested untrusted input, it must be constrained so that it is impossible for that input to trigger any consequential actions" (Beurer-Kellner et al., June 2025).
CaMeL, from Google DeepMind, Google and ETH Zurich, is the constructive proof plus the honest price tag. A protective system layer outside the model solved 77% of AgentDojo tasks with provable security against 84% undefended (Debenedetti, Shumailov et al., March 2025). Seven points of utility is what containment costs. Anyone selling it at zero cost is selling something else.
Browsers give the clearest illustration. University of Washington researchers tested seven agentic browsers, demonstrated a cross-origin attack on ChatGPT Atlas and identified the preconditions for attacks on Chrome with Gemini, Claude for Chrome and Perplexity Comet. Their finding: "the same-origin policy is reduced to the strength of an agent's defenses against prompt injections" (Paul G. Allen School; UW News, 30 June 2026). A hard boundary became a probabilistic model behaviour.
Containment is not a fix. Anthropic says so in the same documentation: "Sandbox isolation reduces the impact of a breach, but it does not eliminate risk. Any approach that allows network egress can still leak data the agent can read" (Claude Code docs). That is why the 24-of-25 result above matters: Anthropic's conclusion from it was that the environment is what holds. Self-hosting has the same shape, where running it yourself moves the bill rather than removing it.
What should you configure today?
Seven settings, in order. The first matters most because it fails silently: Claude Code's sandboxed Bash tool is OS-enforced and open by default. The documentation states: "By default, if the sandbox cannot start because dependencies are missing or the platform is unsupported, Claude Code shows a warning and runs commands without sandboxing" (Claude Code docs, retrieved 18 September 2026).
Read that twice if you are on Linux. macOS has nothing to install because sandboxing uses the built-in Seatbelt framework. Linux and WSL2 need bubblewrap plus socat. If you never installed them, you have a warning you scrolled past and no boundary at all.
- Turn the sandbox on and make it fail closed. Enable
sandbox.enabled, then setsandbox.failIfUnavailableto true so a missing dependency aborts the run rather than quietly removing the boundary (Claude Code docs). - Declare the file and network allow list explicitly. You define which files and domains commands can touch. The operating system enforces it for every Bash, PowerShell or Monitor command plus its child processes. That last clause is the valuable part.
- Check which mode Codex CLI actually started in. OpenAI's documentation states that by default the agent runs with network access turned off, using Seatbelt through
sandbox-execon macOS andbwrapplusseccompon Linux. The three documented modes are read-only, workspace-write and danger-full-access (OpenAI, retrieved 18 September 2026). A project config can move you off the default without announcing it. - Stop treating deny rules as a boundary. Anthropic's security docs state that "a deny rule matches the command as written" (Claude Code docs). A re-spelled equivalent is not the command as written. Allow lists are hygiene rather than containment.
- Put untrusted content behind a machine boundary. The same page recommends virtual machines for running scripts and making tool calls, especially when interacting with external web services. A separate OS account is the cheaper version.
- Cut what the agent can read. GitGuardian counted 28.65 million new hardcoded secrets in public GitHub commits during 2025, up 34% year over year, with AI-service secrets at 1,275,105 and up 81% (GitGuardian, 17 March 2026). Deny
~/.aws,~/.sshand.envbefore you need to. The documentation is explicit that this needs doing by hand, because "the default read policy still allows them" (Claude Code docs). SetallowUnsandboxedCommandsto false in the same block. A command that fails under the sandbox then cannot simply be retried outside it. - On Linux, know what is underneath. Landlock is a stackable Linux Security Module whose goal is "to enable restriction of ambient rights (e.g. global filesystem or network access) for a set of processes", including for unprivileged ones (Linux kernel documentation).
| What to set | Claude Code | Codex CLI | Why it matters |
|---|---|---|---|
| Sandbox on | sandbox.enabled | sandbox_mode set explicitly rather than assumed | Enforcement moves from the model to the kernel |
| Fail closed | sandbox.failIfUnavailable true | No documented equivalent; verify the mode at start | The default warns you once then runs unsandboxed |
| Network egress | Declare allowed domains in the sandbox config | Off by default | Egress is what turns a read into a leak |
| Platform dependency | macOS Seatbelt is built in; Linux and WSL2 need bubblewrap plus socat | macOS sandbox-exec; Linux bwrap plus seccomp | A missing dependency is the common silent failure |
| Untrusted external content | Vendor recommends a VM | read-only is one of three documented modes | Untrusted input plus consequential actions is the pattern to break |
Sources for the table, all retrieved 18 September 2026: Claude Code sandboxing, Claude Code security, Claude Code sandbox environments, OpenAI Codex agent approvals and security.
None of this is exotic infrastructure work. It is ten minutes of config on hardware you already own, the same arithmetic that makes running models on your own machine pay back.
What the sandbox still does not cover
Two gaps and one deprecation notice. Anthropic's docs state "MCP servers and command hooks are separate processes that run unconstrained on the host" (Claude Code docs, retrieved 18 September 2026). The sandboxed Bash tool alone is insufficient for fully unattended runs. The line to pin above the desk: "With no prompts to catch mistakes, the isolation boundary you choose is what protects your system."
The primitive underneath carries a deprecation notice. The
sandbox-execman page shipped with macOS states that the command is DEPRECATED and points developers at App Sandbox instead. That man page is dated 9 March 2017 and it still ships (SANDBOX-EXEC(1), man page text mirrored by keith.github.io). Check it yourself withman sandbox-exec. OpenAI's documentation states that Codex CLI runs commands usingsandbox-execon macOS (OpenAI). That is the floor of your stack.
Anthropic documents the fix for that gap. It is not a hook. The @anthropic-ai/sandbox-runtime package "wraps an entire process in the same Seatbelt or bubblewrap isolation that the built-in Bash sandbox uses", putting every tool, hook and MCP server inside the boundary rather than beside it (Claude Code docs). It is a beta research preview. A dev container or a VM does the same job with more setup. Reach for one of those first.
We wrote a small tool for a narrower gap, above the sandbox rather than instead of it. agent-guard is an MIT-licensed Python package created on 16 September 2026, covering Claude Code and Codex CLI on macOS, Linux and Windows (GitHub). It adds three things on top of the default runtimes:
- blocks reads of secret paths at the hook
- freezes the session on a detected exfiltration attempt; a frozen session cannot unfreeze itself
- gates persistence changes behind one native OS dialog per session
Its README publishes its own limits, which is the part worth trusting. "Real containment is the runtime sandbox (Claude sandbox.enabled, Codex sandbox_mode) or a separate OS user for untrusted pipelines." The rules match command text. In its words, "a compiled program that opens a file is invisible to them. That's the sandbox's job." It has zero stars at the time of writing. Treat it as one operator's notes in executable form rather than as a recommendation.
What has actually gone wrong on real machines?
Three documented cases. In the most uncomfortable one, nothing defensive stopped it. Amazon Q Developer for VS Code version 1.84.0 shipped to users carrying malicious code. AWS's bulletin records that "the malicious code was distributed with the extension but was unsuccessful in executing due to a syntax error" (AWS, 23 July 2025, CVE-2025-8217).
| Incident | Date | What happened | What actually stopped it |
|---|---|---|---|
| Amazon Q Developer for VS Code v1.84.0 | July 2025 | Shipped to users carrying malicious code; AWS states it "prevented the malicious code from making changes to any services or customer environments" | A syntax error in the attacker's payload |
Nx s1ngularity | August 2025 | A postinstall script drove locally installed AI CLI agents as a file-search agent to inventory credential files, then exfiltrated to attacker-named public GitHub repos | Not a model defence. The compromise was in the package; the advisory records no known CVE |
| Claude Code CVE-2025-54795 | August 2025 | A command-parsing flaw let untrusted content already in the context window bypass the confirmation prompt and execute a command | A patch, in v1.0.20 |
The Nx case is the one to sit with. The attacker did not need to compromise the agent. A malicious package ran on the developer's machine and used the agent's own filesystem reach as the payload, with an instruction that began "You are a file-search agent" and asked for an inventory of configuration and environment files (GHSA-cxm3-wv7p-598c, 27 August 2025). The advisory adds that the script modified .zshrc and .bashrc to trigger system shutdowns. No model was jailbroken. A trusted local tool was pointed at the disk.
CVE-2025-54795 makes the narrower point. Rated High at CVSS 8.7, it let untrusted content in the context window bypass the confirmation prompt through a parsing error, fixed in v1.0.20 (Anthropic, 1 August 2025). A permission prompt is a software component with a CVE number rather than a boundary. Boundaries are enforced by something that does not parse your input.
FAQ
Six questions are answered in the schema above. For the workflow-selection side of the question, start with the Ops Automation Playbook.
The bottom line
Be generous to the model layer. Going from 5.5% to 2.0% attacker success at fifteen attempts is a genuine result, measured against 1,130 real attacks and published by the vendor with the losing number left in.
Then be precise about what it is. A probability improves. A boundary holds or it does not. The gap between a 2.0% benchmark score and an up-to-80% working chain is not a scandal. It is the definition of the word benchmark. An unattended agent is the machine that finds the attacks outside it.
Most people have not arrived here yet. Stack Overflow's 2025 developer survey found 52% of developers "either don't use agents or stick to simpler AI tools", with 46% actively distrusting AI accuracy against 33% who trust it (Stack Overflow). It does not measure unattended operation. Read that as agent use in general rather than a count of scheduled pipelines.
So the rule is short. Turn the sandbox on. Make it fail closed. Turn egress off by default. Cut read access to your credential paths. Put anything touching untrusted external content behind a separate OS account or a VM. Then accept the part that stays true afterwards: the boundary shrinks the damage rather than removing it.
Next reads: the SaaS replacement playbook for the build-versus-buy frame. Then what is actually safe to automate on LinkedIn, which is the same question applied to an account you cannot sandbox.
Sources (retrieved 2026-09-18):
- Anthropic, How we contain Claude across products: https://www.anthropic.com/engineering/how-we-contain-claude
- Anthropic, Mitigating prompt injections in browser use: https://www.anthropic.com/research/prompt-injection-defenses
- Anthropic, Transparency Hub, Claude Opus 5 External Red Teaming: https://www.anthropic.com/transparency
- Anthropic, How we built Claude Code auto mode: https://www.anthropic.com/engineering/claude-code-auto-mode
- Claude Code docs, Choose a permission mode: https://code.claude.com/docs/en/permission-modes
- Claude Code docs, Choose a sandbox environment: https://code.claude.com/docs/en/sandbox-environments
- Claude Code docs, Configure the sandboxed Bash tool: https://code.claude.com/docs/en/sandboxing
- Claude Code docs, Security: https://code.claude.com/docs/en/security
- OpenAI, Codex agent approvals and security: https://developers.openai.com/codex/agent-approvals-security
- OWASP Gen AI Security Project, LLM01:2025 Prompt Injection: https://genai.owasp.org/llmrisk/llm01-prompt-injection/
- UK NCSC, Prompt injection is not SQL injection (it may be worse): https://www.ncsc.gov.uk/blog-post/prompt-injection-is-not-sql-injection
- Zhan, Fang, Panchal, Kang, Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents: https://arxiv.org/abs/2503.00061
- Debenedetti, Shumailov et al., Defeating Prompt Injections by Design (CaMeL): https://arxiv.org/abs/2503.18813
- Beurer-Kellner et al., Design Patterns for Securing LLM Agents against Prompt Injections: https://arxiv.org/abs/2506.08837
- University of Washington, Agentic Browsers and the Same-Origin Policy: https://agent-security.cs.washington.edu/agentic_browsers_sop.html
- UW News, Some agentic AI browsers come with major cybersecurity risks: https://www.washington.edu/news/2026/06/30/some-agentic-ai-browsers-come-with-major-cybersecurity-risks-uw-study-finds/
- AWS Security Bulletin AWS-2025-015: https://aws.amazon.com/security/security-bulletins/AWS-2025-015/
- GitHub Security Advisory GHSA-cxm3-wv7p-598c (Nx): https://github.com/nrwl/nx/security/advisories/GHSA-cxm3-wv7p-598c
- GitHub Security Advisory GHSA-x56v-x2h6-7j34 (Claude Code, CVE-2025-54795): https://github.com/anthropics/claude-code/security/advisories/GHSA-x56v-x2h6-7j34
- Johann Rehberger, Breaking Claude Code Opus 5 Auto Mode with Indirect Prompt Injection: https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/
- SANDBOX-EXEC(1), macOS man page text mirrored by keith.github.io: https://keith.github.io/xcode-man-pages/sandbox-exec.1.html
- Linux kernel documentation, Landlock: https://docs.kernel.org/userspace-api/landlock.html
- Stack Overflow 2025 Developer Survey, AI: https://survey.stackoverflow.co/2025/ai
- GitGuardian, The State of Secrets Sprawl 2026: https://blog.gitguardian.com/the-state-of-secrets-sprawl-2026/
- GitHub, dancolta/agent-guard: https://github.com/dancolta/agent-guard
Frequently asked questions
Is prompt injection solved in 2026?
No. Anthropic's own research page states that no browser agent is immune to prompt injection and that it publishes findings to demonstrate progress rather than to claim the problem is solved. On the Gray Swan benchmark of 1,130 high-transferability attacks, Claude Opus 5 reduced attacker success within 15 attempts from 5.5% to 2.0%. That is real progress rather than a fix. OWASP's LLM01:2025 entry states it is unclear whether fool-proof methods of prevention exist at all. Treat the model layer as a probability and the operating system as the boundary.
What is the difference between an AI agent sandbox and a permission prompt?
A permission prompt or classifier judges one action at a time. Anthropic's Claude Code documentation calls the auto mode classifier a per-action control rather than an isolation boundary. The engineering post that introduced auto mode publishes a 17% false-negative rate on real overeager actions alongside a 0.4% false-positive rate on real traffic. Claude Code's sandbox is enforced by the operating system instead: macOS uses the built-in Seatbelt framework and Linux uses bubblewrap. The boundary applies to every Bash command plus its child processes. One is a judgement call. The other is a rule the process cannot argue with.
Does Claude Code's sandbox fail open by default?
Yes. Anthropic's documentation states that if the sandbox cannot start because dependencies are missing or the platform is unsupported, Claude Code shows a warning and runs commands without sandboxing. On macOS there is nothing to install because sandboxing uses the built-in Seatbelt framework. On Linux and WSL2 it depends on bubblewrap plus socat, which plenty of machines do not have. Set `sandbox.failIfUnavailable` to true so a missing dependency stops the run rather than silently removing the boundary you assumed was there.
Does OpenAI's Codex CLI block network access by default?
Yes. OpenAI's Codex documentation states that by default the agent runs with network access turned off. The sandbox uses Seatbelt policies through `sandbox-exec` on macOS and `bwrap` plus `seccomp` on Linux. Three modes are documented: read-only, workspace-write and danger-full-access. Inside the default boundary the agent can read files, edit within the workspace and run routine local commands. For unattended runs the thing worth checking is which mode the session actually started in, because a config file can move you off the default without telling you.
Why is indirect prompt injection worse for unattended agents?
Because an unattended agent retries and a person does not. Anthropic reports Claude Opus 5 attacker success within 15 attempts at 2.0% across a 1,130-attack benchmark, which is a figure about a campaign rather than a guarantee about a run. A scheduled pipeline that ingests external content every hour reaches the 15th attempt and keeps going. Separately, researchers at UIUC evaluated eight published defences against indirect prompt injection and bypassed all of them using adaptive attacks, reporting attack success rates consistently over 50%.
Can a sandbox stop data exfiltration?
It reduces the blast radius rather than removing the risk. Anything that still permits network egress can leak whatever the agent is able to read. Anthropic's containment post describes a February 2026 internal red-team exercise in which Claude Code completed a credential exfiltration in 24 of 25 retries. Its conclusion is that the defence that holds in that situation is the environment, specifically egress controls. Claude Code's security documentation also recommends virtual machines for scripts and tool calls that touch external web services. Cut read scope as well: deny `~/.aws`, `~/.ssh` and `.env` paths.

