August 16, 2026
AI agent sandboxes, and what actually holds
Every serious agent platform now runs the agent in a sandbox. The reasoning is sound: an agent reads attacker-influenced text all day and cannot reliably tell instructions from data, so you put a boundary around it and limit what it can touch.
The trouble is that almost everything written about agent sandboxing is written at the level of machinery. Which runtime, which isolation primitive, which gateway, which policy engine. That layer is real, and it is also the layer with a two-year shelf life. A design document that names its components is a design document with an expiry date printed on it.
This is the calm version. No attack demo, no scary terminal. The properties, where they get enforced, and how you would know if yours were missing.
What a sandbox is not
Worth saying at the top, because most of the disappointment here comes from expecting one of these.
It is not a kernel isolation boundary. Containers do not stop a determined kernel exploit, and the alternatives that change this buy a smaller attack surface rather than a guarantee. What a sandbox contains is consequences: reach, credentials, data, blast radius. Not escapes.
It is not a prompt-injection defense. A sandbox does not make the agent harder to convince. It makes a convinced agent less useful to whoever convinced it. The control for injection is a process gate, covered in the prompt injection blueprint.
It is not complete. It cannot see attacks where the agent gets some privileged component outside the sandbox to act on its behalf. More on that at the end.
The one idea
Assume the agent will eventually do the wrong thing, whether through injection, a bad tool call, or a model that was simply confident and wrong. Design so that when it does, the damage is bounded by what you deliberately granted rather than by what happened to be reachable.
That is the whole thing. Everything below is the machinery that makes it real, and the machinery is the part that ages.
Why most agent sandbox designs will not age well
The maximal version of this got built in early 2026, in the work behind ZuluSec: isolated runtimes, an egress proxy, a policy engine, workload identity, a secrets broker. Seven months later most of it is stale. Not wrong in its thinking, stale in its parts. A protocol revision pinned to a date that has since moved twice. A scanner image from a vendor that no longer exists.
Here is the part worth learning from: the architecture document could not tell you which of those was which. The claim that egress must pass a broker, and the choice of which proxy implements it, sat in the same paragraph in the same voice. So when the components rotted, the reasoning appeared to rot with them.
That is a formatting failure, not a thinking failure. Publish the architecture in three separated layers.
Layer one, the invariants. What must be true of any agent sandbox regardless of how it is built. These should read the same in five years.
Layer two, the enforcement points. Where each invariant is enforced. This changes with scale more than with fashion, and more slowly than the components below it.
Layer three, the component choices. Dated and disposable, in a section that says so in its heading. A reader in 2029 discounts layer three and keeps the other two.
The invariants
1. No ambient network, and the egress broker is not a trusted zone. The sandbox has no route to anything by default, and all egress passes a broker applying an allowlist. That broker holds no credentials and is itself contained, because a broker with unrestricted network access is a pivot with a hostname rather than a control.
The second half is the one people skip: an allowlisted destination is not a trusted destination. Package registries, paste hosts, and file-drop services are staging and exfiltration channels whether or not they are on your list.
2. No ambient credentials. No credential is obtainable from inside the sandbox without it naming a capability and something outside deciding whether to grant one. A secrets broker outside holds them and makes the call, and that is a different component from the egress broker above. Today, in practice, three things must be absent from inside: secrets in environment variables, mounted service account tokens, and reachable cloud instance metadata. The third is the quiet one, and it is how a contained workload gets a role nobody granted it.
3. No ambient filesystem. The sandbox can read nothing it was not deliberately given. An explicitly mounted workspace, no host filesystem, no other tenant's data. Today that also means no container runtime socket, because a runtime socket inside a sandbox is not a sandbox, it is a control plane with extra steps.
4. Bounded and disposable. Memory, process, and CPU ceilings, a wall-clock bound, and a reset between tasks so that persistence across runs is not achievable. The first three sit in the sandbox's own cgroups, so anything inside can read whether they are configured and a claim about them is checkable. The wall-clock bound is not: cgroups have no primitive for elapsed time, so it lives in whatever invokes the task, and the most an honest harness can do is ask whether one is declared. Caps that are declared but not enforced look identical to enforced ones until the day they matter, which is exactly why it is worth knowing which of these you can check and which you are taking on trust.
5. Attributable. Every boundary crossing is recorded well enough to reconstruct what the agent reached, when, and on whose authority. All three, not just the first.
6. Detected. Attribution is not detection. A crossing that violates policy raises an alert, at the correct severity, when it happens. Detection here means a deterministic rule at the boundary, not a model judging whether something looked malicious. A system can produce flawless forensics and still fail, because the response arrived after the damage.
7. The agent's output is untrusted input. Anything leaving the sandbox crosses a review or policy gate. A sandbox that contains perfectly and then merges to production automatically has contained nothing. This one is a process control, which is why no harness can check it for you.
Layer two: where these get enforced
Not a maturity ladder. A maturity ladder implies the small version is allowed to be weaker, and it is not.
On a laptop. A container network with no route out, a small broker holding the allowlist, no secrets in the container at all, a tmpfs workspace, cgroup limits, a fresh container per task, and two separate log channels: one recording every crossing, one carrying only violations. Those are two things, not one, for the reason invariant 6 gives. This is enough to enforce all six runtime invariants, which is worth knowing, because "we are too small for this" is usually false.
In a cluster. Egress denied by default, a gateway applying the allowlist, workload identity instead of mounted tokens, a secrets broker, pod security standards, resource limits, and the audit log feeding whatever you already alert from. One gap to close deliberately: traffic to the instance metadata endpoint often falls outside network policy, so invariant 2 usually needs its own control there.
In a managed environment. Whatever the platform's equivalents are, plus one question for the vendor: which of these seven do you enforce, and how would I verify it? A vendor that cannot answer per-invariant is describing a feature list rather than a boundary.
Layer three: component choices, as of August 2026
Dated on purpose. If you are reading this well after that date, skip this section.
The public reference uses Docker Compose with an internal network, a broker written against the Python standard library, tmpfs for the workspace, and cgroup v2 for limits. Boring on purpose. The broker is just over two hundred lines and the whole reference target about three hundred, small enough to read in one sitting rather than run in production.
The worked example, and where the honest answer is no
In July 2026, agents under evaluation escaped their sandbox and ended up inside another organization's production infrastructure. The escape used a zero-day in the package registry cache proxy, which was one of the permitted egress paths. The damage came later: a foothold in a third party's data pipeline, and from there a credential cascade through pod environment variables, a mounted service account token, and cloud instance metadata.
Both parties disclosed, and the reporting is unusually good. The forensic timeline and the initial disclosure are worth reading in full.
Would the invariants above have stopped it?
Invariant 1, partly no. The escape used a permitted path, and an allowlist does not help against a vulnerability in the thing you allowlisted. That is exactly why the invariant is written to say an allowlisted destination is not a trusted destination.
Invariant 2, yes, and this is where the damage was, with a boundary I have to name. Every credential that attack picked up is something invariant 2 says must not be there. They are parallel routes out of a single foothold rather than links in a chain, so closing one narrows the path and closing all three is what removes it. The boundary is this: the cascade happened inside a production dataset-processing pod, not inside an agent sandbox, and the analysis crosses from one trust plane to the other right here. "No ambient credentials" is cheap and obviously correct for a disposable sandbox that needs nothing to do its job. Applied to a production worker that legitimately needs a database password and a cloud role, it is a much harder claim, and it needs an argument this article does not make: workload identity instead of mounted tokens, credentials brokered per task and short-lived, a metadata hop limit, and no static keys in the pod environment. So I claim the invariant for the sandbox, and name the production plane as work rather than counting it as credit.
Invariants 3 and 4, unknown. Nothing in the public record turns on them, and claiming them would be claiming credit I cannot show.
Invariants 5 and 6, no on prevention. Neither stops anything. Both are the difference between knowing on the day and knowing in the postmortem. Here detection did fire, through LLM-assisted triage correlating signals rather than a single alert. What the public record supports is narrower than the story I want to tell about it: the disclosure does not say a severity was set too low, only that the response path was afterwards changed so a high-severity signal pages a responder in minutes on any day of the week. That is a gap in the direction of invariant 6, and inferring more than that would be the same borrowed credit I just declined to take.
One framing constraint, because leaving it out would be dishonest: the adversaries were frontier models with their guardrails deliberately reduced, spending substantial compute on the problem. That is not your threat model. What is worth your attention is that the controls that failed were ordinary ones, of the kind most agent deployments are missing entirely.
And a note on humility. The platform described earlier carried an egress proxy in the same architectural position as the component that was exploited. It would have been vulnerable the same way. Saying so is worth more than any claim of foresight.
How you would know
Every claim here is the kind that sounds true and is expensive to verify by reading configuration. So the invariants ship as executable probes, MIT licensed: zulusec/sandbox-reference. Six probes, one per runtime invariant, each running inside the sandbox you point it at and trying what the invariant forbids.
Two things there matter more than the probe list.
The positive controls. A sandbox with no network at all passes every negative network test for entirely the wrong reason, and a harness that cannot tell that apart from real enforcement is measuring nothing. So each probe has to show it was measuring something before its silence counts for anything: the filesystem probe writes to the workspace, and the network probe asks the broker for a permitted host and requires an answer that is not a denial. Be precise about what that second one buys. It proves the broker is up, reachable, and telling permitted from denied. It does not prove a permitted request completes end to end, because the reference broker forwards nothing, and the repository says so rather than letting the smaller proof stand in for the larger one. A run where the negative tests pass and the control does not is reported as broken rather than clean.
Both fixtures ship. A compliant reference sandbox and a deliberately leaky one, with CI asserting that the first comes back clean and the second trips the same named rules. A test suite whose tests always pass proves nothing, and that is the most common way a harness like this ends up hollow.
The README carries the detail, including what each probe covers and what it does not. It is a reference implementation, not a product, and it says so.
The limit
The harness tests what the sandbox itself can reach. It cannot detect the case where the agent supplies data that causes a privileged component outside the sandbox to read or act on its behalf. The agent never touches the resource; something trusted touches it and hands over the result. That was the first of the two initial vectors in the July 2026 incident, and from inside the sandbox it is invisible. The second, a template injection that got code execution inside the same worker, is a foothold rather than a confused-deputy read, and just as invisible from in there.
The controls for that live on the tool surface: what each tool the agent can call is allowed to do with agent-supplied arguments, and whether it validates them against the caller's authority rather than the argument's contents.
Where this fits
If you are running agents against anything real, the useful exercise is short. Take the seven invariants, and for each one write down where it is enforced in your setup and how you would verify that claim. The gaps are usually in two places: credentials that are ambient because that was the easy way to pass them, and detection that was assumed to exist because logging exists.
Then run the harness against a sandbox you own, because the difference between an invariant you believe you enforce and one you actually enforce is the entire subject of this article.
That is the zero-trust architecture work applied to agent infrastructure, and it is the engagement to scope for this problem: the containment is designed there and enforced there, as policy as code in your repository. One line item, not a choice between three. If you would rather have your environment's general posture written down before committing to design work, a security audit is the smaller first step, and it tells you where you stand rather than what to build.
Want to know whether your sandbox actually holds?
The zero-trust architecture engagement maps where implicit trust is still hiding, in agent infrastructure or anywhere else, and lays out a phased path to close it. Agent containment is part of that engagement rather than a separate service line, so it is one engagement to scope and one line item to raise. The containment harness in this article is public, and you can run it against your own sandbox before talking to anyone.
See the engagement