September 2, 2026
A reference architecture for agents that work the way your people do
Most organisations running coding agents today cannot answer four questions about any given run. Where did it execute. Whose credential did it use. What did it reach. What is left behind. The agents are real and the work is real; the platform underneath them is usually a developer's laptop, a shared token, and hope.
This is a reference architecture for the other way round. It describes a platform ZuluSec built over four days spanning the end of August, on the sandbox blueprint's invariants, to find out what those invariants make possible once you stop treating them as constraints. The implementation stays private. The shape is the point, because the shape is what any organisation can have.
What it feels like to use
The home page asks one question: what do you want done? You pick a workspace and answer in plain words. A terminal opens and you watch an agent do it, in a machine that did not exist a moment ago and will not exist once it is finished. When it pushes a commit or opens a pull request, the service shows your name, because it was done as you.
Everything else on the platform is that same question with the words already written:
- Notes are a repository of markdown drawn as a graph of what links to what, with a reader beside it, and an agent that can be asked about any of it.
- Runbooks are the procedures in that repository shown as cards. Each has a button that hands the procedure to an agent with the repository cloned.
- The queue is the open issues at the connected tracker, read with your own account, each one something an agent can be given.
- Code is the repository as work in progress: ask for a change and a sandbox branches, commits, pushes, and opens a pull request that waits for a person.
And a workspace can work on its own. A release note every Friday and a link check every night are two scheduled runs, each with its own words and its own clock, each run as the person who owns it, with the same boundary and the same record as a session somebody is watching.
Any agent works the same way. Claude Code, Codex, and Gemini CLI are declarations: how to install one, what credential it signs in with, where it reads its instructions, and what is worth keeping when the session ends. Swapping one agent for another changes none of the boundary.
That is the art of the possible in one paragraph: agents that work the way your people do, as your people, inside a line you drew, with a record you can read. The rest of this article is how the pieces fit.
The five planes
Look at the diagram at the top. Five planes, each with one job, and a session moves through all of them.
Identity. Local accounts and roles, plus an OpenID Connect provider, so the same people sign in to the platform and to the applications around it. Connected apps are the other half: a person links their account at the git service once, through ordinary OAuth2, and from then on anything the platform does there is done as them. There is no platform-wide service account. An agent whose credential nobody has saved does not start.
Workspace. A repository and an agent, presented as whatever the work is. Instructions come in three layers: the platform's, which every session carries; the workspace's; and the person's own. Where each lands inside the sandbox depends on the agent, so one set of words serves whichever agent a person runs. What an agent learns in a session comes back the next time that person works in that workspace, and only for that person, because notes steer an agent the way instructions do and one person's should not turn up in another's sandbox.
Execution. A session is a microVM with its own kernel, booted from a read-only image that has the agents installed into it. It gets a bounded workspace disk, a read-only input disk, an output disk with a ceiling, a fixed memory and CPU budget, and a wall clock enforced from outside. When the agent finishes, or nobody has been attached for ten minutes, or the hour is up, the machine is destroyed. Agents are installed into the image and never into a running machine, and an updater has nowhere to write, because a disposable sandbox has no business updating itself.
Boundary. The session has no network device. It has two roads out, and neither is a network interface. The first is an egress proxy on the host that allows a destination by exact hostname and port, resolves names on the host so the guest never resolves anything, refuses any name that points at private address space, and records every attempt whether it was allowed or not. The second is the platform itself, over a channel that only that one sandbox can open. On that channel the session presents a token minted for it alone, naming one repository at one service and dying with the session, and the platform swaps it for the person's real account at the edge. The person's credential for the service is never inside the sandbox. Nobody else's is either.
Evidence. Every consequential act appends to a hash-chained, append-only log: sessions starting and stopping, each outbound connection and its verdict, each use of somebody's credential at a service, each scheduled run added or refused, each host allowed or revoked, each image rebuilt, each instruction set, each administrative change. A session's start record carries the digest of the image and kernel it booted, the hosts it was allowed, and the names of the secrets it was given, never the values. One command walks the chain and says whether it is intact.
Two flows through the planes
A runbook, handed to an agent. A person opens the runbooks view of a workspace and presses the button on "rotate the staging certificates." Identity confirms who they are and that they may open this workspace. Workspace assembles what the agent needs to know: the platform's standing instructions, the workspace's, this person's own, and the runbook text as the first prompt. Execution builds the disks and boots a fresh machine with the repository cloned into the workspace disk. Boundary opens the proxy with exactly the hosts this agent's vendor needs plus what the workspace named, and the platform channel with a token for this one repository. The agent works; the person watches the terminal. It pushes a branch and opens a pull request, and the service records the push and the pull request as that person. The machine is destroyed, the disk that carried the agent's credential is overwritten and unlinked, and Evidence holds the start, every connection, the credential use, and the stop, under one run id.
A nightly run nobody watches. At two in the morning the scheduler starts the "check every link in the docs" run as its owner. Everything above happens again, with two differences. There is no terminal and no permission prompts, by design: an agent that stops to ask with nobody there does nothing at all, and the boundary is what makes not asking defensible. And an agent with no way to answer once and exit cannot be scheduled; the workspace page says so rather than hanging. In the morning there is a pull request with the broken links fixed, a transcript, and the evidence. Or there is a failed run with its transcript and evidence, and the next night comes round again.
The rules that make it hold
Five decisions carry the architecture, and every one of them is portable to a platform you build yourself.
The boundary, not the prompt. Nothing here relies on the agent behaving. It relies on what the agent can reach, what it holds, and what is written down. An agent that follows a malicious instruction produces a bad pull request, which is a nuisance, rather than an incident. The prompt injection blueprint is the longer argument for why that is the only place the control can live.
Everything belongs to somebody. No shared service account, no platform credential an agent could spend on someone else's behalf. Every act at an external service is attributable to a person because it was done with that person's account, at the edge, by the platform.
Disposable by default. The machine, its home directory, and the disks that carried anything sensitive are gone at the end. Persistence is a deliberate act: what the agent learned is taken out on purpose, per person, and the work product leaves as a push.
Agents are declared, not built in. A manifest says how to install one and what it needs. The boundary does not know which agent is inside it, which is what makes it a boundary.
The honest list is part of the design. The platform's own documentation has a section called "What is not isolated." The agent's vendor credential is inside the sandbox, because the agent needs it; the proxy checks names and ports, not content; a permitted host is permitted for everything it serves. Writing those down is not a weakness in the architecture. It is the architecture telling you where to look.
Components, as of September 2026
Dated on purpose, following the blueprint's own rule: invariants and enforcement points age slowly, component choices do not. The build under this article is Rust, with Firecracker microVMs on KVM for execution, a self-hosted git service reached through the platform, an HTTP CONNECT proxy for egress, and a browser terminal bridged to the guest's serial console. Firecracker's own design document says it does no network filtering, which is exactly why the egress plane exists as a separate thing. Roughly twelve thousand lines, one node, one organisation.
Three things a production deployment adds, named here so nobody mistakes the reference for the finished article: running the VMM under Firecracker's jailer, which its design document says production should always do; an alert on a refused connection, so that detection means knowing on the day rather than in the postmortem; and an external anchor for the evidence log, so a truncated tail is as detectable as an edited middle.
Where to start
You do not need all five planes on day one, and the order matters.
- Start at the boundary. One sandbox with no network device and an egress proxy that allows by name is most of the safety and none of the product. The public harness tells you whether what you built actually enforces it.
- Then the credential edge. Get the person's service credential out of the sandbox and behind a per-session token. This is the change that turns "an agent did something" into "this person did something, through an agent."
- Then the evidence. Hash-chain the log before there is much in it.
- Then the workspace shapes and the scheduler. These are the product. They are also the easy part once the three above are true, which is the whole argument of this article.
Where this fits
Deciding where the lines go, what enforces each one, and how you would know is the zero-trust architecture work applied to agent infrastructure. This shape is one answer. Whether it is your answer depends on what your agents already reach today, and a security audit is the smaller first step if you would rather have that written down before designing anything.
Want this shape in your environment?
The zero-trust architecture engagement is where an agent platform's boundaries get decided and enforced: what a session may reach, whose credential it carries, what is written down when it crosses a line, and a phased path from the agents your team already runs to this shape. Containment is one line item inside that engagement, not a separate product.
See the engagement