August 15, 2026
Prompt injection in agent workflows, explained
Prompt injection gets talked about as a trick: someone hides "ignore your previous instructions" in a web page and the chatbot says something embarrassing. That framing makes it sound like a content problem, which makes it sound like something a filter can fix.
It is not a content problem. It is a missing boundary, and once you see where the boundary is missing, the whole thing stops being a curiosity and becomes an architecture question you can actually answer.
This is the calm version. No attack demo, no scary terminal. Just the model, and what to build.
The one idea
An agent reads everything as one stream of text. The instructions you gave it, the file it just opened, the web page it fetched, the error message it got back: all of it arrives in the same context window, in the same format, with nothing marking which part is trusted.
Every other system you secure has a line between the control plane and the data plane. A SQL database knows the difference between the query and the values, once you use parameters. A shell knows the difference between the command and the argument, once you quote it. That line is what makes injection preventable in those systems.
An agent has no such line. There is no parameterized version of a prompt. So anything the agent reads can attempt to become an instruction, and the model decides, on vibes, whether to follow it.
Where the untrusted text actually comes from
Look at the diagram at the top. The left box is the part people underestimate. In a development workflow the agent reads far more than a chat message:
- Issue and pull request comments, including from outside contributors
- Web pages and documentation it fetches while researching
- Dependency contents: a README, a changelog, a package description, a comment in vendored code
- Tool output: logs, stack traces, test failures, API responses
- The repository itself: filenames, commit messages, code comments
- Documents and tickets pulled in through connected tools
Every one of those is written by someone who is not you. A CI log can contain an attacker-controlled string, because the attacker controlled an input that got logged. That is the actual attack surface, and it is much larger than "a user typed something clever."
Why the obvious defenses do not hold
Telling the model to ignore injected instructions. This is the most common attempt and the weakest. You are asking the same component that cannot tell instructions from data to please tell instructions from data. It raises the effort required and it stops nothing determined. A trust boundary cannot be created by requesting one.
Filtering for suspicious phrases. Injection does not need a magic phrase. It can be polite, plausible, and phrased as helpful context. Filters catch the examples in blog posts.
Detection by the agent. Worth having, worth not relying on. Here is a real example from ZuluSec's own build loop, which is worth telling because of how it turned out. An agent working on this site flagged what it believed was an injected instruction telling it to conceal something from the operator. It refused to comply and reported it. On investigation, the text was a genuine message from the tooling, and the agent's account of where it came from was wrong.
The detection was unreliable in both directions: it fired on something benign, and it misdiagnosed the source. What held was not the detection. What held was the standing rule that an instruction to hide something from the operator is never followed no matter which channel it arrives on, and that anything suspicious gets surfaced rather than resolved quietly. The agent did not have to be right for the outcome to be right.
That is the whole lesson in miniature. Design so the model's judgment is not the control.
What actually holds
Assume the agent will eventually follow a malicious instruction. Build so that when it does, nothing irreversible happens.
The agent proposes, it does not dispose. Its output is a diff, not an action. A compromised agent writes a bad pull request, which is a nuisance. A compromised agent with deploy credentials is an incident. Keep the destructive verb out of the agent's hands and in a pipeline that a human triggers.
Nothing reaches production without a gate the model does not control. Tests, checks, and a human reading the diff. The important word is diff: a reviewer looking at a change sees what it does, whatever the agent believed it was doing.
Note that merging is not shipping, and being precise about which one you are gating is worth the pedantry. ZuluSec's own reference repository merges first and reviews after, with each round of review findings landing as its own pull request, because review there is by a single practitioner and nothing in that repository reaches production. Where agent-written code does reach production, the gate belongs before the deploy, and it should be a human who triggers it.
Least privilege, and short-lived. The token the agent holds should be read-only wherever reading is enough, scoped to what the task needs, and expired soon. Most of the damage in a bad agent run comes from credentials that were broader than the job.
Allowlist the reach. Which tools it can call, which hosts it can talk to. Injection usually needs an exit: a place to send what it collected, or a system to change. Removing the exit removes most of the value in compromising the agent at all.
Keep the security decision deterministic. If the model decides whether something is a finding, an injected instruction can change the finding. If a check with a fixed rule decides, and the model only orchestrates the check and writes the report, then the injected instruction has nothing to move. This is the difference between an agent that does the work and an agent that builds and runs the thing that does the work.
There is a stricter tier of the same idea, and it is the one to reach for when the data itself must never reach a model. Put the orchestration into the tooling as well, so the agent builds what runs and then stays out of the loop entirely. Nothing the tooling reads goes back to a model, which means an injected instruction has no channel to arrive on in the first place. It costs flexibility, and in a regulated environment that is a trade worth making.
Log the run so it can be read afterward. What was fetched, what was called, what was changed. Not because it prevents anything, but because the first question after a strange result is what the agent actually saw.
The principles underneath it
- Anything the agent reads may be hostile. Including your own logs.
- The model is not a trust boundary. Put the boundary somewhere deterministic.
- Bound the blast radius, not the input. You cannot enumerate the bad inputs. You can enumerate what the agent is allowed to do.
- Never let the agent hide from the operator. Concealment is the one instruction that is always refused, whatever the source.
Where this fits
If you are letting an agent write code that reaches production, this is the design question underneath the whole thing, and it is answerable. The controls are ordinary: least privilege, a review gate, deterministic checks, an audit trail. Nothing here is exotic. What is new is remembering to apply them to a component that reads attacker-influenced text all day and cannot tell the difference.
That is how the automation engineering work is built and what it is built to withstand, and a security audit is the place to start if you want to know where you are exposed today.
Want automation built so this cannot bite you?
Automation engineering builds the tooling with the gates already in place: deterministic checks, least privilege, a review step, and an audit trail. There is a public reference implementation you can read before talking to anyone.
See how it is built