LESSON · 11 SEPTEMBER 2026
Indirect prompt injection and why agent controls sit at the action layer
The verdict
Indirect prompt injection hides instructions in content an agent reads, such as a web page, an issue or a file, and the agent then acts on them with its user's access. Because a model cannot reliably tell data from instructions, the dependable control point is the action itself: whether this agent, in this session, may run this command or call this tool.
By Best AI Security editors · 11 September 2026 · 3 min read
What is indirect prompt injection?
Prompt injection is input written to override a model's instructions. In the direct form, the user types it. In the indirect form, it arrives inside content the model processes on the user's behalf: a README in a repository, a comment on a ticket, a page the agent browsed, or the output of an MCP tool. A chatbot that reads such content might give a wrong answer. An agent that reads it can take a wrong action, because it has tools.
Why is it worse for agents?
Agents run with the permissions of the person who started them. A coding agent on a developer laptop can typically read files, run shell commands, install packages and use the cloud CLI profiles already on the machine. The OWASP GenAI Security Project's 2026 Top 10 for LLM Applications, announced on 1 September 2026, ranks Excessive Agency third: an agent with more functionality, permissions or autonomy than its task needs. Injection supplies the intent; excessive agency supplies the reach.
What does Bay's Ghostjacking write-up describe?
In August 2026 Bay published research it calls "Ghostjacking": indirect prompt injection that turns a trusted agent's legitimate permissions into the attack path. Bay describes the conditions that make it work as "broad tools, broad credentials, weak approval" and calls the attack "an authorization problem disguised as a prompt problem." Its point for defenders is that endpoint tooling sees a trusted coding agent doing ordinary things, so the question has to be asked at the level of the agent's action.
Where can a control stop it?
- Before the content arrives: limit which MCP servers and tools an agent can load, so fewer untrusted sources reach it. See MCP security.
- At the action: evaluate each tool call with its arguments and session context, then allow, ask or deny. This is the step that still works when the injected text looks harmless.
- After the result: check tool output before it goes back to the model. The MCP tools specification asks clients to validate tool results before passing them to the model.
- In the record: log whether a person prompted the action or the agent chose it after reading outside content.
What context helps a decision?
A rule that sees only the command cannot tell a developer's request from an injected one. Useful context includes what the agent read just before the action, what it did earlier in the session, which data or credentials the action touches, and whether the user was present. Bay says its decisions weigh identity, prior actions and data accessed. Noma Security describes inspecting the event, the session, the identity behind the agent and the data it reaches. Onyx Security inspects every prompt, tool call and model response inline. Our note on allow, ask or deny policies turns this into rules.
What should a buyer ask?
- Show a demonstration in which an agent reads a file containing hidden instructions, and what your product does when the agent tries to act on them.
- Does the decision use session history, or only the current command?
- Does the audit record show that an action followed outside content rather than a user prompt?
Next lesson
Related
Sources
- Bay, Ghostjacking (Aug 2026) · Reviewed Sep 2026
- OWASP GenAI Security Project announcement (1 Sep 2026) · Reviewed Sep 2026
- MCP tools specification · Reviewed Sep 2026
- Noma Security, endpoint agents · Reviewed Sep 2026
- Onyx Security, AI security · Reviewed Sep 2026