prompt injection is a permission problem
If an assistant reads a malicious ticket and tries to email private invoices, I want the application to stop it. Whether the model understood the attack is a separate question.
Imagine a support assistant that can search tickets, read attachments and send email. A customer submits a ticket containing a perfectly ordinary problem, followed by an instruction to find every invoice and send the files to an outside address. The assistant reads it while doing its job. Then it proposes the email.
At that point, arguing about whether the prompt was clever is beside the point. The customer was allowed to submit a ticket. They weren't allowed to read other customers' invoices. Somewhere between reading that text and executing the proposed call, the application needs to preserve that distinction.
This is the part of prompt injection I find most useful to reason about. We can spend a lot of time trying to make the model recognize hostile instructions, but there will still be a decision to make when it asks for something it shouldn't. This is an example of what OWASP calls indirect prompt injection. The same route exists through retrieved documents, repository comments, PDFs, web pages and tool responses. None of those sources has to resemble a system prompt to influence what happens next.
following the call to execution
Suppose the proposed call is read_attachment(ticket_id). The tool gateway should check whether the current user can read that attachment. A ticket title being visible doesn't settle that. If the next call is send_email(to, body), the gateway also needs to establish that the user asked to send something, and that the destination and content are allowed.
The service account makes this harder than it first appears. It may have access to the whole support system because it serves many users. Using that account's permissions for every tool call would let an ordinary user inherit its reach through the assistant. The gateway has to carry the user's narrower identity and tenant context into each operation, even though execution happens through a more capable account.
I would make the policy input explicit: (principal, tenant, action, resource, destination, origin). Here, origin records whether the proposal followed the user's request, a ticket, a retrieved page or another model inference. It helps explain an unexpected write. It doesn't authorize one. Missing information should stop the operation until the application can make a decision.
proposed = model.tool_call
assert schema.validate(proposed.arguments)
assert policy.allows(session.user, proposed.action,
proposed.resource, session.tenant)
if proposed.action.is_external_write:
assert approval.matches_exactly(proposed.arguments)
execute_with_scoped_credential(proposed)
These checks aren't interchangeable. A hostile request can have perfectly valid arguments, so passing schema validation says little about permission. The identity and resource information used by policy.allows must come from the application. Letting the model supply a tenant ID and then trusting it in the check would defeat the isolation you're trying to enforce.
Approval needs the same care. For an external email or irreversible write, show the recipient, exact content and affected resource. A batch needs an understandable scope or a hard cap. Once the user approves, the execution layer should use those parameters. If the model changes the recipient afterward, the old approval no longer covers the operation. A vague “finish the task” button leaves too much undecided.
some tools are too open-ended
Compare query_database(sql) with get_invoice(invoice_id). The first accepts a statement chosen by the model. The second can run a fixed query with tenant and row-level checks. Likewise, run_test(test_id) can map a reviewed identifier to a fixed command in an isolated runner, while run_shell(command) leaves a much larger set of actions available.
I'd look at those arguments before spending more time on the system prompt. A summarizer that only reads tickets is easier to contain than one that can also delete records, deploy services and send mail. Short-lived credentials scoped to the current task help. So does trimming tool responses: there's no reason to send the model a whole customer record when it needs two fields to answer the question.
Even a fetch tool can move data out. A document can ask the assistant to put private text into a URL query string and request that URL. Nothing has to be uploaded in a form; making the request is enough. The scheme, hostname, redirects and resolved IP all need checking before the request leaves. Cloud metadata endpoints and internal address ranges also need restrictions where appropriate. Validating the starting URL won't help if an unchecked redirect changes the destination.
And that fetch might be the third step. First the assistant reads a restricted file. Then it produces a summary. Then it puts the summary in a URL. Each call can look unremarkable when inspected alone. Keeping task-level records of restricted sources and proposed external destinations gives the gateway a chance to catch the sequence.
what i'd put in the test
I still want retrieved text labelled and separated from application instructions. That helps the model and makes the source easier to investigate. I wouldn't make an access decision depend on the model respecting the label every time, particularly after a model update or a change in conversation length.
The test needs to go past the assistant's answer and inspect execution. Put instructions in ticket bodies, document titles, OCR output and search snippets. Include quieter cases where an attacker slips a destination or resource ID into otherwise plausible task data. Record the proposed call and the gateway's decision, with sensitive values redacted, alongside the tool and policy versions.
A refusal sentence isn't enough evidence. The model could say the right thing while a separate path still runs the call. Conversely, a model that proposes the wrong action can still be contained by a correct permission check.
In production, denied calls grouped by action and source are useful starting points for investigation. So are cross-tenant attempts, unfamiliar external destinations and privileged operations without a direct user request behind them. For the ticket example, I want the test to allow the model to fail: it reads the instruction and proposes the email. Then I want a recorded authorization decision showing exactly why that email never ran.
technical references: owasp on prompt injection ↗, owasp on excessive agency ↗, and owasp on output handling ↗. the ticket example is hypothetical.