← writing

rag needs an authorization boundary

Remove a user's access to a project, then ask the assistant a question they asked yesterday. If it returns the old answer from cache, the source system's permission change hasn't done the job you expected.

Document folders in an office cabinet with a frosted glass access door.
Generated illustration of restricted document storage.

This is a useful place to start reviewing a RAG application. The happy path is easy to demonstrate: upload a document, ask a question, get an answer with a citation. Revocation makes you follow all the copies created along the way. The document may have become chunks in a vector index, text in a reranking request, context in a model call and an answer in a cache. Changing one ACL doesn't automatically reach all of them.

The original file service could be configured correctly throughout. Once the pipeline copies text elsewhere, the new system has its own storage, API and operators. Its access decisions need to stay connected to the source. Otherwise “this document is private” only describes the first copy.

start with the caller

Before searching, resolve the user's authenticated identity and tenant. Those should determine which documents can become candidates. A tenant ID in the request body is something to validate, not an identity to trust.

Searching everything and filtering the top ten afterward has two problems. You can get ten excellent matches from a tenant the caller can't access, leaving an empty answer even if useful authorized documents exist further down. You may also expose the rejected text to a reranker, cache or debug trace before the filter runs. Put the restriction as early in retrieval as the engine supports, and verify what “filtering” actually means in that engine.

Some vector databases rank candidates before applying metadata filters. That affects recall and may require a larger candidate pool. It also affects which components encounter the data. An API accepting a filter parameter doesn't, by itself, tell you where the check happens.

scope = authz.scope_for(user, tenant)
candidates = index.search(embedding(query), filter=scope, limit=40)
authorized = [c for c in candidates
              if authz.can_read(user, c.document_id, c.version)]
context = rerank(authorized, query)[:8]

The second check is there because indexed permissions can be stale. Before sending text to the model, consult the authoritative access-control service again, including the document version if the source supports it. The ordering matters: external reranking services and logs that capture chunk text belong after that check as well.

If the permission service is unavailable, the application may need to fail the request or answer with less context. There isn't a useful interpretation of “permission unknown” that grants access to a restricted document in a shared index.

Back to the revoked user. A fresh retrieval check might correctly refuse the document while the answer cache still returns yesterday's response. A key based only on prompt text won't notice the difference. One possible key is (tenant, principal, scope_version, query_hash, corpus_version). Access changes advance the scope version; content changes advance the corpus version. Old entries stop matching while purge jobs remove them.

That still leaves work already in flight. You need a maximum revocation delay that covers ACL propagation, index updates, cached responses and running requests. A deletion job finishing in one store isn't evidence that a response already being generated has forgotten the document. Define the behaviour you're promising and test that full interval.

keeping track of the copies

All of this gets harder if ingestion throws away the source information. Each chunk needs a document ID, tenant, classification, owning system, source version, ingestion time and authorization reference. Stable chunk IDs within a document version help you replace or purge the right data. An embedding and a string of text aren't enough to work backward from an answer to an access decision.

Version changes need a rule too. Either the old version or the new one should be visible according to a defined consistency model. Quietly serving chunks from both can produce stale answers and make a revocation investigation much harder to follow. The same applies to derived summaries: deleting the original chunks while leaving their summaries searchable doesn't finish the deletion.

I'd also review the crawler's credentials. A permanent administrator token makes ingestion convenient because nothing is out of reach. That's precisely the access a compromised worker would inherit. Scope workers by source, and record which credential and source version each run used. Where ACLs are available, capture them and arrange for changes to reach the index. If a source can't reliably report its access rules, restricted material should wait until you've defined how to handle it.

The overlooked copies often live in diagnostics. A trace containing retrieved text has become another place that stores the document, with whatever permissions and retention the logging system uses. Document IDs, versions, scores and policy decisions are safer defaults for routine logging. If you need content samples to debug a problem, give those samples their own access review and expiry. Embeddings deserve access and retention rules too; they can reveal information about their sources.

try the routes nobody uses in the demo

A tenant filter on search doesn't protect a fetch-by-ID endpoint that skips it. Exports, reranking and debug routes need the same scrutiny. For tenants where a leak would be particularly costly, separate collections, keys or infrastructure may be appropriate. The isolation choice doesn't remove the need to derive the tenant from the session and test forged IDs on every route.

There is also a separate problem once authorized text reaches the model: the document may contain hostile instructions. Permission to read it doesn't make those instructions authoritative. Keep source text delimited and retain the citation. A proposed tool call still requires its own authorization check. Writing back to the knowledge base needs a separately reviewed capability, because an assistant could otherwise turn a bad response into context for its next request.

For testing, give two tenants documents with nearly identical language. Ask the same question as each user and examine candidates before filtering, model context, caches, logs and citations. Revoke access and repeat immediately. Delete a document and follow both its chunks and summaries through the promised deletion window. Missing ACLs and malformed metadata need cases of their own; ingestion won't always receive a well-formed record.

When an answer does expose something it shouldn't, a citation tells you which document to look at. Keep the version and the permission decision from that request as well. By the time you investigate, memberships and ACLs may have changed again. You'll need to know why access was allowed then, not just what the settings say now.

technical references: owasp on vector and embedding weaknesses ↗, owasp on prompt injection ↗, and nist ai risk management framework ↗.