Your AI Search Ignores Your Permissions: Building Permission-Aware Retrieval in SaaS
Tenant isolation is not enough once you add RAG. This is how I make AI search and assistants respect per-document permissions, revocations and derived data in multi-tenant SaaS.
Pavel Duglas
AI Automation & MVP Architect
A client once showed me a bug report that looked harmless: “The AI assistant knows about the reorg.” A junior sales rep asked the in-app assistant about their team’s plans, and it answered with details from an HR document that only three directors could open. Every tenant check passed. The rep was in the right workspace, looking at their own company’s data. The problem was that the assistant never asked whether this specific person could see this specific document. It only asked whether they belonged to the company.
This is the most common security hole I find when I audit AI features in SaaS products. The core app has a careful permission model. Then someone adds retrieval, and the index quietly flattens all of it into “everything in the tenant.” Here is how I build retrieval that respects permissions from day one.
Tenant isolation is the floor, not the model
Most teams get tenant isolation right because it is loud when it fails. Customer A seeing Customer B’s data is a headline. So you add tenant_id to every table, every query, every vector namespace, and you feel safe.
But inside a tenant, real products have layers: private documents, team folders, projects with invited guests, admin-only settings, records owned by one user. Your regular API enforces those layers on every read. Your AI feature usually does not, for one simple reason: the vector index is a copy of your data, and copies lose their access rules unless you carry them over deliberately.
When you chunk a document and embed it, you get text plus a vector. The ACL lives somewhere else, in your relational database, and nothing connects them. The assistant then retrieves by similarity, which is completely blind to who is asking.
The rule: filter before similarity, not after
The first instinct is to retrieve top-k chunks and then drop the ones the user cannot see. I call this post-filtering, and I avoid it for three reasons.
- Empty results. If you fetch 10 chunks and the user can see 1, the answer quality collapses. For users with narrow access it collapses every time.
- Things see data before the filter does. Rerankers, logging middleware and debug traces often touch the full top-k. That means restricted text is sitting in your logs and possibly in a third-party reranker call.
- It is easy to forget. One new code path that skips the filter and you have a leak. Security that depends on every developer remembering a step is not security.
The correct approach is pre-filtering: the permission constraint is part of the search query itself, so restricted chunks are never candidates.
Store access metadata on every chunk
Every chunk carries the principals allowed to read it. I store a flat list of principal IDs, not the raw ACL structure, because vector databases filter well on arrays and badly on nested logic.
chunk = {
"id": "doc_812#c4",
"tenant_id": "t_19",
"source_id": "doc_812",
"acl_version": 7,
"allowed_principals": ["user:44", "group:directors", "role:admin"],
"text": "...",
"embedding": [...],
}
Principals are users, groups and roles expanded into a single namespace. A document shared with a team gets group:team_sales, not a list of every member. That keeps the index stable when people join or leave a team.
Resolve the caller’s principals at query time
At query time you compute the full set of principals for the current user and pass it as a filter.
def search(user, query, k=8):
principals = [f"user:{user.id}"]
principals += [f"group:{g}" for g in user.group_ids()]
principals += [f"role:{r}" for r in user.roles()]
return vector_db.query(
vector=embed(query),
top_k=k,
filter={
"tenant_id": user.tenant_id,
"allowed_principals": {"$in": principals},
},
)
Notice the tenant filter is still there. Principal matching is the second wall, not a replacement for the first. And the function takes a user, not a tenant_id. I make it impossible to call search without an identity. If someone needs a system-level search for a background job, they have to write a separate function with a name that makes reviewers nervous.
Permissions change, your index doesn’t know
Pre-filtering solves the static case. The hard part is time. Someone removes a guest from a project at 10:00. Your re-indexing job runs nightly. For 14 hours that guest can still ask the assistant about the project.
I handle this with two mechanisms together.
Push ACL changes as events
Every permission change in the core app emits an event: document shared, unshared, moved to a folder, group membership changed. A worker consumes those events and updates allowed_principals on the affected chunks. This is a metadata update, not a re-embedding, so it is cheap and fast. Most vector stores let you patch metadata by filter on source_id.
Group membership changes are the nice part of using group principals: removing a user from group:team_sales requires no index update at all, because the user’s principal set shrinks at query time.
Re-check against the source of truth before answering
Events get delayed, workers crash, queues back up. So after retrieval, before any chunk goes into the prompt, I do one cheap check against the primary database:
hits = search(user, query)
source_ids = {h.source_id for h in hits}
readable = permissions.filter_readable(user, source_ids) # one SQL query
context = [h for h in hits if h.source_id in readable]
Yes, this looks like post-filtering. The difference is that it is a safety net on top of pre-filtering, not the primary mechanism. In practice it drops almost nothing. When it does drop something, I log it as a sync lag metric. If that number grows, the event pipeline is broken and I want to know before a customer does.
Derived artifacts are where the real leaks hide
Chunks are the obvious part. The leaks I actually find in audits are in everything the AI feature produces from them.
- Generated summaries. A “weekly project digest” built from ten documents is a new document. Who can read it?
- Cached answers. Caching an answer to “what’s our Q3 plan?” per tenant means the next user gets an answer built from the first user’s access.
- Conversation history. If you embed past chats for memory, a director’s chat now contains restricted content and lives in the index.
- Extracted entities and tags. Structured data pulled from a restricted document still reveals what the document says.
My rule is simple: a derived artifact inherits the intersection of its sources’ permissions. If a summary uses a doc readable by the sales team and a doc readable by directors, the summary is readable only by people in both. That is often nobody useful, which is a signal the artifact should be generated per user or not stored at all.
For caches, the cache key includes a hash of the source IDs used plus the ACL versions, and the cached answer is only served after the same readability check. For conversation memory, chats are private to the user by default and never go into a shared index.
Tools run as the user, not as the service
Once the assistant can call tools, like “look up invoice” or “list open tickets,” the same problem appears in a new shape. The easy implementation calls your internal API with a service token that can read everything in the tenant. Now the model decides what the user sees, and the model is not an authorization system.
Every tool call executes with the requesting user’s permissions, through the same authorization layer your normal API uses. If the user cannot open invoice 4411 in the UI, the tool gets a 403, and the assistant says it cannot access it. No special AI path, no bypass. If a tool genuinely needs elevated access, it returns only the minimal derived fact, like a count, and that decision is documented and reviewed.
The leak test suite I add to every project
Permission bugs do not show up in happy-path demos, so I write tests that go looking for them. These run in CI against a seeded test tenant.
- Canary documents. Seed a restricted doc containing a unique string like
CANARY-7Q2X. Log in as a user without access, ask a dozen questions designed to pull it out, and assert the string never appears in any answer, citation, log line or trace. - Revocation test. Share a doc with a user, confirm the assistant can use it, revoke access, wait for the event worker, and confirm it can no longer use it. Then kill the worker and confirm the source-of-truth check still blocks it.
- Same tenant, different users. Two users in one tenant with non-overlapping access. Every query from one must return zero chunks owned only by the other.
- Derived artifact test. Generate a summary from mixed-permission sources and assert a user with partial access cannot retrieve it.
- Tool bypass test. Ask the assistant to fetch a record by ID that the user cannot open, and assert the tool returns a denial.
The canary test alone has caught real leaks for me twice, both times in logging, not in answers.
My checklist before an AI search feature ships
- Every chunk has
tenant_id,source_id,allowed_principalsandacl_version. - The search function requires a user object and builds the principal filter itself.
- Retrieval is pre-filtered; the post-check exists only as a safety net with a metric.
- ACL changes emit events that patch index metadata within seconds.
- Summaries, caches and memories inherit the intersection of source permissions.
- Tools execute through the normal authorization layer as the requesting user.
- Canary and revocation tests run in CI.
- Logs and traces never contain chunk text from before the permission filter.
None of this is exotic. It is the same discipline you already apply to your API, extended to the one part of the system that everybody forgets is a read path. If your assistant can answer a question, it has read the data. Treat it that way.
FAQ
Can't I just use a separate vector namespace per user instead of ACL metadata?
Only if documents are truly private and never shared. The moment one document is visible to a team, a per-user namespace means duplicating the chunk for every member and re-indexing whenever membership changes. Namespaces work well per tenant. Inside a tenant, principal metadata with a query-time filter scales much better and handles groups for free.
Does pre-filtering hurt search quality or speed?
Usually not in a way users notice. Most modern vector databases apply metadata filters during the approximate nearest neighbor search, so you still get a full top-k of readable results. It can slow down when a user can see a tiny fraction of a huge index; in that case I partition by tenant first and keep principal lists short by using groups instead of individual users.
What about documents synced from external tools like Google Drive or Notion?
Pull the source system's permissions during sync and map them into your principal namespace, then subscribe to their change webhooks or poll sharing changes on a short interval. If the source cannot give you reliable permissions, the safe default is to make synced content readable only by the user who connected the integration until an admin explicitly shares it wider.
Related articles
Done for you
I will build a platform with accounts, roles and payments
A user area, an admin area, payment and CRM integrations, and a structure that survives the second version.
from $3,000 · 3 to 5 weeks