Direct Answer: Treat RAG Permissions as a Request-Time Control
RAG permission enforcement is the set of technical and operational controls that determines whether a particular user may retrieve, see, cite, or generate an answer from a particular document. It must operate at request time, not only when a document is first indexed. As of 28 September 2026, a defensible enterprise design should combine source-system authorization, identity-aware retrieval, tenant and group filters, answer-level controls, provenance, and continuous audit evidence. The model must never be treated as the primary permission boundary. Retrieval-augmented generation expands what a model can access, but it does not automatically preserve the access restrictions of the underlying systems.
Also worth reading: How Should Enterprises Design AI Agent Permissions Without Exposing Sensitive Data? · How Should Enterprises Secure Knowledge Exchange When SaaS and AI Retrieval Meet? · How Should Enterprises Enforce Row-Level Security in RAG Pipelines in 2026?
A useful rule is: authorize before retrieval, filter during retrieval, validate before generation, and record after the response. Source connectors should pass document identifiers and access metadata without exposing unauthorized content to the orchestration layer. The retrieval service should apply identity, tenant, role, group, classification, purpose, and jurisdiction constraints before ranking candidates. The generation layer should receive only the authorized evidence and attach citations that allow users and auditors to verify the result. This approach is stronger than asking a large language model to “remember not to reveal” restricted information, because probabilistic instructions are not a reliable authorization mechanism.
Why Existing RAG Security Controls Are Not Enough
Traditional enterprise permissions often stop at the application or database query. RAG changes that model by copying or indexing material into vectors, chunks, caches, logs, and prompts. Once data is transformed, an application may retrieve text without making another call to the originating system such as SharePoint, Google Drive, Confluence, Salesforce, or a document repository. If the index is shared across teams, the danger is not merely that an answer is inaccurate; it is that information governed by ACLs, legal holds, regional restrictions, or confidentiality labels crosses a trust boundary.
The most common failure pattern is disconnected enforcement. An administrator configures access in the source system, a connector imports the file, and an engineering team assumes those settings will follow automatically. In practice, indexing jobs may run with a service account that can see every document, while the chatbot runs under an account that lacks the source document's group membership. Vector similarity then ranks semantically relevant text without considering whether the requester is allowed to see it. Search relevance is therefore only one part of retrieval; authorization is a prerequisite to relevance.
Security guidance from Oracle, OWASP, Wiz, CSO Online, AWS, and materials concerning FedRAMP and the Indian law-enforcement or intelligence ecosystem all point toward the same broad requirement: AI systems need ordinary identity, data, and monitoring controls rather than a separate security philosophy. “Innocent Generative AI” describes RAG as a technique in which relevant documents are retrieved at inference time, but that description does not make the retrieval layer inherently permission-aware. Enterprises should assume that any system holding or processing regulated or confidential information requires controls comparable to those applied to the source application, plus controls for prompts, outputs, logs, and model interactions.
Reference Architecture: From User Request to Auditable Answer
A secure RAG request usually travels through seven stages. First, the application authenticates the user through an enterprise identity provider and establishes tenant, role, group, and session attributes. Second, an authorization service evaluates the request against source-system policies and rejects access when identity or context cannot be verified. Third, the retrieval planner translates those attributes into a constrained query, such as tenant ID, document ACLs, permitted groups, classification labels, and approved regions. Fourth, the vector or keyword search executes inside that boundary. Fifth, a second policy check examines each candidate, because vector databases may not consistently enforce row-level or document-level filters across every backend.
The generation stage should receive a minimal, filtered evidence set rather than a broad search result. A practical threshold is to return no more than the number of passages needed for the question—for example, 4 to 12 chunks in many enterprise assistants—while measuring whether additional passages improve answer quality. More context is not automatically safer or more accurate; excessive context increases exposure, latency, and token cost. The answer should include source names, document titles, dates, and links where policy allows, and it should state when the evidence is insufficient rather than filling gaps with model knowledge. After generation, the platform should log the user, model, query policy, retrieved document IDs, denied candidates, source version, response, and timestamp without storing sensitive text where that creates a new risk.
A useful architectural distinction is between filter-time and post-retrieval enforcement. Filter-time enforcement prevents unauthorized content from entering the model's context. Post-retrieval enforcement can suppress citations or outputs, but it is too late if the restricted text has already been sent to the model or recorded in telemetry. Defense in depth still matters: use both, but treat filtering before generation as the principal control. For high-risk data, the generation service itself should be unable to access unrestricted corpora, and the retrieval service should have no credentials broad enough to bypass user-level policy.
Connector and Index Design Choices That Affect Permissions
Connectors determine whether source permissions survive ingestion. A connector should support incremental synchronization, source version tracking, deletion propagation, and preservation of ACLs. It should also distinguish a document's owner from its current readers, because access can change after indexing. Group membership, sharing links, external guests, retention labels, legal holds, and regional residency can all matter. A connector that imports only text and a filename discards the metadata required to make a sound authorization decision later.
There are three broad index patterns. A shared index with strict metadata filters is economical but depends on correct filter logic and database configuration. A per-tenant or per-security-domain index reduces the blast radius and makes deletion easier, but it increases operational overhead. A per-user or per-group index gives strong isolation, yet it creates large provisioning and cache-management costs. Many enterprises use a hybrid design: separate indexes for highly regulated tenants or domains, with filtered shared indexes for lower-risk material. The correct choice depends on sensitivity, expected query volume, change frequency, and the cost of reproducing source permissions in the new system.
Permission metadata should be minimized and normalized. A group named “Finance–EMEA” in the source may become a different identifier in the identity provider, and stale mappings can accidentally broaden or narrow access. A nightly synchronization is often inadequate for documents whose permissions change during the day. Enterprises should define a maximum acceptable staleness, measure it, and use events or short intervals for high-risk repositories. As a conservative starting point, sensitive permission changes should propagate within 15 minutes, while emergency revocations may require immediate session termination and a short-lived retrieval cache. These are design targets, not universal compliance guarantees.
| Feature | Shared filtered index | Per-tenant or per-domain index | Per-user or per-group index |
|---|---|---|---|
| Isolation | Relies on consistent metadata filters | Strong tenant or domain separation | Strong user or group separation |
| Operational cost | Lowest | Moderate | Highest |
| Permission complexity | High query-time dependency | Lower cross-tenant risk | Low retrieval ambiguity |
| Best fit | General knowledge with stable ACLs | Regulated or larger customers | Highly sensitive or small deployments |
| Main weakness | Filter or configuration error can expose data | More indexes and synchronization work | Index explosion and cache overhead |
| Typical answer | Use only with tested row-level controls | Common enterprise compromise | Use selectively for exceptional domains |
An access-control list answers whether a subject may use a resource, but enterprise RAG often needs more than a single yes-or-no rule. A requester may be allowed to read a document but not export it, cite it to another team, use it for a model-training purpose, or access it from another country. A practical policy can evaluate action, subject, resource, tenant, purpose, environment, and time. Classification labels may require redaction before content reaches an external model, and purpose limitations may prevent a general assistant from reusing a restricted record for an unrelated workflow.
Provenance is both a user feature and a security control. Every factual statement should be traceable to one or more authorized source passages, with a link that causes a fresh permission check when opened. This prevents a copied citation from becoming a side channel. The platform should not display document titles, snippets, counts, or access-denied distinctions that reveal restricted information. A user should generally see “no authorized evidence is available,” not a detailed explanation that a document exists in another department. Logging can retain more diagnostic detail for approved security personnel, subject to retention and access policies.
The model must also be instructed to work only from the supplied authorized context, but that instruction is a quality safeguard rather than the security boundary. Tests should include direct requests, indirect requests, prompt injection embedded in documents, attempts to reveal system prompts, cross-tenant queries, role changes, revoked links, and requests for hidden citations. A model may comply with malicious text in a retrieved document, so untrusted content should be treated as data, not as an operator. A useful evaluation target for a mature deployment is zero unauthorized retrieval in a defined adversarial test set, with 100% of generated factual claims linked to an authorized source or explicitly marked as model-generated.
Practical Implementation Plan and Measurable Thresholds
Implementation should begin with a data and permission inventory, not with a vector database. Identify the top 10 to 20 repositories, classify their sensitivity, document who can search and export, and record whether ACLs are user-based, group-based, role-based, or combination-based. Then choose a small pilot containing at least one ordinary repository and one restricted repository. Establish success criteria before connecting production data: no cross-tenant retrieval, no response containing an unauthorized source, acceptable answer quality, predictable latency, and complete audit records.
The next step is to preserve identity and policy context through the request. Use short-lived tokens, avoid shared administrator credentials for user queries, and separate ingestion identity from search identity. Implement source deletion and permission-change propagation, then test them rather than assuming connectors support them. Add a policy decision point that can deny by default when a group mapping is missing. After retrieval, compare the candidate set with a separate authorization result and fail closed if the two disagree. Finally, test the complete path with a user who changes role, loses group membership, or attempts a cross-tenant query.
Measure both security and usefulness. Useful metrics include unauthorized-retrieval rate, unauthorized-answer rate, stale-permission rate, time to revoke access, percentage of answers with verifiable citations, citation correctness, and the proportion of requests denied by policy. A reasonable initial target for a new production system is 100% negative testing for known cross-tenant and role-escalation cases, at least 95% retrieval-policy evaluation coverage during pilot, and revocation propagation within 15 minutes for high-risk content. These thresholds must be adjusted for law, contract, and risk appetite; they are engineering starting points rather than legal safe harbors.
Alternatives, Trade-offs, and Cost
Enterprises can build a policy-aware RAG service internally, buy a managed RAG platform, or use a general assistant connected to repositories through an existing security platform. Internal development offers maximum control over indexing, models, and data residency, but it requires scarce identity, security, connector, and evaluation expertise. Managed platforms reduce time to deployment and often provide connectors, observability, and role-based administration, yet customers must still verify tenant separation, data retention, training use, regional processing, and the vendor's subcontractor terms. General enterprise search tools may already have stronger ACL ecosystems than a new RAG product, but their citation quality and conversational behavior may be less suited to domain questions.
Open-source RAG systems can be attractive because they expose retrieval and connector components for inspection. They are not automatically safer, however, and a pluggable connector model increases the number of places where permissions can be lost. Licensing, patching, support, and operational maturity also matter. A small proof of concept may cost little in platform fees but still require weeks of engineering and security review. Production budgets commonly range from several thousand dollars per month for a focused internal deployment to tens or hundreds of thousands of dollars annually for a managed, high-scale enterprise platform, depending on document volume, embedding and inference usage, connectors, retention, support, and compliance work. The largest cost is often authorization integration and testing rather than the model itself.
Pricing should be compared on control features, not token price alone. Ask whether pricing includes per-user seats, indexed documents, retrieval requests, storage, private networking, audit exports, SSO, SCIM, encryption-key options, and regional deployment. Verify whether the vendor charges separately for connectors, fine-tuning, evaluations, or long-term logs. A cheaper service that cannot enforce source ACLs may be more expensive because it requires a separate security layer or creates unacceptable risk. Conversely, a premium product is not justified if its policies remain generic and administrators cannot inspect denied retrievals.
Common Mistakes and When to Act
The most damaging mistake is treating ingestion-time access as permanent. Permissions change when employees leave teams, documents move between repositories, contractors end engagements, and regulatory holds expire. The second mistake is connecting repositories with a service account that bypasses user permissions, then relying on a prompt to compensate. The third is returning a broad candidate set and asking the model to remove forbidden passages. The fourth is measuring only answer quality; a fluent response can still contain restricted information. The fifth is logging full prompts and retrieved content without applying the same protection used for source systems.
Other weak patterns include using tenant names without verified tenant identity, trusting group names supplied by a client, ignoring document-level sharing links, and failing to distinguish “not found” from “not allowed.” Security teams should act immediately when a deployment handles regulated records, cross-tenant data, external users, or agent actions that can write back to systems. A pilot can tolerate a narrow test corpus and manual review, but production should not proceed until revocation, deletion, audit, and incident-response procedures have named owners.
There is no need to wait for a perfect model to begin permission work. In fact, the first useful step is to build a representative permission test suite and run it against the retrieval pipeline. If the organization cannot answer basic questions—who can access this repository, how quickly is revocation effective, what happens when a user changes group, and who can read the logs—it is not ready to add more data sources. When external processing is considered, obtain documented answers about model training, subprocessors, retention, geographic processing, and breach notification rather than relying on marketing language.
Minimum Governance Standard for Production
A production standard should require an owner for source permissions, an owner for identity mapping, an owner for retrieval policy, and an owner for model and output controls. Document the default-deny behavior, exception process, data classification rules, retention periods, and review cadence. Review high-risk policies at least quarterly and after major connector, identity, model, or repository changes. Keep an auditable record of policy versions because an answer generated under an old rule may be difficult to interpret after access changes.
The standard should also state what the system will not do. For example, it may not use restricted sources for general answers, expose denied document names, or allow a model to call a write-capable tool without a separate approval policy. Explicit limitations are more reliable than broad claims of enterprise readiness. Security, privacy, legal, and data owners should approve them together, because a technically correct policy can still be unacceptable under contractual or regulatory obligations.
Ultimately, RAG permission enforcement is a continuous data-control problem rather than a one-time feature. The best design makes the secure path the ordinary path, keeps authorization close to the source and retrieval layers, and preserves enough provenance to prove what happened. No model can repair an index that contains content the requester was never allowed to retrieve. The practical objective is not perfect natural-language reasoning; it is a system where unauthorized material is excluded before generation, permitted answers are traceable, and administrators can demonstrate both enforcement and accountability.