ArchitectureAzure Copilot

Permission-aware RAG with Azure AI Search: carrying document access from source to answer

A grounded assistant is only as discreet as its retrieval step. If the search index returns a chunk the user could not open at the source, the model will quote it fluently. This page sets out where the permission check belongs in an Azure AI Search and Azure OpenAI design, the four enforcement approaches Microsoft documents, how to keep permissions current after they change, and how to prove trimming works before go-live.

Reviewed 8 min read

On this page
  1. Where retrieval-augmented answers leak, and where the check must sit
  2. Reference architecture: the path of one trimmed question
  3. Four enforcement approaches documented for Azure AI Search
  4. Capturing permissions at indexing time
  5. Filtering each query by the caller's identity and group claims
  6. Revocation latency and other permission-sync failures
  7. Persona and negative tests before go-live
  8. Logging who saw which source through the assistant
  9. Questions and answers
  10. Sources

Where retrieval-augmented answers leak, and where the check must sit

Retrieval-augmented generation adds a new reader to every document: the index. Indexers usually read sources with a service identity that can see far more than any single employee. If nothing records who may read each document, the search step returns the best-matching chunks whoever asked, and the model turns them into a confident answer with a citation.

The permission check therefore belongs at retrieval, before any content enters the prompt. Instructing the model not to reveal restricted material does not work: it cannot tell which passages are restricted, and a prompt is not an access control. Two secondary leak points need the same discipline: answer caches shared across users, and conversation logs that store retrieved text where a wider group can read it.

Reference architecture: the path of one trimmed question

Indexing is covered further down. This is the query path, where the caller's identity has to travel with every retrieval call.

Sign inUser tokenQuestion and tokenQuery with user contextPermitted chunks onlyPrompt with chunksGrounded answerAnswer and citations01User and clientapp02Microsoft EntraID03Orchestrator04Azure AI Search05Azure OpenAIdeployment
  1. User and client app

    Signs the user in and sends the question with the user's token, never a group list built in the browser.

  2. Microsoft Entra ID

    Issues the token that identifies the user and, with Microsoft Graph, their group memberships.

  3. Orchestrator

    Carries the user's identity into every search call and builds the prompt from returned chunks only.

  4. Azure AI Search

    Drops every chunk whose permission metadata does not match the caller.

  5. Azure OpenAI deployment

    Answers from the permitted chunks and never sees permission fields.

  1. User and client app to Microsoft Entra IDSign in
  2. Microsoft Entra ID to User and client appUser token
  3. User and client app to OrchestratorQuestion and token
  4. Orchestrator to Azure AI SearchQuery with user context
  5. Azure AI Search to OrchestratorPermitted chunks only
  6. Orchestrator to Azure OpenAI deploymentPrompt with chunks
  7. Azure OpenAI deployment to OrchestratorGrounded answer
  8. Orchestrator to User and client appAnswer and citations
Conceptual sequence for one question; the user context is either the user's token in a request header or a group filter built on the server.

Capturing permissions at indexing time

  1. Grant access through groups, then index group identifiers

    Wherever possible, grant access at the source through Entra ID groups and store group identifiers in the index. Individual user entries multiply with staff turnover and turn every joiner or leaver into a re-indexing event.

    Output
    A group-based access design for each source
  2. Add permission fields to the index schema

    Create string collection fields for user and group identifiers, make them filterable and not retrievable, and enable the index's permission filter option when you use a native approach2. Keeping them out of results means identifiers never reach the prompt.

    Output
    Index schema with permission fields
  3. Project permissions onto every chunk

    When a skillset splits documents for vector search, map the permission fields through index projections so each chunk carries them. A chunk without permission fields cannot be matched to the right caller2.

    Output
    Chunked index where every row is trimmable
  4. Record where each permission came from

    Store the source path and the time permissions were last read next to each document. During an incident, this tells you which documents were exposed and since when.

    Output
    Traceable permission metadata
  5. Keep indexing inside your network boundary

    If sources sit behind private endpoints, set the indexer to the private execution environment; otherwise it can fail silently and leave an empty or partial index4.

    Output
    Indexer configuration reviewed with network owners

Filtering each query by the caller's identity and group claims

With a native approach, the orchestrator passes the signed-in user's token in the x-ms-query-source-authorization request header, and Azure AI Search compares its user, group and scope claims with the stored permission metadata before returning anything1. The application still needs a reader role on the index; the user token narrows what each query may return.

With security filters, your code builds the filter. The orchestrator resolves the user's memberships on the server, typically through Microsoft Graph, and applies a search.in expression over the group field, which Microsoft recommends over long chains of equality checks3. Never accept a group list from the browser, and resolve transitive memberships: for people in a large number of groups, Entra ID tokens carry an overage indicator instead of the full list.

Either way, the same identity must reach every retrieval, including follow-up searches in multi-turn conversations and agent tools that query the index.

Revocation latency and other permission-sync failures

Query-time enforcement can only reflect permissions already synchronized into the index1. The gap between a change at the source and the next sync is how long a removed user can still retrieve the content.

Inherited permission changes are missed

Early signalA library or folder is locked down, yet its files still appear in answers.

MitigationFor SharePoint, item-level changes are picked up on indexer runs in recent preview API versions, but parent-scope changes need a permissions resync or a document reset, so build that call into your change process2.

The sync schedule ignores sensitivity

Early signalHR or legal content is re-indexed on the same relaxed schedule as public policies.

MitigationSchedule indexers per source by sensitivity and alert on failed runs, because a failed run quietly extends the exposure window.

No emergency removal path

Early signalSomeone reports a leaked document and the only option is to wait for the next run.

MitigationKeep a runbook to delete or reset specific document keys at once, and rehearse it before go-live.

Access survives in caches and transcripts

Early signalCached answers or chat logs are readable by people who never had access to the source.

MitigationCache per user or not at all, and protect conversation logs as strictly as the most sensitive source they could contain.

Persona and negative tests before go-live

Trimming is proven by trying to break it. Create test accounts that mirror real roles and run the same question set under each.

0 of 6 checked

Logging who saw which source through the assistant

For audit, log for each answer the user's object ID, the query, the document keys returned and cited, and the index version, in a workspace at least as restricted as the content. That is how you answer the question a security team will eventually ask: who could have seen this document through the assistant, and when.

Questions and answers

Should SharePoint content be indexed or queried remotely?

Index it when you need custom chunking, enrichment or one index over several sources, and accept that permissions must be synchronized. Microsoft also documents a remote SharePoint knowledge source that queries SharePoint at question time through the Copilot retrieval API, so SharePoint's own permissions and labels apply without a copy2. Remote retrieval removes sync latency but gives you less control over ranking and chunking.

What if some users belong to hundreds of groups?

Resolve memberships on the server and filter with search.in over a group field, which keeps response times practical where long equality chains would not3. Reduce the problem at the source too, by granting access through a few purpose-made groups instead of every team group a person joins, and test latency with your largest real membership lists.

Can sensitivity labels replace access-list trimming?

They solve a related but different problem. Labels describe how sensitive content is and which protection policies apply, while access lists say who may open a specific item. Azure AI Search can ingest Microsoft Purview labels and evaluate them at query time in preview, and many designs use both: access lists to decide visibility, labels for extra restrictions and for showing how sensitive an answer's sources are1.

Can the assistant query the index with its own identity instead of the user's?

Only when every document in the index may be seen by everyone in the audience. A service identity knows nothing about the person asking, so it returns the same results to all users. As soon as any source contains restricted documents, each query has to carry the user's identity, either as a token in the request or as group filters built on the server.

Sources

  1. Document-level access control in Azure AI Search — Microsoft Learn · checked 10 October 2026
  2. Use a SharePoint indexer to ingest permission metadata — Microsoft Learn · checked 10 October 2026
  3. Security filter pattern in Azure AI Search — Microsoft Learn · checked 10 October 2026
  4. How to configure network isolation for Microsoft Foundry — Microsoft Learn · checked 10 October 2026

More in Azure Copilot

Back to Azure Copilot

Next step

Have your document permissions reviewed before you index them

Send us the sources you plan to ground the assistant on and how access is granted today. We will reply with the enforcement approach that fits each source and the persona tests to run first.

Review my retrieval design