top of page

Amazon Quick RAG Access Control Moves Permission Checks to Query Time

13 hours ago
13 min read

Amazon Quick changed RAG access control by adding a second permission check before enterprise content reaches the model. The October 7 announcement targets a persistent security gap: indexed permissions can become outdated between synchronization cycles.

The new Amazon Quick RAG access control design combines fast filtering inside a search index with real-time verification against the original data source. AWS describes support for enterprise knowledge drawn from systems such as Microsoft SharePoint, Google Drive, and Atlassian Confluence.

That distinction matters because retrieval-augmented generation, or RAG, generates answers using passages retrieved from connected information sources. If retrieval admits an unauthorized passage, the model can expose its contents through a summary, comparison, or indirect answer.

The primary contest is therefore not AWS against another vendor. It is source-authoritative verification against the widely used practice of copying permissions into an AI index and trusting that copy.

AWS says its two-stage approach narrows the interval between a permission change and its enforcement inside an AI response. However, the announcement does not eliminate identity, configuration, latency, auditing, or connector risks. It changes where enterprises should draw the retrieval security boundary.

Amazon Quick RAG Access Control Adds a Second Gate

The important change is not another enterprise connector. It is a permission decision made after retrieval candidates are found and before their text reaches the model.

Many RAG systems ingest both content and access-control lists, or ACLs, during a scheduled crawl. An ACL records which users or groups can access a particular resource. The system stores those permissions as metadata beside indexed document passages.

When someone submits a question, the retrieval layer searches the index and filters results using that stored metadata. This arrangement is efficient because both relevance ranking and permission filtering happen near the vector index.

The weakness is time. An indexed ACL represents the permissions observed during the last successful synchronization. It does not necessarily describe who can open the document at the moment of a query.

AWS introduced its real-time ACL design as an additional control above that indexed filtering. The first stage still uses stored ACL data to reduce the candidate set. The second stage asks the connected source whether the user currently has access.

Only passages that survive both stages become context for the large language model. Context is the retrieved information supplied to the model while it prepares an answer.

This ordering is crucial. The system does not rely on the model to recognize confidential material or remove it after generation. It attempts to exclude unauthorized passages before generation begins.

AWS illustrates the process with Google Drive. Quick first performs semantic search, which retrieves passages by meaning rather than exact keyword matches. It applies the ACLs stored in the index to produce a smaller group of candidate documents.

Quick then calls Google Drive APIs to validate those candidates. AWS says the service uses administrator-provided service account credentials to create user-specific access tokens through impersonation.

Google Drive remains the authoritative source for each candidate’s permissions. A document that fails the live check is removed, even if the indexed ACL still indicates access.

That sequence preserves much of the speed advantage of an index. Checking every document in a large repository through a remote API would produce substantial latency and request volume. Checking only a narrowed candidate set creates a more practical security and performance balance.

SharePoint follows the same broad pattern, although its identity flow differs. AWS documentation describes pre-retrieval filtering followed by delegated verification of the user’s current SharePoint rights.

For an ACL-enabled SharePoint knowledge base, Quick prompts the user to sign in when protected content becomes relevant. The service then uses a delegated token to validate access to each candidate document.

According to the SharePoint ACL flow, that sign-in is generally a one-time step. The associated refresh token lasts approximately 90 days.

The documentation also specifies delegated access for reading site items, files, the user’s basic profile, and maintaining authorized access. Those scopes deserve review because real-time verification depends on functioning identity delegation.

This is more than a connector refresh. AWS is assigning different responsibilities to two layers. The index handles fast candidate selection, while the source system provides the final permission answer.

That architecture turns stale ACL metadata from the sole decision-maker into an initial filter. It can still affect which candidates are considered, but it no longer has the final word for supported real-time flows.

Cached Permissions Became the Weak Link in Enterprise RAG

Enterprise RAG inherits every complicated permission rule inside its source systems, then adds synchronization and identity mapping as new failure points.

A typical company repository rarely has one simple access policy. SharePoint can combine sites, groups, inheritance, exceptions, and explicit grants. Google Drive can include personal files, shared drives, direct sharing, group membership, and organization-wide settings.

Confluence adds spaces, pages, group membership, and inherited restrictions. One company can use all three systems while also connecting OneDrive, Amazon S3, and internal web applications.

Amazon Quick currently documents integrations for S3, Confluence, Google Drive, OneDrive, SharePoint, and authenticated web content. Its data access integrations use several authentication patterns, including OAuth and service accounts.

Replicating access rules into one normalized index requires a connector to interpret each source correctly. It must preserve user identities, nested groups, inheritance, deny rules, and changes made after the previous crawl.

A mapping defect can grant access too broadly. A delayed synchronization can preserve access after an employee changes roles. A failed crawl can leave content current while its permission representation remains old.

The problem becomes more serious when employees treat an AI assistant as a shortcut across repositories. A conventional interface reveals files individually, often with familiar folder or site boundaries. A RAG assistant combines evidence across sources into one answer.

That synthesis improves usefulness, but it also changes the exposure pattern. A user does not need to know that a restricted document exists. A broad question can retrieve a passage and convert it into a concise statement.

An answer might combine a public project update with a confidential budget, personnel decision, or acquisition plan. Even a partial disclosure can reveal information that the original interface would have hidden.

Post-generation filtering offers a weak remedy because the model has already received the passage. Guardrails can identify categories such as personal data or unsafe content, but they do not automatically understand every company’s document permissions.

The correct place for document authorization is before generation. That principle also matters for citations, follow-up questions, summaries, exports, and actions triggered by agents.

AWS’s announcement focuses on the lag between synchronized permissions and live source state. Consider an employee removed from a confidential strategy group shortly after a scheduled crawl.

A replicate-and-filter system might continue recognizing the employee’s old membership until the next successful synchronization. Event-driven updates can reduce that interval, but they do not cover every permission change across every platform.

AWS specifically notes that some changes, such as Confluence group membership updates, do not always produce a usable event. A connector cannot immediately react to an event it never receives.

Source systems also evolve. A new sharing method or policy type can outpace the connector’s translation logic. The index might then misrepresent permissions until the connector receives an update.

Real-time verification changes the dependency. The AI layer still needs working integration logic, but the final decision comes from the system already responsible for the resource.

This is why the announcement pressures teams building custom RAG stacks. They must now justify why a replicated permission snapshot is sufficient when a major cloud platform offers source validation during retrieval.

It also pressures enterprise buyers to ask more precise questions. “Does the product support ACLs?” is no longer enough because indexed ACL filtering and live authorization provide different guarantees.

A useful evaluation should identify the source of truth, the identity passed during retrieval, the timing of checks, the treatment of failures, and the evidence available to auditors.

The broader lesson also applies to personal and team knowledge systems. A well-designed AI knowledge base needs boundaries that match the information it connects, not merely the interface presenting it.

The Two-Stage Mechanism Trades Simplicity for Fresher Decisions

AWS improves permission freshness by accepting a more complex retrieval path with additional identity, API, and operational dependencies.

The first stage exists for scale. Quick searches the vector index and applies synchronized ACL metadata before contacting a source platform.

This step limits live calls to documents that are both semantically relevant and apparently accessible. Without that reduction, every question might trigger permission requests across a much larger corpus.

The second stage exists for correctness. Quick checks the candidate documents through the relevant source API and drops any candidate that the user cannot currently access.

This hybrid model resembles a coarse filter followed by an authoritative decision. The coarse filter controls cost and latency. The final decision addresses revoked access and imperfect permission replication.

The model only receives passages approved by the live check. That design reduces the chance that unauthorized material enters prompts, generated answers, citations, or downstream model processing.

The mechanism also clarifies what “real time” means in this context. It does not mean Quick constantly synchronizes every permission. It means the system validates selected documents while processing a query.

This approach can reflect revocation sooner than a scheduled crawl. AWS says changes appear in AI responses within moments instead of waiting hours or days for synchronization.

That timing is a company claim, not an independently measured service-level guarantee. Actual behavior will depend on the connected platform, token state, API availability, connector configuration, and the specific knowledge-base mode.

The architecture creates several operational questions. A source API can throttle requests, return transient errors, or experience an outage. A delegated token can expire or lose required consent.

Enterprises need to know how Quick handles each condition. A secure default should fail closed, meaning uncertain permissions exclude the document instead of allowing it.

Failing closed protects confidentiality, but it can reduce answer quality or produce no result during an identity failure. Users may interpret that absence as missing knowledge rather than a security decision.

Observability therefore becomes essential. Administrators need records showing which source was checked, which identity was used, whether verification succeeded, and why a document was excluded.

Latency deserves equal attention. One remote permission check may be inexpensive, but an answer can depend on several documents from multiple repositories.

Parallel verification can reduce waiting time, although it can increase burst traffic toward connected APIs. Sequential verification controls concurrency but can make an assistant feel slow.

Caching a successful live decision can improve performance, but caching reintroduces a freshness interval. AWS’s public article does not provide enough detail to assess every cache, timeout, retry, or rate-limit policy.

Identity mapping remains another difficult boundary. The querying identity in Amazon Quick must correspond to the identity recognized by Google Workspace, Microsoft Entra, or another source.

Service-account impersonation can preserve user-specific decisions when configured correctly. It also introduces credentials, delegation policies, audit trails, and administrative privileges that security teams must examine.

The source still matters more than the vector store, but the integration becomes security-sensitive infrastructure. A mistake in impersonation or token handling can undermine the value of live checking.

AWS documentation for custom Bedrock data sources illustrates an important limitation. Its custom ACL documentation says those sources use customer-provided ACL metadata rather than real-time source verification.

The same documentation makes an even sharper distinction. ACL-aware filtering is not an authentication boundary because Bedrock cannot verify the identity context supplied by the calling application.

Applications must authenticate users upstream and pass verified identity information. Enterprises should not treat metadata filtering alone as complete authorization.

For custom sources, the application supplies allow and deny entries with each document. Bedrock applies them before retrieval, and deny entries override allow entries.

However, those permissions are only as current and accurate as the customer’s ingestion process. There is no authoritative source API for Bedrock to consult when the custom connector defines the ACL itself.

That caveat prevents an overly broad reading of the AWS announcement. Real-time verification is a connector-specific capability, not a universal property of every Bedrock knowledge-base configuration.

The architecture is still meaningful. It establishes a better target for supported repositories while documenting that custom implementations retain more responsibility.

Real-Time Checks Do Not Turn Bedrock Into the Security Boundary

The new layer reduces one exposure window, but enterprises still own authentication, configuration, source governance, testing, and incident detection.

AWS presents source-authoritative verification as protection against stale or incorrectly mapped ACL data. That claim is reasonable for permission changes successfully evaluated through supported source APIs.

It does not mean every access-control problem disappears. The system can only enforce the permissions that the source returns for the identity and resource it checks.

If the source itself grants access too broadly, Quick will respect that broad grant. If an administrator places confidential information in a widely shared folder, real-time verification will not infer a stricter business policy.

The same issue applies to inherited permissions. Source authority improves technical consistency, but it cannot determine whether an inherited grant was appropriate.

Organizations still need access reviews, least-privilege policies, offboarding procedures, and ownership rules for shared repositories. RAG can expose weak source governance faster because it makes scattered content easier to find.

Authentication is another independent control. Bedrock’s documentation explicitly warns that ACL-aware filtering does not authenticate end users. The calling application must establish identity before supplying user context.

That warning matters because a trustworthy permission check against an untrustworthy identity proves little. A malicious or defective application could pass another user’s identifier unless upstream controls prevent it.

Enterprises should test the complete path, starting with sign-in and ending with the generated response. Tests should cover revoked access, group changes, inherited permissions, explicit denies, token expiration, API failure, and knowledge-base recreation.

SharePoint introduces a configuration constraint worth noting. AWS says ACL management must be enabled while creating the knowledge base and cannot be changed afterward.

A team that omitted the setting must create another knowledge base. That requirement can affect rollout plans, reindexing, acceptance tests, and change management.

Required Microsoft permissions also need scrutiny. The administrator-managed setup can require directory and group-reading rights, plus access to selected or broader SharePoint sites.

The delegated verification application requests separate permissions for reading files and site content. Security teams should distinguish those two applications and understand which credentials support ingestion versus query-time checks.

Custom connectors demand another test program. Incorrect ACL field casing, a missing list, or mismatched user email can silently remove documents from retrieval.

AWS says such retrieval failures close access rather than reporting an authorization error. That behavior protects data, but it complicates diagnosis because users may simply receive fewer results.

Content security extends beyond permissions. Authorized documents can contain malicious instructions intended to manipulate a model, a risk commonly called indirect prompt injection.

A correct ACL does not make a document safe. It only establishes that the user may access it. Enterprises still need content controls, model safeguards, tool restrictions, and monitoring.

AWS mentions Bedrock Guardrails, grounding checks, and configurable safety policies alongside the ACL architecture. Those controls address different risks and should not be treated as replacements for authorization.

The company’s own Generative AI Lens has warned that reconstructing complicated ACLs through metadata creates engineering effort and possible permission gaps. It recommends careful selection of managed or custom approaches.

That guidance supports the motivation for query-time checks. It also reinforces the need to examine implementation details rather than accept a broad label such as “permission-aware RAG.”

Independent validation remains limited. AWS supplied the architecture, documentation, and customer example, but no public benchmark compares leakage rates, latency, API overhead, or failure behavior.

Mondelēz International provides the announcement’s main customer signal. AWS says the company has deployed Amazon Quick for more than 35,000 employees across four regions.

Jamahl Wiggins, a senior M365 innovation specialist at Mondelēz, said real-time access control helped satisfy security and compliance reviewers. The statement shows enterprise demand, although it does not substitute for an independent security assessment.

Buyers should request evidence from their own environment. A representative pilot needs real group structures, frequent permission changes, sensitive content, and controlled attempts to retrieve revoked information.

Teams should also measure false denials. A system that never leaks because it frequently drops authorized content can still fail as a knowledge product.

Useful acceptance metrics include authorization accuracy, retrieval completeness, added latency, token renewal failures, throttling rates, and the percentage of unanswered questions caused by verification.

The strongest conclusion is therefore narrower than the marketing message. Amazon Quick RAG access control gives supported deployments a fresher authorization decision, while leaving the surrounding security system intact and necessary.

Three Signals Will Show Whether the Design Holds at Enterprise Scale

The next test is whether source-authoritative verification remains accurate, observable, and responsive across real repositories and connector types.

The first signal is documented connector coverage. AWS’s announcement names SharePoint, Google Drive, and Confluence as central enterprise sources, while its detailed examples focus on Google Drive and SharePoint.

Buyers should watch for source-specific documentation that explains which connectors perform live checks. Documentation should also distinguish admin-managed, user-managed, and custom configurations.

That distinction matters because similarly named knowledge bases can carry different authorization behavior. One Google Drive configuration might use user authorization, while another relies on a service account and impersonation.

If AWS publishes consistent verification semantics across more connectors, the case for a common enterprise security model becomes stronger. If coverage remains narrow, teams will still operate mixed assurance levels.

The second signal is operational evidence. Enterprises need latency distributions, throttling behavior, timeout handling, retry rules, fail-closed semantics, and logs that connect each answer to its authorization checks.

Real-time verification is convincing during a normal request. Its credibility depends on what happens when Microsoft Graph, Google Drive, or another source responds slowly or not at all.

A mature implementation should make these failures visible without exposing sensitive document names. Administrators should be able to separate missing content, failed retrieval, and denied authorization.

AWS can strengthen confidence by documenting audit events and service limits. Customer case studies can help when they include measured behavior rather than only governance approval.

The Mondelēz deployment creates an important reference point because AWS reports more than 35,000 employees across four regions. Future details about adoption, reliability, and support operations would make the example more informative.

If large customers report stable performance under frequent permission changes, the architecture gains practical support. If they require broad exemptions or frequent troubleshooting, its operational burden becomes clearer.

The third signal is how competitors and internal platform teams respond. Query-time authorization can become a standard procurement requirement for enterprise RAG rather than an optional security feature.

Vendors may expose similar source validation, provide permission-aware retrieval through native enterprise search, or argue that synchronized indexes can deliver equivalent assurance with shorter latency.

Custom RAG teams face the same choice. They can add source calls, rely on carefully synchronized ACL metadata, isolate security domains into separate indexes, or query an existing permission-aware search system.

Each route has a tradeoff. Live checks add dependencies, replicated ACLs create freshness risk, separate indexes increase operational complexity, and inherited enterprise search can constrain retrieval design.

The market response will reveal whether source-authoritative verification becomes a baseline or remains a premium architecture for highly sensitive repositories.

For enterprise buyers, the immediate action is straightforward. Ask every RAG provider where the final document authorization decision occurs.

Then revoke access to a sensitive file and query for its contents before the next scheduled synchronization. Repeat the test through direct questions, summaries, citations, and follow-up prompts.

Review the logs when access fails. Confirm whether the system contacted the authoritative source, which identity it presented, and whether the document ever entered model context.

Amazon Quick RAG access control raises the standard by moving the final check closer to the source and closer to query time. The design deserves attention because it addresses a concrete exposure window.

Its lasting value will depend on connector coverage, transparent failure behavior, and measurable performance under real enterprise load. Can your current RAG system answer those same authorization questions with evidence rather than assurances?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page