top of page

Amazon Quick Compliance Trades Open-Ended AI for Provably Complete Lease Reviews

5 days ago
14 min read

Amazon published an Amazon Quick compliance architecture that can sweep thousands of leases without trusting an open-ended chat agent to decide what counts.

The design, called the Adjudicated Query pattern, connects Amazon Quick to a bounded Model Context Protocol server over a deterministic rules engine. The language model handles conversation, but rules and structured data determine which leases were evaluated and which conditions failed.

That division challenges the usual enterprise chatbot model. Retrieval-augmented generation can locate a relevant clause, yet it cannot guarantee that every applicable document entered a portfolio-wide answer. Amazon instead treats complete coverage as a database and rules problem.

The result is less autonomous than an agent that improvises its own analysis. It is also easier to defend. Each finding can point back to a lease, a clause, a rule version, an extracted value, and the expected value.

This is not merely a new way to search contracts. It is a proposal for deciding where generative AI should stop when an incomplete answer creates legal, financial, or regulatory exposure.

Amazon Quick Compliance Now Comes With a Completeness Receipt

The central change is that Amazon Quick can present a conversational compliance answer backed by evidence that the full eligible population was checked.

AWS published the lease compliance design on October 2, 2026. The post includes a reference architecture and a deployable AWS Cloud Development Kit sample.

The running example asks a deceptively simple question: Which leases fail a defined compliance requirement?

A conventional chat agent might search lease text, retrieve several relevant passages, and summarize the apparent exceptions. That response can be useful, but its fluent wording does not establish population coverage.

The Adjudicated Query pattern changes the execution path. Amazon Quick remains the conversational front door, while a bounded MCP server exposes approved operations to the agent.

Model Context Protocol, or MCP, is an interface that lets AI applications call tools and retrieve contextual resources. The official MCP architecture separates the AI host from servers that provide specific capabilities.

“Bounded” is the important word in Amazon’s design. The server does not give the model unrestricted access to arbitrary code or a general-purpose database connection.

Instead, it offers narrow compliance operations. Those operations sit above a deterministic rules engine, which evaluates predefined conditions against structured lease data.

The architecture diagram also places an Amazon Aurora store beneath both the chat experience and an Amazon Quick Sight dashboard. A Lambda-hosted MCP server mediates chat requests through Amazon API Gateway and Amazon Cognito.

Amazon Bedrock appears only where the workflow needs model reasoning. That detail expresses the pattern’s governing principle: use probabilistic AI for language tasks, then use deterministic components for exhaustive evaluation.

A returned answer can therefore include more than a list of suspicious leases. AWS shows a response containing sample noncompliant findings, count totals, a synthetic-data caveat, and a dashboard link.

Those totals form a completeness receipt. The receipt records the evaluated population and the number of resulting findings, giving reviewers a way to challenge the scope.

A separate findings dashboard provides one row for each lease-rule pair. Each row carries the lease identifier, the rule that fired, the extracted value, and the expected value.

A detail view then connects the finding to the underlying text. It displays the verbatim lease clause beside the applicable rule, its version, its citation, and the compared values.

That presentation turns a chat answer into the beginning of a review trail. The user can move from a portfolio claim to an individual result and then to the source language.

The sample remains a reference implementation, not evidence from a disclosed production portfolio. AWS uses synthetic data in its illustrated response, so the screenshots do not prove real-world accuracy or throughput.

However, the architectural change is concrete. The chat agent no longer carries sole responsibility for interpreting scope, applying every rule, calculating totals, and explaining the result.

It delegates those jobs to components that can expose their inputs and outputs. That makes Amazon Quick compliance an orchestration problem rather than a prompt-writing exercise.

Why Retrieval Alone Cannot Prove Every Lease Was Checked

Semantic relevance and complete coverage answer different questions, even when both systems return convincing text.

Retrieval-augmented generation, or RAG, searches a document collection for passages related to a user’s request. The model then uses a limited selection of those passages to compose its response.

That process works well when someone wants to locate a renewal clause in one lease. It can also summarize unusual wording or compare a small number of identified provisions.

Portfolio compliance demands a different guarantee. The system must establish which documents belong in scope, apply each relevant rule, and record the outcome for every required lease-rule combination.

A retriever ranks passages by relevance. It does not naturally prove that every lease contributed a result.

Increasing the retrieval limit does not turn semantic search into exhaustive evaluation. Large portfolios can exceed a model’s usable context, while repeated boilerplate can crowd out less common clauses.

Document boundaries also matter. A retrieved passage might omit an amendment, exhibit, or definition that changes how the clause should be interpreted.

The weakness becomes clearer when a user asks for a negative. “Show every lease that lacks a required term” requires evidence about documents where no matching clause was found.

Search systems are optimized to retrieve what exists. Proving that something does not exist across thousands of documents requires a defined population and a recorded check against each member.

A plausible answer can therefore be incomplete without looking obviously wrong. That is a dangerous failure mode because the interface rewards readability while hiding omitted records.

The Adjudicated Query pattern assigns scope to structured data. The system can select an eligible lease population through explicit filters, then pass that population to the rules engine.

Every rule evaluation can produce a recorded state. A lease can pass, fail, require review, or remain unevaluated because a required value is missing.

Those distinctions matter. Treating “not found” as “compliant” would conceal extraction failures, while treating every missing value as a violation could overwhelm reviewers.

The completeness receipt gives users a basic reconciliation mechanism. If the portfolio contains a defined number of eligible leases, the result should account for that same population.

That does not guarantee semantic correctness. A rule can still encode the wrong policy, and an extracted value can still misrepresent a clause.

It does establish procedural coverage. Reviewers can ask whether the expected population was processed, whether all active rules ran, and whether any records ended in an unresolved state.

This is the primary opponent in Amazon’s design: open-ended model judgment versus bounded, auditable execution.

The contrast does not make generative AI useless. The model remains valuable for interpreting a natural-language question, collecting required parameters, and explaining structured results.

It can also help a user refine scope. Someone might ask about active retail leases in selected jurisdictions, then narrow the answer to renewals occurring during a particular period.

However, the agent should not silently invent the legal meaning of “active,” “retail,” or “compliant.” Those definitions belong in governed fields, approved rules, or an explicit clarification step.

That boundary is central to defensible lease compliance automation. The model translates between people and the system, but it does not become the policy system.

This separation resembles effective knowledge blending. Source language, structured facts, and governed calculations remain distinct, while the interface connects them for the user.

The practical payoff is not a more eloquent answer. It is an answer whose scope can be counted, whose findings can be inspected, and whose governing logic can be named.

The Adjudicated Query Pattern Moves Authority Outside the Model

Amazon’s mechanism works because the language model requests an adjudicated result rather than generating the result from retrieved prose.

The word “adjudicated” signals that another component settles the compliance question under explicit rules. The model can ask for that decision, but it cannot alter the decision procedure during the conversation.

A typical interaction starts in the Amazon Quick chat agent. The user describes a portfolio question in ordinary language, perhaps asking for leases that violate a notice requirement.

The agent identifies an approved MCP operation and supplies the required parameters. Those parameters can include rule identifiers, dates, jurisdictions, lease categories, or other governed filters.

The MCP server validates the request before passing it onward. A narrow tool contract can reject missing, malformed, or unauthorized parameters instead of letting the model improvise around them.

The rules engine then applies a deterministic test. Given the same data, rule version, and parameters, it should return the same evaluation result.

That repeatability is important during review. A compliance team can reproduce an earlier answer even after the chat session ends.

The underlying Aurora store supplies structured values and identifiers. It also provides a place to retain rule evaluations, findings, and provenance beyond a model’s temporary context.

Amazon Quick Sight presents the resulting records as a dashboard. That gives analysts a filterable view that does not depend on conversational phrasing.

The chat interface and dashboard therefore become two views over the same adjudicated findings. One explains and navigates the results, while the other supports inspection across rows and filters.

AWS’s detail-view illustration adds another layer. A reviewer can see the source clause beside the rule that fired, including the rule version and citation.

Rule versioning matters because compliance policy changes. An answer should identify which policy definition governed the evaluation at that moment.

Without that identifier, a team cannot explain why the same lease passed last quarter and failed after a policy update. It also cannot reproduce an earlier report fairly.

A rule citation provides policy context. It can connect a technical condition to an internal control, contractual standard, or governing requirement.

The extracted value shows what the system believed the lease said. The expected value shows the threshold or condition used during comparison.

Together, those elements create a defensible chain: source text, structured interpretation, approved rule, deterministic comparison, and reported finding.

The AWS CDK sample also changes how teams can evaluate the idea. CDK defines cloud infrastructure in code, allowing builders to deploy repeatable stacks rather than assemble the reference manually.

The official CDK guide explains how applications synthesize infrastructure definitions into deployable AWS resources. That model supports review and version control for the sample architecture.

Infrastructure as code does not make the compliance logic correct. It does make the environment easier to reproduce, inspect, and remove after testing.

Identity remains part of the mechanism. The reference architecture routes requests through Cognito and API Gateway before they reach the Lambda-hosted MCP server.

That path creates places to authenticate users, authorize operations, limit requests, and log access. Each control still requires configuration aligned with the organization’s policies.

The agent should never become an authorization shortcut. A user who cannot access a lease through the dashboard should not retrieve its clause through chat.

The same rule applies to aggregate results. A total can reveal restricted information even when it hides individual rows.

Teams therefore need access controls at several layers: source documents, structured records, rule execution, findings, dashboards, and conversational responses.

The mechanism is more involved than connecting a folder to a chatbot. That complexity is the cost of making compliance answers inspectable.

It is also the pattern’s strongest argument. High-stakes automation should reveal where policy, computation, model reasoning, and human judgment enter the result.

Lease Compliance Automation Still Depends on Extraction Quality

Deterministic rules cannot rescue an incorrect structured value, so the architecture moves risk rather than eliminating it.

The rules engine evaluates the data it receives. If the system extracted a notice period incorrectly, a perfectly executed rule can still produce the wrong finding.

That creates a critical distinction between procedural completeness and substantive correctness. The completeness receipt can prove that every eligible record was processed, but not that every record was understood correctly.

Lease language makes this problem difficult. A requirement may appear in the main agreement, an amendment, an exhibit, or a definition referenced from another section.

Dates can depend on commencement conditions rather than a printed calendar value. Renewal terms can combine an initial period, optional extensions, and deadlines calculated from another event.

Numeric values can also carry qualifiers. A lease might specify different thresholds by year, location, use category, or operating condition.

A flat field cannot safely represent every variation. The data model needs explicit states for ambiguity, conflicts, missing documents, and unresolved dependencies.

Source citations help reviewers detect these problems. A finding should lead directly to the clause and surrounding context used for extraction.

Yet citation is not validation. A model can point to the correct paragraph while still interpreting its effect incorrectly.

Organizations need field-level evaluation before relying on lease compliance automation. Tests should measure errors separately for dates, options, monetary values, notice periods, and policy-specific classifications.

The test set should include difficult material. Scanned pages, tables, handwritten changes, amendments, unusual templates, and poor optical character recognition can expose failures hidden by clean samples.

Teams should also test correlated mistakes. Multiple model calls do not provide independent assurance when they share similar training patterns or receive the same incomplete context.

Human review should focus on consequential and uncertain cases. A system can route missing values, conflicting amendments, low-confidence extractions, and unusual clauses into a queue.

The rules themselves require equivalent scrutiny. A deterministic implementation can consistently apply an incorrect policy.

Each rule needs an owner, an approval history, an effective date, and tests covering expected passes and failures. Changes should be reviewed like production code.

Organizations should preserve earlier rule versions instead of overwriting them. Historical reports need the logic that produced them.

They should also record the evaluated population before running the sweep. Otherwise, later data changes can make the original completeness claim impossible to reconstruct.

The NIST AI framework emphasizes governance, measurement, and management across an AI system’s lifecycle. Those practices fit this architecture better than a one-time accuracy benchmark.

Operational metrics should include extraction correction rates, unresolved records, rule failures, access denials, and reviewer overrides. Aggregate accuracy alone can hide concentrated errors in high-risk fields.

Latency and scale remain open questions as well. The AWS post describes sweeping thousands of leases, but the reference publication does not disclose a customer benchmark across a real portfolio.

Actual performance will depend on stored data quality, rule complexity, database capacity, concurrency, model use, and the number of lease-rule pairs.

The sample screenshots use synthetic data. They illustrate the user experience, not validated production outcomes.

That limitation does not invalidate the pattern. It defines the next testing requirement.

A serious pilot should compare the system against a labeled portfolio and an existing review process. It should measure both missed violations and unnecessary escalations.

False negatives create hidden exposure. False positives consume legal and operational time, potentially erasing the efficiency gained from automated screening.

The best deployment target is therefore not immediate autonomous judgment. It is a controlled workflow that finds review candidates, proves coverage, and keeps source evidence within reach.

Bounded MCP Tools Reduce One Risk but Create New Control Points

A narrow MCP server limits agent freedom, yet every exposed operation still expands the system’s security and governance surface.

MCP makes tool integration easier by giving agents a standard way to discover and invoke capabilities. That convenience can become risky when servers expose broad actions or accept loosely validated arguments.

Amazon’s bounded approach reduces that risk. A compliance agent needs approved query and adjudication functions, not arbitrary SQL, shell access, or unrestricted document retrieval.

A small tool set is easier to review. Security teams can identify which operations exist, what each operation accepts, and which data it can return.

Input validation is essential because natural-language requests can contain ambiguous or hostile content. The server should treat model-generated arguments as untrusted input.

Authorization must happen when the tool runs, not only when the user opens Amazon Quick. A valid session does not imply permission to access every lease or rule.

The architecture’s use of Cognito and API Gateway offers enforcement points. However, builders still need to map identities, groups, leases, portfolios, and permitted operations correctly.

Logging also requires care. Audit records should capture who requested a sweep, which scope and rule versions were used, when it ran, and what result identifier was returned.

Logs should avoid duplicating sensitive lease text unnecessarily. A complete audit trail does not require copying confidential clauses into every infrastructure log.

Prompt injection remains relevant even with deterministic rules. A malicious clause could contain text designed to influence a model that extracts or explains the document.

Bounding the MCP tool prevents such text from rewriting the rules engine. It does not automatically stop the model from producing a misleading narrative around a valid result.

The interface should distinguish generated explanation from adjudicated output. Counts, rule identifiers, and finding states should come directly from the controlled service.

Generated prose should not silently change “unresolved” into “compliant.” It should preserve uncertainty expressed by the structured result.

Tool descriptions also deserve review. Agents select tools partly from their names and descriptions, so unclear metadata can cause routing errors.

Schema evolution creates another control point. Adding a field or changing an enumeration can break the assumptions built into rules, dashboards, and model prompts.

Teams should version tool contracts and test backward compatibility. A compliance report must not change meaning because an MCP schema shifted without coordinated review.

Availability matters too. If the rules service fails, the agent should report that no adjudicated answer is available.

It should not fall back to an unbounded model-generated judgment unless the interface labels that result clearly and policy permits it.

The system also needs limits on portfolio sweeps. Expensive or large operations may require pagination, asynchronous execution, quotas, or explicit approval.

A user should receive a stable job identifier rather than wait for a chat session to retain the entire process state.

Results should remain accessible through governed storage and the dashboard. The chat transcript should not become the sole system of record.

These controls make the pattern less magical than many agent demonstrations. They also make it more credible for regulated work.

The wider enterprise AI market often emphasizes how many actions an agent can perform. Amazon’s proposal makes the opposite argument: trust grows when the agent’s authority is deliberately small.

That principle extends beyond leases. Insurance policies, vendor contracts, safety inspections, and regulatory filings all combine natural-language evidence with rules that demand complete application.

The reusable idea is not an AWS service list. It is the separation of conversational flexibility from decision authority.

Three Signals Will Test the Adjudicated Query Pattern

The pattern will matter if real deployments prove complete coverage, manageable review costs, and durable governance beyond the reference sample.

The first signal is production evidence from diverse lease portfolios. Buyers should look for disclosed evaluations across scanned documents, amendments, tables, jurisdictions, and drafting styles.

Useful evidence will separate population coverage from extraction accuracy. It will also report false negatives, false positives, unresolved records, and human corrections by field.

If deployments consistently reconcile every eligible lease while keeping critical extraction errors low, the case for Amazon Quick compliance becomes stronger.

If teams can prove coverage but still need to reread most documents, the architecture will function mainly as a better review queue.

The second signal is mature rule lifecycle management. Enterprises need approvals, effective dates, test cases, citations, rollbacks, and historical reproducibility for every rule.

A finding should retain the exact rule version used during evaluation. Updating a policy should create a new governed version instead of changing earlier results silently.

Watch whether AWS or its partners provide clearer workflows around authoring, testing, approving, and retiring rules. The reference architecture establishes the execution pattern, but operational governance determines whether teams can sustain it.

The third signal is whether bounded MCP designs become a standard procurement requirement for high-stakes agents. Buyers increasingly need to distinguish assistants that explain evidence from systems authorized to make decisions.

The emerging GenAI profile highlights risks specific to generative systems and complements broader AI governance work. Implementations can use that guidance to define testing and oversight expectations.

A bounded tool contract, deterministic adjudication, and a completeness receipt provide concrete controls for that conversation. They make system behavior easier to describe than an agent whose capabilities change with its prompt.

However, the receipt must remain meaningful. It should show the intended population, processed population, exclusions, unresolved records, rule versions, and execution time.

A single total without those details can create false assurance. Completeness depends on scope, and scope depends on data quality and policy definitions.

Organizations evaluating the pattern should begin with one material compliance question. Define the eligible population, encode the rule, label a representative test set, and identify cases requiring legal judgment.

Then compare the automated findings with the current process. Measure reviewer time, corrections, missed conditions, unresolved cases, and the effort needed to explain each result.

Test access boundaries through both chat and dashboard views. Confirm that aggregate answers cannot reveal portfolios outside the user’s authorization.

Finally, rerun the same evaluation after changing a rule or correcting a lease value. The system should update predictably while preserving the evidence behind the earlier result.

That exercise will reveal whether the architecture behaves like a governed compliance system or an impressive conversational demonstration.

Amazon’s most important contribution here is not another contract chatbot. It is a clear boundary around what the chatbot is allowed to decide.

For enterprise buyers, that boundary offers a useful demand: do not accept a confident portfolio answer without a population count, versioned rules, and traceable source evidence.

For builders, the next move is equally concrete. Deploy the sample in a controlled environment, replace synthetic records with a representative test set, and try to break the completeness claim.

Can every excluded lease be explained? Can every finding reach its clause? Can reviewers reproduce the result after the policy changes?

Those questions should guide any Amazon Quick compliance pilot. If the system cannot answer them, it is still search with a persuasive interface.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page