top of page

Amazon Contract Intelligence Platform Takes On RAG's Portfolio Blind Spot

4 hours ago
12 min read

Amazon published a contract intelligence architecture built around eight extracted fields, creating an Amazon contract intelligence platform for questions that RAG-only chat mishandles.

The reference design targets a stubborn enterprise problem. A chatbot can often retrieve one payment clause from one agreement. It becomes unreliable when asked to total values, compare dates, or count unsigned documents across hundreds of contracts.

Amazon's answer is not a larger prompt or a longer context window. It separates document understanding from portfolio analysis. AI agents extract and verify fields, a database handles calculations, and Amazon Quick gives users one interface for both workflows.

That distinction puts retrieval-only contract assistants under pressure. The system treats retrieval-augmented generation, or RAG, as one component rather than the database for every question. The result is a useful test of where enterprise agents belong and where conventional data systems remain essential.

What Amazon Actually Built

The design turns every uploaded contract into both a searchable document and a structured database record.

AWS published the contract intelligence architecture on September 29, 2026. It is a reference implementation, not an announcement of an autonomous legal service or a customer deployment.

A React application provides the user interface. Contract PDFs enter an Amazon Simple Storage Service bucket, which triggers an automated processing pipeline.

The first agent reads each PDF and extracts eight fields. AWS says the implementation uses Anthropic's Claude Sonnet 4.6 and returns structured JSON with a confidence score for every field.

A second agent independently reads the same document. This verifier uses Claude Haiku 4.5, according to AWS, and compares its findings with the extractor's output.

The two agents run through the open-source Strands Agents SDK on Amazon Bedrock AgentCore. AgentCore provides the managed runtime where the agent code executes and scales.

A disagreement does not automatically mean the verifier wins. For signature status, the system calls Amazon Textract as a visual tiebreaker. The verified record then goes into Amazon Aurora PostgreSQL.

The original PDF follows another path. It remains available through a knowledge base for document-specific retrieval, including questions about payment terms or individual clauses.

Amazon Quick sits across both sources. Its Quick Sight capability displays dashboards built from database records. Its conversational interface can route questions toward structured analytics or document retrieval.

The distinction matters because those sources answer different question types. A clause lookup needs relevant text. A portfolio total needs every applicable record and a dependable calculation.

The design also streams pipeline status to the browser through WebSocket connections. Users can see whether a file is being extracted, verified, checked, or stored.

That visibility is more than interface polish. An enterprise workflow becomes easier to inspect when users can see which processing stage produced a delay or disagreement.

AWS says a contract can move through the pipeline in seconds under typical conditions. That remains a design claim, and actual performance will depend on document length, concurrency, model availability, and regional configuration.

The publication follows an earlier AWS contract-management design from January 2026. That version emphasized multiple specialist agents for legal, risk, compliance, and workflow tasks.

The new architecture is narrower and more revealing. It concentrates on extraction accuracy and portfolio analytics, which expose weaknesses that conversational demonstrations often hide.

Why RAG Cannot Total a Contract Portfolio

RAG selects relevant passages, while portfolio analysis requires complete records and controlled calculations.

The original RAG research paired a language model with retrieved external knowledge. That pattern helps a model answer questions without placing an entire source collection inside its prompt.

A typical system splits documents into chunks and creates vector representations for them. When a user submits a question, semantic search retrieves a limited set of closely related chunks.

This mechanism works well when the desired answer exists in a few passages. A question about termination language can retrieve the relevant clause without reading every page again.

The mechanism becomes a liability when the question spans the full collection. Consider a procurement leader asking for the total committed value across every active agreement.

The retriever still selects the chunks that appear most relevant. It does not guarantee that every active agreement contributes one complete, correctly normalized value.

Increasing the number of retrieved chunks does not fully solve the problem. Contracts contain repeated labels, amendments, tables, footnotes, and conflicting dates. Relevant passages can also outnumber the model's usable context.

The missing capability is not conversational fluency. It is coverage.

AWS illustrates the issue with a hypothetical portfolio of 250 contracts. At 10 to 20 pages each, that portfolio can contain as many as 5,000 pages.

A person can eventually inspect every document and maintain a spreadsheet. However, every new agreement or amendment can make the spreadsheet stale.

A RAG assistant can answer more quickly, yet still omit records outside its retrieval window. The answer may sound complete even when the calculation covers only a subset.

This is the main contest behind the Amazon contract intelligence platform: RAG-only chat versus extraction followed by database queries.

Under the extraction approach, every contract passes through the same schema. Values, dates, counterparties, signature states, and other selected fields become rows and columns.

A database can then filter active records, group them by vendor, and calculate totals. It can also return the records behind the result for further inspection.

That architecture does not make RAG obsolete. It assigns RAG a narrower job that matches its strengths.

Document retrieval remains valuable for questions that resist normalization. Payment language, liability provisions, exceptions, and unusual obligations often need their surrounding text.

Structured analytics serves a different layer. It supports questions such as how many agreements have expired, which unsigned contracts carry the largest value, or what renewals approach a selected date.

The two paths can complement each other. A structured query identifies the contracts requiring attention, while retrieval brings back the supporting clauses.

This division also produces a clearer failure model. Retrieval errors affect a document answer. Extraction errors can affect dashboards and every aggregate built from the stored field.

That makes the ingestion pipeline more consequential than the chat interface. The conversational layer is only as trustworthy as the records and retrieval sources behind it.

How the Amazon Contract Intelligence Platform Changes the Data Path

The architecture moves the hardest work from question time to ingestion time.

A RAG-only assistant postpones interpretation until someone asks a question. The Amazon contract intelligence platform interprets selected contract fields when each document enters the system.

That change creates a reusable analytical layer. Once a renewal date has been extracted, verified, and stored, multiple dashboards and questions can use the same normalized value.

The first step is document ingestion through Amazon S3. An upload starts the processing flow without requiring an analyst to open the file manually.

The extraction agent then reads the PDF natively, according to AWS. It produces the eight expected fields and attaches confidence scores that downstream components can inspect.

Confidence scores are signals, not guarantees. They can help rank review work, but they do not establish that an extracted value matches the legal meaning of a clause.

The independent verifier introduces a second reading. Using a different model family member is intended to reduce correlated errors from repeating the same extraction process.

This idea resembles review by two people, but the analogy has limits. Two models from the same provider can still share training patterns, blind spots, and document-handling weaknesses.

AWS evaluated extractor and verifier combinations with 20 contracts. The team hand-labeled eight fields in each contract, producing 160 ground-truth values.

That evaluation suggested the extractor mattered more than the verifier. A more capable extraction model preserved results when paired with a lighter verifier, according to the authors.

AWS also says stronger models did not always provide a meaningful improvement on that dataset. The company recommends testing available models against each organization's contracts and acceptance criteria.

That qualification is important. Twenty contracts form a directional test, not evidence that the selected combination will generalize across industries, languages, or drafting styles.

After verification, the database becomes the system for aggregate questions. Amazon Quick connects to Aurora PostgreSQL and can query live data rather than waiting for a separate export.

Amazon Quick also connects to the knowledge base containing the source documents. Its chat agent can therefore support structured questions and individual document lookups within one interface.

This routing is the architectural mechanism behind how Amazon contract intelligence works. The model does not perform every calculation by reading contract prose during each conversation.

For an aggregate request, the structured source supplies filtered records and calculations. For a document-specific request, the knowledge base retrieves relevant contract text.

AWS presents example outputs based on a sample portfolio containing 20 contracts. The examples include portfolio value, expired-contract totals, and counts of signed and unsigned agreements.

Those numbers illustrate the interface. They are not operational results from a disclosed customer portfolio and should not be read as performance evidence.

Embedded dashboards provide another way to inspect the same data. Users can view portfolio totals, signature status, extraction results, and confidence comparisons without leaving the application.

That shared foundation can reduce disagreements between dashboard and chat outputs. Both interfaces can reference the same verified database records for analytical questions.

The design also preserves access to source material. Analysts do not have to treat a database value as the final word when a contractual decision requires reading the clause.

This hybrid pattern reaches beyond contracts. Insurance claims, compliance filings, leases, and onboarding records can all mix repeatable fields with document-specific language.

It also matches a broader knowledge blending principle. Structured facts and source context serve different purposes, and useful systems need a governed bridge between them.

Verification Matters More Than Another Model Call

The most instructive part of the design is its refusal to let language models settle every dispute.

During testing, AWS found that the verifier sometimes labeled empty signature blocks as signed. Reported confidence reached between 95 and 100 percent in some false-positive cases.

The model recognized words and layout associated with signing. It then treated the presence of a signature field as evidence that somebody had signed it.

That failure exposes a recurring problem with generative AI. A confidence score can describe the model's internal certainty without proving that the underlying conclusion is correct.

AWS responded by adding Textract only when the two models disagreed about signature status. Textract examines visual features to detect handwritten or digital signatures.

The choice creates a three-part control. One model performs extraction, another model checks it, and a specialized computer-vision service resolves a defined class of disagreement.

This is a better fit than asking a third language model to vote. A third model might reproduce the same semantic confusion between a signature line and an actual signature.

The pattern also limits the specialized check to contested cases. That preserves the primary workflow while applying a different technical method where the models show uncertainty.

However, Textract is a tiebreaker for signature presence, not an arbiter for every field. Contract values, renewal rules, dates, and party identities can create different ambiguities.

An amendment might replace an earlier value. An automatic renewal clause can require contextual interpretation. A signature may exist while the agreement remains incomplete for another reason.

Those cases need explicit escalation rules. The AWS post says unresolved disagreements should go to a human reviewer when deterministic services cannot settle them.

That human path is central to responsible adoption. It determines whether automation reduces routine work or simply hides uncertain records inside a database.

The system should preserve each extracted value, its source location, both model outputs, and the final resolution. Otherwise, reviewers cannot reconstruct why a dashboard includes a particular figure.

Organizations also need field-specific thresholds. A wrong contact name and a wrong termination date do not carry the same operational risk.

The NIST AI framework emphasizes testing, evaluation, verification, and validation for AI systems. Contract workflows need those practices at the data-field level.

Teams should measure precision, recall, and exact-match performance by field. They should also track disagreement rates, reviewer overrides, and errors discovered after approval.

Accuracy should be segmented by document type. Native PDFs, scanned pages, tables, amendments, handwritten annotations, and multilingual agreements can behave differently.

The 20-contract AWS evaluation offers a starting methodology. It does not provide enough diversity to establish production reliability for another organization's portfolio.

A useful pilot should sample the documents that create the most risk. That includes unusual templates and low-quality files, not only clean agreements from a standard form.

The design's dual-model verification remains valuable because it makes disagreement observable. It creates a measurable event that can trigger deterministic checks or human review.

Yet agreement between models cannot be treated as ground truth. Two agents can produce the same wrong value, especially when the source text is ambiguous.

The essential control is traceability. Every structured value should lead a reviewer back to the page, clause, and extraction decision that produced it.

Accuracy, Access, and Economics Remain Unproven

The reference architecture defines a credible mechanism, but it does not establish production readiness for every contract portfolio.

The first uncertainty is evaluation scale. AWS used 20 contracts and 160 labeled values when comparing model combinations.

That sample can reveal obvious differences between configurations. It cannot represent the full range of formatting, drafting, scanning, language, and amendment patterns in enterprise agreements.

The second uncertainty is schema coverage. Eight fields can support useful dashboards, but contract operations often depend on more complex obligations and conditional dates.

A renewal date may depend on notice periods. A value may combine committed fees, consumption charges, credits, or indexed increases.

Flattening those terms into one record can create an illusion of certainty. The schema must preserve qualifiers when a field cannot be represented as a simple number or date.

The third uncertainty concerns model change. AWS notes that model availability evolves and advises builders to retest the system before relying on new options.

A replacement model can alter extraction behavior even when the surrounding application remains unchanged. Production teams therefore need fixed evaluation sets and release gates.

The fourth uncertainty is authorization. Contracts can contain confidential pricing, employee information, security obligations, and strategic vendor terms.

AWS says AgentCore supports session isolation, while policy controls can sit outside agent code. The AgentCore documentation describes separate execution environments for user sessions.

However, infrastructure isolation does not automatically create correct business permissions. AWS documentation states that client backends must maintain the relationship between users and session identifiers.

Builders must also enforce access at the document, database, dashboard, and agent-tool layers. A chat interface should not reveal a record that the same user cannot open elsewhere.

The fifth uncertainty is operational economics. The AWS post provides a sample cost model, but commercial terms change and portfolio workloads differ.

Processing costs depend on page counts, model calls, retries, storage, database capacity, concurrency, and the frequency of user queries. Human review can become the largest variable.

A credible business case should measure cost per accepted record, not cost per model invocation. A cheap extraction that creates extensive review work is not economically cheap.

The competitive landscape also offers several implementation routes. Microsoft's contract extraction model returns structured fields from contracts through Document Intelligence.

Google's Document AI likewise converts unstructured documents into structured data for downstream processing.

These products differ in schemas, customization, orchestration, analytics, and cloud integration. Buyers should compare the complete control loop rather than one extraction benchmark.

Amazon's differentiation in this design is the connection between agents, verification, a relational database, embedded analytics, and document Q&A.

That integration can appeal to organizations already operating on AWS. It can also increase architectural commitment across storage, models, databases, analytics, and identity controls.

Teams should test portability before production. Extracted records and provenance data should use documented schemas that remain accessible outside one conversational interface.

They should also define failure ownership. Procurement, legal operations, data engineering, and security teams may each assume another group validates the output.

A workflow needs one accountable owner for schema changes, evaluation thresholds, exception queues, and access reviews. Without that ownership, the system can automate inconsistency.

The larger lesson is that contract intelligence is a data-governance project with an AI ingestion layer. Treating it as a chatbot deployment understates the work.

What Buyers Should Watch Next

Three signals will show whether this architecture becomes a dependable operating system for contract data or remains a persuasive reference demo.

The first signal is evaluation evidence from larger and more diverse contract sets. Buyers need field-level results across scans, amendments, tables, languages, and uncommon drafting patterns.

Published accuracy alone will not be enough. Useful evidence should disclose dataset composition, error categories, review policies, and how often both models agreed on an incorrect value.

Better evidence would strengthen Amazon's central claim that independent verification improves reliability. Persistent correlated errors would weaken the case for the dual-model design.

The second signal is production-grade exception handling. The platform needs configurable review queues, source citations, approval history, and clear routes for unresolved disagreements.

Amazon Quick includes human-in-the-loop capabilities, but organizations must show how those controls fit daily contract operations. Reviewers need context, not another unexplained confidence score.

A mature workflow should allow someone to correct a field, document the reason, and update downstream analytics without losing the original output.

Successful deployments will also separate low-risk automation from high-risk decisions. Portfolio counts can tolerate different controls than termination notices or financial commitments.

The third signal is adoption beyond demonstration workloads. Watch for disclosed customer deployments that connect extraction quality to measurable operational outcomes.

Relevant indicators include review time, correction rates, stale-record frequency, and the percentage of contracts processed without escalation. Those measures matter more than chatbot response speed.

Competition will shape adoption as well. Microsoft and Google already support structured document extraction, while contract lifecycle platforms offer domain-specific workflows.

Amazon must show that Quick and AgentCore reduce integration work without limiting schema flexibility or governance. Competitors must show equally coherent paths from documents to verified portfolio analytics.

The Amazon contract intelligence platform makes its strongest argument through architecture, not model novelty. It recognizes that retrieval, extraction, verification, and calculation are separate jobs.

That recognition gives enterprise teams a practical decision rule. Keep RAG for locating and explaining source passages. Use structured records for counting, sorting, comparing, and totaling.

Then place review controls where a wrong answer would alter a legal or financial decision. No model confidence score should bypass that judgment automatically.

For teams evaluating contract AI, the next step is a representative pilot. Select difficult documents, label the required fields, define acceptance thresholds, and measure reviewer effort.

Ask whether every answer can be traced to its source. Test permissions across users, contracts, dashboards, and chat sessions. Recalculate aggregates after corrections and amendments.

Most importantly, compare the hybrid architecture against the process it replaces. Does it produce fresher portfolio data with fewer hidden errors and a defensible audit trail?

That is the test that matters. An Amazon contract intelligence platform succeeds when it makes the entire portfolio reliably queryable, not merely when one contract produces a convincing answer.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page