top of page

DHS Expands AI for FOIA Processing, but Human Review Is the Real Test

The Department of Homeland Security plans to expand AI-assisted FOIA processing after ending fiscal 2025 with 245,572 pending requests. The department says automation can reduce administrative work and help employees handle a rising volume of public-records requests. The conflict is straightforward: faster processing has value only if the resulting searches, classifications, and redactions remain accurate.

DHS received more than 1 million Freedom of Information Act requests during the fiscal year. It processed nearly the same number, yet its pending inventory still grew from 221,068 requests at the start of the year. According to the department, its formal backlog remained equal to 16 percent of requests received.

Those figures help explain why federal agencies see automation as more than an experiment. However, DHS is considering AI for decisions that influence what records people receive. That makes the plan an accountability test, not merely an office modernization project.

The department already lists a pre-deployment use case for identifying and redacting sensitive content. Customs and Border Protection also uses document-review software that can learn from reviewers and classify records by relevance. Civil-liberties advocates warn that these systems can reproduce existing patterns of excessive redaction while shrinking the space for careful human review.

DHS Is Moving AI Deeper Into the FOIA Workflow

The important change is not that DHS uses software, but that AI is moving closer to consequential review decisions.

DHS has used electronic tools for years to accept requests, manage cases, share files, remove duplicates, and deliver records. Those functions replace manual handling without necessarily deciding what the public can see. The newer use cases reach further into document identification and redaction.

The department’s annual FOIA reporting says its Privacy Office continues to develop and deploy new technology for processing requests. DHS expects to launch additional automation tools during the coming year. It describes the goals as improving efficiency and removing repetitive administrative work.

That description covers a wide range of possible functions. An automated system might populate case fields, detect duplicate files, convert scanned pages into searchable text, or route records to the correct office. Each function can save time without replacing a legal judgment.

Document responsiveness is more sensitive. A responsive record falls within the scope of a requester’s search, even if some information must later be withheld. If a system mistakenly classifies that record as irrelevant, the requester may never know it existed.

Redaction creates another layer of risk. FOIA permits agencies to withhold information under specific statutory exemptions, including protections involving national security, law enforcement, and personal privacy. Applying those exemptions often requires context, legal interpretation, and a judgment about foreseeable harm.

DHS has identified an AI use case intended to streamline the identification and redaction of sensitive material. The department categorized that planned use as not high-impact, according to its public inventory. That designation matters because high-impact systems face stronger risk-management expectations under federal AI policy.

The inventory describes the system as being in pre-deployment. That means the public record does not yet establish how broadly DHS will use it, which components will participate, or what percentage of its suggestions employees will accept.

A separate tool is already part of the picture. DHS lists RelativityOne as a commercial AI system for knowledge retrieval. A Customs and Border Protection privacy analysis describes its role more specifically, including support for timely responses to FOIA requests.

Within CBP’s FOIA division, the platform can learn from reviewer decisions and tag documents as relevant or irrelevant. This technique is commonly called technology-assisted review, meaning software uses examples labeled by people to prioritize or classify a larger collection.

Technology-assisted review does not always mean autonomous release decisions. Reviewers can use predictions to order documents, identify likely duplicates, or focus attention on uncertain classifications. The public documents leave important details unclear, including when employees must verify a tag and how DHS measures errors.

That gap creates the central tension. DHS presents AI as support for overloaded employees, but some documented capabilities can influence which records reach those employees. Efficiency depends on fewer manual steps. Accountability depends on knowing which steps cannot safely be removed.

The Request Numbers Explain Why DHS Is Automating

DHS faces a workload problem large enough to make automation attractive, even before staffing pressure enters the calculation.

The department began fiscal 2025 with 221,068 pending FOIA requests. It then received more than 1 million new requests and processed nearly the same volume. DHS nevertheless finished with 245,572 pending requests, an increase of 24,504 during the year.

A pending request and a formal backlog are not identical. A request becomes backlogged when it remains unresolved beyond the applicable statutory response period. DHS says its backlog represented 16 percent of requests received, even though its total pending inventory exceeded 245,000.

The distinction is useful, but it does not make the operational challenge small. Every pending case still requires tracking, communication, searching, review, or delivery. Complex requests can involve large email collections, immigration files, law-enforcement records, video, or material held across several components.

DHS has long carried an unusually large share of the federal FOIA workload. Its agencies include Customs and Border Protection, Immigration and Customs Enforcement, the Federal Emergency Management Agency, the Transportation Security Administration, and U.S. Citizenship and Immigration Services.

Several of those components maintain records that people need for immediate personal or legal reasons. Immigration requests can involve an individual’s official file. Disaster-related requests can concern government decisions after an emergency. Journalists and oversight organizations seek policy records, contracts, communications, and enforcement data.

The need for speed therefore extends beyond convenience. A delayed record can affect litigation, reporting, immigration proceedings, historical research, or public scrutiny of an active policy.

Federal workforce changes add pressure. More than 600 government information specialists left federal service after the second Trump administration began, according to data cited in the initial report. Nearly 100 of those departures came from DHS.

The government information specialist category includes many employees who process FOIA requests. Departures do not translate directly into an identical reduction in FOIA capacity because job duties vary. Still, the federal workforce data shows a broader contraction affecting agencies that already faced growing records workloads.

Fewer experienced reviewers create a difficult cycle. Remaining employees receive more cases, managers seek faster tools, and reviewers have less time to inspect automated results. The technology intended to relieve pressure can then become harder to supervise properly.

DHS is not alone in this shift. The National Archives’ Office of Government Information Services reported that 18.6 percent of responding agencies used AI or machine learning in FOIA processing. That finding shows adoption is underway, although most agencies still had not reported using these tools.

Other departments are exploring similar functions. The Department of Health and Human Services has described work involving automated workflow features, document-discovery platforms, and an AI proof of concept. Defense officials have also examined agentic, generative, and predictive AI for FOIA operations.

This is becoming a government-wide response to rising demand and constrained staffing. DHS stands out because of its volume, sensitive records, and enforcement responsibilities. A mistake inside this environment can affect both individual requesters and broader public oversight.

The Core Tradeoff Is Speed Versus Reviewable Judgment

AI can accelerate discovery without possessing the legal judgment that FOIA decisions require.

Some tasks offer a relatively clear automation case. Optical character recognition can turn scanned pages into searchable text. Duplicate detection can prevent reviewers from reading the same email thread repeatedly. Language identification and file conversion can prepare records for review.

Search assistance can also help employees find records that a narrow keyword query might miss. Machine-learning systems can identify related concepts, rank likely responsive documents, and group similar material. Those capabilities are valuable when a case includes thousands of files.

Cody Venzke, a senior staff attorney with the American Civil Liberties Union, acknowledged that potential in the original reporting. He said AI might speed document discovery and reduce the chance that broad requests miss responsive material.

His warning concerns what happens next. A redaction is not simply a pattern-matching exercise. Reviewers must identify a legal exemption, apply it to the specific material, and consider whether disclosure would create foreseeable harm.

Context changes the answer. A name might require protection in one record but already be public in another. A passage describing deliberations might contain segregable factual information that should still be released. A law-enforcement technique might be sensitive at one moment and widely known later.

Models can assign labels based on prior examples, but past labels are not necessarily correct. Agency responses reflect legal interpretations, risk preferences, litigation history, and reviewer habits. Training on those responses can reproduce the same tendencies at greater scale.

That concern is especially important because agencies are often accused of over-redacting records. An overly cautious human reviewer can hide too much information in one case. A system trained on thousands of cautious decisions can apply that behavior across thousands of cases.

Abigail Kunkler, a law fellow at the Electronic Privacy Information Center, argued that automation trained on existing responses would reflect existing redaction patterns. Her concern is not merely that an AI model will occasionally fail. It is that the system will standardize an institutional preference for withholding information.

The opposite error also matters. A model might fail to recognize personal data, security information, or legally protected details. That can expose individuals or harm an investigation. Agencies therefore have a rational reason to tune systems toward caution.

However, tuning for caution shifts errors toward non-disclosure. Requesters cannot easily detect a record that a classifier excluded before review. They can challenge visible redactions, but they cannot challenge a document they never received or knew existed.

That asymmetry makes responsiveness classification particularly consequential. A false positive sends an irrelevant document to a reviewer and consumes time. A false negative can silently remove a responsive record from the process.

DHS needs evaluation measures that reflect this difference. A single accuracy score can hide harmful failure patterns. The department should separately track missed responsive records, unnecessary review, incorrect redaction suggestions, and sensitive information exposed.

Testing also needs realistic data. A system evaluated on clean sample documents may perform differently on scanned pages, handwritten notes, email chains, attachments, audio transcripts, or records containing several languages. DHS processes all of these formats across its components.

The strongest use of AI would keep the system in an assistive role. It could rank documents, suggest redactions, and explain which patterns triggered a recommendation. A trained employee would confirm the record’s responsiveness and make the final withholding decision.

That structure preserves human authority, but only if review is substantive. Requiring an employee to press an approval button does not create meaningful oversight. Reviewers need enough time, training, and system information to challenge the recommendation.

DHS Has Not Explained Enough About Its Safeguards

The public inventory identifies AI capabilities but does not provide enough evidence to judge their controls.

DHS categorized its planned sensitive-content identification and redaction use case as not high-impact. That decision limits the risk-management requirements associated with the system. Yet public information does not fully explain why the use case received that classification.

Federal policy definitions matter here. A system can influence access to government information without making a decision about benefits, policing, or personal liberty. That narrower role might place it outside formal high-impact categories.

Still, a low formal classification does not mean a low public-accountability risk. FOIA is one of the main legal mechanisms for examining executive agencies. A tool that changes searches or redactions can shape what journalists, litigants, researchers, and citizens learn about government activity.

The discrepancy between DHS descriptions deserves attention. The common commercial inventory categorizes RelativityOne as a knowledge-retrieval system. CBP’s privacy analysis describes a feature that learns from reviewers and independently tags documents for relevance.

Both descriptions can be technically accurate. Knowledge retrieval includes finding and classifying documents. However, the broader label does not communicate the significance of excluding a record from further review.

Patrick Eddington, a Cato Institute senior fellow, argued that this difference obscures the capability most likely to justify closer review. His criticism points to an inventory design problem: a general category can satisfy disclosure requirements without explaining the operational decision.

DHS should publish the role each tool plays at each stage. The description should distinguish search, prioritization, responsiveness classification, redaction recommendation, final approval, and quality assurance. Combining those functions under “processing” prevents meaningful evaluation.

The department should also explain the human-review threshold. Does every irrelevant tag receive verification, or only a sample? Can the software release a redacted document without line-by-line approval? What happens when the model expresses low confidence?

Audit records are another missing piece. A defensible system should preserve the model version, input collection, classifications, reviewer changes, and final outcome. Those records would help DHS investigate errors and respond to administrative appeals or litigation.

Appeal procedures must account for automated assistance. A requester challenging an inadequate search needs to know whether the agency used machine learning. Courts examining the adequacy of that search need evidence about how the system was configured and supervised.

Procurement adds further uncertainty. DHS posted a notice of intent to award Deloitte a contract for customized FOIA software with AI features. A notice of intent signals the department’s planned purchasing path, not proof of a completed deployment or measured performance.

The planned software follows a federal FOIA technology showcase that included nearly 40 vendors. According to the event summary, almost 85 percent highlighted AI features. These included audio masking, redaction, relevance scanning, field population, and public-facing chatbots.

Vendor interest shows that DHS will have options. It does not show that those tools meet a common accuracy, security, or explainability standard. Product demonstrations rarely reproduce the document quality and legal ambiguity found in contested requests.

Contract terms can turn safeguards into enforceable requirements. DHS can require testing against representative records, disclosure of model changes, incident reporting, audit-log retention, accessibility, and limits on secondary use of government data.

The department can also require independent evaluation before broader deployment. Testing should include civil-liberties specialists, experienced FOIA professionals, records officers, security personnel, and representatives of frequent requesters.

Without those details, the public must infer safeguards from short inventory descriptions and procurement notices. That is not enough for systems intended to support a transparency law.

AI FOIA Redactions Need Measurable Human Oversight

Human review becomes credible only when DHS measures whether people catch the system’s mistakes.

Agencies often answer AI concerns by saying a person remains in the loop. The phrase sounds reassuring, but it describes many different operating models. A reviewer might examine every document, check a small sample, or intervene only after the system flags uncertainty.

DHS should define the review model for each function. Administrative automation can operate with routine quality checks. Responsiveness classification needs stronger validation because false negatives can hide records. Redaction suggestions require legal review before release.

Sampling can support quality assurance, but the sample must reflect risk. Random checks may miss rare documents containing important responsive material. DHS should oversample low-confidence predictions, unusual formats, sensitive topics, and cases involving broad public interest.

The agency also needs a baseline. It cannot show improvement without comparing assisted review against established human procedures. Measures should include processing time, reviewer workload, missed records, reversed redactions, appeal outcomes, and litigation findings.

Speed alone would provide a misleading result. A system that closes cases faster by excluding more documents might appear efficient while reducing disclosure. DHS needs paired measures showing both timeliness and completeness.

Redaction accuracy requires more than counting black boxes on a page. The department should examine whether each withholding has a valid exemption, whether non-exempt information was reasonably separated, and whether the agency applied the foreseeable-harm standard.

Reviewer behavior should be measured too. Automation bias occurs when people give excessive weight to a system’s suggestion. It can increase when reviewers face high workloads or assume that software has already completed the difficult analysis.

Interface design can reduce that risk. A system should show supporting text and confidence information without presenting its answer as settled. Reviewers should record why they accepted or rejected sensitive recommendations.

Training is equally important. Employees need to understand what the model evaluates, where it performs poorly, and how changes in source material affect results. Legal training alone does not explain machine-learning failure modes.

DHS invited roughly 500 employees from its FOIA processing centers to a virtual training program in 2024. Sessions included discussion of AI and machine learning in FOIA processing. That is a useful starting point, but deployment requires continuing, tool-specific preparation.

Independent oversight can test whether the controls survive operational pressure. Inspectors general, the National Archives’ FOIA ombudsman, congressional committees, and courts all have different roles in evaluating agency transparency.

Requesters also provide evidence. Patterns of missing records, inconsistent redactions, or unusually narrow searches can reveal problems that aggregate performance data misses. DHS needs a process for connecting those complaints to system evaluation.

The department should disclose material incidents. If a model omits responsive records or exposes protected information, DHS should document the scope, correction, and preventive changes. Quietly fixing one request would leave the underlying risk intact.

Public reporting does not require releasing sensitive records or proprietary code. DHS can publish test methods, aggregate results, review requirements, and error categories. Those disclosures would let outsiders evaluate the governance without exposing operational secrets.

The broader lesson applies beyond FOIA. An agency cannot compensate for fewer skilled employees simply by placing an approval step after automation. Meaningful oversight requires people with time, authority, training, and evidence.

Three Signals Will Show Whether the Plan Improves Access

The next test is whether DHS publishes operational evidence, not whether it announces additional AI tools.

The first signal is a more detailed AI inventory. DHS should identify which components use each system, which processing stages receive automation, and which decisions require human confirmation. It should also explain the basis for its impact classifications.

Clearer disclosures would strengthen the department’s efficiency argument. They would show that DHS has separated low-risk clerical tasks from decisions affecting access to records. Another general description of “knowledge retrieval” would weaken confidence.

The second signal is the structure of the Deloitte procurement and any resulting task order. The procurement notice should lead to requirements covering accuracy tests, audit trails, human review, model changes, data protection, and incident reporting.

A contract containing measurable controls would turn public promises into vendor obligations. A contract focused mainly on throughput and deployment speed would shift risk toward FOIA employees and requesters.

The third signal is performance over the coming reporting cycle. DHS should disclose whether the pending inventory falls from 245,572 while processing quality remains stable. Appeal reversals, litigation findings, and substantiated search complaints can provide important counterweights to closure totals.

A smaller backlog would support DHS’s case only if automation improves access. Faster responses with broader redactions or missing records would represent administrative progress without stronger transparency.

The public should also watch staffing. The FOIA Ombudsman report shows that AI adoption is spreading across agencies. It does not establish that software can replace experienced records professionals.

If specialist departures continue, DHS may struggle to provide the review its safeguards assume. If the department maintains trained teams and uses automation for preparation and prioritization, the tools have a better chance of reducing repetitive work safely.

The most credible outcome would not be a fully automated FOIA office. It would be a system that finds more responsive records, removes routine handling, and gives employees more time for legal analysis. DHS would then publish enough evidence for outsiders to verify that result.

FOIA automation deserves scrutiny precisely because the technology sits inside a transparency process. A hidden mistake can determine which government actions remain hidden. That makes auditability a core feature, not an optional compliance layer.

DHS now has an opportunity to establish a workable model for AI-assisted public-records processing. The department should define human authority, test realistic failure modes, and report results beyond raw speed.

Requesters, journalists, and oversight groups should ask three questions as deployment advances: Which decisions does AI influence, who checks those decisions, and what evidence shows the checks work? The answers will determine whether DHS reduces delay or merely automates withholding.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page