DHS AI FOIA Redactions Promise Speed, but Risk Deeper Secrecy
DHS plans to use AI to recommend redactions across 100,000 to 140,000 annual public-records requests, despite limited evidence that machines can make reliable disclosure judgments. The DHS AI FOIA redactions initiative promises shorter processing times. It also risks making excessive withholding faster, cheaper, and harder for requesters to challenge.
The Department of Homeland Security says employees will review the recommendations before records leave the agency. That safeguard matters because the Freedom of Information Act, or FOIA, requires contextual legal judgment. A model can recognize names or identification numbers. It cannot independently decide whether releasing a policy discussion would cause a legally recognizable harm.
This is not a contest between technology and paper-based bureaucracy. The real conflict is between processing efficiency and meaningful public access. Federal agencies face record request volumes that human reviewers struggle to manage. Yet the government has not shown that automated recommendations will consistently favor disclosure when the law permits it.
That question carries unusual weight at DHS. The department processed nearly one million requests during fiscal year 2025, according to federal data cited in the federal redaction plan. Only 2.3 percent produced records without redactions. If automation influences even a fraction of that workload, its assumptions will shape what journalists, researchers, attorneys, and ordinary citizens can learn about federal activity.
What DHS AI FOIA Redactions Will Actually Do
The immediate change is not autonomous censorship, but automated recommendations entering a legal review process at enormous scale.
U.S. Customs and Border Protection, an agency within DHS, plans to use Google artificial intelligence and machine-learning tools during FOIA processing. The system will identify content that might qualify for withholding and recommend redactions to government employees.
A redaction removes part of a record before its release, usually by covering text, images, audio, or video. Agencies apply redactions when information falls under a FOIA exemption or another law prohibits disclosure. Common examples include personal data, classified material, law-enforcement information, and certain internal deliberations.
The DHS description says automation should reduce processing times. The department has also said staff members will review each case and ensure that requesters receive accurate and complete information. Those statements indicate that a human remains responsible for the final decision.
However, the planned scale changes the practical meaning of that review. DHS expects to apply AI to request categories representing more than 10 percent of its submissions. The department estimates that range at approximately 100,000 to 140,000 requests.
Reviewing a machine suggestion is not necessarily equivalent to making a fresh decision. A busy employee may begin with the model’s highlighted passages and focus on confirming them. That workflow can create automation bias, the tendency to accept a system’s output because it appears systematic or authoritative.
The distinction becomes critical when a model moves beyond obvious personal information. Detecting a Social Security number follows a relatively stable pattern. Applying an exemption to an email about an unfinished policy requires an understanding of the document’s author, purpose, timing, audience, and relationship to a final decision.
DHS has not publicly established how often its system will handle simple pattern matching instead of those contextual decisions. It has not disclosed comprehensive performance results, error rates, testing records, training examples, or override rates.
Other departments are also experimenting. The Department of the Interior reportedly uses Microsoft technology to locate material that might fall under attorney-client protections. The Department of Health and Human Services has described a proof-of-concept system for reviewing and redacting large document collections. The Justice Department says several components already use automated search and review functions.
The movement is therefore broader than one DHS deployment. Federal agencies are building a layer of machine assistance between government records and the people requesting them. What remains unsettled is whether that layer will help reviewers find releasable information or encourage them to remove more of it.
That uncertainty turns an operational upgrade into a test of government transparency.
Why Federal Agencies Are Automating FOIA Now
AI is entering FOIA because request volumes and record collections have outgrown many agencies’ existing review systems.
FOIA gives the public a right to request records from executive-branch agencies. It does not cover Congress, federal courts, or most offices that directly advise the president. Each covered agency must locate responsive records, examine them, apply any lawful exemptions, and release all reasonably segregable nonexempt material.
That process becomes slow when one request covers thousands of emails, attachments, messages, videos, spreadsheets, and scanned documents. Reviewers must find duplicates, identify responsive material, consult other offices, protect legitimate secrets, and preserve everything that can legally be released.
Across the federal government, agencies processed a record 1.6 million requests during fiscal year 2025. That represented a 9 percent increase over the previous year, according to the Justice Department’s annual FOIA data.
Technology can address several mechanical parts of this workload. Optical character recognition turns scanned pages into searchable text. Deduplication removes repeated email chains. Entity recognition locates names, phone numbers, addresses, and other recurring patterns. Classification systems can prioritize documents that appear relevant to a request.
These uses are not equally risky. Searching a collection for responsive records can broaden access if the alternative is an incomplete manual search. Automatically concealing language affects the substance of the public release.
Federal officials have spent years exploring that distinction. In May 2026, the Chief FOIA Officers Council, the National Archives’ Office of Government Information Services, and the Justice Department held the third NexGen FOIA Tech Showcase.
The government’s technology showcase sought tools for electronic discovery, case processing, accessible releases, public reading rooms, and automated redaction. Proposed systems covered paper files, digital content, video, audio, and structured data.
The event followed similar showcases in 2022 and 2024. It also reflected a recommendation from the federal FOIA Advisory Committee, which encouraged agencies to seek industry information about AI-assisted processing.
Adoption remains limited but is increasing. The National Archives’ latest FOIA assessment found that 20.6 percent of responding federal agencies used AI or machine learning to help search or organize records in 2025. That was up from 18.6 percent in the previous assessment.
The same report found that 80 percent used electronic discovery tools for FOIA searches, compared with 72 percent in 2020. Those figures show that agencies were already automating document work before generative AI became the dominant technology story.
The backlog problem is real. Delayed access can make accurate information useless. A record released after an election, court dispute, policy debate, or public emergency may arrive too late to inform accountability.
Automation can therefore serve openness when it helps staff locate documents, remove duplicates, or identify plainly protected personal data. Faster processing is not an empty benefit.
The danger appears when agencies treat speed as the primary measure of success. A closed request is not necessarily a fulfilled request. A fast response filled with unnecessary black boxes can provide less public value than a slower, carefully reviewed disclosure.
That is why the central performance question cannot be how many pages the system processes per hour. It must be how much lawful information reaches the public, with how many errors, after what level of human scrutiny.
The Efficiency Promise Meets a Disclosure Problem
The same system that can accelerate valid redactions can also scale the government’s existing tendency to withhold too much.
DHS processed 994,992 FOIA requests in fiscal year 2025. It released records without redactions in only 2.3 percent of completed cases. About 40 percent resulted in partially redacted releases.
The remaining requests ended for several reasons. Some received full denials under statutory exemptions. In many cases, the department said it found no responsive records. Other requests were withdrawn, referred elsewhere, or closed for procedural reasons.
Those categories do not prove misconduct. DHS handles immigration, border enforcement, cybersecurity, emergency management, transportation security, and protective operations. Its records routinely contain personal data and information connected to law enforcement or national security.
Still, the baseline matters. DHS is not introducing automated recommendations into a system known for broad, unredacted disclosure. It is introducing them into a system where full releases are already rare.
AI can push that system in either direction. A well-designed model might help employees isolate a protected name while releasing the rest of a paragraph. That would support the law’s requirement to separate exempt information from material the public may receive.
A poorly calibrated model might flag an entire paragraph because one sentence resembles protected language. If employees accept the recommendation, the system would conceal nonexempt context along with potentially sensitive text.
Models also make two different kinds of error. A false negative misses information that should be protected, possibly exposing private or classified material. A false positive marks lawful information for removal.
Government incentives do not treat those errors equally. An accidental disclosure can create an immediate security incident, privacy complaint, or disciplinary problem. An unnecessary redaction usually burdens the requester, who must appeal or sue to challenge it.
That imbalance encourages conservative settings. A system optimized to avoid accidental release will probably flag more material. Human reviewers working under time pressure may then approve those suggestions because withholding feels institutionally safer.
The result can be a ratchet toward secrecy. Each model recommendation appears cautious on its own. Across hundreds of thousands of requests, the cumulative effect could remove large amounts of information that a more deliberate review would release.
Automation also makes consistency possible, but consistency is not automatically fairness. A model can apply a narrow interpretation of the law across an entire agency. It can also reproduce inconsistent decisions from the records used to configure or train it.
FOIA decisions often depend on institutional history. If past reviewers regularly withheld a type of discussion, a system learning from their work may present that practice as the normal answer. The model does not know whether earlier decisions reflected careful legal analysis, outdated guidance, or habitual over-redaction.
That creates a feedback loop. Previous withholdings influence automated recommendations. Reviewers accept those recommendations. The approved output then becomes evidence supporting future automated decisions.
The government can interrupt that loop only through deliberate evaluation. Agencies need test collections containing both exempt and releasable material. Independent reviewers should measure false positives, false negatives, and the amount of nonexempt text lost around valid redactions.
Override data also matters. If employees almost never reject model suggestions, oversight may be nominal. If reviewers frequently add or remove redactions, the system may not be saving as much labor as promised.
DHS has said humans will remain involved. It has not yet shown whether that involvement will be independent, adequately staffed, documented, and subject to quality checks. Until those details become public, the efficiency claim remains incomplete.
AI Cannot Decide Foreseeable Harm by Pattern Alone
A model can detect language patterns, but FOIA requires agencies to connect each withholding decision to a specific, reasonably foreseeable harm.
Congress strengthened that requirement through the FOIA Improvement Act of 2016. An agency generally may withhold exempt information only when it reasonably foresees that disclosure would harm an interest protected by the exemption. It may also withhold information when another law prohibits release.
The foreseeable-harm standard prevents agencies from treating every technically available exemption as an automatic command. A document can fit within an exemption’s boundaries while still lacking a meaningful reason for secrecy.
Justice Department disclosure guidance tells agencies to consider context. Relevant factors can include a record’s age, earlier official disclosures, its sensitivity, and its purpose. Agencies should consult knowledgeable staff when the possible harm is unclear.
That work differs fundamentally from entity detection. A model can find a person’s name. It cannot determine from the name alone whether disclosure would invade privacy, identify a confidential source, reveal official conduct, or repeat information already made public.
The same challenge applies to the deliberative process privilege. That protection can cover certain predecisional communications reflecting agency consultation. It does not allow an agency to hide every draft, recommendation, or internal email.
Researchers have tested machine learning on this problem. Jason R. Baron, Mahmoud F. Sayed, and Douglas W. Oard assembled material from Clinton presidential records and trained classifiers to identify potentially deliberative passages.
Their machine-learning study found that systems performed best when training and evaluation followed consistent reviewer interpretations. Performance weakened when reviewers, record custodians, or subject matter changed.
In later public discussion, Baron described the methods as roughly 70 percent accurate for distinguishing material associated with the deliberative process privilege. Published results under consistent conditions reported F1 measures between 70 and 83 percent. An F1 score combines how many relevant passages a model finds with how often its positive classifications are correct.
Those results are useful for triage. They are not strong enough to support unsupervised legal decisions. At 70 percent performance, a consequential share of passages will be classified incorrectly.
The research also exposes a deeper problem. Human reviewers do not always agree about whether the same language qualifies for withholding. A system trained on one reviewer’s interpretation can struggle when another applies the law differently.
That disagreement is not merely technical noise. It reflects the contextual character of FOIA. Legal judgments turn on what the record represents, what the government has already disclosed, and what concrete harm would follow from release.
DHS therefore needs more than a human positioned at the end of an automated pipeline. Meaningful review requires enough time, authority, and information to reject the model’s recommendation. The employee must also examine whether a narrower redaction would protect the legitimate interest.
Agencies should preserve an audit trail for every automated suggestion. That record should identify which passages the system flagged, which exemption it associated with them, what the employee changed, and why the final decision satisfied foreseeable harm.
Requesters do not need access to sensitive training data or system-security details. They do need intelligible information about how automated tools affected their records. Without that visibility, an appeal becomes harder because the requester cannot distinguish a lawyer’s judgment from a model-assisted default.
A black-box recommendation should never become a black box on the page merely because a human clicked approve.
Human Review Is a Safeguard Only When It Changes Outcomes
The phrase “human in the loop” offers little protection unless agencies measure whether people actually question automated recommendations.
DHS says employees will review AI-assisted cases. The federal FOIA ombuds office similarly warns that AI and machine learning cannot replace professional judgment about exemptions and foreseeable harm.
That is the correct formal position. Its effectiveness will depend on workflow design.
If a reviewer sees the complete document before any proposed markings appear, the person can form an independent view. If the system presents a page already covered with suggested redactions, those markings establish an anchor. The reviewer must then justify removing a protection instead of deciding whether to add one.
The difference may seem procedural, but it changes the direction of doubt. In the first workflow, uncertainty can favor disclosure. In the second, uncertainty favors keeping the model’s black box.
Staffing creates another constraint. An agency may adopt automation precisely because it lacks enough people to handle its workload. If the same shortage leaves reviewers only minutes to inspect each result, formal human approval becomes a throughput checkpoint.
Performance targets can amplify that pressure. Managers may measure requests closed, pages processed, or backlog reduction. Those numbers are easy to count. The amount of nonexempt information preserved through careful review is harder to observe.
An effective oversight program would use several measures. Agencies should compare machine-assisted decisions with independent review samples. They should track how often employees remove proposed redactions, narrow them, or add protections the system missed.
They should also test outcomes across document types. A tool that reliably detects passport numbers in standardized forms may perform poorly on policy emails, handwritten notes, spreadsheets, recorded interviews, or historical documents.
Testing must include the consequences of both error types. False negatives can expose protected information. False positives can deny public access. Treating privacy and security errors as the only meaningful failures would build over-redaction into the evaluation process.
Agencies should publish aggregate results. Useful disclosures would include the number of requests processed with AI, the exemptions involved, sampled error rates, reviewer override rates, and changes in processing time.
Public reporting would not reveal the sensitive content of individual requests. It would establish whether the tool behaves as an assistant or an unacknowledged decision-maker.
Appeals offer another source of evidence. If AI-assisted cases produce more successful administrative appeals, that trend would suggest the system or its review process is withholding too broadly. Agencies should compare those outcomes with conventionally reviewed cases.
Courts will eventually play a role. FOIA litigation often requires agencies to explain the basis for redactions through declarations or a Vaughn index, a document connecting withheld material to claimed exemptions. Automated tools do not reduce that legal burden.
If an agency cannot explain why disclosure would cause harm without referring generally to a model, its justification should fail. The government remains accountable for the final decision, regardless of which vendor supplied the software.
Vendor oversight is therefore part of public accountability. Contracts should permit government testing, preserve logs, document model changes, and prevent sensitive records from being reused for unrelated training. Agencies must also understand whether updates alter behavior after a system enters production.
The strongest version of government AI redaction would help employees locate possible sensitivities while preserving their independent legal judgment. The weakest version would turn over-redaction into a repeatable workflow with a human signature at the end.
DHS has described the first version. The evidence needed to distinguish it from the second has not yet been made public.
Three Signals Will Show Whether the System Serves Openness
The next test is not whether DHS deploys AI, but whether the department proves that the technology releases more lawful information without exposing protected material.
The first signal is operational transparency. DHS should publish a clear description of which request categories use automated recommendations, which models perform the work, and which exemptions they support.
That description should separate basic pattern recognition from contextual legal classification. Locating an identification number presents a different risk from evaluating internal policy discussions. Combining them under one “AI” label would prevent meaningful scrutiny.
The department should also disclose aggregate performance measures. False-positive rates will show how often the system recommends hiding releasable material. False-negative rates will indicate how often it overlooks protected information.
Reviewer override rates will reveal whether humans exercise independent judgment. A near-zero rate could mean the tool is exceptionally accurate. More likely, it would raise questions about automation bias, limited review time, or incentives that discourage employees from changing suggestions.
The second signal is the quality of released records. Processing time should fall, but disclosure should not shrink. Researchers can compare the proportion of full releases, partial releases, and complete denials before and after deployment.
Those categories are imperfect. One deleted word and ten pages of black boxes can both count as a partial release. DHS should therefore examine the share of responsive text released within representative document samples.
Administrative appeals will provide another measure. A rise in successful challenges involving machine-assisted redactions would weaken the department’s case. Stable or improved disclosure outcomes, paired with shorter delays, would support its efficiency argument.
The third signal is enforceable accountability. Requesters need to know when automation materially influenced a response. Agency appeal staff and courts need records showing how each recommendation entered the final decision.
Congress, inspectors general, and the National Archives should examine whether DHS applies the foreseeable-harm standard independently after the model flags a passage. They should also evaluate vendor controls, data security, model updates, and protections against feedback loops.
None of these safeguards requires rejecting automation. FOIA offices need better search, deduplication, document conversion, and review tools. Public access does not improve when responsive records sit untouched in a backlog for years.
However, faster redaction is not the same as faster disclosure. The first removes information. The second gives the public timely access to everything the law permits.
That distinction should guide every evaluation of DHS AI FOIA redactions. The department can validate its approach by publishing audit methods, error measures, override data, and before-and-after disclosure outcomes.
Journalists, researchers, attorneys, and other requesters should watch those indicators instead of accepting lower processing times as proof of success. They should also use administrative appeals when an agency offers generic explanations or appears to withhold more than necessary.
AI can help find the boundary between public and protected information. It should not quietly move that boundary toward secrecy. The decisive question is simple: after DHS accelerates the process, will Americans receive more usable information, or merely receive their blacked-out pages sooner?



