top of page

Government AI Agent Flooding Is Turning Better Access Into a Capacity Crisis

Sep 12
12 min read

Chris Schmitz has documented 84 potential cases of government AI agent flooding, revealing a conflict hidden inside easier access to public services. AI tools are helping people prepare complaints, appeals, petitions, and benefit claims that once demanded hours of specialized work. Agencies must now process that additional demand, often without additional staff or funding.

The immediate temptation is to classify the new material as automated spam. That explanation fits some submissions, particularly repetitive messages or sprawling documents with little relevant information. However, Schmitz argues that most observed cases involve people pursuing something they have a legitimate right to request.

That distinction changes the policy problem. Governments are not simply defending their systems against bots. They are discovering how much demand was suppressed by confusing instructions, unfamiliar legal language, fragmented evidence, and the psychological cost of confronting bureaucracy.

The resulting tension reaches beyond one study. If an AI assistant makes a valid claim easier to submit, higher application volume represents improved access. Yet the same improvement can overwhelm institutions whose budgets quietly depended on eligible people giving up.

What Government AI Agent Flooding Actually Changed

AI has lowered the cost of asking government to act, while leaving the cost of answering largely unchanged.

Schmitz, Lewis Hammond, and Alan Chan define agentic flooding as an AI-enabled increase in the volume or complexity of requests that strains a government service. Their 84-case study covers 11 jurisdictions and several kinds of public administration.

The cases include benefit applications, regulatory complaints, freedom-of-information requests, judicial filings, planning objections, and parliamentary petitions. These systems differ legally, but many share a vulnerable interface: they accept open-ended text from members of the public.

Current flooding does not usually involve a fully autonomous agent navigating every government portal. Instead, a person gives documents or background information to a large language model, which produces a structured submission. The person then reviews, copies, or files that material.

This relatively simple workflow can remove several barriers at once. The model can summarize an official letter, identify a possible appeal route, organize evidence, and draft language that resembles professional correspondence. A smartphone photo can replace hours spent decoding unfamiliar instructions.

The research separates two forms of pressure. Quantitative flooding means an agency receives more requests. Qualitative flooding means each request becomes longer, more detailed, or more complicated to evaluate.

An office can experience either form without a dramatic surge in claimant numbers. Ten short complaints might become ten extensive submissions containing dozens of allegations. Each allegation can require classification, evidence checks, internal consultation, and a written response.

The researchers describe their dataset as evidence of potential cases, not proof that AI caused every recorded increase. That qualification matters because several forces can change public-service demand. Policy changes, economic conditions, public awareness, institutional failures, and new reporting channels can all raise volume.

The timing remains notable. Many of the observed services had relatively stable submission patterns before widely available generative AI arrived. Their workloads then increased as text-generation tools became easier to use.

According to the reported findings, complaints to the United Kingdom’s Housing Ombudsman rose from roughly 2,600 in 2022 to just over 7,000 in 2025. Complaints submitted to the United States Consumer Financial Protection Bureau also increased sharply over that period.

Those figures cannot isolate AI as the cause. They do show why agencies are paying attention. Even partial AI involvement can become operationally significant when every valid submission creates a legal or procedural duty to respond.

The memorable change is therefore not autonomous software attacking government websites. It is ordinary people gaining inexpensive assistance with work that institutions previously expected them to perform alone.

AI Did Not Create the Underlying Claims

The central reversal is that many supposedly excessive requests reveal unmet demand rather than invented demand.

Administrative processes impose learning, compliance, and psychological costs. A claimant must discover a program, interpret eligibility rules, gather documents, meet deadlines, and communicate in an institution’s preferred format.

Every additional requirement can screen out fraud or incomplete cases. It can also discourage eligible people who lack time, confidence, legal knowledge, language fluency, or professional help.

Public agencies have long treated incomplete take-up as a service-design problem. The OECD’s access research says complex rules, information gaps, and cumbersome applications prevent eligible households from receiving social support.

An AI assistant changes that equation by absorbing part of the application effort. It does not create the legal entitlement. It makes the route toward asserting that entitlement easier to follow.

Schmitz summarized the point directly: “The vast majority of cases we find are people who are entitled to claim for something, claiming for that thing.”

This statement does not establish that every submission is correct. Entitlement still depends on evidence, legal criteria, and individual circumstances. It does challenge the assumption that rising volume primarily reflects abuse.

Consider a tenant who receives an inadequate response to a repair complaint. Before generative AI, that person might need to locate the applicable policy, reconstruct months of messages, and explain the failure coherently. The burden can be especially severe during illness, displacement, or financial stress.

With an assistant, the tenant can organize the timeline and produce a clearer complaint. That filing still requires investigation. However, the agency can no longer rely on exhaustion or confusion to keep it out of the queue.

The same mechanism applies to financial complaints. The CFPB’s complaint reporting documents the scale and composition of consumer submissions, although it does not attribute their growth solely to AI.

A consumer disputing incorrect credit information might possess a valid grievance but struggle to describe it using the right categories. AI can convert records and correspondence into a more structured narrative. It can also help the consumer persist after an initial rejection.

This creates an uncomfortable measurement problem. A larger queue can indicate a deteriorating service, attempted manipulation, improved public awareness, reduced administrative burden, or all four together.

Calling every increase flooding therefore risks smuggling a policy judgment into a workload description. From the agency’s perspective, a surge is a capacity threat. From the claimant’s perspective, the same surge can represent newly practical access to a right.

Government AI agent flooding is important because it exposes that difference in perspective. Public systems often measure success through processing speed and manageable caseloads. Citizens care whether valid needs receive fair decisions.

An institution can preserve excellent processing statistics by making applications difficult. That outcome looks efficient internally while producing exclusion externally.

AI agents make this arrangement less stable. They reveal demand that paperwork once concealed, pushing officials to decide whether administrative friction was a safeguard or an unofficial rationing mechanism.

The Access Win Becomes a Capacity Crisis

Removing friction at the public entrance can move the bottleneck directly onto investigators, caseworkers, and adjudicators.

Most government services were not staffed for every eligible person to submit a complete, persistent claim. Their workload forecasts reflect historical behavior, including abandonment, missed deadlines, and applications that never begin.

AI assistance can invalidate those forecasts quickly. More people can file, while each person can submit more detailed material and pursue more stages of review.

The Housing Ombudsman offers a concrete picture of the operational pressure. Its 2025 business plan described unprecedented casework, contact every 25 seconds, and more than 7,000 investigations delivered during the year.

An investigation cannot be scaled like text generation. Staff must establish jurisdiction, examine evidence, contact relevant parties, apply policy, and document a defensible decision. Sensitive cases also require judgment that cannot safely be reduced to pattern matching.

This asymmetry drives the crisis. A claimant can generate a lengthy appeal in minutes. The receiving authority can spend hours determining which passages matter and whether each assertion is supported.

Longer submissions can also create defensive work. Officials must search for relevant facts inside repeated arguments, fabricated citations, or language that sounds legally precise without identifying a valid issue.

That problem resembles the pressure created by low-quality AI security reports. Generating a plausible vulnerability report is cheap, but deciding whether the issue is real still consumes expert attention. Public agencies face an additional constraint because statutes or service rules can require them to respond.

Yet the security analogy has limits. A fabricated bug report offers little social benefit. An AI-assisted housing complaint or benefit appeal can come from a person whom the system was created to serve.

The core opponent is therefore wider access versus fixed administrative capacity. It is not citizens versus civil servants, and it is not humans versus machines.

Caseworkers can welcome clearer applications while still being overwhelmed by their number. Applicants can make legitimate claims while unintentionally contributing to longer queues. Both sides experience a system designed for an earlier level of participation.

Government AI agent flooding also redistributes expertise. Before generative AI, professional advocates and lawyers helped clients convert lived problems into institutional language. Now a general-purpose model can approximate part of that drafting work.

The approximation is inconsistent. Models can clarify arguments, but they can also invent rules or encourage users to overstate weak points. Agencies must distinguish useful assistance from confident error after the submission arrives.

That evaluation becomes harder when AI-generated language makes weak and strong claims look equally polished. Writing quality no longer offers a reliable signal of legal merit or personal effort.

Governments can respond by adding their own automation. Classifiers can route cases, extract issues, detect duplicates, or summarize documents for human reviewers. Those tools can reduce clerical work, but they introduce questions about accuracy, bias, privacy, and appeal rights.

Automated triage also creates a contest between two probabilistic systems. One model expands a citizen’s grievance into formal language. Another compresses that language into categories used by the agency.

Important context can disappear between those steps. The applicant might never know which claim the second system ignored, while the caseworker may never see the original emphasis.

Capacity building must therefore include more than purchasing software. Agencies need clear submission structures, auditable triage, sufficient human review, and procedures for correcting automation errors.

Without those investments, faster access simply moves waiting time deeper into the institution. The form becomes easier, but the decision takes longer.

More Paperwork Does Not Mean More Truth

AI can reveal valid claims and still degrade the information environment used to decide them.

The strongest case for AI assistance rests on legitimate access. The strongest skeptical case concerns reliability, duplication, and strategic use.

Large language models predict plausible text. They do not independently establish that a legal rule applies, a document is genuine, or an applicant’s recollection is accurate. A polished submission can contain false citations or irrelevant arguments without obvious warning signs.

Users may also ask a model to strengthen their position. The resulting text can transform uncertainty into confidence, treat inference as evidence, or generate several overlapping allegations from one event.

These failures matter because public decisions affect housing, income, immigration status, licensing, debt, and other consequential interests. Reviewers cannot assume that professional language reflects professional verification.

Fraud is another uncertainty. Lower application costs benefit eligible claimants, but they also lower the cost of testing false stories across several programs. Better identity controls can limit impersonation without establishing whether the facts inside a claim are true.

Coordinated political participation raises a different issue. AI can help residents understand a planning proposal and express informed objections. It can also produce thousands of near-identical comments that exaggerate the apparent breadth of public engagement.

A government body must decide whether participation should be measured by people, submissions, distinct arguments, or demonstrated local interest. Generative text makes those categories easier to manipulate.

The research does not establish a universal fraud rate, nor does it prove that AI generated every surge in its dataset. Its value comes from identifying a recurring pattern and a credible mechanism.

Causal uncertainty should limit dramatic predictions. It should not become an excuse for inaction. Agencies can measure document length, duplicate language, processing time, abandonment, decision outcomes, and successful appeals without guessing about every user’s tools.

Those measurements should separate workload from merit. A service might receive twice as many claims while approving a similar share. That pattern would differ from a surge dominated by rejected or duplicate submissions.

Agencies should also distinguish AI assistance from automation at scale. A person using a model to organize one appeal creates different risks from a commercial service filing thousands of claims through an automated workflow.

The first resembles expanded access to advice. The second resembles industrialized intermediation, especially when a company profits from successful claims or controls communication with applicants.

Both activities can be lawful. They should not be governed through one blunt category called AI-generated content.

Disclosure requirements sound attractive but create enforcement problems. Applicants may not know which features in a writing tool involve AI. A disclosure label also says nothing about accuracy, authorization, or entitlement.

Detection software offers even less certainty. AI-text detectors can misclassify human writing and provide no dependable basis for denying a public service. Their use could particularly disadvantage people writing in a second language.

The better questions concern provenance and responsibility. Did the applicant authorize the submission? Can the agency verify the identity? Are supporting documents available? Can the applicant correct errors and understand the decision?

These controls focus on the integrity of the process rather than the style of the prose. They remain useful whether a claim was written by a lawyer, an advocate, a model, or the applicant alone.

Government AI agent flooding should therefore be treated as a service-design problem with abuse cases, not merely an abuse problem with occasional legitimate users.

Adding Friction Would Punish Rightful Claimants

The fastest defenses can restore manageable queues by recreating the barriers that AI helped people overcome.

Agencies have several familiar tools for suppressing demand. They can introduce fees, require in-person visits, narrow submission windows, impose strict word limits, or add difficult identity checks.

Each measure can reduce volume. Each can also exclude people with limited money, mobility, time, language proficiency, or access to documents.

That tradeoff is especially severe for financially valuable and administratively complex services. Those programs give people a strong reason to seek assistance, while their rules make unsupported navigation difficult.

Fees can deter speculative requests, but they also price access according to income. In-person requirements can frustrate automated bulk filing, but they burden disabled applicants, rural residents, caregivers, and hourly workers.

Strict templates can improve processing if they ask for relevant facts clearly. They become exclusionary when a technical mistake automatically defeats an otherwise valid claim.

Rate limits also require careful design. Limiting submissions per verified person can slow mass filing. It can harm people facing several separate problems or advocates authorized to represent many clients.

Digital identity provides another option. Strong authentication can establish who authorized an interaction and support fair rate controls. It also creates privacy, accessibility, and implementation challenges.

The alternative is capacity building. Governments can simplify eligibility rules, prefill known data, provide structured application interfaces, and proactively enroll people when entitlement can be established from existing records.

These measures address the underlying administrative burden. They also reduce the incentive to produce expansive narratives because the system asks for specific evidence in a predictable format.

The OECD has argued that linked administrative data and automatic enrollment can improve benefit take-up. That approach removes work from citizens and receiving offices instead of automating a contest over paperwork.

Public agencies can also publish machine-readable rules and official agent interfaces. An application programming interface, or API, is a structured channel through which approved software can exchange defined information.

A carefully designed interface can limit arbitrary document length, request evidence consistently, and return clear status information. It can make automated assistance more accountable without forcing users back to paper.

However, an API must not become the only doorway. People need accessible human channels, especially when their circumstances do not fit standard fields.

Agencies also need legal certainty around AI-assisted processing. A caseworker should know which decisions require human judgment, what information an automated system can access, and how an applicant can challenge an error.

Procurement rules should require activity logs and evaluation against real cases. Officials need to see why a tool routed, summarized, or flagged a submission before relying on its output.

Independent audits should test whether automation changes rejection rates across language, disability, income, or demographic groups. Efficiency gains cannot justify hidden discrimination.

This redesign costs more initially than adding a fee or closing an email inbox. It also aligns capacity with the promise of public access.

The policy choice is revealing. If officials respond to higher legitimate participation by rebuilding obstacles, then the former simplicity of the queue depended on exclusion.

If they redesign services around actual demand, AI assistance can expose weaknesses that governments needed to address anyway.

Three Signals Will Show Which Path Governments Choose

The next phase will be defined by measurable agency responses, not by another jump in model capability.

The first signal is whether public bodies publish evidence linking AI assistance to specific operational outcomes. Raw submission counts are insufficient.

Useful reporting would compare volume, document length, processing time, duplicate rates, approval rates, successful appeals, and unresolved backlogs. It should separate individual assistance from coordinated or commercial bulk filing.

If agencies release that information, the debate can move beyond anecdotes. High approval rates would strengthen the argument that AI is uncovering legitimate demand. High duplication or rejection rates would support narrower controls against low-value submissions.

The second signal is whether governments choose friction or capacity. New fees, closed channels, broad AI restrictions, or mandatory physical visits would show that institutions are protecting workload by reducing access.

Structured forms, stronger staffing, automatic enrollment, auditable triage, and official agent interfaces would point in the other direction. They would show that agencies accept increased participation as a service obligation.

The third signal is whether processing capacity keeps pace with easier filing. Investigation backlogs and decision times will reveal whether governments have only modernized the citizen-facing entrance.

The Housing Ombudsman’s workload and the CFPB’s complaint volumes offer useful reference points, but future reporting needs clearer causal analysis. Researchers should examine policy changes and economic conditions alongside AI adoption.

That work should also track distribution. Faster average decisions can conceal worse outcomes for cases that do not fit automated categories.

For developers, the lesson is that agent quality cannot be measured only by successful submission. A responsible system should preserve evidence, show sources, request confirmation, and avoid turning one grievance into unsupported claims.

Enterprise buyers should ask who remains accountable when an agent communicates with a public body. Authorization, data retention, audit logs, and correction workflows matter more than fluent drafting.

Knowledge workers should treat generated submissions as consequential records. The convenience of producing a formal letter does not remove the obligation to check names, dates, claims, and cited rules.

Public agencies face the hardest assignment. They must protect limited capacity without treating newly capable citizens as attackers.

Government AI agent flooding will test whether digital access was a genuine commitment or merely a convenient slogan. A manageable queue is not a successful service when eligible people stay away because the process defeats them.

The practical question is now unavoidable: when AI reveals the demand hidden behind administrative burden, will governments fund the service people were promised, or restore the burden that kept them silent?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page