EPA AI Impact Assessments Stay Secret as Transparency Groups Appeal
The EPA withheld final assessments for its high-impact AI tools, despite federal rules requiring agencies to document risks before deploying such systems. That refusal has now produced an administrative appeal and a direct test of federal AI transparency.
Public Employees for Environmental Responsibility, known as PEER, and the Free Information Group, or FIG, announced the appeal on September 15, 2026. Their March Freedom of Information Act request sought records about AI in chemical assessments, including screening tools associated with seven chemicals.
EPA identified its public AI inventory and final impact assessments as responsive records. However, it withheld the assessments under the deliberative process privilege and reported finding no responsive records for 13 other categories in the request.
That response creates a conflict larger than one records dispute. Federal policy promotes rapid AI adoption while requiring added controls for systems affecting health, safety, and rights. The EPA AI impact assessments should reveal whether those controls work before algorithmic outputs enter consequential agency decisions.
The unresolved question is not whether EPA uses AI. Its public plans and inventories establish that it does. The question is whether people outside the agency can evaluate the evidence behind its assurances.
The EPA AI Impact Assessments Appeal Targets Final Records
The appeal challenges both EPA’s decision to withhold the assessments and the apparent limits of its records search.
PEER and FIG filed their original request on March 5, 2026. According to their appeal announcement, the request covered EPA’s use of AI in chemical assessments and its screening work involving seven chemicals.
The organizations sought more than a general list of AI projects. They requested documents about tool selection, testing, performance, safeguards, and the role that AI might play in chemical evaluations.
EPA said it located only two categories of responsive material. One was its 2025 AI Use Case Inventory, which the agency already publishes. The other was a set of final impact assessments for high-impact EPA AI tools.
Those assessments were not released. EPA invoked the deliberative process privilege, a protection commonly associated with internal communications created before an agency reaches a decision.
The dispute turns partly on the word “final.” A final assessment can contain analysis, recommendations, or other material connected to internal deliberations. Its title alone does not settle whether every part qualifies for withholding.
However, the title gives the requesters a clear basis for scrutiny. If the assessments record completed testing, adopted controls, final risk determinations, or factual findings, the groups argue that EPA should not treat them as entirely secret deliberations.
FOIA disputes often require agencies to separate exempt passages from reasonably segregable, nonexempt information. That means the central issue is not necessarily full disclosure or total secrecy. EPA may face questions about whether it could release factual sections, summaries, conclusions, or redacted versions.
The search itself is another point of conflict. EPA reportedly found no records responsive to 13 other categories covering broader AI governance and chemical-screening practices.
That result does not prove that records are missing or that EPA conducted an inadequate search. Agencies can discuss programs publicly without creating every document a requester expects.
Still, the gap deserves examination. EPA has published an AI strategy, maintains governance bodies, tracks use cases, and has discussed expanding AI across agency work. A records search returning so little detail creates tension between the agency’s public program and its documentary trail.
PEER Executive Director Tim Whitehouse framed the issue around public accountability. He argued that taxpayers fund these tools and that their outputs can affect health, safety, and rights.
FIG partner Kevin Bell described the tools as policy expressed through code. His concern is that an automated recommendation can shape an agency outcome even when a human formally signs the final decision.
The appeal does not establish that EPA’s systems are unsafe, biased, or controlling chemical decisions. It asks the agency to reveal enough evidence for outsiders to evaluate those possibilities.
That distinction matters. The strongest case for disclosure does not depend on proving algorithmic harm in advance. It rests on the difficulty of identifying errors when the public cannot inspect the assessment, safeguards, or testing process.
Federal Rules Require More Than an AI Inventory
A project list shows that AI exists, while an impact assessment should show whether a consequential system is fit for its assigned role.
EPA’s public AI use inventory defines a use case broadly. It can include systems designed, developed, procured, or used to support agency missions, decisions, services, or public benefits.
The agency’s 2025 inventory was initially released on February 10, 2026, and later updated. EPA also publishes a consolidated report for common forms of AI use.
Inventories provide useful baseline information. They can identify a project’s purpose, development stage, technical category, operating office, vendor involvement, and impact classification.
They do not necessarily expose the evidence used to approve deployment. A short description cannot show whether test data represented real operating conditions, whether error rates differed among populations, or whether reviewers challenged the project team’s assumptions.
OMB Memorandum M-25-21 draws that distinction. It defines high-impact AI as a system whose output serves as a principal basis for decisions or actions with significant effects on rights or safety.
Human involvement does not automatically remove a system from that category. An agency must consider the actual function of the AI output and how heavily a decision process relies upon it.
Under the federal AI policy, agencies must complete an impact assessment before deploying a high-impact AI use case. They must also update that assessment throughout the system’s lifecycle when appropriate.
The required analysis is extensive. It includes the intended purpose, expected benefits, data quality, model capability, potential effects on privacy and civil rights, reassessment procedures, direct costs, and independent review.
Testing must anticipate real-world outcomes rather than rely only on laboratory performance. Agencies also need procedures to detect and address declining performance after deployment.
These requirements make an AI impact assessment more informative than an inventory entry. The assessment should connect a system’s claimed benefit to evidence, limitations, oversight, and an identified person willing to accept residual risk.
That difference explains why the EPA AI impact assessments have become the center of the appeal. EPA’s published inventory confirms activity, but it cannot substitute for the detailed analysis required for high-impact deployments.
OMB’s policy also contains important limits. Not every AI-assisted workflow qualifies as high impact, even when it operates inside a sensitive program.
A search tool that helps scientists locate papers may remain advisory. A model that ranks chemicals for regulatory attention could carry greater weight, depending on how staff use its output.
The same technology can therefore receive different classifications in different workflows. Context, authority, data, and human reliance matter more than the model’s marketing label.
This creates room for reasonable disagreement. EPA might conclude that a tool supports research without serving as the principal basis for a consequential decision. Critics might see the same ranking or recommendation as practically decisive because employees rarely depart from it.
The impact assessment should document that boundary. Without access to its reasoning, outsiders cannot tell whether EPA tested a genuinely advisory system or accepted a narrow definition that understates its influence.
Faster Chemical Review Raises the Cost of Hidden Assumptions
AI can help EPA organize scientific evidence, but speed becomes dangerous when screening choices quietly determine which evidence receives attention.
Chemical assessments require scientists to process large and uneven bodies of information. Studies can differ in species, exposure route, dose, outcome, quality, and relevance to human health.
Machine learning and natural language processing can help classify studies, extract information, identify related records, or prioritize material for review. These functions can reduce repetitive work and help specialists navigate a growing evidence base.
However, screening is not neutral. A model’s ranking can influence which studies reviewers examine first, which records receive closer attention, and which apparent gaps trigger additional research.
A false positive can waste time. A false negative can be more serious because relevant evidence might never reach the expert responsible for evaluating a hazard.
The risk depends on workflow design. A model used as a second search channel creates a different exposure than one used to exclude documents from further review.
This is why broad labels such as “AI-assisted” reveal little. Readers need to know whether the system retrieves, summarizes, ranks, excludes, predicts, or recommends.
They also need to know what data trained or evaluated the tool. Scientific literature includes publication bias, inconsistent terminology, incomplete abstracts, and decades of changing research methods.
Chemical names create additional complexity. One substance can have several names, while closely related compounds can have materially different properties.
EPA’s assessment would ideally explain how the system handles these conditions. It should also describe validation samples, relevant error measures, human review procedures, and the response when performance falls below expectations.
No public evidence cited in the appeal proves that EPA has delegated final chemical risk decisions to an AI system. PEER and FIG say the agency has not explained whether AI tools enter regulatory decision-making or how they are configured.
That verification gap should remain central. It would be inaccurate to describe an experimental screening tool as an automated regulator without evidence showing that role.
It would be equally premature to dismiss the concern because a human remains somewhere in the process. Automation bias, the tendency to accept a system’s recommendation too readily, can influence outcomes even when formal authority stays with an employee.
Staffing and time pressures increase that risk. When reviewers face thousands of records, a prioritized list can become the practical shape of the review.
AI can also improve consistency. A documented model might apply screening criteria more uniformly than teams working under different deadlines.
That potential benefit strengthens the case for measurable evaluation. EPA should be able to show whether the tool finds relevant evidence at an acceptable rate and whether human reviewers can identify its mistakes.
The dispute is therefore not a simple contest between technology and manual review. It concerns transparent automation versus automation whose influence cannot be independently examined.
EPA’s own plans show that the agency sees AI as more than a minor experiment. Its fiscal 2027 budget materials describe AI supporting rulemaking and permitting tasks, including the categorization of tens of thousands of public comments.
The budget requested $202.2 million and 730 full-time-equivalent positions to advance AI capabilities. That figure covers a broader program, not only chemical screening or high-impact tools.
Still, it shows why governance records matter now. Oversight that works for a pilot may not scale automatically across rulemaking, permitting, enforcement, and scientific assessment.
EPA Promises Controls but Keeps the Evidence Internal
EPA’s published governance structure sounds serious, yet the withheld assessments contain the evidence needed to test whether that structure changes operational decisions.
EPA’s AI strategy describes three governance groups. The Executive AI Governance Board sets direction, while an AI subcommittee develops policy and maintains the use-case inventory.
A broader community of practice gives data scientists and program staff a place to exchange feedback and operating lessons. That group is not itself a governing body.
The strategy says EPA incorporates the National Institute of Standards and Technology’s AI Risk Management Framework. It also identifies pilots, monitoring, feedback cycles, training, cybersecurity review, and legal consultation as safeguards.
For high-impact systems, EPA says its AI subcommittee oversees controls intended to address possible negative effects. These are sensible components of federal AI governance.
The strategy does not show how those controls performed for a particular system. It cannot reveal whether an independent reviewer found unresolved gaps, whether testing covered realistic failure modes, or whether leaders accepted risks over technical objections.
That information belongs in project-level records. The impact assessments are therefore where EPA’s general governance promises should meet operational evidence.
Withholding them creates a credibility problem even if the legal exemption ultimately applies. The public is asked to trust a process whose most informative outputs remain unavailable.
EPA could narrow that gap without publishing source code, protected personal data, cybersecurity details, or confidential business information. It could release redacted assessments, evaluation summaries, model cards, failure categories, or plain-language risk decisions.
It could also explain why a specific system qualifies as high impact. If a tool influences health or safety decisions, the agency should identify the decision, the role of its output, and the point where accountable officials intervene.
Conversely, EPA should explain any determination that an apparently sensitive use is not high impact. That would clarify whether the system provides minor administrative support or shapes a consequential outcome.
The federal government’s wider record shows that incomplete disclosure is not unique to EPA. A 2026 Brookings analysis found that agencies documented more than 3,600 AI use cases in 2025.
High-impact systems represented 12.3 percent of the total, or 445 use cases. Yet the analysis found that more than 85 percent of deployed high-impact cases lacked some required public information about risk controls.
Those inventory findings do not prove individual agencies failed to conduct internal assessments. They show that public reporting has not kept pace with the expansion of consequential AI.
The same analysis identified an important policy shift. Earlier guidance covered systems that controlled or significantly influenced outcomes. M-25-21 now focuses on systems whose outputs serve as the principal basis for decisions.
That wording can narrow the covered category. An agency might argue that AI only informs a broader human process, even when it substantially shapes the evidence available to decision-makers.
The risk is an accountability gap. A system can matter enough to alter a workflow but remain below a strict interpretation of “principal basis.”
EPA’s acknowledgment that it holds final assessments indicates that it has classified at least some tools as high impact. However, the available response does not reveal the tools’ names, functions, or deployment status.
This uncertainty should prevent sweeping claims in either direction. The records might document careful testing and restrained deployment. They might also expose gaps, unresolved risks, or unclear responsibility.
Disclosure is valuable precisely because both outcomes remain possible.
The Appeal Tests Secrecy, Not the Accuracy of EPA’s AI
A victory for the requesters would improve visibility, but it would not by itself prove that EPA’s tools are accurate or improperly deployed.
An administrative FOIA appeal asks an agency to reconsider its initial response. The reviewing office can affirm the decision, reverse it, or require a broader search and additional disclosure.
The current challenge remains at that stage. No court has ruled that EPA improperly withheld the assessments, and the appeal announcement presents the requesters’ account of the dispute.
EPA has a legitimate interest in protecting some internal deliberation. Decision-makers need space to test ideas, debate risks, and reject preliminary conclusions without turning every exchange into final policy.
AI assessments can also contain security-sensitive descriptions, personal information, procurement details, or confidential data. A responsible disclosure process must account for those interests.
The harder question is whether those concerns justify withholding final documents in full. Public accountability becomes weak when an agency can disclose that safeguards exist while shielding all evidence showing how those safeguards were applied.
A balanced release could preserve protected passages and expose the decision record. Useful fields would include the tool’s purpose, responsible office, deployment status, evaluation design, known limitations, monitoring schedule, and escalation process.
Independent review deserves particular attention. OMB requires a reviewer who was not involved in development to identify concerns or gaps.
That requirement creates organizational friction by design. Project teams often focus on expected benefits, while an independent reviewer should test assumptions and examine failure modes.
The public does not need every internal comment to evaluate whether this safeguard operated. It needs enough information to know what concerns were raised and how the accountable official resolved them.
There is also a risk that disclosure becomes performative. Agencies can publish polished summaries that omit thresholds, adverse findings, or unresolved disagreements.
Raw technical detail presents the opposite problem. A long document can appear transparent while remaining inaccessible to people affected by the system.
Effective federal AI transparency therefore needs layers. A short summary should explain the decision and its consequences, while technical documentation should support expert review.
Structured inventory fields should connect those materials over time. Stable identifiers would help researchers track whether systems moved from pilot to deployment, changed impact classifications, or received updated assessments.
EPA’s public inventory currently offers a starting point. The requested assessments would add the risk evidence that a project list cannot provide.
Government-wide audits reinforce the need for that connection. The Government Accountability Office found that reported generative AI uses across 11 selected agencies rose from 32 in 2023 to 282 in 2024.
Its federal AI review also described policy compliance, technical resources, budgets, and fast-changing technology as recurring management challenges.
Those findings do not evaluate the disputed EPA tools. They show why relying on internal assurances becomes harder as agencies add systems faster than governance practices mature.
The appeal also raises a practical lesson for companies adopting AI. An inventory, review committee, or policy document does not prove that deployed systems meet stated standards.
Organizations need traceable evidence linking each use case to its data, tests, limitations, monitoring, owner, and risk decision. Otherwise, governance becomes a collection of promises disconnected from operating software.
EPA occupies a particularly sensitive position because scientific assessments can inform rules, exposure limits, permitting choices, and enforcement priorities. Errors may propagate through decisions affecting communities, workers, businesses, and ecosystems.
That does not mean every EPA model requires public source code or unrestricted access to its data. It means scrutiny should rise with the consequence of the output.
Three Signals Will Show Whether Federal AI Oversight Has Teeth
The next test is whether EPA releases meaningful evidence, not whether it publishes another high-level statement supporting responsible AI.
The first signal is EPA’s decision on the administrative appeal. A full reversal would give the requesters access to the assessments, subject to any other applicable redactions.
A partial release may be more likely. Its value will depend on what remains visible after redaction.
A useful response would preserve facts about purpose, evaluation, limitations, independent review, and monitoring. Pages dominated by withholding marks would leave the central accountability question unresolved.
The decision should also address the records search. EPA can explain which offices, systems, custodians, and search terms it examined without exposing protected content.
If the agency locates additional governance or chemical-screening records, that would strengthen the requesters’ argument that the first search was incomplete. If it documents a careful search and still finds nothing, the concern shifts toward whether EPA creates adequate records.
The second signal is the next revision of EPA’s AI inventory. Readers should watch for newly identified high-impact systems, altered deployment stages, stable project identifiers, and links to risk-management summaries.
EPA’s public page says the agency continuously maintains the inventory. Updates can reveal whether the disputed systems remain active and whether chemical assessment projects appear under recognizable descriptions.
Classification changes will matter. A high-impact tool reclassified as lower risk should include a clear rationale tied to its operational role.
The reverse is also important. A pilot that begins shaping regulatory or health-related decisions should trigger stronger testing, monitoring, and disclosure.
The third signal is government-wide enforcement of M-25-21. OMB can request documentation through accountability reviews and annual inventory reporting.
Public summaries of determinations and waivers would show whether federal AI rules carry consequences beyond agency self-reporting. Repeated gaps without corrective action would weaken the policy’s credibility.
Congress, inspectors general, and GAO can also examine whether agencies perform required assessments before deployment. Their work could test records that public inventories cannot fully expose.
For developers, the case illustrates why deployment context matters as much as model design. A classifier that organizes documents can become consequential when its ranking determines which evidence a specialist sees.
For enterprise buyers, it shows the weakness of relying on vendor claims or governance checklists. Buyers need evaluation records tied to their own data, users, and decision process.
Knowledge workers should care because human review does not guarantee independent judgment. Reviewers need time, authority, and accessible evidence to challenge an automated recommendation.
The EPA AI impact assessments dispute is therefore a test of institutional memory as well as public access. Agencies must document why a system was approved, what its reviewers knew, and who accepted the remaining risk.
Without those records, later investigators may struggle to reconstruct how an automated workflow influenced a decision. The same problem appears in private organizations when model changes, staff departures, and scattered documentation erase context.
The appeal will not determine whether all federal AI is trustworthy. It can establish whether the evidence behind a high-impact classification remains visible only to the agency that made it.
Watch the appeal decision, the next EPA inventory, and OMB’s enforcement record. If those signals produce specific evidence, federal AI oversight will become easier to test. If they produce only policy language, the government’s most consequential systems will remain accountable mainly to themselves.



