Agencies Cannot Govern the AI Systems They Cannot Find
- Olivia Johnson

- 1 day ago
- 13 min read
Google News surfaced a blunt conflict for public agencies: AI governance starts with finding systems that often entered government without a clear AI label.
Policies can prohibit sensitive data uploads, require human review, or establish responsible-use principles. None of those controls work if officials cannot identify the tools, embedded features, vendor services, and employee experiments already influencing public work.
That visibility gap has become more urgent as agencies move beyond isolated chatbot pilots. AI now appears inside permitting platforms, translation services, document systems, fraud detection tools, call centers, and ordinary office software. The contest is no longer innovation versus caution. It is documented, continuously monitored adoption versus fragmented use that central leadership discovers too late.
What the Google News Headline Gets Right About AI Inventories
An AI inventory is becoming the operational starting point for public-sector oversight, not another administrative spreadsheet.
The headline distributed through Google News captures a basic governance problem. Before an agency can classify risk, test accuracy, or assign accountability, it must establish which systems exist and how they are used.
That sounds straightforward until officials define what counts as an AI system. A department might disclose a generative chatbot while overlooking automated ranking, document classification, forecasting, transcription, or image recognition inside existing software.
An inventory must therefore cover more than products purchased under an AI contract. It should include internally developed models, vendor-hosted services, optional software features, employee-selected tools, and automated systems inherited through older agreements.
Public agencies also need to document uses, not only products. The same model can help draft an internal memo or recommend whether someone receives a government benefit. Those applications present very different consequences.
A useful inventory record identifies the responsible department, business purpose, affected population, data categories, vendor, model or service, decision authority, human-review process, and current deployment status. It also records whether a person can challenge an output.
This distinction matters because an application list answers only a narrow question: what software did the organization approve? Governance requires a wider answer about where automated analysis changes public work or affects residents.
The federal government’s experience shows the scale involved. A 2026 review of the latest consolidated inventory found more than 3,600 reported AI uses across 41 agencies, compared with about 700 uses in 2023. The federal adoption review also found that reported adoption remained concentrated among larger agencies.
Those figures demonstrate rapid growth, but they do not prove that every deployment has been captured. Reporting rules change, agencies interpret definitions differently, and employees can use outside services without creating a formal procurement record.
That leaves two kinds of visibility gaps. The first involves shadow AI, meaning tools used without central approval or review. The second involves embedded AI, meaning automated features added to products that an agency already owns.
Embedded systems can be particularly hard to track. A vendor can activate a summarization feature, change a model provider, or expand an automated recommendation function during a routine software update.
An annual questionnaire will not reliably detect those changes. By the time the next survey arrives, the feature might already have processed confidential records or shaped hundreds of decisions.
The inventory therefore needs an update mechanism. Procurement events, contract renewals, security reviews, software releases, and employee access requests should all trigger a check for new or changed AI capabilities.
Connecticut offers one public model. The state says it maintains an annual inventory of approved systems and evaluates proposed uses through a multidisciplinary governance process. Its state AI framework brings technology, legal, policy, ethics, and security expertise into review.
An annual publication can improve transparency, but the internal record must move faster than the publication schedule. Residents need a stable public view, while administrators need a live operating record.
That combination turns inventory work into infrastructure. The public document supports accountability. The internal system supports daily decisions about access, testing, procurement, incidents, and retirement.
Google News readers might encounter the inventory argument as a simple call for better recordkeeping. Its deeper implication is more demanding. Agencies must treat AI discovery as a continuing control that follows each system from initial proposal through its final shutdown.
The Visibility Gap Pressures Procurement, Legal, and Program Leaders
Invisible AI shifts risk away from the person selecting a tool and toward officials who remain accountable for its consequences.
Chief information officers cannot solve this alone. They usually manage approved infrastructure, security requirements, and enterprise contracts. They do not necessarily see every web service an employee opens or every automated feature a vendor activates.
Procurement teams face a related limitation. A solicitation might describe analytics, workflow automation, or decision support without using the term artificial intelligence. Standard technology categories can conceal the mechanism that deserves review.
Legal teams often enter after a contract has been drafted. At that point, the agency might lack leverage to demand model documentation, audit access, incident notification, data restrictions, or advance notice of material changes.
Program leaders carry another form of pressure. They understand the workflow and the residents affected, but they might not recognize when a vendor’s feature meets the agency’s definition of AI.
Each group therefore sees only part of the system. Governance fails when everyone assumes another office owns the complete picture.
Procurement is the most practical front door because it can require disclosure before deployment. A vendor should explain which functions use AI, what data they process, which third parties participate, and whether the system changes after purchase.
Contracts should also address updates. A disclosure made during bidding loses value if the vendor later replaces the underlying model, introduces agent-like actions, or expands data retention without a new review.
An AI agent is software that can plan or execute a sequence of actions toward a goal. That makes change control especially important because the system might do more than generate a recommendation.
A conventional chatbot returns text for a person to assess. An agent might retrieve records, modify a case, send a notice, or initiate another process. The operational risk depends on permissions, reversibility, and the number of records affected.
Program offices must describe those consequences in plain language. A technical description of a model does not reveal whether an error delays a routine memo or blocks access to housing assistance.
Maryland has tied inventory work to existing technology governance rather than building a disconnected reporting exercise. Its public AI inventory process describes a catalog connected to state review structures and agency use-case reporting.
That integration is significant. Agencies already maintain procurement, cybersecurity, privacy, records, and vendor-management processes. AI discovery works better when those systems exchange information instead of demanding a separate manual submission.
Legal mandates are adding pressure as well. Maryland law establishes inventory and impact-assessment duties for high-risk systems used by state units. The schedule distinguishes newer procurements from systems acquired earlier.
Yet a deadline cannot produce an accurate inventory by itself. Agencies still need a defensible method for discovering technologies that were purchased before the mandate or classified under another label.
Smaller local governments face the hardest version of this problem. They often use the same cloud services as larger jurisdictions but lack dedicated privacy officers, model evaluators, or AI procurement specialists.
A 2026 Pennsylvania survey cited by Governing found that 81 percent of responding government officials had no generative AI policy. Among those respondents, 62 percent agreed that having one would be valuable.
Policy absence is serious, but limited capacity creates an even larger constraint. A small municipality might know that employees use an AI assistant without having the expertise to evaluate data flows or negotiate specialized contract language.
Shared services can help. States, counties, associations, and larger cities can publish common definitions, intake forms, approved-use patterns, and contract clauses. They can also provide review support for smaller jurisdictions.
This does not eliminate local accountability. It gives officials a usable baseline and reduces the chance that every municipality repeats the same vendor investigation.
The forced response is organizational, not merely technical. Procurement must ask better questions, legal teams must secure continuing disclosure, security teams must map access, and program leaders must own the real-world use.
That pressure will persist because AI adoption changes faster than government budgeting and contracting. A system approved for one purpose can acquire new functions long before the original agreement expires.
Agencies that cannot connect those offices will govern through incomplete snapshots. Agencies that create a shared record can move from scattered awareness to accountable decisions.
The Main Conflict Is Documented Adoption Versus Shadow AI
The central contest is between systems an agency can trace and AI use that spreads through convenience, defaults, and incomplete disclosure.
Shadow AI does not always begin with deliberate rule-breaking. It often starts when an employee uses a familiar public tool to summarize a document, improve an email, translate a notice, or analyze a spreadsheet.
The immediate benefit is clear. The tool reduces repetitive work and produces a result within seconds. The governance cost remains hidden until someone asks what information left the agency or how the output influenced a decision.
Blocking every public service can push use further underground. Employees facing workload pressure may turn to personal devices, private accounts, or products that security teams cannot observe.
A practical alternative combines approved tools, clear boundaries, training, and proportionate monitoring. Employees need to know which data may enter a system, which tasks require human verification, and which decisions cannot be delegated.
Embedded AI creates a second route around formal review. An agency can approve a records platform, office suite, or case-management product before that product contains generative features.
The vendor later adds summarization or recommendation capabilities through an ordinary update. Users see a convenient button, while the original security and privacy reviews remain focused on the earlier product.
This is why a product-based inventory is insufficient. Agencies need a feature-level view where a material function changes data use, decision authority, or exposure to error.
Vendor disclosure is necessary but cannot provide the entire answer. Suppliers have incentives to market AI features broadly, describe technical dependencies narrowly, or treat model changes as ordinary service improvements.
Government customers need contractual notice for changes that affect risk. They also need technical and administrative methods for confirming what employees can access.
Single sign-on logs can reveal access to approved services. Browser and network monitoring can identify some unapproved domains. Expense records, help-desk tickets, surveys, and software inventories can provide additional signals.
None of those methods is complete. Technical scanning can identify a service without explaining its purpose. Employee interviews can capture purpose while missing occasional or unauthorized use.
The best discovery process combines evidence. It compares procurement records, identity logs, network observations, vendor disclosures, program interviews, privacy assessments, and security reviews.
That comparison should produce questions, not automatic accusations. A previously unknown AI feature may be low risk, disabled, or irrelevant to public decisions. Discovery initiates classification.
Classification then determines the response. A writing assistant used with public information requires different controls from a system scoring inspections, employment applications, or benefit claims.
High-risk does not simply mean technically advanced. It describes the consequences attached to an error, the sensitivity of the data, the affected population, and the opportunity for human correction.
This approach resembles asset management in cybersecurity. An organization cannot patch, monitor, or retire a device it does not know exists. AI adds another dimension because a known product can change its behavior without changing its name.
The AI risk framework from the National Institute of Standards and Technology organizes work around governing, mapping, measuring, and managing risk. Discovery supports every part of that cycle.
Mapping requires context about intended use and affected people. Measurement requires tests connected to that context. Management requires an owner who can change, suspend, or retire the system.
A static inventory can still create false confidence. Officials might celebrate a complete spreadsheet while systems drift away from the conditions originally reviewed.
Model drift means performance or behavior changes as data, environments, or system components evolve. Vendor updates can produce similar changes even when an agency never retrains a model.
Continuous inventory does not mean constant manual inspection. It means defining events that require reassessment, such as a new data source, model replacement, expanded user group, changed purpose, incident, or public complaint.
This is the article’s core tradeoff. Easy access helps employees experiment and improve services, but decentralized adoption weakens visibility. Tight central control improves oversight, but excessive friction encourages workarounds and slows useful deployments.
The answer is not maximum restriction. It is a review system fast enough that employees prefer the approved route and rigorous enough that consequential uses receive scrutiny.
That makes responsiveness a governance control. If approval takes months for a low-risk summarization tool, the formal process becomes less relevant to how people actually work.
Leaders should set review lanes based on consequence. Low-risk uses can follow standard conditions, while systems affecting rights, eligibility, enforcement, safety, or employment receive deeper assessment.
Documented adoption wins only when the documented route remains usable. Otherwise, the inventory records the organization leaders wish they had, not the one operating in practice.
Why an AI Inventory Can Still Produce False Confidence
Finding a system is necessary, but an inventory does not prove that the system is fair, secure, accurate, or well governed.
A registry can become a compliance artifact. Departments submit entries, central staff publish a list, and everyone assumes the most important work is complete.
The entry may describe an approved purpose without showing actual use. It may identify a vendor but omit subcontractors, model providers, data sources, or external plugins.
It may also rely on self-reporting. Employees and contractors cannot disclose a feature they do not recognize, and vendors may not expose every dependency in a clear format.
Even accurate entries age quickly. A system can expand from drafting internal text to generating resident communications. A pilot can become a routine workflow without a formal decision point.
These limitations do not make inventories useless. They explain why discovery must connect to ownership, assessment, monitoring, and incident response.
Every listed use should have an accountable business owner. That person does not need to understand every model parameter, but must understand the purpose, consequences, and conditions for stopping the system.
Technical ownership is different. Information technology or a vendor can maintain a platform while the program office remains responsible for how the output enters a public decision.
A meaningful record therefore identifies both roles. Otherwise, the program can blame the technology office while the technology office says it only supplied a tool.
Impact assessments provide the next layer. They examine affected groups, foreseeable harms, data quality, testing, appeal procedures, human oversight, and alternatives to automation.
However, an assessment performed before launch is still a prediction. Real performance depends on user behavior, changing data, operational pressure, and the cases that reach the system.
Monitoring must test those assumptions after deployment. Agencies should track error patterns, overrides, complaints, processing outcomes, security incidents, and differences across relevant groups.
Public transparency adds another constraint. Agencies should disclose enough information for residents to understand consequential uses without publishing security-sensitive details or protected personal data.
The Organisation for Economic Co-operation and Development has emphasized that public trust depends partly on whether people believe governments can regulate emerging technologies responsibly. Its public trust findings connect governance quality with more positive views of public-sector AI.
Transparency alone will not earn that trust. A public list loses credibility if residents later learn that an unlisted system influenced enforcement, employment, or access to services.
Accuracy also matters more than inventory size. A jurisdiction with 200 vague entries is not necessarily better governed than one with 40 precise, current records tied to real controls.
Comparisons between agencies can therefore mislead. A rising count might indicate faster adoption, improved reporting, a broader definition, or all three.
The same caution applies to the federal increase beyond 3,600 reported uses. The number shows activity and improved visibility, but it cannot establish completeness or quality across agencies.
Independent research published in 2026 examined multiple federal disclosure regimes and found that no single system fully describes how government constructs or deploys AI. The finding suggests that even mature reporting structures can fragment relevant information.
Another risk is scope inflation. If an inventory captures every basic automated rule, reviewers may spend limited attention on low-consequence tools while missing systems that materially influence people.
Definitions should be broad enough to prevent evasion but paired with risk tiers. Discovery should cast a wide net, while assessment resources follow consequence.
Agencies must also preserve institutional memory. Staff turnover, reorganizations, and expiring contracts can separate a system from the people who understood its original limits.
The inventory should retain decisions, test results, approved conditions, incidents, and change history. A new owner needs more than a product name and launch date.
Knowledge continuity matters for any organization managing complex AI-supported work. Teams can improve it by maintaining a searchable technical knowledge base that connects policies, evaluations, contracts, and operational evidence.
Still, documentation must not become a substitute for judgment. A well-organized record can show that a review occurred without proving that reviewers asked the right questions.
The most skeptical interpretation deserves attention: inventory programs can make visible AI look controlled while unobserved use continues elsewhere.
Agencies should test that possibility. They can compare the central inventory against network activity, employee surveys, procurement data, and vendor feature lists, then investigate meaningful discrepancies.
They should also give workers a safe reporting channel. Employees are more likely to disclose experimentation when the response focuses on risk correction instead of automatic punishment.
Finally, leaders should publish limitations alongside totals. A credible inventory explains its scope, reporting method, update date, exclusions, and uncertainty.
That honesty does not weaken governance. It prevents a provisional record from being mistaken for a complete map.
What Agencies Should Watch After the Google News Cycle Moves On
The next test is whether governments turn inventory announcements into observable controls that remain accurate after products and workflows change.
The first signal is the quality of public inventories. More jurisdictions will publish lists because laws, executive policies, and public expectations encourage disclosure.
Readers should look beyond the number of entries. Strong records identify the use, responsible agency, purpose, status, risk level, and role of human review in language residents can understand.
New York’s inventory guidance offers one example of a state defining common reporting expectations across government entities. Its inventory guidance focuses on identifying systems in use and supporting public accessibility for selected data.
If public records become more specific and easier to compare over time, the documented-adoption model is gaining ground. If lists remain vague or appear only once, the visibility problem persists.
The second signal is procurement language. Agencies should require vendors to identify AI functions, model providers, training or operational data practices, update mechanisms, and material changes.
The Government Accountability Office reported in April 2026 that selected federal agencies were not systematically collecting lessons from AI acquisitions. Its procurement assessment treated shared learning as an important step for improving future purchases.
Watch whether governments create reusable contract clauses and review findings. That would show that discovery is influencing how agencies buy technology, not merely how they describe it afterward.
Also watch whether contracts give agencies notice and options when systems change. A vendor’s ability to replace a model without meaningful review can make yesterday’s assessment obsolete.
The third signal is evidence of continuous reconciliation. Agencies should compare their inventories with identity logs, security tools, expense data, software records, and employee reporting.
A growing inventory is not automatically a failure. It can indicate that discovery is improving. The meaningful question is whether newly found uses receive owners, classifications, and corrective action.
Public incident reports will provide another clue. An agency with no reported AI problems might have excellent controls, but it might also lack channels for detecting and disclosing them.
Leaders should track near misses as well as confirmed harm. A confidential document uploaded to an unapproved assistant deserves attention even if the vendor did not retain it or no resident experienced an adverse result.
The Google News headline will soon give way to another AI policy fight. The operational question should remain: can an agency name the systems influencing its work, explain their purpose, and show who can stop them?
For public employees, that question should shape tool selection before sensitive information enters a prompt. For technology buyers, it should shape contracts before a vendor becomes difficult to replace.
For residents, inventory quality offers a practical way to judge whether responsible-AI promises have become working controls. Ask whether the published list covers actual uses, identifies human accountability, and records meaningful updates.
Finding AI is not the whole governance program. It is the point where slogans meet evidence. When the next Google News story highlights a failed automated decision, the most revealing fact will be whether officials knew the system existed before it caused harm.


