top of page

National Grid AI Disclosure Order Puts Utility Automation Under State Review

2 hours ago
12 min read

National Grid must disclose every artificial intelligence system used by its New York utilities within 60 days, despite AI already reaching routine utility operations. The September 17 order turns the National Grid AI disclosure into more than a technology inventory. It asks whether utilities can explain systems that influence customer service, maintenance, inspections, and potentially other regulated work.

The New York Public Service Commission opened Case 26-M-0552 to examine AI across public utilities and other regulated entities. Its inquiry covers National Grid subsidiaries and other electric, gas, and water providers. Companies must identify their AI uses alongside the policies, procedures, and protocols that govern them.

The conflict is straightforward. Utilities want automation that can detect faults, prioritize repairs, and process customer requests faster. Regulators need evidence that those systems remain secure, reviewable, and accountable when their outputs affect essential services. The commission is starting with disclosure because neither side can evaluate that tradeoff without a credible inventory.

New York Wants an Inventory, Not an AI Promise

The order changes AI oversight from voluntary governance into a formal regulatory examination.

The commission’s AI utilization case is titled broadly. It examines artificial intelligence in the operations of public utilities and other regulated entities. That scope matters because AI can enter a utility through many routes, not just a centrally approved model.

A utility might build a machine-learning system internally. It might buy software that quietly adds an AI feature during an update. A contractor might use automated analysis to inspect equipment. Customer-service vendors might summarize calls or recommend responses through models that the regulated utility does not directly operate.

New York’s approach therefore begins with “all uses,” according to the 60-day requirement. The utilities must also describe the controls surrounding those uses. A list of product names would not explain who approved each system, what data it handles, or when a person reviews its output.

The commission identified customer service, predictive maintenance, and facility safety inspections as current utility applications. These are materially different risk categories.

A customer-service assistant can provide an incorrect answer about billing or payment assistance. A predictive maintenance model can misclassify equipment risk and influence work priorities. An inspection system can miss visible damage or generate a false alert that sends crews toward the wrong asset.

The consequences also depend on how each output is used. A model that merely highlights an image for human review presents a different risk from software that automatically changes a maintenance schedule. A chatbot that drafts a reply differs from one permitted to initiate a customer-account action.

That distinction is why the inquiry should not be read as a ban on utility AI. The commission has not announced that AI systems are inherently unacceptable. It is asking where they operate, what authority they have, and which safeguards accompany them.

The timing gives utilities a demanding assignment. Sixty days is long enough to collect known deployments, but short enough to expose weak internal tracking. Large organizations often distribute software purchasing, experimentation, and vendor management across many business units.

National Grid is especially important because its name spans several regulated New York entities. The state’s review can therefore test whether AI governance works consistently across corporate structures, service territories, and operating functions.

Other utilities face the same basic challenge. Con Edison, Central Hudson, National Fuel Gas, New York State Electric and Gas, Rochester Gas and Electric, Orange and Rockland, and regulated water providers all depend on complex technology supply chains.

The first outcome will not be a scorecard proving which company uses AI best. It will be a baseline showing whether each utility can locate and describe the systems operating under its responsibility.

Why the National Grid AI Disclosure Reaches Beyond Chatbots

The hardest systems to disclose will be embedded models that employees no longer recognize as separate AI products.

The phrase “artificial intelligence” can describe generative assistants, forecasting models, computer vision, anomaly detection, and optimization software. These technologies do not share one architecture or one failure mode. A useful inventory must classify them by function and consequence.

National Grid’s public projects illustrate that range, although projects announced elsewhere do not establish what its New York subsidiaries currently use. In Britain, the company has described autonomous drone inspections that capture images of transmission infrastructure. AI-assisted analysis can then help assess equipment condition.

One National Grid project uses drone imagery and machine learning to identify corrosion on steel transmission towers. Its inspection program also uses historical data to forecast how corrosion might develop. The company estimates that the process can reduce manual image review and improve maintenance planning.

That is a real operational scenario, but it also demonstrates why disclosure needs context. Regulators would need to know whether a model merely ranks images or determines a condition grade. They would also need its error thresholds, review process, data lineage, and escalation rules.

Data lineage describes where information came from and how it changed before reaching a model. For an inspection system, that could include camera specifications, image timestamps, asset identifiers, prior maintenance records, and human labels used during training.

A model trained on one tower design, climate, or imaging process might perform differently elsewhere. An apparently accurate average can also conceal weaker performance on rare defects. Those defects might matter most because they carry greater safety consequences.

Customer service creates a different set of questions. Generative AI can summarize calls, draft agent responses, or retrieve policy information. It can also produce a confident answer that does not match the customer’s tariff, account status, or regulatory protections.

In that setting, the disclosure should identify the source material available to the model. It should explain whether responses are grounded in approved documents, whether customer data enters an external service, and whether employees can see uncertainty signals.

A reliable internal record matters because policies, prompts, testing results, and incident reports otherwise become scattered across departments. An organized AI knowledge base can support that work, provided it does not replace formal compliance systems.

Predictive maintenance introduces another mechanism. Models can use sensor readings and historical failures to estimate which assets need attention. Their value comes from directing limited crews and maintenance budgets toward higher-risk equipment.

However, a recommendation can become operationally influential before anyone formally delegates authority to it. Employees may learn to trust a consistently useful ranking. Over time, a nominally advisory model can shape decisions almost automatically.

That is known as automation bias, the tendency to favor an automated recommendation even when contrary evidence exists. The risk does not require reckless employees. It can arise when teams face time pressure, complex data, and performance targets.

An adequate National Grid AI disclosure should therefore separate technical permission from practical influence. A model might lack authority to approve work, yet still determine which options employees see first. Interface design can make an “optional” recommendation difficult to challenge.

Vendor software adds another complication. Utilities might know the name of an enterprise platform without knowing every model used inside it. Updates can introduce automated features, while subcontractors can rely on their own AI tools.

The state’s order places accountability back on the regulated company. A utility cannot fully assess customer, safety, or cybersecurity risks if its inventory stops at systems built by its own data scientists.

Utility Innovation Now Has to Compete With Accountable Control

New York is forcing utilities to reconcile AI’s operational benefits with the evidence required for public accountability.

The case for utility AI is not difficult to understand. Electric, gas, and water systems generate large amounts of equipment, weather, customer, and operating data. Models can help identify patterns that people would struggle to review at the same speed.

Inspection tools can reduce hazardous climbing and focus specialists on uncertain cases. Forecasting systems can support staffing and equipment planning. Customer-service tools can help agents locate complicated information during a call.

Utilities also face growing demand and more complicated grid conditions. A National Grid Partners survey released alongside the inquiry said 78 percent of surveyed utility innovation leaders were deploying or operationalizing at least one AI application for interconnection demand. Because the company sponsored that research, the figure should be treated as an industry signal rather than an independent census.

These potential benefits explain why a blanket prohibition would create its own risks. Refusing useful detection tools can leave a utility dependent on slower processes. The relevant question is whether a particular system improves decisions under controls appropriate to its role.

New York’s commission has already shown that it views digital systems through a critical-infrastructure lens. In April, it adopted enforceable utility cybersecurity rules for regulated electric, gas, steam, and water providers.

Those rules require risk-based cybersecurity programs, access controls, authentication practices, intrusion detection, and response planning. They took effect on June 1, 2026. The AI inquiry arrives less than four months later.

The two actions are connected, but they are not interchangeable. Conventional cybersecurity asks whether an attacker can compromise a system or its data. AI governance must also ask whether an authorized system behaves inaccurately, opaquely, or inconsistently during normal use.

A model can fail without being hacked. Its training data can omit important situations. Its output can change after a vendor update. Employees can apply it outside its tested purpose. A correct statistical prediction can still create an unfair result when embedded in a flawed process.

The commission specifically highlighted hallucinations, algorithmic bias, transparency problems, privacy risks, configuration errors, cyberattacks, and functional brittleness. Functional brittleness means a system performs acceptably in familiar conditions but fails sharply when circumstances change.

These concerns align with the federal AI risk framework. NIST organizes AI risk management around governance, mapping, measurement, and management. It treats validity, security, transparency, privacy, and fairness as related characteristics rather than isolated checkboxes.

NIST’s framework remains voluntary, while New York’s inquiry comes from a sector regulator. That difference is significant. A utility filing can connect general risk principles to named systems, accountable executives, operating procedures, and public-service obligations.

Critical infrastructure also changes the acceptable burden of proof. A photo-filter recommendation and an inspection model might use related techniques, but an incorrect result has different consequences. Regulators reasonably expect stronger validation when a system influences safety, reliability, billing, or access to essential service.

The federal government has made a similar distinction. The Department of Homeland Security’s critical infrastructure framework calls for cybersecurity controls, customer-data protection, evaluation of failure modes, and meaningful transparency.

New York’s investigation can move those principles into utility-specific oversight. It can ask who signs off on a model, how often performance is tested, and what happens when the system produces an unacceptable output.

The primary tension is therefore not National Grid against another utility. It is utility innovation against accountable control.

Utilities gain value when models operate across more data and more decisions. Regulators gain confidence when authority is limited, records remain available, and failures trigger corrective action. Those objectives can coexist, but only when a company can show how.

This disclosure process will pressure governance teams as much as engineering teams. Lawyers must interpret the order. Procurement teams must identify vendor features. Security specialists must map data flows. Business owners must explain actual use rather than intended use.

A strong response will connect those perspectives. A weak one will present a polished AI policy alongside an incomplete operational inventory.

Disclosure Alone Cannot Prove an AI System Is Safe

An inventory creates visibility, but it does not validate model performance or prevent harmful deployment.

The commission’s order is a necessary first step because unknown systems cannot be supervised. Yet a complete list can still conceal weak controls. The state will need to distinguish documented governance from tested governance.

A policy might require human review for high-impact outputs. That statement does not reveal whether reviewers have enough time, expertise, or information to disagree with the model. It also does not show whether disagreement affects performance evaluations.

Similarly, a vendor might advertise accuracy without testing the system on a utility’s data and operating conditions. Average benchmark results do not establish reliability during storms, equipment degradation, unusual account disputes, or incomplete sensor coverage.

The commission should therefore examine each system’s intended purpose and its reasonably foreseeable misuse. It should ask which outcomes were tested, which groups or operating conditions were underrepresented, and who owns remediation.

Generative systems require special scrutiny because fluent language can hide factual errors. NIST calls these errors confabulations, meaning confidently presented false or erroneous content. In utility customer service, polished wording can make an incorrect policy answer more persuasive.

Computer-vision systems present a different uncertainty. Their failures may depend on lighting, camera angle, weather, equipment type, or image compression. A human reviewer needs access to the original evidence, not just a model-generated label.

Predictive systems can degrade as equipment, climate patterns, work practices, or sensor networks change. This is model drift, a decline in performance when real-world conditions move away from the data used during development.

The order also raises a confidentiality problem. Detailed system diagrams and vulnerability information can create security risks if disclosed indiscriminately. Vendors may also claim that model documentation contains proprietary information.

The commission will need enough detail to evaluate risk without creating a road map for attackers. That likely requires a careful division between public filings and protected submissions. Excessive secrecy, however, would prevent customers and independent experts from assessing consequential uses.

Public transparency is most important where AI directly affects people. Customers should know when an automated system shapes billing support, complaint routing, service eligibility, or safety communications. They also need a practical way to reach a qualified person.

The inquiry’s breadth creates another uncertainty. “All AI” sounds clear until utilities apply it to decades of analytics software. Traditional statistical models, rules-based automation, machine learning, and generative AI can overlap.

An overly narrow definition would omit influential systems. An unlimited definition could bury regulators in low-risk tools, including spam filters or routine office features. Risk-based classification offers a better path after the initial inventory.

The state can group systems by consequences and autonomy. A low-impact drafting assistant should not receive the same scrutiny as software connected to field operations. A safety system should face stronger evidence requirements than a tool that formats an internal presentation.

National Grid and other utilities might also identify experimental systems that never entered production. Those pilots remain relevant because their data and access arrangements can create risks. However, regulators should distinguish controlled tests from active decision systems.

The most important skeptical point is that disclosure measures organizational awareness, not safety itself. A utility can know exactly what it uses and still apply inadequate tests. Conversely, an incomplete first filing might expose inventory weaknesses without proving that deployed systems have caused harm.

New York should avoid converting the number of disclosed systems into a simplistic ranking. More reported systems might reflect broader adoption, better governance, or both. Fewer systems might indicate restraint, poor discovery, or a narrow interpretation of the order.

The useful comparison will be control quality. Regulators should look for named owners, clear purposes, approved data sources, test results, monitoring schedules, vendor obligations, human override procedures, and incident records.

They should also examine retirement plans. AI governance is incomplete when companies can approve a system but lack a process for disabling it after unacceptable performance.

The inquiry has not yet established that National Grid or another named utility misused AI. It has established that the commission considers current adoption broad enough to warrant structured oversight. Reporting should preserve that distinction as filings emerge.

Three Signals Will Show Whether New York’s Inquiry Has Teeth

The next stage will reveal whether the National Grid AI disclosure becomes an enforceable oversight model or remains a one-time inventory exercise.

The first signal is the quality of the utility filings after the 60-day window. The strongest submissions will classify systems by purpose, data, autonomy, and consequence. They will identify both internally developed models and AI embedded in vendor products.

Watch for whether utilities describe actual workflows. A filing that says AI “supports maintenance” provides little value. A useful disclosure explains what input enters the model, what output it produces, who reviews that output, and which action can follow.

National Grid’s filing will be especially informative because the company has publicly promoted AI-related infrastructure technologies. Regulators and customers will be able to compare that innovation narrative with the governance practices disclosed for its New York operations.

A detailed filing would strengthen the view that major utilities can trace AI across complicated organizations. Broad categories and repeated confidentiality claims would weaken it.

The second signal is the commission’s treatment of risk tiers. The regulator must decide whether every system receives similar review or whether scrutiny increases with potential harm.

A credible risk model would place customer rights, public safety, grid reliability, and sensitive data near the top. It would also account for autonomy. Software that recommends an action presents less direct risk than software permitted to execute it without approval.

Look for requirements involving independent validation, incident reporting, change management, and recurring performance tests. These measures would show that New York intends to supervise AI through its full lifecycle.

Lifecycle oversight matters because models and vendors change. An inventory filed in November 2026 can become stale after a software update, acquisition, new data source, or expanded business use. The commission will need rules for reporting material changes.

The third signal is whether the inquiry produces enforceable safeguards. The commission has said it will evaluate the disclosures and determine whether additional protections and policies are needed. That leaves several possible outcomes.

It could issue general guidance, establish recurring reporting, add AI controls to cybersecurity examinations, or propose binding rules. It could also set specific expectations for customer-facing and safety-related systems.

A move toward recurring reports and risk-based controls would strengthen the article’s central judgment. It would turn disclosure into a continuing accountability mechanism. A report that summarizes industry practices without follow-up obligations would weaken that conclusion.

The state should also clarify remedies. Regulators need options when a system lacks sufficient evidence, exposes customer data, or performs outside approved boundaries. Those options might include corrective plans, additional testing, restricted use, or suspension.

Utilities, meanwhile, should not wait for the commission’s final policy. The disclosure deadline creates an immediate reason to establish one accountable inventory and reconcile it with procurement, cybersecurity, privacy, operations, and customer-service records.

For technology teams, the practical lesson extends beyond New York utilities. AI adoption becomes harder to defend when an organization cannot say where models operate or who accepts their risks. Documentation must follow deployment, including features supplied by third parties.

For enterprise buyers, the inquiry highlights questions that belong in every AI purchase. What data leaves the organization? Can the vendor explain model changes? Does the customer receive incident notices? Can outputs be audited, challenged, and reproduced?

For customers, the key issue is recourse. Automation should not make it harder to correct a bill, report a dangerous condition, or contest a decision. Human review has value only when it is accessible and empowered.

The National Grid AI disclosure order begins with a simple demand: show the regulator where AI is being used. Its lasting importance depends on what New York does with the answer.

Over the next three months, watch the completeness of utility inventories, the commission’s risk classifications, and any proposal for recurring controls. Then ask a harder question of every organization deploying AI in consequential work: if a regulator demanded a complete, defensible map within 60 days, could the organization produce one?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page