top of page

PJM Data Center Reliability Faces a Federal Test After a 3,800 MW Load Drop

2 hours ago
12 min read

PJM data center reliability became a federal concern after 3,800 megawatts of demand vanished from the grid during a normally cleared transmission fault. The July 22 event did not cause a blackout. However, it exposed a conflict between data center protection and regional grid stability.

Data centers in Northern Virginia transferred to backup power after voltage fell during a fault on a 230-kilovolt transmission line. Their systems protected sensitive computing equipment as designed. From PJM’s perspective, however, thousands of megawatts disappeared in two rapid waves.

Regulators now face a difficult tradeoff. Data centers need protection from electrical disturbances, while grid operators need major loads to remain connected through routine faults. Rules designed for ordinary commercial customers do not fully address facilities operating at the scale of power plants.

A Normal Fault Triggered an Abnormal Load Transfer

The July event showed that a routine transmission fault can become a regional balancing problem when many data centers respond together.

The disturbance began in Dominion Energy’s Northern Virginia service territory on July 22, 2026. A fault affected a 230-kilovolt transmission line serving an area with a dense concentration of hyperscale data centers.

The line protection cleared the fault normally. No breaker remained stuck, and the utility did not order customers offline. Yet data center control systems detected the voltage disturbance and transferred substantial demand to onsite backup supplies.

PJM initially lost about 2,970 megawatts of load. A second wave brought the total transfer to approximately 3,800 megawatts, according to the grid operator’s account reported in a reliability review.

PJM’s total load fell from 99,984 megawatts to 96,205 megawatts. That represented a decline of roughly 3.8 percent within a short operating interval.

The grid had already scheduled generators to serve the higher level of demand. When the data centers disconnected, generation temporarily exceeded the remaining load. Frequency and voltage moved as PJM worked to restore the balance.

Area control error measures the mismatch between scheduled and actual power exchanges across a balancing area. The July event pushed that measure sharply positive because PJM suddenly had more generation than its customers required.

Operators dispatched reactive power resources to reduce voltage and adjusted generation. PJM returned its area control error within its required limit after nine minutes. The applicable North American reliability standard allows 30 minutes.

That response matters. The system remained stable, and operators corrected the imbalance well within the required period. The incident therefore should not be described as a blackout or a failure of PJM’s control room.

Dominion also said it did not disconnect the data centers. Their own protection systems initiated the transfer to backup power. That distinction assigns the immediate action to facility controls, not utility-directed load shedding.

The outcome was still serious because the disturbance crossed facility boundaries. Each operator protected its equipment independently, but the combined response created a single grid-scale event.

PJM called it the largest load transfer in its history. Similar disturbances in Northern Virginia had removed about 1,500 megawatts in both 2024 and 2025.

The sequence suggests that July was not an isolated equipment failure. It was a larger recurrence of behavior that grid officials had already observed.

PJM data center reliability is therefore no longer limited to securing enough electricity for future campuses. It also concerns how connected facilities behave during disturbances lasting only fractions of a second.

Why AI Data Centers Behave Differently From Ordinary Loads

AI campuses combine concentrated demand, sensitive electronics, and automated controls that can move faster than conventional grid resources.

Traditional grid planning treats electricity demand as relatively predictable. Consumption rises and falls, but most customers do not disconnect gigawatts of load simultaneously after a brief voltage change.

Hyperscale computing challenges that assumption. One campus can consume hundreds of megawatts, while several campuses may share the same transmission corridor and similar protection settings.

Those facilities contain servers, networking equipment, cooling systems, power converters, batteries, and backup generators. Automated systems monitor incoming electricity and protect equipment against voltage or frequency conditions outside configured limits.

A ride-through requirement defines the voltage and frequency conditions that equipment must tolerate without disconnecting. Generators already face detailed ride-through rules because simultaneous generation losses can destabilize the system.

Large computational loads have not historically faced equivalent national requirements. Grid operators often lack complete information about their internal controls, transfer thresholds, or expected recovery sequence.

That gap becomes important when many facilities use similar equipment. A single voltage sag can trigger a correlated response across several sites, even when each system behaves correctly on its own terms.

The NERC large-load study describes computational demand as unusually fast and difficult to model. Some facilities can reduce consumption within seconds, faster than conventional generators can adjust output.

NERC also documented one North American data center dropping from approximately 450 megawatts to 40 megawatts within 36 seconds. The facility later returned to its previous demand over several minutes.

Such changes affect more than energy supply. They influence frequency regulation, voltage control, operating reserves, and scheduled power exchanges between neighboring regions.

AI workloads add another source of uncertainty. Training runs can create variable or periodic consumption patterns, while cooling systems and power electronics affect reactive power and power quality.

Not every AI data center has the same design. Some facilities use uninterruptible power supplies for critical equipment. Others rely on workload checkpoints that let operators resume a training process after an interruption.

This diversity makes broad assumptions risky. Grid planners need verified models for specific facilities, not a generic profile labeled “data center.”

Operators also need to know what signals can change consumption. Electricity prices, internal equipment protection, computing schedules, backup-system status, and contractual obligations may each produce a different response.

The July transfer demonstrated the most immediate danger. A routine fault caused numerous facilities to interpret the same voltage signal in similar ways.

The result resembles a common-mode failure, where separate systems share a hidden dependency. Here, the shared dependency was the electrical disturbance and the facilities’ sensitivity to it.

PJM officials have argued that the fault should not have caused those facilities to leave the grid. Operating Committee Chair Emanuel Bernabeu said the data centers appeared too sensitive and disconnected too early.

Facility operators have a competing responsibility. They must prevent equipment damage, preserve customer workloads, and meet uptime commitments. Staying connected through a deeper disturbance can increase those risks.

That creates the central tradeoff for PJM data center reliability. A conservative protection setting can help an individual campus while making the larger grid harder to balance.

PJM Data Center Reliability Now Depends on Shared Rules

The regulatory response is shifting large data centers from passive customers toward accountable participants in bulk-system operations.

The Federal Energy Regulatory Commission, or FERC, oversees interstate electricity transmission and wholesale power markets. The North American Electric Reliability Corporation, or NERC, develops enforceable reliability standards under FERC supervision.

FERC directed NERC in July to establish an initial regulatory framework for computational loads. The effort covers large data centers, cryptocurrency facilities, and other electronically controlled demand connected to the bulk power system.

NERC has proposed three foundational standards. CLO-001-1 addresses interconnection studies and modeling. CLO-002-1 covers operational data and communications. CLO-003-1 addresses protection coordination and disturbance monitoring.

These categories target gaps revealed by repeated load-loss events. Planners need accurate facility models before connection. Operators need real-time information after connection. Engineers need coordinated protection settings when disturbances occur.

NERC has also proposed new registered roles for Computational Load Owners and Computational Load Operators. Registration would place qualifying entities inside the formal reliability system rather than treating them like ordinary retail customers.

That change carries practical consequences. Registered entities can face documentation, monitoring, reporting, and compliance duties tied to bulk-system reliability.

NERC accepted comments on its initial standards and registry proposals through September 18, 2026. The organization’s large-load action plan tracks these standards, working groups, and supporting technical guidance.

FERC ordered NERC to submit initial reliability standards and registry criteria by December 31, 2026. NERC must also develop a plan for further standards, including unresolved performance requirements.

The timeline reveals an important limitation. Initial rules can establish modeling, communications, and protection-coordination duties. Detailed minimum ride-through requirements require additional technical work.

PJM does not want to wait for every national rule. The grid operator is evaluating whether to expand its own interconnection reliability requirements for current and future computational loads.

Regional action offers speed. PJM can use its tariff, interconnection process, operating agreements, and transmission-owner coordination to address risks specific to Northern Virginia.

A regional approach also creates complications. Data center operators develop facilities across several power markets. Different technical requirements can increase engineering complexity and make compliance less consistent.

FERC is separately examining how all six regional grid operators integrate large loads. Its June show-cause orders require each operator to justify existing tariffs or propose reforms.

Those orders cover application procedures, transmission studies, cost transparency, co-location, flexible service, and access to adequate generation. They address how facilities connect, not only how they behave afterward.

The two regulatory tracks are connected. Faster interconnection without effective operating standards could add concentrated demand before grid operators understand its response to faults.

Conversely, highly restrictive reliability requirements could delay projects or make valuable flexibility harder to use. Data centers can reduce demand quickly, which can help during shortages when operators coordinate that response.

The question is not whether large loads should ever disconnect. Some disturbances require separation to protect people and equipment. The question is when, how quickly, and with what notice they should do so.

Grid rules must also govern reconnection. If thousands of megawatts return together, the recovery can create a second balancing and voltage challenge.

Reliable integration therefore requires a coordinated sequence. Facilities should tolerate defined disturbances, notify operators when transferring, and return according to an agreed restoration plan.

Protection Rules Must Balance Grid Stability and Equipment Risk

A strict ride-through mandate would reduce sudden load losses, but it could transfer electrical and commercial risk back to data center operators.

The case for ride-through standards appears straightforward after a 3,800-megawatt transfer. If the fault cleared normally, connected facilities should not have reacted like the system was collapsing.

The technical boundary is less simple. Voltage conditions vary by location, duration, phase, and equipment configuration. A safe operating envelope for one campus may not suit another.

Existing facilities present the hardest problem. They were designed under earlier interconnection rules and may contain equipment that cannot meet a future performance curve without modification.

Retrofitting protection systems can require engineering studies, controller changes, testing, and coordination with equipment vendors. Some sites may need additional power-conditioning or energy-storage equipment.

Regulators must decide whether new requirements apply only to future facilities. Grandfathering every existing site would preserve the operational risk in regions already hosting dense clusters.

Applying new rules retroactively would raise cost and jurisdiction questions. Retail service remains partly subject to state regulation, while FERC’s authority centers on interstate transmission and bulk-system reliability.

Standards must also distinguish planned flexibility from uncontrolled disconnection. A data center that reduces demand after a grid operator’s instruction can provide a valuable service.

An automated transfer caused by an unknown local threshold is different. Operators cannot schedule reserves effectively when they do not know which condition will remove a large block of demand.

Data sharing is therefore as important as physical performance. PJM needs validated load models, protection thresholds, backup-transfer logic, restoration procedures, and current facility availability.

Operators may resist sharing sensitive technical details. Campus configurations can reveal operational practices, customer requirements, and security information.

A workable regime will need controlled access and clear reporting boundaries. Grid operators require enough detail for planning without creating an unnecessary repository of sensitive facility data.

Another uncertainty involves ownership. A campus owner, cloud provider, tenant, utility, equipment vendor, and backup-power operator may influence different parts of the electrical response.

The entity signing an interconnection agreement may not directly control server power systems. A national registration framework must identify who carries each operational duty.

There is also a danger of treating all computational facilities alike. Cryptocurrency mines often respond strongly to electricity prices, while cloud facilities usually prioritize uptime.

AI training facilities may tolerate interrupted workloads differently from services handling real-time transactions. Protection and flexibility requirements should reflect measurable electrical behavior.

PJM data center reliability rules must remain performance-based where possible. Operators need predictable outcomes during disturbances, regardless of the computing service inside a building.

The July event does not prove that data centers alone threaten grid collapse. PJM restored control quickly, no customers were involuntarily shed, and the triggering line fault cleared correctly.

It does prove that aggregate load behavior has reached a scale that deserves formal treatment. Repeated events of 1,500 megawatts and then 3,800 megawatts establish a pattern regulators cannot dismiss.

Future reporting should avoid confusing load loss with lost electricity supply. The immediate problem was excess generation after demand disappeared, not an initial shortage of power.

Both directions matter to reliability. A fast demand increase can strain reserves, while a fast decrease can raise frequency and voltage. A synchronized return can produce the first problem immediately after the second.

Rules should address the full cycle, including disturbance detection, ride-through, transfer, operator notification, and staged reconnection. Focusing only on the moment of disconnection would leave the mechanism incomplete.

The Incident Pressures Data Centers, Utilities, and Grid Operators

Responsibility is distributed across the system, so no single technical change can resolve the risk.

Data center owners face the most visible pressure. They may need to disclose equipment behavior, revise protection settings, participate in studies, and accept continuing operational obligations.

Cloud providers will also need to examine whether uptime practices conflict with grid performance. A control scheme optimized for one campus may create external costs when replicated across a regional cluster.

Utilities such as Dominion sit between campuses and the regional operator. They own local transmission and distribution assets, manage interconnection requirements, and observe electrical conditions near each facility.

PJM balances the wider region. It must plan generation, schedule reserves, coordinate transmission, and manage sudden changes across a territory serving 67 million people.

NERC defines the reliability framework, but it depends on regional operators and utilities for technical evidence. The recurring Northern Virginia events provide an unusually clear record for standards development.

FERC must reconcile reliability with another federal priority: connecting large loads faster. AI developers argue that long power-delivery timelines constrain investment and computing capacity.

FERC Commissioner David Rosner See has noted that large and flexible loads can change system operations, planning needs, grid upgrades, and cost allocation. Existing systems were not designed for their current pace and scale.

Speed and reliability do not have to be opposites. Better models and standard data packages can reduce uncertainty during interconnection reviews.

Clear performance requirements can also prevent late redesigns. Developers can account for ride-through, telemetry, and backup-transfer obligations before selecting equipment.

However, rushed approvals can lock in poorly understood behavior. A campus may enter service years before a transmission project or regional standard catches up.

The pressure extends to state regulators and ordinary customers. Transmission upgrades, new generation, and reliability services carry costs that someone must pay.

FERC’s large-load orders specifically emphasize preventing cost shifting and improving transmission-cost transparency. That debate will continue even if the operational problem receives a technical solution.

Data centers can offer system benefits when properly integrated. Their controls can reduce computing demand, move workloads, use batteries, or transfer to onsite generation after a coordinated instruction.

The July incident demonstrates why that flexibility must be visible. Fast response is useful only when operators understand its trigger, size, duration, and return path.

Developers should expect power agreements to demand more than a forecasted peak load. Utilities will increasingly ask how a facility behaves during voltage sags, frequency changes, and communications failures.

Equipment vendors will face similar questions. Power supplies and controllers may need standardized performance settings that account for both computing protection and grid requirements.

Insurers and financing partners may also scrutinize compliance. A campus that cannot satisfy evolving rules could face retrofit obligations or limits on its operating arrangements.

This wider pressure explains why the story matters beyond Northern Virginia. Other regions are attracting similarly concentrated AI infrastructure, including Texas, Ohio, Georgia, and parts of the Midwest.

A standard developed after PJM’s experience could shape projects nationwide. Conversely, fragmented regional rules could turn electrical compliance into another factor in data center site selection.

Three Signals Will Show Whether Regulators Can Close the Gap

The next test is whether regulators convert broad concern into measurable duties before another large disturbance occurs.

The first signal is NERC’s December 31 submission. Initial standards should clearly assign responsibility for modeling, operational communications, protection coordination, and disturbance records.

Weak registry criteria would leave major facilities outside the framework. Overly broad criteria could capture smaller loads that do not create meaningful bulk-system risk.

The second signal is PJM’s treatment of ride-through performance. The grid operator must decide whether regional requirements can precede a complete national standard.

The most useful proposal would define measurable voltage and frequency behavior, transition reporting, and reconnection expectations. It should also address existing campuses rather than only new projects.

The third signal is the industry’s response to the next disturbance. A smaller or better-coordinated transfer would suggest that improved settings and communications are working.

Another multi-gigawatt loss after a normally cleared fault would strengthen the case for mandatory performance standards. A disorderly return from backup power would expose an additional gap.

Readers should also watch whether incident information becomes more transparent. Public operating data helped establish the scale of the July event, but facility-level causes remain less visible.

Transparent aggregate reporting can support accountability without exposing sensitive campus details. It can also help utilities identify shared equipment settings before they produce correlated behavior.

The larger lesson is not that AI data centers consume too much electricity. The immediate warning is that concentrated digital infrastructure can change demand faster than established grid practices anticipate.

PJM successfully controlled the July disturbance, but reliable recovery is not a substitute for prevention. The grid may face a similar event during more difficult operating conditions.

PJM data center reliability now depends on treating large computing campuses as active power-system participants. Their behavior must be studied, communicated, and coordinated like any other grid-scale resource.

For technology buyers and enterprise leaders, the issue reaches beyond utility policy. Cloud availability increasingly depends on agreements among facility operators, utilities, regional grids, and federal regulators. Organizations should ask providers how critical workloads handle regional power events, backup transfers, and staged recovery. They should also preserve operational records outside any single cloud environment through a searchable knowledge base. The next reliability incident will test more than electrical equipment. It will test whether the institutions supporting AI infrastructure can coordinate at machine speed.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page