top of page

Microsoft Research Is Forecasting Space Weather Risks on Power Grids, but Operators Still Need Proof

1 day ago
11 min read

Microsoft Research has built a system for forecasting space weather risks on power grids 30 to 60 minutes before specific hazards appear. The experimental pipeline estimates localized risk across 66,935 substations in the continental United States. That is a sharper target than a broad warning that a geomagnetic storm is approaching.

The system addresses a persistent gap between space-weather forecasts and decisions inside a utility control room. Operators need to know which locations face the most stress, when that stress will arrive, and whether protective action is justified.

Microsoft says its system detected nearly 80% of major events during evaluation. Yet the company also acknowledges that utilities must validate the model with operational data before depending on it. The central question is no longer whether machine learning can produce a continental risk map. It is whether those predictions remain reliable during rare storms, when mistakes carry their highest cost.

Forecasting space weather risks on power grids becomes location-specific

Microsoft’s main advance is the connection between an incoming solar disturbance and a risk estimate for an individual substation.

The September 30, 2026, risk forecasting pipeline starts with solar-wind measurements from the first Sun-Earth Lagrange point, known as L1. This observation point sits upstream in the solar wind, giving forecasters a short view of conditions moving toward Earth.

The system forecasts two measures of geomagnetic activity. The Auroral Electrojet index, or AE, tracks magnetic activity associated with electrical currents around the auroral region. The Disturbance Storm Time index, or Dst, describes the strength of the ring current surrounding Earth during a storm.

Those global indicators do not reveal the complete risk to a particular grid asset. Microsoft therefore combines them with latitude, local geology, ground conductivity, and modeled grid infrastructure. A gradient-boosting model then estimates dB/dt, the rate at which the local magnetic field changes over time.

That rate matters because rapid magnetic changes induce electric fields in the ground. Those fields can drive geomagnetically induced currents, or GICs, through grounded transmission equipment. A GIC behaves much like direct current inside a network designed for alternating current.

The resulting flow can push a transformer core into saturation. According to the federal explanation of transformer saturation, saturation can produce abnormal currents, voltage problems, heating, and greater reactive-power demand. Severe conditions can contribute to equipment damage or service disruption.

This chain explains why a general storm alert has limited operational value. Two substations can face different exposure during the same storm because their latitude, geology, grounding, and connected transmission lines differ. Direction also matters because the induced electric field interacts with line orientation.

Microsoft’s model tries to preserve those local differences while covering an entire continent. It generated estimates for all 66,935 substations in approximately 333 milliseconds during measured inference. That speed would let analysts test many scenarios without waiting for a lengthy simulation.

The output is not a direct measurement of current flowing through every transformer. It is a modeled estimate derived from magnetic-field changes and location-specific inputs. That distinction will matter when utilities evaluate whether the warnings are accurate enough to influence operations.

Still, the work changes the form of a Microsoft space weather forecast. Instead of stopping at “a strong storm is coming,” the pipeline asks where its ground effects are most likely to become dangerous.

The warning window is short, but operationally meaningful

Thirty to 60 minutes is not enough time for major construction, but it can support focused operational decisions.

Space-weather prediction has an unavoidable timing problem. Forecasters can observe an eruption leaving the Sun days before it reaches Earth. However, its most consequential magnetic properties remain uncertain until the disturbance approaches instruments near L1.

The solar wind can also change rapidly. Its interaction with Earth’s magnetosphere depends on speed, density, and magnetic orientation. Those conditions determine how efficiently energy enters the near-Earth environment.

NOAA guidance says measurements near L1 commonly provide only minutes of detailed warning before solar-wind conditions reach Earth. Microsoft is working within that physical constraint, not eliminating it. Its contribution is to turn that limited warning into geographically specific estimates.

A useful forecast could help operators review exposed equipment, increase attention on selected substations, or prepare targeted network adjustments. Microsoft identifies possible actions such as changing reactive-power reserves or temporarily reconfiguring parts of the network.

These actions are not trivial. A conservative response can carry costs, constrain transfers, and complicate normal system operations. A missed event can expose equipment to unusual thermal and electrical stress.

That creates an asymmetric decision problem. False alarms can encourage unnecessary intervention, while false negatives can leave important assets exposed. A useful warning system must therefore provide more than an impressive average detection rate.

It must communicate confidence, timing, geographic scope, and the severity threshold being forecast. It must also fit into existing procedures that define who can act and under what conditions.

North American operators already work under requirements addressing geomagnetic disturbances. The applicable operating standard requires plans, processes, and procedures intended to reduce the effects of such events.

That existing framework gives Microsoft’s research a possible path into practice. The model does not need to replace official alerts or operating plans. It could become an additional decision layer that helps utilities direct attention within those plans.

However, research output and operational authority are different things. An operator must understand a model’s behavior when sensors fail, inputs arrive late, or storm conditions fall outside the training distribution.

The 30-to-60-minute warning window makes those issues more important. Teams cannot spend most of that period debating whether a prediction is trustworthy. Procedures, thresholds, and accountability must be established before the storm arrives.

A Microsoft space weather forecast would therefore face both a technical test and an institutional one. It must produce accurate local estimates, then deliver them in a form that trained operators can use under time pressure.

The model combines physical signals with machine learning

The pipeline’s strength comes from combining global storm forecasts with the local conditions that shape ground-level risk.

Microsoft divided the prediction process into three stages. First, the system uses solar-wind observations to forecast AE and Dst. It also prepares geographic and conductivity features for each modeled substation.

Second, the machine-learning stage combines those inputs to estimate dB/dt. Third, the pipeline converts those predictions into local risk estimates and aggregates them into a continental assessment.

This is a mechanism-focused design rather than a single opaque prediction. Each stage corresponds to a physical step between solar activity and infrastructure exposure. That structure should make errors easier to investigate than one model that jumps directly from solar-wind data to grid risk.

The project used public data. Microsoft identified NASA OMNI solar-wind records, Kyoto World Data Center indices, INTERMAGNET observations, U.S. Geological Survey magnetometer data, and GridSFM-derived infrastructure information.

Local conductivity is especially important. Resistive rock can support stronger geoelectric fields than more conductive ground under comparable magnetic disturbance. Coastal boundaries and complex geological structures can further change the response.

Recent federal conductivity mapping illustrates the scale of that challenge. The United States Magnetotelluric Array gathered data at more than 1,800 stations between 2006 and 2024.

That effort produced a nationwide electrical portrait of the crust and upper mantle. Such measurements can reveal vulnerabilities that simpler, one-dimensional conductivity assumptions might overstate or underestimate.

Microsoft’s location features bring this geographic variation into the forecasting system. They also make expansion beyond the continental United States more complicated. A model cannot simply reuse the same local inputs in another country with different geology and grid topology.

The project also used a system of 50 AI agents to explore features, validation strategies, and model configurations. That detail describes the research workflow, not the forecasting model operating inside a utility.

The distinction matters because agent count does not measure forecast quality. The important evidence comes from held-out performance, event detection, comparison with baselines, and eventual validation against operational measurements.

Microsoft reported an AE root mean square error of 410.2 nanoteslas. Root mean square error summarizes the typical size of prediction errors while placing greater weight on large misses.

For Dst, the reported error was 7.2 nanoteslas. During the most active periods in the 2020-to-2026 evaluation, Microsoft says the model outperformed the Burton equation during 62.2% of individual hours.

The Burton equation is an established empirical method for estimating storm-time Dst from solar-wind inputs. Beating it during many active hours is encouraging, but it does not settle every operational question.

Performance can vary by storm, region, threshold, and forecast horizon. A model can achieve a favorable aggregate error while missing a small number of dangerous extremes. Those tail events matter most to infrastructure operators.

The system’s staged design offers a route for diagnosing such failures. Researchers can examine whether the error began in the solar-wind forecast, the geomagnetic indices, the local dB/dt estimate, or the conversion into infrastructure risk.

Detection rates show promise, not operational readiness

The published results establish a credible research signal, but they do not yet establish dependable grid protection.

Microsoft evaluated the final risk stage at three dB/dt thresholds. The reported detection rate was 76.5% for major events at or above 10 nanoteslas per minute. It reached 81.2% for severe events at or above 20 nanoteslas per minute.

For extreme events at or above 50 nanoteslas per minute, detection fell to 64.1%. That means the model missed more than one-third of the events in the most extreme category during the reported evaluation.

These figures should not be merged into one general accuracy claim. Each threshold creates a different evaluation subset. Event frequency, class balance, location, and timing can all affect the resulting rate.

Microsoft compared the GIC risk stage with simple linear regression because it found no equivalent, widely deployed system for a direct operational benchmark. That makes the comparison informative but limited.

A linear model can test whether the additional machine-learning structure captures useful nonlinear relationships. It does not show that the system performs better than every forecasting product already used by utilities or government agencies.

NOAA has operated versions of its Geospace model for years. That physics-based system predicts regional magnetic disturbances and supports situational awareness for power-grid operators.

The two approaches are not exact substitutes. NOAA’s system models geospace dynamics, while Microsoft’s pipeline connects forecast indices and location features to modeled substation risk. They can be viewed as complementary paths through the same forecasting problem.

Microsoft also reports that adding its Dst prediction to the AE forecast improved severe-event detection by 1.2 percentage points. This gain supports the use of multiple physical indicators, although it is not a large shift by itself.

The harder issue is whether reported detection rates transfer to live operations. Historical datasets can contain gaps, uneven geographic coverage, and very few examples of the rarest storms.

Researchers must also prevent information leakage, which occurs when evaluation data indirectly influence model development. The Microsoft post does not provide enough detail for an outside reader to reproduce every split, threshold, or test condition.

That does not invalidate the results. It means the evidence should be treated as an initial research report rather than an independently replicated performance record.

The current maps are also demonstrations, not records from a live operational event. Microsoft explicitly says further validation with utilities and operational data is necessary before grid use.

That statement defines the appropriate standard. Forecasting space weather risks on power grids requires measured current data, asset details, and operator feedback that public datasets cannot fully supply.

Substation coordinates alone do not describe transformer design, grounding resistance, network state, or asset condition. Microsoft identifies transformer-level modeling as future work, confirming that the present output lacks some equipment-specific detail.

There is also a communication risk. A colored continental map can look more certain than its underlying estimates. Operational displays need calibrated uncertainty, clear timestamps, data-quality warnings, and explanations for sudden changes.

The research is strongest when understood as a prioritization system. It can identify locations that warrant closer analysis. It does not independently establish that a particular transformer will fail.

The May 2024 storm shows why localization matters

Recent experience shows that one geomagnetic storm can create different failures across grids, farms, satellites, and navigation systems.

The Gannon storm reached Earth between May 10 and May 13, 2024. It produced the first G5 geomagnetic conditions in more than two decades and pushed auroras far beyond their usual latitudes.

Its effects were widespread but uneven. NASA’s Gannon storm review recorded tripped high-voltage lines, overheated transformers, disrupted tractor guidance, and changes to some flight routes.

GPS-guided tractors in parts of the American Midwest veered off course or stopped working. These failures arrived during the planting season, when precise positioning supports efficient field operations.

Other systems experienced different problems. Satellite operators managed increased atmospheric drag, navigation signals degraded, and some communications became less reliable.

NASA’s event record says GOES-16 stopped transmitting data for nearly two hours on May 13. It also describes degraded Starlink service and electrical problems affecting the European Space Agency’s Gaia spacecraft.

Yet the storm did not cause a broad public power emergency in the United States. Government reports noted grid irregularities without a major population-level outage.

That contrast is central to space weather grid impact analysis. A severe planetary disturbance does not create uniform damage. Local infrastructure, engineering controls, geography, and operating decisions shape the result.

The March 1989 Hydro-Québec collapse remains the historical warning. Rapid magnetic changes induced currents that contributed to a grid failure, cutting power to more than six million people.

Modern utilities have improved monitoring, standards, and operating procedures since that event. However, electrical networks have also become more interconnected and dependent on satellite timing, digital communications, and precise navigation.

This combination raises the value of localization. A continental warning can tell operators that conditions are dangerous. It cannot identify which substations deserve immediate attention.

Microsoft’s pipeline aims to narrow that gap by estimating where rapid magnetic-field changes will become most relevant. Its 66,935-location output provides much more geographic detail than a national storm category.

However, the May 2024 storm also shows why evaluation cannot focus only on the electric grid. The same solar event can affect agriculture, aviation, spacecraft, radio links, and positioning services.

Those systems create dependencies. Grid operators use communications and timing services, while recovery teams depend on transportation and navigation. A localized electrical forecast becomes more valuable when combined with forecasts for those connected systems.

Microsoft’s project currently focuses on substation risk. It should not be interpreted as a complete model of every economic or infrastructure consequence.

The strongest future systems will likely connect several layers. They will combine storm arrival, regional magnetic response, local conductivity, network topology, asset characteristics, and current operating state.

That is a demanding goal. The May 2024 experience shows why it is worth pursuing and why a single impressive map is not enough.

Three signals will determine whether utilities can use it

The next phase must replace promising retrospective results with utility-grade evidence, uncertainty estimates, and asset-level detail.

The first signal is validation with measured GIC and operational utility data. Microsoft needs to test whether high predicted dB/dt aligns with currents observed at grounded transformers during multiple storms.

This validation should include different regions, geological settings, network configurations, and storm intensities. Strong agreement would support the model’s local risk rankings. Weak agreement would expose missing grid or conductivity features.

Utilities should also examine false alarms and missed events separately. A model that catches most storms but repeatedly identifies the wrong locations will have limited operational value.

The second signal is a move from substation-level output to transformer-level risk. Microsoft says it wants to add asset characteristics so forecasts can distinguish equipment within the same site.

That work would require information that is often sensitive or unavailable in public datasets. Transformer construction, grounding, loading, protection settings, and network connectivity can affect the current experienced by an asset.

Partnerships with utilities will therefore matter as much as further model development. Without those details, the system remains a geographic screening tool rather than an equipment protection system.

The third signal is integration with established forecasting and operating workflows. Researchers should test whether the output improves decisions when used beside official warnings, geoelectric-field maps, and existing procedures.

A successful pilot would define alert thresholds before an event. It would document which actions are available, who approves them, and how uncertainty affects the response.

Integration also needs resilience. The system should report stale inputs, missing observations, and degraded forecast confidence. Operators need a safe fallback when a sensor or data feed fails.

Longer forecast horizons will attract attention, but they should not become the only measure of progress. More warning time has little value if uncertainty grows too quickly or local predictions become unreliable.

Microsoft has proposed temporal-transformer methods for learning longer patterns in solar-wind data. It also wants to expand internationally and connect forecasts with power-flow analysis through GridSFM.

Those directions make technical sense. Each also adds a new validation burden. Another country will bring different geology, grid models, regulations, and data quality.

For now, forecasting space weather risks on power grids is best viewed as a credible research advance with a clearly stated deployment gap. The pipeline is fast, geographically detailed, and grounded in relevant physical signals.

Its most important result is not a claim that machine learning has solved geomagnetic risk. It is a demonstration that broad storm warnings can be translated into thousands of local estimates within seconds.

The next storm will provide another demanding test. Will predicted hotspots match measured currents, and will operators find the warning early enough to act? Those answers will determine whether Microsoft’s system remains a research prototype or becomes part of grid defense.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page