NASA Solar AI Detects Emerging Active Regions Hours Before They Appear
- Sophie Larsen

- Aug 15
- 12 min read
NASA researchers reached Google News with an eye-catching claim: an AI model found solar-eruption warning signs between five and 29 hours early. The model did detect changes before new active regions became visible. However, it did not directly predict a solar flare or coronal mass ejection.
That distinction changes how the news should be understood. Active regions are magnetically complex areas that can produce sunspots, flares, and coronal mass ejections. Predicting their emergence gives forecasters an earlier view of possible trouble, but it does not reveal whether an eruption will occur.
The research team used observations from NASA’s Solar Dynamics Observatory and a long short-term memory network, known as an LSTM. The model searched for changing acoustic patterns beneath apparently quiet areas of the Sun. Its results challenge forecasting systems that begin analyzing risk only after magnetic structures become easy to see.
The result is technically interesting and operationally unfinished. It moves the observation window backward, before a sunspot becomes obvious, but rests on limited testing. That tension matters more than the amplified headline.
What NASA’s Solar AI Actually Predicted
The model forecast the visible emergence of solar active regions, not the eruptions those regions might later produce.
The study came from Spiridon Kasapis, Irina Kitiashvili, Alexander Kosovichev, and John Stefan. Their affiliations connect the work to NASA Ames Research Center and the New Jersey Institute of Technology.
The peer-reviewed active-region study appeared in the Astrophysical Journal Supplement Series in October 2025. An earlier manuscript had circulated publicly since September 2024.
Its target was continuum intensity, a measurement of visible light from the solar surface. Sunspots appear dark because strong magnetic fields suppress normal convective heat transport. Predicting a future intensity decrease therefore offers an indirect signal that an active region is forming.
That is different from predicting an eruption. A newly emerged region can remain relatively quiet, produce modest flares, or generate major space-weather events. Its emergence establishes a possible source region, not a scheduled explosion.
The researchers trained recurrent neural networks on time-series observations from the Solar Dynamics Observatory. An LSTM is a neural-network architecture designed to retain relevant information across a sequence. Here, the sequence described how solar measurements changed over time.
The model received three broad kinds of input. Doppler measurements tracked motion toward or away from the observer. Continuum images measured surface brightness, while magnetograms recorded the line-of-sight magnetic field.
The team transformed those observations into acoustic-power and magnetic-flux sequences. Acoustic power describes the strength of oscillations detected at different locations on the Sun. Changes in those oscillations can reveal activity before emerging magnetic flux produces an obvious surface feature.
The approach resembles listening for machinery behind a wall. The observer cannot see the machine, but changing vibrations can signal that something has started moving. On the Sun, turbulent convection and acoustic oscillations provide the hidden motion.
Earlier work assembled observations of 61 emerging active regions. That dataset helped researchers identify patterns associated with the transition from quiet photosphere to visible magnetic activity. The emergence dataset was designed specifically to study early-warning signals.
The later model predicted intensity values 12 hours into the future during each inference step. Researchers then examined whether its predicted intensity decline crossed their emergence criteria early enough to provide a warning.
Google News aggregation compressed that multilayered process into a claim about hidden eruption signs. The shorthand is understandable because active regions are the birthplaces of many eruptions. It still skips an essential forecasting step.
The accurate summary is narrower. NASA-backed machine learning detected patterns associated with active-region emergence before those regions became visibly established. That result can support eruption forecasting, but it is not itself an eruption forecast.
Why Looking Below the Visible Sun Matters
The research shifts attention from visible magnetic structures toward the quieter motions that precede them.
Most operational analysis becomes more informative after a region has appeared. Forecasters can then measure its size, magnetic complexity, growth rate, and recent flare history. Those observable features help estimate the probability of later activity.
This model asks an earlier question. Can the dynamics of an apparently quiet patch reveal where magnetic flux will emerge before a sunspot forms?
NASA’s Helioseismic and Magnetic Imager, or HMI, makes that question testable. HMI continuously observes the visible solar disk and measures intensity, Doppler velocity, and magnetic fields. Its repeated measurements allow researchers to follow subtle changes instead of comparing isolated snapshots.
The instrument operates aboard the Solar Dynamics Observatory, which launched in February 2010. The mission studies how solar activity forms and drives space weather. Its science objectives explicitly include understanding active-region emergence and improving forecasts.
The research focused on acoustic-power maps derived from HMI Doppler observations. Solar acoustic waves interact with turbulent convection and magnetic structures. Their changing power can therefore act as a proxy for processes that are not yet visually obvious.
The model also used magnetic flux and continuum intensity. During the successful early detections, the emerging regions had reached only 4% to 9.6% of their eventual maximum magnetic flux. The signal was present while most of the region’s later magnetic strength remained undeveloped.
That finding supplies the study’s central mechanism. The network was not identifying a complete sunspot earlier than a human observer. It was recognizing a temporal pattern associated with a future reduction in surface brightness.
The distinction between temporal and static analysis matters. A single image can show whether a location looks unusual at that moment. A sequence can reveal whether several weak measurements are moving together in a recognizable direction.
This is where LSTM models are useful. Their memory mechanism can retain patterns across earlier observations and reduce the importance of irrelevant fluctuations. The model can therefore learn how a region evolves, rather than treating each frame as an independent event.
NASA reported that its Pleiades supercomputer processed the HMI observations and ran the machine-learning workloads. A 2024 computing summary described an active region detected 18 hours before it became visible, without a false positive in that example.
That statement describes a project milestone, not a global performance guarantee. “Without false positives” referred to the highlighted test, not every possible location and time across the Sun.
Still, early active-region prediction would add useful preparation time. Forecasters could assign attention to a developing location before conventional indicators become strong. Solar telescopes with narrow fields could also prioritize promising targets earlier.
Satellite operators would not immediately change operations because an active region might emerge. They need additional evidence about flares, energetic particles, or Earth-directed plasma. Yet earlier awareness can begin a longer chain of monitoring and risk assessment.
The model’s contribution is therefore upstream. It does not replace flare forecasts, coronal mass ejection models, or geomagnetic-storm predictions. It supplies an earlier candidate location for those systems to watch.
The Google News Claim Meets a Five-Region Test
The most important number is not the 29-hour warning, but the five unseen regions used for final testing.
The researchers evaluated their trained models on five active regions excluded from training. The best configuration, called Model 8, successfully predicted all five under an experimental evaluation.
Under a stricter operational setup, it successfully predicted three. Those regions were NOAA active regions 11726, 13165, and 13179.
The reported lead times were not uniform. The model identified AR11726 about 10 hours early, AR13165 about 29 hours early, and AR13179 about five hours early.
Those results demonstrate that useful precursor information existed in the recorded sequences. They do not establish a stable warning horizon that forecasters can expect for every future region.
The model also produced an average root mean square error of 0.11 for active and quiet areas in tested variations. Root mean square error summarizes the typical difference between predicted and observed intensity values. Lower values indicate closer predictions under the study’s normalization.
That metric alone does not answer the operational question. A forecast service must identify emerging regions without flooding forecasters with false alarms across large quiet areas.
It must also work across changing viewing geometry. Measurements near the solar limb, the apparent edge of the disk, suffer stronger projection effects. Instrument conditions, data gaps, and solar-cycle differences can also alter input distributions.
The study’s experimental and operational results illustrate this gap. In the experimental setting, researchers could evaluate whether predictions matched known emergence behavior under controlled conditions. An operational test must act more like a real forecast, with less knowledge about what happens next.
Three successes from five held-out cases are promising for exploratory work. They are insufficient for declaring a dependable global warning system. A small sample can be strongly influenced by the selected regions, thresholds, and definitions of success.
There is also a distinction between prediction horizon and warning time. The network was trained to forecast continuum intensity 12 hours ahead. Repeated forecasts and the timing of threshold crossings produced reported emergence warnings extending as far as 29 hours.
That does not mean the network generated a single precise 29-hour forecast of a visible sunspot. It means its evolving predictions crossed the researchers’ emergence criteria that far before the reference time.
Headline compression often removes those details. A reader scanning Google News can easily infer that the model anticipated a dangerous eruption 29 hours before it occurred. The underlying paper supports a more cautious conclusion.
The model anticipated the visible formation of one active region 29 hours before its recorded emergence. Active regions are associated with elevated eruption risk, but the model did not say that region would produce a particular flare.
The difference resembles storm formation and tornado prediction. Detecting an environment where a storm will develop is useful. It does not predict whether that storm will generate a tornado, where the tornado will travel, or whom it will affect.
That limitation does not make the work unimportant. Forecasting chains improve when their earliest stages start sooner. However, each stage needs separate validation against the outcome it claims to predict.
The study’s strongest evidence concerns active-region emergence. Claims about solar eruptions remain downstream implications that future systems must test directly.
Early Emergence Detection Pressures Existing Forecast Workflows
A reliable pre-emergence signal would force space-weather forecasting to monitor quiet regions, not only visible active ones.
Current workflows already combine human analysis, physical measurements, statistical methods, and machine learning. Once an active region appears, forecasters can evaluate magnetic complexity and previous activity.
Pre-emergence prediction changes where that process starts. It directs computing and observational resources toward locations that still look quiet in ordinary intensity images.
The immediate pressure falls on data pipelines. A research model can process carefully prepared observations after an event. An operational service must receive, calibrate, and analyze full-disk data continuously.
HMI generates measurements across the visible Sun at regular intervals. NASA’s public SHARP products, meaning Space-weather HMI Active Region Patches, package magnetic and intensity information around already identified regions.
The model’s objective sits partly before that patch exists. A production system would need to scan quiet-Sun tiles, maintain their histories, and update emergence probabilities as new observations arrive.
That raises a class-imbalance problem. At any moment, most candidate locations will not produce a large active region. Even a low false-positive rate can create many alerts when applied across the entire disk and many time steps.
Forecast designers must therefore evaluate precision, recall, false-alarm rate, and warning time together. Maximizing lead time is not enough if alerts become too frequent to trust.
Researchers must also choose what an alert means. It could mark a predicted intensity decline, a likely magnetic-flux emergence, or a probable NOAA-designated active region. Those outcomes overlap, but they are not interchangeable.
The approach competes less with one company than with a conventional observational sequence. Traditional workflows wait for clear surface magnetic evidence, then estimate flare risk. The new route tries to identify a developing source before that evidence fully emerges.
This is the article’s main opponent: early acoustic inference versus visible-region monitoring. The new model extends the first route but still depends on the second for confirmation.
Other machine-learning projects address later stages. Some classify whether established active regions will flare. Others estimate whether a flare will accompany a coronal mass ejection or produce energetic particles.
Those systems often use SHARP parameters, magnetograms, or X-ray histories after relevant activity is measurable. They answer questions closer to the final hazard, but usually begin later in the physical sequence.
The NASA Ames model moves earlier while accepting greater uncertainty about the final outcome. A region can emerge without generating a damaging event. The earlier the forecast begins, the more branching possibilities remain.
That tradeoff should guide operational design. A pre-emergence alert belongs in an attention-ranking system, not an automatic shutdown command for spacecraft or infrastructure.
For example, an alert could ask observatories to preserve higher-cadence data for a location. It could also trigger additional model runs as magnetic flux strengthens. Later evidence would determine whether the situation deserves an operational warning.
This tiered response resembles anomaly detection in other technical systems. An early signal opens an investigation. It does not independently establish the severity or cause of the anomaly.
The model could also improve scientific observation. High-resolution solar instruments cannot always monitor every promising location at once. Earlier targeting would increase the chance of capturing the entire formation process.
Complete sequences are valuable because solar eruptions remain difficult to explain. Researchers need observations from before emergence through magnetic evolution and, sometimes, eruption. Missing the first stage limits causal analysis.
The practical value may therefore arrive first in research operations. Better instrument targeting and data collection could precede any public forecasting product.
What the Results Still Cannot Establish
The model has not shown that its precursor patterns generalize across the full Sun, multiple instruments, and an entire solar cycle.
Five unseen active regions provide an initial test, not a representative operational benchmark. Solar regions differ in size, latitude, magnetic structure, growth rate, and viewing angle.
A model can perform well on a small held-out set yet struggle when conditions shift. This issue, called distribution shift, occurs when operational inputs differ from the training data.
Solar activity itself follows an approximately 11-year cycle. Observations near maximum activity contain more crowded and complex magnetic environments than observations near minimum. Validation should cover both conditions.
The researchers also need larger quiet-Sun control sets. An emergence model must distinguish genuine precursors from ordinary acoustic variability across many locations that never develop notable magnetic activity.
False positives deserve particular attention. A scientific case study can highlight a successful warning. An operational forecaster needs the number of incorrect alerts produced while searching for that success.
The NASA computing summary highlighted an 18-hour detection without false positives. The peer-reviewed paper provides broader methodological evidence, but it does not convert that example into a universal no-false-alarm claim.
Reproducibility is another requirement. Independent teams should apply the method to comparable HMI sequences using documented preprocessing, tiling, thresholds, and evaluation rules.
The underlying observation source is well established. NASA’s SHARP dataset contains vector magnetic fields, line-of-sight fields, continuum intensity, Doppler velocity, and related products.
However, a shared data source does not guarantee comparable results. Choices about normalization, temporal windows, active-region labels, and quiet controls can strongly affect model performance.
The model’s interpretability also remains limited. Researchers linked its predictions to changes in acoustic power, but an LSTM does not automatically explain which physical process caused every prediction.
A useful scientific system should show where and when influential signals appeared. It should also distinguish acoustic precursors from instrumental artifacts or patterns correlated with the labels.
Physics-informed validation can help. Researchers can compare the model’s high-salience regions with simulations of rising magnetic flux and known helioseismic signatures. Agreement would strengthen the physical interpretation.
Cross-instrument testing would be valuable too. HMI offers a long, consistent record, but future operational systems cannot assume one instrument will remain available indefinitely. A portable method needs stable behavior across replacement or complementary sensors.
Most importantly, active-region emergence is only one link in the hazard chain. The next links include flare production, coronal mass ejection direction, energetic-particle acceleration, and interaction with Earth’s magnetic field.
Errors compound along that chain. An early region alert can be correct while a later eruption forecast is wrong. A flare prediction can also be correct while the event poses little risk to Earth.
The article’s title direction therefore overstates what has been established. “Solar-eruption signs” suggests evidence tied directly to eruptions. The verified signal was tied to emerging active regions that can become eruption sources.
This is not just semantic caution. Users act differently on a source-region alert and a hazard forecast. Satellite operators, grid planners, and aviation services need calibrated probabilities tied to specific impacts.
For readers tracking AI claims through Google News, the lesson is broader. A model’s target variable matters as much as its accuracy. Predicting a precursor does not equal predicting every downstream event associated with that precursor.
Three Signals Will Show Whether the Model Can Become Operational
The next phase must prove scale, control false alarms, and connect early emergence warnings to later hazards.
The first signal is a much larger prospective test. Researchers should run the model on incoming observations without selecting known emergence cases in advance.
A prospective evaluation would record every alert, every missed active region, and every quiet location incorrectly flagged. It should span enough time to include changing solar conditions and viewing geometries.
Success would mean the five-region result survives a realistic full-disk test. Failure would suggest that the original cases captured a narrower pattern than the headline implied.
The second signal is a transparent false-alarm analysis. Researchers should publish precision, recall, warning-time distributions, and false alarms per observing period.
An alerting system must show whether longer lead times come at the cost of lower reliability. That tradeoff determines whether forecasters can integrate the model without adding distracting noise.
A useful benchmark should also separate small emergences from large, space-weather-relevant regions. Correctly detecting many weak regions may raise overall accuracy without improving preparation for consequential events.
The third signal is integration with downstream flare and eruption models. A pre-emergence warning becomes more valuable if later measurements update the same region’s evolving risk.
That integrated system could begin with acoustic evidence, add magnetic-flux growth, and then incorporate field complexity or flare history. Each stage would refine the probability rather than treating the first alert as a final answer.
Such integration would strengthen the central claim if it improves warning time without degrading calibration. It would weaken the claim if early alerts add little beyond forecasts made after visible emergence.
There is no need to wait for a perfect end-to-end system before using the research. Observatories can test early targeting, and scientists can collect more complete emergence sequences now.
Public communication needs greater discipline. A Google News headline can make “predicting a precursor” sound like “predicting an eruption.” Readers should check the model’s actual output, evaluation sample, and operational definition.
The underlying result remains worth following. A neural network identified evolving signals while active regions were still mostly hidden, including three operationally successful cases with five-to-29-hour lead times.
The open question is no longer whether those signals exist in selected observations. It is whether they recur reliably across the Sun and improve the forecasts that matter to people and machines.
Watch for prospective full-disk testing, complete false-alarm statistics, and links to later eruption probabilities. Those three results will determine whether this Google News story marks an operational advance or a promising research milestone.


