AgeNet-SHAP Maps Regional Brain Age, but It Is Not an Alzheimer’s Risk Test
- Ethan Carter

- 4 hours ago
- 11 min read
Google News surfaced an AI study that maps 145 brain regions, despite a crucial limit: the model does not predict who will develop Alzheimer’s disease.
Researchers at Washington University School of Medicine in St. Louis developed AgeNet-SHAP using structural MRI data from 825 participants. Their system estimated brain age and ranked the regions influencing each estimate. It also found that regional patterns differed across cognitively normal participants, people with mild cognitive impairment, and people with Alzheimer’s disease.
That combination sounds close to an individual risk assessment. It is not one. The model connected MRI-derived brain volumes with age, diagnostic groups, and clinical severity inside a research dataset. It did not establish a screening threshold, predict future diagnoses, or prove that accelerated regional aging causes Alzheimer’s disease.
The real advance is narrower and more useful. AgeNet-SHAP turns a single brain-age estimate into an interpretable map of anatomical contributions. The tension lies between that research insight and the much stronger clinical meaning readers may attach to the word “risk.”
What the Google News Headline Leaves Out
AgeNet-SHAP maps associations within MRI data, not a person’s future probability of developing Alzheimer’s disease.
The underlying peer-reviewed study appeared in Biology Methods and Protocols in 2025. Its authors are Gauri Darekar, Taslim Murad, Hui-Yuan Miao, Deepa S. Thakuri, Ganesh B. Chand, and the Alzheimer’s Disease Neuroimaging Initiative.
The researchers analyzed baseline T1-weighted structural MRI scans. This common MRI sequence provides anatomical contrast that helps software measure the volumes of different brain tissues and structures.
Their experimental dataset contained 825 participants between 55.1 and 94 years old. Of those participants, 684 came from ADNI-2, while 141 came from other phases of the Alzheimer’s Disease Neuroimaging Initiative.
The ADNI-2 group included 206 cognitively normal participants and 478 people classified within the mild cognitive impairment or Alzheimer’s continuum. The latter group included 171 with early mild cognitive impairment, 152 with late mild cognitive impairment, six with unspecified mild cognitive impairment, and 149 with Alzheimer’s disease.
The remaining ADNI phases contributed 28 cognitively normal participants and 113 participants with mild cognitive impairment or Alzheimer’s disease. These data came from the ADNI research program, a long-running public-private project built to study biomarkers and disease progression.
Researchers segmented every MRI into 145 regions of interest. Those measurements covered gray matter, white matter, and cerebrospinal fluid areas. They normalized regional volumes to reduce differences caused by head size.
AgeNet then learned a relationship between those measurements and chronological age. It used four hidden neural-network layers containing 256, 128, 64, and 32 units. The researchers trained and tested it through ten-fold cross-validation, repeating the process five times.
Cross-validation separates a dataset into multiple portions, repeatedly training on some portions and testing on another. This approach reduces dependence on one favorable split, although it does not replace validation in an independent clinical population.
The team compared AgeNet with Lasso regression, ridge regression, and support vector regression. According to the paper, the neural network produced stronger correlations between predicted and chronological age than those conventional models.
That result establishes AgeNet as the team’s preferred age-estimation model. It does not establish superior Alzheimer’s diagnosis, because predicting chronological age was the model’s primary training task.
This distinction matters when a short headline moves through Google News. “Brain age” is a model-generated estimate based on patterns in reference data. It is not a measured biological clock, and it is not interchangeable with dementia status.
A younger estimated brain age does not guarantee cognitive health. An older estimate does not diagnose disease. The output summarizes how closely selected anatomical features resemble patterns associated with age in the training data.
The study therefore changed the granularity of the question. Instead of asking only whether a brain looks older than expected, researchers can ask which combination of regions drove that answer.
That is valuable for hypothesis generation. It also creates a larger interpretation burden, because an explanation of a model remains an explanation of that model. It is not automatically an explanation of human biology.
Explainable AI Replaces One Score With a Regional Map
The central mechanism is SHAP, which assigns each measured region a contribution to AgeNet’s prediction.
Many brain-age systems compress an MRI into one number. Researchers can compare that predicted age with chronological age to calculate a brain-age gap. A positive gap means the model sees a pattern resembling an older brain.
That summary has obvious appeal. It is also blunt. Two people can receive similar global estimates even when different anatomical changes drive their results.
AgeNet-SHAP tries to expose those differences. SHAP, short for Shapley additive explanations, estimates how each input influences a model’s output relative to a reference. Its logic comes from a method for distributing credit among cooperating contributors.
For this study, each contributor was a regional brain-volume feature. A larger absolute SHAP value meant that a region had a stronger influence on the age prediction for a particular participant.
The team first tested the method with semi-simulated data from 187 cognitively normal participants. Researchers deliberately perturbed selected regional measurements, giving them a known answer against which to evaluate the explanation methods.
AgeNet was paired with three interpretation approaches: SHAP, local interpretable model-agnostic explanations, and layer-wise relevance propagation. The authors report that AgeNet-SHAP identified all intentionally perturbed regions as important predictors.
That test supports the pipeline’s technical validity under controlled conditions. It shows that the method can recover deliberately introduced signals when researchers know which features were altered.
The simulation does not show that every highly ranked region in real patient data represents an Alzheimer’s mechanism. Real anatomy contains correlated measurements, biological variation, scanner effects, age effects, and disease-related changes that cannot be cleanly separated.
The paper’s multivariate approach is still important. Brain structures do not age independently, and neurodegenerative disease does not respect a simple one-region-at-a-time model. Volumes across connected or developmentally related areas often change together.
Traditional analyses frequently test each region separately. That process can identify areas associated with age or diagnosis, but it can miss how combinations of measurements jointly affect a prediction.
AgeNet-SHAP instead asks how the neural network uses all 145 features together. This approach captures nonlinear relationships that a simple linear model might not represent.
The tradeoff is that SHAP values depend on the trained model, selected baseline, feature correlations, and available data. A high value means a feature influenced AgeNet’s output. It does not mean that feature independently caused accelerated aging.
This is a wider issue for explainable medical AI. A systematic SHAP review found growing use of SHAP and related methods in Alzheimer’s detection research. It also emphasized explainability as a requirement for assessing model trustworthiness.
Trustworthiness needs more than a visually intuitive heat map. Researchers must test whether explanations remain stable across scanners, demographic groups, preprocessing methods, and external datasets.
The AgeNet-SHAP study provides a framework for that work. It does not complete it.
Its strongest contribution is methodological: it connects a reasonably simple neural network with participant-level regional explanations, then relates those explanations to diagnostic and clinical measures.
That creates a path from prediction toward interpretation. It also reveals why the word “explainable” requires care. The system explains which measurements affected its estimate, not why a disease developed.
Alzheimer’s Severity Appears as a Distributed Pattern
The disease signal grew broader from mild cognitive impairment to Alzheimer’s disease, rather than collapsing into one decisive structure.
When the researchers compared diagnostic groups, mild cognitive impairment showed moderate regional differences relative to cognitively normal participants. Alzheimer’s disease produced stronger and more widely distributed differences.
That progression supports the paper’s main biological interpretation. Advanced disease involves a network of anatomical changes, so a multivariate regional model can reveal patterns that a global brain-age score obscures.
The study also examined Clinical Dementia Rating Sum of Boxes scores. CDR-SB combines ratings across six cognitive and functional domains, producing a measure of clinical impairment severity.
Individualized AgeNet-SHAP features showed associations with CDR-SB across the Alzheimer’s continuum. In other words, regional contributions to predicted brain age varied alongside clinical severity.
Several implicated areas align with established Alzheimer’s research. The authors discuss temporal and limbic structures associated with memory, alongside frontal, parietal, occipital, and subcortical regions.
They also found negative correlations between clinical severity and features involving the superior parietal lobule and occipital fusiform gyrus. The researchers interpret the broader results as evidence that both normal aging and disease-specific processes contribute to observed regional patterns.
That language is more precise than saying the model “maps Alzheimer’s risk.” The analysis was largely cross-sectional, meaning it compared measurements and clinical status around a baseline visit. It did not follow every cognitively normal participant until an outcome became known.
Risk prediction requires a different study design. Researchers would need to start with people who do not have Alzheimer’s dementia, generate predictions, and observe who later develops impairment or disease.
They would also need to specify a time horizon. A five-year conversion estimate is clinically different from a lifetime association with disease.
The model would then require calibration, meaning predicted probabilities must match observed outcomes. Discrimination alone, such as separating groups already known to differ, would not be enough.
AgeNet-SHAP did not perform that complete exercise. Its age predictions and regional severity associations are research findings, not validated individual probabilities.
This gap does not make the results unimportant. It defines where they fit in the evidence chain.
Regional brain age can help researchers form more specific hypotheses. A global estimate might indicate atypical aging, while a regional profile can show whether influential features cluster in memory-related structures or appear more broadly.
Other teams are moving in a similar direction. A 2023 regional brain-age model used ridge regression trained on 3,418 healthy controls and tested on 651 independent controls.
That earlier work achieved a mean absolute error of seven years in its test group. It highlighted structures including the nucleus accumbens, inferior temporal gyrus, thalamus, brainstem, and caudate nucleus.
Its simpler statistical approach offered direct estimates of each region’s contribution. AgeNet-SHAP pursued nonlinear prediction and participant-level feature attribution in an Alzheimer’s-focused dataset.
The comparison establishes the main opponent in this story: interpretable regional research versus clinically validated personal risk prediction.
Regional explanations promise more anatomical detail. Clinical prediction demands prospective evidence, stable performance, useful thresholds, and proof that the result improves decisions.
AgeNet-SHAP advances the first route. Headlines can easily imply that it has completed the second.
Why This Is Not Yet an Alzheimer’s Risk Test
The model’s explanations are scientifically interesting, but its dataset and task do not support clinical screening claims.
The first limitation is cohort dependence. Every experimental scan came from ADNI, a research program with structured enrollment, imaging protocols, and extensive clinical characterization.
ADNI has generated an unusually valuable dataset. However, its participants do not automatically represent everyone who receives an MRI in routine care.
A clinical population can differ in education, race, ethnicity, income, coexisting illness, medication exposure, scanner access, and disease stage. Those differences can alter both measured anatomy and model performance.
External validation must therefore test the complete pipeline on data collected elsewhere. Cross-validation within ADNI reduces overfitting risk, but every fold still comes from the same broader research environment.
The second limitation is the model’s target. AgeNet learned to estimate chronological age, not future Alzheimer’s onset. SHAP then explained that age estimate.
Researchers subsequently compared those explanations across diagnostic groups and clinical severity. This is a legitimate exploratory analysis, but it cannot convert an age-prediction model into a validated disease-risk calculator.
The third limitation is modality. Structural MRI captures anatomy, including regional volume loss. Alzheimer’s disease also involves molecular pathology that amyloid and tau biomarkers can measure through positron emission tomography, cerebrospinal fluid, or blood tests.
The authors explicitly identify PET, electroencephalography, and diffusion tensor imaging as future extensions. Combining modalities might separate general aging from disease-specific processes more effectively.
The fourth limitation concerns correlated features. Regional volumes influence one another through anatomy, development, degeneration, and measurement procedures. SHAP must distribute importance across those related inputs.
Different feature groupings, baselines, or preprocessing pipelines can change that distribution. Two explanations can therefore differ even when their models produce similar age estimates.
The fifth limitation is causal interpretation. A region’s SHAP value answers a computational question: how did this feature influence this model’s output for this input?
It does not show that shrinking or preserving that region will alter disease progression. It also cannot determine whether a structural difference is a cause, consequence, correlate, or compensation.
The sixth limitation is clinical usefulness. A research model must outperform or complement existing assessments before entering practice. That means comparison with cognitive testing, clinician judgment, blood biomarkers, genetic information, and established imaging markers.
A newer patch-based framework illustrates another regional route. It trained separate three-dimensional neural networks on bilateral patches from ten subcortical structures.
That system combined regional predictions into a global age estimate and paired it with cognitive scores to create a proxy for cognitive reserve. Its reported age-estimation correlation reached 0.983 with a mean absolute error of 2.92 years.
Those numbers should not be ranked directly against AgeNet-SHAP without aligned datasets and evaluation rules. Different cohorts, age distributions, preprocessing choices, and test procedures can produce materially different results.
The broader competition is not a leaderboard contest. Researchers are testing several ways to gain anatomical specificity without losing generalizability.
Some systems use whole-brain images and generate saliency maps. Others operate on segmented volumes, voxels, or predefined patches. Simpler regression models sacrifice some flexibility but make coefficients easier to inspect.
No single design has resolved the central tradeoff. More complex models can capture nonlinear interactions, yet complexity increases the work required to validate explanations.
For clinicians, a false sense of precision is a direct risk. A color-coded regional map can look definitive even when uncertainty remains high.
For patients, an “older” regional estimate could cause distress without providing an actionable diagnosis. Conversely, a reassuring estimate might discourage appropriate evaluation despite cognitive symptoms.
Medical AI therefore needs communication safeguards alongside technical validation. Reports should distinguish predicted age, age gap, diagnostic association, severity association, and prospective disease risk.
Google News readers rarely receive those distinctions from a headline alone. The underlying evidence supports a regional map of model contributions and disease-associated patterns. It does not support self-diagnosis or treatment decisions.
Three Signals Will Show Whether Regional Brain Age Matters
Independent validation, prospective prediction, and added clinical value will determine whether AgeNet-SHAP becomes more than a research framework.
The first signal is replication outside ADNI. Researchers need to run the frozen pipeline on independent cohorts collected at different hospitals, with different scanners and broader participant demographics.
A convincing result would preserve age-estimation performance and produce reasonably stable regional patterns. Large performance drops or shifting feature rankings would weaken claims that the system captures generalizable biology.
This test must cover the full workflow. Repeating only the neural network while changing segmentation, normalization, or reference data would make it difficult to identify why results changed.
The second signal is prospective conversion prediction. A future study should enroll cognitively normal people or patients with mild cognitive impairment, generate baseline regional profiles, and track clinical outcomes.
Researchers could then test whether AgeNet-SHAP features predict conversion beyond chronological age, cognition, hippocampal volume, and established biomarkers.
That comparison is critical. A statistically significant association has limited value if it adds no useful information to existing assessments.
Prospective work should report sensitivity, specificity, calibration, and performance across practical time horizons. It should also explain how missing scans, scanner changes, and participant dropout affect the results.
The third signal is decision-level benefit. Even an accurate risk estimate matters only if it improves a real clinical choice.
Researchers must identify the intended use before validation. Possible uses include selecting participants for trials, prioritizing biomarker testing, tracking structural change, or supporting specialist assessment.
Each use requires a different threshold and tolerance for error. A research-enrollment tool can accept tradeoffs that would be inappropriate for screening asymptomatic adults.
The strongest future system may not rely on MRI alone. Structural patterns could be combined with cognitive assessments and molecular biomarkers while retaining regional interpretability.
That combination would also test whether AgeNet-SHAP captures distinct information. If its regional features merely restate hippocampal atrophy or clinical severity, a simpler measure may be preferable.
Researchers should also publish explanation-stability analyses. A regional map should not change dramatically because of a minor preprocessing variation or a slightly different reference sample.
Documentation will matter as much as model architecture. Clinical teams need clear definitions of the training population, exclusion criteria, scanner requirements, uncertainty, and failure modes.
The paper’s authors present their approach as a foundation for diagnostics, prognostics, disease-progression research, and personalized medicine. Those are future applications, not completed validations.
For developers, the lesson extends beyond neuroimaging. Explainability tools can reveal how a model constructs an answer, but they do not automatically validate the answer’s real-world meaning.
For knowledge workers following medical AI through Google News, the safest workflow is to preserve the headline, primary paper, cohort details, and stated limitations together. A searchable AI knowledge base can help keep those layers connected instead of reducing a study to one claim.
The next regional brain-age headline should prompt three questions. Was the model tested outside its original cohort? Did it predict a future outcome? Did it improve a decision beyond existing measures?
Until all three answers are clear, AgeNet-SHAP is best understood as an interpretable research map. It shows how distributed anatomy contributes to an AI-generated age estimate and how those patterns relate to Alzheimer’s severity.
That is a meaningful step. It is not yet a personal Alzheimer’s forecast. Readers, clinicians, and AI builders should judge the next wave of Google News coverage by whether the evidence finally closes that gap.


