top of page

Retinal AI Can Predict Systemic Risk, but Prediction Is Not Diagnosis

Google News has surfaced a Cureus review arguing that artificial intelligence can turn retinal findings into warnings about diseases beyond the eye. The conflict is immediate. Researchers can extract cardiovascular, metabolic, kidney, and neurological signals from retinal images, yet most systems remain risk estimators rather than diagnostic tools.

That distinction matters because the retina offers a rare, noninvasive view of blood vessels and neural tissue. A standard fundus photograph captures the eye’s interior surface, including the optic disc, macula, and retinal vessels. Algorithms can analyze patterns that clinicians might not quantify consistently during routine examinations.

The idea already has substantial research behind it. Google researchers published a major cardiovascular study in 2018, while newer systems have targeted stroke and broader cardiometabolic risk. However, the evidence also exposes the central tension: predicting an outcome in a curated dataset does not establish clinical benefit in everyday care.

A retinal model can produce an impressive score without explaining what action should follow. It can also perform differently across cameras, clinics, age groups, and patient populations. Precision medicine begins only when a prediction improves a real decision for the person being screened.

What Google News Put Back in Focus

The important development is not a single new diagnostic product. It is the growing case for treating retinal images as systemic health data.

The Cureus review highlighted through Google News brings together a field often called oculomics. Oculomics studies relationships between measurable eye features and health elsewhere in the body. Artificial intelligence expands that approach by detecting complex image patterns that resist manual measurement.

The biological premise is credible. Retinal vessels share characteristics with microvasculature elsewhere in the body. Changes in vessel width, branching, tortuosity, hemorrhage, or perfusion can reflect vascular injury and chronic metabolic stress.

The retina also contains neural tissue connected to the central nervous system. That relationship has encouraged research into associations with stroke, cognitive decline, and other neurological conditions. These associations do not mean an eye photograph reveals every disease directly.

Deep learning provides the computational mechanism. A deep learning model learns statistical representations from many labeled images instead of relying only on predefined vessel measurements. During training, it adjusts internal parameters to reduce errors between its predictions and known clinical outcomes.

A landmark retinal risk study trained models using data from 284,335 patients. Researchers then tested the models on independent datasets containing 12,026 and 999 patients.

The system estimated age within a mean absolute error of 3.26 years. It classified sex with an area under the receiver operating characteristic curve, or AUC, of 0.97. AUC measures how well a model separates two outcome groups across possible decision thresholds.

The model also estimated smoking status with an AUC of 0.71. Its systolic blood pressure estimate had a mean absolute error of 11.23 millimeters of mercury. For major adverse cardiac events, the reported AUC was 0.70.

Those figures established that retinal photographs contain more systemic information than conventional examination had extracted. They did not establish that the algorithm should replace a blood pressure cuff, laboratory test, or cardiovascular examination.

The paper’s attention maps suggested that the model used the optic disc and blood vessels for different predictions. Attention maps highlight image regions that influence an algorithm’s output. They offer clues about model behavior, but they are not complete causal explanations.

This is why the current Google News moment deserves careful framing. The research direction has moved beyond an isolated demonstration. Teams are now testing retinal prediction across diseases, populations, imaging devices, and clinical settings.

Yet the central claim remains narrower than some headlines imply. Retinal AI can detect statistical signals associated with systemic health. Whether those signals improve outcomes depends on validation, workflow design, and responsible clinical follow-up.

Retinal AI Moves From Risk Factors to Future Events

The field is shifting from estimating familiar risk factors toward predicting events that have not happened yet.

Early work often asked whether a retinal image could recover information already available elsewhere. Age, smoking status, blood pressure, and glycated hemoglobin provided convenient labels. Success showed that the retina carried related information, although it did not prove added clinical value.

Event prediction raises the stakes. A model that forecasts cardiovascular disease or stroke must remain accurate over time. It must also contribute information beyond established risk calculators, medical history, physical examination, and laboratory testing.

A 2025 pragmatic study evaluated automated retinal photography and AI-based cardiovascular risk assessment in two Australian primary care clinics. The system generated a retinal-predicted cardiovascular disease score from non-mydriatic fundus photographs. Non-mydriatic imaging captures retinal photographs without routinely dilating the pupil.

Among 361 participants, 339 received a risk score. That represented a 93.9 percent imaging success rate. The study also reported satisfaction of approximately 90 percent among end users.

The model’s real value, however, depended on predictive performance rather than convenience alone. In a separate UK Biobank validation cohort, the retinal score recorded an AUC of 0.672 for ten-year cardiovascular events. The conventional World Health Organization score recorded an AUC of 0.693.

Those results were broadly comparable, but neither score approached perfect discrimination. The primary care trial also found only moderate correlation between the retinal and WHO scores.

The comparison reveals an important nuance. A retinal model can approximate conventional risk assessment using a photograph, which could help where laboratory access is limited. That does not automatically make it superior to collecting the established clinical variables.

The study’s training process further complicates interpretation. Researchers trained the retinal model using WHO cardiovascular risk scores as its reference labels. Those scores incorporated age, sex, smoking, systolic blood pressure, diabetes status, and cholesterol.

The model therefore learned to reproduce a clinical risk construct before researchers tested it against future events. It was not originally trained only from adjudicated heart attacks and strokes. That difference affects what its output represents.

Stroke research has taken another route. DeepRETStroke uses retinal images to identify patterns associated with silent brain infarction, or brain injury detected by imaging without a recognized clinical stroke. The system also estimates five-year stroke risk.

Researchers pretrained DeepRETStroke with 895,640 retinal photographs. They tested silent brain infarction detection across external datasets and reported AUC values from 0.751 to 0.792.

For five-year incident stroke prediction, external-dataset AUC values ranged from 0.728 to 0.895. The stroke prediction study included data from several countries and ethnic groups, strengthening the validation design.

The wide performance range still deserves attention. It shows that results change with population, clinical context, prevalence, and data collection. A high result in one cohort cannot be treated as a universal operating specification.

Google News coverage can compress these studies into a simple promise: one photograph predicts systemic complications. The scientific record supports a more useful conclusion. Retinal images provide a scalable risk signal whose reliability depends heavily on where and how it is used.

The Real Contest Is Capability Versus Clinical Validity

Retinal AI is advancing faster as a predictive capability than as a validated clinical service.

Capability answers whether a model can identify an association. Clinical validity asks whether that association remains accurate in the intended population and setting. Clinical utility asks a harder question: does using the result improve patient care?

These stages are often blurred in public discussion. A retrospective model can receive images, produce scores, and outperform a baseline. That result remains far removed from proving fewer strokes, earlier treatment, or better long-term survival.

The distinction becomes clearer when comparing systemic prediction with autonomous diabetic retinopathy screening. The latter has a defined target, a defined patient group, and an established referral pathway. Its output addresses whether retinal disease above a specified threshold is present.

The US Food and Drug Administration classifies relevant systems as retinal diagnostic software devices. Its definition covers software that analyzes digital fundus images to identify referable retinal disease. The agency’s device classification describes these products as prescription devices.

AEYE-DS, for example, analyzes fundus images to detect more-than-mild diabetic retinopathy in adults with diabetes. Its cleared indication concerns an eye disease visible through retinal imaging. It does not grant broad authority to diagnose systemic cardiovascular or neurological disease.

EyeArt follows a similarly bounded model. Its indication covers automated detection of more-than-mild and vision-threatening diabetic retinopathy. The software checks image quality, analyzes images, and returns condition-specific results.

Systemic prediction lacks that same clinical simplicity. A high cardiovascular score could trigger blood pressure measurement, lipid testing, a primary care visit, or further imaging. Each choice changes costs, benefits, and the risk of unnecessary intervention.

The model also needs a useful decision threshold. AUC summarizes ranking performance across thresholds, but clinicians must eventually choose a specific cutoff. That cutoff determines sensitivity, specificity, false alarms, and missed cases.

A tool with moderate AUC can still help when it reaches people who otherwise receive no assessment. The same tool can add little when comprehensive clinical data already exist. Context determines whether convenience produces meaningful value.

This capability-versus-validity contest is the article’s central reversal. More diseases becoming predictable from retinal data does not make deployment easier. Every added prediction creates another requirement for evidence, interpretation, consent, and follow-up.

Health systems must decide who owns the result. An optometrist may acquire the image, while a primary care physician manages cardiovascular risk. Neurology may become involved when a stroke signal appears.

The patient also needs a clear explanation. A risk estimate is not a diagnosis, and an algorithmic alert can create anxiety without confirming disease. Communication becomes part of the system’s safety, not an optional interface detail.

That makes retinal AI less like a universal scanner and more like a clinical coordination technology. The model is only one component. Cameras, operators, records, referral rules, clinicians, and patients determine the final outcome.

What the Numbers Do Not Show

Reported accuracy can conceal weaker performance in the people who most need reliable risk assessment.

The Australian primary care study offers a direct warning. Its retinal system tended to estimate conventional scores less accurately among higher-risk individuals. Subgroup analysis identified older men as a group needing particular improvement.

The researchers connected that limitation to the development data. The UK Biobank contains a relatively healthy volunteer population with fewer high-risk participants. A model trained there can face distribution shift when deployed among sicker patients.

Distribution shift occurs when real-world patients or images differ from the model’s training data. The difference may involve age, ethnicity, disease prevalence, camera hardware, lighting, image quality, or referral patterns.

Camera changes alone can matter. The UK Biobank used a device producing 45-degree macula-centered photographs. The Australian trial used a portable camera producing 50-degree images. Different fields of view and image processing can alter the features available to the model.

Age and sex can also act as shortcuts. A study using 3,000 Qatar Biobank participants found that deep models predicted age and sex accurately. It also found that those variables mediated predictions of several cardiometabolic factors.

The Qatar Biobank study reported an age mean absolute error of 2.78 years and a sex AUC of 0.97. Blood pressure, glycated hemoglobin, relative fat mass, and testosterone showed weaker performance.

Mediation creates an interpretability problem. A model might appear to recognize a disease-related retinal signature while partly reconstructing age or sex. That can still support prediction, but it changes the biological story attached to the result.

Another risk comes from labels. If a model learns from an existing risk calculator, it can inherit that calculator’s assumptions and omissions. It might automate a conventional score without discovering a distinct retinal biomarker.

Models trained on diagnoses recorded in health systems face other biases. Access to care influences who receives testing and which conditions enter the record. The algorithm can learn patterns of healthcare delivery alongside patterns of disease.

Image quality can exclude patients unevenly. Cataracts, small pupils, poor fixation, and other ocular conditions can make photographs ungradable. A system’s headline accuracy usually applies only after quality control removes unusable images.

False positives carry downstream consequences. They can prompt laboratory testing, specialist referrals, imaging, or patient concern. False negatives can provide reassurance when established screening would have identified risk.

Calibration therefore matters alongside discrimination. A calibrated model assigns risks that match observed event rates. If a model labels a group as having ten percent risk, roughly that proportion should experience the outcome within the stated period.

Calibration can deteriorate across locations or over time. Changes in cameras, treatments, population health, or clinical documentation can shift the relationship between images and outcomes. Health systems need monitoring after deployment, not only validation before launch.

Privacy introduces a separate concern. Retinal images can encode age, sex, cardiovascular characteristics, and other sensitive attributes. Patients may consent to eye screening without expecting their photograph to support unrelated systemic predictions.

That expanded use requires transparent consent and governance. Organizations should specify which models analyze an image, what outputs enter the medical record, and who can access them. Secondary research use deserves equally clear controls.

Clinicians also need documentation that connects a score to evidence. Model cards, subgroup results, validation cohorts, camera compatibility, and known failure modes should be available during procurement. A single accuracy figure is not enough.

Teams evaluating medical evidence can preserve study details through a structured AI knowledge base. Such documentation supports review, but it cannot substitute for clinical validation or regulatory clearance.

The skeptical reading is not that retinal AI lacks value. It is that the least visible parts of deployment often determine safety. Dataset composition, calibration, consent, referrals, and monitoring matter as much as the neural network.

Precision Medicine Needs a Care Path, Not Just a Score

A retinal prediction becomes precision medicine only when it changes an appropriate decision for a specific patient.

The strongest near-term use case may involve opportunistic screening. People already visit optometry clinics, diabetes programs, and primary care offices that use fundus cameras. One image could support eye screening and a carefully validated systemic risk assessment.

That approach can reduce collection friction. Retinal photography is noninvasive and can be automated. Portable equipment may extend assessment into settings where laboratory testing or specialist access remains limited.

However, convenience should not determine the medical claim. A clinic must define the eligible population, image protocol, output, and follow-up. It must also specify when conventional testing overrides the algorithm.

Consider a patient receiving diabetes eye screening. The camera captures gradable retinal photographs, and an approved system assesses diabetic retinopathy. A research-stage model might separately estimate cardiovascular risk from the same images.

The second output should not quietly appear as a diagnosis. The clinic would need an established threshold, patient consent, clinician review, and a response plan. That plan might begin with standard cardiovascular measurements rather than immediate specialist referral.

The care pathway should also address disagreement. A retinal score may classify a patient as high risk while a conventional calculator gives a lower result. Clinicians need evidence explaining whether the retinal model adds independent information.

Multimodal systems may eventually resolve part of that problem. These models combine retinal features with demographics, medical history, laboratory results, or genetic information. The goal is to use complementary signals instead of forcing one image to answer every question.

A multimodal risk study examined retinal, genomic, and clinical information in people with type 2 diabetes. Its premise was that each data source offers a different view of cardiovascular risk.

This direction fits precision medicine better than the universal-eye-scan narrative. Retinal data can contribute another phenotype, meaning an observable biological characteristic. It does not need to replace established measurements to become useful.

The best system may therefore ask a narrower question. Does adding retinal information improve decisions for patients whose conventional risk remains uncertain? That target supports clearer trials and more defensible clinical action.

Prospective impact studies are essential. Researchers should compare care with and without retinal AI, then measure referrals, completed follow-up, treatment changes, adverse events, and patient outcomes. Model accuracy alone cannot answer those questions.

Economic and operational effects matter too, even without discussing specific prices. A high false-positive rate can overwhelm referral capacity. A low image success rate can burden staff and exclude patients with common ocular conditions.

Workflow integration must avoid alert fatigue. Clinicians already receive many electronic warnings, and poorly targeted alerts are often dismissed. Retinal risk results should appear only when they support a defined decision.

The output should communicate uncertainty. A risk range, confidence information, and main limitations can be more responsible than a dramatic red warning. Patients should know whether the tool has been validated for people like them.

Human review will remain important even when image acquisition becomes automated. Clinicians must consider symptoms, family history, medication, prior disease, and contradictory evidence. The algorithm cannot see every relevant part of the patient’s condition.

Retinal AI also creates a coordination opportunity for eye-care professionals. Optometrists and ophthalmologists already identify signs linked with hypertension, diabetes, and vascular disease. Validated tools could standardize escalation while preserving clinical judgment.

Primary care teams would remain central because systemic risk management extends beyond the image. They can confirm measurements, assess competing conditions, and choose prevention strategies. A prediction without that bridge has limited value.

The field should judge success through completed care, not generated scores. A system that identifies risk but loses patients during referral has failed operationally. One that supports timely confirmation and appropriate treatment has delivered clinical utility.

Three Signals to Watch After the Google News Attention

The next phase will be decided by external validation, prospective outcomes, and tightly defined regulatory claims.

First, watch for validation across genuinely independent health systems. The strongest studies will include different countries, ethnic groups, disease prevalence, and camera models. They should publish subgroup performance, calibration, image rejection rates, and decision thresholds.

Multi-country validation strengthened the DeepRETStroke evidence, but its reported performance still varied across datasets. That variation should be treated as useful information. It shows where the model travels well and where adaptation remains necessary.

A credible deployment study should also lock the model before evaluation. Researchers should avoid repeatedly adjusting it against the final test population. Independent replication by teams outside the developer organization would strengthen confidence further.

If broad external performance remains stable, the case for retinal biomarkers becomes stronger. If results fall sharply in older, high-risk, or underrepresented groups, the field must narrow its intended uses.

Second, watch for prospective trials that measure patient outcomes or concrete care changes. Researchers need to show more than correlation with an existing score. They should track whether retinal alerts lead to completed assessments and beneficial interventions.

Relevant measures include confirmed diagnoses, appropriate referrals, treatment initiation, and event reduction. Studies should also report unnecessary testing, missed cases, patient anxiety, and clinician workload.

Randomized or carefully controlled designs would be especially informative. They can separate the effect of the algorithm from the effect of simply giving clinics more attention and resources.

If retinal screening improves follow-up without creating excessive false alarms, the precision medicine argument gains support. If it only adds another score to the medical record, enthusiasm should cool.

Third, watch the wording of regulatory submissions and clinical guidance. The FDA’s existing retinal AI examples focus primarily on eye disease, including diabetic retinopathy. Systemic prediction requires different intended-use statements and evidence.

The agency maintains a public list of AI medical devices, but inclusion does not imply approval for every possible use. Each device has a specific indication and clinical context.

A narrow claim may prove more realistic than a broad systemic screening label. Developers might target a defined population, disease, risk horizon, camera, and referral pathway. That specificity can make evaluation and safe implementation easier.

Google News will likely keep surfacing studies that connect retinal images with additional conditions. Readers should look past the number of predicted diseases. The decisive questions concern validation, incremental value, and clinical action.

For developers, the challenge is building traceable systems that behave consistently across changing data. Enterprise buyers need evidence tied to their own population and hardware. Clinicians need outputs that support a defined next step.

Patients need something simpler: an honest explanation of what the image can reveal and what it cannot. They should know that a retinal risk signal is neither a confirmed diagnosis nor a substitute for comprehensive care.

The emerging evidence supports cautious optimism. Retinal images contain meaningful systemic signals, and artificial intelligence can extract some of them at scale. Several studies now extend beyond basic risk factors into cardiovascular events and stroke.

Yet prediction remains the beginning of the workflow. Precision medicine requires validated decisions, responsible communication, and follow-up that improves outcomes. Without those elements, the eye may become a window into risk without opening a path to better care.

As more retinal AI stories appear through Google News, ask three questions. Was the model validated outside its development setting? Did its use improve a patient decision? Does its regulatory claim match the headline?

Those checks separate an interesting research result from a dependable medical tool. They also provide the clearest way to follow this field without confusing algorithmic confidence with clinical certainty.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page