top of page

SAFE-WAIT Sepsis AI Trial Expands, but Five NSW Hospitals Will Test Its Real Value

Sep 15
12 min read

NSW Health is expanding the SAFE-WAIT sepsis AI trial to five emergency departments after a pilot accelerated antibiotics for certain high-risk patients by 83 minutes.

The rollout moves SAFE-WAIT beyond its original six-month trial at Westmead Hospital. It will test whether a locally developed clinical model remains useful across different hospitals, patients, and working conditions.

That expansion is the real story. Sepsis algorithms often look promising inside their development hospitals, yet disappoint when exposed to new populations and clinical workflows. SAFE-WAIT now faces that harder test.

The project also targets an unusual gap in emergency care. Many deterioration systems monitor patients after admission, when laboratory results and repeated observations are available. SAFE-WAIT focuses on people still waiting for their first medical review.

Its primary opponent is not another AI vendor. It is the conventional threshold-based alert, which generally reacts after predefined clinical signs appear. SAFE-WAIT aims to identify risk earlier without overwhelming nurses with false alarms.

The SAFE-WAIT Sepsis AI Trial Moves Beyond Westmead

NSW Health is turning a promising single-hospital pilot into a two-and-a-half-year evaluation across five emergency departments.

The participating sites are Westmead, Auburn, Blacktown, Mount Druitt, and Nepean hospitals. The expanded project is scheduled to begin at the start of 2027.

SAFE-WAIT received almost A$500,000 through the state’s Translational Research Grants Scheme. The project shares a broader A$4.9 million government investment supporting high-impact medical research.

Emergency physician Associate Professor Amith Shetty led the NSW Health clinical team that developed the system. Researchers from Western Sydney University, the University of Sydney, and other health organizations have contributed to its evaluation.

The model uses information recorded during a patient’s initial triage assessment. Inputs include demographic details, vital signs, and relevant risk factors available before a physician examines the patient.

It then assigns a low, moderate, or high sepsis risk. A dashboard presents those categories through a traffic-light display, giving nurses a prioritized view of the waiting room.

A high-risk result does not diagnose sepsis. It prompts clinical reassessment, which can include repeated observations, further investigation, escalation, fluids, or medication when clinically appropriate.

That distinction matters because sepsis is difficult to recognize early. Sepsis occurs when the body’s response to infection causes organ dysfunction, but its earliest signs can resemble less dangerous conditions.

Traditional confirmation can also depend on information that is unavailable during triage. Laboratory results, imaging, and a physician’s complete assessment often arrive after the patient has already waited.

SAFE-WAIT attempts to use the sparse information available at arrival. It recalculates risk when new vital signs enter the system, allowing staff to follow changes while a patient remains unreviewed.

According to the NSW Health announcement, the initial Westmead pilot ran for six months during 2024. The agency says staff identified and treated some high-risk patients earlier.

The upcoming evaluation has a more demanding purpose. It must show that SAFE-WAIT can operate reliably outside the hospital where the live pilot occurred.

That shift changes the evidence standard. The question is no longer whether the dashboard can run or produce encouraging associations at Westmead. It is whether other emergency departments can obtain consistent clinical value.

Why the Waiting Room Is the Critical Test

SAFE-WAIT concentrates on the period when clinical information is thinnest, patient volume is visible, and deterioration can remain hidden.

Westmead Hospital recorded more than 80,000 emergency attendances during the year cited by Western Sydney Local Health District. Its waiting room can contain 40 to 60 patients simultaneously.

A triage nurse must identify the most urgent cases from that crowded field. Existing triage categories help allocate priority, but sepsis does not always present with an immediately obvious pattern.

Some patients arrive with clear instability and receive rapid attention. Others appear less urgent before their condition changes, creating a monitoring problem while they await medical review.

The Western Sydney account says SAFE-WAIT identified high-risk patients about 90 minutes before traditional sepsis triggers showed corresponding symptoms.

The system’s developers designed it as decision support for nurses, not autonomous clinical control. Its practical value therefore depends on what happens after a risk category appears.

A red signal has limited value if staffing shortages prevent reassessment. It can also be harmful if frequent false alerts pull nurses away from patients with other urgent conditions.

The first pilot reported a more useful operational signal. High-risk patients in triage categories two and three received antibiotics 83.6 minutes earlier than comparable low-risk patients after adjustment.

That result aligns with NSW Health’s public summary of an average 83-minute improvement. However, the detailed study contains an important complication.

High-risk patients placed in the less urgent triage categories four and five received antibiotics 186.37 minutes later. The researchers identified a significant interaction between triage level and SAFE-WAIT risk.

This finding does not establish that SAFE-WAIT caused those delays. The study was observational, and treatment timing reflects disease presentation, triage decisions, workflow, and clinician judgment.

It does show why an average can conceal important differences. An effective warning system must help the patients most likely to be overlooked, not only those already receiving urgent triage classifications.

The pilot also associated high-risk status with physician review 3.5 minutes earlier. That change was statistically significant, although its practical importance is smaller than the antibiotic result.

For moderate-risk patients in triage categories four and five, physician review took 10.62 minutes longer. Again, the interaction with existing triage decisions deserves attention during expansion.

These mixed effects make the multicenter trial necessary. A dashboard can only improve care when nurses understand it, trust it appropriately, and can act without disrupting other priorities.

How SAFE-WAIT Challenges Traditional Sepsis Alerts

The model trades the precision of a strict rule for earlier and more sensitive detection, creating both its strongest advantage and its main operational risk.

Traditional sepsis alerts often use fixed thresholds. A rule can activate when specific vital signs, laboratory values, or organ dysfunction criteria cross predefined limits.

That approach is transparent, but it can miss nonlinear relationships. A collection of individually modest abnormalities can indicate risk even when no single measurement crosses a strict threshold.

SAFE-WAIT uses machine learning to combine available triage data. Its system architecture processes blood pressure, pulse, respiratory rate, temperature, oxygen saturation, demographics, and electronic medical record data.

The pilot architecture used an event-driven cloud design for near-real-time inference. Its dashboard remained separate from the hospital’s electronic medical record, which introduced an extra interface into the nursing workflow.

Researchers compared SAFE-WAIT with the Western Sydney Local Health District’s existing sepsis alert. That conventional alert used criteria related to systemic inflammatory response syndrome, commonly called SIRS.

In the reported analysis, SAFE-WAIT reached a sensitivity of 0.86. Sensitivity measures the share of eventual sepsis cases identified by the system.

The existing alert recorded sensitivity of 0.13 in the same analysis. SAFE-WAIT therefore captured far more patients who later met the study’s confirmed sepsis definition.

That improvement came with a false-positive rate of 0.42. The conventional alert’s false-positive rate was 0.02, making it substantially more selective.

The competing approaches consequently identify different populations. The conventional alert appears to flag fewer but generally sicker patients after clearer warning signs emerge.

SAFE-WAIT casts a wider net earlier. Among patients classified as high risk, 24 percent developed confirmed sepsis, according to the project’s preprint.

That leaves many high-risk patients who did not meet the study’s eventual sepsis outcome. Some might still have required monitoring or treatment, but every alert consumes clinical attention.

The pilot analyzed 108,401 eligible patient encounters. Of these, 104,904 received only a SAFE-WAIT risk category, 1,208 triggered only the existing sepsis alert, and 2,289 triggered both systems.

Older patients were more frequently classified as high risk. Thirty-eight percent of patients aged 65 or older entered SAFE-WAIT’s high-risk group.

Most high-risk patients, 94 percent, received triage category two or three. This concentration suggests that SAFE-WAIT often agreed with nurses that a patient needed relatively prompt attention.

Still, the model added useful separation within those broad categories. Researchers reported that every 4.8 high-risk flags identified one additional patient with sepsis, septic shock, or another defined adverse outcome.

SAFE-WAIT also categorized most patients who developed septic shock as high risk at triage. The median onset of low blood pressure indicating shock followed roughly 92 minutes later.

Those results support the model’s proposed mechanism. It combines limited signals before a conventional threshold reveals obvious deterioration.

Yet prediction and intervention are different achievements. A better risk score does not automatically prove fewer deaths, shorter hospital stays, or safer antibiotic use.

The Pilot Shows Speed, Not Yet Better Outcomes

SAFE-WAIT has evidence of earlier recognition and treatment, but it has not yet established a causal improvement in patient outcomes.

The most detailed public evaluation is a SAFE-WAIT preprint. A preprint has not completed the peer-review process, so its estimates require cautious interpretation.

The study compared patients across SAFE-WAIT risk categories and against the existing alert. Researchers adjusted several treatment analyses for factors including age, gender, presentation time, and illness severity.

Moderate-risk patients received antibiotics 58.81 minutes earlier than low-risk patients in the adjusted analysis. High-risk patients received them 116.99 minutes earlier.

Those figures describe associations between assigned risk and treatment timing. They do not come from random assignment to AI-supported care versus usual care.

Sicker patients can receive faster treatment because clinicians independently recognize their condition. The model can also influence care, making it difficult to separate prediction quality from staff response.

The study reported higher adjusted odds of intensive care admission and in-hospital death among high-risk patients. That pattern indicates successful risk stratification, not harm caused by the model.

High-risk status carried adjusted odds of 1.77 for intensive care admission. It carried adjusted odds of 2.26 for in-hospital mortality.

A model that correctly identifies vulnerable patients should place more eventual adverse outcomes in its high-risk group. However, those associations do not show that its alerts prevented any outcome.

The expanded trial should therefore measure more than accuracy. Treatment timing, escalation rates, intensive care use, hospital length of stay, mortality, and unintended antibiotic exposure all matter.

Antibiotic speed requires careful interpretation. Early antimicrobial treatment is central to sepsis care, but unnecessary antibiotics carry individual and population risks.

A false-positive alert should not automatically trigger medication. NSW Health describes SAFE-WAIT as a prompt for nursing action and clinical reassessment, preserving human review before treatment.

Workflow evidence matters equally. The original implementation used a separate read-only web application rather than full electronic record integration.

That design can support controlled testing, but it asks clinicians to monitor another screen. Usage rates, missed alerts, response times, and overrides should form part of the multicenter evaluation.

The technical paper also reported no implemented monitoring for data drift during the pilot. Data drift occurs when incoming patient information changes from the data used to develop the model.

Different hospitals can introduce such changes immediately. Their patients, documentation practices, equipment, staffing patterns, and emergency workloads do not perfectly match Westmead’s.

A system can therefore retain its software function while losing clinical accuracy. Continuous monitoring and planned recalibration are essential for detecting that decline.

Latency presents another risk. SAFE-WAIT depends on information moving from triage systems through an inference pipeline and back to the dashboard.

Delayed or incomplete records can make an apparently real-time score clinically stale. The architecture paper acknowledged that processing latency and high data volumes could leave cases underestimated.

These are not reasons to reject the project. They are the implementation questions that separate a research model from dependable hospital infrastructure.

Sepsis AI Has a Difficult Record Outside Its Home Hospital

The history of clinical sepsis algorithms shows why external validation matters more than an impressive result at one site.

The best-known caution involves the Epic Sepsis Model, a proprietary system deployed across hundreds of United States hospitals.

An independent JAMA validation examined 38,455 hospitalizations at Michigan Medicine. Sepsis occurred in 2,552 of them.

At the evaluated threshold, the model missed 67 percent of sepsis cases. It also generated high-risk scores for 18 percent of all hospitalizations.

Its hospitalization-level area under the receiver operating curve was 0.63. This measure reflects how effectively the model ranks patients with and without the outcome.

The Epic model identified only 7 percent of sepsis patients who had not already received timely antibiotics. Clinicians would have evaluated eight alerted patients to find one eventual sepsis case.

That failure illustrates two recurring problems. Models can perform worse outside their development setting, and excessive alerts can weaken staff attention.

SAFE-WAIT differs from Epic’s model in important ways. It targets emergency waiting rooms, uses information available around triage, and presents graded risk rather than an isolated diagnostic declaration.

It was also developed inside the health system conducting the trial. The published architecture describes its inputs and implementation more openly than many proprietary hospital algorithms.

Those distinctions do not remove the validation problem. They make the five-site rollout the correct next experiment.

A 2026 systematic review evaluated 53 studies covering more than seven million patient admissions. Its findings capture both promise and uncertainty.

The best-performing machine learning models reached a pooled area under the curve of 0.88. Pooled sensitivity was 77.2 percent, while specificity reached 84.7 percent.

However, positive predictive value was only 34.2 percent. In practical terms, many patients receiving positive predictions did not develop the defined sepsis outcome.

The review also found extreme variation among studies and a high risk of bias in most included research. Its authors called for standardized validation and prospective pragmatic trials.

That evidence places SAFE-WAIT in a broader transition. Healthcare has many retrospective sepsis models, but far fewer systems tested prospectively inside ordinary clinical workflows.

Australia has moved particularly cautiously. A 2024 clinical AI review described most Australian hospitals as clinical AI-free outside medical imaging.

The authors cited trust, privacy, bias, governance, and regulation among the barriers. They also argued that structured implementation could let hospitals test clinical AI without skipping safety controls.

SAFE-WAIT represents that implementation approach. It began with a silent evaluation, progressed to a live pilot, and now moves toward multicenter testing.

The sequence is important. A silent trial calculates scores without exposing them to clinicians, allowing researchers to assess behavior before alerts influence care.

A live pilot then reveals workflow effects. Multicenter deployment tests transportability, which measures whether performance survives changes in people, settings, and data.

The NSW expansion will not automatically settle every issue. Five hospitals in one state still represent a narrower environment than national deployment.

It can provide something more valuable than another benchmark, however. It can show whether a locally built model creates repeatable clinical improvements under different operational pressures.

Five Hospitals Will Test Reliability, Workload, and Equity

The next phase must determine whether SAFE-WAIT remains accurate, actionable, and fair when its operating environment changes.

The first signal to watch is performance by hospital. Sensitivity, specificity, positive predictive value, and false-alert volume should be reported separately for each site.

Similar results across Westmead, Auburn, Blacktown, Mount Druitt, and Nepean would strengthen the case for transportability. Wide variation would suggest local recalibration or workflow redesign.

The evaluation should also publish subgroup performance. Age already influenced high-risk classifications, and other patient characteristics can change model behavior.

Relevant analysis includes sex, language background, comorbidities, pregnancy status, and Indigenous status when governance and sample sizes permit responsible reporting.

A model can show acceptable overall accuracy while serving smaller groups poorly. Hospitals need confidence that one population does not absorb more missed cases or unnecessary alerts.

The second signal is the relationship between alerts and action. Researchers should track how often nurses view scores, repeat observations, request medical review, or escalate care.

Those measures can reveal whether the system changes behavior. They can also identify automation bias, which occurs when users defer too readily to a computerized recommendation.

A low-risk classification deserves special scrutiny. Nurses must retain authority to escalate care when clinical judgment conflicts with the dashboard.

Likewise, a high-risk result should initiate reassessment rather than automatic treatment. This structure limits unnecessary intervention while preserving the benefit of earlier attention.

Alert burden will be decisive. The Westmead analysis reported a 0.42 false-positive rate, which is substantial even when higher sensitivity is clinically desirable.

The practical question is not whether every false positive can disappear. It is whether the resulting workload remains manageable and produces enough valuable reassessments.

Researchers should publish the number of alerts per nursing shift and per patient. They should also report response times, repeated alerts, dismissals, and periods of nonuse.

If alert volume climbs while response rates decline, the trial’s central claim weakens. Stable engagement with earlier care would strengthen it.

The third signal is patient outcome evidence. The initial headline concerns faster antibiotics, but the expanded evaluation should follow the complete clinical pathway.

Important outcomes include septic shock, intensive care admission, mortality, hospital length of stay, readmission, and antibiotic exposure among patients without confirmed sepsis.

A reduction in treatment time without better outcomes can still represent useful process improvement. It does not justify claims that the model saves lives.

Conversely, a small outcome improvement across a large emergency population can be meaningful. The trial needs enough statistical power and an appropriate comparison design to detect it.

The two-and-a-half-year duration creates an opportunity to study performance over time. It should capture seasonal infections, workload changes, staffing turnover, and evolving documentation.

It also makes drift monitoring unavoidable. Investigators should define when performance triggers review, recalibration, or temporary withdrawal.

Governance will matter alongside statistics. Clinical teams need clear responsibility for model updates, incident reporting, access controls, and communication with patients.

The system handles sensitive health information, even if clinicians see only a risk display. Data movement, retention, security, and vendor access require documented controls.

Transparent reporting would make this trial relevant beyond NSW. Hospital leaders elsewhere need implementation evidence, not just another claim about algorithmic accuracy.

The strongest result would combine reliable cross-site performance with manageable alert volume and measurable improvements in care. Anything less should narrow the system’s intended use.

What Comes Next for SAFE-WAIT

Three developments will determine whether SAFE-WAIT becomes durable clinical infrastructure or remains an encouraging Westmead experiment.

First, watch the site-by-site validation after the 2027 restart. Consistent sensitivity and treatment effects across all five hospitals would support broader deployment.

A large decline outside Westmead would weaken the model’s central advantage. It would indicate that local data, staffing, or workflow contributed heavily to the original result.

Second, watch how the teams manage false positives and nursing workload. A useful system must earn attention repeatedly without competing destructively with other emergency priorities.

Public reporting should connect each risk category with clinical actions and outcomes. That evidence will reveal whether the traffic-light interface changes care appropriately.

Third, watch whether the trial demonstrates patient benefit beyond faster antibiotics. Lower rates of septic shock or other adverse outcomes would materially strengthen the case.

The absence of an outcome difference would require a narrower conclusion. SAFE-WAIT might still improve monitoring or care processes, but it should not be presented as a proven lifesaving intervention.

The trial’s most important contribution may be its testing discipline. Clinical AI needs staged evaluation, human oversight, subgroup analysis, drift monitoring, and transparent failure criteria.

For developers, the lesson is that prediction accuracy is only the beginning. Integration, latency, actionability, and staff trust determine whether a model functions during real care.

For hospital buyers, the project offers a practical evaluation checklist. Ask where the model was trained, how it performs locally, and what workload each alert creates.

For patients, the central safeguard is equally direct. SAFE-WAIT should help clinicians notice deterioration, while leaving diagnosis and treatment decisions with trained professionals.

The SAFE-WAIT sepsis AI trial now has an opportunity to provide rare multicenter evidence from emergency waiting rooms. Its value will depend on results, not the promise attached to AI.

The question for the next two and a half years is measurable: can SAFE-WAIT preserve earlier recognition across five hospitals without adding unsafe treatment or unmanageable alert fatigue?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page