top of page

HHS AI Clinical Trials Push Targets Fixed Phases With Adaptive Evidence

7 days ago
10 min read

HHS launched a five-year program that challenges the fixed phases, duplicated infrastructure, and delayed decisions built into conventional drug trials. The HHS AI clinical trials initiative combines predictive models, continuous analysis, and automated operations under one research framework.

The centerpiece is SURPASS, short for Simulation-augmented, Real-time Platform Adaptive Seamless Trials. ARPA-H, the Advanced Research Projects Agency for Health, will operate the program alongside three supporting projects for trial sites, health data, and patient navigation.

The immediate promise is speed. The harder task is proving that adaptive software can produce evidence as dependable as the processes it replaces. That conflict puts regulators, drug sponsors, clinical research organizations, trial sites, and statistical teams under pressure at the same time.

HHS AI Clinical Trials Move Beyond Fixed Phases

SURPASS is not simply adding an AI assistant to an existing trial. It proposes a different operating model for generating clinical evidence.

HHS announced SURPASS on September 30, 2026. According to the department’s clinical trial program, the initiative will combine predictive computational models, shared infrastructure, common control groups, and real-time analysis.

The program is designed to run for five years. It is expected to begin soliciting proposals later in the fall from teams spanning statistics, AI, trial design, operations, and regulation.

HHS did not disclose the program’s funding in its initial announcement. It also did not explain what regulatory flexibility successful participants would receive, according to the first SURPASS coverage.

Those omissions matter because SURPASS reaches beyond administrative automation. Its three technical areas address trial design, statistical inference, and operations.

The first area is a phaseless design engine. It would use digital twins and related predictive models to simulate possible clinical and operational outcomes before a study begins. A digital twin is a computational representation of a patient or process used to explore likely outcomes.

The word “phaseless” does not mean clinical testing would lose its safety gates. It describes a more continuous structure that can reduce rigid handoffs between traditionally separated trial phases.

The second area is a continuous inference engine. It would analyze accumulating data using statistical rules designed to remain valid during repeated reviews. Researchers could then add, stop, or adjust treatment arms without waiting for one fixed analysis date.

The third area is an agentic operations layer. In this context, an agent is software that can complete a defined sequence of operational tasks with limited human direction. HHS wants these systems to support trial setup, data cleaning, treatment-arm onboarding, and dataset construction.

SURPASS therefore combines several established ideas rather than relying on a single new model. Adaptive trials already permit predetermined changes based on accumulating evidence. Platform trials can test several treatments under a shared protocol. Computational models already inform some drug-development decisions.

The proposed change is to place those elements inside one persistent system. That system would learn during a trial instead of treating design, recruitment, analysis, and reporting as mostly separate stages.

HHS says drug and biologic development often lasts more than a decade, can cost up to $2 billion on average, and fails more than 90 percent of the time. These figures came from the agency’s announcement and should be understood as its justification for the program.

The useful question is not whether the current process is slow. Few participants dispute that point. The question is whether a continuous platform can remove delay without moving uncertainty into harder-to-audit models.

That distinction creates the article’s central tension. Faster operations are valuable, but clinical evidence must remain interpretable after a trial changes direction.

Four Projects Attack Four Different Bottlenecks

HHS is treating slow trials as a systems problem involving design, sites, data, and patients, not just a shortage of better algorithms.

SURPASS covers trial design and execution. Three complementary ARPA-H projects address conditions that determine whether an adaptive study can operate at national scale.

STACK targets clinical site capacity. HHS says the project will use AI to accelerate site activation and help research-naive health facilities become capable trial sites.

Site activation includes contracts, staff preparation, technology setup, regulatory review, and sponsor qualification. A faster statistical design provides little benefit if participating hospitals take months to begin enrolling patients.

STACK also carries a geographic argument. More qualified sites could let patients join studies closer to home rather than traveling repeatedly to a major academic center.

That expansion introduces a practical test. New sites must collect consistent data, protect participants, and follow complex protocols. Automating paperwork cannot substitute for experienced investigators, trained coordinators, or reliable clinical systems.

COMMONS addresses data infrastructure. HHS describes it as a privacy-by-design system for consent and regulatory-grade data access.

Regulatory-grade data must be traceable, controlled, and suitable for decisions about treatment safety and effectiveness. Ordinary data aggregation does not meet that standard automatically.

COMMONS is intended to connect information needed for enrollment, predictive modeling, and trial operations. If implemented well, it could reduce repeated data requests and make eligible patients easier to identify.

It could also expose one of the initiative’s biggest implementation challenges. Hospitals and research networks store information in different formats, under different governance rules, and with uneven data quality.

Consent presents another challenge. A patient’s permission to contribute information must remain understandable when data moves across institutions or supports new computational analyses.

CINCH focuses on the patient side. The project is intended to help people contribute real-world data, coordinate their care, and identify trials that fit their circumstances.

Real-world data comes from routine care and daily life rather than a tightly controlled study visit. It can include electronic health records, claims, laboratory results, registries, and information collected by connected devices.

Those sources can widen the picture of treatment effects. They can also contain missing information, inconsistent measurements, and patterns created by unequal access to care.

CINCH must therefore do more than match keywords in a patient record. A useful navigation system must explain why a study appears relevant and keep clinicians involved in eligibility decisions.

Together, SURPASS, STACK, COMMONS, and CINCH form a layered program. SURPASS redesigns the trial. STACK expands where it can run. COMMONS supplies governed information. CINCH connects patients with the system.

This structure separates the announcement from narrower clinical AI projects. HHS is not promising that one prediction model will select successful drugs. It is trying to create shared machinery for repeated evidence generation.

That broader ambition also increases execution risk. The four projects depend on one another, but they involve different institutions, technical standards, and legal obligations.

A failure in any layer can limit the others. Better matching cannot repair an unavailable site. A faster site cannot compensate for unusable data. A strong model cannot rescue a poorly governed protocol.

The Real Contest Is Speed Versus Verifiable Evidence

The primary conflict is not AI against human researchers. It is the promise of continuous adaptation against the need for stable, reproducible conclusions.

Traditional phase boundaries create delays, but they also organize responsibility. Sponsors know which evidence must exist before progressing. Reviewers can identify when a protocol changed and which dataset supported a decision.

A continuous platform weakens those familiar boundaries. That can eliminate duplicated work, but it makes statistical planning and auditability more important.

Adaptive platform trials illustrate the tradeoff. They can test several interventions under one master protocol, remove ineffective arms, add new candidates, and share control groups.

During the COVID-19 pandemic, platform trials produced important treatment evidence while conventional programs struggled with speed. Their success showed that adaptable infrastructure can answer urgent questions efficiently.

However, adaptation must be specified before investigators see results that could influence their choices. Unplanned changes can create bias, inflate false-positive risk, or make a treatment effect harder to interpret.

A major platform trial review describes the design, conduct, oversight, and reporting requirements needed for these studies. Its central lesson is that flexibility requires more planning, not less.

Shared controls need similar discipline. A common control group can reduce the number of participants assigned to standard treatment and improve efficiency across several comparisons.

Problems emerge when a new treatment arm is compared with patients enrolled much earlier. Medical practice, disease patterns, diagnostic methods, and patient characteristics can change over time.

Researchers call those earlier participants nonconcurrent controls. Using them can increase statistical power, but it can also introduce time-related bias.

SURPASS proposes an always-valid inference engine to address repeated analysis and adaptation. HHS has not yet published the engine’s statistical specification, validation benchmarks, or regulatory acceptance criteria.

The same evidence burden applies to digital twins. A model can simulate likely outcomes before enrollment, helping teams test assumptions or identify weak protocol choices.

A simulation remains an estimate. It reflects the patients, variables, and relationships represented in its training data.

If a digital twin underrepresents older adults, rural patients, pregnant people, or patients with multiple conditions, the resulting trial may become efficient around an incomplete population.

Validation must therefore examine more than average prediction accuracy. Sponsors and regulators will need to know where a model fails, how uncertainty is calibrated, and whether performance changes across sites.

They will also need a process for model updates. A continuously improving model sounds attractive, but changing software during a trial can complicate reconstruction of past decisions.

Every consequential recommendation should be linked to a specific model version, input dataset, rule, and human approval. Without that record, speed can weaken accountability.

HHS has already stated that AI-supported biomedical research should emphasize reproducibility, transparent methods, and unbiased review. Its AI strategy also identifies regulatory submissions supported by AI evidence as a measurable outcome.

SURPASS will test whether those principles survive contact with an operational platform. Public documentation will be especially important because HHS wants the resulting tools, standards, and regulatory examples to benefit the wider research system.

The goal should not be to make every trial adaptive. Some questions fit stable, conventional designs. Others involve outcomes that cannot be measured quickly enough to guide real-time changes.

Long-term survival, delayed adverse effects, and slowly developing conditions do not become immediate because software analyzes incoming data faster.

The design must follow the medical question. Otherwise, the program risks treating adaptation as a target rather than a method.

Operation TrialBlazer Raises the Competitive Stakes

SURPASS is part of a larger federal attempt to make the United States a faster place to develop medicines, especially as China’s clinical research capacity grows.

HHS introduced Operation TrialBlazer in June 2026 as a department-wide clinical research initiative. It coordinates work across ARPA-H, FDA, NIH, the National Cancer Institute, and federal health technology officials.

The initiative’s political argument is direct. HHS says early-stage research has moved overseas and that the United States risks losing associated investment, expertise, and drug-development activity.

China provides the most visible comparison. IQVIA data showed that Chinese companies’ share of global trial starts increased from 1 percent in 2009 to 30 percent in 2024.

The United States still holds major advantages, including a large pharmaceutical market, experienced regulators, leading research hospitals, and deep biotechnology financing.

Yet trial speed affects where companies test products and where investors place capital. A scientific program that moves slowly can lose opportunities even when its institutions remain respected.

SURPASS attacks design and operational delay. Other TrialBlazer measures focus on regulatory review and submission requirements.

The FDA has launched an Expedited Investigational New Drug pilot for first-in-human studies. Selected sponsors can work with qualified research institutions and submit components of an application on a rolling basis.

Rolling submission allows regulators to review completed sections before the entire package is ready. FDA also wants some institutional review and site activation work to proceed alongside application development.

The agency says clearer phase-appropriate chemistry and manufacturing expectations can save companies six to 12 months. Those estimates appear in its development roadmap, which also covers quantitative dose selection and late-stage master protocols.

For later development, FDA has clarified when one rigorous pivotal investigation plus confirmatory evidence can support approval. It has also revised draft guidance for basket, umbrella, and platform trials.

These policies matter because ARPA-H cannot create regulatory acceptance by itself. SURPASS can fund tools and demonstrations, but FDA determines whether resulting evidence supports a medical product decision.

The relationship between the agencies will shape adoption. Sponsors will hesitate to rebuild trial programs around a new system without clear evidence that regulators can evaluate its outputs.

The opposite risk also exists. A policy drive for faster domestic development could create pressure to interpret uncertainty generously.

The strongest defense is transparent validation. Faster review should mean earlier coordination, clearer requirements, and less duplicated work. It should not mean weaker standards for safety or effectiveness.

TrialBlazer’s federal initiative explicitly says modernization should preserve scientific rigor. That principle now needs observable implementation.

This is why the program’s success should not be measured only by shorter timelines. HHS must also track protocol deviations, representativeness, adverse-event detection, reproducibility, and agreement between model predictions and clinical outcomes.

Domestic competitiveness adds urgency, but it does not resolve the evidence problem. The United States gains little from faster trials if other regulators, clinicians, or patients do not trust their conclusions.

What HHS Must Prove Next

The next milestones must show that SURPASS can produce validated tools, regulator-ready evidence, and broader patient access at the same time.

The first signal is the structure of ARPA-H’s solicitation. Applicants need measurable objectives for speed, cost, participant burden, statistical validity, and population coverage.

A solicitation centered mostly on automation features would weaken the program’s claim. A strong one will require prospective validation against conventional approaches and clear failure criteria.

Teams should also state how models will be audited across demographic groups and participating sites. Aggregate accuracy is not enough for decisions that determine who enters a trial or which treatment continues.

The second signal is regulatory alignment. FDA participation must advance from general support for innovative designs to concrete methods for evaluating SURPASS outputs.

Sponsors need guidance on acceptable digital-twin validation, model changes, shared controls, continuous analysis, and documentation. They also need to know when a conventional design remains necessary.

Public examples would help. A regulator-reviewed protocol, simulation plan, and analysis framework could make the program useful beyond its funded teams.

Regulatory alignment would strengthen the initiative if it produces clear standards before large trials depend on them. Persistent ambiguity would slow adoption, regardless of the software’s technical performance.

The third signal is operational evidence from STACK, COMMONS, and CINCH. HHS should report whether new sites activate faster, whether they recruit representative populations, and whether patients can understand the consent process.

A larger site count is not automatically meaningful. The sites must enroll participants, maintain data quality, retain staff, and protect patients.

COMMONS should be judged by the availability of traceable regulatory-grade data, not the volume of records connected. CINCH should be judged by appropriate matches, completed enrollment, and patient comprehension.

These measures would reveal whether HHS AI clinical trials expand access or simply accelerate work at institutions already equipped to participate.

The program should also publish negative findings. A digital twin that performs poorly in one disease area can still provide valuable evidence about where the approach does not belong.

That transparency would help distinguish scientific infrastructure from a technology procurement campaign. It would also support the reusable standards HHS says it wants to create.

For clinical teams, the near-term task is preparation. Organizations considering participation should map their data provenance, model governance, consent language, and decision logs now.

Knowledge from statisticians, clinicians, data engineers, regulatory specialists, and site operators must remain connected throughout the study. Teams can use a searchable knowledge base to preserve assumptions, decisions, and model documentation across those groups.

The initiative’s most important outcome will not be a claim that AI made trials faster. It will be evidence that researchers reached dependable answers sooner, with fewer unnecessary burdens and no hidden loss of rigor.

HHS has defined the ambition. ARPA-H must now show its methods, FDA must clarify the evidentiary path, and funded teams must publish results that outsiders can reproduce. Will the first SURPASS projects expose their validation rules as clearly as they promote their speed?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page