Trump Medicare AI Pilot Promised Faster Reviews. WISeR’s Denials Tell a Different Story
The Trump Medicare AI pilot began reviewing selected services across six states in January 2026, promising faster decisions and less waste. Nine months later, WISeR faces a harder question. Can software help protect Medicare without turning financial savings into delayed or denied care?
The Wasteful and Inappropriate Service Reduction Model, known as WISeR, brings private technology vendors into Original Medicare’s authorization process. Those companies use artificial intelligence and other technologies to examine whether requests meet existing Medicare coverage rules.
The program does not permit an algorithm to issue a final denial by itself. CMS requires a licensed clinical reviewer to confirm every non-affirmation, its term for a rejected authorization request. However, newly released records describe backlogs, technical failures, slow responses, and substantial differences among states.
The central conflict is larger than whether a particular algorithm works. CMS pays participating companies a portion of spending avoided after certain requests or claims are rejected. Critics argue that this structure rewards denial volume, while CMS says audits and quality adjustments discourage improper decisions.
That clash places WISeR between two familiar problems. Medicare needs tools to detect unnecessary services and fraud. Patients also need protection from automated systems that can make administrative errors faster and harder to challenge.
The Trump Medicare AI Pilot Brought Prior Authorization Into Six States
WISeR is not a nationwide replacement for Medicare coverage decisions, but it is a consequential test of technology-assisted prior authorization inside Original Medicare.
CMS launched the model on January 1, 2026. Participating organizations began accepting authorization requests on January 5 for services delivered from January 15 onward.
The program operates in Arizona, New Jersey, Ohio, Oklahoma, Texas, and Washington. CMS selected one technology company for each state: Zyter, Genzeon, Innovaccer, Humata Health, Cohere Health, and Virtix Health.
WISeR is scheduled to continue through December 31, 2031. CMS describes it as a voluntary model because providers can choose whether to request authorization before delivering an included service.
That description requires context. A provider who skips prior authorization does not escape review. The resulting claim becomes subject to prepayment medical review, meaning Medicare can examine it before releasing payment.
For providers, the practical choice is therefore between two review pathways. One occurs before treatment, while the other occurs after treatment but before reimbursement.
The model initially covers a narrow group of services that CMS associates with questionable utilization, patient-safety concerns, or documented fraud. Examples include skin and tissue substitutes, epidural steroid injections, electrical nerve stimulator implants, and knee arthroscopy for osteoarthritis.
It also includes certain incontinence devices and services connected with diagnosing or treating impotence. CMS delayed the inclusion of deep brain stimulation and percutaneous image-guided lumbar decompression until a later performance year.
The official WISeR model says the program does not change Medicare’s coverage or payment policies. Instead, participating companies must apply existing national or local coverage requirements earlier in the claims process.
AI’s precise role can vary by participant. The technology can organize clinical records, compare documentation with coverage criteria, identify missing information, and help prioritize cases for review.
CMS says technology can affirm eligible requests, but it cannot independently finalize a rejection. Every non-affirmation requires review by an appropriately licensed clinician.
That distinction matters because descriptions such as “AI denies Medicare care” can imply a completely autonomous process. WISeR formally retains a human decision point, although the quality and independence of that review remain important questions.
The program is also a test of institutional design. CMS is assessing whether outside technology companies can perform reviews faster and more consistently than traditional administrative processes.
If WISeR succeeds, it can provide a template for expanding technology-assisted utilization management. If it fails, it can demonstrate how software, fragmented operations, and financial incentives compound one another.
The first months have already shown why that distinction matters. The controversy is not simply about AI accuracy. It concerns the entire chain through which a recommendation becomes a delayed treatment, rejected request, resubmission, or appeal.
Why CMS Targeted These Medicare Services
WISeR addresses a real spending problem, but the concentration of that spending complicates claims about what the pilot itself can save.
CMS created the model after identifying services with unusually high spending growth or documented vulnerability to inappropriate billing. Skin substitutes provide the clearest example.
These products are used in wound care and can be medically necessary. However, Medicare spending on them rose sharply as average prices and utilization increased.
A KFF analysis of complete traditional Medicare claims data found that WISeR-covered services accounted for $12.3 billion in Part B spending during 2024. That represented 5.3 percent of total traditional Medicare Part B spending.
The same services accounted for $2.4 billion, or 1.1 percent of Part B spending, in 2019. Spending therefore increased by roughly 400 percent over five years.
Skin substitutes drove most of the increase. They represented $10.3 billion, or 83.4 percent, of spending on WISeR services during 2024.
Their average price per service rose from about $2,300 in 2019 to $21,200 in 2024. That was an increase of 820 percent.
This growth gives CMS a strong reason to intervene. A system that pays submitted claims without timely scrutiny can become an attractive target for abusive billing.
Prior authorization can also prevent patients from receiving procedures that lack evidence for their specific condition. Unnecessary treatment carries medical risks, including infection, complications, and avoidable recovery time.
CMS has previous evidence that targeted authorization can reduce spending. Its application materials cite selected outpatient procedures whose Medicare spending fell from $51 million during the second half of 2019 to $35 million during the corresponding period in 2023.
Yet the broader numbers create an important counterpoint. According to KFF’s claims analysis, only about 207,500 beneficiaries used a WISeR service within the six pilot states during 2024.
That figure represented 19.7 percent of the roughly 1.1 million traditional Medicare beneficiaries who received an included service nationwide. It also represented a small portion of all Medicare beneficiaries.
More importantly, CMS introduced a separate nationwide payment change for skin substitutes at the start of 2026. The agency estimates that policy will reduce spending on these products by nearly 90 percent.
Price reform, rather than prior authorization, may therefore produce most near-term savings in the program’s largest spending category. That makes WISeR’s independent financial effect difficult to isolate.
A fall in skin-substitute spending would not automatically validate the authorization technology. Evaluators must separate savings created by lower prices from savings associated with fewer inappropriate services.
The same issue affects denial statistics. A high denial rate can indicate that a contractor found widespread noncompliance. It can also signal confusing documentation rules, an aggressive review threshold, or poor implementation.
Raw approval percentages cannot resolve those possibilities. Reliable evaluation requires information about initial decisions, resubmissions, reversals, appeals, patient outcomes, and differences among services.
CMS has a legitimate obligation to protect public money. Critics do not need to deny that responsibility to question whether WISeR’s design protects patients equally well.
The pressure falls most heavily on clinicians and beneficiaries. A contractor can process a request as a record, but a delayed record may represent a patient waiting with an open wound or persistent pain.
The Payment Formula Turns Savings Into the Main Conflict
WISeR’s most controversial feature is not its use of AI. It is the decision to connect vendor compensation directly to averted Medicare spending.
Participating companies do not receive ordinary fixed payments for processing requests. They can earn a portion of the estimated spending prevented by qualifying non-affirmations and denied claims.
For prior authorization, CMS calculates a regional benchmark for the service. It then applies a discount, a participant payment rate, and a quality multiplier.
A rejected request becomes eligible for payment only if it is unique and remains rejected after resubmission or appeal. The participant must also comply with reporting and performance requirements.
Prepayment review follows a related structure. A vendor may receive compensation when its review results in denial of a claim that fails Medicare’s existing coverage criteria.
CMS can withhold or recover that payment when an appeal succeeds. The company also bears the cost of handling unlimited resubmissions, and it receives only one payment per beneficiary for the request.
These provisions create meaningful safeguards. They reduce the value of issuing a weak denial that a provider can readily overturn.
Every proposed rejection also requires human clinical review. CMS can audit decisions, demand corrective action, withhold compensation, or remove a poorly performing participant.
However, the payment methodology still starts with avoided spending. A request that receives approval does not generate the same savings-based reward.
That asymmetry fuels criticism from physicians, hospitals, lawmakers, and digital-rights advocates. They argue that the business model points contractors toward finding more reasons to reject requests.
The quality multiplier does not fully eliminate that concern. CMS guidance assigns participants with aggregate quality scores from 85 to 100 percent a 100 percent multiplier.
Scores from 60 to 84 percent produce a 95 percent multiplier. Scores below 60 percent produce a 90 percent multiplier.
In other words, weak aggregate performance can reduce payments without necessarily eliminating them. CMS retains stronger enforcement options, but the ordinary formula applies a limited financial reduction.
This structure creates a classic principal-agent problem. Medicare wants vendors to identify only services that fail established coverage criteria. Vendors earn money when reviews produce qualifying avoided spending.
The goals overlap when a request is clearly inappropriate. They diverge when the documentation is incomplete, the coverage standard is ambiguous, or the patient’s circumstances require judgment.
Technology can intensify that divergence. A system optimized to surface questionable requests may increase reviewer attention on borderline cases, even without issuing autonomous denials.
A clinician technically remains responsible for the final decision. Still, the software can determine which evidence appears first, which cases receive scrutiny, and how uncertainty is framed.
Human oversight is therefore not a complete answer. Reviewers can inherit automation bias, meaning they place excessive confidence in a computer-generated recommendation.
Time pressure adds another risk. When request volumes rise, a reviewer may lack enough time to independently reconstruct a complex clinical case.
CMS argues that its safeguards align contractor behavior with accuracy. Critics answer that the underlying incentive remains strongest when care is not authorized.
Both positions contain testable claims. CMS should be able to publish decision accuracy, reversal rates, timeliness, and service-specific outcomes for every participant.
Without that information, observers cannot determine whether the payment formula rewards careful review or merely tolerates errors as a business expense.
Early Denial Rates and Delays Challenge WISeR’s Efficiency Claim
The first operational records show that technology did not produce consistently fast or predictable decisions across the six-state pilot.
The Electronic Frontier Foundation obtained about 1,000 pages of WISeR records after filing a Freedom of Information Act lawsuit against CMS. The documents cover vendor agreements, reporting rules, operational updates, and provider feedback.
The released CMS records run through portions of the program’s opening months. They do not provide a complete final evaluation, but they reveal problems hidden by top-level descriptions.
According to reporting based on those documents, thousands of authorization requests exceeded the program’s three-day processing target. Several hundred remained unanswered at the end of March.
One pending request was reportedly 83 days old. Records also described system outages, incorrectly categorized submissions, and communication failures between vendors and Medicare Administrative Contractors.
Ohio providers reported particular frustration with Innovaccer’s systems. Correspondence showed that the company had warned CMS before launch that it needed more time to complete the planned implementation.
Some providers returned to paper processes when digital portals did not function reliably. That outcome directly conflicts with WISeR’s promise to reduce administrative work through technology.
Provider feedback gave the delays a clinical dimension. Some offices described patients crying because painful procedures remained stuck in the authorization process.
Such reports do not prove that every pending request should have been approved. They do show that processing speed is a patient-safety issue, not merely a vendor performance statistic.
Washington hospitals documented similar effects earlier in the year. Data from 16 hospitals suggested that some procedure timelines grew from roughly two weeks to between four and eight weeks.
At the University of Washington Medical System, urgent authorizations that previously took one day reportedly required 15 to 20 days. Standard requests had previously taken about three days.
The system also reported nearly 100 patients waiting for epidural steroid injections. These procedures can address severe pain, even when they are not emergency interventions.
The Washington hospital findings covered early implementation and a limited group of institutions. They should not be generalized automatically to every vendor or state.
Still, the variation itself is significant. A federal demonstration should test a reproducible operating model, not six unrelated implementations with incomparable results.
High early rejection rates also demand careful interpretation. Reports based on released data described substantial non-affirmation levels in Washington during the opening months.
Texas data reportedly showed an initial approval rate near 62 percent before physician review raised it to about 84 percent. If accurate, that change illustrates the importance of human intervention.
It also raises a harder question. A system that flags too many legitimate requests can burden clinicians even when human review eventually corrects the recommendation.
False positives are not harmless. Each one can trigger document retrieval, phone calls, peer review, rescheduling, or a delayed treatment decision.
CMS has said turnaround times improved after early problems. Reported averages later fell to about 1.7 days for prior authorization and slightly more than three days for prepayment review.
Improvement would show that some failures were launch problems rather than permanent design defects. However, averages can conceal outliers and differences across services, vendors, and patient groups.
The meaningful metric is not simply whether the mean falls below three days. CMS must identify how many requests miss the target and how long the slowest cases remain unresolved.
It should also distinguish delay from denial. A request left pending can block care as effectively as a formal rejection, while providing fewer opportunities for appeal.
Human Review Does Not Solve the Transparency Problem
WISeR’s safeguards operate after private systems analyze patient records, yet the public still lacks enough information to evaluate those systems.
CMS says every non-affirmation receives human clinical review. Beneficiaries and providers also retain existing appeal rights when a submitted claim is denied.
Providers can resubmit authorization requests without a numerical limit. They may also seek peer-to-peer clinical review during the resubmission process.
These safeguards separate WISeR from a fully automated denial engine. They matter because clinical coverage decisions often require context that a rules-based system cannot capture.
Yet formal human involvement does not reveal what happens before that review. CMS has not provided a standardized public description of each vendor’s models, data sources, validation procedures, or error thresholds.
The six companies may use “AI” in very different ways. One may extract information from clinical notes, while another may classify cases or generate a suggested coverage determination.
Those functions carry different risks. Document extraction can miss a crucial statement. Classification can reproduce biases in historical decisions. Generative systems can introduce unsupported information.
The public record also offers limited evidence about subgroup performance. Medicare beneficiaries vary by age, disability, language, geography, race, and access to specialists.
A system can show acceptable aggregate accuracy while performing poorly for a smaller population. That is why error rates should be reported across meaningful patient and provider categories.
Privacy presents another unanswered question. Vendors receive sensitive medical information to perform reviews, creating additional access points and operational dependencies.
Business associate agreements can establish legal duties, but they do not substitute for public security testing. The consequences of a breach extend beyond billing records.
EFF’s lawsuit sought records on accuracy testing, bias, hallucinations, privacy, and audit procedures. Its central concern is not that any use of software is inherently unacceptable.
The concern is that an opaque system can influence access to government benefits without enough evidence for outsiders to assess reliability.
This transparency gap also makes accountability harder. When a patient experiences a delay, responsibility may be distributed among a provider, vendor, contractor, CMS office, and clinical reviewer.
Each party can point to another part of the workflow. The beneficiary still experiences the system as a single barrier.
Medicare Advantage offers a relevant warning. Prior authorization is common in that program, and federal investigators have previously found instances where plans denied requests that met Medicare coverage rules.
WISeR does not duplicate Medicare Advantage. CMS controls its coverage standards, requires clinical confirmation, and maintains direct oversight of participants.
Nevertheless, it imports familiar utilization-management tools into Original Medicare. That makes past problems an appropriate benchmark for evaluating safeguards.
The strongest defense of WISeR is therefore not the phrase “human in the loop.” It would be measurable evidence that human review catches errors before patients experience harm.
CMS should report how frequently technology recommends rejection, how often clinicians disagree, and how many initial decisions change after resubmission. Appeal outcomes should be connected to individual vendors and services.
Timeliness data should include distributions, not only averages. Accuracy audits should explain their sampling methods and publish enough results for independent analysis.
The agency should also separate missing-document cases from substantive medical-necessity decisions. Those categories impose different burdens and require different remedies.
A documentation problem may be solved by better forms or electronic record integration. A clinical disagreement requires stronger review and appeal protections.
Transparency will not eliminate every dispute. It will make disagreements specific enough for patients, providers, and policymakers to evaluate.
WISeR’s Next Tests Will Decide Whether the Model Expands
The next phase should be judged through three signals: vendor-level outcomes, enforcement actions, and evidence of expansion beyond the original service list.
The first signal is a complete performance dataset for each participant. CMS needs to publish approval, non-affirmation, resubmission, reversal, appeal, and processing-time results.
These numbers should be broken down by service and state. A national aggregate would hide whether one contractor or clinical category drives most of the problems.
CMS should include the percentage of technology recommendations changed by human reviewers. That measure would show whether automation filters routine work or generates excessive false positives.
A low disagreement rate is not automatically reassuring. Reviewers can defer too readily to software, so independent audits must assess final decisions against Medicare’s coverage rules.
Patient impact also belongs in the dataset. Evaluators should track treatment abandonment, rescheduling, worsening conditions, and complaints linked to processing delays.
If CMS provides detailed evidence and outcomes improve, the case for WISeR becomes stronger. If data remain partial, the model’s efficiency claims will remain difficult to verify.
The second signal is enforcement. CMS has several remedies available, including corrective action plans, payment withholds, recoupment, and termination.
The agency’s willingness to use them will show whether quality requirements constrain vendors in practice. Rules matter less when poor performance produces only modest payment reductions.
Enforcement should also be timely. A contractor should not continue creating backlogs for months while CMS waits for annual reconciliation.
Public corrective actions can help providers understand what failed and whether procedures changed. Confidential supervision cannot build trust in a program affecting access to federal benefits.
The third signal is expansion. CMS has said it may add services, regions, or participants during later performance years.
Released records reportedly discussed possible future categories such as air ambulance transportation, MRI scans, and certain high-cost Part B drugs. Broader coverage would expose more beneficiaries to the model.
Expansion before credible evaluation would weaken CMS’s claim that WISeR is a controlled experiment. A pilot should generate evidence before becoming a template.
Congress will remain part of that decision. Lawmakers have already proposed restricting funding, while physician and hospital groups have requested stronger safeguards or an end to the program.
Political opposition alone does not prove that the model cannot work. Fraud prevention and inappropriate-service reduction remain legitimate public goals.
However, CMS carries the burden of showing that its chosen mechanism improves on existing review systems. Cost savings cannot stand alone when the process controls access to medical care.
The Trump Medicare AI pilot therefore presents a governance test for other public-sector AI programs. Agencies increasingly use private technology to sort cases, recommend decisions, and target enforcement.
Those systems often operate before a formal human decision. Their greatest influence may come from defining which cases appear suspicious and which evidence receives attention.
For developers, the lesson is that model accuracy is only one part of deployment quality. Queue design, failure recovery, logging, escalation paths, and user communication can determine real-world outcomes.
For enterprise buyers, WISeR shows why incentive design belongs in technical procurement. A model can meet a benchmark while the surrounding contract rewards the wrong operational behavior.
For patients and clinicians, the immediate task is more practical. They should document submission dates, response times, resubmissions, and communications when a WISeR service is involved.
They should also distinguish a prior authorization non-affirmation from a denied claim. The available response and appeal pathway depends on where the case sits in the payment process.
WISeR has not yet established that AI-assisted review is inherently harmful or inherently efficient. It has shown that deploying the technology changes who controls delays, documentation, and financial risk.
The decisive evidence will come from corrected decisions, patient outcomes, and CMS enforcement, not from broad promises about AI. Until those results are public, expansion deserves close scrutiny.
Readers tracking the Trump Medicare AI pilot should watch the underlying records rather than any single denial percentage. Ask whether CMS publishes comparable vendor data, acts when standards fail, and waits for evidence before broadening WISeR. Those choices will reveal whether the program remains a bounded experiment or becomes the foundation for wider automated review. The core question is straightforward: can Medicare verify savings without making older and disabled Americans bear the cost of an unproven administrative system?



