top of page

Nearly 4 in 10 CFOs Expect Gen AI to Pay Off Within Two Years

PYMNTS reached Google News with a striking result: nearly four in ten surveyed CFOs expect very positive generative AI returns within two years. That confidence sounds like a victory for enterprise AI. The underlying data tells a more complicated story.

Finance leaders are seeing stronger results from focused applications. They are also extending the time needed to embed generative AI across entire organizations. The apparent contradiction separates early operational value from a complete financial payoff.

The enterprise CFO survey covered 60 finance chiefs at United States companies with at least $1 billion in annual revenue. PYMNTS Intelligence conducted the survey from December 15 through December 19, 2025. Its findings describe a small, concentrated group of large enterprises rather than every American business.

Those CFOs reported better results across tasks ranging from automated workflows to customer responses. They also identified fewer drawbacks than they had six months earlier. Yet their expected timelines for full integration nearly doubled.

That tension matters more than the headline percentage. CFOs are not simply asking whether a model can draft text, analyze records, or produce a forecast. They are deciding whether AI can operate inside governed workflows and create durable economic value.

The primary contest is therefore not AI optimists against AI skeptics. It is the promise of quick, targeted wins against the reality of slow organizational change. The winners will connect those two timelines without treating productivity estimates as audited returns.

What the Google News Headline Actually Reveals

The survey shows growing confidence in specific AI applications, not proof that enterprise-wide transformation is nearly complete.

PYMNTS found that CFO expectations for generative AI ROI had become more dispersed by December 2025. A larger group anticipated very positive returns within one or two years. Another group expected a much longer path to a full economic payoff.

That distribution is more informative than a single average. It suggests finance leaders have stopped treating every generative AI project as the same investment. A document assistant and a company-wide forecasting system carry different costs, risks, and implementation requirements.

The distinction also explains why favorable ROI expectations can rise while integration forecasts get longer. A team can save time on research or reporting before its company redesigns an entire financial process. Early value does not require immediate institutional transformation.

PYMNTS reported that satisfaction reached between 80% and 100% across the use cases already deployed by surveyed firms. Ratings improved especially for code generation, automated workflows, product innovation, and real-time customer responses. Those figures describe respondents’ assessments, not independently audited financial returns.

The surveyed CFOs also reported fewer perceived drawbacks. The average fell from 6.92 drawbacks in July 2025 to 4.23 in December. Among the largest companies, the average dropped from 7.52 to 3.44.

Several operational concerns declined with it. The share citing errors and output problems fell from 80% to 35%. Integration and implementation concerns dropped from 70% to 45%, while cost and maintenance concerns fell from 78.3% to 23.3%.

These changes support a credible interpretation. Experience is helping large companies solve some problems that looked daunting during early pilots. Teams are establishing controls, choosing narrower use cases, and learning where human review remains necessary.

However, a reduction in reported concerns does not make those concerns disappear. A 35% error concern remains material inside financial reporting, compliance, or treasury. A 45% integration concern still indicates substantial work across systems and data.

The study also relies on self-reported perceptions from 60 executives. That design can identify changes in sentiment and practice, but it cannot establish causation. It does not prove that generative AI produced a particular margin improvement.

The headline should therefore be read as evidence of maturing expectations. Nearly four in ten CFOs foresee strong near-term returns because they can separate useful deployments from broad transformation. Their optimism has become more selective, not less serious.

This is the first important shift in the PYMNTS CFO survey. Finance leaders increasingly believe generative AI works in defined settings. They are now confronting the harder question of whether those settings can become repeatable, controlled operating systems.

Quick AI Wins Are Colliding With Slow Integration

Generative AI ROI appears fastest when companies contain the task, define the baseline, and keep responsibility with accountable employees.

A targeted use case has visible boundaries. A finance team can measure how long a recurring report took before deployment and compare that baseline with later results. It can also record corrections, review time, and the cost of operating the model.

Enterprise integration has no such simplicity. It touches access permissions, data definitions, approval chains, software contracts, employee roles, and regulatory obligations. A model can perform well while the surrounding process remains unsuitable for automation.

This difference helps explain the longer timelines in the PYMNTS research. By December 2025, respondents expected full integration to take nearly twice as long as they had anticipated between July and September. More experience produced greater confidence in use cases and greater caution about scaling.

Deloitte found a similar divide in its AI ROI findings. Its research included 1,854 executives across Europe and the Middle East, supported by 24 detailed interviews.

Among generative AI users in that research, 15% reported already achieving significant, measurable ROI. Another 38% expected such returns within one year of investing. Yet 62% said a typical AI use case reached satisfactory ROI within two to four years.

The distribution underscores a crucial measurement problem. “Significant” and “satisfactory” returns can represent different thresholds. A company can recognize useful productivity gains before recovering the full costs of implementation and organizational change.

Agentic AI makes the timing problem sharper. An AI agent is software that can plan and perform multistep actions under defined permissions. Deloitte found that 57% of respondents were using agentic AI, but only 10% reported significant current ROI.

Half of those agentic AI users expected returns within three years. Another third expected a three-to-five-year timeline. Greater autonomy raises the potential value, but it also increases integration and control requirements.

Finance functions illustrate why. A model can summarize variance explanations with limited access and mandatory review. An agent that changes a forecast or initiates an adjustment needs permissions, validation rules, logs, and an accountable owner.

Companies must also decide what counts as value. PYMNTS found that improved customer experience and margins were leading ROI measures by December. Cost reduction, revenue expansion, and lower capital expenditures were also gaining attention.

Headcount reduction became a less common yardstick. That change weakens the simplest automation narrative. CFOs increasingly see value through faster cycles, improved decisions, and better service rather than direct labor removal alone.

Those benefits remain difficult to compare. Hours saved are not automatically converted into profit. Faster analysis matters only when employees redirect that time toward useful decisions or additional work.

A credible generative AI ROI model must therefore account for accepted outputs, review effort, failure rates, and downstream business effects. It should include infrastructure, integration, governance, training, and vendor management. Ignoring those costs creates an attractive but incomplete result.

This is why knowledge access matters in everyday deployments. Employees need reliable context before AI can produce useful work. A searchable AI knowledge base can reduce retrieval friction, but human verification still protects consequential decisions.

The early wins are real when teams measure them carefully. The collision begins when leaders assume those wins will scale without redesigning workflows. Every additional system, permission, and stakeholder introduces another condition for success.

CFOs Now Have to Defend Both Speed and Control

Finance leaders face pressure to fund AI quickly while preserving the controls that make financial systems trustworthy.

CFOs traditionally challenge investments that lack measurable returns. Generative AI places them in an unusual position. They must act as capital gatekeepers while competitors, boards, and operating teams demand faster adoption.

That pressure is already visible. A July 2026 survey reported that 92% of CFOs and senior finance executives felt pressure to demonstrate acceptable AI returns. The finance leader pressure research covered 1,505 executives in four countries.

Half of its respondents said their AI agents had delivered only limited measurable ROI. Just 7% said their organizations prioritized AI governance over adoption speed. These results expose the risk of letting deployment schedules outrun accountability.

Only 44% felt fully confident that they could explain an AI agent’s actions to an auditor or regulator. Another 30% said internal controls had not been updated during the previous year. Those gaps matter when software begins influencing transactions or financial judgments.

Auditability means a company can reconstruct what a system did, which information it used, and who approved the result. Conventional software follows predefined logic. Generative systems can produce variable outputs, so organizations need stronger records around prompts, sources, revisions, and approvals.

The same survey found that only 28% required documented audit logs showing how an AI agent reached decisions. Another 23% said responsibility for an AI error would be unclear or assigned to nobody. Those are governance deficiencies, not minor technical details.

A model error in an internal draft has a manageable consequence. An unexplained error in tax, reporting, or compliance can create financial and legal exposure. The control environment must match the impact of the action.

This does not mean every use case needs the same restrictions. Low-risk summarization can use lighter review than capital allocation or external reporting. Risk-based controls let companies move quickly where consequences remain contained.

CFOs also have to resolve ownership. Technology teams can operate models, but they cannot define every acceptable finance outcome. Finance teams understand materiality and approval requirements, while security teams understand access and data exposure.

Successful governance joins those responsibilities before deployment. The business owner defines the outcome and tolerance for error. Technology teams maintain the system, and risk functions test whether controls operate as intended.

The PYMNTS findings suggest large companies are improving at this work. Their reported concerns about errors, costs, and implementation declined substantially over six months. Rising effectiveness ratings indicate that controlled exposure can build institutional confidence.

However, confidence can become dangerous when it replaces testing. Self-reported satisfaction does not reveal how often employees silently correct outputs. It may also exclude abandoned pilots, which can make active deployments appear stronger.

Organizations should track the full funnel. That includes attempted tasks, completed outputs, accepted outputs, revisions, escalations, and failures. A model that produces drafts quickly but demands extensive checking may shift work instead of eliminating it.

The pressure on CFOs will intensify as budgets rise. Bain surveyed more than 100 global CFOs and found that 83% planned to increase enterprise AI spending by more than 15% over two years. Its AI spending plans also showed that 42% expected increases above 30%.

Only 15% to 25% had scaled AI across finance functions, according to Bain. Overall, 31% were satisfied with their outcomes. Satisfaction exceeded 60% among companies in the top quartile of AI maturity.

That gap turns operational maturity into the decisive factor. Buying more access does not create mature data, workflows, or controls. CFOs must fund the organizational work that converts a model into a dependable capability.

What the Numbers Still Do Not Prove

The evidence supports cautious optimism, but it does not establish a universal two-year deadline for generative AI ROI.

The PYMNTS sample is narrow by design. Every respondent led finance at a United States company generating at least $1 billion annually. Large enterprises have resources, data, and technical teams that smaller organizations may not possess.

The sample also contains only 60 CFOs. One respondent represents about 1.7 percentage points. A small change in answers can therefore produce a visible movement in reported shares.

That limitation does not invalidate the study. It means readers should avoid translating its results into a prediction for every business. The survey is best treated as a view into large-enterprise finance leadership.

The timing language also deserves care. Respondents were asked when returns would become “very positive,” not when every implementation cost would be recovered. Different executives can apply different thresholds to that phrase.

PYMNTS reported a broader portfolio of ROI measures, including margins, customer experience, cost reduction, and revenue. That approach reflects how businesses actually create value. It also makes comparisons harder unless each company defines its calculation consistently.

A customer experience gain can matter financially, but the connection requires evidence. Leaders need to show whether better service changed retention, conversion, support costs, or another measurable outcome. Otherwise, the benefit remains an intermediate indicator.

The survey’s timeline creates another caution. PYMNTS collected responses during five days in December 2025. The headline entered the current Google News cycle later, but the underlying responses are a snapshot.

Models, contracts, and enterprise strategies have changed since that fieldwork. Some projects have advanced, while others have encountered cost or security problems. A current headline should not make the survey date invisible.

Independent studies also describe sharply different levels of success. A Wharton and GBK Collective enterprise adoption study surveyed 801 enterprise decision-makers in 2025. It found that 72% formally measured generative AI ROI.

Three out of four respondents in that study reported positive returns. Four out of five expected investments to pay off in roughly two to three years. Those findings reinforce optimism while extending the expected horizon beyond the headline’s two-year window.

Wharton’s research also found that 82% used generative AI at least weekly, while 46% used it daily. Regular use had clearly become common among enterprise decision-makers. Adoption, however, is not the same measure as economic impact.

The study identified a human constraint as well. Eighty-nine percent agreed that generative AI enhanced employee skills, yet 43% saw a risk of declining skill proficiency. AI can improve output while weakening expertise if employees stop practicing essential tasks.

Finance leaders must account for that possibility. A faster analyst who retains judgment represents durable value. A faster process that erodes the ability to detect errors creates a hidden liability.

Other research finds notable operational gaps even where returns look favorable. One 2026 finance survey reported 97% AI adoption among respondents and positive returns for more than 75% of investments. Yet 45% still spent more than 60% of their time on manual tasks.

That finance adoption gap shows why broad claims need context. A finance department can report positive AI returns while continuing to rely on spreadsheets, manual reconciliations, and labor-intensive reporting.

The most skeptical interpretation is that optimistic surveys measure expectations and selected wins more readily than total transformation. The most favorable interpretation is that targeted value arrives before visible organizational change.

Both can be true. Generative AI can produce worthwhile returns without remaking an entire finance department. The mistake is treating early success as proof that every process should now be automated.

CFOs should demand evidence at the use-case level and aggregation at the portfolio level. Failed experiments belong in the calculation alongside successful deployments. Excluding them inflates the apparent performance of the overall program.

They should also distinguish recurring improvements from one-time gains. A model might help clear a backlog, but that benefit will not repeat indefinitely. Durable ROI requires ongoing work, lower recurring costs, better decisions, or sustained revenue effects.

The headline is useful because it captures a real turn in executive sentiment. It becomes misleading only when stripped of sample limits, measurement choices, and longer integration timelines. Those details define whether the two-year expectation is credible.

Three Signals Will Test the Two-Year Payoff

The next phase will be judged by audited operating results, production-scale adoption, and evidence that controls are keeping pace.

The first signal is a change in how companies report AI value. Productivity estimates and employee surveys helped during experimentation. Boards will increasingly expect measures tied to margins, revenue, working capital, error reduction, or completed transaction cycles.

A stronger disclosure would include operating costs and human review. It would separate gross time savings from realized financial value. It would also identify whether the result came from generative AI, conventional automation, or broader process redesign.

If companies begin reporting repeatable financial outcomes, the near-term PYMNTS forecast gains support. If disclosures remain limited to usage and estimated hours saved, the verification gap will persist.

The second signal is movement from pilots into controlled production. Bain found that only 15% to 25% of surveyed CFOs had scaled AI across finance. That leaves a large distance between experimentation and normalized operations.

Production deployment should mean more than organization-wide access to a chatbot. It requires integration with approved data, defined tasks, monitoring, escalation rules, and named owners. Employees also need clear instructions for challenging unreliable outputs.

Scaling rates will reveal whether the strongest use cases can survive real operational complexity. A rise in governed, repeatable deployments would strengthen the two-year thesis. A growing inventory of stalled pilots would weaken it.

The third signal is whether governance metrics improve alongside adoption. The July finance survey found low use of documented agent logs and widespread uncertainty about accountability. Those gaps become more consequential as systems receive greater autonomy.

Watch for updated internal controls, mandatory audit trails, tested incident plans, and explicit ownership for AI errors. These measures do not guarantee positive ROI. They reduce the chance that one failure erases gains from many successful tasks.

Governance improvements would also make financial measurement more trustworthy. A company cannot calculate value accurately if it does not record system behavior, corrections, and exceptions. Controls create the evidence needed to support an ROI claim.

The Google News headline therefore marks a checkpoint, not a finish line. Nearly four in ten CFOs expect very positive returns within two years, while the same research shows full integration taking longer.

That split is rational. Companies do not need to transform every workflow before receiving value from selected applications. They do need disciplined measurement before treating that value as a durable enterprise return.

For developers, the lesson is to design for traceability, review, and bounded actions. For enterprise buyers, it is to evaluate workflow outcomes rather than model demonstrations. Knowledge workers should examine whether AI improves their judgment or only accelerates first drafts.

The decisive question is no longer whether generative AI can help finance teams. It already does in many defined tasks. The question is whether organizations can turn those gains into repeatable, governed financial performance.

Over the next two years, watch what CFOs report after implementation costs, corrections, and control work enter the ledger. That evidence will determine whether today’s optimism becomes an operating advantage or another forecast that moved faster than reality.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page