What OpenAI Says It Learned From Building Finance Around AI - and Where Promise Meets Control
OpenAI has set two unusually aggressive finance goals: a zero-day close and automated, continuously updated forecasting. Yet CFO Sarah Friar says neither goal has been reached. That admission gives the openai what question a useful tension. The company is describing a working transformation, not announcing a finished autonomous finance department.
The distinction matters because OpenAI has access to its own models, ChatGPT Work, Codex, technical experts, and substantial organizational attention. Its finance team still found that installing AI tools was the easy part. Redesigning decisions, data access, approvals, and accountability proved harder.
That puts OpenAI’s claims against a less flattering industry record. Gartner found finance AI adoption barely moved between 2024 and 2025. BCG found that many finance teams reported modest or unmeasurable returns. OpenAI’s experience argues that the missing ingredient is not another isolated assistant. It is a redesigned operating model that keeps humans responsible for every consequential result.
OpenAI What Changed Inside the Finance Function
OpenAI’s finance project shifts the target from automating tasks to rebuilding the path between evidence and decisions.
Friar published her five lessons on August 10, 2026, nearly two years after joining OpenAI as CFO. She describes inheriting a small finance organization supporting a company growing at an exceptional pace. Even there, closing and forecasting still depended on recurring manual work.
OpenAI’s two ambitions frame the entire project. A zero-day close would offer a continuously reconciled and traceable view of the company’s financial position. Continuous forecasting would update that view as operating conditions, evidence, and assumptions changed.
The company is still building toward both goals. That qualification is central, because zero-day close is an ambition rather than a reported production result. OpenAI has not published independent measures of its close time, forecast accuracy, exception rate, or financial return.
Still, its finance lessons reveal meaningful operational changes. The team is moving from static spreadsheets and presentation decks toward live systems built over approved business data. Those systems connect analysis, supporting evidence, scenarios, and follow-up questions.
OpenAI also created IR-GPT, a custom GPT grounded in approved investor-relations material. The tool prepares initial responses to investor diligence questions that previously required extensive searching, drafting, consistency checks, and review coordination.
According to Friar, a strong first draft can now appear in seconds instead of hours. The investor-relations team still checks the answer, adds context, and ensures consistency across investors. The model accelerates preparation, while employees retain responsibility for the response.
A finance hackathon helped produce that system and other custom GPT projects for procurement and tax. OpenAI paired finance employees with sales engineers, then asked participants to bring recurring work they wanted to change. That structure turned experimentation into workflow design within one day.
The shift also reaches planning. One finance employee without prior coding experience reportedly used Codex to build an advertising forecast tool. It converts monthly projections into weekly and daily plans while accounting for weekdays and holidays.
The tool compares forecasts and keeps figures connected to the approved model. Marketing leaders can inspect changes, consider additional spending, and see where returns appear to weaken. The example remains a company-reported case, but it illustrates the project’s real destination.
OpenAI is not trying to place a chatbot beside an unchanged spreadsheet process. It wants live financial systems that preserve sources, expose assumptions, and help people act before a reporting period ends.
That is the first major lesson. An AI-native finance function is not defined by how often employees prompt a model. It is defined by whether the operating cycle changes without obscuring ownership.
A Continuous Forecast Starts With the Close
The most important mechanism is a connected evidence layer, because faster language generation cannot repair fragmented financial records.
A conventional monthly close pulls information from several places. Actual spending may sit in the general ledger, while purchase orders live elsewhere. Accruals remain in spreadsheets, and explanations often hide inside messages or presentation notes.
OpenAI wants to connect approved spending plans, general-ledger actuals, purchase orders, accruals, and transaction details in a continuously reconciled view. Reconciliation means matching related financial records and identifying differences that require investigation.
In the proposed workflow, AI prepares an initial variance explanation and flags exceptions. Finance verifies the figures, applies judgment, and completes the final sign-off. The close remains, but the scramble to reconstruct events after month-end begins to shrink.
That distinction prevents a common mistake. Generative AI can draft a confident explanation for a variance, but confidence does not establish that the underlying records agree. The system needs traceable source data before generated commentary becomes useful.
A continuously refreshed forecast builds on the same foundation. OpenAI says its workflows combine statistical models, sales conversations, account-level evidence, operating data, and finance judgment. Interactive tools then place the baseline, supporting evidence, and alternative scenarios in one view.
Suppose a new customer commitment does not appear in the approved forecast. The system can surface the evidence and show how an adjustment affects the quarter or year. Finance can inspect the assumption before deciding whether to change the baseline.
This process is different from letting a model silently revise an official plan. The system recommends, explains, and compares. An authorized finance professional decides whether the organization’s approved view should change.
The approach aligns with broader industry evidence. A finance AI study from McKinsey found that organizations are applying AI across planning, controls, working capital, and cost management. It also reported that weak integration often prevents pilots from scaling.
McKinsey surveyed 102 CFOs and found 44 percent used generative AI across more than five use cases in 2025. That was up from 7 percent in the previous survey. However, nearly two-thirds of respondents in a broader survey had not begun scaling AI across their enterprises.
The contrast supports OpenAI’s mechanism. Companies can accumulate use cases without rebuilding the data and decision path connecting them. Activity expands, but the operating system remains fragmented.
OpenAI’s finance model begins with a consequential decision and works backward. Teams map the required data, tools, approvals, and handoffs. They then identify where AI can analyze information, coordinate work, or complete a bounded step.
That method also changes how organizations manage context. Important evidence often spans financial systems, operational tools, contracts, sales notes, and internal discussions. A well-governed knowledge blending workflow can connect those sources while preserving their origin and relevance.
The benefit is not simply a faster report. Leaders receive a more current view of what changed, why it changed, and which action could alter the result. Finance spends less time assembling a briefing and more time challenging its assumptions.
Yet OpenAI has not disclosed enough data to compare its new workflows with the old ones. There is no published forecast-error series or close-cycle benchmark. The mechanism is credible, but its measured effect remains unavailable.
That gap separates an instructive operating model from a proven industry standard. Other CFOs can copy the sequence of work, but they cannot assume OpenAI’s results will transfer automatically.
Finance Professionals Are Becoming Builders
OpenAI’s clearest organizational bet is that domain experts should shape their own tools instead of waiting for a central engineering queue.
Friar says every member of her finance team is building custom dashboards and tools with ChatGPT Work and Codex. That statement does not mean every accountant becomes a software engineer. It means finance expertise can travel further into the system design process.
OpenAI cites internal research showing that 40 percent of specialized AI use by finance professionals involves work outside traditional finance. It says another 22 percent involves engineering-related tasks. Those figures describe usage patterns, not job replacements or verified productivity gains.
The advertising forecast example shows how the builder model works. The employee understood the monthly planning problem, the approved model, and the decisions marketing leaders needed to make. Codex helped translate that knowledge into a tool with smaller planning intervals.
The advantage is proximity. A central software group can build a technically polished application while missing a key approval rule or business assumption. A finance professional already understands where the work becomes ambiguous and which exceptions matter.
However, proximity also creates new governance requirements. A useful prototype can become operational before security, documentation, testing, and maintenance responsibilities are clear. A spreadsheet macro created similar problems, but AI can increase the speed and scale of informal development.
OpenAI addresses this with a combination of bottom-up experimentation and top-down strategy. Employees identify recurring problems, while leaders direct resources toward workflows with material business value. Technical specialists help participants build and test their ideas.
This is more disciplined than purchasing seats and hoping usage produces value. Access remains necessary because employees need room to discover relevant problems. Structured events, approved data, and named priorities turn that access into an operating program.
The model also changes the relationship between finance and IT. IT still owns infrastructure, identity, security, and integration expertise. Finance contributes the accounting logic, operating context, control requirements, and decision criteria.
BCG reached a similar conclusion after surveying more than 280 finance executives experienced with AI. Its AI ROI research found stronger results when finance teams collaborated with IT and vendors. Fully staffing dedicated teams also improved reported success rates.
BCG found that 40 percent of surveyed CFOs did not know which AI functions their existing vendors already offered. That creates two opposite risks. A company might rebuild commodity capabilities, or it might accept a generic product that cannot represent its unique decision context.
OpenAI’s approach suggests a practical boundary. Use existing platforms for common controls and infrastructure, but customize workflows where company-specific data and judgment create material value. Algorithmic forecasting often falls into the second category.
The builder model also pressures traditional finance software suppliers. SAP, BlackLine, and other enterprise vendors already embed AI into reporting, reconciliation, payment prediction, and close management. Their advantage lies in established integration and control layers.
OpenAI’s model challenges them from another direction. If finance employees can rapidly create interfaces over approved company data, the value moves from packaged screens toward trusted context and flexible workflows. Vendors must make customization easier without weakening auditability.
Central technology teams face pressure as well. Their role moves away from receiving every dashboard request and toward establishing safe building blocks. They must define data access, evaluation methods, reusable components, logging, and escalation paths.
The change will not eliminate specialized finance roles. It raises the value of people who can connect financial judgment with system behavior. The strongest builders will understand both why a number matters and how the supporting workflow can fail.
That makes training more demanding than prompt instruction. Employees need to test outputs, recognize data gaps, document assumptions, and know when a model should not act. Building speed without those habits can create faster mistakes.
Stronger Controls Are the Price of More Speed
OpenAI’s most important reversal is that broader AI access requires more explicit control, not less finance oversight.
The company’s IR-GPT example makes that tension visible. A system grounded in approved material can prepare a response quickly, but an investor answer carries legal, reputational, and consistency risks. OpenAI therefore keeps human review at the center.
Friar says CFOs should define which data an AI system can access and which actions it can take. They should also specify approval requirements, escalation conditions, and responsibility for changes to an approved baseline.
Every output should connect to a reliable source. Forecasts should include explanations, while changes to approved figures require finance authorization. These principles sound familiar because they extend existing financial controls into AI-mediated work.
The challenge lies in implementation. A source link does not guarantee that a generated conclusion faithfully represents the source. A human approval step also provides little protection if reviewers treat it as a routine click.
Effective oversight needs role-based access, activity logs, version history, evaluation tests, and clear separation between preparation and authorization. High-risk actions should require stronger evidence and additional review. Lower-risk drafting tasks can use lighter controls.
OpenAI also points to usage limits, model-routing rules, budget controls, and approval thresholds. Model routing selects different models according to task requirements, cost limits, or risk. These controls let CFOs manage AI as a variable operating expense.
However, the company has not disclosed its control testing, error rates, audit findings, or incident history. Its recommendations describe the intended design. They do not independently establish that every internal tool consistently meets that standard.
Industry evidence supports this skepticism. Gartner’s adoption survey found that 59 percent of finance leaders used AI in 2025. That was only one percentage point above the previous year.
Gartner attributed slower growth partly to complexity, data problems, and talent constraints. Ninety-one percent of respondents reported low or moderate impact during initial adoption. More mature users were substantially more likely to report stronger results.
Knowledge management was the most common use case among AI-using finance organizations, at 49 percent. Accounts-payable automation followed at 37 percent, while error and anomaly detection reached 34 percent.
Those figures suggest that organizations favor bounded workflows with identifiable inputs and review points. Fully autonomous forecasting or close management remains harder because the consequences spread across reporting, capital allocation, compliance, and executive decisions.
The openai what narrative becomes less convincing if readers interpret “AI-native” as “AI-controlled.” Friar’s actual model preserves finance ownership. AI gathers, drafts, compares, and flags, while people approve consequential outputs.
That arrangement also creates a measurement problem. Human review time can erase apparent automation gains. Errors discovered late can add rework, delay decisions, and weaken trust. A model that produces an answer quickly is not necessarily reducing total cycle time.
Controls should therefore appear inside the ROI calculation. Teams need to count review labor, remediation, integration, monitoring, and ongoing evaluation. Otherwise, an efficient demonstration can hide an expensive production workflow.
There is also a concentration risk when many processes depend on the same model provider or context layer. A model update can alter behavior across forecasting, procurement, tax, and investor relations. Finance teams need regression testing before accepting changed outputs.
The same principle applies to data permissions. A custom assistant should not retrieve sensitive records merely because they might improve an answer. Access must reflect the employee’s role, the task’s purpose, and the organization’s retention policies.
OpenAI’s argument is strongest when read as a control design, not an autonomy claim. Faster preparation increases the need for traceability because more outputs can reach reviewers in less time. Speed raises the volume of judgment calls.
AI ROI Needs a Workflow Scorecard
OpenAI wants CFOs to measure completed, dependable work rather than seats, prompts, or token consumption.
Friar proposes four questions for each workflow. Did AI complete work that mattered? What did the result cost after employee time, review, and rework? Was it usable, and did it improve speed or decisions?
That framework rejects activity metrics. A rising number of users can indicate adoption, but it cannot show whether the close improved. Token usage can reveal consumption without establishing business value.
Finance teams need operational measures tied to each process. A close scorecard can track cycle time, automated reconciliation, exceptions requiring review, and the time needed to explain variances. A forecasting scorecard can track accuracy, refresh frequency, scenario time, and decision quality.
These measures should use a credible baseline. Teams need to compare the new workflow with the previous process across similar periods and conditions. Otherwise, seasonal changes or staffing differences can be mistaken for AI gains.
Quality thresholds also matter. A first draft that arrives in seconds creates limited value if analysts must rebuild it. The relevant figure is accepted work after review, not raw output volume.
BCG found that only 45 percent of surveyed executives could quantify returns from finance AI initiatives. Among those measuring returns, one-third reported less than 5 percent. The median reported return was 10 percent, below the 20 percent many targeted.
Only about one in five respondents reported returns of at least 20 percent. BCG also found stronger results when organizations linked AI to broader finance transformation instead of treating it as a standalone project.
One example involved a consumer-goods finance team that built a driver-tree model connecting operational and financial measures. A generative AI interface then answered questions over that structure. BCG reported that the system cut report-generation time by 50 percent.
The same foundation later supported algorithmic forecasting delivered 30 percent faster. Those figures describe one consulting case rather than a universal result. Still, they show why shared data and workflow redesign can create cumulative value.
The cheapest model can also be more expensive in practice. A weaker system may require repeated attempts, more checking, and additional corrections. A higher-quality model may reduce total labor enough to offset greater usage costs.
This is why value per unit of intelligence is more useful than model cost alone. The denominator should include model usage, implementation, employee time, review, and rework. The numerator should reflect accepted work or improved decisions.
CFOs should also distinguish productivity capacity from booked financial return. Saving analyst time does not automatically reduce expenses or create revenue. The organization must redirect that capacity toward other valuable work for the gain to become economic.
Forecast accuracy has a similar complication. A more accurate forecast creates value only when leaders use it to change inventory, hiring, marketing, infrastructure, or capital allocation. Measurement must connect the prediction to a decision.
The same principle applies to faster close cycles. Earlier figures can improve planning and risk response. Yet speed has little value if corrections arrive later or reviewers lose confidence in the numbers.
OpenAI’s scorecard therefore pressures both buyers and vendors. Buyers must define acceptable work before deployment. Vendors must support evaluations that reflect workflow outcomes, not only technical benchmarks.
The company has not released its own scorecard results. Readers should treat the framework as a proposed discipline rather than evidence of achieved return. OpenAI’s next step is to publish enough operational data to test its claims.
Three Signals Will Show Whether OpenAI’s Model Travels
The next test is whether OpenAI can turn an internal operating philosophy into measurable, repeatable results under real financial controls.
The first signal is published progress toward zero-day close and continuous forecasting. OpenAI should disclose cycle time, reconciliation coverage, exception rates, forecast accuracy, and revision frequency. Improvement across several periods would strengthen its claim that workflow redesign produces dependable value.
A lack of metrics would weaken the story. OpenAI does not need to reveal confidential financial details, but it can publish normalized improvements or process measures. Without them, outsiders cannot separate an advanced prototype from a stable operating system.
The second signal is evidence that governance scales with employee-built tools. OpenAI describes broad building activity across its finance team. The important question is whether those systems use consistent permissions, testing, source traceability, version control, and approval policies.
Independent audit evidence would be especially useful. So would a disclosed evaluation framework for IR-GPT, forecasting tools, and other high-impact workflows. Significant incidents or uncontrolled tool growth would challenge the builder model.
The third signal is how enterprise finance vendors and large customers respond. SAP, BlackLine, and other established suppliers already provide AI-assisted reconciliation and reporting. Their next releases will show whether configurable, context-rich workflows become standard product expectations.
Deloitte’s CFO priorities reinforce the pressure. Its survey covered 200 North American finance chiefs at companies with significant annual revenue. The research found technology transformation rising among executive priorities for 2026.
OpenAI’s five lessons offer a sharper test than broad adoption. Give employees access, redesign complete workflows, let domain experts build, preserve accountability, and measure dependable work. Each step depends on the others.
Access without direction creates scattered experiments. Workflow redesign without reliable data produces polished uncertainty. Employee building without governance creates operational risk, while controls without useful metrics can freeze promising work.
The openai what story is therefore not about removing finance professionals from financial decisions. It is about moving their judgment closer to live data while automating the preparation surrounding it.
For CFOs, the immediate action is to choose one consequential workflow and define its evidence, approvals, baseline, and outcome measures. Then test whether AI reduces total effort without weakening control. What would your finance team rebuild first if every generated answer had to remain traceable, reviewable, and economically measurable?



