top of page

CSBS AI Supervisory Framework Turns Voluntary Guidance Into an Examiner Playbook

5 days ago
12 min read

CSBS released its AI supervisory framework on September 16, giving state examiners a common playbook despite leaving adoption to individual agencies. The framework covers banks and nonbank financial companies supervised at the state level. It also arrives as federal regulators reconsider how older model and third-party guidance should apply to newer AI systems.

That contrast defines the story. The CSBS AI supervisory framework is voluntary and creates no new legal obligations. Yet it gives examiners specific questions, document requests, review procedures, and risk-tiering tools that can shape real examinations.

The federal approach is moving in a different direction. Revised federal model risk guidance expressly excludes generative and agentic AI from its scope. CSBS is giving state examiners a broader method for finding those systems, understanding their use, and deciding when deeper scrutiny is appropriate.

The framework therefore matters less as a new rule than as an operating manual. Financial institutions now have a clearer preview of what an examiner may ask about AI inventories, governance, vendors, customer impacts, and autonomous actions.

What the CSBS AI Supervisory Framework Changes

CSBS has converted broad AI governance principles into a practical examination process that state agencies can adopt or adapt.

The CSBS framework describes itself as a discretionary tool for state examiners. It helps them identify AI use, assess associated risks, and decide when existing supervisory resources should support a deeper review.

That limited description is important. CSBS did not announce a binding national standard for every state-regulated institution. Each state financial regulatory agency will decide how much of the framework enters its supervisory program.

The framework was approved by two CSBS committees in August 2026. CSBS then released it publicly on September 16. Its intended audience includes examiners overseeing state-chartered banks and state-licensed nonbank financial companies.

The package contains several connected resources. Its Core Examiner Guide supplies initial scoping questions, a document request list, and review procedures. Those procedures cover governance, oversight, AI inventories, use cases, generative AI, and other emerging applications.

An Examiner Work Program adds detail for conducting the review. Separate nonbank supplements address third-party risk, model risk, and consumer protection at nonbank financial services companies.

An optional AI Use Case Risk Tiering Worksheet supports case-by-case assessment. Instead of treating every AI tool as equally dangerous, examiners can consider the purpose, exposure, complexity, customer consequences, and controls surrounding a specific deployment.

The Source Support Document identifies the materials used to develop the framework. CSBS says those materials include the National Institute of Standards and Technology’s AI Risk Management Framework and Treasury’s AI terminology work.

That foundation connects state supervision with established risk-management concepts. The NIST AI framework, for example, organizes AI risk work around governing, mapping, measuring, and managing risk.

The CSBS package brings those concepts into an examination setting. An institution may need to show where AI is used, who approved it, which data it touches, and how someone monitors its performance.

Examiners can also ask how the institution controls vendor systems. Relevant records can include contracts, data-use terms, testing documents, risk assessments, and examples of customer-facing output.

For agentic systems, the review becomes more operational. An agentic system can take actions toward a goal with limited human direction. Examiners may examine its permitted actions, human checkpoints, activity logs, reversibility, and emergency shutdown controls.

That level of specificity changes preparation. A general AI policy will not answer every likely question. Institutions need evidence connecting written governance to actual systems, vendors, decisions, and customer outcomes.

The framework also reaches uses that may never fit a narrow mathematical definition of a model. A customer-service assistant, employee research tool, fraud workflow, or autonomous process can still create privacy, cybersecurity, compliance, and operational risks.

CSBS has effectively supplied a map from AI discovery to deeper examination. That map creates the article’s central tension: state agencies retain discretion, but supervised firms now know what evidence those agencies can request.

Why State-Regulated Firms Are Now on Notice

The immediate pressure falls on institutions that cannot produce a reliable inventory of their own AI use.

An AI inventory is a structured record of deployed or approved AI systems. It usually identifies each system’s owner, purpose, data, vendor, risk level, users, controls, and review history.

Many institutions already maintain model inventories and vendor lists. Those records may not capture every AI-enabled feature inside productivity software, customer platforms, security tools, or outsourced business processes.

The discovery problem grows when vendors add AI through routine product updates. A bank may approve a software service for one function, then receive summarization or automated decision features months later.

Employee adoption creates another gap. Staff can use public AI assistants, browser extensions, or embedded copilots before compliance teams classify the activity. Blocking every tool does not solve the governance question when authorized products also introduce AI.

The CSBS AI supervisory framework pushes institutions to connect these fragments. Examiners can begin with scoping questions and documents, then identify areas that deserve closer review.

A financial institution should expect questions about accountability. Examiners may want to know which executive or committee owns AI risk and who can suspend a system when problems appear.

That responsibility cannot remain purely technical. A credit tool can implicate fair-lending requirements. A customer assistant can create disclosure, privacy, or unfair-practices concerns. A fraud system can affect transaction access and error resolution.

The U.S. Government Accountability Office documented this wider risk landscape in its financial AI review. It identified potential benefits alongside biased decisions, data-quality failures, privacy concerns, and new cybersecurity threats.

The GAO also found that financial regulators primarily oversee AI through existing laws, guidance, and risk-based examinations. CSBS follows that pattern. Its framework helps examiners connect an AI use case with existing supervisory authorities and established risk categories.

That structure pressures nonbank firms as well as banks. Mortgage companies, money transmitters, consumer lenders, and other licensed businesses often operate across several states.

A multi-state company can face different implementation choices from different agencies. One regulator may incorporate the full framework, while another uses selected questions or existing examination procedures.

The practical response is not to build a separate governance program for every jurisdiction. Firms need a defensible evidence base that can support several supervisory approaches.

That base starts with the inventory, but it cannot end there. Each material use case needs a named owner, an approved purpose, documented limits, and controls matching its potential impact.

Third-party systems require particular attention. A vendor’s security questionnaire may not explain how its AI feature was tested, whether customer data trains another model, or how outputs change after updates.

Institutions also need records of human oversight. Saying that a person remains involved does not explain what the person reviews, when intervention occurs, or whether the person can reverse an automated action.

For knowledge work, documentation is easier when policies, approvals, vendor materials, and testing records remain searchable together. A maintained AI knowledge base can support that work, although it does not replace formal compliance systems.

The institutions under greatest pressure are not necessarily those with the most AI. They are those unable to explain where AI operates, why it operates, and which controls follow its risk.

That distinction reflects a risk-based approach. A low-impact drafting assistant should not receive the same review as an automated lending decision or autonomous payment action.

However, an institution must first identify both systems. Without discovery and inventory, it cannot make a credible risk distinction.

AI Banking Supervision Now Faces a Consistency Test

The main contest is between supervisory flexibility and the need for consistent expectations across states and federal agencies.

CSBS deliberately preserves examiner discretion. The framework accounts for an institution’s size, complexity, risk profile, and actual use of AI.

That tailoring can protect smaller institutions from disproportionate compliance work. A community bank with limited AI exposure should not need the same program as a complex financial group deploying models across major business lines.

State flexibility also reflects the structure of American financial supervision. State agencies oversee different combinations of banks, lenders, money transmitters, mortgage companies, and other regulated businesses.

The tradeoff is predictability. When every agency decides how to use the framework, institutions cannot assume that examinations will follow an identical scope or threshold.

One examiner may treat the Core Examiner Guide as a preliminary screen. Another agency could build its document requests and examination work around the entire package.

This concern is not theoretical. CSBS Chair Rhoshunda Kelly has previously called for consistent implementation of supervisory expectations and clearer coordination between state and federal regulators.

The federal position creates another layer. In April 2026, the Federal Reserve, Office of the Comptroller of the Currency, and Federal Deposit Insurance Corporation revised their joint model risk guidance.

The revised federal guidance emphasizes a risk-based approach tailored to model exposure, organizational size, and operational complexity. It covers development, validation, monitoring, governance, inventories, documentation, and vendor products.

However, that guidance says generative and agentic AI are novel and rapidly evolving. It therefore excludes those systems from the document’s formal scope.

This exclusion does not mean banks can use generative AI without controls. The guidance says broader risk-management and governance practices should determine appropriate oversight for tools outside its scope.

It does create a boundary problem. A traditional statistical system may fall squarely under federal model risk guidance. A generative assistant connected to the same workflow may require controls drawn from other frameworks.

CSBS addresses the discovery side of that gap. Its examiner guide explicitly includes generative AI and emerging uses within initial review procedures.

The two approaches are not direct contradictions. Federal agencies narrowed one specific body of model guidance, while CSBS created a broader examiner framework that points back to existing supervisory resources.

Still, firms must translate the distinction into operations. They need to know when a system is a covered model, an AI-enabled process, a third-party service, or several of those categories together.

Vendor dependencies make the classification harder. A financial institution may not receive access to training data, underlying code, evaluation methods, or complete performance evidence.

Federal guidance says proprietary vendor components can create validation challenges. It still expects organizations to understand conceptual soundness, design, development data, performance, and continuing fitness for purpose.

Generative AI vendors may not provide all that information. Their products can also change more frequently than conventional banking models, sometimes through updates controlled entirely by the provider.

CSBS gives state examiners a route for examining that uncertainty. Its nonbank supplements and document requests can connect vendor oversight with consumer protection, operational controls, and model risk where applicable.

The federal agencies are also revisiting third-party oversight. Their September 2026 third-party proposal seeks a principles-based approach tailored to individual relationships.

That proposal is nonbinding, and public comments remain part of the process. It nevertheless shows that vendor oversight is changing alongside AI supervision.

Institutions therefore face several moving layers. They must track state implementation, federal model boundaries, developing third-party guidance, and laws governing the underlying financial activity.

The best defense is a control system that follows the use case rather than its label. Calling software a copilot, assistant, algorithm, or workflow should not decide the entire risk analysis.

Purpose, authority, data access, customer impact, and reversibility offer more durable criteria. Those factors remain relevant even when a vendor changes the underlying model.

A Practical Playbook Still Leaves Legal Uncertainty

The framework improves examiner readiness, but it does not resolve the authority, enforcement, or standardization questions surrounding AI banking supervision.

CSBS states that the framework is discretionary. It also says each state agency controls whether and how the materials enter its supervisory program.

That limits claims about immediate nationwide effect. Publication does not mean every state examiner will begin using every worksheet during the next examination cycle.

The framework also creates no new legal obligation by itself. An examiner finding can still connect AI activity to existing laws, safety and soundness concerns, consumer protection requirements, or established supervisory expectations.

This distinction matters for regulated firms. A document labeled voluntary can influence how examiners gather evidence and evaluate whether another binding requirement has been violated.

The practical burden may arrive before any formal enforcement action. Institutions can receive broader document requests, face follow-up questions, or need specialist staff to explain complex systems.

Smaller firms may struggle with that preparation. They depend heavily on core processors and technology vendors, but often have less bargaining power to obtain detailed AI documentation.

A vendor may provide a standard audit report without revealing model behavior, prompt handling, subcontractors, or update controls. The institution remains responsible for understanding risks it can only partly observe.

Risk tiering introduces another uncertainty. The optional worksheet encourages proportional review, but risk classifications depend on assumptions about impact, autonomy, data, and control effectiveness.

Two reviewers can reasonably classify the same use case differently. A drafting assistant might appear low risk until employees enter confidential customer data or rely on fabricated legal analysis.

Customer-facing output adds further complexity. A chatbot may not make formal credit decisions, yet inaccurate answers can affect complaints, fees, account access, or consumers’ understanding of their rights.

Agentic systems raise higher operational stakes. Human approval is meaningful only when the reviewer has enough information, time, authority, and expertise to stop a harmful action.

Logs also need substance. A record showing that an AI system acted does not prove that the decision was appropriate or that its reasoning can be reconstructed.

The framework’s reliance on existing supervisory resources is sensible, but it can produce overlapping reviews. One AI deployment may touch cybersecurity, privacy, third-party risk, operational resilience, consumer compliance, and model governance.

That overlap can improve coverage when teams coordinate. It can also produce duplicated requests or conflicting control expectations when responsibilities remain unclear.

Evidence quality presents another challenge. Institutions should distinguish vendor claims from independent testing, internal evaluation, production monitoring, and customer outcomes.

A demonstration supplied by a provider does not establish performance in a bank’s environment. Tests should reflect the institution’s users, data, workflows, failure modes, and legal obligations.

Generative outputs require evaluation beyond conventional accuracy. Reviewers may need to assess fabrication, harmful instructions, data leakage, inconsistent treatment, prompt attacks, and resistance to unauthorized actions.

Continuous monitoring also becomes harder when systems change quickly. A control assessment performed before launch can lose relevance after a model, data source, prompt, or integration changes.

This is why inventories must include lifecycle triggers. Institutions should define which changes require renewed testing, legal review, vendor assessment, executive approval, or suspension.

The CSBS AI framework explained as a simple compliance checklist would miss that point. Its value comes from structuring inquiry, not certifying a system as permanently safe.

There is also no evidence yet that all state agencies will interpret the package consistently. Early examination practice will reveal whether the framework produces convergence or another layer of jurisdictional variation.

Institutions should avoid two opposite mistakes. They should not treat a discretionary framework as an immediate binding rule, and they should not dismiss it as irrelevant guidance.

The more accurate reading sits between those positions. CSBS has created a common supervisory vocabulary and a reusable process. Adoption, interpretation, and enforcement remain decentralized.

Three Signals Will Show Whether the Framework Matters

The framework’s real influence will become visible through state adoption, examination practice, and coordination with developing federal guidance.

The first signal is formal or operational adoption by state agencies. Public notices, revised examination manuals, new document requests, and examiner training will show where the framework becomes active.

A state does not need to issue a new regulation to make the framework consequential. Incorporating its questions into routine examinations can change how firms prepare and how examiners identify higher-risk uses.

Broad adoption would strengthen the case for building one enterprise-wide evidence system. Fragmented adoption would increase the need to map governance records against jurisdiction-specific expectations.

The second signal is the content of actual examinations. Institutions should watch whether examiners focus first on inventories and governance or move quickly into testing particular use cases.

Requests for vendor contracts, model documentation, output samples, incident records, and risk assessments will reveal which sections carry the most practical weight.

Attention to agentic controls would be especially significant. Questions about permitted actions, human checkpoints, audit logs, reversibility, and shutdown mechanisms would indicate deeper operational scrutiny.

Examinations will also reveal how reviewers treat informal employee use. A framework centered only on approved enterprise systems could miss substantial data and decision risks.

The third signal is federal coordination. Federal agencies have narrowed current model risk guidance while preparing additional work on AI and third-party relationships.

The OCC said in April that agencies planned further information gathering focused on AI, including generative and agentic systems. Any resulting request or guidance can clarify where the federal approach overlaps with CSBS.

A coordinated approach would reduce uncertainty for state-chartered banks that also face federal supervision. Divergent definitions or documentation standards would create more translation work.

The GAO’s continuing attention to credit-union oversight offers another useful indicator. Its recommendation for broader model risk guidance shows that supervisory coverage remains uneven across institution types.

None of these signals requires firms to wait. The framework already identifies preparation work that supports sound governance regardless of a particular agency’s adoption decision.

Institutions can verify their inventories, assign accountable owners, classify use cases, and document approval criteria. They can also test whether vendor records answer the questions an examiner is likely to ask.

High-impact systems deserve outcome monitoring linked to their real purpose. A fraud tool should be evaluated against fraud outcomes and customer disruption, not only technical benchmarks.

Customer-facing systems need representative output review. Teams should preserve evidence of failures, corrections, complaints, and changes rather than recording only successful demonstrations.

Agentic deployments require clearly bounded authority. Organizations should define which actions need approval, which can be reversed, and who can disable the system during an incident.

Boards and senior managers do not need to master every model architecture. They do need accurate reporting on material uses, unresolved risks, incidents, vendor limitations, and accepted exceptions.

The same discipline helps employees. Clear rules should distinguish approved systems from prohibited uses and explain how workers can report unexpected behavior.

The CSBS AI supervisory framework will matter most if it changes these operating habits. A well-organized policy folder alone will not demonstrate control over deployed systems.

For technology vendors, the message is equally direct. Financial institutions will increasingly request clearer documentation about data handling, testing, system changes, subcontractors, logs, and human intervention.

Vendors that cannot provide credible evidence can become harder to approve, even when their products perform well. Procurement teams need contractual rights that support ongoing oversight after deployment.

For financial institutions, the next step is a focused readiness review. Can the organization identify every material AI use, connect it to an owner, and produce evidence supporting its risk classification?

Can it explain what a vendor system does without relying entirely on sales materials? Can it show what happens when the system fails or takes an unauthorized action?

Those questions capture the practical meaning of the new framework. State implementation remains uncertain, but the expected evidence is becoming clearer.

The next examination may not use every CSBS document. It can still ask the same underlying question: does the institution understand and control the AI operating inside its business?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page