top of page

The Real AI Power Test: Can Governments Inspect, Intervene, and Walk Away?

Sep 2
12 min read

Google News surfaced a sharper sovereign AI test this week: governments must prove they can inspect critical systems, intervene during failures, and leave suppliers.

That test comes from a newly published systematic review by researchers Raghu Raman and Prema Nedungadi. Their work challenges a familiar government playbook built around national models, domestic data centers, and localization rules.

The researchers call their alternative “credible sovereignty,” meaning demonstrable control over deployed AI rather than legal authority or domestic branding alone. The distinction puts governments opposite a difficult reality: most national AI programs still depend on external clouds, chips, models, software, and expertise.

A country can require public data to remain inside its borders while renting the infrastructure that processes it. It can commission a national language model while depending on foreign accelerators and proprietary development tools. It can regulate frontier models without possessing enough independent expertise to evaluate them.

The resulting conflict is not simply domestic technology versus foreign technology. It is declared control versus operational control.

That conflict matters to public agencies, regulated industries, developers, and enterprise buyers. When an AI service changes behavior or becomes unavailable, legal jurisdiction does not automatically provide technical access, migration capacity, or a usable replacement.

What the Google News Report Actually Changed

The new research turns AI sovereignty from a political label into a set of operational tests.

The original Google News report covered a study published on August 24, 2026. The study examines how control operates across infrastructure, data, models, procurement, and institutional accountability.

The authors define credible sovereignty through three capabilities: inspection, intervention, and accountability.

Inspection asks whether authorized institutions can understand and audit a deployed system. That includes access to relevant documentation, evaluation results, incident records, data governance controls, and system behavior.

Intervention concerns what happens after officials identify a problem. A government needs the practical ability to suspend a service, change its configuration, limit a feature, transfer a workload, or replace the provider.

Accountability asks whether responsibility remains identifiable and enforceable across the supply chain. That question becomes difficult when agencies, cloud providers, model developers, integrators, and subcontractors each control different components.

These capabilities sound basic. Yet each can disappear inside a modern AI service.

An agency might receive extensive compliance documentation without gaining access to the model or evaluation environment. It might possess a contractual suspension right but lack the technical capacity to continue essential services elsewhere.

A regulator might demand corrective action from a local deployer while the underlying model provider operates from another jurisdiction. Export restrictions or licensing changes can also affect infrastructure that domestic law does not control.

The study therefore separates formal authority from actual leverage. A sovereign claim becomes credible only when institutions can produce evidence that control survives deployment, incidents, and supplier conflict.

This framework changes the sovereign AI debate because it rejects a single ownership test. Domestic ownership can support control, but a national flag on a model does not answer who maintains its dependencies.

The study also rejects the idea that sovereignty requires complete self-sufficiency. Few countries can reproduce every chip, cloud layer, foundation model, cybersecurity service, and evaluation tool.

Instead, dependencies must remain visible, contestable, and replaceable. A government can rely on external technology while retaining meaningful authority, provided it can verify conditions and act when those conditions fail.

That is the central shift behind the headline. Sovereign AI is no longer measured by what a government launches. It is measured by what that government can do after launch day.

A National AI Label Does Not Guarantee Control

Control at one layer of the AI stack can coexist with deep dependence at every other layer.

The peer-reviewed sovereignty study reviewed English-language scholarship published from 2020 onward. Its authors initially assessed 152 records and retained 88 for close reading.

They used topic modeling as a structured aid, then paired it with close analysis. The resulting framework identifies four broad governance logics involving infrastructure, geopolitical blocs, European regulation, and community claims in the Global South.

The paper’s most useful contribution is its layered model. AI sovereignty can strengthen at the data layer while remaining weak at the compute layer.

Consider a public health agency using a domestically hosted assistant. Patient records might stay in a national data center, satisfying a residency requirement.

However, the assistant’s model could depend on software maintained abroad. Its accelerators might require externally controlled firmware, while its safety evaluations rely on tools unavailable to the agency.

The deployment is locally hosted but not fully inspectable. A licensing dispute, export restriction, or terminated support agreement could also make intervention difficult.

The reverse arrangement presents different problems. A government might use an openly documented model while hosting it on a foreign hyperscale cloud.

Officials can inspect parts of the model, but they still depend on the provider for identity controls, logging, storage, networking, and operational continuity. Moving the workload may require months of engineering work.

Model weights create another control point. Weights are the learned numerical parameters that determine how a trained model responds.

An agency using a closed model generally cannot inspect those parameters directly. Even access to the weights would not automatically reveal why a complex model produced a specific answer.

Meaningful inspection therefore requires several forms of evidence. These can include model documentation, evaluation access, change records, system logs, data provenance, security controls, and independent testing.

Intervention also requires more than an emergency stop button. Suspending an unsafe service is useful, but public agencies often provide functions that cannot simply disappear.

A government needs continuity planning, substitute systems, exportable data, compatible interfaces, and people who know how to operate the replacement. Otherwise, its right to terminate becomes a right to disable its own service.

This is where “walk away” becomes the hardest test. Exiting a provider requires technical portability, contractual permission, budget capacity, and a viable alternative.

A procurement team may negotiate data export while overlooking application logic, model-specific prompts, evaluation suites, vector indexes, or identity integrations. Those dependencies can make migration prohibitively slow.

The same lesson applies to enterprise buyers. A contract stating that customers own their data does not establish whether they can reproduce the service elsewhere.

Teams should distinguish legal ownership, practical exportability, and operational substitutability. Each answers a different question about control.

For knowledge workers, the issue is equally concrete. AI services increasingly hold meeting records, research, decisions, and working context.

A structured personal knowledge base can make source material easier to preserve and review. However, organizations must still examine export formats, model dependencies, access controls, and retention terms.

Sovereignty is therefore a chain property. A deployment remains controllable only when its critical dependencies support inspection, intervention, and credible exit.

Procurement Contracts Are Becoming AI Policy

Governments often gain more operational leverage through procurement clauses than through national model announcements.

AI governance debates usually emphasize legislation, safety institutions, and computing investment. Procurement receives less attention because it looks administrative.

Yet contracts determine the rights available during an actual failure. They decide whether officials receive incident reports, documentation, audit access, export assistance, and advance notice of material changes.

A useful procurement test begins before deployment. Agencies need to identify which components they must inspect and what evidence the supplier must provide.

A general promise of transparency is insufficient. The contract should specify documentation, evaluation records, logging, security evidence, subcontractor disclosures, and response timelines.

The second test concerns intervention rights. Agencies should define conditions for suspending features, limiting access, preserving records, or requiring corrective measures.

Those rights must connect to technical controls. A contract cannot support rapid intervention if only the supplier can access the relevant configuration.

The third test is exit readiness. That includes machine-readable exports, documented interfaces, migration assistance, deletion verification, and a timeline for continued access during transition.

Portable data alone may not recreate an AI workflow. Agencies should also identify prompts, retrieval indexes, evaluation cases, access policies, and process logic needed by a replacement.

The European Commission has started translating these ideas into measurable cloud procurement. Its 2026 cloud framework evaluates providers through 48 criteria across eight categories.

Those categories cover strategic, legal, data and AI, operational, supply-chain, technological, security, and environmental considerations. The framework also assigns Sovereignty Effectiveness Assurance Levels.

In April 2026, the Commission awarded four sovereign cloud contracts worth up to €180 million over six years. It said the multi-provider structure was intended to increase resilience and avoid dependence on one supplier.

The procurement shows what operationalization can look like. Sovereignty becomes a collection of defined requirements rather than an unrestricted political claim.

It also shows that foreign technology and credible control are not mutually exclusive. One selected consortium uses an environment based on Google Cloud technology but operated by European companies under specific controls.

That arrangement does not settle every sovereignty question. It illustrates how governments can evaluate external dependencies rather than treating their presence as an automatic failure.

Still, procurement clauses have limits. A small agency can demand audit rights while lacking qualified staff to exercise them.

A government can require workload portability while maintaining no alternative environment. It can request incident reports without possessing the expertise needed to interpret them.

Contracts create options, not capacity. Those options become meaningful only when governments fund technical teams, evaluation infrastructure, cybersecurity operations, and migration exercises.

This distinction also matters for vendors. Buyers increasingly need evidence that a system remains governable after integration.

Providers that document interfaces, preserve exportability, support independent testing, and disclose critical dependencies can reduce procurement risk. Claims about national hosting alone will face greater scrutiny.

The emerging market signal is clear. Public-sector AI contracts will increasingly treat auditability and exit readiness as product capabilities rather than legal appendices.

Regulation Can Be Strong on Paper and Weak in Practice

A government cannot enforce what its institutions cannot independently examine, reproduce, or challenge.

The European Union offers the clearest test of the difference between regulatory power and operational capacity.

The EU AI Act gives regulators meaningful authority over covered systems and general-purpose AI models. It requires technical documentation from model providers and establishes additional duties for models presenting systemic risks.

Those duties include model evaluations, adversarial testing, systemic-risk assessment, incident reporting, and cybersecurity measures. The law also gives relevant authorities routes to request information and evaluate compliance.

These are significant inspection and accountability tools. They create enforceable obligations that voluntary commitments cannot match.

However, legal access does not automatically produce technical understanding. Regulators need secure facilities, evaluation methods, specialized staff, computing resources, and procedures for handling confidential information.

Frontier models introduce a particular challenge. Evaluations can become obsolete as models receive updates, tools, new system prompts, or different deployment settings.

A result obtained during predeployment testing may not capture behavior inside a public service. Inspection must therefore include continuous monitoring and deployment-specific evidence.

Intervention presents another gap. A regulator may restrict a model, yet affected agencies and businesses still need replacement services.

A rapid withdrawal could protect users from one risk while disrupting hospitals, benefits systems, courts, or other essential operations. Governance must account for both safety and continuity.

The law also operates across a fragmented supply chain. A foundation-model provider controls core capabilities, a cloud operator controls infrastructure, and an integrator configures the final application.

A public agency controls the use case but may not see changes made farther upstream. Responsibility can become ambiguous when harmful behavior emerges from interactions between those layers.

Effective enforcement needs clear duties at each control point. It also needs evidence that can travel between organizations without becoming meaningless or exposing protected information.

This is why the credible sovereignty study treats accountability as a separate capability. Inspection can reveal a problem without showing who must fix it.

Governments must also resist confusing extensive paperwork with visibility. A vendor can deliver thousands of pages while withholding the evidence needed to test an important claim.

Useful documentation should connect system properties to verifiable artifacts. Evaluation results need methods, conditions, limitations, and change histories.

Incident reports require consistent definitions. Otherwise, providers can classify similar failures differently, making comparison and trend analysis difficult.

Independent evaluation can reduce this information imbalance. Yet evaluators also depend on access, expertise, and stable testing environments.

The Commission acknowledged this capacity challenge in 2026 when discussing the need for more top-tier evaluation expertise inside Europe. Many leading third-party model evaluators have historically operated outside the region.

The broader lesson extends beyond Europe. Countries can copy legal language without replicating institutional capacity.

A regulation promising access, audits, and corrective action remains incomplete if agencies cannot use those powers under operational pressure. Rules and capacity must develop together.

Sovereign AI Can Replace Foreign Dependence With Domestic Concentration

Reducing reliance on global platforms does not guarantee competition, accountability, or wider access at home.

Sovereign AI programs often begin with a reasonable concern. Governments do not want critical public systems exposed to foreign legal demands, commercial decisions, or geopolitical restrictions.

Domestic infrastructure can reduce some of those risks. Local expertise can improve language coverage, cultural knowledge, security response, and institutional confidence.

However, sovereign systems require large capital investments and scarce technical skills. These requirements can concentrate control among a small number of companies or state-linked institutions.

Smaller agencies may lose choice if one approved domestic provider becomes the default. Universities, nonprofits, and local businesses can also face limited access to computing resources.

The label changes, but the dependency persists. A foreign gatekeeper becomes a domestic gatekeeper.

This creates a different accountability problem. Governments may apply weaker scrutiny to national champions because those firms support strategic goals.

Officials can also treat criticism of a domestic provider as opposition to national competitiveness. That dynamic makes independent evaluation more important, not less.

The credible sovereignty framework addresses this risk by asking “for whom” control exists. A state can gain authority while communities and smaller institutions lose meaningful participation.

Data sovereignty makes that conflict visible. A government might claim national control over datasets containing Indigenous languages, community records, or culturally sensitive knowledge.

Yet national jurisdiction does not erase the rights of affected communities. Credible governance requires consent, participation, contestability, and remedies.

Developing economies face additional constraints. Building a domestic language model does not eliminate reliance on imported chips, cloud software, security tools, or external expertise.

Localization rules can raise costs without producing substitute capacity. They can also limit access to useful cross-border services when domestic alternatives remain immature.

The right response is not unrestricted dependence. It is selective control based on workload sensitivity and realistic institutional capacity.

Highly sensitive government systems may justify stricter hosting, access, and supply-chain requirements. Lower-risk applications can use interoperable external services with audit and exit safeguards.

This risk-based approach avoids treating every AI workload as strategically identical. A public chatbot providing office hours does not require the same controls as a system affecting benefits or criminal justice.

The latest government AI outlook found that 32 of 36 surveyed OECD countries reported AI training programs for government workers.

Training is a necessary foundation, but general AI literacy is only the beginning. Credible sovereignty needs specialists who can evaluate models, negotiate contracts, secure infrastructure, and test migrations.

Governments should also examine how capacity is distributed. A central AI unit may become highly capable while local agencies remain unable to challenge vendors.

Shared evaluation services can help, provided agencies retain clear authority and access. Common procurement templates can also raise minimum standards without requiring every institution to build its own framework.

The skeptical conclusion remains important. The new research proposes a compelling test, but it has not yet validated that test against a large set of real deployments.

The authors explicitly present their indicators as an initial operationalization. Future research must determine which measures predict continuity, accountability, and effective intervention.

A supplier can satisfy a checklist while remaining difficult to replace. Governments will need exercises, incidents, and comparative evidence to test whether documented sovereignty survives contact with reality.

Three Signals Will Show Whether Governments Can Walk Away

The next phase of sovereign AI will be judged through enforcement, migration tests, and evidence from deployed public systems.

The first signal is whether regulators use their inspection powers against leading general-purpose AI providers.

Documentation requests alone will reveal little. The meaningful test involves independent evaluations, access to relevant evidence, and corrective measures tied to specific findings.

Successful inspection would strengthen the case that legal authority can become operational control. Repeated delays or dependence on provider-run tests would expose the capacity gap.

The second signal is whether major public buyers conduct real portability exercises.

A migration clause has limited value until an agency exports its data, reconstructs a workflow, transfers access policies, and operates the replacement. Governments already test disaster recovery and cybersecurity response; supplier exit deserves similar treatment.

A successful exercise does not require abandoning the original provider. It needs to demonstrate that essential services can continue under documented conditions.

Failed migrations would still provide useful evidence. They would reveal hidden dependencies in formats, interfaces, evaluation suites, identity systems, or staff knowledge.

The third signal is whether procurement frameworks produce genuine supplier diversity.

Awarding several contracts is only the starting point. Agencies must distribute workloads, preserve interoperability, and prevent one platform from becoming unavoidable through accumulated integrations.

Diversity should also include evaluation and assurance providers. A government remains exposed if every approved system depends on one testing organization or one proprietary benchmark.

These signals matter beyond public policy. Enterprise buyers face the same concentration risks, although they operate under different legal duties.

Technology leaders should ask whether they can audit important system behavior, restrict unsafe functions, preserve decision records, and move critical workflows. These are governance questions and architecture questions at the same time.

Developers should treat portability as an engineering property. Standard interfaces, documented dependencies, reproducible evaluations, and separable data layers reduce the cost of future intervention.

Knowledge workers should preserve access to the source material behind AI-generated answers. That practice supports review when models change and reduces dependence on one provider’s memory system.

Google News brought attention to a research paper, not a completed government reform. The paper’s value lies in giving officials and buyers a sharper standard.

The real AI power test is not whether a country can announce a national model. It is whether institutions can inspect a system when claims conflict, intervene before harm spreads, and maintain essential work after leaving a supplier.

Over the next several months, watch what governments test rather than what they launch. Do regulators gain direct evidence, do agencies rehearse migrations, and do procurement systems preserve meaningful alternatives?

Those actions will show whether sovereign AI represents operational authority or another layer of branding. Readers evaluating any critical AI service can apply the same question now: if the provider’s behavior, terms, or availability changed tomorrow, could your organization understand the change, respond effectively, and keep working somewhere else?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page