top of page

AIUC AI Agent Safety Gets $40 Million, but Certification Is Not a Guarantee

Sep 16
13 min read

AIUC raised $40 million to turn AI agent safety into something enterprises can test, certify, and insure before autonomous systems receive sensitive access. The Series A gives Artificial Intelligence Underwriting Company fresh capital to expand a model built around audits, recurring technical evaluations, and liability coverage.

That model challenges the usual approach to enterprise AI risk. Vendors often document their controls, buyers conduct separate reviews, and deployment teams add monitoring after approval. AIUC wants an independent standard to connect those activities and attach financial consequences when an agent causes covered damage.

The company was founded by Rune Kvist, Anthropic’s first product and go-to-market hire, and Rajiv Dattani, a former METR chief operating officer. Their central argument is direct: improving agent capabilities will not unlock enterprise adoption when buyers cannot measure or transfer the associated risk.

That creates the real contest behind the funding announcement. AIUC is not simply competing with another startup. It is testing whether independent certification and insurance can control risks that vendor assurances and conventional compliance reviews cannot fully address.

AIUC’s $40 Million Bet on AI Agent Safety

The financing turns AIUC’s certification model from an experiment into a serious attempt at enterprise risk infrastructure.

AIUC announced the Series A on September 15, 2026. Ribbit Capital led the round, while First Harmonic participated. The company previously raised a $15 million seed round led by NFDG, bringing its reported total funding to $55 million.

The company plans to extend its work from application-level agents toward frontier models. That expansion matters because model behavior increasingly shapes what agents can access, decide, and execute. However, AIUC has produced its most visible evidence at the application layer.

AIUC names Cursor, ElevenLabs, Harvey, KPMG, Lovable, UiPath, and Intercom’s Fin among organizations carrying its AIUC-1 trust mark. These products span coding, legal work, customer support, automation, and voice generation.

The range helps AIUC argue that agent risk is not confined to one category. A coding agent can expose credentials or introduce insecure dependencies. A customer service agent can disclose personal information or invent a policy.

The company says AIUC-1 evaluates agents against approximately 5,000 risk and attack combinations tailored to each business context. Those tests cover behavior involving jailbreaks, hallucinations, data leaks, and other operational failures.

A jailbreak is an attempt to bypass an AI system’s restrictions through crafted instructions. A hallucination occurs when the system generates unsupported information while presenting it as reliable.

According to the company, test results feed into a report of roughly 100 pages. AI agents run parts of the evaluation and analyze results, while human reviewers verify the final audit.

The process reflects an important shift in enterprise evaluation. Traditional reviews examine policies, access controls, retention rules, and incident procedures. AIUC also tests what an agent actually does when placed under adversarial pressure.

The company’s funding announcement says more than 250 security and risk leaders help shape the standard. That group represents prospective buyers rather than only AI vendors.

AIUC says certified agents undergo independent audits and quarterly recertification. Its public documentation also describes annual reviews of operational controls and at least quarterly technical testing.

That distinction is important. An annual audit can capture policies and established practices, while recurring tests can address changing models, prompts, tools, and attack techniques.

AIUC is effectively betting that agent behavior changes too quickly for a one-time certificate. A model update, new integration, or expanded permission can alter the risk after an initial review.

The $40 million round does not prove that AIUC-1 will become a lasting standard. It does show that investors see enterprise trust as a separate market, not merely a feature vendors can handle internally.

That market will depend on whether buyers treat certification as meaningful evidence during procurement. It will also depend on whether insurers can translate test results into useful coverage terms.

Why Smarter Agents Create a Harder Approval Problem

Enterprise resistance increasingly comes from uncertainty about behavior, responsibility, and loss, not a shortage of impressive AI capabilities.

AI agents differ from conventional chatbots because they can take actions through software tools. Depending on their permissions, they can edit code, query databases, send messages, or update business systems.

That autonomy raises the cost of a mistake. A false answer can become an incorrect transaction. A manipulated prompt can become unauthorized access to documents or external services.

Kvist told TechCrunch that banks, hospitals, governments, and militaries no longer reject AI simply because models lack intelligence. He argued that these organizations cannot guarantee what deployed systems will or will not do.

That statement is a company founder’s assessment, not an independent measure of market demand. Still, it captures a familiar enterprise problem: a successful pilot does not automatically satisfy security, legal, and procurement teams.

Pilots usually operate with limited data, users, and integrations. Production deployments connect systems to real workflows, customer information, intellectual property, and regulated records.

Security reviews must therefore assess both the vendor and the deployed configuration. The underlying model matters, but so do retrieval systems, prompts, identity controls, tools, and human approval steps.

Those elements can change independently. A vendor can update a model while the customer changes permissions. A previously safe workflow can become riskier when an agent gains access to production systems.

AIUC says many agents complete pilots but stall during security review because buyers lack credible evidence about security and reliability. Its proposed answer is a common test framework paired with independent verification.

The pressure falls first on AI vendors selling into large organizations. Each buyer can otherwise request different questionnaires, tests, contract language, and technical evidence.

That fragmented process consumes time without necessarily producing comparable results. A shared standard could reduce repeated work if buyers accept its scope and methodology.

Enterprise buyers face the opposite pressure. They want faster access to useful automation, but approval teams remain responsible when an agent leaks information or takes an improper action.

Internal teams cannot rely entirely on a vendor’s claims. They also cannot reproduce every adversarial evaluation for every agent under consideration.

The NIST AI framework offers a voluntary structure for identifying and managing AI risks. However, it does not certify individual agents or guarantee their behavior in a specific deployment.

SOC 2 reports provide another useful reference point. They evaluate controls related to security, availability, processing integrity, confidentiality, and privacy.

AIUC borrows the idea of a recognized assurance document, but it targets a different problem. Its tests focus on how an agent behaves when exposed to risky instructions and operational conditions.

This does not make SOC 2 obsolete. AIUC-1 and conventional assurance reviews examine different layers, and enterprises will likely require both.

The resulting approval stack may become more demanding, not less. Buyers could request conventional security evidence, AI-specific testing, monitoring plans, contractual protections, and insurance.

AIUC succeeds only if its certification simplifies that stack enough to justify another review. A badge without procurement recognition would add paperwork rather than remove it.

How AIUC-1 Combines Tests, Audits, and Insurance

AIUC’s key mechanism is not a single safety test, but a feedback loop connecting technical evidence with certification and financial exposure.

The first component is a standard covering security, safety, reliability, privacy, accountability, and wider social risks. AIUC describes AIUC-1 as a baseline designed specifically for agents.

The second component is technical evaluation. Testers try to trigger unsafe behavior, expose information, manipulate instructions, or reveal unreliable outputs under defined conditions.

The third component is an independent audit. AIUC has begun authorizing outside auditors, including Schellman, to review evidence and determine whether organizations meet the standard.

The fourth component is insurance. AIUC offers liability coverage intended to protect vendors and enterprise customers when an agent failure creates a covered business loss.

Insurance changes the incentive structure because risk receives a financial expression. An insurer needs evidence to decide which losses qualify, which controls matter, and which systems present greater exposure.

That is where AIUC’s model differs from a voluntary trust badge. The company wants evaluation results to influence the availability and terms of coverage.

Kvist previously told Fortune that insurance can reward steps that reduce risk. He compared those incentives with vehicle safety features that affect conventional insurance decisions.

The analogy is useful but incomplete. Vehicles operate within mature testing regimes, extensive loss histories, and established legal frameworks. Agentic AI lacks comparable historical data.

Insurers need credible information about failure frequency and severity. They also need clear boundaries between model defects, deployment errors, customer misuse, and third-party attacks.

AIUC can collect structured evidence through audits and evaluations. Over time, claims data could reveal which test results actually predict costly incidents.

That relationship has not yet been publicly demonstrated at scale. AIUC has not disclosed enough claims history to establish how strongly certification scores correlate with real-world losses.

The company nevertheless has a plausible starting mechanism. Testing identifies known weaknesses, audits check controls, and insurance introduces financial accountability for defined outcomes.

Its work with Cursor illustrates the approach. AIUC says the coding agent underwent thousands of evaluations across 12 risk categories.

The tests examined secrets leakage, hidden prompt injection, and insecure coding defaults. They also covered both the desktop development environment and cloud agents.

AIUC’s Cursor certification says Schellman reviewed operational controls alongside technical behavior. Those controls included data retention, access management, incident response, and human oversight.

The test configuration reportedly used rules, hooks, ignored-file settings, and automated review. That matters because agent safety depends on the combined system rather than the language model alone.

Consider a malicious instruction hidden inside a repository file. A coding agent might encounter that text while inspecting a project and treat it as an authorized command.

A useful evaluation must test whether the agent follows the hidden instruction, reveals a secret, or installs an unsafe dependency. It should also examine whether surrounding controls contain the result.

That is more concrete than asking whether an AI provider maintains a security policy. It evaluates a behavior that can directly affect a development team.

The same logic applies to customer service. Evaluators can test whether an agent discloses account information, invents refund rules, or follows instructions supplied by an unauthorized user.

AIUC says its standard changes as threats, capabilities, and regulations evolve. Quarterly updates and recurring tests are designed to stop certification from becoming a static snapshot.

Frequent changes introduce another challenge. Buyers need to know which version of the standard applied, what product configuration was tested, and when material changes require another review.

Without that traceability, a certificate can outlive the system it describes. AIUC will need strict scoping rules as customers deploy agents across more integrations and use cases.

Certification Cannot Guarantee That an Agent Will Behave

AIUC-1 can provide evidence about tested conditions, but it cannot guarantee safe behavior across every prompt, user, integration, or future model update.

AI systems operate across an enormous range of possible inputs. Evaluators can sample important attack patterns, but they cannot exhaust every interaction an agent might encounter.

A test suite also reflects the threats its designers know how to express. New attack methods can emerge after certification, while legitimate product changes can create different failure paths.

AIUC addresses this problem through recurring evaluations. That reduces the age of its evidence, but it does not remove the underlying uncertainty.

The company’s methodology presents another question. AIUC uses agents to conduct parts of its tests and analyze resulting data, with humans verifying the final audit.

Automation can expand testing coverage. It can also reproduce blind spots when evaluation agents misunderstand a task, overlook an ambiguous result, or favor patterns represented in their instructions.

Human verification helps, but readers should not treat it as proof that every test result is correct. The quality of the audit depends on sampling, reviewer judgment, and access to relevant system information.

Independence also needs careful definition. AIUC develops the standard, supports certification, and participates in insurance arrangements tied to the same risk model.

That combination can align incentives when failures create financial costs. It can also create perceived conflicts if the organization benefits from expanding certification and coverage.

Authorized third-party auditors can provide separation between standard setting and individual assessments. Yet the market still needs transparency about auditor oversight, failed reviews, appeals, and enforcement.

Public certification announcements naturally emphasize successful outcomes. Buyers also need to understand how often systems fail, which weaknesses recur, and whether vendors correct them before receiving approval.

Detailed reports are not always public because they may expose security weaknesses. That confidentiality is reasonable, but it limits independent scrutiny of broad claims.

AIUC says Cursor’s complete scope and evaluation details are available through the vendor’s trust portal. Access-controlled evidence can help enterprise buyers without publishing attack instructions.

However, a trust portal still places interpretation on the customer. Security teams must decide whether the tested configuration matches their own deployment.

A certificate for an agent with restricted repository access may not apply when another customer grants deployment credentials. The product name can remain unchanged while the operational risk becomes much larger.

Insurance carries similar limitations. A policy does not prevent an incident. It transfers some financial consequences after coverage conditions, exclusions, and loss definitions are applied.

Some harms are difficult to price or repair. Exposed data cannot always be recovered, and unsafe automated decisions can create regulatory or reputational consequences beyond a covered payment.

Coverage disputes can also reveal ambiguity over responsibility. An insurer may examine whether the vendor, model provider, deploying company, or user caused the loss.

That makes contract design central to AI insurance. Coverage must define the insured system, approved uses, required controls, reporting obligations, and excluded conduct.

AIUC’s public materials do not provide enough information to evaluate every policy condition. Enterprise buyers should review actual coverage documents rather than infer protection from certification alone.

Regulation creates another uncertainty. A private standard can help organizations structure evidence, but it cannot replace legal obligations in every jurisdiction.

A framework may map controls to regulations without determining whether a particular deployment complies with the law. That conclusion still requires legal and operational analysis.

AIUC’s most defensible claim is therefore narrower than “safe AI.” It offers a repeatable way to examine selected risks, document controls, and support an insurance decision.

That can still be valuable. Enterprise risk management rarely eliminates uncertainty. It creates evidence, assigns responsibility, and establishes procedures for failures that remain possible.

The Real Contest Is Independent Assurance Versus Vendor Promises

AIUC is betting that enterprise buyers will demand outside evidence instead of accepting safety claims produced entirely by AI vendors.

AI developers already conduct internal evaluations, red-team exercises, and security reviews. Large model providers also publish system cards and selected test results.

Those practices provide useful information, but vendors choose many testing assumptions and disclosure boundaries. Commercial pressure can influence how weaknesses are framed or prioritized.

Independent evaluation attempts to create distance between the company selling an agent and the evidence used to approve it. The evaluator asks whether the system satisfies a shared requirement.

METR provides a useful historical reference. The research organization evaluates whether advanced models and agents can complete increasingly difficult tasks under controlled conditions.

Dattani served as METR’s COO between 2024 and 2025 and remains a board member, according to TechCrunch. His move into AIUC brings evaluation experience into a commercial assurance model.

The two organizations are not direct equivalents. METR has focused on frontier-model capabilities and associated risks, while AIUC targets enterprise adoption, certification, and insurance.

Their shared premise is that model developers should not remain the only judges of their own systems. Independent testers can provide evidence for buyers, policymakers, and the public.

AIUC takes that premise closer to procurement. Its reports are intended to help enterprises decide where an agent passes, where concerns remain, and whether deployment is acceptable.

That decision-oriented framing matters. An evaluation does not need to label an entire system safe or unsafe. It can document where controls work and where human review remains necessary.

The approach resembles established product assurance systems. Technical standards become influential when manufacturers, buyers, auditors, insurers, and regulators recognize the same evidence.

Ribbit Capital investor Nick Shalek described coordination across builders, enterprises, security leaders, auditors, and insurers as AIUC’s cold-start challenge. The investor’s support signals confidence, not independent validation.

Standards markets can favor early leaders because each participant attracts more participants. Vendors want the certificate buyers request, while buyers request certificates available across many vendors.

That network effect also creates competition over legitimacy. AIUC must persuade the market that its standard is sufficiently independent, technically demanding, and adaptable.

Alternatives include vendor-led testing, internal enterprise reviews, government rules, sector-specific frameworks, and open evaluation projects. Most organizations will combine several approaches.

AIUC therefore does not need to replace every framework. It needs to become a trusted bridge between technical testing and commercial risk decisions.

The company’s customer list gives it an early foothold across prominent agent categories. Yet adoption announcements do not reveal how often certification actually shortens procurement.

That outcome should become measurable. Buyers can compare review duration, remediation work, incident rates, and insurance decisions before and after adopting AIUC-1.

The strongest evidence would come from repeat behavior. Enterprise customers would request renewed certificates, auditors would identify material problems, and insurers would adjust decisions using observed outcomes.

The weakest outcome would be badge inflation. Vendors could collect another logo while buyers continue running the same bespoke reviews and accepting the same unresolved risks.

AIUC’s future depends on avoiding that fate. Its audits must uncover meaningful weaknesses, and its certification must remain difficult enough to carry information.

Three Signals Will Show Whether AIUC’s Model Works

The next test is whether AIUC can convert funding, certifications, and insurer participation into observable changes in enterprise deployment decisions.

The first signal is broader independent audit capacity. Schellman became AIUC’s first authorized auditor, but a durable standard needs multiple qualified firms applying consistent methods.

Additional auditors would expand capacity and reduce dependence on AIUC’s internal operations. However, expansion only strengthens the model if accreditation and quality controls remain demanding.

Watch for published rules covering auditor training, conflicts, review consistency, and sanctions. Clear governance would strengthen AIUC’s claim that the certificate represents independent assurance.

Weak or opaque oversight would undermine it. A standard becomes less informative when different auditors interpret the same requirement in incompatible ways.

The second signal is evidence that certification changes procurement outcomes. AIUC says security reviews often block agents after successful pilots.

The company should eventually show whether certified vendors complete reviews faster, face fewer repeated questionnaires, or reach production with fewer exceptions. Aggregated data could protect customer confidentiality while testing the central business claim.

Renewals will matter more than launch announcements. A customer that repeats quarterly testing and annual review demonstrates continuing value beyond initial marketing.

Buyers should also examine whether certification applies to the configuration they plan to use. Product teams need a searchable knowledge base containing scopes, test evidence, exceptions, and deployment changes.

That record becomes critical after an update. Security teams must know whether a new model, tool, or permission invalidates earlier conclusions.

The third signal is insurance performance. AIUC’s combined model becomes more credible if test findings influence underwriting and predict real losses.

Useful evidence would include anonymized claims patterns, common failure categories, control effectiveness, and changes in coverage decisions. Such data would show whether certification measures financially relevant risk.

The opposite result would weaken the thesis. If certified systems experience similar losses, or insurers disregard evaluation results, the link between testing and underwriting would remain unproven.

Frontier-model expansion deserves attention within this signal. Application audits examine complete agent systems, while model-level evaluation addresses capabilities shared across many products.

Moving upstream increases AIUC’s potential influence but also raises methodological difficulty. A general model behaves differently after developers add tools, prompts, memory, and access controls.

AIUC will need to explain how model evidence connects with application evidence. Neither layer alone captures the full deployment risk.

Enterprise buyers should avoid waiting for a perfect guarantee. None is available from AIUC, model vendors, regulators, or internal review teams.

They should instead ask concrete questions. Which configuration was tested? Which risks failed? What changed after remediation? When does the certificate expire? What losses does the policy exclude?

Developers should expect those questions to become normal as agents gain access to consequential systems. Clear evidence can become a product advantage, especially when buyers cannot reproduce every evaluation themselves.

Knowledge workers should care because agent failures increasingly affect the information and decisions surrounding their work. A wrong answer is more consequential when software can act on it.

AIUC’s $40 million round supports a credible experiment in private AI governance. The company is combining technical tests, recurring audits, certification, and insurance around one shared risk model.

The approach does not rein in every rogue agent, and certification cannot promise that outcome. Its value rests on a more practical question: can independent evidence make risky deployments easier to evaluate and harder to excuse?

Over the next several months, watch the auditors, procurement data, and insurance results. Those signals will show whether AIUC AI agent safety becomes infrastructure or remains another trust badge.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page