AIUC raised $40 million to certify AI agents for enterprise safety, using thousands of tests to assess jailbreaks, hallucinations, and data leaks.

Artificial Intelligence Underwriting Company (AIUC), a startup founded by an early Anthropic employee and the former COO of AI-safety organization METR, has raised $40 million to build an independent testing and certification layer for AI agents.
The Series A was led by Ribbit Capital, with participation from First Harmonic, according to TechCrunch. The round brings AIUC’s total funding to $55 million, following a previously reported $15 million seed round backed by Nat Friedman’s NFDG, Emergence, Terrain, and Anthropic co-founder Ben Mann.
AIUC is targeting a growing enterprise concern: whether an AI system can be trusted to operate inside business workflows, not simply whether it can complete tasks. Its proposed AIUC-1 standard is modeled partly on SOC 2, the widely used framework for assessing controls around security and business processes.
AIUC was founded by Rune Kvist, an early Anthropic employee, and Rajiv Dattani, who served as METR’s COO from 2024 to 2025 and remains on the organization’s board, according to TechCrunch. The founders are positioning the company as an independent evaluator for companies that build or deploy AI agents.
The startup says it has assembled a consortium of about 250 security and risk leaders to help define what enterprise buyers want to know before purchasing an agent. Those discussions inform the tests AIUC uses to evaluate systems, Dattani told TechCrunch.
AIUC says its testing service runs an agent through approximately 5,000 scenarios covering areas such as jailbreaks, hallucinations, and data leakage. The process produces a report of roughly 100 pages, identifying where the system behaves reliably and where buyers should apply restrictions or additional controls.
The tests themselves are run with the help of AI agents, and AI is also used to analyze the resulting data. According to Kvist, people verify the final audit. That human review is important because an automated evaluator can inherit the same blind spots, misinterpret model behavior, or fail to recognize a subtle security issue.
The company’s pitch reflects a change in how businesses are evaluating enterprise AI. An agent that performs well on a benchmark may still be unsuitable for a finance, healthcare, government, or customer-service workflow if it can expose confidential information, follow malicious instructions, or take an action outside its authorization.
Kvist told TechCrunch that many organizations are no longer holding back solely because AI models lack capability. Instead, they need assurances that systems will behave within commitments made to customers and regulators. AIUC’s approach is intended to provide evidence that can be reviewed before deployment, rather than relying only on a vendor’s internal testing.
The company lists Cursor, Lovable, Harvey, and ElevenLabs among its customers. The source does not provide details about the scope of those relationships, which products were assessed, or whether the companies have completed AIUC-1 certification. Those adoption signals should therefore be treated as company-reported rather than independently verified customer results.
AIUC’s business model also differs from traditional insurance, despite the word “underwriting” in its name. The evidence describes a testing and certification service, not coverage for losses caused by an AI system. Its immediate product is an assessment layer intended to support procurement and deployment decisions.
Dattani’s background at METR gives AIUC a connection to an established area of AI evaluation. METR has worked with frontier AI labs to measure whether agents can reliably complete tasks. Its research has also been used in investigations involving model behavior, including work connected to OpenAI’s Hugging Face incident.
AIUC says its focus is broader and more enterprise-facing. Rather than concentrating primarily on task performance, its AIUC-1 process is designed to examine security and reliability failures that can emerge when agents interact with real business data, tools, and users.
That distinction matters for buyers. A system may successfully complete a task in a controlled evaluation but still require narrow permissions, monitoring, human approval, or a restricted operating environment in production. Independent testing could help procurement teams compare systems, although a certification cannot guarantee that an agent will remain safe after a model update, a tool change, or a new attack technique.
The broader idea of outside evaluation is also gaining attention among frontier-model companies. Anthropic CEO Dario Amodei has called for greater caution in frontier development and suggested that independent evaluators could observe and verify safety work at AI labs. AIUC is not proposing to embed evaluators at customer sites, but its founders are pursuing a related principle: organizations should have access to independent evidence before trusting powerful systems.
The funding is a confirmed business event reported by TechCrunch, as are AIUC’s founders, its AIUC-1 standard, and the company’s stated testing process. The reported customer names, consortium size, test count, and report length come from AIUC’s own disclosures through the interview and should be understood as company claims.
There is not yet enough evidence to determine how well AIUC-1 predicts incidents in production, how consistently different evaluators would score the same agent, or whether enterprises will treat the certification as a procurement requirement. The source also does not identify public audit results, failure rates, remediation outcomes, pricing, or the proportion of tests that agents pass.
Those gaps are significant. Security certifications become useful when buyers understand their scope and when assessments are repeatable over time. AI agents can change behavior after model updates, tool integrations, prompt changes, or shifts in the data they can access. A one-time report may therefore be less valuable than continuous evaluation tied to deployment controls.
The first signal will be whether AIUC publishes details from completed AIUC-1 assessments, including methodology, scoring criteria, and examples of failures discovered during testing. Buyers will also want to know how the company handles model updates and whether certification must be renewed when an agent’s tools or permissions change.
Another important signal is independent validation. Evidence from enterprise customers, security researchers, or auditors unaffiliated with AIUC would help establish whether the tests identify real-world risks rather than only synthetic failures.
The market will also reveal whether AIUC’s approach becomes part of procurement. Adoption by large banks, hospitals, government agencies, or major software platforms would suggest that third-party agent evaluation is becoming a standard control. Conversely, if companies use certification as a marketing badge without publishing meaningful results, its value to builders and buyers will remain limited.
AIUC is addressing a real bottleneck in enterprise AI: organizations need a defensible way to decide where an agent can act autonomously and where it needs limits. A structured external assessment could make those decisions more concrete than broad claims about model safety or general capability.
But certification should be treated as evidence, not permission to remove safeguards. The strongest version of this market will combine independent tests with live monitoring, least-privilege access, human escalation, and repeat assessments after material system changes. AIUC’s funding gives that model room to develop; its credibility will depend on whether it can demonstrate that the reports improve deployment decisions in practice.