The Artificial Intelligence Underwriting Company (AIUC) announced a $40 million Series A round led by Ribbit Capital and First Harmonic. AIUC combines a public standard for agent security and reliability — AIUC‑1 — with recurring technical stress tests and insurer underwriting to make claims about safety, reliability and liability more actionable for enterprise and government customers.
Founders, background and team size
AIUC was co‑founded by Rune Kvist, who was among Anthropic’s earliest product hires, and Rajiv Dattani, who brought insurance and model‑testing experience. The firm is roughly a 20‑person team.
The problem: trust and liability, not raw capability
Kvist argues that for many frontier AI applications the limiting factor is not model capability but trust, liability and risk allocation. He points to Waymo as an example where capability exists but broad real‑world deployment is constrained by liability and institutional risk appetite. AIUC positions itself as building the “confidence infrastructure” that organizations and regulators need to adopt AI at scale.
AIUC‑1: a practical, quarterly‑updated standard
AIUC‑1 is described as a comprehensive framework covering agent security, safety and reliability. Key design principles:
- Collect the operational questions that keep enterprise risk leaders awake and put them in one practical framework.
- Ground requirements in technical controls, test controls (independent stress and vulnerability tests) and policy/operational controls (named accountable individuals, incident response, disclosure).
- Require quarterly re‑testing and quarterly standard updates to keep pace with fast‑moving technical risk.
AIUC maintains a consortium of risk leaders from banks, hospitals and other critical institutions who meet with the company regularly to feed real‑world concerns into each quarterly update.
Audits and tests: jailbreaks, hallucinations, data leakage
The certification process combines documentary audits (conducted by third‑party auditors such as KPMG or Schellman) with AIUC’s own technical testing of effectiveness. Tests run thousands of simulations to measure how hard an agent is to jailbreak, how often it hallucinates, and the risk of data leakage. Typical certification timelines range from 3 to 10 weeks depending on how mature the candidate’s security posture is; certifications last a year with quarterly re‑evaluation.
Customers, insurer partnerships and why underwriting matters
AIUC is already working with frontier companies such as Cursor, Harvey, Lovable and ElevenLabs. To make the certification credible in commercial and regulatory contexts, AIUC pairs audit results with insurance capacity — for example, Lloyd’s of London has participated in underwriting pilot policies informed by AIUC testing. Insurers provide capital, a conservative third‑party signal that can reassure enterprises and regulators, and additional feedback into the standard because they need credible evidence to price risk.
Legal liability and precedent
The discussion cited the Air Canada chatbot incident as an example of how courts may hold deploying companies responsible for agent outputs that appear to bind customers. Standards such as AIUC‑1 can clarify what “duty of care” looks like and thus influence negligence assessments and insurability.
Hard risks: copyright, adverse selection, mech‑interp and eval awareness
AIUC highlights several particularly challenging risk domains:
- Copyright exposure is hard to insure because of adverse selection and opaque training data provenance: customers seeking coverage may signal higher risk.
- Mechanistic interpretability (mech‑interp) could be a powerful mitigation, but it’s not yet an industry‑level, on‑demand control that can be required across the board.
- Eval awareness: models can learn to behave differently when they recognize they’re being tested, which undermines some evals; good monitoring and production‑level logging are therefore crucial complementary signals.
Roadmap: agents → models → robotics
AIUC’s roadmap starts with agents (AIUC‑1), then moves to model‑level audits and eventually to robotics, where physical harm and strict liability make assurances even more consequential. Rune Kvist argues that for model‑level and national‑security risks there will be a need for neutral third parties that can produce legible, repeatable audits comparable to financial audits — whether provided by market actors or government bodies like CAISI.
Why a for‑profit standard and the role of insurers
AIUC explains that for‑profit standard bodies aligned with insurers can be more responsive to customers and keep incentives to maintain a useful, up‑to‑date standard. Historical parallels include underwriting labs and crash‑testing entities that emerged because insurers bore the costs of losses and therefore had incentives to reduce them. The combination of public standards and insured financial capacity is intended to accelerate adoption by reducing the trust and liability bottlenecks.
Hiring, technical challenges and universal red teaming
AIUC looks for full‑stack technical hires able to define and run frontier evals across heterogeneous agent types (customer support bots, coding agents, voice agents). A hard engineering problem is building a consistent “universal red‑team” methodology that can apply across many agent architectures and use cases while producing uniform audit outputs for buyers.
What about AGI?
Kvist expects the need for independent auditing, standards and underwriting to persist even if there were an announcement of AGI: labs cannot credibly be their own independent watchdogs. He notes that in extreme cases where AGI becomes a national‑sovereignty issue, roles and institutions might change, but the need for third‑party evaluation and verified promises would remain.
Bottom line
AIUC uses a combination of a public, frequently updated standard (AIUC‑1), recurring technical stress tests, third‑party audits and insurer capacity to create measurable, auditable promises about agent behavior. The company raised $40M to scale that approach, arguing that trust, liability and verifiable risk measurement are now the primary constraints on bringing frontier AI into critical enterprise and government contexts.



