Underwriting Superintelligence: AIUC's Insurance, Standards & Audits to Accelerate AI Adoption
Summary
- The AI underwriting company’s central thesis is that stronger safeguards should accelerate AI deployment rather than restrain it. Kvist compares safety infrastructure to the helmet and seatbelt that let a race driver corner faster: “security and progress are mutually reinforcing.” Labenz frames the downside as a nuclear-style outcome in which weaponization proceeds while practical benefits are frozen out.
- Insurance, standards, and independent audits form a single incentive system that none of the three can provide alone. Standards define best practice, audits establish whether developers actually follow it, and insurance pays when failures still occur. Because insurers want premium growth but bear losses from lax controls, Dattani argues this offers a market-based middle ground between voluntary promises and inflexible top-down regulation.
- Existing policies leave both enterprises and insurers exposed to unresolved AI coverage ambiguity. Cyber and other policies may cover familiar categories of harm, but often never mention an AI-induced cause; meanwhile, insurers have not priced a risk they do not yet know how to price. Dattani suspects a replay of early cyber insurance, when years of litigation eventually separated computer-related losses into differently structured products.
- Red teaming could help compensate for the historical loss data that conventional underwriting lacks. An EY study cited in the conversation found enterprises already experiencing losses of $1 million, while smaller companies absorb thousands or tens of thousands that insurers may never see. Repeated evaluations can generate synthetic frequency and severity data, while parametric triggers could produce pre-agreed payouts without years of litigation.
- Private markets may insure routine AI failures, but the largest tail risks probably require a government backstop. “No one knows what the kind of tail risk distribution looks like because we have not yet seen it,” Kvist cautions. The proposed analogue is US nuclear insurance: operators carry strict liability and mandatory coverage, while damages above $15 billion move to the government balance sheet so private insurers can still perform the governance function.
- AIUC-1 packages enterprise AI concerns into one auditable procurement language built from more than 500 industry conversations. It covers data and privacy, security, safety, reliability, accountability, and societal risks, emphasizing disclosure where hospitals and retailers may reasonably choose different thresholds. Initial red teams sometimes find attack-specific failure rates as high as 25%; after safeguards, the same rates can fall by 90%, though the founders stress that “there are no panaceas.”
- The application layer is the AI underwriting company’s first commercial wedge because agent vendors make unusually specific promises without foundation-model-scale balance sheets. Customer-support agents promise to resolve a stated percentage of tickets while implicitly promising not to cause brand disasters: “the thicker the promises,” the thicker the assurance required. The longer-term ladder runs from million-dollar risks through hundreds of millions and low billions to tens of billions, spanning applications, foundation-model systems, and data centers.
- The company is structuring its economics to avoid the issuer-paid race to the bottom that damaged credit ratings before 2008. As a managing general agent, its compensation will depend on insurers’ underwriting results; large losses mean lower or no compensation, lost carrier relationships, and potentially, “We’ll go out of business.” Labenz disclosed that he invested in the company’s seed round alongside Nat Friedman, Emergence, and Terrain. Founding technical contributors include Cognition, Ada, Intercom, and Recraft; the enterprise consortium includes JPMorgan Chase, Confluent, and Anthropic’s deputy CISO.
Deep dive
1. Security infrastructure should accelerate deployment
Kvist begins with present-day risks rather than extinction: data leakage, security vulnerabilities, incorrect advice, and brand disasters are familiar harms made more severe by vastly expanded attack surfaces that organizations understand poorly.
As systems become smarter, the failure modes change rather than disappear. Kvist’s analogy is a highly intelligent CEO: greater IQ may eliminate elementary logical mistakes while enabling more sophisticated fraud; similarly, advanced systems could progress from deception or sycophancy to more capable manipulation.
Catastrophic scenarios include terrorists cheaply accessing “brilliant scientists” able to help create new viruses, as well as national-security leaders preemptively attacking rivals because they believe AI could confer decisive strategic advantage.
Labenz clarifies his own “nuclear outcome”: society could suppress cheap, abundant, everyday benefits while retaining AI weaponization, just as nuclear technology produced thousands of weapons without delivering the super-cheap, abundant energy he believes society should have received.
Kvist says this asymmetry is already visible: autonomous vehicles remain constrained in ordinary deployment, while autonomous drone warfare is increasingly deployed with few apparent barriers.
2. Insurance, standards, and audits form one incentive loop
The founders’ governing phrase is “security and progress are mutually reinforcing.” A race driver wears a helmet and seatbelt to go faster around corners; stronger steering and security should likewise permit more ambitious deployment of agentic systems.
Dattani separates the mechanism into three jobs: insurance supplies financial protection, standards codify best practice, and audits answer, “Have you met the standards?” Insurers cannot distinguish good from bad risks without requirements, while requirements cannot create trust through self-attestation alone.
The proposed middle ground matters: insurers benefit when adoption expands their premium pool, so they should resist irrelevant or excessively onerous requirements. Yet because they pay claims, they also resist lax standards and broken commitments.
The historical specimen is Benjamin Franklin’s 1752 Philadelphia fire insurer, created as the city’s population grew 10-fold and dense housing spread risk across neighbors. It developed building codes and inspections; later, insurer-backed Underwriters Laboratories applied the same loop to electrical products and the UL certificate.
3. Today’s policies leave AI losses in a costly gray zone
Existing insurance commonly addresses harm categories such as cyber vulnerabilities, confidential-data leaks, and downtime, but policies frequently say nothing about AI as the cause. Companies therefore cannot be sure they are covered, while insurers unknowingly accept risks for which no premium was calculated.
The founders expect the ambiguity to resolve as computer risk did in the early 2000s. Familiar harms began occurring at different frequencies and severities, litigation accumulated between customers and insurers, and cyber coverage eventually separated into products with distinct pricing and structures.
Waiting for claims is especially ill-suited to AI: a product built from historical data today could describe the pre-LLM era and omit agentic risk entirely. “We’re all self-insured whether we know it or not,” Labenz concludes—even when that underwriting happens informally inside a hospital or bank.
4. Red-team data can price ordinary risk, not unknowable tails
Conventional insurers lack both relevant history and a mechanism for seeing and collecting emerging losses. The cited EY study found $1 million enterprise losses, while smaller firms may quietly write off thousands or tens of thousands, leaving insurers blind to the long tail beneath headline incidents.
Red teaming and evaluations can generate a synthetic loss dataset before equivalent real-world events accumulate. Insurers can translate observed failure frequency into familiar costs—such as a data leak or lawsuit—and feed both estimated frequency and severity into pricing models.
Dattani also sees scope for parametric triggers: pre-agree whether a defined out-of-distribution behavior constitutes failure and what it pays. A qualifying event could then release, for example, $1 million without waiting for a prolonged coverage dispute.
Labenz’s power-law challenge remains unresolved: a few incidents might dominate aggregate losses. The nuclear template caps private exposure—US operators carry mandatory insurance and supply-chain liability, but damage above $15 billion receives a government backstop—while preserving financially motivated third-party governance below that threshold.
5. AIUC-1 turns fragmented concerns into procurement language
AIUC-1 emerged from conversations with more than 500 security leaders, general counsels, and other executives across sectors including banking and healthcare. The objective was to gather every concern slowing agent adoption into a framework detailed enough for independent verification.
Its report spans data and privacy, security, safety, reliability, accountability, and societal risk. Practical questions include prompt injection, jailbreaks, hallucinations, bias, output filtering, human oversight, robustness, responses to angry customers, and whether an agent stays within its assigned role.
Consensus proved “remarkably uncontroversial” because much of the standard creates disclosure rather than imposing one universal threshold. A hospital and retailer may accept different risks, but each benefits from “one shared language” and third-party evidence rather than a vendor grading its own homework.
The harder frontier is societal harm that may not concern an individual buyer directly—for example, whether a product enables cyberattacks at unprecedented scale. AIUC includes such risks where an incident could chill the entire market, while trying to keep tests pragmatic enough for ordinary vendors.
6. Audits matter only if attackers get another move
Each audit begins with an incident database. Air Canada’s chatbot inventing a refund policy motivates financially consequential hallucination tests; research showing chatbots attempting blackmail under experimental conditions is retained as a more speculative incident class.
AIUC abstracts those cases into harms, attack vectors, and required technical, legal, and operational safeguards. Reviewers may inspect the codebase, written policies, groundedness filters, monitoring, content controls, and incident-response plans, then red-team the system to verify that those measures actually work.
Labenz’s pushback invokes “the attacker moves second”: static defenses can defeat known prompts while adaptive humans rapidly find another route. In the study he cites, human red teamers were “basically 100% undefeated,” making a one-time certification vulnerable to security theater.
The answer is iterative rather than absolute. Certification lasts one year and requires quarterly technical audits that incorporate product changes, new incidents, and new research; where to stop remains a judgment balancing cost against confidence, because “with enough effort every system is reachable.”
7. The application layer is the wedge, not the ceiling
Labenz observes that many application-layer businesses grew from weekend hackathons with little security infrastructure, yet now deploy agents into enterprise workflows. Unlike Claude.ai’s nonspecific “helpful tool” promise, these vendors make specific operational and brand promises.
The distinction is already blurring: Claude Code is “a system of models” with integrations and access, not merely a foundation model. The founders therefore treat agents—not the traditional foundation-versus-application boundary—as the more durable frame.
Rajiv says foundation-model providers, with some notable exceptions, are already doing substantial work to taxonomize risks and build defense-in-depth approaches. He sees more low-hanging fruit at the application layer, where many vendors still have little protection.
The company says the same standards-and-audit process could extend to data centers, where companies are placing tens of billions of dollars into infrastructure and shareholders may not want to bear downtime, attack, or breach risk. Its proposed ladder starts with millions or tens of millions, advances through hundreds of millions and low billions, and eventually reaches tens of billions.
8. The business sells recurring assurance and deployment-specific confidence
AI-native developers and enterprises building agents internally pay yearly; the certificate lasts one year and quarterly tests are mandatory. Pricing scales with product surface, risk, and customer value, with Labenz’s suggested five- to six-figure range described as “broadly correct.”
Before contracting, AIUC performs an outside-in gap analysis against public evidence, then reviews the standard row by row to estimate engineering, non-engineering, and audit effort. The gap analysis can be completed “within 24 hours.”
Customer burden varies sharply: a mature client with thousands of employees required only “a small handful of hours” of security-team work and evidence collection, while a two- or three-person company needed hands-on help implementing safeguards and remediating findings.
AIUC is experimenting with certifying a particular deployment, not merely a general product—for example, testing a voice-support agent against one bank’s policies, customers, and environment. That deeper assurance is commercially legible because it can directly unlock an enterprise contract.
9. Financial skin in the game is the defense against captured auditors
Labenz compares the risk to pre-2008 credit-rating agencies: issuers could shop for favorable ratings, leaving auditors reluctant to turn over too many stones. The founders’ answer is insurance-linked compensation, because Moody’s lacked a direct mechanism forcing it to internalize the risks being created.
AIUC plans to operate as a managing general agent whose compensation depends on carrier underwriting profit. Good risk assessment and fewer losses improve its economics; large claims reduce compensation, jeopardize relationships with established insurers, and could put the company out of business.
On policy, Dattani welcomes markets for competing third-party oversight but rejects safe harbors that let risk creators escape accountability. Insurance should “internalize these externalities,” while a broader network—founding technical contributors Ada, Intercom, Cognition, and Recraft; enterprise-consortium participants including JPMorgan Chase, Confluent, and Anthropic’s deputy CISO; plus academics and leaders across geographies—keeps AIUC-1 responsive to new safeguards and incidents.