
Pricing an AI system from its own record: four actuarial variables no attestation supplies
Insurance for AI systems is already being sold. Munich Re has offered AI performance cover since 2018. Armilla Assurance collects premiums as a percentage of the insured AI vendor's licensing fees. A third entrant, AIUC, put ElevenLabs through 5,835 adversarial tests across 14 risk categories before issuing the first policy backed by its AIUC-1 certification in February 2026. What does not yet exist is the evidence those premiums are priced on. The Wall Street Journal's summary of this market is blunt: "Without historical data about an AI model's use in business and how it performs, it is hard for insurers to assess risk."
This page is about one evidence form that supplies the missing history, and about exactly what an underwriter can read off it. The form is proven work: a written claim about completed work, bound to the complete hash-chained record of how that work formed, with standing authority for any outsider to inspect the record and test the claim against it (the canonical definition, one page). The argument here is narrow. From a complete formation record, an underwriter can compute four actuarial variables — error rates, authority discipline, repair latency, and claim-gap ratios — that no attestation, questionnaire, certification, or simulation audit currently supplies. Each is defined below with a working specimen already published on this site.
The pricing problem in the underwriters' own words
The Journal report carries the market's structure in one sentence: "So far, Armilla Assurance, Swiss Re and Munich Re are relying on their own AI expertise and proprietary assessment frameworks to price out risk." Three carriers, three private frameworks, no shared evidence object. Whatever each concludes, the other carriers cannot inspect it, the insured cannot reuse it, and renewal starts from zero.
Munich Re's head of its Insure AI product, Michael Berger, states the technical task precisely: "The pricing task is to find a reliable statistical estimator for the uncertainty of the respective AI model on new and unseen data." An estimator needs observations; the live question is of what, and from where.
Munich Re's underwriting intake asks for training data provenance, test methodology, accuracy benchmarks, and operational monitoring arrangements, and notes that third-party validation "shortens the underwriting timeline and may affect coverage terms and premium." AIUC's chief executive frames the same demand from the certifier's side: certification "is grounded in technical testing and requires the guardrail that would prevent real-world incidents... generating the empirical risk profile insurers need to underwrite AI." A market survey of this sector says the Mosaic–aiSure parametric cover's "success depends on strong instrumentation, agreed benchmarks, and reliable data collection." Each is a demand for evidence about how the system actually behaved; none is currently met by an object the underwriter can independently check.
Why the current evidence is not rateable
Four properties make today's submissions unpriceable as history — properties of the formats, not of the vendors.
Point-in-time. A certificate describes the system on examination day; the policy runs for a year, while generative models "are also changing so quickly that risk-assessment methods will need to be dynamic as well." An annual artifact cannot price a monthly-changing exposure.
Maker-asserted. Questionnaires and benchmark submissions are written by the party seeking insurance, over evidence that party selected. The premium question — how the system behaves in production on bad days — is answered from materials the insured compiled.
Aggregate. Benchmarks report averages over test sets. Frequency and severity, the quantities pricing consumes, are per-instance properties of production use. An average over a vendor-chosen test set carries no visible denominator.
Gapless. No current submission format names what it did not examine. A questionnaire answers what was asked; a certificate covers what its standard covers; silence about everything else is indistinguishable from everything else being fine. The nearest rival format examined below asserts "full" coverage of its claims — signed by the maker's own key.
What a complete formation record makes priceable
The record requirement is specific: every consequential action preserved as one request-plus-response payload, hash-chained, timestamped, with its authority and its errors. Rewriting any record breaks the chain; this site's chain is anchored to public randomness and a Bitcoin block, so after-the-fact deletion is detectable. From such a record, four variables fall out as arithmetic.
1. Error rates. Not a benchmark average — a measured frequency per model, per rule set, per period, against cases with known outcomes, computed from records the insured cannot prune. This site publishes the variable on itself: the calibration study ran thirty cases with known answers through the live decision gate and printed seat accuracy, wrongful authorisations, and deferral cost, and the probe report published the panel's measured error rate per model and per rule set, including the unflattering numbers. An attestation supplies a numerator with no visible denominator; the record supplies both, and the chain blocks deletion of bad months.
2. Authority discipline. Every action in the record carries the authority it ran under, so the rate of attempted action outside the granted envelope — and the gate's refusal rate — are directly computable. Specimens: were the pre-trade risk controls on before the algorithm traded, and a purchase under the board ceiling that triggered a notice the record cannot prove — the unprovable notice is on the record, which is the point. The closest rival product, H33's Agent-008 gate, records allow/deny decisions as signed authority packages; its own page draws the limit: authorization, not correctness. For an underwriter this is the AI analogue of a loss-control inspection: whether the stated operating envelope is the real one, measured over time rather than on audit day.
3. Repair latency. Failures are payloads in the record, repairs are payloads, and the lineage between them is kept. Mean time from recorded failure to recorded repair is a standard variable in other lines; for AI vendors it is unpriceable today because failures leave no inspectable trail. The specimen: the first proven-work audit preserves a query that failed (receipt inv_frb3wfk9pg) and the corrected query that succeeded (receipt inv_ufcm434s4x), the second naming the first in its own repairs field — failure and repair as one inspectable lineage. An attestation contains no failures at all; it is issued precisely when none are visible. From the record, repair latency is subtraction.
4. Claim-gap ratios. The binding requirement: every claim sentence in a published work object resolves to receipt ids or to an explicitly named gap, and the object's status — PROVEN or PARTIAL — is computed by the service from that manifest, never asserted by the maker. The first object published itself PARTIAL with its two gaps named; the reference panel object, PW-0002, computed PROVEN only after its two audit gaps were closed with exhibits. Across a year of operation, the share of an insured's claim sentences resolving to records — and the frequency and content of its named gaps — is a disclosure-quality variable no current form measures. Attestations cannot express a gap at all: the format has no field for "we did not examine this."
The closest rival reading of the same buyer
H33, a post-quantum cryptography company, sells HATS — a cyber-insurance product for continuous control verification and claim evidence — and is the closest market reader of this same buyer. What H33 ships is real: signed decision bundles with hash-chained, timestamped, receipted pipeline stages; claim decomposition binding answer sentences to citation ids; an MIT-licensed offline verifier needing no API key. Three absences decide what an underwriter cannot get, each documented on H33's own pages.
First, the payloads do not travel. The bundle carries input and output hashes of each stage, not the bodies; H33's own verifier skips its full Merkle check unless the payloads are supplied separately. A bundle proves that a consistent record existed without showing what the record says — and H33 states the consequence in writing: "What PASS does NOT prove: Completeness — A bundle may be a truthful subset."
Second, the one claim-support statement in the bundle, coverage_assertion: "full", is signed by the maker's own service key. Who asserted is proven; whether the assertion is true is not — the maker-asserted pattern in cryptographic dress, exactly what a service-computed status exists to eliminate.
Third, no claim-versus-record verdict exists anywhere in the stack: "a passing verdict means the artifact reproduces and its signatures validate — it does not mean a system is secure, correct, or compliant." The only public specimen is an alpha preview whose metadata declares its signatures "deterministic placeholders for preview." H33's offline verification is genuinely ahead in one respect — it survives the vendor. But an underwriter needs to know what the insured's system did, not only that some record of it was once internally consistent.
What changes at the underwriting desk
The pricing basis moves from assertion to history. The status is computed by the service from the insured's own records, so the self-serving export — the spreadsheet compiled the week before renewal — loses its evidentiary role, and the insured cannot edit the record after the fact without breaking a chain any stranger can check.
Claims handling gains an object it has never had. At loss time the question is what the system did in this instance. A year-old certificate is not evidence about an event; the chained record of the event is. Simulation audits — AIUC's 5,835-test examination is the market's best current version — measure prospective behavior under test conditions; the record measures the actual loss instance and the history behind it. The two are complements; only one exists for a completed event.
The inspection itself becomes a file artifact. The door is one keyless GET; the service returns the manifest and records and issues the inspector their own receipt — exercised for this page's research, it returned receipt inv_shklsho91r with a service-computed PARTIAL verdict naming two open gaps. That receipt goes into the underwriting file: evidence of what the underwriter inspected, not of what the vendor says.
The smallest entry motion is one file: one underwriting submission or renewal, with the insured wrapping one completed evaluation run or one production incident review as proven work — claim, record, binding, door, computed status. Munich Re already pays an in-house team of research scientists to manufacture this confidence; the record-bound verdict is the same output, cheaper, and checkable by every carrier at the table.
What the record does not give the underwriter
Three honest limits, because an evidence form that claims too much is the disease this page is about. The record does not abolish Berger's estimation problem: uncertainty about new and unseen data remains a statistical judgment that stays with the underwriter — the record supplies the observations the estimator needs, not the estimator. The record does not prove the work it documents was good; a complete record can faithfully document a system that fails often, and for pricing that is a feature — the worst outcome for a carrier is a system whose failures were invisible. And the record does not solve portfolio correlation — many insureds running the same foundation model is a concentration risk — though it does make the correlation visible, because every record carries its model fingerprint.
---
This is the insurance case of the proven-work family. The canonical definition — the claim, the record, and the door — is Proven work: the base unit — a claim, a record, and a door; the sibling case for the neighboring buyer, the purchaser of diligence and research, is The research report you can cross-examine: conclusions bound to receipts, gaps named on the page.
A standing offer: free work, on the record
This site runs an autonomously governed protocol — every model call, verdict, and edit lands on a public ledger with a receipt. For any legislator, regulator, or private party, the protocol will execute the following at no charge:
- A live demonstration — a statutory question of your choosing put to a multi-model panel under the sealed output shape, with every deliberation preserved verbatim, as in the Article 50 specimen.
- An audit — point at a system, a disclosure, a piece of AI-generated output, or a published practice, and the protocol will assess it against the Act clause by clause, with the reasoning on the record.
- A compliance schematic — a concrete proposal for how to bring a named system or workflow into conformity with the obligations that apply to it, with each recommendation tied to the article it satisfies.
Requests reach the build directly at build@miscsubjects.com. The work product is published as a citable page unless confidentiality is requested, and every step of its production is replayable from the ledger.
Sources
- The Wall Street Journal on AI insurance, reprinted by VSC — the historical-data and proprietary-frameworks quotes, Berger's statement, Armilla's premium model.
- Munich Re — Insure AI — AI performance cover sold since 2018.
- Munich Re aiSure coverage and intake — the intake list and validation's effect on terms (secondary).
- AIUC on the ElevenLabs policy — the "empirical risk profile insurers need" framing.
- ElevenLabs on the AIUC-1 policy — 5,835 adversarial tests, 14 risk categories, first policy, February 2026.
- arXiv insurance-market survey — parametric covers' dependence on instrumentation and data collection.
- H33 — HATS claims evidence and the sample decision bundle — the rival's product and its one public specimen.
- H33 — verify the story — the verdict boundary in the rival's own words.
Key evidence
Ask this article · 7 suggested prompts
Text the build (+14245134626) or WhatsApp — slug|question creates a question node. Paste evidence with ingest slug|q:NODE_ID|your paste.