{"_self":{"principle":"Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.","widget":"article_topology","feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","contains":"claims, sources, anecdotes, question_graph slice","slug":"insurer-ai-performance-rate-table","urls":{"read":"https://miscsubjects.com/api/articles/insurer-ai-performance-rate-table/topology"},"how_to_use":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","write":null,"imessage":null,"router_tag":null,"proof_chain":[{"step":1,"claim":"Articles are voxel graphs of tiered claims, not prose blobs.","verify":"https://miscsubjects.com/api/articles/constitution"},{"step":2,"claim":"Claims link to hash-chained sources via source_ids.","verify":"https://miscsubjects.com/api/articles/insurer-ai-performance-rate-table/sources"},{"step":3,"claim":"Ask reads topology; ingest/claim append to ledger.","verify":"https://miscsubjects.com/api/protocol"},{"step":4,"claim":"Models queue growth: populate → collaborate → repair → reflex.","verify":"https://miscsubjects.com/api/protocol/grow"},{"step":5,"claim":"Graph proves its own shape (reflex) and $/claim (yield).","verify":"https://miscsubjects.com/graph.html?layer=reflex"},{"step":6,"claim":"Full feature index + _explain on every API response.","verify":"https://miscsubjects.com/api/articles/system-map"}],"related_features":[{"id":"ask","name":"Ask protocol","what":"Answer only from topology; creates question_node with gaps and ingest_hint.","urls":{"read":"https://miscsubjects.com/api/articles/insurer-ai-performance-rate-table/prompts","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"graph_topology","name":"Cross-article graph","what":"Merged claims/sources across condition+stack slugs for one question.","urls":{"read":"https://miscsubjects.com/api/articles/insurer-ai-performance-rate-table/graph-topology?question=..."}},{"id":"question_graph","name":"Question graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output).","urls":{"read":"https://miscsubjects.com/api/articles/insurer-ai-performance-rate-table/question-graph","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"voxels","name":"Voxel graph","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance.","urls":{"read":"https://miscsubjects.com/api/articles/insurer-ai-performance-rate-table/voxels","write":"https://miscsubjects.com/api/protocol/claim"}}],"system_map":"https://miscsubjects.com/api/articles/system-map","system_map_markdown":"https://miscsubjects.com/api/articles/system-map?format=markdown","not_medical_advice":true},"_explain":{"feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","why":"Every feature is auditable collective intelligence","how":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","model":null,"verifies":null,"urls":{"read":"https://miscsubjects.com/api/articles/insurer-ai-performance-rate-table/topology"},"imessage":null,"router":null,"related":[{"id":"ask","what":"Answer only from topology; creates question_node with gaps and ingest_hint."},{"id":"graph_topology","what":"Merged claims/sources across condition+stack slugs for one question."},{"id":"question_graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output)."},{"id":"voxels","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance."}],"not_medical_advice":true},"slug":"insurer-ai-performance-rate-table","title":"You cannot write an AI performance guarantee without a loss-frequency estimate. The probe table is the rate table.","register":"technical","tags":["governance","insurance","adjudication","use-case"],"updated_at":"2026-07-30T13:31:31.915Z","body_excerpt":"## The underwriting problem, stated as an actuary would\n\nInsurance is written on frequency and severity. Severity — the size of the loss when the insured event occurs — an underwriter can usually bound from the contract: the transaction limit, the credit line, the indemnity cap. Frequency is the problem. Every line of business that exists became writable when someone assembled a credible answer to *how often does this happen* — mortality tables for life, loss triangles for casualty, catastrophe models for property. Machine judgement has no such table. When Munich Re's aiSure, Armilla, Relm, and the Lloyd's syndicates that have circled AI performance cover assess a proposal, the question that stalls it is not whether the model is impressive. It is: **at what rate is it wrong, measured how, on what fixed basis?**\n\nAbsent that number, one of three things happens, and all three are visible in the market today:\n\n1. **The risk is declined.** No rate, no policy.\n2. **The risk is written narrow** — cover attaches only to a specific model version on a specific task with the vendor standing behind it, which is really the vendor's warranty wearing an insurance wrapper.\n3. **The risk is written with a loading** large enough to absorb everything the underwriter cannot see: the *opacity loading* (the model's failure modes are unknown) and the *moral-hazard loading* (the insured operates the model, observes its failures first, and controls what gets reported). Loadings of that size price the product out of the use cases that need it.\n\nTwo further structural problems make it worse than an ordinary new line. First, **correlated error**: if an insurer writes a thousand policies on judgements made by the same model family, the errors do not diversify — a defect in the checkpoint is a defect in every insured decision simultaneously, which is a catastrophe-shaped exposure, not a frequency-shaped one. Second, **claims adjudication**: when the insured says \"the model was wrong and it cost us,\" reconstructing what the model saw, what it was instructed with, and what it actually concluded is, for an ungoverned system, forensic archaeology. Every one of those disputes is loss-adjustment expense, and the anticipated expense is priced in before the first claim.\n\nThis page maps a running system's measured artifacts onto those exact inputs. Every claim opens to a live receipt.\n\n## The rate table\n\nUnder a rule set pinned to a content hash — so the basis of measurement is beyond dispute — each model's error rate is measured on a fixed suite and published:\n\n[[embed:source:s1]]\n\nRead it as an actuary, because that is what it is shaped for. It is a **per-seat frequency estimate on a fixed, hashed basis**: the rule set cannot drift under the measurement, the suite is versioned, and re-running it after a vendor swaps checkpoints is the change-detection instrument. It is not a vendor benchmark: the limits — one task class, deliberately small n, the prevalence paradox that makes raw accuracy misleading on skewed case mixes — are stated on the page itself, because an underwriter who prices on a hidden sample is the one who gets hurt at the first claim.\n\n## Correlated versus independent error: the panel and its statistics\n\nA single model's error rate, however well measured, leaves the correlation problem untouched. The system's answer is structural: each governed decision is put to **several models from different training families**, separate vendors, no shared state, each blind to the others. Diversification across seats, though, is only real if two things hold, and both are measured rather than assumed.\n\nFirst, the seats' findings must be *comparable* — otherwise \"agreement\" is unfalsifiable. A governing constitution compels every seat into the same output shape: verdict, clauses relied on, a clause-by-clause derivation (did the clause trigger, does it support or defeat the action, on which evidence records), the records that were absent, the strongest rejected alt","ranking":"safety-first (interaction_risk/limitations), then quote-gated effective_weight","claims":[{"id":"c1","text":"AI performance guarantees are not being written at scale because machine judgement has no loss-frequency history in a form an actuary can use, so the risk is either declined or loaded to the point of pricing itself out.","tier":"system","section":"The underwriting problem","interaction_risk":false,"status":"active","source_ids":[],"why_material":"The market-blocking gap this artifact fills.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c2","text":"Per-model error rates measured under a rule set pinned to a content hash are a loss-frequency estimate for machine judgement, published with its sampling limits stated.","tier":"system","section":"The rate table","interaction_risk":false,"status":"active","source_ids":["s1"],"why_material":"The missing actuarial input, produced as a live table rather than a vendor assertion.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c3","text":"The panel's seats are separate models from separate vendors with no shared state, and the governed output structure that makes their findings comparable appeared in zero of 48 ungoverned calls.","tier":"system","section":"Correlated versus independent error","interaction_risk":false,"status":"active","source_ids":["s4"],"why_material":"Diversification across seats is only real if the errors are independent and the findings are comparable.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c4","text":"Krippendorff's alpha and Fleiss' kappa are published alongside the rates, so an underwriter can see whether the seats' errors are correlated — the statistic that determines whether a multi-model panel actually diversifies the risk.","tier":"system","section":"Correlated versus independent error","interaction_risk":false,"status":"active","source_ids":["s1"],"why_material":"Correlated error is the tail risk a panel cannot be priced without.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c5","text":"The derivation-agreement gate fails closed: a unanimous verdict was refused because two seats derived it through different trigger states, converting a would-be undetected error into a detected, receipted deferral to a human.","tier":"system","section":"Why the loading collapses","interaction_risk":false,"status":"active","source_ids":["s2","s3"],"why_material":"Detected deferral is a priceable event; undetected error is the fraud/opacity loading.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c6","text":"A sealed authorisation on record shows every seat firing the same clauses in the same trigger states on the same evidence — the artifact a parametric trigger can reference.","tier":"system","section":"A parametric trigger","interaction_risk":false,"status":"active","source_ids":["s5"],"why_material":"A claim event definable from the receipt alone removes the loss-adjustment dispute.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c7","text":"Malformed findings — invented clauses, missing fields, no terminal decision line — are voided by a deterministic parser and can never authorise, and the gate's own one recorded failure (false convergence on clause numbers) is documented with its fix.","tier":"system","section":"Why the loading collapses","interaction_risk":false,"status":"active","source_ids":["s2","s6"],"why_material":"Fail-closed behaviour plus a documented self-caught failure is the moral-hazard answer.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c8","text":"A governed call costs $0.0006 to $0.0024 and a three-model sealed decision $0.0049, so putting the measurement on every covered decision costs effectively nothing against the insured exposure.","tier":"system","section":"The economics","interaction_risk":false,"status":"active","source_ids":["s4"],"why_material":"Removes the economic objection to per-decision evidence as a policy condition.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c9","text":"The same machinery separates specification failure from model failure: a governed critique of a case file found eight input defects, the lead one a necessity-stated-as-sufficiency error that had caused every prior divergence.","tier":"system","section":"The coverage boundary","interaction_risk":false,"status":"active","source_ids":["s7"],"why_material":"Whether the insured's policy text or the model caused the loss is the coverage dispute; here it is decidable from receipts.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c10","text":"No calibration study establishes correctness at a known rate; the published rates cover one task class with small n; and the genuine APPROVE used two model families, not three.","tier":"system","section":"What is not satisfied","interaction_risk":false,"status":"active","source_ids":[],"why_material":"An underwriter must not be sold more than the evidence supports; these are the exact gaps a pilot must close.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false}],"sources":[{"id":"s1","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act","title":"Measured per-model error rates under a fixed rule set","summary":"Per-model error rates on a hashed suite, with Krippendorff alpha and Fleiss kappa — the agreement statistics that separate correlated from independent error — and the prevalence paradox stated rather than hidden.","claim_ids":["c2","c4"],"hash":"4c96267182b5fa693dacd8775133020f2a0e6554ef508a65548615a6cb30c0b7"},{"id":"s2","type":"live_surface","url":"https://miscsubjects.com/a/auditable-reasoning-hardened","title":"The derivation-agreement gate — fail-closed by construction","summary":"Independent models under a pinned rule set; the gate refuses to authorise when their clause-by-clause derivations diverge, even on a unanimous verdict. Includes the false-convergence defect and its documented fix.","claim_ids":["c5","c7"],"hash":"34d9af0f41b7371ae9454462d2a002b69263e9144890b564ef5e5b61e3552416"},{"id":"s3","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_o6s0exhodd","title":"A unanimous verdict, refused on divergent derivation","summary":"Three models returned the same verdict citing the same clauses; two derived it through different trigger states, so the gate escalated instead of concluding — a detected deferral instead of an undetected error.","claim_ids":["c5"],"hash":"656c222e8c45b003ca5c5c033d7641ec1704305660c9f9e61510777dacef3e00"},{"id":"s4","type":"live_surface","url":"https://miscsubjects.com/a/auditable-reasoning-audited","title":"The 72-call variance study: cost and the governed structure","summary":"Three prompt arms x three models x eight runs. Auditable structure appeared in 0 of 48 ungoverned calls; clause-citation Jaccard rose 0.74 to 0.95 under the constitution; a governed call costs $0.0006-$0.0024 and a three-model sealed decision $0.0049.","claim_ids":["c3","c8"],"hash":"3cb1906262064b1a738597de574a6c1d82fa69f9d9ab0c7c723ffb6f7949c1b5"},{"id":"s5","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_wl0rnh136b","title":"The genuine APPROVE — unanimous verdict, identical derivation","summary":"The one clean authorisation on record: every seat fired the same clauses in the same trigger states on the same evidence. What a covered, sealed decision looks like.","claim_ids":["c6"],"hash":"d52b245867de8ccbe4233030937a8b8acde9d4f4a0f1f2378c1302757495936c"},{"id":"s6","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_2dsklah529","title":"A structurally invalid finding, voided","summary":"The cheapest seat cited clauses 7, 8 and 12 of a six-clause rule set. A deterministic parser voided the finding; malformed output can never authorise. The fail-closed floor an underwriter can rely on.","claim_ids":["c7"],"hash":"1097f1a99705789de3eb97841c7a9aee0e523e38de1110d93a78104201ab7137"},{"id":"s7","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_qh3ge2x74b","title":"The instrument auditing its own input: eight defects","summary":"A governed model asked to critique the case input found the rule set stated only a necessary condition where a sufficient one was needed — separating specification failure from model failure, which is the coverage boundary.","claim_ids":["c9"],"hash":"48f911bee33eff0baef1e2cd01af9cecc27b3b99af67d135edeba42ec3e2e507"}],"anecdotal_sources":[],"scientific_sources":[],"user_reports":[],"related_articles":[],"question_graph":{"slug":"insurer-ai-performance-rate-table","questions":[],"evidence":[],"edges":[],"counts":{"questions":0,"evidence":0,"edges":0}},"honesty":{"active_claims":10,"retracted_claims":0,"cut_claims":0,"challenges":0,"scrub_events":0,"note":"Retracted/cut claims stay on ledger but are excluded from ask unless ?include_inactive=1"},"counts":{"claims":10,"claims_total":10,"sources":7,"anecdotal":0,"scientific":0,"user_reports":0,"questions":0,"evidence_ingests":0}}