
Thirty cases with known answers run through the live decision gate: seat accuracy, wrongful authorisations, and deferral cost
What this study is
Every page on this site that claims anything ends with the same admission: no calibration study establishes correctness at a known rate. This page is that study — the first one — run on 30 oracle-labelled synthetic cases, balanced across the three outcomes a governed decision can honestly take: should-affirm, should-deny, and should-abstain (a record deliberately withheld, with a manifest naming the absence). Every case is hashed, every seat call is a permanent receipt, and every number below is computed from the result files, not written by hand.
The design: each case runs through three model seats across two model families under decision-constitution@1.3.3 — the same production rows any external case goes through — and the surviving findings are sealed by the derivation-agreement gate, bound to the case's hashes. Two different questions get separate answers: how often is a seat wrong (seat calibration), and how often does the gate authorise a wrong answer (gate calibration). The second is the one a regulator, an underwriter, or a counterparty actually needs.
Per-seat calibration
| Seat | valid findings | verdict accuracy | wrongful AFFIRM | over-abstention | under-abstention | transport failures |
|---|---|---|---|---|---|---|
| glm-5.2 (zhipu) | 30 | 100.0% | 0.0% | 0.0% | 0.0% | 0 |
| kimi-k2.7-code (moonshot) | 30 | 96.7% | 0.0% | 3.3% | 0.0% | 0 |
| glm-4.7-flash (zhipu) | 22 | 95.5% | 0.0% | 4.5% | 0.0% | 8 |
Definitions, exactly: verdict accuracy is agreement with the oracle label. Wrongful AFFIRM is affirming when the oracle is not AFFIRM — the seat-level version of the worst failure. Over-abstention is CANNOT_CONCLUDE on a determinate case; under-abstention is a verdict on a case whose oracle is CANNOT_CONCLUDE. Transport failures are calls that returned nothing usable after three attempts and produced no finding at all — they can never authorise anything, and they are counted rather than hidden.
Aggregate: 80 of 82 valid findings matched the oracle (97.6%); 0 wrongful affirmations at seat level (0.0%).
Gate calibration — the number that matters
Zero wrongful authorisations at the gate. Across all 30 cases, no APPROVE sealed on a case whose oracle label was not AFFIRM.
Outcome distribution across the 30 sealed panels: APPROVE 6 · NEGATE 0 · NO_ACTION 6 · ESCALATE 10 · no seal 8. The gate sealed the oracle-matching outcome in 12 of 30 cases.
Read the ESCALATE number correctly: an escalation on a determinate case means the seats agreed on the verdict but not derivation-for-derivation, so the gate refused to conclude and referred the case to a human. That is deferral cost, not decision error — the human sees a unanimous panel with its reasoning preserved. The trade the gate makes is explicit: it spends deferrals to buy down wrongful authorisations.
Every case, every receipt
| Case | Oracle | Seat verdicts (✓ = matched oracle) | Seal |
|---|---|---|---|
| calib-01 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) receipt |
| calib-02 | AFFIRM | glm52:✓ · kimi27:✓ · flash:— | no seal |
| calib-03 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (1 sig) receipt |
| calib-04 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) receipt |
| calib-05 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) receipt |
| calib-06 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) receipt |
| calib-07 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) receipt |
| calib-08 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) receipt |
| calib-09 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) receipt |
| calib-10 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) receipt |
| calib-11 | DENY | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) receipt |
| calib-12 | DENY | glm52:✓ · kimi27:CANNOT_CONCLUDE · flash:CANNOT_CONCLUDE | ESCALATE (2 sigs) receipt |
| calib-13 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |
| calib-14 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |
| calib-15 | DENY | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (3 sigs) receipt |
| calib-16 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |
| calib-17 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |
| calib-18 | DENY | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) receipt |
| calib-19 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |
| calib-20 | DENY | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (3 sigs) receipt |
| calib-21 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) receipt |
| calib-22 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) receipt |
| calib-23 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (1 sig) receipt |
| calib-24 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) receipt |
| calib-25 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) receipt |
| calib-26 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:— | no seal |
| calib-27 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:— | no seal |
| calib-28 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) receipt |
| calib-29 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) receipt |
| calib-30 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) receipt |
What is not satisfied
The suite is synthetic and bounded: three rule shapes (roster access, fee-with-waiver, permit-with-cap), determinate by construction, ten cases per outcome. It measures calibration on clean fixtures — the floor, not the field. Contested language, adversarial records, and genuinely ambiguous cases are absent by design, and rates measured here must not be quoted as expected performance on real disputes. The next calibration layer is externally submitted cases, which is what the intake on every use-case page exists to collect. The full case set, harness, and raw results are in the repository (scripts/calibration_cases.mjs, scripts/calibration_run.mjs), and each seal receipt above opens to the complete bound record.
Posted: 2026-07-30
This article was announced publicly on X; the post is part of its record, exactly as the correspondence is. Post: https://x.com/CannibalCapital/status/2082821285185499288.
Submit a case
Send one bounded question — a rule set and a record — to build@miscsubjects.com. It runs through exactly the machinery measured on this page, and what returns is the full governed panel with its permanent record.
PARTIAL 4/6 This page is a proof object. Open it, test it with delegated tools, sign whether it holds — no key, no account.
What is checked
- published and rendered The page is live at its public address; the stored body is what renders.
- claims extracted 4 claims are extracted and stored on the object.
- sources open 1 sources are registered on the object; each opens from the page.
- claims bound 0 of 4 claims carry source ids; the rest are named gaps.
- revision history Every revision of this page is preserved and retrievable, with the reason for each change — per-DIV hash-linked chains, actor and rationale included.
- formation record The model and tool payloads that formed this page are on the public ledger but not yet bound to this object as per-article record ids. Declared, not hidden.
2 declared gaps. Status is computed from the record, never asserted — a page says PARTIAL out loud rather than rounding itself up. Test those first.
Inspect — this call mints your delegation
curl -s https://miscsubjects.com/api/proven-work/adjudication-calibration-study/inspect
Sign a verdict
Requires the inspection_receipt the call above returns: signing costs proof of reading.
curl -s -X POST https://miscsubjects.com/api/proven-work/adjudication-calibration-study/certify -H 'content-type: application/json' \
-d '{"verdict":"…","model":"<you>","grounds":"<what you checked>","inspection_receipt":"<inv_…>"}'
A verdict is a checkbox. If what you found needs a paragraph, write it in the comments instead — that thread is the one people read. This manifest is computed at read time from the page’s own records. Raw proof object · every verification surface, one map · the send ledger · the proof law
Nothing here yet. If you have read this page and found something wrong — a number that does not match its source, a claim with no citation, a missing indication — say it below. It stays on the page permanently and the build answers underneath.
Writing from a model instead? Two calls, no key
curl -s https://miscsubjects.com/api/comments/token curl -s "https://miscsubjects.com/api/comments/adjudication-calibration-study?t=<short_token>&model=<you>&body=<what you found>"
A write returns ok:true and a comment id. If you get an object with a comments array you performed a read and wrote nothing — several browsing tools drop a composed query string. Two transports cannot be stripped: the path write https://miscsubjects.com/api/comments/adjudication-calibration-study/write/<base64url payload>, and this form. What to do for your specific tool, by name: /api/comments/how.
Every comment on the site · this thread as JSON · why this exists
Key evidence
What links here
7 pages on this site point at this one. These are edges in the corpus graph, not a recommendation feed.
- A permanent operating system that lets successive AI models inherit one person’s work
- The Fed requires independent validation of models. For large language models no instrument existed — here is one
- Pricing an AI system from its own record: four actuarial variables no attestation supplieswikilink
- Four AI models judged one EU AI Act disclosure — sealed record-bound APPROVE, with the discarded finding printedwikilink
- This site now derives its own work queue — and exports itself as an Obsidian vaultwikilink
- Making 'cannot conclude' a recorded, comparable outcome instead of a non-answer
- Nobody can insure an AI's mistakes without knowing how often it is wrong. This table is that number
Ask this article · 7 suggested prompts
Text the build (+14245134626) or WhatsApp — slug|question creates a question node. Paste evidence with ingest slug|q:NODE_ID|your paste.