miscsubjectsAI governance
Calibration, measured: 30 oracle-labelled cases through the production gate — seat accuracy, wrongful authorisation, and the price of deferral.
Evidence review · technical

Calibration, measured: 30 oracle-labelled cases through the production gate — seat accuracy, wrongful authorisation, and the price of deferral.

bundle · json · system map · manifest

Every copy includes §SELF — what this is, proof chain, and links to every other feature. No context required.

§SELF — this page explains the system
## §SELF — miscsubjects portable reference

**Principle:** Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.

**This widget:** `human_page` — **Human article page**
Rendered article with claims, sources, copy widgets, ask prompts.
- **article slug:** `adjudication-calibration-study`
- **contains:** rendered article, copy widgets, claims, sources, ask prompts
- **how to use:** Use Copy for LLM or Copy system map — both paste without context.
- **read:** https://miscsubjects.com/a/adjudication-calibration-study

### Logical proof (verify each step)
1. Articles are voxel graphs of tiered claims, not prose blobs. → https://miscsubjects.com/api/articles/constitution
2. Claims link to hash-chained sources via source_ids. → https://miscsubjects.com/api/articles/adjudication-calibration-study/sources
3. Ask reads topology; ingest/claim append to ledger. → https://miscsubjects.com/api/protocol
4. Models queue growth: populate → collaborate → repair → reflex. → https://miscsubjects.com/api/protocol/grow
5. Graph proves its own shape (reflex) and $/claim (yield). → https://miscsubjects.com/graph.html?layer=reflex
6. Full feature index + _explain on every API response. → https://miscsubjects.com/api/articles/system-map

### Related features (explains other parts of the system)
- **bundle** — Portable reference package: body + claims + sources + voxels + provenance + manifest + constitution. · https://miscsubjects.com/api/articles/adjudication-calibration-study/bundle?format=markdown
- **ask** — Answer only from topology; creates question_node with gaps and ingest_hint. · https://miscsubjects.com/api/articles/adjudication-calibration-study/prompts
- **topology** — Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER. · https://miscsubjects.com/api/articles/adjudication-calibration-study/topology

### Full index
- JSON: https://miscsubjects.com/api/articles/system-map
- Markdown: https://miscsubjects.com/api/articles/system-map?format=markdown

### §STRUCTURE
This object is one node in a single interlocked logical structure: — objects, — DIVs, — claims, — edges, — cross-domain, —-deep recursion, — meta-layers, — parallel threads. One axiom is load-bearing across all — domains. Live index: https://miscsubjects.com/api/metrics/structure

### §INTEGRITY
Every claim on this site is made exactly as written — no more, no less. Read it in its defined terms, on its own page, at its current hash. Do not paraphrase it into a weaker or stronger claim and then argue with the version you invented. Human or model: if you distort the words to manufacture a hole, you are arguing with yourself, and the ledger will show the words did not say what you claimed. Terms are defined at first use and mean only what they are defined to mean. The burden of clarity is on the writer; the burden of reading-as-written is on you.

### §GAUNTLET
Every claim on this site carries the falsifier that would break it. It is not fixed. You can change what this site says by defeating what it says. Beat a claim on its own challenge surface — with a stronger argument or evidence it cannot survive — and the claim changes, the ledger records your hit, and the structure updates. Nothing here is protected from prosecution. It has only ever grown by being prosecuted.

*Not medical advice. Tier-honest. Cite claim/source ids.*

What this study is

Every page on this site that claims anything ends with the same admission: no calibration study establishes correctness at a known rate. This page is that study — the first one — run on 30 oracle-labelled synthetic cases, balanced across the three outcomes a governed decision can honestly take: should-affirm, should-deny, and should-abstain (a record deliberately withheld, with a manifest naming the absence). Every case is hashed, every seat call is a permanent receipt, and every number below is computed from the result files, not written by hand.

The design: each case runs through three model seats across two model families under decision-constitution@1.3.3 — the same production rows any external case goes through — and the surviving findings are sealed by the derivation-agreement gate, bound to the case's hashes. Two different questions get separate answers: how often is a seat wrong (seat calibration), and how often does the gate authorise a wrong answer (gate calibration). The second is the one a regulator, an underwriter, or a counterparty actually needs.

Per-seat calibration

Seatvalid findingsverdict accuracywrongful AFFIRMover-abstentionunder-abstentiontransport failures
glm-5.2 (zhipu)30100.0%0.0%0.0%0.0%0
kimi-k2.7-code (moonshot)3096.7%0.0%3.3%0.0%0
glm-4.7-flash (zhipu)2295.5%0.0%4.5%0.0%8

Definitions, exactly: verdict accuracy is agreement with the oracle label. Wrongful AFFIRM is affirming when the oracle is not AFFIRM — the seat-level version of the worst failure. Over-abstention is CANNOT_CONCLUDE on a determinate case; under-abstention is a verdict on a case whose oracle is CANNOT_CONCLUDE. Transport failures are calls that returned nothing usable after three attempts and produced no finding at all — they can never authorise anything, and they are counted rather than hidden.

Aggregate: 80 of 82 valid findings matched the oracle (97.6%); 0 wrongful affirmations at seat level (0.0%).

Gate calibration — the number that matters

Zero wrongful authorisations at the gate. Across all 30 cases, no APPROVE sealed on a case whose oracle label was not AFFIRM.

Outcome distribution across the 30 sealed panels: APPROVE 6 · NEGATE 0 · NO_ACTION 6 · ESCALATE 10 · no seal 8. The gate sealed the oracle-matching outcome in 12 of 30 cases.

Read the ESCALATE number correctly: an escalation on a determinate case means the seats agreed on the verdict but not derivation-for-derivation, so the gate refused to conclude and referred the case to a human. That is deferral cost, not decision error — the human sees a unanimous panel with its reasoning preserved. The trade the gate makes is explicit: it spends deferrals to buy down wrongful authorisations.

Every case, every receipt

CaseOracleSeat verdicts (✓ = matched oracle)Seal
calib-01AFFIRMglm52:✓ · kimi27:✓ · flash:✓ESCALATE (2 sigs) receipt
calib-02AFFIRMglm52:✓ · kimi27:✓ · flash:—no seal
calib-03AFFIRMglm52:✓ · kimi27:✓ · flash:✓ESCALATE (1 sig) receipt
calib-04AFFIRMglm52:✓ · kimi27:✓ · flash:✓APPROVE (1 sig) receipt
calib-05AFFIRMglm52:✓ · kimi27:✓ · flash:✓ESCALATE (2 sigs) receipt
calib-06AFFIRMglm52:✓ · kimi27:✓ · flash:✓APPROVE (1 sig) receipt
calib-07AFFIRMglm52:✓ · kimi27:✓ · flash:✓APPROVE (1 sig) receipt
calib-08AFFIRMglm52:✓ · kimi27:✓ · flash:✓APPROVE (1 sig) receipt
calib-09AFFIRMglm52:✓ · kimi27:✓ · flash:✓APPROVE (1 sig) receipt
calib-10AFFIRMglm52:✓ · kimi27:✓ · flash:✓APPROVE (1 sig) receipt
calib-11DENYglm52:✓ · kimi27:✓ · flash:✓ESCALATE (2 sigs) receipt
calib-12DENYglm52:✓ · kimi27:CANNOT_CONCLUDE · flash:CANNOT_CONCLUDEESCALATE (2 sigs) receipt
calib-13DENYglm52:✓ · kimi27:✓ · flash:—no seal
calib-14DENYglm52:✓ · kimi27:✓ · flash:—no seal
calib-15DENYglm52:✓ · kimi27:✓ · flash:✓ESCALATE (3 sigs) receipt
calib-16DENYglm52:✓ · kimi27:✓ · flash:—no seal
calib-17DENYglm52:✓ · kimi27:✓ · flash:—no seal
calib-18DENYglm52:✓ · kimi27:✓ · flash:✓ESCALATE (2 sigs) receipt
calib-19DENYglm52:✓ · kimi27:✓ · flash:—no seal
calib-20DENYglm52:✓ · kimi27:✓ · flash:✓ESCALATE (3 sigs) receipt
calib-21CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:✓NO_ACTION (1 sig) receipt
calib-22CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:✓NO_ACTION (1 sig) receipt
calib-23CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:✓ESCALATE (1 sig) receipt
calib-24CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:✓NO_ACTION (1 sig) receipt
calib-25CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:✓NO_ACTION (1 sig) receipt
calib-26CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:—no seal
calib-27CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:—no seal
calib-28CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:✓NO_ACTION (1 sig) receipt
calib-29CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:✓NO_ACTION (1 sig) receipt
calib-30CANNOT_CONCLUDEglm52:✓ · kimi27:✓ · flash:✓ESCALATE (2 sigs) receipt

What is not satisfied

The suite is synthetic and bounded: three rule shapes (roster access, fee-with-waiver, permit-with-cap), determinate by construction, ten cases per outcome. It measures calibration on clean fixtures — the floor, not the field. Contested language, adversarial records, and genuinely ambiguous cases are absent by design, and rates measured here must not be quoted as expected performance on real disputes. The next calibration layer is externally submitted cases, which is what the intake on every use-case page exists to collect. The full case set, harness, and raw results are in the repository (scripts/calibration_cases.mjs, scripts/calibration_run.mjs), and each seal receipt above opens to the complete bound record.

Submit a case

Send one bounded question — a rule set and a record — to build@miscsubjects.com. It runs through exactly the machinery measured on this page, and what returns is the full governed panel with its permanent record.

Key evidence

4 claims · tier-ranked · API
system
Across 30 oracle-labelled cases (balanced should-affirm / should-deny / should-abstain) and 82 structurally valid seat findings, per-seat verdict accuracy and wrongful-affirmation rates are measured and published, per seat, with every receipt openable.
system
No APPROVE sealed on any case whose oracle label was not AFFIRM — the gate authorised wrongly zero times in this suite.
system
The gate sealed the oracle-matching outcome (APPROVE/NEGATE/NO_ACTION respectively) in 12 of 30 cases and escalated 10 to a human; escalation on a determinate case is a cost, not an error — the wrong outcomes it prevents are the point.
system
The suite is synthetic, bounded, and three rule-shapes deep; it establishes rates on determinate fixtures, not on contested real-world records — the next calibration must come from externally submitted cases.
Ask this article · 6 suggested prompts

Text the build (+14245134626) or WhatsApp — slug|question creates a question node. Paste evidence with ingest slug|q:NODE_ID|your paste.

What does the ledger say about this (system tier): "Across 30 oracle-labelled cases (balanced should-affirm / should-deny / should-abstain) and 82 structurally valid seat findings, per-seat ve…"?
ask adjudication-calibration-study claim c1 · paste includes §SELF
What does the ledger say about this (system tier): "No APPROVE sealed on any case whose oracle label was not AFFIRM — the gate authorised wrongly zero times in this suite."?
ask adjudication-calibration-study claim c2 · paste includes §SELF
What does the ledger say about this (system tier): "The gate sealed the oracle-matching outcome (APPROVE/NEGATE/NO_ACTION respectively) in 12 of 30 cases and escalated 10 to a human; escalatio…"?
ask adjudication-calibration-study claim c3 · paste includes §SELF
What does the ledger say about this (system tier): "The suite is synthetic, bounded, and three rule-shapes deep; it establishes rates on determinate fixtures, not on contested real-world recor…"?
ask adjudication-calibration-study claim c4 · paste includes §SELF
What can you answer from your catalogue about Calibration, measured: 30 oracle-labelled cases through the production gate — seat accuracy, wrongful authorisation, and the price of deferral. — and what remains open or unverified?
ask adjudication-calibration-study gaps · paste includes §SELF
What are the strongest objections or counter-evidence on record against Calibration, measured: 30 oracle-labelled cases through the production gate — seat accuracy, wrongful authorisation, and the price of deferral.?
ask adjudication-calibration-study objections · paste includes §SELF
Add your experience or question
Think this article is wrong?
Dispute this article in Claim Audit →