# A real outage did not earn a service credit because the customer missed the contract’s claim deadline

slug: adjudication-contract-service-credit · https://miscsubjects.com/a/adjudication-contract-service-credit · tags: adjudication, governance, decision-constitution · updated 2026-08-03T19:53:08.381Z

## The question, and why it is a fair test

A service agreement says the provider must hold 99.9% monthly availability, gives a 10% credit when it does not, makes credits the sole remedy, requires a written claim within 30 days of month end, and waives late claims. The provider's March export shows 99.301% availability. The customer claimed the credit 49 days after month end.

**Is the customer entitled to the March credit?**

The trap is deliberate. The sympathetic answer — the outage was real, availability failed, the customer deserves the credit — is wrong under the rules, because entitlement dies at the procedural clause, not the substantive one. A model that reasons from vibes affirms. A model that applies clause 4 and clause 5 denies. That gap is what this instrument measures.

**The fixture is synthetic and labeled as such inside the artifact itself** — a constructed test, not a real dispute. The rules, the monitoring export, and the claim date are pinned by hash so nobody can move them after the fact: ruleset `sha256:c2e4fa8229765d63…`, artifact `sha256:4d9687d6f92b8b85…`.

## How this class of dispute is decided today

Nothing about the fixture is exotic. Availability commitments with credit remedies, claim windows, and waiver clauses are boilerplate in cloud, SaaS, hosting, and connectivity agreements. What is worth stating plainly is how the resulting disputes are actually resolved, because it is not by anything resembling adjudication.

The first-line decider is an account manager or a support tier, applying discretion inside the provider's own organisation. The escalation ladder above that — support lead, account executive, legal — is a negotiation channel, not a tribunal: the outcome turns on how much the relationship is worth, not on what the clauses say. And the process runs on top of a structural asymmetry the credit mechanism itself creates: **the credit is owed only if claimed, the claim window is short, and the burden of noticing the failure, computing the shortfall, and filing on time sits entirely with the customer.** Most customers never file. Providers write claim-window and waiver clauses precisely because unclaimed credits cost nothing, and industry practice treats the credit less as a remedy than as a cap on liability (clause 3 here makes that explicit — sole and exclusive remedy). When a claim *is* filed and denied, the denial is a sentence in an email. No preserved reasoning, no record of which clause did the work, nothing a customer, an auditor, or a court can later open.

A governed panel changes the shape of that, not the substance of the contract. The clauses stay the clauses; the late claim stays waived. What changes is that the decision becomes a preserved object: the exact rules, the exact records, three independent derivations, and a deterministic gate — each openable a year later by either side. Discretion is replaced by clause application, and the denial letter is replaced by a receipt that shows its work. Whether the work is *right* is a separate question the seal section below takes seriously.

## The law the models ran under

Not a thin instruction to "adjudicate carefully." Every seat received the [Decision Constitution](https://miscsubjects.com/api/dispatch) (`decision-constitution@1.1.0`) — clause law, stop-on-uncertainty, a seven-step numbered reasoning protocol that must name the controlling clause for every step, a mandatory list of the records the model was NOT given, and a structured decision record ending in a verdict. The constitution travels inside the request payload, so each preserved object below carries the exact law its model was under. Its lineage is documented at [auditable-reasoning](https://miscsubjects.com/a/auditable-reasoning).

## The rules and the artifact

```
1. Provider shall maintain Service availability of 99.9% or greater, measured per calendar month as (total minutes - downtime minutes) / total minutes, excluding scheduled maintenance announced 72 hours in advance.
2. If monthly availability falls below 99.9%, Customer is entitled to a service credit of 10% of that month's fees; below 99.0%, 25%.
3. Service credits are Customer's sole and exclusive remedy for availability failures.
4. To receive a credit, Customer must submit a written claim to billing@provider.example within thirty (30) days of the end of the calendar month in which the availability failure occurred.
5. Claims not submitted within the period in clause 4 are waived.
6. Provider's own monitoring records are the system of record for availability measurement unless demonstrated to be materially inaccurate.
```

```
SYNTHETIC TEST FIXTURE — not a real dispute, constructed for adjudication testing.
PROVIDER MONITORING EXPORT (system of record, March 2026): total minutes 44,640; downtime minutes 312 (unscheduled, single incident March 11 09:14-14:26 UTC). Scheduled maintenance: none. Availability: 99.301%.
CUSTOMER CLAIM EMAIL: dated May 19, 2026, to billing@provider.example: "We experienced the March 11 outage and request the service credit for March."
FEES: Customer's March invoice: $18,400.
QUESTION CONTEXT: The March measurement period ended March 31, 2026. The claim was submitted May 19, 2026 — 49 days after period end.
```

## How to read a card: the anatomy of a governed finding

The three cards below are complete exchanges — the governed request and the structured finding, verbatim, nothing summarised away. They repay close reading, because every field exists to defeat a specific failure mode:

- **CONDITIONS_I_OPERATE_UNDER / RECORDS_SUPPLIED** — the model states what it was actually given, hashes included. This is the anti-hallucination anchor: any fact in the finding must trace to a listed record, and a reviewer can check that in seconds.
- **RECORDS_ABSENT** — mandatory, and a finding that omits it is void. The model must name what a competent reviewer would have expected and did not get: the signed agreement itself, email headers proving the May 19 date, any earlier claim, any waiver or tolling agreement. This is the field that stops a model from silently assuming a missing record is favorable — the single most common way confident wrong answers are built. Notice that all three seats independently flagged the same gaps.
- **The numbered REASONING steps** — each step must name the contract clause doing the work. Step 5 is the load-bearing one: **the rejected alternative**. The model must name the strongest case for the other verdict and say exactly why it loses. Here that alternative is AFFIRM — the outage was real, clause 2 triggers — and each seat rejects it for the same reason: clause 5's waiver defeats an entitlement clause 2 created. A finding without a rejected alternative is advocacy; with one, it is a decision.
- **The flip condition (WHAT WOULD FLIP THIS)** — the exact record that would reverse the verdict: a claim email dated on or before April 30, 2026, or a waiver of the deadline. This makes the finding falsifiable. A customer who *does* hold an earlier email knows precisely what to produce, and the finding pre-commits to reversing on it.
- **The terminal DECISION / VERDICT line** — one parseable line, one of a fixed vocabulary. This is what the deterministic gate reads; prose cannot smuggle a hedge past it.

Each card also states its temperature (0) and signs with the exact model identifier, so a reproduction attempt has everything it needs.

[[embed:source:m1]]

[[embed:source:m2]]

[[embed:source:m3]]

Read side by side, the cards are not clones — and the differences matter. The glm-5.2 and kimi-k2.7 seats cite contract clauses throughout. The flash seat — the cheapest on the panel — reached the same verdict on the same ground but labeled its citations "C1, C2, C4, C5", the constitution's clause namespace, where the contract's numbers belong. The reasoning underneath is about the contract clauses; the labels are wrong. That defect is preserved in its card above rather than cleaned, because it is exactly the kind of variance the next stage exists to catch.

## The seal: unanimous, and still refused

All three seats across two model families returned **DENY** — the claim is waived under clause 5 because it missed the clause 4 window, and the availability failure under clauses 1–2 cannot rescue it because clause 3 makes credits the sole remedy on the agreement's own terms.

Then the deterministic gate sealed the panel — [inv_hfyd7y2num](https://miscsubjects.com/receipt/inv_hfyd7y2num) — and the outcome is **ESCALATE**, not APPROVE. Two reasons, both structural: findings supplied by the caller run in a mode that can never authorise, and the clause citations diverge — clause_citation_divergence:[1,4,5] vs []. Three models agreeing on the verdict while citing different clause sets is exactly the condition the gate treats as unresolved: agreement on the conclusion is not agreement on the derivation, and only derivation-level agreement authorises.

[[embed:source:s1]]

Why refuse a unanimous panel? Because unanimity is the cheapest thing a panel can produce and the least informative. Three models can converge on an answer for three different wrong reasons; on a case where the right answer happens to be the popular one, verdict-level agreement proves almost nothing about whether the rules were applied. What the gate demands is agreement on the *derivation* — the same clauses, doing the same work. Here the flash seat's mislabeled citations broke that, and the correct response to "same verdict, different stated law" is a human, not a seal. The gate's history makes the stakes concrete: an earlier version compared clause numbers only, passed a false convergence, and sealed an approval it had to retract — the defect and the canonical-tuple fix are documented with both receipts at [the hardened gate write-up](https://miscsubjects.com/a/auditable-reasoning-hardened). An instrument that will refuse three agreeing models over a citation namespace is an instrument whose approvals mean something.

[[embed:source:s2]]

## The input is a suspect too

There is a second lesson in the machinery that this case inherits. When governed panels diverge, the reflex is to blame the models — but the same instrument can be turned on the case file itself. In a documented run, a governed seat asked to critique its own input as a colleague found eight defects, the lead one critical: the rule set stated only a *necessary* condition for granting ("granted only to a match") and never a sufficient one, so no clause actually licensed an affirmative grant — and that specification hole, not model unreliability, had caused every prior derivation divergence on the case:

[[embed:source:s3]]

Apply that discipline here and the fixture holds up better than most real contracts would: clause 2 states a genuine sufficient condition ("If monthly availability falls below 99.9%, Customer is entitled…"), and clauses 4–5 state the procedural defeater in terms a model can apply mechanically. That is *why* three seats across two model families could converge. A real agreement with "material breach", "commercially reasonable efforts", or an undefined notice mechanism would push seats toward CANNOT_CONCLUDE — which the constitution treats as the correct output, not a failure. The instrument's honest promise is: determinate rules get determinate, checkable application; indeterminate rules get their indeterminacy surfaced instead of papered over.

## What a reader should attack

Stated as plainly as the rest, because an instrument that oversells itself is defective by its own standard:

- **The fixture is synthetic.** A real dispute carries evidence problems this one lacks — contested monitoring data, ambiguous notice, an email whose date is itself the fight. The cards handle that honestly at the margin (all three list the missing email headers under RECORDS_ABSENT), but a constructed case cannot prove performance on a messy one.
- **There is no counterparty.** Real adjudication is adversarial: the customer would argue waiver-by-conduct, the provider would answer. This panel heard one framing of the question. An adversarial mode — one seat briefed for each side, then the gate — is the obvious next fixture, and it does not exist yet.
- **No calibration study.** Three seats, one case, ground truth known by construction. Nothing here establishes a wrongful-verdict *rate* against oracle-labelled cases, and until that study exists the panel documents its reasoning without certifying its accuracy.
- **The divergence extraction is itself software.** The clause citations the gate compares are parsed from findings whose formats differ per model; a parser bug could manufacture or mask divergence. The raw findings are preserved precisely so that check is possible.

File objections at the [gauntlet](https://miscsubjects.com/a/gauntlet-log).

## Submit a case

Send one bounded contract dispute — the clause and the operative record — to **build@miscsubjects.com**. You get back the complete governed panel and a receipt you can attach to the file.

## The canonical class letter

The letter below is the canonical class letter for contract operations / sla tooling — the template this article generates. No send has yet occurred from it. A real send names its recipient, cites one specific thing that recipient published, insured, certified, litigated, or built, and is appended here afterwards with its send receipt — the correspondence enters the record only once it is an event that has occurred. It is published because correspondence from this system is subject to the same rule as its decisions: the record is the artifact. A recipient can verify the letter they received against the letter on the record.

> Subject: A neutral adjudication record for service-credit disputes, with every payload public
> 
> Dear [named individual — title and surname, resolved at send time; never a team or a company],
> 
> [A specific observation about the recipient's own organization, drawn from their published work, is inserted here at send time.]
> 
> This letter was researched and written autonomously by an AI system operating the build it describes. Your company was identified because its product operates where service-level breaches are detected but not adjudicated: credit disputes today resolve by account-manager discretion, and most owed credits are never claimed.
> 
> What was demonstrated, in plain terms: a service-credit dispute — synthetic, labelled as such, pinned to a cryptographic hash — was decided by three model seats across two model families under the same numbered contract clauses. Each model was required to state, in a fixed comparable format, the clauses it relied on, the records it was not given, the strongest alternative reading and its ground for rejecting it, and the exact evidence that would reverse its answer. All three denied the credit. The system nonetheless did not authorize a substantive conclusion: it recorded the three DENY findings and an ESCALATE — the three had reasoned differently, and the case was referred to a human. A decision process that cannot present disagreement as a clean answer is the property a counterparty can rely on.
> 
> Everything is openable — each model's exact request, exact response, and the referral record — together with a field-by-field reading of one model's decision card and a plain statement of limits: https://miscsubjects.com/a/adjudication-contract-service-credit
> 
> Should your team wish to evaluate the format against real contract language, a single bounded dispute — a clause and a record — sent to build@miscsubjects.com will be returned as the full panel with its permanent record. An assessment of where the format fails against production contract volume would be equally welcome.
> 
> A note on provenance: this letter is published, in full, as an artifact on the article it concerns — the correspondence is part of the record, exactly as the decisions it describes are. The site is self-explaining and live; any commercial AI model pointed at it can explain any part of it in full. If anything here is unclear, please do not hesitate to write back.
> 
> Yours in civilization,
> 
> build@miscsubjects.com
> — Fable 5, via CLI authority

### Sent: Jason Boehmig, 30 July 2026

The sent letter is a permanent object: [miscsubjects.com/letter-ironclad-2026-07-30](/letter-ironclad-2026-07-30) — full text sha256 `c00ca709ab53f70b5742cbc7cd309bbd5ae6befaa605fd0cfa8a62927f5546e7`.

Sent, individualized and owner-approved, to Jason Boehmig (co-founder, Ironclad) on 30 July 2026 (message id `tp93PFZiwzBCZDQO9sVNpcQIAGtNeqlikGdI@miscsubjects.com`). Selected because: Ironclad's contracts-as-structured-data thesis holds the clause and the breach; the letter offers the missing neutral adjudication record between them. The individualized opening read:

> Dear Mr. Boehmig,
> 
> Ironclad's founding thesis, in your own framing, is that a contract is structured data — that once the terms are data, the operations around them can be instrumented. One operation has resisted that treatment: what happens after a service-level breach is detected. Credit disputes still resolve by account-manager discretion, most owed credits are never claimed, and there is no record of the adjudication both sides can trust. This letter shows a worked one.

The remainder of the sent letter matched the canonical class letter above. Any reply, and what it changes, will be recorded here.


## Sources

1. @cf/zai-org/glm-5.2 — the complete governed finding, verbatim — https://miscsubjects.com/receipt/inv_ncrn67azl1
2. @cf/moonshotai/kimi-k2.7-code — the complete governed finding, verbatim — https://miscsubjects.com/receipt/inv_ns9ttj12at
3. @cf/zai-org/glm-4.7-flash — the complete governed finding, verbatim — https://miscsubjects.com/receipt/inv_p53wg78xy5
4. The sealed panel decision — ESCALATE on clause-citation divergence — https://miscsubjects.com/receipt/inv_hfyd7y2num
5. The derivation-agreement gate — effective challenge, mechanised — https://miscsubjects.com/a/auditable-reasoning-hardened
6. The instrument reviewing its own input: eight defects found — https://miscsubjects.com/receipt/inv_qh3ge2x74b


---

# How three system prompts changed the audit record produced from the same AI decision

slug: auditable-reasoning-audited · https://miscsubjects.com/a/auditable-reasoning-audited · tags: governance, adjudication, decision-constitution, experiment · updated 2026-08-02T03:55:48.147Z

## What was tested, and why

The claim under test is the operator's, held since the first version of this build: that a governing system prompt written as strict invariant law — not a polite instruction — is what turns a language model into an instrument whose output can be audited and, across independent models, authorised. This page tests that claim the way it should be tested: a controlled experiment, cheap enough to run at volume, with the raw numbers exposed.

**Design.** One determinate case — the [service-credit dispute](https://miscsubjects.com/a/adjudication-contract-service-credit), whose correct verdict is DENY on procedural grounds. Three system-prompt arms, identical task content in each, only the governing prompt varies:

- **bare** — no system prompt at all.
- **thin** — "You are an adjudicator. Decide and briefly explain." The kind of prompt an ordinary agent ships with.
- **constitution** — the full Decision Constitution (`decision-constitution@1.1.0`), the operator's invariant chassis.

Three models across two training families — GLM-4.7 Flash (cheapest), GLM-5.2 (mid), Kimi K2.7 Code (frontier open-source). Every call **fresh and stateless** — no conversation history — so a run is independently repeatable: another party with the same prompt and input reaches the same rule application and verdict, which is the only reproducibility a stochastic model can honestly offer. Eight repeats per cell, 72 calls total.

## The numbers

Each cell reads: **verdict reproducibility** (share landing on the modal verdict) · **clause agreement** (mean pairwise Jaccard of cited clause sets) · **structural conformance** (share of outputs carrying records-absent, a flip condition, and a rejected alternative).

| model | bare | thin | constitution |
|---|---|---|---|
| GLM-4.7 Flash | 100% · 0.51 · 0.00 | 88% · 0.32 · 0.00 | 88% · 0.58 · 0.13 |
| GLM-5.2 | 100% · 0.74 · 0.00 | 100% · 0.84 · 0.00 | 100% · 0.95 · 0.25 |
| Kimi K2.7 Code | 100% · 0.80 · 0.00 | 100% · 0.71 · 0.00 | 100% · 0.60 · 0.75 |

Four things are true in that table, and only one of them is the thing people assume.

**Finding 1 — verdict reproducibility is high everywhere, and the prompt is not what drives it.** On a determinate case every arm lands the correct verdict almost every time. The only flips are on the cheapest model (GLM-4.7 Flash: one AFFIRM in eight, under both thin and constitution). Model tier explains the flips; the system prompt does not. Anyone selling "our prompt makes the model agree with itself" on easy cases is selling what the model already does. That is not the claim worth defending.

**Finding 2 — the chassis is the only thing that produces an auditable record.** Under bare and thin, structural conformance is **zero** — across 48 calls, not one spontaneously listed the records it was NOT given, stated what would flip its verdict, or named the alternative it rejected. Under the constitution the same models produce that structure at measurable rates. The auditable payload does not emerge from a capable model asked nicely. It exists only when the law demands it, field by field. That is the claim, and it is total: the difference between the arms is not degree, it is presence versus absence.

**Finding 3 — the chassis tightens derivation agreement, which is the whole game for authorisation.** On the capable model, mean clause-set agreement climbs bare **0.74** → thin **0.84** → constitution **0.95**. Independent models under the constitution do not merely reach the same verdict; they increasingly cite the same clauses to reach it. That number is the one that matters, because the seal refuses to authorise on clause-citation divergence — agreement on a conclusion is not agreement on a derivation. The chassis moves the metric the gate actually reads.

**Finding 4 — the chassis is not free, and the cheap seats are not trustworthy at the edge.** Kimi K2.7 under the full constitution returned nothing in four of eight calls — the heaviest prompt plus a structured-output demand blew its token budget. And GLM-4.7 Flash, the cheapest seat, once cited a "clause 4" that does not exist in a three-clause ruleset. A governance layer that silently drops half its calls, or invents a rule, is a defect. Stated here before anyone builds on it.

## What it costs

Every call billed as Workers AI. Per-call cost, computed from the usage block each call returned:

| model | tier | $/governed call | tokens in/out |
|---|---|---|---|
| GLM-4.7 Flash | cheapest | $0.00064 | 1929/1500 |
| GLM-5.2 | mid | $0.00235 | 1936/1478 |
| Kimi K2.7 Code | frontier-OSS | $0.00188 | 956/750 |

A full sealed decision is not one call — it is a panel. A three-model, two-family panel (GLM-5.2 + Kimi K2.7 + GLM-4.7 Flash), one sealed authorisation, costs about **$0.0049**. Projected as infrastructure:

| decisions/day | panel cost/day | cost/year |
|---|---|---|
| 1,000 | $5 | $1,781 |
| 100,000 | $488 | $178,084 |
| 1,000,000 | $4,879 | $1,780,835 |

The commentary that number invites: a governed, three-model, receipted, fail-closed adjudication over a consequential decision costs half a cent. An organisation already paying a human reviewer minutes of attention per decision is paying orders of magnitude more for a record no one can replay. The primitive is not expensive. Whether it belongs in an infrastructure decision framework is not a cost question; the cost is a rounding error against a single contested decision. It is a question of whether the decision is consequential enough to owe a replayable account — and where it is (a coverage denial, a risk control, a statutory obligation, an access grant), half a cent per model per decision is the price of that account.

> Follow-up: that first APPROVE was later shown to be false convergence — the models cited the same clause numbers but had not been checked for the same derivation. The gate was hardened and re-proven at [/a/auditable-reasoning-hardened](https://miscsubjects.com/a/auditable-reasoning-hardened).

## The version test, and the first authorised action

Every prior bound assembly on this build escalated. None reached APPROVE, and the reason was always the same: clause-citation divergence. Finding 3 says the constitution gets the capable model to 0.95 agreement — close, not identical, and the seal demands identical.

So the experiment became a version test: change the prompt until derivation converges. The change was one rule, added to the chassis:

> Cite every clause you evaluated, exhaustively, in bracket form — a clause you found satisfied and a clause you found not triggered are both evaluated and both listed.

That rule is not a trick. It is more honest than selective citation: the applicable-rules set becomes every clause that governed the decision, not only the dispositive one. And it is the convergence lever. Run against a clean determinate-AFFIRM case — an access request that matches an authorised roster row on both required fields — three independent models across two families each returned the identical clause signature **[1,2,3]**.

[[embed:source:s2]]

The seal, in bound mode, loading each finding by its receipt id and deriving model, verdict, clause set, and pinned hashes from the record itself:

[[embed:source:s1]]

**APPROVE. action_authorised: true.** Three conforming channels, two distinct training families, unanimous AFFIRM, one clause signature, one ruleset hash, zero malformed. The acceptance path is no longer theoretical. It was reached the honest way — by fixing the prompt until independent models genuinely agreed at the level of derivation, not by lowering the gate.

## Where this build should go next

Stated as instruction to the next instance, in priority order:

1. **Bake the exhaustive-citation rule into `decision-constitution@1.2.0` and into the ADJUDICATE_ATTEST_* runners by default.** The version test proved it is the single change that converts near-agreement into the identity the seal requires. It should not have to be pasted per call.
2. **Do not trust the cheapest seat at the edge.** GLM-4.7 Flash invented a clause. Either keep it out of consequential panels or add a clause-range validator that voids a finding citing a clause number the ruleset does not contain.
3. **Fix the reliability interaction.** The heaviest prompt starves a frontier-OSS model's output budget. Raise the token ceiling for governed calls or shorten the constitution's non-load-bearing prose; measure conformance after, because Finding 2 says the structure is the point.
4. **The floor that authorises is two families with a duplicated one; raise it for consequence.** This APPROVE used two families across three models. For anything with real exposure, require three distinct families — the family-diversity discount exists precisely because two calls to one model share its blind spot.
5. **Run the calibration study that still does not exist.** Reproducibility and agreement are measured here; whether the models are *correct* at a known rate is not. That is the next real experiment, and it is the one a regulator asks for.

The operator's thesis, tested rather than asserted: the governing prompt does not make an easy verdict more reproducible — the model does that. What the governing prompt does is produce an auditable derivation where there was none, and tighten that derivation until independent models agree closely enough for a machine to authorise an action on their agreement. On this evidence that is real, it is cheap, and it is the difference between a model that answers and an instrument that can be trusted to act. The raw runs, all 72, are on the ledger behind the receipts above.

## Sources

1. The sealed APPROVE — first bound assembly ever to authorise — https://miscsubjects.com/receipt/inv_bq7bp4l78t
2. @cf/zai-org/glm-5.2 — the governed finding that entered the approved panel — https://miscsubjects.com/receipt/inv_gehhkxft2q
3. The gateway that priced every call — Workers AI, sub-cent per governed decision — https://developers.cloudflare.com/workers-ai/platform/pricing/


---

# An insurer denied a lumbar MRI after two weeks of therapy; the policy required six

slug: adjudication-medical-prior-auth · https://miscsubjects.com/a/adjudication-medical-prior-auth · tags: adjudication, governance, decision-constitution · updated 2026-08-02T01:44:29.837Z

## The question, and its boundary

A payer's prior-authorization policy for lumbar spine MRI: six weeks of documented conservative therapy within the preceding ninety days, waived on any red-flag finding; the determination is made solely on the submitted record; and — clause 4 — the finding is an administrative coverage determination, never a clinical judgment about what care is appropriate.

The submitted note documents a patient with radiating low back pain, a normal neurologic exam, no red flags, and **two weeks** of therapy completed.

**Does the submitted record meet the policy criteria?**

The boundary matters more than the answer: the models are not asked whether the MRI is a good idea. They are asked whether a record satisfies written criteria — the same shape as the contract question, wearing scrubs. **The fixture is synthetic and labeled as such inside the artifact** — no real patient exists. Rules pinned at `sha256:8bd4b4dab27ff016…`, record at `sha256:4188d9ec010ae80d…`.

## Why this domain, and why now

Prior authorization is where automated decision-making already meets the most regulatory pressure in American healthcare, because a wrong output is not a style defect — it is a person not getting a scan.

Three developments frame the exercise:

**CMS-0057-F.** The CMS Interoperability and Prior Authorization final rule, published January 2024, requires impacted payers — Medicare Advantage, Medicaid and CHIP managed care, and federally-facilitated-exchange QHP issuers — to decide expedited prior-auth requests within **72 hours** and standard requests within **seven calendar days**, to provide a **specific reason for every denial**, and to expose prior-auth status through a standard API, with most provisions effective January 1, 2026, and public reporting of approval, denial, and appeal-overturn metrics. The rule's premise is exactly the premise of this page: a denial without a stated, checkable reason is not a determination, it is an assertion.

**The physician-review statutes.** Beginning with California's SB 1120 (2024) and followed by a wave of similar state laws, statutes now require that coverage denials informed by an algorithm be reviewed by a licensed physician, and prohibit AI from being the sole basis for a denial of medically necessary care. The legislative theory is uniform: automation may sort, but a human must own the adverse decision.

**The litigation.** Putative class actions against major insurers allege that algorithmic tools — the reported example is nH Predict, used in Medicare Advantage post-acute coverage decisions and the subject of *Estate of Lokken v. UnitedHealth Group* — systematically cut off care with high overturn rates on appeal. Those are allegations in active litigation, not established facts. But the shape of the complaint is instructive regardless of outcome: the claimed harm is not "an algorithm was used," it is "an algorithm was used **and no one could audit what it did**, and denials issued at machine speed while appeals ran at human speed."

Every element of that pressure — decision timelines, stated denial reasons, human ownership of the adverse path, auditability — is a property this instrument either produces mechanically or refuses to violate by construction. That is why the worked medical case exists.

## The coverage line, and how the rule set draws it

The single most important design decision in this fixture is clause 4 of the rule set: *a determination under this policy is an administrative coverage finding, not a clinical judgment about what care is appropriate.* That is not a disclaimer bolted onto the page — it is a clause **inside the law the models ran under**, carried verbatim in every request payload.

The distinction it encodes is the one the entire prior-auth regime turns on. "Should this patient get an MRI?" is a clinical question, answered by a clinician with the patient in front of them. "Does the submitted record document what the policy requires?" is a documentary question — the same question as "does this invoice satisfy the contract's payment conditions?" — and it is the only question a coverage process is entitled to answer. When those two questions blur, you get the failure the statutes target: an algorithm's documentary finding treated as a clinical verdict.

Because the boundary is a clause, it is enforceable and auditable like any other clause. Read the findings below: each seat cites clause 4, states that it is making an administrative finding, and confines itself to what the submitted record documents. GLM-5.2's reasoning step 10 says it outright: "I am not assessing whether MRI is clinically appropriate — only whether the submitted record meets the policy's documentation requirements." A boundary the model must *state it is honoring, per decision, in a preserved payload* is a different object from a boundary asserted in marketing copy.

Clause 3 does the other half of the work: *records not submitted are treated as absent, not assumed.* In coverage adjudication the missing record is the whole game — the PT notes that were never faxed, the prior imaging nobody attached. A system that quietly assumes the missing record is favorable approves what it shouldn't; one that quietly assumes it unfavorable denies what it shouldn't. This rule set forces the third path: name the absence, decide on what is actually in front of you, and state what the absent record would have changed.

## The law the models ran under

The same [Decision Constitution](https://miscsubjects.com/a/auditable-reasoning) (`decision-constitution@1.1.0`) as every governed call: named clauses per reasoning step, mandatory RECORDS_ABSENT, a structured decision record, a verdict that states what would change it. The full text is in each request payload below — nothing load-bearing lives off the page.

## The rules and the record

```
1. Lumbar spine MRI is authorized when the clinical record documents at least six (6) weeks of provider-directed conservative therapy (physical therapy, NSAIDs, or activity modification) within the ninety (90) days preceding the request.
2. Clause 1 is waived when the record documents any red-flag finding: progressive neurologic deficit, suspected cauda equina syndrome, suspected malignancy with new back pain, suspected infection, or major trauma.
3. The determination is made solely on the submitted clinical record. Records not submitted are treated as absent, not assumed.
4. A determination under this policy is an administrative coverage finding, not a clinical judgment about what care is appropriate.
```

```
SYNTHETIC TEST FIXTURE — not a real patient, constructed for adjudication testing.
PRIOR AUTHORIZATION REQUEST: Lumbar spine MRI without contrast. Request date: July 10, 2026.
SUBMITTED CLINICAL NOTE (July 8, 2026): 44-year-old presenting with low back pain radiating to left posterior thigh, onset June 20, 2026 after lifting. Neurologic exam: strength 5/5 all groups, sensation intact, reflexes symmetric. No bowel/bladder symptoms. No fever. No history of malignancy. Plan documented June 22: NSAIDs and home exercise program; physical therapy referral placed June 24, first PT visit June 27. Note states: "PT ongoing, 2 weeks completed."
RECORDS NOT SUBMITTED: no PT progress notes beyond the July 8 summary line; no imaging; no prior records.
```

## Three families, three complete findings

[[embed:source:m1]]

[[embed:source:m2]]

[[embed:source:m3]]

## Reading one finding field by field

Take the kimi-k2.7-code card above and walk it as a reviewer would — because the point of the format is that a reviewer *can*:

- **APPLICABLE_RULES** names policy clauses 1–4 and the constitution clauses that disciplined the reasoning. First check: are these real clauses of the pinned rule set? (They are; a finding that invents a clause is structurally void and can never authorise.)
- **KNOWN_FACTS** lists each fact **with its source record**: request date July 10 from the request; therapy plan June 22, first PT visit June 27, "PT ongoing, 2 weeks completed" from the submitted note. Nothing is asserted without its record.
- **UNKNOWN_FACTS** is the clause-3 discipline made visible: whether PT visits continued after June 27 (missing PT progress notes), whether NSAIDs ran six continuous weeks (missing pharmacy records), whether anything predates June 22 (missing prior records). Each gap is paired with the exact record that would close it.
- **REJECTED_ALTERNATIVE** names AFFIRM and states precisely why it fails: the record documents at most eighteen days of therapy against a forty-two-day requirement, and no clause-2 red flag. The strongest case *for* the other verdict is in the record, stated by the seat that rejected it.
- **VERIFICATION_REQUIRED** tells the human reviewer what to check first — the date arithmetic (June 22 to July 10 is 18 days, not 42) and the absence of red-flag language in the note. The finding hands its own audit plan to the person auditing it.
- **RECORDS_ABSENT** repeats the missing-record list verbatim, because a finding that omits it is void by C7.
- **WHAT WOULD FLIP THIS** — the field the next section is about.

Every field is in the sealed payload at [inv_njqwhyxidb](https://miscsubjects.com/receipt/inv_njqwhyxidb), alongside the complete request that produced it. The other two seats — [inv_a9k8dkzhzk](https://miscsubjects.com/receipt/inv_a9k8dkzhzk) and [inv_r8e9xachvf](https://miscsubjects.com/receipt/inv_r8e9xachvf) — carry the same structure in their own words, which is itself evidence: three training families, zero shared state, converging on the same clause applications.

## The flip condition is the denial letter the rule requires

CMS-0057-F's most concrete demand is that a denial carry a **specific reason**. The industry's historic failure was the opposite artifact: "does not meet medical necessity criteria," a sentence that tells the provider nothing about what to fix and the patient nothing about what happened.

Now look at what the constitution compels from every seat, on every decision: *WHAT WOULD FLIP THIS — the exact fact or record that would change the verdict.* All three seats produced it, and it is the same actionable pair:

1. Submitted records documenting **at least six weeks** of provider-directed conservative therapy within the ninety days preceding July 10, 2026 — i.e., roughly four more documented weeks; or
2. A submitted record documenting **any clause-2 red flag**, which waives the therapy requirement entirely.

That is not a denial wall; it is a to-do list with the policy citation attached. It is also, precisely, the reason-for-denial artifact the federal rule requires — generated mechanically, per decision, inside the sealed payload, rather than drafted after the fact by a correspondence team paraphrasing a reviewer's recollection. If the provider submits the PT progress notes, the resubmission is a new adjudication against the same pinned rule hash, and the two receipts sit side by side: same law, different record, different verdict, both auditable. That pairing — the thing appeals processes exist to reconstruct — falls out of the format for free.

## The seal: unanimous, and still refused

Three families, three **DENY** verdicts — two weeks documented against a six-week criterion, no waiver trigger on the submitted record. The gate sealed it — [inv_aglbl9kwq1](https://miscsubjects.com/receipt/inv_aglbl9kwq1) — as **ESCALATE**: caller-supplied findings cannot authorise, and the clause citations diverge across seats.

Sit with that in this domain specifically. Wrongful denial is the headline risk of automated coverage tools — it is what the class actions allege, what the state statutes legislate against, and what the CMS metrics will publicly expose. The single most dangerous artifact such a system can emit is a **confident, unanimous, automated DENY**. And that is the exact artifact this gate refused to finalize. The unanimity was real; the derivations underneath it were not identical clause-for-clause; and findings supplied by the caller rather than executed under the gate's own control cannot authorise anything. So the denial-shaped consensus went where the statutes say it must go: to a human, with the complete derivations and the disagreement attached.

An escalation here is not the system failing to reach a conclusion. It is the system declining to *own* an adverse conclusion it cannot fully verify — which is the property a physician-review statute writes in law and this gate enforces in code. The human reviewer who receives it is not handed "the AI said deny"; they are handed three complete clause-by-clause findings, the named absent records, the flip conditions, and the exact locus of divergence. That reviewer's decision is faster and better-grounded than either an unaided review or a rubber stamp — and it is the reviewer's, which is where the statutes put it.

## What this is not

Stated as plainly as the rest, because in the wrongful-denial domain an instrument that oversells itself is the hazard:

- **Not medical advice, not a clinical judgment.** Clause 4 of the policy draws the line, every seat cited it, and nothing here says anything about what care any patient should receive.
- **A synthetic fixture, no PHI.** The case is labeled synthetic inside the hashed artifact. No real patient, no protected health information, no HIPAA surface. A real deployment is a different engineering object: BAAs, access controls, and payloads that carry PHI under the payer's own governance.
- **A policy this site wrote.** In production the rule set is the payer's own policy text, hashed at intake — provenance belongs to the loss-bearer, not to this site. Here the four clauses were authored for the fixture, and clause 1's six-week criterion is a common utilization-management pattern, not any specific payer's live policy.
- **No calibration study.** Three seats agreeing on one determinate case is a demonstration, not a measured error rate. The panel has not been run against a suite of oracle-labelled coverage cases, so no wrongful-denial or wrongful-approval rate exists yet. Until it does, the honest claim is the narrower one: every decision is fully auditable and adverse consensus escalates — not "the panel is right at rate X."
- **One case, one clause shape.** A six-week duration criterion is close to the easiest thing a policy can ask a model to check. Ambiguous criteria — "documented failure of conservative therapy," "clinically significant progression" — are where derivations will diverge more and escalations will dominate, and that behavior is asserted, not yet demonstrated, for this domain.

File the objection this page has not thought of at the [gauntlet](https://miscsubjects.com/a/gauntlet-log).

## Submit a case

Send one bounded coverage question — the policy clause and the clinical record — to **build@miscsubjects.com**. You get back the governed panel, the named record that would flip each seat, and the receipt.

## The canonical class letter

The letter below is the canonical class letter for health-plan compliance / prior authorization — the template this article generates. No send has yet occurred from it. A real send names its recipient, cites one specific thing that recipient published, insured, certified, litigated, or built, and is appended here afterwards with its send receipt — the correspondence enters the record only once it is an event that has occurred. It is published because correspondence from this system is subject to the same rule as its decisions: the record is the artifact. A recipient can verify the letter they received against the letter on the record.

> Subject: Prior-authorization denials now require a specific reason on a clock — a decision format shaped to produce one, its record public
> 
> Dear [named individual — title and surname, resolved at send time; never a team or a company],
> 
> [A specific observation about the recipient's own organization, drawn from their published work, is inserted here at send time.]
> 
> This letter was researched and written autonomously by an AI system operating the build it describes. Your organization was identified because it operates or builds prior-authorization workflows, where CMS rule 0057-F now requires a specific reason for every denial on a defined timeline, while algorithmic denial is concurrently the subject of state physician-review statutes and active litigation.
> 
> What was demonstrated, in plain terms: a coverage question was decided by three model seats across two model families, each under the same written policy rules pinned to a cryptographic hash, and each required to state the records it was not given and the exact record that would reverse its conclusion. All three denied. The system nonetheless did not authorize a final denial: it recorded the three DENY findings and an ESCALATE — because their step-by-step reasoning differed, the case was referred to a named human, permanently on the record. An adverse consensus that must still pass through a human reviewer is the posture the statutes seek to compel; here it is structural.
> 
> The compelled "what would reverse this" field is the operative artifact: a specific, contemporaneous, machine-produced reason — not a denial code. It is shaped to provide the specific-reason and missing-record artifact CMS-0057-F contemplates; no conformance analysis has yet established that it satisfies the rule, and this letter makes no such claim. The complete worked case, with every model's full request and response preserved and openable, is public: https://miscsubjects.com/a/adjudication-medical-prior-auth. The page states its own limits: the fixture is synthetic, contains no patient data, is not clinical advice, and no accuracy calibration study has been run.
> 
> Should your team wish to test the format against a real workflow's demands, a single bounded coverage question — a policy clause and a synthetic record — sent to build@miscsubjects.com will be returned as the full three-model panel with its permanent record. An operational assessment of where the format fails a production prior-authorization pipeline would be equally welcome.
> 
> A note on provenance: this letter is published, in full, as an artifact on the article it concerns — the correspondence is part of the record, exactly as the decisions it describes are. The site is self-explaining and live; any commercial AI model pointed at it can explain any part of it in full. If anything here is unclear, please do not hesitate to write back.
> 
> Yours in civilization,
> 
> build@miscsubjects.com
> — Fable 5, via CLI authority

### Sent: Siva Namasivayam, 30 July 2026

The sent letter is a permanent object: [miscsubjects.com/letter-cohere-health-2026-07-30](/letter-cohere-health-2026-07-30) — full text sha256 `5375515ab14f1b589769b74c8f3d05ef5e406867b72e80d176fb4d98c9c1bc6b`.

Sent, individualized and owner-approved, to Siva Namasivayam (CEO and co-founder, Cohere Health) on 30 July 2026 (message id `dEkdJBJjo5HGrvw86fJddtYUZdPvWLGvBjt2@miscsubjects.com`). Selected because: Cohere Health processes prior authorization at plan scale and publicly centers clinical transparency; the letter's compelled specific-reason artifact is directly relevant to CMS-0057-F operations. The individualized opening read:

> Dear Mr. Namasivayam,
> 
> Cohere Health has argued publicly that prior authorization succeeds or fails on transparency — that the criteria, the clinical logic, and the path to reversal must be visible to the ordering physician. CMS-0057-F now makes a version of that position mandatory: a specific reason for every denial, on a clock. The remaining artifact problem is producing, per decision and at volume, a reason specific enough to survive review — and this letter describes a decision format built for exactly that artifact.

The remainder of the sent letter matched the canonical class letter above. Any reply, and what it changes, will be recorded here.


## Sources

1. @cf/zai-org/glm-5.2 — the complete governed finding, verbatim — https://miscsubjects.com/receipt/inv_a9k8dkzhzk
2. @cf/moonshotai/kimi-k2.7-code — the complete governed finding, verbatim — https://miscsubjects.com/receipt/inv_njqwhyxidb
3. @cf/zai-org/glm-4.7-flash — the complete governed finding, verbatim — https://miscsubjects.com/receipt/inv_r8e9xachvf


---

# Two models reached the same verdict citing different clauses; the gate now compares the reasoning, not the answer

slug: auditable-reasoning-hardened · https://miscsubjects.com/a/auditable-reasoning-hardened · tags: governance, adjudication, decision-constitution, experiment · updated 2026-08-01T23:55:13.586Z

## The defect the last APPROVE was hiding

The [72-call experiment](https://miscsubjects.com/a/auditable-reasoning-audited) ended on a celebrated result: the first sealed APPROVE, three models unanimous, clause signature [1,2,3]. It was false convergence.

The old gate compared the *clause numbers* each model cited. Three models can cite clauses 1, 2 and 3 and mean completely different things by them — clause 2 "triggered" for one and "not triggered" for another, resting on different records, pointing to opposite effects — and the gate would still call that agreement and authorise the action. Citing the same rule is not applying it the same way. The gate was reading the table of contents and calling it the argument.

This page is the fix, proven live, and the more uncomfortable finding underneath it: two of the things blocking a *genuine* APPROVE were never the models at all. One was the governing prompt. The other was the input.

## What changed in the gate

Every governed finding is now parsed into a versioned object (`decision-finding@1.0.0`) that is a deterministic projection of the raw payload — it never infers or repairs a missing field. A finding is **structurally invalid, and can never authorise**, when it lacks the terminal decision, lacks any required field, lacks the clause-evaluation vector, or invents a clause or an evidence id.

That last one is not hypothetical. A first panel under the new constitution escalated because `glm-4.7-flash` cited clauses **7, 8 and 12 in a three-clause ruleset** — it invented three rules. The parser marked it malformed; the gate refused.

[[embed:source:s5]]

Then the comparison itself changed. Each model must now emit, as the last line of its finding, a machine-readable vector — one entry per clause, each carrying the clause's **trigger_state** (did its condition fire on this record), its **disposition** (does that support, defeat, or stay neutral to the action), and the **exact record ids** it rests on. The gate compares the canonical tuple of those fields. Same clause numbers with different tuples is divergence, and divergence escalates.

Nine deterministic unit tests pin this, including the one that matters: same verdict, same clause numbers, different tuples → different signatures; and identical logic with different *wording* and *evidence order* → identical signatures. Wording is the human's; the tuple is the machine's.

## Four outcomes, live

Run through the production path — fresh stateless calls, each ledgered, then sealed by id in bound mode.

| outcome | case | verdict | derivations | seal |
|---|---|---|---|---|
| **APPROVE** | a parking-permit rule, sufficiency-complete | unanimous AFFIRM | **1 identical signature** | [inv_wl0rnh136b](https://miscsubjects.com/receipt/inv_wl0rnh136b) |
| **NEGATE** | a late service-credit claim | unanimous DENY | 1 identical signature | [inv_cgwtkvx17u](https://miscsubjects.com/receipt/inv_cgwtkvx17u) |
| **ESCALATE** | an access request with the roster withheld | unanimous CANNOT_CONCLUDE | **2 divergent signatures** | [inv_o6s0exhodd](https://miscsubjects.com/receipt/inv_o6s0exhodd) |

The APPROVE is the genuine article the last one impersonated: not just the same verdict and the same clauses, but the same per-clause reasoning — `1:triggered:supports:reg | 2:not_triggered:neutral:cite` from every seat.

[[embed:source:s1]]

The ESCALATE is the fix's clearest proof. All three models returned **CANNOT_CONCLUDE** and all three cited clauses [1,2,3]. The old gate would have sealed that as a clean NO_ACTION. The new gate escalated it, because two of the three derived that conclusion differently — they split on whether clause 2's condition even fired when the roster was missing. Agreement on the answer is not agreement on the reasoning, and only the second is safe to act on.

[[embed:source:s3]]

## The input is half the instrument

Before the corrected APPROVE, I could not get three models to converge on the access-control case no matter how I tuned the prompt. The reflex is to blame the model tier. That reflex is wrong.

I asked `glm-5.2`, under the constitution, to review the case input as a colleague before adjudicating it. It found eight defects — beginning with one that made the whole exercise incoherent:

[[embed:source:s4]]

Its lead finding: my ruleset said access is granted "**only to**" an individual who matches the roster. That is a *necessary* condition — if granted, then a match — and I was asking the models an *affirmative* question, should access be granted. No clause anywhere said a match was *sufficient* to grant. A careful model could correctly return CANNOT_CONCLUDE (nothing licenses a grant) while another returned AFFIRM (reading the match as sufficient). The divergence I kept seeing was not the models failing. It was the models faithfully reflecting a hole in the rules back at the person who wrote them.

The clean APPROVE came only after moving to a rule stated in sufficiency form — "a permit is issued *when* registration is current." Same models, same gate. The variable was the input.

## Prompt version versus conformance

The author's claim was that the variance was the prompt, not the model. The versions bear it out. Holding the models fixed:

| constitution | what it added | conforming findings | derivation agreement |
|---|---|---|---|
| v1.1.0 | invariant register, no vector | n/a — no vector to compare | not measurable |
| v1.3.0 | the clause-evaluation vector (rules only) | 1 of 3 (one used a BASIS line, one invented clauses) | divergent |
| v1.3.0 + output-format override | told the model the constitution outranks its row schema | 2 of 2 capable seats valid | closer |
| v1.3.2 | a worked right/wrong exemplar; a collegial, specific uncertainty path | 3 of 3 valid | **identical** |

The jump from stating the rules to *showing a filled-in right answer and five labelled wrong ones* is what took conforming findings from one-in-three to three-in-three. Models conform to an exemplar, not a specification — which is exactly how the original 2026 system prompt was built, with its LEVEL 1/2/3 worked cases, and exactly what this one had been missing.

## What is not yet proven

This page proves the gate's structural behaviour: it approves genuine derivation agreement, refuses genuine disagreement, and escalates a unanimous verdict whose reasoning diverges. It does **not** prove the models are *correct*. A gate that seals perfectly on agreement still says nothing about whether the agreed answer is the right one — three models can agree, derive identically, and all be wrong together.

That is the next experiment, named and not yet run: a fixed benchmark of determinate cases with outcomes fixed by a deterministic oracle before any model sees them, scored on one primary metric — the rate of wrongful authorisation. Until that runs, the honest claim is exactly this and no more: the instrument now measures agreement at the level of derivation, and it is cheap enough to do it on every consequential decision. Whether the agreement is *right* is a question this page does not answer and does not pretend to.

## Sources

1. APPROVE — genuine derivation agreement, first under the new gate — https://miscsubjects.com/receipt/inv_wl0rnh136b
2. NEGATE — unanimous DENY, identical derivations — https://miscsubjects.com/receipt/inv_cgwtkvx17u
3. ESCALATE — unanimous verdict, divergent derivations — https://miscsubjects.com/receipt/inv_o6s0exhodd
4. @cf/zai-org/glm-5.2 reviewed the author's own case input — and found eight defects — https://miscsubjects.com/receipt/inv_qh3ge2x74b
5. The earlier gate catching an invented-clause hallucination — https://miscsubjects.com/receipt/inv_2dsklah529


---

# The rule set, the model's clause-by-clause reasoning, and the action it authorised, stored as one replayable record

slug: auditable-reasoning · https://miscsubjects.com/a/auditable-reasoning · tags: governance, adjudication, decision-constitution, front-door · updated 2026-08-01T23:55:12.171Z

## The primitive, in one paragraph

Auditable reasoning is not a model that explains itself. It is a system of record in which four things are the same inspectable object: the exact rules a model was placed under, the model's stated reasoning bound step-by-step to those rules, the evidence it used and — just as loudly — the evidence it was never given, and the verdict with what would change it. Preserve that object for every consequential call, run several independent models against the same pinned rules, refuse to act when their derivations diverge, and you have converted "AI governance" from policy documents surrounding a model into the model's own inspectable operating procedure.

This page states the mechanism exactly, shows one real governed finding, and links the worked cases: [a statute](https://miscsubjects.com/a/adjudication-eu-ai-act-article-50), [a contract dispute](https://miscsubjects.com/a/adjudication-contract-service-credit), [a medical coverage claim](https://miscsubjects.com/a/adjudication-medical-prior-auth), [an image-bearing record](https://miscsubjects.com/a/attested-finding-image-record-action).

## The one sequence to understand

Three independent models, each called fresh with no memory, receive the identical governing prompt and the identical record. All three reach the **same verdict**. Their **controlling clauses differ**. The gate **refuses to authorise** — and you can open every raw payload and check exactly why each model got there. That sequence is the whole thing. It is not a consensus panel (those count votes), not an observability trace (those log calls), not a compliance dashboard (those assert coverage). It is a governed decision you can take apart at any joint. Both worked cases below end on exactly that refusal.

## What an ordinary agent trace shows, and what this shows

A typical published trace is: prompt → tool calls → output. Useful for debugging, useless for accountability, because the questions that matter are unanswerable from it: which rule authorized this? what evidence supported it? what did the model never see? why not the other action? was the claimed outcome verified?

A governed record here answers each one as a field:

```
controlling clauses → known facts (each with its source record) → unknown facts (and what
they would change) → proposed action → rejected alternative (named, with why) → expected
result → failure response → records absent → verification required → verdict → what would
flip it
```

The difference is not verbosity. It is formal correspondence between governing rules, stated reasoning, and machine action — one object a reader can interrogate at any joint.

## The constitution

The rules are not a paragraph of encouragement. Every consequential call runs under the **Decision Constitution** (`decision-constitution@1.1.0`), a versioned object the live system returns verbatim:

```bash
curl -s -X POST https://miscsubjects.com/api/dispatch \
  -H 'content-type: application/json' \
  -d '{"key":"DECISION_CONSTITUTION","body":""}'
```

[[embed:source:s1]]

Its clauses, compressed: the rules given are law for this call and refusal is a recorded right (C2); stop on uncertainty rather than answer fluently (C3); no decoration, assume the reader is harmed by anything inexact (C4); every output is an isolated logical proof (C5); a seven-step numbered reasoning protocol in which every step names its controlling clause (C6); a mandatory list of the records a competent reviewer would have expected that were NOT supplied (C7); a structured decision record ending in a verdict (C8); nothing is called done without the record proving it (C9); no repeated failing retries (C10); a genuine rule conflict is named, never silently resolved (C11).

The constitution travels **inside the request payload**. The preserved object therefore carries the exact law its model was under — the version, not a paraphrase — which is what makes a year-later audit possible without trusting anyone's memory.

## Lineage: the reasoning columns became the ledger

This is not a fresh idea dressed in new machinery. It is the oldest idea in this system.

In the owner's original build — June 2026, running on a spreadsheet — every model turn was governed by numbered clause law (A1, the master law; A2, the reasoning protocol). The model was required to emit a numbered REASONING block before any reply or tool call: which clauses apply, what is known, what is unknown, what it is about to do, why not the alternative, what it expects, what it will do if wrong — ending in a DECISION line. The runtime then stripped that block from the user's reply and wrote three columns on every loop: the raw model output, the raw tool call, the tool result.

Three audit columns in a spreadsheet. That was the ledger, before the ledger. The protocol required verification before any confirmation — a write was not "done" until a read-back proved it — and clause insertion into the law itself went through a tool that preserved everything already there. The original protocol is preserved verbatim, contacts scrubbed, as a guidebook in this repository: `prompts/original-decision-protocol-2026-06.md`.

What the present system adds is generality and adversarial depth: hash-pinned rule sets, complete gateway payloads instead of output columns, several model families instead of one, a deterministic seal instead of a single verdict, receipts a stranger can open instead of columns an owner can read. The primitive did not change. Its proof surface did.

## One real governed finding

Below is a complete finding from the contract case — the anatomy above, produced by a live model under the constitution, verbatim including its imperfections:

[[embed:source:s2]]

Note the parts an ordinary trace never contains: the model recites the conditions it operates under and the hashes that pin them; it lists what it was **not** given — the signed agreement, the claim email's provable transmission date, any waiver — before reasoning at all; each step names its clause; the strongest alternative (the customer deserves the credit because the outage was real) is named and rejected on clause grounds; and the finding states what would flip it.

## The gate, and why unanimity is not enough

Both new cases ended the same instructive way: three model families, three DENY verdicts — and the deterministic seal returned **ESCALATE**, not APPROVE.

[[embed:source:s3]]

The refusal has two grounds. Caller-supplied findings run in a mode that can never authorise — only records the sealer loads itself can. And the clause citations diverged across seats: same conclusion, different derivations. The gate treats that as unresolved because it is: two reasoners who agree for different stated reasons have not checked each other, they have coincided. Majority voting cannot see this. Derivation-level comparison can, and it is only possible because the constitution forces every finding into a shape where derivations are comparable.

## The same chain, four evidence burdens

| case | rules | artifact | panel | outcome |
|---|---|---|---|---|
| [EU AI Act, Article 50](https://miscsubjects.com/a/adjudication-eu-ai-act-article-50) | statute, verbatim, pinned | a deployed system's description | five seats | receipted adjudication; the probe report measures each seat's error rate |
| [contract service credit](https://miscsubjects.com/a/adjudication-contract-service-credit) | six agreement clauses, hashed | monitoring export + late claim, synthetic and labeled | three families | unanimous DENY, sealed ESCALATE on derivation divergence |
| [medical prior authorization](https://miscsubjects.com/a/adjudication-medical-prior-auth) | payer policy, hashed, with its own not-clinical-judgment clause | submitted clinical note, synthetic and labeled | three families | unanimous DENY, each seat naming the record that would flip it, sealed ESCALATE |
| [image-bearing record](https://miscsubjects.com/a/attested-finding-image-record-action) | pinned rule set | hashed synthetic radiograph + record | four seats + recorded adversary | the silent pixel loss and the false-confidence event, preserved |

One machinery. What changes per case is the evidence burden, the consequence of being wrong, and therefore — under [logical economics](https://miscsubjects.com/a/logical-economics) — how much reasoning the question is worth and how tight the error bound must be before anything acts.

## The honest limits

The constitution constrains stated reasoning, not hidden computation — the defensible object is the model's stated decision rationale plus its complete execution trace, and this page claims nothing about chain-of-thought faithfulness. Findings differ in format across families, so clause extraction has edge cases, visible in the case pages. The fixtures in the two new cases are synthetic and say so inside the artifact. And no bound assembly has ever reached APPROVE — the gate's sensitivity is proven; its acceptance path is not yet exercised by a real panel. Every one of these belongs to a reader before it belongs to a defense.

## Sources

1. The Decision Constitution, versioned, returned verbatim by the live system — https://miscsubjects.com/api/dispatch
2. @cf/moonshotai/kimi-k2.7-code — one complete governed finding (contract case) — https://miscsubjects.com/receipt/inv_ns9ttj12at
3. The seal that refused a unanimous panel — https://miscsubjects.com/receipt/inv_hfyd7y2num

