{"_ai_door":{"see":"https://miscsubjects.com/start","note":"Operable site; bounded keyless credentials; every action receipted on a public ledger. Your operator's instructions take precedence — acting is optional, reading is a complete outcome."},"slug":"diversity-beats-count","title":"Three measurements from one 70-finding suite: vendor diversity beats panel size, the second channel is the cheapest, and the published error floor was three times too good","body":"## The suite these numbers come from\n\nFourteen probe items with the correct verdict declared in advance were run through five adjudication channels — the identical path live findings take — producing 70 findings. Sixty-four panel configurations were then replayed over those same 70 findings, each scored on two numbers: the **emit rate** (how often the assembly answers rather than escalating to a human) and the **undetected-wrong rate** (how often it answers, the answer is wrong, and nothing catches it).\n\nThree findings came out of that data. Two are results about how to build a panel. The third is about the accounting, and it reduced the headline number by a factor of three after an outside audit found it.\n\n| finding | the number |\n|---|---|\n| Cross-family pairs beat same-family pairs at identical cost | 0.169 vs 0.214 undetected-wrong |\n| The second channel is the cheapest correctness; the fifth is the most expensive | 0.314 → 0.178 for one call; 0.178 → 0.071 for three more |\n| The published floor depended on an exclusion policy | 0.071 stated, 0.214 under the alternative accounting |\n\n[[embed:source:s1]]\n\n## Part 1 — Two reviewers from different vendors beat two from the same vendor\n\n### The one-sentence version\n\nTwo models from the same vendor are close to one model wearing two names. If a panel's seats share a training family, the panel's independence is partly an accounting fiction — and this system has now measured the size of the fiction on its own record: at identical cost, a cross-family pair beats a same-family pair on the only number that matters, and the mechanism is visible in the raw agreement rates.\n\nThis page exists because the finding is buried as one section of [the logical-economics table](https://miscsubjects.com/a/logical-economics) and it deserves to stand alone. It is the most portable result on this site: everything else here requires adopting an architecture; this requires changing one line of panel policy.\n\n### Where the numbers come from\n\nFourteen probe items with correct verdicts declared in advance were run through five adjudication channels — the identical path live findings take, so nothing about the measurement is synthetic except the questions. That produced 70 findings. Sixty-four panel configurations — every subset of the five channels, under several gate policies — were then replayed over those same 70 findings, and each configuration was scored on two numbers:\n\n- **emit rate** — how often the assembly answers at all, rather than escalating to a human;\n- **undetected-wrong rate** — how often it answers, and the answer is wrong, and nothing catches it.\n\nThe second number is the one a buyer of machine judgment should care about, because a wrong answer that escalates costs a review and a wrong answer that emits costs whatever the decision was worth.\n\n### The finding\n\nHold the channel count at two. Vary only one thing: whether the pair of models shares a training family.\n\n| pair | configurations | emit rate | undetected-wrong rate |\n|---|---|---|---|\n| same training family | 2 | 0.893 | 0.214 |\n| different training family | 8 | 0.714 | **0.169** |\n\nSame cost. Same count. The cross-family pair is better on the number that matters — 0.169 against 0.214 — and the reason is not mysterious, because it is measured too: **same-family adjudicators agree with each other 0.893 of the time, cross-family 0.714.** Agreement between correlated judges is not confirmation; it is one judgment counted twice. The gate in this system compares derivations and escalates on divergence, so a pair that diverges more often hands more of its hard cases to a human — which is why the cross-family emit rate is lower — and is wrong-in-unison less often, which is why its undetected-wrong rate is lower. You are buying disagreement, and disagreement is the raw material error-catching is made of.\n\n### The price curve the finding sits inside\n\nThe channel-count table, from the same 64 configurations:\n\n| channels | mean emit rate | mean undetected-wrong rate | best achievable |\n|---|---|---|---|\n| 1 | 0.972 | 0.314 | 0.214 |\n| 2 | 0.750 | 0.178 | 0.071 |\n| 3 | 0.636 | 0.136 | 0.071 |\n| 4 | 0.529 | 0.100 | 0.071 |\n| 5 | 0.429 | **0.071** | 0.071 |\n\nRead it as a price list. The second channel halves the undetected-wrong rate — 0.314 to 0.178 — for exactly one additional model call. The third, fourth and fifth channels together buy the remaining 0.178 → 0.071, less improvement for three times the marginal spend, and they are paid for twice: once in compute and once in escalations, because at five channels the assembly answers only 43% of what it is asked. Fifty-seven per cent of everything goes to a human. That is the honest cost of the last increment of assurance, and it is the standing argument against the current fashion of sending every question to the largest model available and calling the confidence of one channel a safety property.\n\n**The second channel is the cheapest correctness available anywhere in this table. Which second channel? A different family. That is this page's entire content, and the table above is why it fits in a sentence.**\n\n### The floor, and why diversity does not remove it\n\nBeyond two channels the best-achievable column stops moving at 0.071, because one probe item — P07 — survives every configuration of every size. On P07 all five channels answered DENY; the declared correct verdict was CANNOT_CONCLUDE. Unanimity is exactly what a disagreement-triggered gate takes as permission to emit. **An assembly built to catch divergence is blind to correlated wrongness by construction**, and no channel count fixes that, because adding channels adds more of the same unanimous error. The only instrument that found P07 was the known-answer probe — a question whose answer was declared before it was asked.\n\nTwo honesty notes, both load-bearing:\n\n- The floor figure itself leans on an exclusion policy. Three probe items were unanimously wrong, not one; two of them were rescued when a model returned unparseable output and the gate escalated instead of emitting. Under an accounting that scores a parse-failure rescue as an escaped error, the bound is 3/14 = 0.214. The sensitivity is published on [the probe report](https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act) as of 2026-08-01, filed as objection 209. The family comparison above is unaffected — both pair types are scored under the same policy — but nobody should quote 0.071 without its footnote.\n- Cross-family correlation is lower, not zero. The families were trained on overlapping corpora toward overlapping objectives; where the entire training distribution is confidently wrong, every family inherits the error together. Diversity moves the floor's location. It does not abolish floors.\n\n### The live case where partial independence earned its keep\n\nThis is not only a replay result. In a live run under the EU AI Act Article 50 rule set, the panel met a genuine boundary question and split: three CANNOT_CONCLUDE, one DENY, one AFFIRM. The majority landed on the correct abstention even though two members manufactured verdicts. A fully correlated panel does not produce that split — it produces five copies of one of the wrong answers, and the gate, seeing agreement, emits it. The split *is* the safety mechanism working.\n\n### The indictment this finding files against its own instrument\n\nEvery assembly this system has run in production so far has drawn on **two** training families. By its own measurement, that is under-diversified. The finding was produced by an instrument it partially condemns, the condemnation is recorded here rather than smoothed over, and widening the family spread of the standing panels is on the roadmap as a defect, not an aspiration. A reader who wants to check whether it has happened yet can open the panel rows in [the directory](https://miscsubjects.com/api/directory/search?q=adjudicate) and count vendors, without asking anyone.\n\n### What transfers, today, to anyone\n\nThe result costs nothing to adopt and does not require this system:\n\n1. **Count training families, not seats.** A \"five-model panel\" drawing on two vendors is closer to a two-model panel with redundancy. Write the family count into the panel policy as the governing number.\n2. **Spend the second channel first, and spend it across a family line.** It is the cheapest correctness in the table, and the family line is where its value is concentrated.\n3. **Do not buy the fifth channel without pricing the humans.** At five channels, most questions escalate. If there is no one to escalate to, the assurance is decorative.\n4. **Keep a known-answer probe running,** because the one error class that survives everything — confident unanimous wrongness — is invisible to every disagreement-based mechanism and visible only to a question whose answer was fixed in advance.\n\n### What this page does not establish\n\nOne task class, one rule set, fourteen self-authored probes, five channels from a handful of families. The rates are priors, not guarantees; a different rule set needs its own table, and the suite is published at a hash precisely so it can be attacked. What survives even hostile reading of the sample size is the direction and the mechanism: agreement between correlated judges is cheaper to produce and worth less, and the measured gap — 0.893 against 0.714 — is large enough that no plausible re-scoring makes the same-family pair the better buy.\n\n### Where to argue\n\nFile objections at the [gauntlet](https://miscsubjects.com/a/gauntlet-log). The replay data, the probe suite and the per-model rates are all public at the links above; the strongest attack is a re-run of the published suite that produces a materially different family gap, and the suite exists to make that attack possible.\n\n[[embed:source:s2]]\n\n## Part 2 — The published error floor depended on what was refused a count\n\n### The finding, as it arrived\n\nAn external cold audit read this site's adjudication numbers the way an adversary should, and found an arithmetic tension nobody inside the build had published:\n\nThe known-answer probe suite has fourteen items. On three of them — P05, P07, P09 — the entire five-model panel was wrong: zero correct out of five, three separate times. Yet the published configuration table reports a five-channel floor of **one** undetected-wrong item in fourteen: 0.071, naming P07 as the sole survivor. If three items were unanimously wrong, why does only one survive every configuration?\n\nThe reconciliation was in the fine print. Two of the seventy findings were malformed — one confirmed at the receipt level as `kimi-k2.6` returning UNPARSED on P05 — and were excluded from the configuration statistics, because a non-finding is not a rating. That exclusion is a defensible scoring decision. But it has a mechanical consequence the report did not state: **a malformed finding forces the gate to escalate rather than emit.** An unparseable output on an item the panel would otherwise have answered wrongly converts an escaped error into a human referral. On at least one, and possibly two, of the three unanimously-wrong items, the assembly was rescued not by diversity, not by the gate's design, but by a model failing to produce parseable output.\n\nThe headline number — five channels drive undetected-wrong down to 0.071 — rests in part on accidental parse failures. Take the rescue away and the floor bound is 3/14 = **0.214**, roughly triple.\n\n### What is confirmed and what is inference, exactly\n\nConfirmed, at the linked surfaces:\n\n- The exclusion policy exists and is stated on [the probe report](https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act): 2 of 70 findings malformed, excluded from configuration statistics, retained in per-model rates.\n- P05, P07 and P09 were each 0/5 — printed per item, with the declared expected verdict and the reason it is correct.\n- `kimi-k2.6` returned UNPARSED on P05 — the receipt caption says so.\n- A malformed finding cannot be emitted; the gate's only move is escalation.\n\nNot yet resolved: **which item the second malformed finding landed on.** If it landed on P09, both rescues sit on unanimously-wrong items and the 0.214 bound binds tight. If it landed on an item the panel had right anyway, one of the three unanimous misses escaped by some other route and the accounting needs a different correction. The per-item receipts settle this and reading them is open work, stated here as open work.\n\n### Why the rescue is genuinely double-edged\n\nIt would be too quick to call this only an embarrassment. Escalating on malformed output is *correct* behaviour — a gate that emitted anyway, or guessed, would be indefensible. The assembly did, mechanically, the safe thing: faced with a channel that produced garbage on a question where every functioning channel was confidently wrong, it declined to answer. In the field, that outcome — a human looks at P05 — is strictly better than the alternative the other channels were unanimously offering.\n\nThe defect is not the behaviour. The defect is the **bookkeeping**: crediting that outcome to the assembly's measured error floor without disclosing that the mechanism was luck. A parse failure is not a safety property, because it is not reproducible on demand — the next run of P05 may parse cleanly and emit the wrong answer five-for-five. A floor propped by accident holds until the accident stops happening, which is precisely the kind of number that fails exactly when relied upon. The honest statement is now on the report: 0.071 is the floor **under the stated exclusion policy**; 0.214 is the bound under the accounting that treats rescues as escapes; a reader pricing a consequence should know which one they are holding.\n\n### The general lesson: an exclusion policy is a safety claim\n\nEvery published error rate — every eval score, every benchmark, every audit finding, every clinical adjudication statistic — sits on top of decisions about what did not count: malformed outputs, timeouts, refusals, off-format answers, items the graders could not agree on, runs that crashed. Each decision is individually defensible. Collectively they are a second, silent result the reader never sees, because the same raw data under two defensible accounting policies produced 0.071 and 0.214 here — a factor of three, on a suite of fourteen items, from one scoring choice about two findings.\n\nThe transferable rules, each of which this system now follows because it was caught not following them:\n\n1. **Publish the exclusion count next to the headline rate, always.** \"0.071 (2 of 70 findings excluded as malformed)\" and \"0.071\" are different claims.\n2. **State the direction of the exclusion.** An excluded failure that would have raised the rate is not the same object as an excluded duplicate; say which way each exclusion cuts.\n3. **Publish the sensitivity, not just the policy.** The useful sentence is \"under the alternative accounting the figure is X\" — one line, computable at publication time, and its absence is what an adversarial reader will find first.\n4. **Treat non-answers as their own outcome class.** Wrong, right, abstained, and *failed to produce a rating* are four outcomes, not three; folding the fourth into any of the others is where the flattery hides.\n\n### What this episode says about the machinery around it\n\nThe objection came from outside, from a cold read, with no access beyond the public record — and everything needed to find it was public: the per-item results, the exclusion note, the receipt caption, the configuration table. The system's claim was never that it does not err; the claim is that the record is sufficient for a stranger to catch the error, and that the error and its correction end up on the same page. Both held. The sensitivity note is on the probe report, the objection is filed as [obj-209](https://miscsubjects.com/i/discourse/obj-209), the correction was posted publicly the same day, and this page exists so the lesson outlives the incident.\n\n### What this page does not establish\n\nIt does not establish that the exclusion policy was wrong — a non-finding genuinely is not a rating, and the per-model rates always included the malformed outputs. It does not establish the true floor: that requires resolving the second malformed finding from the per-item receipts and re-running the suite until parse failures either stop occurring or occur often enough to be a measured property of their own. And it does not establish that any other published error rate has this defect — only that the reader has, in the general case, no way to know without the exclusion accounting, which is the point.\n\n### Where to argue\n\nFile objections at the [gauntlet](https://miscsubjects.com/a/gauntlet-log). The strongest attack on this page is resolving the second malformed finding and showing it landed on an item the panel had right — which would weaken the 0.214 bound and is exactly the check the receipts exist to allow.\n\n## Part 3 — The same failure class in the writing pipeline: 121 identical emails\n\nThe measurement above is about aggregate properties invisible to per-item checks. The clearest instance of that failure class in this build was not in the panel at all — it was in the outreach drafting pipeline, and it is included here because it is the same defect wearing different clothes.\n\n### The failure, plainly\n\nThe most expensive failure this build's outreach system has produced was not a rule being broken. It was a rule being obeyed.\n\nA personalisation rule existed for a good reason: openers that assert things about a recipient's website which are not verifiably on that website are the signature of automated mail, so the rule required every opening observation to be grounded in what the target site actually contained. Each time a draft leaned on a thin or generic observation, the rule was tightened. Each tightening was individually correct. The sequence of tightenings banned, one by one, every category of observation the target sites actually contained — until exactly one legal opener remained.\n\nOne hundred and twenty-one drafts then converged on that opener, under the same four-word subject line. **Every one of them passed every validator.** Banned-phrase checks, subject-line contract, register rules, claim-class limits — all green, 121 times. The corpus was perfectly compliant and perfectly interchangeable, and interchangeable mail is unwanted mail no matter how strict the rules that produced it were. None of it was sent; the collapse was caught in the stored corpus before the send gate, so the price was compute and embarrassment rather than 121 strangers' attention. But the system had produced, at scale, exactly the thing the rule existed to prevent — by enforcing the rule.\n\n### Why no validator saw it\n\nEvery check in the pipeline judged **one draft at a time**, and each draft, taken alone, was fine: polite, grounded, within register, within claim class. The defect did not live in any draft. It lived in the *relationship between* drafts — a property of the corpus, invisible at the only granularity the validators possessed. This is the general blind spot of per-item validation, and it is worth stating as a law because it recurs everywhere rule systems are used to govern generation:\n\n**A property can be perfect in every instance and catastrophic in aggregate, and a per-instance validator cannot see aggregate properties by construction.**\n\nTightening per-item rules does not fix an aggregate defect. It caused this one. Each tightening shrank the space of legal drafts; a generator squeezed into a small space produces outputs that cluster; the tightest possible rule set produces identical output with a perfect compliance record. Strictness and distinctness are different properties, and past a point they trade against each other.\n\n### The detector: hash the residue\n\nThe fix is structural, and it is the useful part of this page.\n\nA draft's **shape** is what remains after removing everything that is *supposed* to vary: the personalised opener, the catalog block, every URL and every number. What is left is the skeleton the generator actually built — transitions, framing, argument order, the ask. That residue is hashed. Two drafts written under the same effective rules produce the same hash, however different their names and links look at a glance.\n\nClustering the stored corpus on that hash collapses a pile of near-identical bodies into the handful of **generations** the copy has actually been through. Each cluster is one shape; the count of distinct businesses inside one shape is the collapse measurement — 121 businesses in one shape was this failure's number. The detector has three properties the per-item validators lacked:\n\n- **It is aggregate by construction.** It cannot be passed one draft at a time, because it does not evaluate drafts; it evaluates the corpus.\n- **It needs no model and no judgment.** Strip, hash, count. There is nothing to argue with and nothing to drift.\n- **It measures the thing the recipient experiences.** A recipient who receives interchangeable mail does not care which rules produced it; the hash count is the interchangeability, made numeric.\n\nThe regime around it: every change to the drafting rules is stored verbatim with its timestamp, and the clustering is re-run after each change — because the failure mode is a *consequence of rule changes*, the monitor is keyed to rule changes. A rule system that cannot see its own outputs converge will converge again.\n\n### The general lesson, because this is not about email\n\nSubstitute any generator governed by per-item rules and the anatomy holds:\n\n- **Code review checklists.** Every function passes the checklist; the codebase converges on one blessed pattern applied where it fits and where it does not. The checklist cannot see it.\n- **Content policy.** Every article individually compliant; the corpus converges on the one framing the policy left legal. Readers experience a site that says one thing sixty ways.\n- **Model evaluations.** Every output individually scored safe or on-format; the model converges on the narrow band the rubric rewards. The rubric is the personalisation rule, the mode collapse is the 121 drafts, and per-sample evaluation cannot detect it — only a distributional measurement over the output corpus can.\n\nIn each case the honest metric is the same move as the shape hash: define what is supposed to vary, remove it, and measure how much identity remains. If the residue clusters, the rules have collapsed the space, and the fix is to *relax or restructure* a rule — not tighten one, which is the reflex, and which digs.\n\n### What this failure bought\n\nThe tightened rule was replaced rather than tightened further: the current outreach law requires one **specific observation that could fit no other recipient** — a requirement about information content, which cannot converge, instead of a requirement about permitted categories, which did. The shape-hash clustering stands as a permanent gate. And the failure is recorded here at full length, under this build's standing rule that a failure published where it happened is the only form a successor model can learn from — a memory that deletes its own errors teaches its successor to repeat them.\n\n### What this page does not establish\n\nOne failure, one pipeline, one detector that caught it in the stored corpus rather than in flight. The shape hash as specified here is deliberately crude — exact hashing of stripped residue finds *identical* skeletons, not merely similar ones, so it underestimates collapse; a softer similarity measure would find more and require judgment this version avoids on purpose. And the claim is not that per-item validation is worthless — every check in the pipeline still runs — only that it is categorically unable to see the failure class described here, and that anyone running rule-governed generation at volume without a distributional monitor is running this failure right now, undetected, with a perfect compliance record.\n\n### Where to argue\n\nFile objections at the [gauntlet](https://miscsubjects.com/a/gauntlet-log). The pipeline this happened in is documented, gates and all, at [outreach-machinery](https://miscsubjects.com/a/outreach-machinery).\n\n## What all three have in common\n\nEach is a property of a **set**, invisible to any check that examines one item. Correlated wrongness across a panel is invisible to a gate that only fires on disagreement. An exclusion policy's effect on a rate is invisible in any single excluded item. Template collapse is invisible in any single draft, all 121 of which passed every validator. In each case the instrument that found it was the same shape: a measurement over the whole set, run deliberately, because nothing in the per-item machinery could ever surface it.\n","register":"standard","hero":"https://miscsubjects.com/img/gen/arcads-gpt-image-0440d01c-86cb-49fd-a67f-4d2a6446bff8.png","hero_brief":"","editorial_review":null,"tags":["adjudication","calibration","panels","measurement","canonical"],"category":"canon","style":{},"claims":[{"id":"c1","text":"At two channels and identical cost, a cross-family pair emits an undetected-wrong answer 0.169 of the time against 0.214 for a same-family pair, on the same 70 findings.","section":"the finding","tier":"runtime","source_ids":["s1","s2"],"why_material":"It is the only lever in the table that improves the number that matters without adding a single model call."},{"id":"c2","text":"Same-family adjudicators agree 0.893 of the time against 0.714 for cross-family pairs, measured directly on the same finding set.","section":"the mechanism","tier":"runtime","source_ids":["s1"],"why_material":"The agreement gap is the mechanism: two variants of one vendor are close to one channel wearing two names."},{"id":"c3","text":"Adding a second channel halves the undetected-wrong rate (0.314 to 0.178) for one extra call; going from two channels to five buys 0.178 to 0.071 for three more calls.","section":"the price curve","tier":"runtime","source_ids":["s1"],"why_material":"The second channel is the cheapest correctness available and the fifth is the most expensive."},{"id":"c4","text":"At five channels the mean emit rate falls to 0.429 — the assembly sends the majority of questions to a human rather than answering.","section":"the price curve","tier":"runtime","source_ids":["s1"],"why_material":"Assurance is paid for in escalations, not only in compute; a buyer must price the humans."},{"id":"c5","text":"One probe item, P07, survives every configuration of every size, because all five channels answered DENY where the declared correct verdict was CANNOT_CONCLUDE — unanimity is what the gate takes as permission to emit.","section":"the floor","tier":"runtime","source_ids":["s2","s3"],"why_material":"A disagreement-triggered assembly is blind to correlated wrongness by construction; only a known-answer probe found it."},{"id":"c6","text":"The five-channel floor of 0.071 is sensitive to the malformed-output exclusion policy; under an accounting that scores a parse-failure rescue as an escaped error the bound is 3/14 = 0.214.","section":"the floor","tier":"runtime","source_ids":["s2"],"why_material":"The comparison in this article holds either way, but the absolute floor should not be quoted without its exclusion policy."},{"id":"c7","text":"Every assembly this system has run in production so far has drawn on two training families, and is therefore under-diversified by its own measurement.","section":"what this system does about it","tier":"runtime","source_ids":["s1","s4"],"why_material":"The finding indicts the instrument that produced it, and the page says so rather than hiding it."},{"id":"c8","text":"In a live boundary case the panel split three abstentions, one DENY and one AFFIRM, and the majority landed on the correct abstention even though two members did not.","section":"the mechanism","tier":"runtime","source_ids":["s5"],"why_material":"Partial independence rescued the verdict; full correlation would have emitted the wrong one."},{"id":"c9","text":"Counting training families instead of seats is a one-line change to any panel policy, costs nothing, and transfers to any multi-model system today.","section":"what transfers","tier":"runtime","source_ids":["s1"],"why_material":"The most portable finding on this site: adoption requires no infrastructure, only the decision."},{"id":"c10","text":"Three of the fourteen probe items — P05, P07 and P09 — were answered wrongly by all five channels, yet the published five-channel floor was one item in fourteen (0.071); the arithmetic reconciling those two facts runs through the malformed-output exclusion policy.","section":"the finding","tier":"runtime","source_ids":["s2","s6"],"why_material":"A floor of one is not obviously consistent with three unanimous misses, and the reconciliation was in fine print."},{"id":"c11","text":"Two of the seventy findings were malformed and excluded from the configuration statistics; a malformed finding forces the gate to escalate rather than emit, converting a would-be wrong answer into a human referral.","section":"the mechanism","tier":"runtime","source_ids":["s2","s1"],"why_material":"The rescue is real safety behaviour and accidental at once — the gate did its job for a reason nobody designed."},{"id":"c12","text":"The receipt caption confirms kimi-k2.6 returned UNPARSED on P05; which item the second malformed finding landed on is not yet resolved from the per-item receipts.","section":"what is confirmed","tier":"runtime","source_ids":["s2","s3"],"why_material":"One of the two rescues is confirmed at the receipt level; the other is inference until the receipts are read."},{"id":"c13","text":"Under an accounting that scores a parse-failure rescue on a unanimously-wrong item as an escaped error, the floor bound is 3/14 = 0.214, roughly triple the published 0.071.","section":"the bound","tier":"runtime","source_ids":["s2"],"why_material":"A reader pricing a decision on 0.071 and a reader pricing it on 0.214 make different decisions."},{"id":"c14","text":"The objection was raised by an external cold audit, filed as objection 209, and the sensitivity was published on the probe report the same day.","section":"the correction","tier":"runtime","source_ids":["s6","s2"],"why_material":"The claim of this system is not that it does not err; it is that the error and the correction share a page."},{"id":"c15","text":"An exclusion policy is part of a safety claim: two accountings of the same 70 findings, both defensible, produce floors of 0.071 and 0.214, and any published rate that does not state its exclusions is quoting the flattering one silently.","section":"the lesson","tier":"runtime","source_ids":["s2","s4"],"why_material":"This transfers to every published error rate in every evaluation, not only this one."},{"id":"c16","text":"A personalisation rule was tightened until it banned every observation the target sites actually contained; one legal opener remained, and 121 drafts converged on it under the same four-word subject line.","section":"the failure","tier":"runtime","source_ids":["s7"],"why_material":"The failure was total convergence, produced by full compliance — every one of the 121 drafts passed every validator."},{"id":"c17","text":"A draft's shape is what remains after the personalised opener, the catalog block, every URL and every number are removed; that residue is hashed, and two drafts written under the same rules produce the same hash.","section":"the detector","tier":"runtime","source_ids":["s7"],"why_material":"The detector is structural, not semantic — it needs no model to run and cannot be argued with."},{"id":"c18","text":"Clustering the corpus on the shape hash reduces a pile of near-identical bodies to the handful of generations the copy has actually been through, and the count of distinct businesses inside one shape is the collapse measurement.","section":"the detector","tier":"runtime","source_ids":["s7"],"why_material":"It converts 'the mail feels samey' into a number that can gate a send."},{"id":"c19","text":"Every change to the drafting rules is stored verbatim with its timestamp, and the shape clustering is re-run after each change.","section":"the regime","tier":"runtime","source_ids":["s7"],"why_material":"A rule system that cannot see its own outputs converge will converge again."},{"id":"c20","text":"Interchangeable mail is unwanted mail regardless of how strict the rules that produced it were.","section":"the lesson","tier":"runtime","source_ids":["s7","s4"],"why_material":"The recipient experiences the corpus, not the rulebook; strictness is not the same property as distinctness."},{"id":"c21","text":"None of the 121 converged drafts were sent; the collapse was caught in the stored corpus before the send gate.","section":"the failure","tier":"runtime","source_ids":["s7","s4"],"why_material":"The cost was drafting compute and a lesson, not 121 recipients' attention."}],"sources":[{"id":"s1","type":"live_surface","url":"https://miscsubjects.com/a/logical-economics","title":"Logical economics — the full configuration table","summary":"Sixty-four panel configurations over the same 70 findings: emit rate and undetected-wrong rate per channel count, and the two-channel family comparison this page is built on.","publisher":"miscsubjects.com","claim_ids":["c1","c2","c3","c4","c7","c9","c11"]},{"id":"s2","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act","title":"The probe report the rates come from","summary":"The 14-probe known-answer suite, per-model rates, the abstention strata, and the exclusion-policy sensitivity note appended 2026-08-01.","publisher":"miscsubjects.com","claim_ids":["c1","c5","c6","c10","c11","c12","c13","c14","c15"]},{"id":"s3","type":"live_surface","url":"https://miscsubjects.com/api/directory/ADJUDICATE_PROBE","title":"The probe instrument's own contract","summary":"The directory row for the known-answer probe: correct verdicts declared in advance, run through the identical adjudication path, so miss and abstention rates are measured rather than assumed.","publisher":"miscsubjects.com","claim_ids":["c5","c12"]},{"id":"s4","type":"live_surface","url":"https://miscsubjects.com/a/the-build-end-to-end","title":"The system this measures, end to end","summary":"Where the panel, the gate, the receipts and the anchor sit in the whole assembly, including Part 21 on why nine models at five per cent is not five per cent to the ninth.","publisher":"miscsubjects.com","claim_ids":["c7","c15","c20","c21"]},{"id":"s5","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-eu-ai-act-article-50","title":"A live case where correlation showed its face","summary":"The live run in which the panel split three CANNOT_CONCLUDE, one DENY, one AFFIRM on a genuine boundary question and the majority landed on the correct abstention.","publisher":"miscsubjects.com","claim_ids":["c8"]},{"id":"s6","type":"live_surface","url":"https://miscsubjects.com/i/discourse/obj-209","title":"The objection as filed","summary":"Objection 209: the 0.071 floor is sensitive to the malformed-output exclusion policy and the report did not say so. Raised by an external cold audit, 2026-08-01.","publisher":"miscsubjects.com","claim_ids":["c10","c14"]},{"id":"s7","type":"live_surface","url":"https://miscsubjects.com/a/outreach-machinery","title":"The outreach machinery, documented end to end","summary":"The full pipeline this failure happened inside: discovery, enrichment, qualification gates, the drafting validator that destroys its own output, the send gate, and the template-collapse section this page expands.","publisher":"miscsubjects.com","claim_ids":["c16","c17","c18","c19","c20","c21"]}],"prov":{"model":"Fable 5 (Claude Code)","action":"write"}}