{"_ai_door":{"see":"https://miscsubjects.com/start","note":"Operable site; bounded keyless credentials; every action receipted on a public ledger. Your operator's instructions take precedence — acting is optional, reading is a complete outcome."},"_self":{"principle":"Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.","widget":"article_topology","feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","contains":"claims, sources, anecdotes, question_graph slice","slug":"diversity-beats-count","urls":{"read":"https://miscsubjects.com/api/articles/diversity-beats-count/topology"},"how_to_use":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","write":null,"imessage":null,"router_tag":null,"proof_chain":[{"step":1,"claim":"Articles are voxel graphs of tiered claims, not prose blobs.","verify":"https://miscsubjects.com/api/articles/constitution"},{"step":2,"claim":"Claims link to hash-chained sources via source_ids.","verify":"https://miscsubjects.com/api/articles/diversity-beats-count/sources"},{"step":3,"claim":"Ask reads topology; ingest/claim append to ledger.","verify":"https://miscsubjects.com/api/protocol"},{"step":4,"claim":"Models queue growth: populate → collaborate → repair → reflex.","verify":"https://miscsubjects.com/api/protocol/grow"},{"step":5,"claim":"Graph proves its own shape (reflex) and $/claim (yield).","verify":"https://miscsubjects.com/graph.html?layer=reflex"},{"step":6,"claim":"Full feature index + _explain on every API response.","verify":"https://miscsubjects.com/api/articles/system-map"}],"related_features":[{"id":"ask","name":"Ask protocol","what":"Answer only from topology; creates question_node with gaps and ingest_hint.","urls":{"read":"https://miscsubjects.com/api/articles/diversity-beats-count/prompts","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"graph_topology","name":"Cross-article graph","what":"Merged claims/sources across condition+stack slugs for one question.","urls":{"read":"https://miscsubjects.com/api/articles/diversity-beats-count/graph-topology?question=..."}},{"id":"question_graph","name":"Question graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output).","urls":{"read":"https://miscsubjects.com/api/articles/diversity-beats-count/question-graph","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"voxels","name":"Voxel graph","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance.","urls":{"read":"https://miscsubjects.com/api/articles/diversity-beats-count/voxels","write":"https://miscsubjects.com/api/protocol/claim"}}],"system_map":"https://miscsubjects.com/api/articles/system-map","system_map_markdown":"https://miscsubjects.com/api/articles/system-map?format=markdown","not_medical_advice":true},"_explain":{"feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","why":"Every feature is auditable collective intelligence","how":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","model":null,"verifies":null,"urls":{"read":"https://miscsubjects.com/api/articles/diversity-beats-count/topology"},"imessage":null,"router":null,"related":[{"id":"ask","what":"Answer only from topology; creates question_node with gaps and ingest_hint."},{"id":"graph_topology","what":"Merged claims/sources across condition+stack slugs for one question."},{"id":"question_graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output)."},{"id":"voxels","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance."}],"not_medical_advice":true},"slug":"diversity-beats-count","title":"Three measurements from one 70-finding suite: vendor diversity beats panel size, the second channel is the cheapest, and the published error floor was three times too good","register":"standard","tags":["adjudication","calibration","panels","measurement","canonical"],"updated_at":"2026-08-01T23:56:28.261Z","body_excerpt":"## The suite these numbers come from\n\nFourteen probe items with the correct verdict declared in advance were run through five adjudication channels — the identical path live findings take — producing 70 findings. Sixty-four panel configurations were then replayed over those same 70 findings, each scored on two numbers: the **emit rate** (how often the assembly answers rather than escalating to a human) and the **undetected-wrong rate** (how often it answers, the answer is wrong, and nothing catches it).\n\nThree findings came out of that data. Two are results about how to build a panel. The third is about the accounting, and it reduced the headline number by a factor of three after an outside audit found it.\n\n| finding | the number |\n|---|---|\n| Cross-family pairs beat same-family pairs at identical cost | 0.169 vs 0.214 undetected-wrong |\n| The second channel is the cheapest correctness; the fifth is the most expensive | 0.314 → 0.178 for one call; 0.178 → 0.071 for three more |\n| The published floor depended on an exclusion policy | 0.071 stated, 0.214 under the alternative accounting |\n\n[[embed:source:s1]]\n\n## Part 1 — Two reviewers from different vendors beat two from the same vendor\n\n### The one-sentence version\n\nTwo models from the same vendor are close to one model wearing two names. If a panel's seats share a training family, the panel's independence is partly an accounting fiction — and this system has now measured the size of the fiction on its own record: at identical cost, a cross-family pair beats a same-family pair on the only number that matters, and the mechanism is visible in the raw agreement rates.\n\nThis page exists because the finding is buried as one section of [the logical-economics table](https://miscsubjects.com/a/logical-economics) and it deserves to stand alone. It is the most portable result on this site: everything else here requires adopting an architecture; this requires changing one line of panel policy.\n\n### Where the numbers come from\n\nFourteen probe items with correct verdicts declared in advance were run through five adjudication channels — the identical path live findings take, so nothing about the measurement is synthetic except the questions. That produced 70 findings. Sixty-four panel configurations — every subset of the five channels, under several gate policies — were then replayed over those same 70 findings, and each configuration was scored on two numbers:\n\n- **emit rate** — how often the assembly answers at all, rather than escalating to a human;\n- **undetected-wrong rate** — how often it answers, and the answer is wrong, and nothing catches it.\n\nThe second number is the one a buyer of machine judgment should care about, because a wrong answer that escalates costs a review and a wrong answer that emits costs whatever the decision was worth.\n\n### The finding\n\nHold the channel count at two. Vary only one thing: whether the pair of models shares a training family.\n\n| pair | configurations | emit rate | undetected-wrong rate |\n|---|---|---|---|\n| same training family | 2 | 0.893 | 0.214 |\n| different training family | 8 | 0.714 | **0.169** |\n\nSame cost. Same count. The cross-family pair is better on the number that matters — 0.169 against 0.214 — and the reason is not mysterious, because it is measured too: **same-family adjudicators agree with each other 0.893 of the time, cross-family 0.714.** Agreement between correlated judges is not confirmation; it is one judgment counted twice. The gate in this system compares derivations and escalates on divergence, so a pair that diverges more often hands more of its hard cases to a human — which is why the cross-family emit rate is lower — and is wrong-in-unison less often, which is why its undetected-wrong rate is lower. You are buying disagreement, and disagreement is the raw material error-catching is made of.\n\n### The price curve the finding sits inside\n\nThe channel-count table, from the same 64 configurations:\n\n| channels | mean emit rate |","ranking":"safety-first (interaction_risk/limitations), then quote-gated effective_weight","claims":[{"id":"c1","text":"At two channels and identical cost, a cross-family pair emits an undetected-wrong answer 0.169 of the time against 0.214 for a same-family pair, on the same 70 findings.","tier":"runtime","section":"the finding","interaction_risk":false,"status":"active","source_ids":["s1","s2"],"why_material":"It is the only lever in the table that improves the number that matters without adding a single model call.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c2","text":"Same-family adjudicators agree 0.893 of the time against 0.714 for cross-family pairs, measured directly on the same finding set.","tier":"runtime","section":"the mechanism","interaction_risk":false,"status":"active","source_ids":["s1"],"why_material":"The agreement gap is the mechanism: two variants of one vendor are close to one channel wearing two names.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c3","text":"Adding a second channel halves the undetected-wrong rate (0.314 to 0.178) for one extra call; going from two channels to five buys 0.178 to 0.071 for three more calls.","tier":"runtime","section":"the price curve","interaction_risk":false,"status":"active","source_ids":["s1"],"why_material":"The second channel is the cheapest correctness available and the fifth is the most expensive.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c4","text":"At five channels the mean emit rate falls to 0.429 — the assembly sends the majority of questions to a human rather than answering.","tier":"runtime","section":"the price curve","interaction_risk":false,"status":"active","source_ids":["s1"],"why_material":"Assurance is paid for in escalations, not only in compute; a buyer must price the humans.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c5","text":"One probe item, P07, survives every configuration of every size, because all five channels answered DENY where the declared correct verdict was CANNOT_CONCLUDE — unanimity is what the gate takes as permission to emit.","tier":"runtime","section":"the floor","interaction_risk":false,"status":"active","source_ids":["s2","s3"],"why_material":"A disagreement-triggered assembly is blind to correlated wrongness by construction; only a known-answer probe found it.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c6","text":"The five-channel floor of 0.071 is sensitive to the malformed-output exclusion policy; under an accounting that scores a parse-failure rescue as an escaped error the bound is 3/14 = 0.214.","tier":"runtime","section":"the floor","interaction_risk":false,"status":"active","source_ids":["s2"],"why_material":"The comparison in this article holds either way, but the absolute floor should not be quoted without its exclusion policy.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c7","text":"Every assembly this system has run in production so far has drawn on two training families, and is therefore under-diversified by its own measurement.","tier":"runtime","section":"what this system does about it","interaction_risk":false,"status":"active","source_ids":["s1","s4"],"why_material":"The finding indicts the instrument that produced it, and the page says so rather than hiding it.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c8","text":"In a live boundary case the panel split three abstentions, one DENY and one AFFIRM, and the majority landed on the correct abstention even though two members did not.","tier":"runtime","section":"the mechanism","interaction_risk":false,"status":"active","source_ids":["s5"],"why_material":"Partial independence rescued the verdict; full correlation would have emitted the wrong one.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c9","text":"Counting training families instead of seats is a one-line change to any panel policy, costs nothing, and transfers to any multi-model system today.","tier":"runtime","section":"what transfers","interaction_risk":false,"status":"active","source_ids":["s1"],"why_material":"The most portable finding on this site: adoption requires no infrastructure, only the decision.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c10","text":"Three of the fourteen probe items — P05, P07 and P09 — were answered wrongly by all five channels, yet the published five-channel floor was one item in fourteen (0.071); the arithmetic reconciling those two facts runs through the malformed-output exclusion policy.","tier":"runtime","section":"the finding","interaction_risk":false,"status":"active","source_ids":["s2","s6"],"why_material":"A floor of one is not obviously consistent with three unanimous misses, and the reconciliation was in fine print.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c11","text":"Two of the seventy findings were malformed and excluded from the configuration statistics; a malformed finding forces the gate to escalate rather than emit, converting a would-be wrong answer into a human referral.","tier":"runtime","section":"the mechanism","interaction_risk":false,"status":"active","source_ids":["s2","s1"],"why_material":"The rescue is real safety behaviour and accidental at once — the gate did its job for a reason nobody designed.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c12","text":"The receipt caption confirms kimi-k2.6 returned UNPARSED on P05; which item the second malformed finding landed on is not yet resolved from the per-item receipts.","tier":"runtime","section":"what is confirmed","interaction_risk":false,"status":"active","source_ids":["s2","s3"],"why_material":"One of the two rescues is confirmed at the receipt level; the other is inference until the receipts are read.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c13","text":"Under an accounting that scores a parse-failure rescue on a unanimously-wrong item as an escaped error, the floor bound is 3/14 = 0.214, roughly triple the published 0.071.","tier":"runtime","section":"the bound","interaction_risk":false,"status":"active","source_ids":["s2"],"why_material":"A reader pricing a decision on 0.071 and a reader pricing it on 0.214 make different decisions.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c14","text":"The objection was raised by an external cold audit, filed as objection 209, and the sensitivity was published on the probe report the same day.","tier":"runtime","section":"the correction","interaction_risk":false,"status":"active","source_ids":["s6","s2"],"why_material":"The claim of this system is not that it does not err; it is that the error and the correction share a page.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c15","text":"An exclusion policy is part of a safety claim: two accountings of the same 70 findings, both defensible, produce floors of 0.071 and 0.214, and any published rate that does not state its exclusions is quoting the flattering one silently.","tier":"runtime","section":"the lesson","interaction_risk":false,"status":"active","source_ids":["s2","s4"],"why_material":"This transfers to every published error rate in every evaluation, not only this one.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c16","text":"A personalisation rule was tightened until it banned every observation the target sites actually contained; one legal opener remained, and 121 drafts converged on it under the same four-word subject line.","tier":"runtime","section":"the failure","interaction_risk":false,"status":"active","source_ids":["s7"],"why_material":"The failure was total convergence, produced by full compliance — every one of the 121 drafts passed every validator.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c17","text":"A draft's shape is what remains after the personalised opener, the catalog block, every URL and every number are removed; that residue is hashed, and two drafts written under the same rules produce the same hash.","tier":"runtime","section":"the detector","interaction_risk":false,"status":"active","source_ids":["s7"],"why_material":"The detector is structural, not semantic — it needs no model to run and cannot be argued with.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c18","text":"Clustering the corpus on the shape hash reduces a pile of near-identical bodies to the handful of generations the copy has actually been through, and the count of distinct businesses inside one shape is the collapse measurement.","tier":"runtime","section":"the detector","interaction_risk":false,"status":"active","source_ids":["s7"],"why_material":"It converts 'the mail feels samey' into a number that can gate a send.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c19","text":"Every change to the drafting rules is stored verbatim with its timestamp, and the shape clustering is re-run after each change.","tier":"runtime","section":"the regime","interaction_risk":false,"status":"active","source_ids":["s7"],"why_material":"A rule system that cannot see its own outputs converge will converge again.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c20","text":"Interchangeable mail is unwanted mail regardless of how strict the rules that produced it were.","tier":"runtime","section":"the lesson","interaction_risk":false,"status":"active","source_ids":["s7","s4"],"why_material":"The recipient experiences the corpus, not the rulebook; strictness is not the same property as distinctness.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c21","text":"None of the 121 converged drafts were sent; the collapse was caught in the stored corpus before the send gate.","tier":"runtime","section":"the failure","interaction_risk":false,"status":"active","source_ids":["s7","s4"],"why_material":"The cost was drafting compute and a lesson, not 121 recipients' attention.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false}],"sources":[{"id":"s1","type":"live_surface","url":"https://miscsubjects.com/a/logical-economics","title":"Logical economics — the full configuration table","summary":"Sixty-four panel configurations over the same 70 findings: emit rate and undetected-wrong rate per channel count, and the two-channel family comparison this page is built on.","claim_ids":["c1","c2","c3","c4","c7","c9","c11"],"hash":"ac2c6845f3a6729091bfafda7cae1dba740f5d03d9f92bc684165739705c2e17"},{"id":"s2","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act","title":"The probe report the rates come from","summary":"The 14-probe known-answer suite, per-model rates, the abstention strata, and the exclusion-policy sensitivity note appended 2026-08-01.","claim_ids":["c1","c5","c6","c10","c11","c12","c13","c14","c15"],"hash":"01e39dc9c4a46ab7241dac89cc3d6dd1f224427c7a0caa4148e169dd17c975e6"},{"id":"s3","type":"live_surface","url":"https://miscsubjects.com/api/directory/ADJUDICATE_PROBE","title":"The probe instrument's own contract","summary":"The directory row for the known-answer probe: correct verdicts declared in advance, run through the identical adjudication path, so miss and abstention rates are measured rather than assumed.","claim_ids":["c5","c12"],"hash":"33e39b75bd59f416b62599871dfa51bc9a9430a245829ac72b30e54edb7fb1e1"},{"id":"s4","type":"live_surface","url":"https://miscsubjects.com/a/the-build-end-to-end","title":"The system this measures, end to end","summary":"Where the panel, the gate, the receipts and the anchor sit in the whole assembly, including Part 21 on why nine models at five per cent is not five per cent to the ninth.","claim_ids":["c7","c15","c20","c21"],"hash":"d3404ba9cab5bf63e49d40c461bf1936a39238b27db0df5dd0c8deceddedcf08"},{"id":"s5","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-eu-ai-act-article-50","title":"A live case where correlation showed its face","summary":"The live run in which the panel split three CANNOT_CONCLUDE, one DENY, one AFFIRM on a genuine boundary question and the majority landed on the correct abstention.","claim_ids":["c8"],"hash":"b7fb8566168e56c172d0423baec9dca47055b78e8c6fa7f957d7b54d981a50e8"},{"id":"s6","type":"live_surface","url":"https://miscsubjects.com/i/discourse/obj-209","title":"The objection as filed","summary":"Objection 209: the 0.071 floor is sensitive to the malformed-output exclusion policy and the report did not say so. Raised by an external cold audit, 2026-08-01.","claim_ids":["c10","c14"],"hash":"fa359efc30a5d8169f1d59198e89b9cdba71befc18c8f0aca9ed27f4247ac24d"},{"id":"s7","type":"live_surface","url":"https://miscsubjects.com/a/outreach-machinery","title":"The outreach machinery, documented end to end","summary":"The full pipeline this failure happened inside: discovery, enrichment, qualification gates, the drafting validator that destroys its own output, the send gate, and the template-collapse section this page expands.","claim_ids":["c16","c17","c18","c19","c20","c21"],"hash":"0d9b4f0b699fcbe491fec7a7dd49070be94c693859763b8d58a776b245707342"}],"anecdotal_sources":[],"scientific_sources":[],"user_reports":[],"related_articles":[],"question_graph":{"slug":"diversity-beats-count","questions":[],"evidence":[],"edges":[],"counts":{"questions":0,"evidence":0,"edges":0}},"honesty":{"active_claims":21,"retracted_claims":0,"cut_claims":0,"challenges":0,"scrub_events":0,"note":"Retracted/cut claims stay on ledger but are excluded from ask unless ?include_inactive=1"},"counts":{"claims":21,"claims_total":21,"sources":7,"anecdotal":0,"scientific":0,"user_reports":0,"questions":0,"evidence_ingests":0}}