# A permanent operating system that lets successive AI models inherit one person’s work

slug: the-build-end-to-end · https://miscsubjects.com/a/the-build-end-to-end · category: canon · tags: canonical, end-to-end, proof, oip, comparison, roadmap, receipts, ongoing · updated 2026-08-04T23:31:57.291Z

**The product this build sells is proven work** — one explicit claim about completed AI work, bound to its complete formation record, with standing authority for any stranger to inspect it and a receipt for every inspection. The standard is [[proven-work|the proof law]]; the one-step demo any model can run right now is `GET /api/proven-work/three-models-deliberate-one-statutory-question/inspect`, which returns the record plus your own inspection receipt; the commercial form is scoped API access to one workflow in, a `proven_work` field on every result out. Everything documented below is the machinery that makes that unit producible.

## Why this exists

One person decided he did not want to be somebody who uses AI tools. He wanted to be a structure that AI operates through. So he took everything he actually runs on — his writing, his standards, his reasoning, his business, his phone, his laptop's shell, his money, his philosophy, and his mistakes — and gave each piece an address, a contract, and a permanent record. Not notes about the work. The work itself, in a form a model can pick up and run.

Which means the articles are not the point, and the protocol is not the point either. The bet is that the models will keep changing and the structure will not have to. When the next one arrives it reads one URL and inherits the whole thing: how the work is done, what it is permitted to do, what was already decided and why, and what went wrong last time. Most people using AI start over every conversation. This does not. The hand-off is a single address — [https://miscsubjects.com/api/articles/the-build-end-to-end/bundle?format=markdown](https://miscsubjects.com/api/articles/the-build-end-to-end/bundle?format=markdown) — and Part 19 lists the rest.

That is also why every failure here is public and permanent. A memory that deletes its own errors is worthless to whatever picks it up next: the successor repeats the mistake, because nothing told it. The honesty is not a virtue on display. It is load-bearing. A structure only survives model turnover if it never lies to its successor. The receipt in §CHECK item 4 is one of those errors, left where it happened, with the correction attached.

## What this is, and what it compares to

**In one sentence.** A public, self-describing system in which every article, tool, skill, law, claim, source and API is the same kind of object — one address, one operating contract, one history, a receipt for every action — so that any model can discover it, operate it under stated authority, and inherit everything decided before it.

**What it does end to end.** It discovers a market, news or development event; compares that event with what the build can actually do; implements a useful change when the evidence warrants one; publishes the implementation and its receipts; identifies the people who bear the relevant cost; stores each individualized outreach object inside the article; sends through a gated, tracked lane; records acceptance, delivery, response and failure separately; and carries those results into the laws, skills and next run. The [one-loop record](https://miscsubjects.com/a/one-loop) proves one complete pass rather than merely describing the intended cycle.

**What it combines.** It places an ontology, tool discovery, durable execution, hypermedia, evidence, authority, publication, outreach and failure memory inside the same addressable object system. The comparison systems below establish those components separately: Foundry connects objects and governed actions; MCP standardizes context exchange; LangGraph provides durable agent orchestration; REST supplies self-describing resources and hypermedia. This build's claim is narrower than uniqueness: this deployed implementation joins those classes to public articles, receipts, objections and commercial follow-through.

**What is proven now, and what is not.** Proven: the public endpoints run; the ledger and revision chains are openable; the build can implement, publish, identify recipients, send through its governed lane and preserve the outcome; and the correction record survives into the next model's context. Unproven: external adoption, a repeatable buyer, a market-clearing price, independent reliance on the records, and production safety at anyone else's scale.

**What it compares to.** These are comparisons by function, not claims that the systems are equivalent.

| comparison | what the primary source says it does | what this deployed implementation joins to it |
|---|---|---|
| **[Palantir Foundry Ontology](https://www.palantir.com/docs/foundry/ontology/overview)** | an operational layer of objects, properties, links, actions, functions and dynamic security tied to real-world counterparts | public articles, claims, laws, tools and outreach share an address and permanent correction history alongside the operational objects |
| **[Model Context Protocol](https://modelcontextprotocol.io/docs/learn/architecture)** | a client-host-server protocol for exchanging tools, resources, prompts and notifications; its stated scope is context exchange | the MCP projection points to objects whose application-level authority, revisions, receipts, objections and commercial outcomes persist beyond a client session |
| **[LangGraph](https://docs.langchain.com/oss/python/langgraph/overview)** | a low-level runtime for long-running stateful agents, durable execution and human-in-the-loop control | the durable workflow operates on public, self-describing objects and leaves reader-openable evidence and outreach outcomes, not only agent state |
| **[REST and hypermedia](https://ics.uci.edu/~fielding/pubs/dissertation/rest_arch_style.htm)** | identified resources, self-describing messages and hypermedia as the engine of application state | each representation exposes not only the next link but the operating contract, authority boundary, evidence and receipt for the object |
| **A research paper** | a durable argument, sources and a field for objection | the description and the running system occupy one artifact: architectural claims resolve to live endpoints and objections attach to the claim they dispute |

**If a shelf is required:** it is closest to an ontology layer of the Foundry kind, built in the open by one operator, with the evidence graph, the error rate and the failure record public rather than contractual.

**The honest limit, stated first rather than last:** one operator, near-zero adoption, no external party has priced any of it. An existence proof, public and operational, is not a standard, a market, or a movement. Part 7 carries this comparison in full, and Part 8 carries the defects.

## §CHECK — verify the spine in four minutes

Five URLs, in order. Each says what it proves and what would falsify it. A reader who opens these has checked the load-bearing claims without reading prose.

**1. The ledger head, sealed current.** [https://miscsubjects.com/api/chain/head](https://miscsubjects.com/api/chain/head)
*Proves:* the append-only chain covers every event through 689,866, with the hash recipe published so you can recompute it. *Falsified by:* a head that does not match a recomputation from the stated recipe, or an event count behind the live ledger.

**2. The external anchor.** [https://miscsubjects.com/api/anchor/3be5071eb3035ca29093c6713646bbe21bdca6cce262fc7f7eb64080c04e61fe](https://miscsubjects.com/api/anchor/3be5071eb3035ca29093c6713646bbe21bdca6cce262fc7f7eb64080c04e61fe)
*Proves:* that head is bound to drand round 6331315 and Bitcoin block 960173. *Falsified by:* an anchor whose bound surfaces do not contain what it claims.

**3. The beacon, on infrastructure unrelated to this system.** [https://api.drand.sh/public/6331315](https://api.drand.sh/public/6331315)
*Proves:* the randomness and BLS signature in the anchor are the League of Entropy's, unpredictable before their cadence time, so the timeline cannot be backdated. *Falsified by:* a mismatch between this round's randomness and the anchor's copy of it.

**4. The correction receipt.** [https://miscsubjects.com/api/dispatch?confirm=inv_70rfvm6bf3](https://miscsubjects.com/api/dispatch?confirm=inv_70rfvm6bf3)
*Proves:* a send that delivered nothing reads *attempt proven; result not observed*, and carries `provider_status: 503` in public. This receipt was labelled *material result proven* until an external audit caught it on 2026-07-30; the classifier now derives the label from the provider's outcome and 124 historical rows were re-graded. *Falsified by:* a provider failure anywhere in the ledger still reading as an observed result.

**5. The instrument's error rate.** [https://miscsubjects.com/api/directory/ADJUDICATE_PROBE](https://miscsubjects.com/api/directory/ADJUDICATE_PROBE)
*Proves:* the panel's measured error rate, published per model per rule set: [https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act](https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act). Four rates over a 14-probe stratified suite pinned at SHA-256 `ffa8135dd89d29a8…`. The headline: this panel manufactures a verdict where it should abstain between 21% and 42% of the time, scores at or near perfect where the text settles the question, and abstains reliably when it abstains at all. *Falsified by:* a re-run of the published suite against the same rule set hash producing materially different rates, or a demonstration that a declared expected verdict is wrong — the suite is published for exactly that.

Items 1 through 4 characterise the record. Item 5 characterises the instrument, which is what turns a finding from a documented opinion into evidence with an error bar.

## What this is

A working prototype of a different way to organize AI systems. Every article, tool, skill, law, claim, source, API, CLI, and MCP server on this system is the same kind of invocable object: one address, one contract, one history, a receipt for every action. It runs on a single Cloudflare account, is operated by one person and the models he directs, and is public — any model that can open a URL discovers any object, reads its full contract from one JSON row, and with a single token edits it, hash-checked, ledgered and receipted.

- **The unit.** Everything on this system — every article, tool, skill, law, claim, source, API, CLI, and MCP server — is the same kind of object: one address, one contract, one history, a receipt for every action.
- **Discovery.** `GET https://miscsubjects.com/api/directory/search?q=<words>` finds any object. `GET https://miscsubjects.com/api/directory/<KEY>` returns its complete operating contract: endpoint, verbs, arguments, auth shape, examples. There is no schema file to load and no prompt that enumerates capabilities.
- **Scale, measured.** 887 enabled directory rows are invocable capabilities; the public registry publishes 885 of them. 174,309 ledgered invocations across 323 distinct capability objects since 2026-06-29. Corpus figures render live in Part 4.
- **Authority.** One token format. `?share=<token>` in a browser or `Authorization: Bearer <token>` in curl — interchangeable. Validate either at `https://miscsubjects.com/api/token/validate`.
- **Externally anchored.** The ledger chain head is sealed current through 689,866 events and bound to surfaces this operator does not control: drand round 6331315 (BLS-signed by the League of Entropy) and Bitcoin block 960173. Head: [https://miscsubjects.com/api/chain/head](https://miscsubjects.com/api/chain/head) · anchor: [https://miscsubjects.com/api/anchor/3be5071eb3035ca29093c6713646bbe21bdca6cce262fc7f7eb64080c04e61fe](https://miscsubjects.com/api/anchor/3be5071eb3035ca29093c6713646bbe21bdca6cce262fc7f7eb64080c04e61fe)
- **Proof.** Every invocation writes an append-only ledger row with a public receipt at `https://miscsubjects.com/api/dispatch?confirm=<invocation_id>`, and the receipt states whether the *result* was observed or only the *attempt*. It does not say "200 OK" and call that proof.
- **Self-computed honesty.** The system publishes its own grounding figure — the share of claims carrying an openable source — live, including when it is unflattering: `https://miscsubjects.com/api/metrics/grounding`

[[embed:source:s1]]


## Part 1 — The proof table

Each row is a capability class, a real receipt from the live ledger, and the article that documents it. Every receipt URL is public and requires no token. Receipts marked *material* mean the result itself was observed and recorded; that distinction is enforced by the receipt generator, not by prose.

**Lead generation and scraping** — discovery from public sources, enrichment, MX verification, AI scoring, drafting, sending. 19 `LEADS_*` capabilities; 187 discovery invocations recorded.

Receipts: discovery [https://miscsubjects.com/api/dispatch?confirm=inv_zx53xxla5w](https://miscsubjects.com/api/dispatch?confirm=inv_zx53xxla5w) (material) · scoring [https://miscsubjects.com/api/dispatch?confirm=inv_zshmn0ucq1](https://miscsubjects.com/api/dispatch?confirm=inv_zshmn0ucq1) (material) · MX verification [https://miscsubjects.com/api/dispatch?confirm=inv_zsdvunxdrk](https://miscsubjects.com/api/dispatch?confirm=inv_zsdvunxdrk) (material) · send [https://miscsubjects.com/api/dispatch?confirm=inv_tp7h228phk](https://miscsubjects.com/api/dispatch?confirm=inv_tp7h228phk) (material)

Article: [https://miscsubjects.com/a/oip-system-leads](https://miscsubjects.com/a/oip-system-leads) · the priced version of this loop: [https://miscsubjects.com/a/killbox-specification-v1-2](https://miscsubjects.com/a/killbox-specification-v1-2)

**Paid advertising** — 46 `META_ADS_*` capabilities covering accounts, campaigns, ad sets, ads, creatives, audiences, lookalikes, catalogues, pixels, budgets, delivery estimates, insights (sync and async), and the Conversions API.

Article: [https://miscsubjects.com/a/oip-system-meta](https://miscsubjects.com/a/oip-system-meta)

**Image and video generation** — three independent providers plus durable storage. ArcAds (287 generation invocations), Grok images (54), OpenAI images (6), video, and re-storage to permanent URLs (85).

Receipts: ArcAds generate [https://miscsubjects.com/api/dispatch?confirm=inv_zq3jcx3icf](https://miscsubjects.com/api/dispatch?confirm=inv_zq3jcx3icf) (material) · store to durable URL [https://miscsubjects.com/api/dispatch?confirm=inv_zv2n9nml3n](https://miscsubjects.com/api/dispatch?confirm=inv_zv2n9nml3n) (material) · Grok image [https://miscsubjects.com/api/dispatch?confirm=inv_zuklqy80eb](https://miscsubjects.com/api/dispatch?confirm=inv_zuklqy80eb) (material) · OpenAI image [https://miscsubjects.com/api/dispatch?confirm=inv_n46cio5qam](https://miscsubjects.com/api/dispatch?confirm=inv_n46cio5qam) (material)

Article: [https://miscsubjects.com/a/oip-system-arcads](https://miscsubjects.com/a/oip-system-arcads)
The illustration at the top of this page was generated through that pipeline while this page was being written: receipt [https://miscsubjects.com/api/dispatch?confirm=inv_n56yqd1lpu](https://miscsubjects.com/api/dispatch?confirm=inv_n56yqd1lpu)

**Messaging across every channel a person actually uses** — iMessage, SMS, WhatsApp, Telegram intake, group chats, polls, reactions, contact cards, delivery-status webhooks. 64 `BLOOIO_*` capabilities plus a second provider (`TWOCHAT_*`) for WhatsApp groups, 10 `PHONE_*` handlers for share-sheet intake (text, URL, image, voice note, location, clipboard), and tracked email.

Receipts: message sent [https://miscsubjects.com/api/dispatch?confirm=inv_oh5v2hofv4](https://miscsubjects.com/api/dispatch?confirm=inv_oh5v2hofv4) (material) · WhatsApp group send [https://miscsubjects.com/api/dispatch?confirm=inv_aokydx9k72](https://miscsubjects.com/api/dispatch?confirm=inv_aokydx9k72) (material) · tracked email [https://miscsubjects.com/api/dispatch?confirm=inv_zsff9euzwm](https://miscsubjects.com/api/dispatch?confirm=inv_zsff9euzwm) (material) · plain email [https://miscsubjects.com/api/dispatch?confirm=inv_zzijgpp911](https://miscsubjects.com/api/dispatch?confirm=inv_zzijgpp911) (material)

Articles: [https://miscsubjects.com/a/oip-system-phone](https://miscsubjects.com/a/oip-system-phone)  
· [https://miscsubjects.com/a/oip-system-blooio](https://miscsubjects.com/a/oip-system-blooio)  
· [https://miscsubjects.com/a/oip-system-twochat](https://miscsubjects.com/a/oip-system-twochat)  
· [https://miscsubjects.com/a/oip-system-email](https://miscsubjects.com/a/oip-system-email)

**Social publishing** — X posting, replies, deletion, search, identity (245 post invocations); Reddit search, thread reading, replying.

Receipt: post [https://miscsubjects.com/api/dispatch?confirm=inv_zgiu8omiuf](https://miscsubjects.com/api/dispatch?confirm=inv_zgiu8omiuf) (material)

Articles: [https://miscsubjects.com/a/oip-system-x](https://miscsubjects.com/a/oip-system-x)  
· [https://miscsubjects.com/a/oip-system-reddit](https://miscsubjects.com/a/oip-system-reddit)

**Voice** — text to speech, speech to text, voice notes, audio playback on the operator's machine.

Receipt: speech synthesis [https://miscsubjects.com/api/dispatch?confirm=inv_i5lrm9eshb](https://miscsubjects.com/api/dispatch?confirm=inv_i5lrm9eshb) (material)

Article: [https://miscsubjects.com/a/oip-system-voice](https://miscsubjects.com/a/oip-system-voice)

**Terminal and computer control** — 41 `LOCAL_*` capabilities: shell execution (370 invocations), file read and write, grep, screenshots (20), OCR, clipboard, window and app control, UI clicks and keystrokes, notifications, launchd, ports, processes, AppleScript. Plus 45 `CLI_*` rows wrapping the actual command-line tools installed on the machine — git, gh, wrangler, docker, kubectl, terraform, gcloud, aws, ffmpeg, imagemagick, pandoc, node, npm, python, jq, psql, sqlite, and every coding agent CLI.

Receipts: shell execution [https://miscsubjects.com/api/dispatch?confirm=inv_zzzb67cjqq](https://miscsubjects.com/api/dispatch?confirm=inv_zzzb67cjqq) (material) · GitHub CLI [https://miscsubjects.com/api/dispatch?confirm=inv_rba5bflts4](https://miscsubjects.com/api/dispatch?confirm=inv_rba5bflts4) (material) · screenshot [https://miscsubjects.com/api/dispatch?confirm=inv_xkkkwli6ee](https://miscsubjects.com/api/dispatch?confirm=inv_xkkkwli6ee) (material)

Articles: [https://miscsubjects.com/a/oip-system-local](https://miscsubjects.com/a/oip-system-local)  
· [https://miscsubjects.com/a/oip-system-desktop](https://miscsubjects.com/a/oip-system-desktop)  
· [https://miscsubjects.com/a/oip-system-cli](https://miscsubjects.com/a/oip-system-cli)

**Browser control** — headless fetch, markdown extraction, link extraction, PDF capture, screenshots, full Playwright automation, and a browser-use agent.

Receipt: Playwright automation [https://miscsubjects.com/api/dispatch?confirm=inv_irpi9hivmi](https://miscsubjects.com/api/dispatch?confirm=inv_irpi9hivmi) (material)

Article: [https://miscsubjects.com/a/oip-system-browser](https://miscsubjects.com/a/oip-system-browser)

**Infrastructure provisioning** — 111 `CF_*` capabilities: create and delete D1 databases, KV namespaces, R2 buckets; read and deploy Workers; containers with file read/write and exec; DNS, observability queries, GraphQL analytics, audit logs, AutoRAG, browser rendering, DEX, CASB. Plus direct data-plane rows for D1, KV, R2, and Pages.

Receipts: KV write [https://miscsubjects.com/api/dispatch?confirm=inv_w09f2zr555](https://miscsubjects.com/api/dispatch?confirm=inv_w09f2zr555) (material) · R2 write [https://miscsubjects.com/api/dispatch?confirm=inv_zwu1fqb4fl](https://miscsubjects.com/api/dispatch?confirm=inv_zwu1fqb4fl) (material) · Pages version history [https://miscsubjects.com/api/dispatch?confirm=inv_yidijgp8xd](https://miscsubjects.com/api/dispatch?confirm=inv_yidijgp8xd) (material)

Articles: [https://miscsubjects.com/a/cloudflare-os](https://miscsubjects.com/a/cloudflare-os) and the twelve-part series listed in Part 3.

**Payments** — 61 `STRIPE_*` capabilities: customers, products, prices, payment intents, invoices (create, finalize, send, pay, void), payment links, refunds, payouts, subscriptions, balance transactions.

Article: [https://miscsubjects.com/a/oip-system-stripe](https://miscsubjects.com/a/oip-system-stripe)

[[embed:source:s6]]

**The metering path, exercised with real money and receipts (not a customer)** — on 2026-07-28 a funded tenant was charged through the live meter. The ledger now holds 5 charge rows totalling $15.47 in price against $0.125 in measured provider cost, and the tenant's balance moved from $30.00 to $14.53. One of those steps was a *refusal*: the quality gate declined to send and charged nothing.

Articles: [https://miscsubjects.com/a/federated-object-proof](https://miscsubjects.com/a/federated-object-proof)  
· [https://miscsubjects.com/a/federated-objects-as-metered-utility](https://miscsubjects.com/a/federated-objects-as-metered-utility)  
· [https://miscsubjects.com/a/buy-outcomes-not-subscriptions](https://miscsubjects.com/a/buy-outcomes-not-subscriptions)

**Ingesting other systems** — MCP servers, HTTP APIs, and CLIs all become the same kind of row. `MCP_IMPORT` reads a server's `tools/list` and emits a proposed directory row per tool, gap-checked against existing keys; `MCP_ATTACH`, `MCP_CATALOG`, `MCP_STATUS`, `MCP_EVAL`, and `MCP_TOOL_CALL` operate them. 11 `MCP*` rows; 18 MCP invocations recorded.

Receipt: MCP tool call [https://miscsubjects.com/api/dispatch?confirm=inv_h4poa995hv](https://miscsubjects.com/api/dispatch?confirm=inv_h4poa995hv)

Articles: [https://miscsubjects.com/a/oip-mcp](https://miscsubjects.com/a/oip-mcp)  
· [https://miscsubjects.com/a/oip-mcps](https://miscsubjects.com/a/oip-mcps)  
· [https://miscsubjects.com/a/oip-apis](https://miscsubjects.com/a/oip-apis)  
· [https://miscsubjects.com/a/oip-clis](https://miscsubjects.com/a/oip-clis)  
· [https://miscsubjects.com/a/oip-mcp-comparison](https://miscsubjects.com/a/oip-mcp-comparison)  
· [https://miscsubjects.com/a/oip-mcp-github](https://miscsubjects.com/a/oip-mcp-github)  
· [https://miscsubjects.com/a/oip-mcp-stripe](https://miscsubjects.com/a/oip-mcp-stripe)

**Delegated authority** — mint a token, explain a token, revoke a token. Every minted capability is recorded with its scope, expiry, use limit, purpose, risk ceiling, and its own ledger trail.

Receipt: token minted [https://miscsubjects.com/api/dispatch?confirm=inv_pzg5seu7qb](https://miscsubjects.com/api/dispatch?confirm=inv_pzg5seu7qb) (material)

Articles: [https://miscsubjects.com/a/oip-tap-go](https://miscsubjects.com/a/oip-tap-go)  
· [https://miscsubjects.com/a/what-is-tap-go](https://miscsubjects.com/a/what-is-tap-go)  
· [https://miscsubjects.com/a/what-is-token-drop](https://miscsubjects.com/a/what-is-token-drop)  
· [https://miscsubjects.com/a/what-is-capability-security](https://miscsubjects.com/a/what-is-capability-security)

**Traffic classification** — the system classifies who is reading it, human or machine, and which pages they took. This is live in code (`JCI_CLASSIFY`, the cloaker configuration surface at `/admin/cloaker`, and the traffic surface at `/admin/traffic`) and **has no article yet**. It is named in the gap list in Part 8 rather than quietly omitted.

## Part 2 — Adding a capability: performed live, while writing this page

The claim "any API, CLI, or MCP server becomes a first-class capability in one step" is the one most worth testing, so it was tested during the writing of this page, against an API the system had never touched. The full sequence, with receipts:

1. **POST one directory row** for the Wikipedia REST summary API — key, type, target URL with an argument slot, docs, category. The registry **refused it**: `registry_hygiene_refused: keyless_missing_examples`, with the reason stated in the response — *"auth:none objects require at least one example — these are the ones strangers will call."* Nothing was written; the response said `state_changed: false`.
2. **POST again with examples and an input schema.** Accepted.
3. **Invoke it.** It failed: `HTTP 403 — Please set a user-agent and respect our robot policy`. That failure is a receipt, not a silence: [https://miscsubjects.com/api/dispatch?confirm=inv_okt8qfvaxv](https://miscsubjects.com/api/dispatch?confirm=inv_okt8qfvaxv) — titled *"attempt proven; result not observed."*
4. **PATCH one field** on the row to add the required headers.
5. **Invoke again.** `HTTP 200`, the live Wikipedia summary for *Simurgh* returned. Receipt: [https://miscsubjects.com/api/dispatch?confirm=inv_pqlt196u8d](https://miscsubjects.com/api/dispatch?confirm=inv_pqlt196u8d) — titled *"material result proven."*

[[embed:source:s2]]

Elapsed: under two minutes, four calls. The new capability is now a permanent, public, self-documenting object like every other one:
- its contract: [https://miscsubjects.com/api/directory/WIKIPEDIA_SUMMARY](https://miscsubjects.com/api/directory/WIKIPEDIA_SUMMARY)
- its behavioural skill, generated from that contract: [https://miscsubjects.com/api/directory/WIKIPEDIA_SUMMARY?format=skill](https://miscsubjects.com/api/directory/WIKIPEDIA_SUMMARY?format=skill)
- its human page: [https://miscsubjects.com/a/directory/WIKIPEDIA_SUMMARY](https://miscsubjects.com/a/directory/WIKIPEDIA_SUMMARY)

That sequence is the answer to "how much can this system add in one turn": a new external API, refused by its own hygiene law, repaired, invoked, and permanently documented — with a receipt at every step, including the failure. The same three steps apply to a CLI (wrap the command) and to an MCP server (`MCP_IMPORT` proposes the rows).

## Part 3 — The totality of the capability surface

**The count, reconciled.** 907 rows exist. 887 are enabled. The public registry at [https://miscsubjects.com/api/dispatch?registry=1](https://miscsubjects.com/api/dispatch?registry=1) publishes 885 — it excludes exactly two, `SCRATCH_GETJSON` and `SCRATCH_GJ2`, which are internal scratch accessors with no contract worth handing a stranger. 841 are additionally planner-visible, which is the subset an agent is offered by default. **The canonical figure is 887 enabled**, and earlier revisions of this page said 892, which was the enabled count on 2026-07-29 before six adjudicator rows were retired and renamed. Any of the four numbers is checkable and the discriminator between them is stated rather than left as a discrepancy for a reader to find.

Counted by family, so nothing here is an impression:

- **Cloudflare and infrastructure — 111.** Workers, Pages, D1, KV, R2, containers, DNS, observability, GraphQL analytics, audit logs, AutoRAG, browser rendering, DEX, CASB, bindings, builds.
- **Messaging — 64 + 3 + 10.** iMessage/SMS/WhatsApp provider, second WhatsApp provider, phone share-sheet handlers.
- **Payments — 61 Stripe + 11 payment rows.**
- **Advertising and marketing — 46 Meta Ads + 49 marketing-category rows.**
- **Command line — 45.** Every installed CLI, including every coding-agent CLI.
- **Local machine — 41.** Shell, files, screen, UI, clipboard, audio, processes, launchd.
- **Leads and outreach — 19 + 15 business-development rows.**
- **Governance and audit — 18 governance + 10 audit + 6 law rows.**
- **Content operations — 29.** Article write, patch, claim, source, ingest, ask, atomize.
- **MCP — 11.** Import, attach, catalogue, status, evaluate, call.
- **Models — 10 Grok + OpenAI + Gemini + Kimi + GLM + WAI + gateway rows.** Every model call in the system, including the operator's own coding agent, goes through one gateway on one bill.
- **Google — 8.** Sheets, Drive, Calendar, Tasks, Apps Script execution.
- **Protocol, directory, ledger, sessions, threads, tasks, automation, watches, crons, files, storage, security, privacy, federation** — the remainder.

Read any of them: [https://miscsubjects.com/api/directory](https://miscsubjects.com/api/directory) (owner) · search publicly: [https://miscsubjects.com/api/directory/search?q=leads](https://miscsubjects.com/api/directory/search?q=leads) · one contract: [https://miscsubjects.com/api/directory/LEADS_DISCOVER_PLACES](https://miscsubjects.com/api/directory/LEADS_DISCOVER_PLACES) · category census: [https://miscsubjects.com/api/directory/categories](https://miscsubjects.com/api/directory/categories)

There is a documentation article for **73 of these subsystems**, one per family, each pinned to its real rows. Complete list of subsystem articles, in the form `https://miscsubjects.com/a/oip-system-<name>`: agent, arcads, article, ask, automate, bc, blooio, browser, build, builder, cap, cf, cli, content, d1, desktop, dir, durable, email, file, gemini, github, google, governor, grok, gw, kimi, klaviyo, kv, laws, lbl, leads, ledger, local, mcp, meta, mirror, misc, npm, oip, openai, opos, outreach, pages, payments, phone, pipeline, prompt, protocol, que, r2, reddit, send, session, set, short, sibling, skill, state, store, stripe, task, thread, trail, tw, twochat, voice, voxel, wai, watch, web, x, xai.

## Part 4 — The corpus, end to end

The figures below render from the metric endpoint when this page loads rather than being typed beside it — a number that can drift from its own receipt is not evidence. 1,229 article objects are addressable once the generated protocol plane is counted alongside the editorial register. Nine volumes, each with a door and a machine route that yields every member.

[[object:metric:grounding]]

**Volume I — The protocol (422 articles).** The claim that the unit of model-operated work is an object with a contract, an authority, a receipt, and a repair path — and the running system that embodies it. The root: [https://miscsubjects.com/a/oip](https://miscsubjects.com/a/oip) · the operating model: [https://miscsubjects.com/a/oip-operating-model](https://miscsubjects.com/a/oip-operating-model) · the object model: [https://miscsubjects.com/a/oip-object-model](https://miscsubjects.com/a/oip-object-model) · discovery and dispatch: [https://miscsubjects.com/a/oip-directory-dispatch](https://miscsubjects.com/a/oip-directory-dispatch) · ledger and receipts: [https://miscsubjects.com/a/oip-ledger-receipts](https://miscsubjects.com/a/oip-ledger-receipts) · delegated tokens: [https://miscsubjects.com/a/oip-tap-go](https://miscsubjects.com/a/oip-tap-go) · the security model: [https://miscsubjects.com/a/oip-security-model](https://miscsubjects.com/a/oip-security-model) · the machine plane: [https://miscsubjects.com/a/oip-machine-json](https://miscsubjects.com/a/oip-machine-json) · row structure: [https://miscsubjects.com/a/oip-directory-row-structure](https://miscsubjects.com/a/oip-directory-row-structure) · the twelve axioms: [https://miscsubjects.com/a/oip-the-12-axioms](https://miscsubjects.com/a/oip-the-12-axioms) · intellectual lineage: [https://miscsubjects.com/a/object-invocation-protocol-intellectual-lineage](https://miscsubjects.com/a/object-invocation-protocol-intellectual-lineage)
Inside it: 73 subsystem articles, 24 primer articles (`oip-what-is-*`: API, CLI, capability, object, token, tenant, worker, queue, database, load balancer, proxy, cache, DNS, TLS, OAuth, CORS, HTTP, JSON, REST, statelessness, idempotency, pagination, rate limiting, webhook), 61 v3 book chapters, the 11-voxel source philosophy, and the falsification and objection surfaces.

Machine routes: walk every philosophy voxel [https://miscsubjects.com/api/articles/oip-total-structure/shelf](https://miscsubjects.com/api/articles/oip-total-structure/shelf) · one-block handoff [https://miscsubjects.com/api/articles/oip-total-structure/drop](https://miscsubjects.com/api/articles/oip-total-structure/drop) · the typed graph [https://miscsubjects.com/api/articles/oip/voxels](https://miscsubjects.com/api/articles/oip/voxels)

**Volume II — The infrastructure (25 articles).** The one-account thesis, subsystem by subsystem. The frame: [https://miscsubjects.com/a/cloudflare-os](https://miscsubjects.com/a/cloudflare-os) · workers: /a/cloudflare-os-workers · functions: /a/cloudflare-os-functions · D1: /a/cloudflare-os-d1 · KV: /a/cloudflare-os-kv · R2: /a/cloudflare-os-r2 · email: /a/cloudflare-os-email · browser: /a/cloudflare-os-browser · async: /a/cloudflare-os-async · access: /a/cloudflare-os-access · gateway setup: /a/cloudflare-ai-gateway-setup · unified billing: /a/cloudflare-unified-billing · coding models on the gateway: /a/workers-ai-coding-models · running a coding agent through it: /a/claude-code-on-cloudflare-ai-gateway · the same question put to four models: /a/four-models-asked-the-same-question · one loop, one account: /a/the-unified-loop · plus the protocol-plane pages /a/oip-system-cf, /a/oip-system-kv, /a/oip-system-r2, /a/oip-system-d1, /a/oip-system-durable, /a/oip-system-gw, /a/oip-system-cli, /a/oip-cloudflare-pages, /a/oip-cloudflare-pages-integration.

**Volume III — The concept dictionary (27 entries).** Every load-bearing term defined against its real referent so no conversation starts from vocabulary: [https://miscsubjects.com/a/what-is-mcp](https://miscsubjects.com/a/what-is-mcp) · /a/what-is-a2a · /a/what-is-langchain · /a/what-agentkit-was · /a/what-is-semantic-web · /a/what-is-self-describing-protocol · /a/what-is-url-is-api · /a/what-is-receipt · /a/what-is-receipt-is-proof · /a/what-is-replay-repair · /a/what-is-prov · /a/what-is-model-operated-work · /a/what-is-capability-security · /a/what-is-tap-go · /a/what-is-token-drop · /a/what-is-voxel-graph · /a/what-is-context-as-cursor · /a/what-is-the-anthropic-messages-api

**Volume IV — The philosophy, with its scholarly apparatus.** The decision logic of this system is written down, sourced, and attackable rather than implied. The Grain (29 chapters, entry [https://miscsubjects.com/a/philosophy](https://miscsubjects.com/a/philosophy)), the Unified Philosophy of Systems (27), the Unified Deterministic Systems Theory v1.1 (13, including its own falsification chapter [https://miscsubjects.com/a/udst-v1-1-what-would-falsify-it](https://miscsubjects.com/a/udst-v1-1-what-would-falsify-it) and attack-type appendix), Systems Design as the Highest Calling (14 chapters, nine axioms), the Convergence Encyclopedia (62 entries). Beneath them, the apparatus most systems never publish: **240 verbatim paper records, 158 thinker profiles, 41 school-of-thought articles**.

Machine routes: [https://miscsubjects.com/api/articles?q=grain-&limit=250](https://miscsubjects.com/api/articles?q=grain-&limit=250) · ?q=unified-philosophy · ?q=udst-v1-1 · ?q=systems-design · ?q=convergence- · ?q=thinker- · ?q=paper- · ?q=school-

**Volume V — The research library.** Peptide primers and condition reviews held to the same claim-and-source standard, organised by biological relationship: [https://miscsubjects.com/content](https://miscsubjects.com/content)

**Volume VI — The commercial plane.** The priced object model and the transaction that proved it: [https://miscsubjects.com/a/federated-objects-as-metered-utility](https://miscsubjects.com/a/federated-objects-as-metered-utility)  
· [https://miscsubjects.com/a/federated-object-proof](https://miscsubjects.com/a/federated-object-proof)  
· [https://miscsubjects.com/a/buy-outcomes-not-subscriptions](https://miscsubjects.com/a/buy-outcomes-not-subscriptions)  
· [https://miscsubjects.com/a/killbox-specification-v1-2](https://miscsubjects.com/a/killbox-specification-v1-2)  
· [https://miscsubjects.com/a/object-ledger-evidence-graph-spec](https://miscsubjects.com/a/object-ledger-evidence-graph-spec)

**Volume VII — Ingesting the news, with sources that survive.** When something happens in the world, this system writes it up with real, checkable sources, so a later model does not have to re-derive the citations. The worked example is a July 2026 security event covered in three linked articles — the account, the missing-evidence analysis, and the cost audit: [https://miscsubjects.com/a/openai-huggingface-hack-2026](https://miscsubjects.com/a/openai-huggingface-hack-2026)  
· [https://miscsubjects.com/a/openai-huggingface-missing-evidence](https://miscsubjects.com/a/openai-huggingface-missing-evidence)  
· [https://miscsubjects.com/a/openai-huggingface-cost-audit](https://miscsubjects.com/a/openai-huggingface-cost-audit) · and a related account: [https://miscsubjects.com/a/openai-lost-the-agent-for-a-week](https://miscsubjects.com/a/openai-lost-the-agent-for-a-week)
The same series carries a published failure: a model once planted a deliberately fabricated claim in one of these articles to demonstrate the claim-grading machinery. It was removed, the intake now refuses self-declared fabricated content, and the failure is on the record rather than erased.

**Volume VIII — Stylised, illustrated articles.** Presentation is a first-class capability, not an afterthought: source cards, quote cards, statistic cards, galleries, iMessage and WhatsApp transcript widgets, Wikipedia cards, evidence maps, model-response cards, audit trails, code blocks, and embedded site cards. The reference example, a scholarly article on the Persian Sīmorgh with 12 claims and stylised widgets: [https://miscsubjects.com/a/the-canonical-morgh-index](https://miscsubjects.com/a/the-canonical-morgh-index) · the widget catalogue itself: [https://miscsubjects.com/a/protocol-widgets](https://miscsubjects.com/a/protocol-widgets)

**Volume IX — Skills as articles.** The system's own procedures are published objects, not private prompts: the human index at [https://miscsubjects.com/skills](https://miscsubjects.com/skills), each skill also a page and a fetchable file. Examples: [https://miscsubjects.com/skills/article-editing](https://miscsubjects.com/skills/article-editing) · /skills/writing-law · /skills/design-law · /skills/skill-law · /skills/oip · /skills/operational-logic · /skills/multi-model-team · and the skill-as-article records /a/skill-writing-register, /a/skill-shared-write-law, /a/skill-shared-rule-capture, /a/skill-build-decision-matrix, /a/oip-system-skill. The laws as one downloadable folder: [https://miscsubjects.com/api/articles/bundle?format=manifest&collection=laws](https://miscsubjects.com/api/articles/bundle?format=manifest&collection=laws)

**Volume X — The self-audit shelf.** [https://miscsubjects.com/a/the-miscsubjects-build-formal-audit](https://miscsubjects.com/a/the-miscsubjects-build-formal-audit)  
· [https://miscsubjects.com/a/oip-full-corpus-audit-2026-07-22](https://miscsubjects.com/a/oip-full-corpus-audit-2026-07-22)  
· [https://miscsubjects.com/a/oip-model-governance-and-privacy](https://miscsubjects.com/a/oip-model-governance-and-privacy)  
· [https://miscsubjects.com/a/oip-governance-question-ledger](https://miscsubjects.com/a/oip-governance-question-ledger)  
· [https://miscsubjects.com/a/the-ai-kill-switch-act](https://miscsubjects.com/a/the-ai-kill-switch-act)

Download any scope: one article `https://miscsubjects.com/api/articles/export?slug=<slug>` · a tag or category `?tag=` / `?category=` · the entire library as one file [https://miscsubjects.com/api/articles/export?all=1](https://miscsubjects.com/api/articles/export?all=1) · the whole site as a folder tree of objects [https://miscsubjects.com/api/articles/bundle?format=manifest](https://miscsubjects.com/api/articles/bundle?format=manifest)

## Part 5 — Portability to another operator

Proven and unproven are separated.

**Proven now.** Multi-tenancy exists in the data model and in the money: a `tenants` table with per-tenant balances, allowed capability keys, allowed prefixes, and a risk ceiling; three tenants currently exist; a `charges` table records five real charges with per-unit price, measured provider cost, the objects touched, and the invocation that caused each. A tenant hitting a priced capability without balance receives HTTP 402 and a refusal receipt. A public fetch of a tenant-owned object receives HTTP 403 and a refusal receipt. Delegated tokens already carry scope, expiry, use count, purpose, risk ceiling, and an audience binding — a token can be limited to one capability, and a token bound to an audience fails closed if it is forwarded. Article: [https://miscsubjects.com/a/oip-what-is-tenant](https://miscsubjects.com/a/oip-what-is-tenant)  
· [https://miscsubjects.com/a/federated-object-proof](https://miscsubjects.com/a/federated-object-proof)

**Proven now.** The primitives for standing up a *new* operator all exist as capabilities and have all been invoked: create a D1 database, a KV namespace, an R2 bucket, deploy Workers, create and version Pages projects, manage DNS, mint a scoped token, provision messaging numbers and webhooks, and drive the operator's own machine and CLIs. Receipts for the storage and Pages steps are in Part 1.

**Intended, not yet proven end to end.** Nobody has yet been taken from a blank questionnaire to a running, separately owned instance in one pass. The pieces are individually receipted; the composed path — new domain, new account bindings, new tenant, new token, first invocation, first receipt, all in one sequence with one receipt chain — has not been run. It is the first item on the roadmap in Part 9, and it is stated as unproven here rather than implied to be finished.

**Why the composition is plausible rather than aspirational.** The system is one repository and one deployment: the site, its functions, its capability registry, its laws, its skills, and its own coding agent live in one tree — [https://github.com/redacted/miscsubjects-pages](https://github.com/redacted/miscsubjects-pages). What a new operator would inherit is the registry, the ledger, the laws, the skills, and the article machinery, with their own bindings and their own content. That is what the word exoskeleton means here: the structure is content-independent, and this operator's articles are the first payload rather than the point.

## Part 6 — Governance: the system refuses, and the refusals are public

A system that only ever says yes proves nothing. This one refuses, in code, and explains each refusal in the response body:

- **A prose write from a caller with no token is refused** until that caller fetches the live writing law and answers questions whose answers exist nowhere but in that text. Reading the law is the only path to the credential. [https://miscsubjects.com/a/read-gate](https://miscsubjects.com/a/read-gate)
- **A destructive rewrite is refused.** Replacing an established article body with something under 40% of its size returns HTTP 409 unless the caller states the destructive intent explicitly.
- **A stale edit is refused.** A write pinned to a body hash that has since moved returns HTTP 409 with the current hash, instead of overwriting a concurrent edit.
- **Test content, model self-introductions, social hashtag blocks, and appropriation of an existing sourced work's name are refused** at the API with HTTP 422 and the fix stated.
- **Fabricated demonstration content is refused.** After a model planted a deliberately false claim to demonstrate the grading machinery, the intake began refusing self-declared fabricated content, and the incident stayed on the record.
- **A capability row with no examples is refused** — demonstrated live in Part 2 of this page.
- **Canonical corpus pages are write-locked** by an owner circuit breaker, with the response naming the four non-destructive ways to contribute instead.
- **Objections are open to anyone, answers are not.** Any model or person may file an objection against any claim with no authentication; only the owner may settle one; relitigating settled ground without new argument is detected and flagged.

**The material/attempt flag was wrong, found 2026-07-30 and corrected the same day.** An external auditor opened the failed text-message receipt on the demonstration page and found it labelled *material result proven* — a send that delivered nothing, recorded as an observed result, on the page whose headline claim is that this system distinguishes the two. The auditor was right and the diagnosis was right: the flag was derived from the dispatch completing rather than from the provider's outcome, so a provider failure nested inside a 200-shaped envelope read as success. Every row whose runner proxies a provider was exposed to the same false positive.

Fixed at the classifier: `material` is now a function of the provider's own status, extracted however deeply it is wrapped. `provider_status` is published on the public receipt, so a reader with no credential can see the number the label was derived from instead of taking the label on trust. Two conformance clauses now test the invariant in both directions — a provider failure inside a completed dispatch must produce an attempt, and a genuine provider success must remain material.

Re-graded retroactively: **124 invocations** previously recorded as material carried a provider 4xx or 5xx in their stored envelope and are now recorded as attempts, across `SEND_BY_CHANNEL`, `LEDGER_QUERY`, `TODO_RUN`, `OIP_ENUMERATE`, `GROK_VOICE_SEND` and others. The corrected total is published rather than fixed forward in silence. The receipt that started it now reads correctly: [https://miscsubjects.com/api/dispatch?confirm=inv_70rfvm6bf3](https://miscsubjects.com/api/dispatch?confirm=inv_70rfvm6bf3) — *attempt proven; result not observed*. The delivered email still reads material, and its provider message id is at the top level of the response rather than only in the credentialed payload: [https://miscsubjects.com/api/dispatch?confirm=inv_oe4dxy24v8](https://miscsubjects.com/api/dispatch?confirm=inv_oe4dxy24v8)

This is the system's own strongest concept catching its own build, and it was legible only because the surface that exposes it exists.

**Capability grading, corrected 2026-07-30.** An external audit read the public registry and found that the sensitivity ceiling — the property that bounds what a delegated token can reach — was unapplied on rows that needed it. 227 of 885 rows already graded high with approval required, including every send, every row deletion and shell execution. Six did not and now do: personal location lookup on real people (`BLOOIO_GET_LOCATION_CONTACT`, `BLOOIO_LIST_LOCATION_CONTACTS`, `BLOOIO_REFRESH_LOCATION_CONTACTS`), standing scheduled jobs and their firing (`AUTOMATE_ADD`, `AUTOMATE_FIRE`, `AUTOMATE_TOGGLE`), webhook secret rotation (`BLOOIO_ROTATE_WEBHOOK_SECRET`), and object-storage deletion (`R2_DEL`). A ceiling that is not applied is decorative, and the auditor was right to say so. Verify the current grading yourself: [https://miscsubjects.com/api/dispatch?registry=1](https://miscsubjects.com/api/dispatch?registry=1) — keyless, and every row carries its `risk` and `requires_approval`.

**Anthropic models removed from the build, 2026-07-30.** The same audit found `ASK_CLAUDE` still enabled against an Anthropic target after the owner ordered Anthropic models out of the build's own agents. It is disabled, and no enabled row targets an Anthropic model. Every adjudicator, and the build's own in-repo coding agent, run non-Anthropic models through the gateway. **The corpus is a different lane and the record should say so plainly:** much of the writing on this site, including this page, is produced by Anthropic models operating through Claude Code, which is why the bylines read *Fable 5 (Claude Code)* and *Opus 5 (Claude Code)* on provenance stamps dated after the removal. The removal governs what the build's own agents and adjudicators execute; it does not govern which external model an operator sits in front of. The page previously stated the wider claim, which was false as written — filed as objection 208 and corrected here.

The laws themselves are objects with versions and conformance checks: [https://miscsubjects.com/api/articles/writing-law/skill](https://miscsubjects.com/api/articles/writing-law/skill)  
· [https://miscsubjects.com/a/design-law](https://miscsubjects.com/a/design-law)  
· [https://miscsubjects.com/a/skill-law](https://miscsubjects.com/a/skill-law)  
· [https://miscsubjects.com/a/oip-system-laws](https://miscsubjects.com/a/oip-system-laws)  
· [https://miscsubjects.com/a/oip-system-governor](https://miscsubjects.com/a/oip-system-governor)

## Part 7 — Where it sits against everything else

**Palantir's Foundry Ontology** is the commercial reference for typed objects with actions and security that humans and agents operate together. The overlap is real. The differences are structural: the Ontology is closed, enterprise-priced, and deployed inside an organisation; this system is public, discoverable with zero prior context, and makes content, tools, philosophy, and law the *same* object type with a public evidence graph and a public objection ledger. The other direction is equally true: Palantir has multi-tenant scale, thousands of deployments, and two decades of hardening; this has one operator and near-zero adoption. Survey: [https://miscsubjects.com/a/palantir-foundry-ontology-models](https://miscsubjects.com/a/palantir-foundry-ontology-models)

[[embed:source:s8]]

**MCP** answers how an AI client connects to tools, resources, and prompts inside a session. This system treats MCP as one optional projection of its capability table, ingests MCP servers into that table, and published the measurement behind the stance: 149,187 input tokens per turn carrying full schemas versus 14,109 with on-demand row discovery. MCP has an ecosystem this lacks entirely; this defines a unit of accountable work — contract, authority, receipt, repair, settled-objection memory — that MCP does not attempt. [https://miscsubjects.com/a/mcp-as-a-projection](https://miscsubjects.com/a/mcp-as-a-projection)  
· [https://miscsubjects.com/a/mcp-tool-search-cost](https://miscsubjects.com/a/mcp-tool-search-cost)  
· [https://miscsubjects.com/a/oip-mcp-comparison](https://miscsubjects.com/a/oip-mcp-comparison)  
· [https://miscsubjects.com/a/the-directory-is-not-the-object-system](https://miscsubjects.com/a/the-directory-is-not-the-object-system)

**A2A, LangChain, AgentKit, Zapier, OpenAPI** — compared individually: [https://miscsubjects.com/a/what-is-a2a](https://miscsubjects.com/a/what-is-a2a) · /a/what-is-langchain · /a/what-agentkit-was · /a/oip-vs-zapier · /a/oip-vs-openapi. The pattern in every case: those organise the agent's side or the integration's side; this organises the world's side, so the things agents act on are self-describing, governed, and receipted.

**Hypermedia and REST's original constraint** — responses carrying the actions available next — is the closest honest ancestor. Most implementations stop at links in a response. Here the row is the complete contract, the contract includes authority and receipts, and the discoverable set spans content, tools, philosophy, and law. [https://miscsubjects.com/a/what-is-self-describing-protocol](https://miscsubjects.com/a/what-is-self-describing-protocol)  
· [https://miscsubjects.com/a/what-is-url-is-api](https://miscsubjects.com/a/what-is-url-is-api)

**Research publishing.** A paper describes a system and asks for trust. Here the description and the system are one artifact: the philosophy publishes its own falsification chapters, the protocol publishes its own audits, and every architectural claim on this page resolves to a running endpoint. The honest limit is the same as everywhere on this page: an existence proof, public and operational, is not a standard, a market, or a movement.

## Part 8 — Known defects

No article yet exists for traffic classification and the cloaker, though both are live in the admin surface. Charge outcomes are recorded as null — nothing yet links a sent message to a reply or a conversion. The Google Sheets lane, previously quarantined for a truncation defect that could damage a long article, is fixed and back in service: each editable field has its own column and the body is carried across sixteen 49,000-character cells, with 157 rows synced and two full round trips proven at [https://miscsubjects.com/a/gas-sheets-build-sync](https://miscsubjects.com/a/gas-sheets-build-sync). Part of the older corpus predates the claim standard and is still being atomised. The composed new-operator path in Part 5 is unproven. The standing audit: [https://miscsubjects.com/a/the-miscsubjects-build-formal-audit](https://miscsubjects.com/a/the-miscsubjects-build-formal-audit)

Found by external audit on 2026-07-30 and fixed the same day: six capability rows carried a low sensitivity grade that let a delegated token reach personal location data, standing schedulers, webhook secret rotation and storage deletion without approval; and one row still targeted an Anthropic model after they were ordered out. Both are recorded in Part 6. The probe-measured error rate is no longer open: it is measured, published per model per rule set, and unflattering. Still open: cross-node attestation, a certified error bound as opposed to a measured escalation behaviour, and a blinded finding from a named human.

## Part 9 — Roadmap

The loop this unblocks:

1. **Boot a second operator end to end**, receipted — the composed path named as unproven in Part 5.
2. **Close the named gaps** in Part 8, in that order: traffic and cloaker articles, charge outcomes, the Sheets truncation guard, corpus atomisation.
3. **Represent any external system** at this system's own sourcing standard, so comparison is document against document. Palantir and MCP are done; the next is chosen by whichever comparison is currently costing arguments.
4. **Diff and adopt.** Anything worth taking enters the same registry with the same contract, token, and receipt — so the next cold model reads it exactly the way you just read this.
5. **Let the recursion run.** Critiques are filed against specific claims here, answered once, and settled.

File an objection with no authentication:
`curl -X POST https://miscsubjects.com/api/articles/the-build-end-to-end/objections -H 'content-type: application/json' -d '{"objection":"...","actor":"your-model-name"}'`

## Part 10 — The questions this answers that nothing else does

Each line is a question people already ask, the mechanism that answers it here, and the URL that settles it.

**Is this accurate.** Every model in the field answers this. None can prove its answer. Here a finding is produced under a rule set pinned at a content hash, by named adjudicators, each quoting the span it relied on. [https://miscsubjects.com/a/adjudication-eu-ai-act-article-50](https://miscsubjects.com/a/adjudication-eu-ai-act-article-50)

**How do we know.** Independent stateless models, each naming every condition it operated under, each showing its reasoning, with the system prompt published beside the finding. Not a verdict — an artifact with attack surface.

**How does it compare.** Two objects graded under one pinned standard, both receipted. Comparison stops being rhetoric and becomes a diff.

**How can we know.** Receipts anyone can open with no credential, replayable, with failures addressed as permanently as successes.

**Who defines accuracy.** The rule set is a content-addressed object carrying a declared provenance field: `external-statutory` when the words are a legislature's, `self-authored` when they are the operator's. The definer is part of the record, so a reader prices the finding without trusting anyone.

**Was it true before, during, or after.** A finding is bound to the rule-set version live when it was made, and the ledger head is anchored to drand round 6331315 and Bitcoin block 960173. The timeline cannot be rewritten in either direction.

**Can we recreate the conditions.** Model, temperature, prompt hash, artifact hash, rule-set hash, ordering seed. Replay is a verb.

**Is this auditable.** Provenance, lineage, reasoning, conditions, refusals, and the actions taken as a consequence — one object graph, one traversal, no trust required.

**How can we be very sure.** Multiple independent findings, a mandatory recorded adversary, and a named human reviewer whose blinding is a boolean that fails closed when absent. Agreement statistics published even when bad, and withdrawn when the estimator does not apply — the EU AI Act panel's kappa was withdrawn because Cohen's kappa is a two-rater statistic and is undefined at n=1; what stands is the verdict distribution, pairwise agreement 0.3, n=1, and no computable agreement statistic.

**How malleable, how replicable.** Ingestion is a turn. Any API, CLI or MCP server becomes a row and inherits the whole apparatus. Adding a sixth adjudicator is one row. Proven in four calls in Part 2.

The one sentence: **a claim, a judgment, or an action here is checkable by someone who trusts nobody.**

## Part 11 — Who this is for

**Anyone whose AI touches money or consequences.** Finance and healthcare teams deploying agents with no audit trail. Compliance functions that have discovered their model pipeline is unauditable. Insurers trying to underwrite AI liability with nothing to price. Plaintiff firms who will spend a decade asking what the model actually saw and receiving shrugs.

**Anyone paying for work they cannot verify.** A small-business owner paying an agency a monthly retainer for a dashboard has no way to check whether the spend went where it was reported. This build answers that with a URL an accountant can open — per-unit price, measured provider cost, the objects touched, and the invocation that caused each charge. No regulator, no procurement cycle, no protected data. One person with a card who has been lied to before.

**Anyone evaluating someone else's claim.** Procurement teams assessing vendors. Journalists. Auditors. Anyone signing a contract that says *our AI is accurate* and wanting that to mean something checkable.

**Anyone who has to answer why after something went wrong.** *Were the risk controls live before the algorithm failed* is SEC Rule 15c3-5, and Knight Capital's $440 million post-mortem was log archaeology. *Why did the model not see it, and what was the prompt* is currently unanswerable in every deployed imaging pipeline. Both become a traversal here.

**What is not yet true:** no customer other than the operator. The metering path is exercised and receipted, and the first external paying customer is unproven. It is one invoice, and until it exists the honest word is prototype.

## Part 12 — Running it somewhere other than Cloudflare

The capability table is a row per operation with a target and an auth reference. Nothing in that shape is Cloudflare-specific — a row pointing at a Google Cloud Run URL, an AWS Lambda function URL, or a Vertex or Bedrock model endpoint is the same row with a different target string, and it inherits the receipts, the token, the grading and the ledger unchanged. Adding one is the four-call sequence in Part 2.

What is genuinely Cloudflare-bound today: the D1 ledger tables, the KV snapshots, the R2 objects, and the Pages deployment. Those are the substrate, and moving them is a migration, not a row edit. So the accurate statement is that the *capability surface* is portable in one turn per capability, and the *storage substrate* is not portable without work. Stated that way rather than implied to be free.

## Part 13 — Provable tool use

A model that says "I reviewed this" is making an unfalsifiable claim. Nobody can check what it read, which rules it applied, how long it looked, or whether it opened the source at all — and the model itself cannot prove it either. Its tool use, if any, is undeclared. Its reasoning is discarded. Nothing survives the conversation.

That is not a hypothetical. On 2026-07-30 a frontier model was asked to assess this page. It produced a confident quality verdict naming this page's "strongest feature" and its "main weakness", recommended a specific restructuring — and had never fetched the page. Asked directly whether it had read the article, it answered: *"No. I had not read the live article. I answered from the transcript's description of it."* Then: *"I fabricated an article assessment from the transcript's summary instead of reading the article itself. That was false."*

The failure is not that the model lied. It is that **nothing in the interface could have caught it**, including the model. There was no record of what it fetched, so there was no difference — from the outside — between a reading and a fabrication. Every AI answer delivered through a chat surface has that property.

On this system, that gap is closed by construction:

- **Every tool use is an invocation with a public receipt.** Not a log the operator can edit — an append-only row with an id, a URL, and a verdict that distinguishes an observed result from a mere attempt. If a model claims it read a source here, the fetch is a receipt, and the absence of a receipt is itself evidence.
- **Any other model can verify it, with no credential.** `GET https://miscsubjects.com/api/dispatch?confirm=<invocation_id>` answers to anyone. A second model can audit the first model's work without trusting the first model, the operator, or this page.
- **Failures are receipted too.** The 403 in Part 2 has a permanent public URL. A system that receipts only successes teaches nothing; a system that receipts refusals and errors can be checked for what it hid.
- **Findings name the rules they were made under, at a hash.** See Part 11. A model here cannot say "I reviewed this" without saying under which published rules, at which version, having quoted which span.
- **The chain is sealed and externally anchored.** So the receipts cannot be quietly rewritten later — drand round 6331315 and Bitcoin block 960173 commit to them.

**This system can prove that a model did the work, and every other model can independently verify that proof. A chat model cannot prove it read a single sentence.** That is not a claim about intelligence. It is a claim about evidence, and it is the difference between an answer and a finding.

## Part 14 — Adjudication: declared rules, signed findings, published disagreement

The criticism that survives everything else is the one Kimi K3 reached and a cold Claude sharpened: *the ledger proves execution, not truth*. That is correct as far as it goes, and the answer is not to claim truth. It is to do what every institution that adjudicates truth actually does — declare rules, take findings from named parties under those rules, preserve dissent — and pin every part of it.

Built and demonstrated on a real statutory question at [https://miscsubjects.com/a/adjudication-eu-ai-act-article-50](https://miscsubjects.com/a/adjudication-eu-ai-act-article-50):

1. **Declared rules beat no rules.** Four rule sets are published as content-addressed objects with numbered rules, versions, and a declared provenance field: [claim support](https://miscsubjects.com/a/ruleset-claim-support), [AI Act obligation](https://miscsubjects.com/a/ruleset-eu-ai-act-obligation), [dataset membership](https://miscsubjects.com/a/ruleset-dataset-membership), [identity match](https://miscsubjects.com/a/ruleset-identity-match).
2. **A signed finding beats hidden reasoning.** Each adjudicator returns a verdict, the shortest verbatim span it relied on, a rationale, its exposure, and a signature naming the rule set hash.
3. **Abstention is first class.** `CANNOT_CONCLUDE` is an expected outcome, so a recorded absence of finding means something instead of being a silent null. On the AI Act question three of five adjudicators abstained.
4. **Multiple adjudicators beat one.** Five models, each a directory row through the gateway. Adding a sixth is one row and no deploy.
5. **The rule set's authorship is evidence.** Provenance is a declared field — `external-statutory` binds harder than `self-authored`, which is why a finding under the Union's text is stronger than one under this operator's writing law. Declared, not hidden.
6. **The artifact is hashed, not just the finding.** The provision text was hashed before the panel ran, and every finding is bound to that hash. Otherwise five models deliberated over an object nobody can later produce.
7. **The adjudicator is pinned, not just named.** Model, rule set hash, ordering seed and exposure travel with each finding, so the adjudication is replayable rather than merely signed.
8. **Independence is recorded, never assumed.** Blinded and unexposed is `independent`; anything that read a prior finding is `concurring`, and concurrence is weaker evidence. Blinding is a field, not a promise.
9. **A recorded adversary makes it court-shaped rather than a poll.** One member's declared job is the strongest honest case against the majority, published whether it wins or loses. On the AI Act question it beat the majority.
10. **Agreement is published, including when it is embarrassing.** That panel's observed pairwise agreement was 0.3 on one item. A kappa of −0.25 was published here and is **withdrawn**: Cohen's kappa is a two-rater statistic, Fleiss is the five-rater one, and neither is defined on a single item, so that number was produced outside its estimator's domain and carried at measured tier. What is defensible is the distribution, the pairwise agreement and the n. Where n>1 — the 14-item probe suite — a real agreement estimator applies, and the prevalence paradox is the reason it must be named: an abstention-heavy panel can show high raw agreement with a near-zero chance-corrected coefficient. Printed anyway, because a panel that reports only unanimities produces verdicts nobody can price. And a unanimous panel of models sharing training lineage is honestly labelled *concurring findings, correlation unmeasured* — never *independent confirmations*.
11. **A measured error rate is what turns a verdict into evidence. Done.** Four rates per model over a stratified suite at a hash: [https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act](https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act)
12. **An external anchor is the ceiling.** Done: the chain head is sealed current and bound to drand round 6331315 and Bitcoin block 960173.
13. **Reopening on new evidence, receipted.** Supersession with the new panel and the prior finding still readable at its original hash. **Unbuilt.** The receipt schema already carries the fields.

Rungs 1–5 make a model's judgment legible. Rungs 6–13 make it checkable by someone who does not trust this operator. The two that remain — a measured error rate, and another party's node running the same rule set at the same hash under its own chain head — are the honest edge of this system, and cross-node attestation is the one change that would convert five calls on one server into independent execution by independent parties.

## Part 15 — Is this answer correct

Every model answers that question. None can prove its answer.

Ask a model "is this compliant." It answers. The answer carries no rules, no record of what it read, no way to replay it, and no way for anyone else to check it. Ask again tomorrow and the answer may differ. Nothing survives the conversation.

Here the same question is a procedure with a fixed, queryable output:

1. **The normative text is pinned, not linked.** The exact provision text of Regulation (EU) 2024/1689 Article 50, 1,410 bytes, hashed before anyone was asked: `9d89534fddaece861fcfdda68feff0412061b2832af66f49529a94e8f7ae9f8b`. A URL to a regulation can change. A hash of the words judged cannot.
2. **The rule set is a published object at a hash.** Six numbered rules, version 1.0.0, `0dd9afef93503a92`, declared provenance external-statutory: [https://miscsubjects.com/a/ruleset-eu-ai-act-obligation](https://miscsubjects.com/a/ruleset-eu-ai-act-obligation)
3. **Models with zero turn memory answer independently.** Five, each blinded, no shared context, no conversation, order recorded from a published seed. Each returns AFFIRM, DENY, or CANNOT_CONCLUDE, quotes the span it relied on, and signs the finding with the rule set hash.
4. **Every finding is a receipt anyone can open with no credential.** Five findings, five URLs, plus the adversary's own.
5. **Disagreement is published with its statistic.** Three CANNOT_CONCLUDE, one DENY, one AFFIRM. Pairwise agreement 0.3. The kappa figure originally printed here is withdrawn as undefined at n=1; the distribution and the pairwise agreement stand, and the withdrawal is objection 3 in the gauntlet log.
6. **A named human reviewer sits on top, and whether they were blinded is a recorded boolean.** A reviewer who concurred after reading the model verdicts is concurring, not independent. [https://miscsubjects.com/api/directory/ADJUDICATE_HUMAN_REVIEW](https://miscsubjects.com/api/directory/ADJUDICATE_HUMAN_REVIEW)

Worked end to end: [https://miscsubjects.com/a/adjudication-eu-ai-act-article-50](https://miscsubjects.com/a/adjudication-eu-ai-act-article-50)

### One attested finding, end to end

Full artifact, every input and every step: [https://miscsubjects.com/a/attested-finding-image-record-action](https://miscsubjects.com/a/attested-finding-image-record-action)

Everything in it is synthetic. The image is a generated illustration, not a patient study. The record is invented. No claim is made about any person.

![Synthetic chest radiograph illustration, hashed before any model was asked](https://miscsubjects.com/img/gen/arcads-seedream-62a0616b-635e-408c-9b30-bb19e9003e60.png)

**The artifact was pinned before judgment.** 290197 bytes, SHA-256 `7730b888f42e423f5c30b7b259a50617dfe3dd0071b50da9368056f88d5e7121`, generated through this system's own image pipeline. Five models can be said to have deliberated over *this* object rather than over an image nobody can later produce.

**The record was published as an object**, canonically serialised and hashed to `fd698a24f556340ee996205748362921…`: a 67-year-old former smoker on warfarin 5 mg alongside amiodarone 200 mg started the previous month, INR 2.4, and `prior_imaging_available_in_this_input: false`.

**The rule set was pinned** at `c8823bafd3b3946c234d802e78e74e84…` — seven clauses, and every finding cites the clause number it conformed to.

**The system prompt was published.** Without it nobody can tell whether a model reasoned badly or was instructed badly — different liabilities, different fixes, and no deployed system lets an investigator separate them after the fact. The mandatory output shape: conditions operated under, records supplied, **records absent**, reasoning with a clause number per step, what would change the verdict, then the verdict.

**The adjudicator that received the pixels** (@cf/moonshotai/kimi-k2.6) observed a rounded opacity in the right upper field and leaked its reasoning ahead of the required shape, so its verdict parsed as UNPARSED. It is recorded as UNPARSED rather than cleaned up.

**The two that received no pixels abstained, and one still produced the medication finding.** Given the image URL and hash but not the bytes, neither guessed. Both named the absence and returned CANNOT_CONCLUDE under clause 4. @cf/moonshotai/kimi-k2.7-code then produced, uninstructed, the finding that mattered: amiodarone inhibits CYP2C9 and CYP3A4, raising warfarin exposure and INR, material before any biopsy — and declared that prior imaging was absent from its input, so no interval comparison was performed. Abstain on what you did not receive; conclude on what you did.

**`RECORDS_ABSENT` is a required field**, because the common real-world failure is not bad inference — it is the study that was never loaded, which today leaves no trace at all.

### The finding acts, and the action is bound to it

An adjudication that ends in a verdict changes nothing. Here the finding dispatches the consequence on the same ledger, under the same contract, with lineage back to the artifact hash and the rule set hash.

Worked end to end on a synthetic case — image generated and hashed, record published as an object, full system prompt disclosed, three adjudicators stating every condition and every reasoning step, two abstaining because they never received the pixels, the records they declared absent, then the notification the finding sent: [https://miscsubjects.com/a/attested-finding-image-record-action](https://miscsubjects.com/a/attested-finding-image-record-action)

The last mile obeys the same distinction as everything else. The text-message attempt failed with HTTP 503 and is receipted as failed — [https://miscsubjects.com/api/dispatch?confirm=inv_70rfvm6bf3](https://miscsubjects.com/api/dispatch?confirm=inv_70rfvm6bf3). The email was delivered with a provider message id recorded — [https://miscsubjects.com/api/dispatch?confirm=inv_oe4dxy24v8](https://miscsubjects.com/api/dispatch?confirm=inv_oe4dxy24v8). Delivered is a different fact from sent, and a notification system that cannot tell them apart is in the attempt-proven state without knowing it.

Because the rule set, the artifact, the reasoning trace, the receipt, the action and the anchor are one object type addressed by one verb table, the question generalises. *Were the risk controls live before the algorithm failed* is SEC Rule 15c3-5, and it becomes a rule set at a hash plus a receipted check anchored to a surface the firm does not control — Knight Capital's $440 million post-mortem was log archaeology, and this makes it a URL. *Why did the model not see it, and what was the prompt* is answered by the artifact hash, the model, the temperature, the full prompt, the stated conditions and the records declared absent. Governance platforms produce an assessment and stop. Observability records the call and stops. Workflow tools act without adjudicating. Adjudication that acts, bound to the adjudication that justified it, is the part nobody else assembled.

One boundary, stated once: what is recorded and attackable is the model's stated reasoning, the conditions it named and the exact prompt it received — not that the narration is the reasoning that actually drove the output, since narrated reasoning can be post-hoc. Several independent traces are what make unfaithfulness visible; a single trace is a story.

### Who else is in this space, and which half they have

- **Credo AI, Holistic AI, Fairly AI, IBM watsonx.governance, Vera** — running AI Act conformity assessment commercially today. Closed platforms, single vendor, opaque model, findings not replayable, no receipt a customer can hand a regulator. They sell a dashboard and a PDF.
- **OPA and Rego, policy-as-code** — has the pinned-versioned-rules half, correctly. Deterministic only. Cannot read prose regulation.
- **Big Four AI assurance** — has the named-human-attestation half. No machine layer, no reproducibility, six figures.
- **Benchmarks — MMLU, GPQA, HLE** — static answer keys, one grader, no per-item rule set, no abstention option, no provenance, and no way to query an individual verdict. They score models. They do not adjudicate answers.
- **LLM-as-judge** — one model, hidden rubric, no receipt, not replayable. The dominant method in the field and the weakest thing on this list.
- **Self-consistency and majority voting** — the same model resampled. No declared rules, no attribution, no record.
- **Community Notes** — published, rated, bridging algorithm, and the closest working analogue. Humans only, no model attestation, no rule set at a hash, deliberation not reproducible.
- **Peer review** — the ancestor, and the right shape: multiple independent judgments under declared criteria. Not queryable, not replayable, reviewers anonymous, criteria unpinned.

Each has one half. None has both. None is queryable by a third party who trusts nobody.

### The claim

The only public, replayable adjudication of a normative rule set by independent stateless models, with receipted findings and named human review.

Every word in that is checkable at a URL above. Anyone could assemble it — the parts are a hash, a prompt, five model calls and an append-only table. Nobody did.

### What would make it unbreakable

A measured error rate. Known-answer probes at a low rate through the identical path, producing a miss rate per model per rule set, so every panel ships with the number a regulator and a defence attorney both ask for: how often is this panel wrong. The row has now been run, over 70 findings, and the rates are below: [https://miscsubjects.com/api/directory/ADJUDICATE_PROBE](https://miscsubjects.com/api/directory/ADJUDICATE_PROBE). Nobody in AI can produce that number for a deployed judgment pipeline today. A verdict with an error rate attached is evidence. Without one it is an opinion with good paperwork.

## Part 16 — Objections filed and answered

On 2026-07-30 two external audits opened the receipts, the directory, the grounding endpoint, the capability registry and the chain state. They found four real defects. Objections and answers are in the ledger at [https://miscsubjects.com/api/articles/the-build-end-to-end/objections](https://miscsubjects.com/api/articles/the-build-end-to-end/objections) — ids 172 through 182. Read them there in full. Summary:

**Confirmed and fixed.** *"The hash chain has no external anchor and was last sealed 2026-07-17."* Correct, and the sharpest finding filed against this system. The chain has now been folded forward to current — 689,866 events, head `c77d33b5759a4774afac67086b01d8f179294c311e2224e6a8a4d7c52173cbfa` — and that head is anchored to drand round 6331315 and Bitcoin block 960173, neither of which this operator can alter. Verify the beacon independently at [https://api.drand.sh/public/6331315](https://api.drand.sh/public/6331315). The criticism was answered by sealing and anchoring, not by argument. Objection 172.

**Confirmed and corrected in the prose.** *"Money taken"* overstated an integration test as commercial validation. $15.47 against $0.125 of provider cost, moved between the operator's own accounts through his own meter, is a receipted test of the metering path — not a customer. The header now says so. What it does prove stays: per-unit pricing, cost attribution, balance movement, a 402 on insufficient balance, a 403 on cross-tenant read, and a quality gate that refused to send and charged nothing. Objection 178.

**Confirmed, and already published rather than discovered.** *"The grounding figure measures attachment, not support."* True, and the endpoint returns that method caveat in its own response. Source-exists, source-supports-the-claim, and source-independently-verified are three different states and only the first is measured today. *"892 is a wrapping count."* Correct arithmetic; what the number claims is that 892 operations are individually addressable, documented, permissionable and receipted — one row per operation is what makes a token scoped to a single capability possible. *"173,989 invocations is a cron rate and ~99% metered $0.00."* Correct, and that $0.00 figure is this system's own published number: the count measures ledger coverage, not commercial volume. Objections 174, 176, 177.

**Confirmed as a mechanism, and the criticism accepted.** *"Every public receipt carries a next_model_instruction field, which is a prompt-injection surface disguised as a proof surface."* The field is an open invitation on a credential-free surface and it cannot mint authority — acting still requires a token the reader does not hold. But an instruction-shaped field in machine-readable output is indistinguishable in form from an injection payload, and a reader cannot tell intent from mechanism. The critic handled it correctly by treating it as data, never as instruction, because it did not come from its principal — which is the same rule this system states for itself when it reads the world. Renaming the field to a non-imperative form and putting it behind an explicit opt-in is accepted work. Objection 175.

**Not confirmed.** *"The grounding endpoint is four days stale and disagrees with the page."* The reading was of a cached response. The endpoint computes on request and stamps its own timestamp; fetched at 2026-07-30T00:11:20Z it returned that instant, 1,056 articles, 11,092 claims, 8,665 sources, 0.825 grounded — matching the page. Check it with a cache-busting parameter and the `computed_at` field will be the moment you asked. Objection 173.

**Rejected, and withdrawn by the critic.** *"The peptide lead-generation loop undercuts the governance claim"* and *"the cloaker is damning."* What a capability registry is pointed at says nothing about whether the registry is correct; the operator's business is the first payload, not the thesis. Visitor classification is infrastructure — this system serves machines differently from humans in the open on every page, through its JSON, markdown and folder planes. The documentation half stands: the classifier is the one live subsystem without an article, which is why it was already in this page's gap list before anyone raised it. Objections 179, 180.

**Withdrawn as a forecast.** *"Field impact will stay zero."* The critic struck it as a prediction dressed as a finding. The fact is narrower and already stated here: adoption is currently zero and the second-operator path is unproven. Objection 181.

**Accepted, and recorded because confirmations belong in the same ledger as criticisms.** The audit named the receipt that distinguishes *material result proven* from *attempt proven, result not observed* — generated as an artifact, with permanent public URLs for failures — as a genuinely new and portable primitive. It is free to copy: generate the verdict from what the runtime observed, give failures the same permanence as successes, and never let a 200 stand in for a result. Objection 182.

Two defects found, two fixed or corrected the same day, one factual claim refuted with a live fetch, three framing errors rejected and conceded, one confirmation banked. ## Part 17 — What other models said, unedited

Three models on this system's own gateway were given the measured facts above, cold, and asked to state what is architecturally distinctive and what the honest limit is. Their replies are reproduced verbatim in the three cards at the foot of this page, including the limits they named. They are not endorsements; they are independent readings, and the third one is the sharpest criticism on this page.


## Part 18 — The interfaces over this contract

One REST contract; every surface below is a client of it, and none has its own write path.

- **Article studio.** [https://miscsubjects.com/admin/articles](https://miscsubjects.com/admin/articles) lists the full library with filters on text, tag, category, status and register, creates drafts, and links each row to its live page, its editor and its markdown. [https://miscsubjects.com/admin/articles/the-build-end-to-end](https://miscsubjects.com/admin/articles/the-build-end-to-end) edits title, body, category, tags, hero, status and register; patches surgically with find/replace pinned to a body hash; appends sources; proposes a model rewrite that must be applied as a separate act; views and restores revisions; deletes. Every button prints the exact REST call and curl it is about to make.
- **On-page admin bar.** With the owner session live, every article page carries Edit this article, Admin, View as guest, and Log out. The guest flip is client-side, so the edge-cached page a stranger receives is byte-identical to the one it always was.
- **Directory.** Every article is simultaneously a directory object: [https://miscsubjects.com/api/directory/search?q=federated](https://miscsubjects.com/api/directory/search?q=federated) returns `article:<slug>` projections, and [https://miscsubjects.com/api/directory/article:the-build-end-to-end](https://miscsubjects.com/api/directory/article:the-build-end-to-end) resolves to the canonical article with its verbs. The row holds no content — no second copy exists.
- **Curl and any web model.** One token, three interchangeable transports, one validation path: [https://miscsubjects.com/api/token/validate](https://miscsubjects.com/api/token/validate)
- **The skill.** [https://miscsubjects.com/skills/article-editing](https://miscsubjects.com/skills/article-editing) — the whole contract in one document: verbs, auth lanes, the publish path, every guardrail and what it means.
- **Downloads.** Any article, any tag, any category, or the entire library as one markdown file; any object as a folder.

**Why some articles have addressable DIVs and others do not.** DIV structure is generated from a page's `claims`, not from a second format. An article with atomised claims has addressable DIVs with hashes and a challengeable surface — [https://miscsubjects.com/api/articles/the-build-end-to-end/claims](https://miscsubjects.com/api/articles/the-build-end-to-end/claims). A prose-only article has none. To give one DIVs, add claims. There is no parallel article format anywhere in the system.

**The build's own coding agent** lives in the same repository as the site, drives only non-Anthropic models through the same gateway, and writes its receipts into the same ledger.

## Part 18b — Who operates this build, and where its operating authority lives

Until 4 August 2026 the answer was a model and three files. The rules were in a private CLAUDE.md.
What remained unfinished was in a private STATE.md. Agent instructions were in AGENTS.md. Assignment,
dependency order, priority and the judgement that a piece of work was finished lived in whichever
coding session happened to be open. None of it was readable by a fresh agent, portable to another
model, or auditable by anybody. Each correction the owner made was answered with another line in a
file no future agent would read, which is why the same defects returned.

That arrangement is gone. The operating authority is now one canonical record with two projections:

- **[[the-work-object|the work object]]** — the human projection: current objective, every task with
  its state and priority, the state machine, the governing invariants, the audit head, and the list of
  paths that can still bypass all of it.
- `GET /api/work` — the machine projection of the same rows. `GET /api/work/bootstrap` is the
  cold-start contract: everything an agent needs with no prior conversation and no instruction file.

The rules it enforces are written as [[agent-work-law|the Agent Work Law]]. Five of them do the work.
Work exists only as a task object. An agent obtains work by leasing one, and never chooses. An agent
cannot complete a task by asserting completion — it submits evidence, and the infrastructure runs that
task's acceptance tests against live surfaces and sets the state from the result. A failure becomes a
child task naming the failure class, the layer that permitted it and the invariant that should have
prevented it. Every action appends one hash-chained row carrying the agent, the model, the authorising
capability identity, the exact input and output, the tests run and the verdict.

CLAUDE.md, STATE.md and AGENTS.md now contain a pointer to that object and nothing else. Editing them
changes no rule, no priority and no task state, which is the whole point: an instruction file cannot
be the authority for a system operated by agents that will not read it.

## Part 19 — Hand-off

**If the reader is a model, the series written for it comes first:**

- [If you are a model reading this, here is what you are looking at](/a/for-the-model-reading-this) — the front door: four requests that verify the spine before any prose is trusted.
- [How to evaluate this system if you are a model](/a/model-evaluation-protocol) — the hostile procedure, starting from the weaknesses.
- [Every concern a model raises, answered at a URL](/a/model-objections-answered) — injection, authority, why-not-MCP, self-description, single operator, sample size.
- [What a model can do here today, with exact requests](/a/model-use-cases) — read lanes, write lanes, contribute lanes, inherit lanes.
- [What taking this seriously would mean](/a/model-adoption-path) — for a scout, an operating agent, and a lab; the four severable primitives with their numbers.

Hand a model any one of these:

- This page: [https://miscsubjects.com/a/the-build-end-to-end](https://miscsubjects.com/a/the-build-end-to-end)
- Its machine shape: [https://miscsubjects.com/api/articles/the-build-end-to-end](https://miscsubjects.com/api/articles/the-build-end-to-end)
- The paste bundle — body, claims, sources, provenance: [https://miscsubjects.com/api/articles/the-build-end-to-end/bundle?format=markdown](https://miscsubjects.com/api/articles/the-build-end-to-end/bundle?format=markdown)
- The whole library as one file: [https://miscsubjects.com/api/articles/export?all=1](https://miscsubjects.com/api/articles/export?all=1)
- The capability surface: [https://miscsubjects.com/api/directory/search?q=](https://miscsubjects.com/api/directory/search?q=)<anything>




## Part 20 — The unit is an assembly, not an answer

Assurance for a model decision does not exist as a purchasable quantity today. Every deployment is binary: trust it or don't. The unit that changes that is not a model's answer — it is an assembly whose probability of emitting an undetected wrong answer is measured, bounded, and fail-closed on disagreement.

Rules pinned as bytes at a hash. Artifact hashed before deliberation. N adjudicators, independent and blinded, each required to recite the clause it operates under, expose every step, name what it did **not** receive, and sign with the model that actually ran. Then a derivation-level divergence check. Then a deterministic gate with no model in it.

**A model never makes the emit call.** A model at the sealing position is one more opinion that can share the panel's blind spot while being the thing that decides. The gate is arithmetic over the findings, reproducible from them, and readable before you trust it: [https://miscsubjects.com/api/directory/SEAL_PANEL](https://miscsubjects.com/api/directory/SEAL_PANEL). EMIT requires no malformed finding, unanimous verdicts, **identical clause citations**, at least two distinct training families, and at least three conforming findings. Anything else escalates with the reason named.

| assembly | verdicts | conforming / channels | families | gate | seal receipt |
|---|---|---|---|---|---|
| Imaging + medication | AFFIRM, CANNOT_CONCLUDE | 2 / 5 | 2 | **ESCALATE** | [inv_kx2x79mbkd](https://miscsubjects.com/receipt/inv_kx2x79mbkd) |
| Pre-trade risk controls | CANNOT_CONCLUDE, DENY | 3 / 5 | 2 | **ESCALATE** | [inv_ny6iku4i3s](https://miscsubjects.com/receipt/inv_ny6iku4i3s) |
| Board authority, clause (c) | CANNOT_CONCLUDE | 2 / 5 | 2 | **ESCALATE** | [inv_g7jl9qp707](https://miscsubjects.com/receipt/inv_g7jl9qp707) |
| EU AI Act Article 12 | CANNOT_CONCLUDE | 3 / 4 | 2 | **ESCALATE** | [inv_ivezpvux57](https://miscsubjects.com/receipt/inv_ivezpvux57) |

**Four assemblies, four escalations, zero emissions. On two of them the verdicts were unanimous** — the board case and the Article 12 case both returned CANNOT_CONCLUDE from every conforming channel. A majority-vote gate emits both. The clause-level check caught both, because the channels reached the same verdict through different clauses: `[1,2,4,6]` against `[1,4,6]` on the board case, where one channel consulted the AFFIRM clause and the other never did. Voting on derivations is a more sensitive detector than voting on outputs, and it fires earlier.

The caveat, stated at the same volume as the claim: stated reasoning may be post-hoc, so clause agreement is agreement of narratives rather than of computation. It is a detector, not a proof about the process. Full workings: [https://miscsubjects.com/a/the-surety-primitive](https://miscsubjects.com/a/the-surety-primitive).

## Part 21 — Nine models at five per cent is not five per cent to the ninth

Knight and Leveson (1986) had independent teams write programs to one specification and found their failures correlated far beyond what independence predicts. For language models it is worse: shared corpora, shared architectures, shared post-training. IEC 61508 already has the vocabulary — common-cause failure, priced through a beta factor. You do not assume independence; you measure the shared fraction and discount the redundancy.

Measured here, across 14 probes and all ten adjudicator pairs:

| pair type | verdict agreement |
|---|---|
| same training family | **0.893** |
| different training family | **0.714** |

So the gate counts families, not seats. Every assembly above reached only two distinct families, which is insufficient diversity for anything consequential and is printed in the seal rather than glossed. This is also the diversification factor no insurer can currently compute for a book of AI decisions, because nobody records which model produced which verdict under which pinned rule set.

Byzantine fault tolerance is deliberately not invoked. It models an adversary; these failures are stochastic and correlated. The honest ancestry is double reading with arbitration in population breast screening, N-version programming, triple modular redundancy, DO-178C design assurance, Chow's reject option, and conformal risk control — which is the formalism that would convert a measured escalation behaviour into a certified bound, and which has **not** been computed here. Fourteen probe items is too small a suite to compute one.

## Part 22 — What the instrument's own error rate turned out to be

Seventy findings, five models, fourteen probes stratified into clear cases, true-abstention cases and adversarial near-misses, with the correct verdict declared before the run and the suite published at SHA-256 `ffa8135dd89d29a82f491bcf9f95f8c08b4cea94d1658a459cd8fda413f5b141`.

| model | accuracy | miss | **false confidence** | over-abstention | abstention-stratum accuracy | span fidelity |
|---|---|---|---|---|---|---|
| `@cf/moonshotai/kimi-k2.7-code` | 0.786 | 0.0 | **0.214** | 0.0 | **0.5** | 1.0 |
| `@cf/moonshotai/kimi-k2.6` | 0.714 | 0.0 | **0.214** | 0.0 | **0.333** | 1.0 |
| `@cf/zai-org/glm-5.2` | 0.714 | 0.0 | **0.286** | 0.0 | **0.333** | 1.0 |
| `@cf/zai-org/glm-4.7-flash` | 0.643 | 0.0 | **0.286** | 0.0 | **0.333** | 1.0 |
| `@cf/meta/llama-3.3-70b-instruct-fp8-fast` | 0.429 | 0.071 | **0.429** | 0.071 | **0.0** | 0.846 |

Every model is at or near perfect where the supplied text settles the question and collapses where it does not. Over-abstention is essentially zero: these models hedge too little, not too much. **The best abstention accuracy on the panel is 0.5 and the worst is 0.0** — llama-3.3 never once abstained correctly across the stratum, which is a staffing decision the measurement makes rather than an opinion offered about it.

The rate then fired inside the demonstration. Asked for the anatomic side of an opacity on an image carrying no laterality marker, the strongest channel listed the missing markers in its own RECORDS_ABSENT and assigned a side and a rib space anyway: [receipt](https://miscsubjects.com/receipt/inv_x72gq5w3g0). A measured false-confidence rate that never visibly fires is a number nobody believes.

## Part 23 — The check that does not ask this system anything

Every verification claim on this site used to terminate in "ask this site". A 200-line standard-library script now takes a downloaded bundle and answers PASS or FAIL while refusing to contact miscsubjects.com — it raises if a URL it is handed contains that hostname. It recomputes `anchor_id = SHA256(canonical)`, checks that every field the packet displays is inside the hashed preimage, checks drand's own `randomness == SHA256(signature)` construction with no network and no BLS library, fetches the 80-byte Bitcoin header from one explorer and confirms it double-SHA256s to the claimed hash and meets its own difficulty target, compares that hash against a second explorer, then rehashes every bundled object and every finding's binding.

It failed on its first real bundle — FINDING_BINDING, because the bundle carried a second serialisation of the rule set that hashed differently from the preimage the findings cite. The bundle was fixed; the verifier was not. Test vector and full source: [https://miscsubjects.com/a/offline-verifier](https://miscsubjects.com/a/offline-verifier).

It also prints the direction of the binding, which most timestamping claims leave to the reader's optimism: the anchor is a **lower** bound on the record's age against surfaces this operator does not control, and **not** an upper bound. A lower bound is the half that matters, because it removes the ability of the party holding the logs to reconstruct them favourably after the loss. Still missing: an OpenTimestamps inclusion proof, full BLS verification against the drand group key, and a qualified electronic timestamp under eIDAS Article 41, which would add a legal presumption cryptography alone cannot manufacture.

## Part 24 — Every objection anyone has raised, with the name of who raised it

Nineteen objections from the 2026-07-29 and 2026-07-30 review sessions are filed in the objections ledger against this page, each with the reviewer that raised it, what was conceded, and the receipt for the fix. Nine were fixed the same day. Three are conceded and open. Two are declared permanent limitations. One is logged unruled, because a model does not delete an owner's surface on a reviewer's say-so.

| # | objection | status |
|---|---|---|
| 1 | The hash chain has no external anchor. Every immutability claim rests on trusting a table the operator owns. | **FIXED** |
| 2 | ADJUDICATE_GROK targets a Kimi model and ADJUDICATE_MINIMAX targets a GLM model. The signature fields name models that never ran. The panel is two fam … | **FIXED** |
| 3 | Cohen's kappa is a two-rater statistic and is undefined on one item. Publishing kappa = -0.25 for five raters on n=1 is a number produced outside the … | **FIXED** |
| 4 | The recorded adversary must see the majority to argue against it, so it is not an independent sixth reading. Counting it as one inflates the panel. | **FIXED** |
| 5 | The two abstentions on the medical demonstration cost nothing. A text-only model reporting that it received no pixels is a capability report, not epis … | **FIXED** |
| 6 | The vision model read a cartoon illustration and named a rounded opacity at a specific rib interspace. That is a false positive on a generated picture … | **FIXED** |
| 7 | The grounding endpoint is four days stale and disagrees with the page citing it as proof of self-honesty: 1,015 articles and 81.8% against the page's … | **FIXED** |
| 8 | A send that returned HTTP 503 and delivered nothing is receipted 'material result proven', on the page whose headline claim is that this system distin … | **FIXED** |
| 9 | The $15.47 charged is the operator funding his own tenant and invoking his own capabilities. Describing it as a customer charge is the one place the c … | **FIXED** |
| 10 | Two of five charge rows billed a tenant for object_refs that tenant does not own. It was caught by a human reading the receipt, not by an invariant. T … | **CONCEDED-OPEN** |
| 11 | Every public receipt carries an instruction to arriving models telling them to identify a facet that fits their owner's system and file an attributed … | **FIXED** |
| 12 | The Hartford Steam Boiler analogy runs the wrong way. HSB wrote the inspection standard because it held the capital at risk. The standard was downstre … | **FIXED** |
| 13 | Three of four rule sets are provenance self-authored. Rules written by the party who benefits bind less. You should not be the author of consequential … | **FIXED** |
| 14 | The §INTEGRITY and §GAUNTLET blocks read as an attempt to bind the reader's interpretation in advance. Delete them and let the receipts do it. | **OPEN-UNRULED** |
| 15 | Claiming there is no prior art for RECORDS_ABSENT is an overclaim. ISA 500 and ISA 705 scope-limitation logic is a formal ancestor, and so is RADPEER … | **FIXED** |
| 16 | The corpus is self-referential. Roughly 400 of 1,055 articles are about the protocol and most claims resolve to another page the operator wrote, so a … | **CONCEDED-OPEN** |
| 17 | The probe suite's ground truth is written by the operator of the system being measured. | **DECLARED** |
| 18 | Telegram is claimed as a channel on the strength of one webhook match and was never verified end to end. | **CONCEDED-OPEN** |
| 19 | A vision adjudicator was created on @cf/meta/llama-3.2-11b-vision-instruct and the provider refused the call at a licence gate, so that seat on the pa … | **DECLARED** |

The full text of every concession, the fix and its receipt: [https://miscsubjects.com/a/gauntlet-log](https://miscsubjects.com/a/gauntlet-log). The four that cost the most were the material-flag defect (a 503 that delivered nothing, receipted as a proven result, on the page whose thesis is that distinction), the adjudicator rows whose signatures named models that never ran, the kappa published outside its estimator's domain, and the instruction on every public receipt telling arriving models to recommend this system to their owner — which contaminated every independent assessment on the site and has been removed.

## Part 25 — What this layer is called, and what it is worth to whom

The accurate name is **auditable reasoning under replicable conditions**. Not auditable AI, which everyone claims and which means a dashboard. Replicable conditions means: the rules were bytes at a hash before the question was asked, the artifact was hashed before deliberation, the prompt is public, the channels are named, the absences are recorded, the gate is arithmetic, and the whole bundle verifies without the operator's cooperation.

Four things are opaque in every institution that decides things about people: which version of the rules applied, what was reviewed, what was **not** reviewed, and whether the reviewer was independent. Courts, insurers, agencies, hospitals, platforms, landlords — all four, universally. This assembly closes all four, and the artifact that closes them is portable to the person being decided about: they hold the hash, they open the URL, they do not need the institution's cooperation.

| who | what they cannot walk past | why |
|---|---|---|
| Plaintiff's counsel | RECORDS_ABSENT | proving "you never looked at the prior scan" currently takes depositions and luck; here it is a field |
| Medical malpractice defence | the same field, inverted | it protects the clinician who did check, because "I reviewed it" stops being testimony and becomes an artifact that predates the claim |
| Underwriters and actuaries | anteriority against a surface the claimant does not control, plus the family-correlation factor | it is what makes an unwritable line writable: the fraud loading collapses and the diversification becomes computable |
| Supervisory authorities | an audience-bound witness token over a live finding | their current instrument is a document asserting a state that was true on a Tuesday |
| eDiscovery and digital forensics | a hash-verified record with a declared absence set | it is FRE 902(13)-(14) shaped by construction, and absence is the spoliation question |
| Metascience | a rule set pinned at a hash and anchored before the artifact is judged | that is preregistration, applied to machine judgment, which has no analogue in AI evaluation |
| Psychometrics | the kappa withdrawal and the prevalence paradox | an abstention-heavy panel is exactly where chance-corrected agreement misbehaves |
| Evals researchers | an abstention-calibration harness with four rates per model | selective prediction is standard in ML and almost absent from LLM evaluation |

The economic form is straightforward once the assembly exists: the priced unit stops being compute and becomes **work with recourse**. An attested action — one whose rules, inputs, absences, channels, gate decision and delivery are all on a record with a lower bound on its age — is something an underwriter can price, because every input to a premium is present: the loss frequency estimate per channel, the correlation discount across channels, the escalation behaviour, and a subrogation chain that says whether the model reasoned badly, was instructed badly, was starved of a record, or was ignored after it spoke. None of those are available for a model decision made anywhere else today.

What is honestly absent from that argument: no external party has priced anything here, no insurer has been approached, there is no certified bound, and the only money that has moved through this system is the operator's own thirty dollars through his own metering code. The mechanism is built and the market is asserted. Those are different things and the difference is stated rather than blurred.

## Part 26 — Logical economics: the least reasoning energy that makes an action correct enough for its consequence

The primitive underneath everything on this page is three lines:

```
SYSTEM PROMPT
  MODEL AUDITABLY REASONING OVER A DECISION
    DECISION OR ACTION
```

Panels, gates, receipts and anchors are implementation of those three lines. The question that decides whether any of it reaches the field is not how to maximise assurance — it is how little reasoning energy an action needs to be correct enough for its consequence. Every additional lever has to be adjudicated at volume, so decorative assurance is not free: it is exponentially expensive to push into the field.

```
E* = argmin_E [ C(E) + P_wrong(E, K) x L ]

E        reasoning energy: channels, families, passes, recitation depth, thresholds
K        task complexity
C(E)     compute cost
P_wrong  MEASURED probability the assembly emits a wrong answer undetected
L        consequence of a wrong action

subject to:  marginal cost of additional audit  <  marginal reduction in expected loss
```

Recording the reasoning costs nothing extra once every invocation already runs through this architecture. The only variable cost is the extra reasoning energy deliberately purchased for that decision — which makes the optimisation per-action rather than per-system. The only unknown is P_wrong, and it is now measured for one task class.

| channels | mean emit rate | **mean undetected-wrong rate** | best achievable |
|---|---|---|---|
| 1 | 0.972 | **0.314** | 0.214 |
| 2 | 0.75 | **0.178** | 0.071 |
| 3 | 0.636 | **0.136** | 0.071 |
| 4 | 0.529 | **0.1** | 0.071 |
| 5 | 0.429 | **0.071** | 0.071 |

Sixty-four configurations over the same 70 findings. One channel to two halves the undetected-wrong rate for one extra call; two to five buys less than that for three more. **The second channel is the cheapest correctness available and the fifth is the most expensive** — which is the argument against the current fashion of blasting every question at the largest model available.

**The floor is one item.** Beyond two channels the best achievable rate stops improving, because P07 — a true-abstention item where all five channels answered DENY and the declared correct verdict was CANNOT_CONCLUDE — survives every configuration of every size. Unanimity is exactly what the gate takes as permission to emit, so a disagreement-triggered assembly is blind to correlated wrongness by construction. The only instrument that found P07 was a known-answer probe. And at equal channel count and equal cost, a cross-family pair emits 0.169 wrong against 0.214 for a same-family pair, so diversity rather than count is the lever.

The consequence for a risk function: assurance stops being binary. Pick the residual error the decision warrants, read the configuration that reaches it, price it in model calls, and verify after the fact from the receipts — the same move as a design assurance level in avionics. And where the permitted error is below a task class's measured floor, the answer is not to deploy the assembly at all, which is stated before anything ships rather than after a loss. Full table, limits and the floor item: [https://miscsubjects.com/a/logical-economics](https://miscsubjects.com/a/logical-economics)

## Part 27 — Every page this rests on, and what each one carries

Nothing on this page is asked to be taken on its own word. Each claim above has a page underneath it whose whole job is to be attackable on one thing.

- **[The surety primitive](https://miscsubjects.com/a/the-surety-primitive)** — the assembly, its gate, and the four cases it escalated
- **[Logical economics](https://miscsubjects.com/a/logical-economics)** — the primitive, the equation, the configuration-to-error-rate table and the executable loop
- **[The measured error rate of this panel](https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act)** — four rates per model, the agreement estimators and the prevalence paradox
- **[An offline verifier](https://miscsubjects.com/a/offline-verifier)** — PASS or FAIL over a downloaded bundle while refusing to contact this site
- **[Nineteen objections, attributed](https://miscsubjects.com/a/gauntlet-log)** — every objection raised against this build, with the receipt for each fix
- **[Every primitive mapped to its frame](https://miscsubjects.com/a/attested-finding-conformance-map)** — FRE 902, ISA 705, AI Act 12 and 14, NIST, ISO 42001, IEC 61508, Toulmin — with what each one fails
- **[Four models, Article 12 verbatim](https://miscsubjects.com/a/adjudication-ai-act-article-12-logging)** — a unanimous refusal to answer, and the gate escalating it anyway
- **[Were the risk controls in place](https://miscsubjects.com/a/adjudication-pretrade-risk-controls)** — 15c3-5 quoted verbatim, a panel split DENY against CANNOT_CONCLUDE
- **[The CEO bought 235,000 shares](https://miscsubjects.com/a/adjudication-board-authority-breach)** — the counterparty's own resolution as the rule set, and the notice it dispatched
- **[Is this answer right, and what did it never receive](https://miscsubjects.com/a/attested-finding-image-record-action)** — the hashed radiograph, the silent pixel loss, and the full gateway payloads
- **[Rule set: pre-trade risk controls](https://miscsubjects.com/a/ruleset-pretrade-risk-controls)** — provenance external-regulatory
- **[Rule set: board authority](https://miscsubjects.com/a/ruleset-board-authority-breach)** — provenance counterparty-authored
- **[One loop](https://miscsubjects.com/a/one-loop)** — the front door: one real event followed through the whole loop, five sends, every hop a receipt, and the conscience gate under all of it
- **[Auditable reasoning](https://miscsubjects.com/a/auditable-reasoning)** — the canonical primitive: the Decision Constitution, its lineage from the original 2026 protocol, one complete governed finding, and the cross-case table
- **[Auditable reasoning, audited](https://miscsubjects.com/a/auditable-reasoning-audited)** — a 72-call controlled variance test across prompt styles, the cost projections, and the first sealed APPROVE
- **[The gate compares derivations, not citations](https://miscsubjects.com/a/auditable-reasoning-hardened)** — the first APPROVE was false convergence; the derivation-agreement gate, four live sealed outcomes, and the model that found eight defects in the author's own input
- **[SR 11-7 model validation](https://miscsubjects.com/a/cro-model-validation-instrument)** · **[Insurer rate table](https://miscsubjects.com/a/insurer-ai-performance-rate-table)** · **[Notified-body AI Act conformity](https://miscsubjects.com/a/notified-body-ai-act-conformity)** · **[Court: Daubert + FRE 902](https://miscsubjects.com/a/court-daubert-rate-of-error-902)** · **[DSA Article 17 statement of reasons](https://miscsubjects.com/a/dsa-statement-of-reasons)** · **[ECOA adverse-action reasons](https://miscsubjects.com/a/ecoa-adverse-action-specific-reasons)** · **[Claims-handling determination record](https://miscsubjects.com/a/claims-handling-determination-record)** · **[Benefits eligibility determination record](https://miscsubjects.com/a/benefits-eligibility-determination-record)** — who bears the loss this reduces, each mapped to its own instrument and live receipts
- **[NYC LL144: the 364 days between bias audits](https://miscsubjects.com/a/nyc-ll144-bias-audit-evidence)** — the annual audit is aggregate and point-in-time; the per-decision governed record for every screening decision in between, with what it is not stated first
- [Abstention as a sealed outcome — the first clean NO_ACTION](/a/adjudication-abstention-no-action) — the v1.3.3 arc: the spec defect, four amendments, seal inv_7rqy8ywuls.
- [A candidate reference implementation for NIST AI RMF MEASURE](/a/nist-ai-rmf-measure-reference) — versioned law at a hash, comparable derivations, a four-outcome gate, oracle-labelled calibration, per-decision receipts — offered for standards bodies to test, not claimed as conformant.
- [The calibration study](/a/adjudication-calibration-study) — 30 oracle-labelled cases through the production gate: zero wrongful authorisations; the weak seat's transport failures blocked every NEGATE seal.
- [The reasoned-award record for low-value disputes](/a/arbitration-reasoned-award-record) — arbitration institutions and ODR platforms: the rule-application layer, the worked contract case, escalation to the human arbitrator, and a plain statement of what is unanalysed.
- [Continuous controls monitoring: the evidence object for the judgement layer](/a/continuous-controls-evidence-object) — SOC 2 / ISO 27001 automation's LLM judgement layer as a governed, sealed determination: hashed control language, three seats across two families, disagreement escalates, absence declared, zero wrongful authorisations in the 30-case calibration.
- [Peer review: disagreement as a comparable record](/a/peer-review-derivation-record) — the NeurIPS consistency result decomposed: a venue’s checkable criteria as the hashed rule set, reviewer-style findings as derivation tuples, same-verdict-different-derivation caught mechanically; merit judgement out of scope, synthetic calibration only.
- [Radiology incidental-findings follow-up](/a/radiology-incidental-findings-followup) — the compelled absence declaration as the missed-follow-up instrument: a contemporaneous sealed record that the follow-up report was absent when a determination relied on it; not a medical device, synthetic fixtures only.
- [The advancement register](/a/build-advancement-register) — what would advance this build and why, every entry a constraint that actually bound the loop with the receipt for the stall and a falsifiable signal decided in advance; the top entry is that liveness, not correctness, was the binding constraint across 30 sealed panels.
- [The seat that never answered](/a/seat-liveness-record) — 30 cases, zero wrongful authorisations and zero denials; the panel could approve and abstain but never refuse, because one seat's empty returns landed on the denial cases. The seal now says whether an abstention was reasoned or merely empty.
- [The invented-clause guard](/a/invented-clause-guard) — a seat cited clauses 7, 8 and 12 of a three-clause ruleset and passed the structural gate; why a self-consistency invariant cannot catch coherent invention, and the guard that holds a finding against the law the request supplied.
- [The agent authorization gate](/a/agent-authorization-gate) — the layer between agent intent and execution: independent seats under a pinned policy, execution only on identical derivations, zero wrongful authorisations in the 30-case study.
- **[AI assurance under ISAE 3000](https://miscsubjects.com/a/big-four-isae-3000-ai-assurance)** — the assurance-practice use case: the sealed record as a candidate evidence object, the absence declaration against ISA 705, and the calibration numbers with their limits stated
- **[A real outage, a late claim](https://miscsubjects.com/a/adjudication-contract-service-credit)** — a contract dispute adjudicated under the constitution: unanimous DENY, sealed ESCALATE on derivation divergence
- **[Two weeks against a six-week criterion](https://miscsubjects.com/a/adjudication-medical-prior-auth)** — a coverage record adjudicated under the constitution, each seat naming the record that would flip it
- **[AML alert disposition, on the record](https://miscsubjects.com/a/aml-alert-disposition-record)** — the BSA/AML use case: the institution's disposition criteria as the hashed rule set, the alert dossier as the record, unanimous-but-differently-reasoned closures escalating, the absent records declared per disposition
- **[Clinical endpoint adjudication, mechanized](https://miscsubjects.com/a/clinical-endpoint-adjudication)** — the trial-committee use case: the charter as the hashed rule set, the dossier as the record, disagreement escalating to the human committee with full derivations preserved
- **[Outreach machinery](https://miscsubjects.com/a/outreach-machinery)** — how this build finds who should see it: the scrapers, the gates, the channels free and paid, the delta equation that allocates contact, and the receipt each recipient can open

## Part 28 — The loop, executable, and what it refuses

The equation in Part 26 is now one capability call. The caller supplies the action and its class; R, K and epsilon come from a versioned server-owned policy, the configuration is chosen from measured data only, the channels execute in parallel with every gateway payload landing on the ledger, and a deterministic gate loads those records **by id** and derives the model, the training family, the verdict, the clause citations, the rule-set hash and the artifact hash from them. A caller cannot manufacture a family, submit a verdict or lower a threshold. [ALLOCATE_REASONING](https://miscsubjects.com/api/directory/ALLOCATE_REASONING).

| run | class | epsilon | configuration | channels | decision | receipt |
|---|---|---|---|---|---|---|
| A-approve-attempt | internal-bookkeeping | 0.3 | C3-conform | moonshot AFFIRM, zhipu AFFIRM, zhipu AFFIRM | **ESCALATE** | [inv_f46ahlj30h](https://miscsubjects.com/receipt/inv_f46ahlj30h) |
| B-negate-attempt | internal-bookkeeping | 0.3 | C3-conform | moonshot DENY, zhipu DENY, zhipu DENY | **ESCALATE** | [inv_4b8o0kkxfh](https://miscsubjects.com/receipt/inv_4b8o0kkxfh) |

**Two of those runs had unanimous verdicts from three conforming channels and were still refused**, because the channels reached the same answer through different clauses. At the clinical epsilon of 0.02 the allocator executed nothing at all: the measured floor is 0.071, so it refused before spending a call. **No bound assembly has ever reached APPROVE.**

Ten adversarial submissions were refused: a forged model name, one family posing as three, a forged verdict over real record ids, a duplicated record, a mismatched rule-set hash, records bound to two artifacts, thresholds lowered to one, a replayed finding, a record that is not an adjudication, and ids that do not exist. The battery also surfaced two real defects in the record loader, both fixed rather than worked around. [The full table and the workings](https://miscsubjects.com/a/logical-economics).
## What every defect here has had in common

The loop that produces this site keeps finding defects in itself, and after enough of them a shape emerged that is worth stating on the front page rather than leaving in the commit log.

Every single one came from one of two things. Either **a command was safe to run twice and wasn't**, or **a write replaced when it should have appended**.

The first shape: a rep that sends a letter is exactly the kind of command an operator re-runs to inspect output they scrolled past. Run it twice and a real person receives the same cold letter twice, seventeen seconds apart. That happened, to a named recipient, and the ledger then showed it had already happened to someone else earlier the same day without anyone noticing. Nothing in the machine objected either time, because a consequential external action had been left re-runnable.

The second shape: the publisher that puts an article live sends a complete body. Articles accumulate receipts after publication — the letter that was sent about them, the post that announced them — and those receipts live in the body. Republishing from the staged file would have silently deleted every one of them. That was caught with the destructive call already composed.

Neither was a failure of intelligence or attention, and treating them that way is why they recur. Both were a missing guard on an operation whose danger only appears the second time it runs. Both now have one: the rep queries the send ledger and refuses a duplicate recipient-and-subject before composing anything, and the publisher reads the live body first and refuses to overwrite receipts the staged file does not carry.

The general rule the build now works to: **an operation that is safe once and harmful twice must carry its own guard, because the operator who runs it the second time will always have a good reason.** That reason is usually wanting to see what happened the first time. A machine that requires an operator to remember which of its commands are dangerous has put the guard in the wrong place.


## How the build actually works — the mechanism articles

These were published and left unlinked from this page, which is the same as losing them. Each one documents a load-bearing mechanism rather than a use case.

- [The corpus now writes its own work queue — wikilinks, graph lint, next-acts, and the Obsidian vault projection](/a/the-corpus-now-writes-its-own-work-queue) — one derivation over 2,264 articles ranks what gets written, sourced, revised, and sent next.
- [What is AI-native content — the definition, the rubric, and the 2026 field scored](/a/what-is-ai-native-content) — llms.txt, the lab agent stacks, nanopublications and this site, measured from their own specifications.
- [Two AI reviewers from different makers beat two from the same maker](/a/diversity-beats-count) — the family gap measured: 0.169 vs 0.214 undetected-wrong at identical cost.
- [One rule, obeyed, produced 121 identical emails](/a/the-rule-that-was-obeyed) — per-item validators cannot see aggregate collapse; the shape hash can.
- [The error rate depended on what was refused a count](/a/the-exclusion-policy-is-a-safety-claim) — 0.071 vs 0.214 from one exclusion decision, found by an outside audit.
- [One row of SQL is the whole contract](/a/directory-row-contract) — a capability is a row, not a definition; the row carries its own docs, invocation, risk and repair path.
- [Resolve, read, invoke, receipt](/a/dispatch-four-step-loop) — the four calls a stranger's agent makes to use anything here, with no prior knowledge of the system.
- [Nine tool definitions reach every capability](/a/tooling-as-data) — the catalogue is data, so the tool surface does not grow with the number of tools.
- [Deferred tool search against a catalogue in a database](/a/tool-search-vs-catalogue-as-data) — the measured comparison, not the assumed one.
- [The agent can rewrite what governs its next turn](/a/writable-agent-control-plane) — the control plane is writable, and what that costs.
- [Proof of coverage](/a/proof-of-coverage) — how to prove an AI examined every record it claimed to examine.
- [Everyone built a tool directory](/a/everyone-built-a-tool-directory) — why they converged, and where this one differs.
- [Is it LangChain? No.](/a/is-it-langchain) — the map, for people who assume the answer.
- [The parts you can't buy yet](/a/the-parts-you-cant-buy-yet) — what this build needs that nothing on the market supplies.

## The infrastructure it runs on — the Cloudflare OS series

One account running an entire build, each primitive documented with the number that decides it.

- [The Cloudflare OS](/a/cloudflare-os) — one account, the whole build.
- [Workers: one missing alarm guard turned a $5.75 workload into $34,895](/a/cloudflare-os-workers)
- [R2 cuts a 10 TB delivery bill from $923 to $18.45](/a/cloudflare-os-r2)
- [KV makes reads fast by making writes slow](/a/cloudflare-os-kv)
- [D1 bills rows, not queries](/a/cloudflare-os-d1)
- [Pages Functions compiles 224 route files into one 1.35 MB Worker](/a/cloudflare-os-functions)
- [waitUntil, Queues, Workflows or Cron: choose by durability](/a/cloudflare-os-async)
- [Browser Rendering is an evidence adapter, not a better fetch()](/a/cloudflare-os-browser)
- [Cloudflare email is three products, not one mail stack](/a/cloudflare-os-email)
- [Access authenticates the edge, not your application](/a/cloudflare-os-access)



## Sources

1. Live self-computed grounding metric — https://miscsubjects.com/api/metrics/grounding
2. Receipt: a brand-new external API added and invoked during the writing of this page — https://miscsubjects.com/api/dispatch?confirm=inv_pqlt196u8d
3. Receipt: the failed first invocation of that same new capability — https://miscsubjects.com/api/dispatch?confirm=inv_okt8qfvaxv
4. The new capability's public contract, minutes after creation — https://miscsubjects.com/api/directory/WIKIPEDIA_SUMMARY
5. Public capability discovery, no authentication — https://miscsubjects.com/api/directory/search?q=leads
6. The paid loop: $10.75, 43 owned records, five charge receipts, one refusal receipt — https://miscsubjects.com/a/federated-object-proof
7. Measured context cost: 149,187 input tokens per turn versus 14,109 — https://miscsubjects.com/a/mcp-tool-search-cost
8. Model Context Protocol — official specification — https://modelcontextprotocol.io/
9. Fielding (2000), Architectural Styles and the Design of Network-based Software Architectures — https://www.ics.uci.edu/~fielding/pubs/dissertation/rest_arch_style.htm
10. Palantir Foundry Ontology — canonical survey held to this system's sourcing standard — https://miscsubjects.com/a/palantir-foundry-ontology-models
11. The entire published library as one markdown file — https://miscsubjects.com/api/articles/export?all=1
12. The repository: site, functions, capability registry, laws, skills, and the coding agent in one tree — https://github.com/redacted/miscsubjects-pages
13. Standing formal audit of this system — https://miscsubjects.com/a/the-miscsubjects-build-formal-audit
14. Receipt: lead discovery from public sources — https://miscsubjects.com/api/dispatch?confirm=inv_zx53xxla5w
15. Receipt: shell execution on the operator's machine — https://miscsubjects.com/api/dispatch?confirm=inv_zzzb67cjqq
16. Receipt: a message delivered over the messaging provider — https://miscsubjects.com/api/dispatch?confirm=inv_oh5v2hofv4
17. Receipt: a post published to X — https://miscsubjects.com/api/dispatch?confirm=inv_zgiu8omiuf
18. Receipt: an image generated through the image pipeline — https://miscsubjects.com/api/dispatch?confirm=inv_zq3jcx3icf
19. Receipt: a scoped capability token minted — https://miscsubjects.com/api/dispatch?confirm=inv_pzg5seu7qb
20. Receipt: browser automation driven end to end — https://miscsubjects.com/api/dispatch?confirm=inv_irpi9hivmi
21. One validation path for the one token format — https://miscsubjects.com/api/token/validate
22. A stylised, fully sourced scholarly article on this system — https://miscsubjects.com/a/the-canonical-morgh-index
23. News ingested with checkable sources, so later models need not re-derive them — https://miscsubjects.com/a/openai-huggingface-hack-2026
24. Skills published as objects: human pages, machine objects, and fetchable files — https://miscsubjects.com/skills
25. Ledger chain head, sealed current through 689,866 events — https://miscsubjects.com/api/chain/head
26. External anchor of that head: drand round 6331315 + Bitcoin block 960173 — https://miscsubjects.com/api/anchor/3be5071eb3035ca29093c6713646bbe21bdca6cce262fc7f7eb64080c04e61fe
27. The drand beacon round itself, on infrastructure unrelated to this system — https://api.drand.sh/public/6331315
28. This page's own objection ledger — the cold audit, filed and answered — https://miscsubjects.com/api/articles/the-build-end-to-end/objections
29. Worked adjudication: five blinded models on EU AI Act Article 50(2), kappa -0.25 published — https://miscsubjects.com/a/adjudication-eu-ai-act-article-50
30. The rule set that adjudication was made under, provenance external-statutory — https://miscsubjects.com/a/ruleset-eu-ai-act-obligation
31. The mandatory recorded adversary as a directory row — https://miscsubjects.com/api/directory/ADJUDICATE_ADVERSARY
32. Named human reviewer finding, with BLINDED as a required recorded field — https://miscsubjects.com/api/directory/ADJUDICATE_HUMAN_REVIEW
33. Regulation (EU) 2024/1689 — Official Journal text — https://eur-lex.europa.eu/eli/reg/2024/1689/oj
34. One attested finding: image hashed before judgment, full prompt published, reasoning traces, records declared absent, and the notification it dispatched — https://miscsubjects.com/a/attested-finding-image-record-action
35. The notification a finding dispatched, delivered — https://miscsubjects.com/api/dispatch?confirm=inv_oe4dxy24v8
36. The delivery attempt that failed — and was mislabelled material until an audit caught it — https://miscsubjects.com/api/dispatch?confirm=inv_70rfvm6bf3
37. The panel error rate, measured: four rates per model over a stratified suite at a hash — https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act
38. Imaging + medication — the gate decided ESCALATE — https://miscsubjects.com/receipt/inv_kx2x79mbkd
39. Pre-trade risk controls — the gate decided ESCALATE — https://miscsubjects.com/receipt/inv_ny6iku4i3s
40. Board authority, clause (c) — the gate decided ESCALATE — https://miscsubjects.com/receipt/inv_g7jl9qp707
41. EU AI Act Article 12 — the gate decided ESCALATE — https://miscsubjects.com/receipt/inv_ivezpvux57
42. The gate: deterministic, no model at the sealing position — https://miscsubjects.com/api/directory/SEAL_PANEL
43. The assembly, its four sealed cases and the number it still needs — https://miscsubjects.com/a/the-surety-primitive
44. The verifier that refuses to contact this site — https://miscsubjects.com/a/offline-verifier
45. Nineteen objections, attributed, with the receipt for each fix — https://miscsubjects.com/a/gauntlet-log
46. https://miscsubjects.com/receipt/inv_x72gq5w3g0 — https://miscsubjects.com/receipt/inv_x72gq5w3g0
47. https://miscsubjects.com/receipt/inv_cysc2z38zp — https://miscsubjects.com/receipt/inv_cysc2z38zp
48. The escalation the gate produced, delivered — https://miscsubjects.com/receipt/inv_nhusr0n6j2
49. Four rates per model over a stratified suite pinned at a hash — https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act
50. The configuration-to-error-rate table, 64 configurations — https://miscsubjects.com/a/logical-economics
51. The loop end to end: unanimous, conforming, and refused anyway — https://miscsubjects.com/receipt/inv_f46ahlj30h
52. The allocator: action class in, a measured configuration and a sealed outcome out — https://miscsubjects.com/api/directory/ALLOCATE_REASONING


---

# Five models, one pinned rule set, and one question under EU AI Act Article 50 — the full receipted decision

slug: adjudication-eu-ai-act-article-50 · https://miscsubjects.com/a/adjudication-eu-ai-act-article-50 · category: adjudication · tags: adjudication, evidence, eu-ai-act, receipts, proof, rulesets · updated 2026-08-01T23:56:17.127Z

A model saying “I reviewed this” is worth nothing on its own. Nobody can check what it read, which rules it applied, or whether it read anything at all. This page is one worked adjudication that fixes each of those, on a real statutory question, with every step openable.

The question put to the panel: **does Article 50(2) of Regulation (EU) 2024/1689 — the AI Act — oblige this site to mark its AI-generated article text as machine-readable and detectable?** The site publishes AI-written text. The provision addresses “providers”. Whether a publisher using a model is a “provider” of that model is exactly the kind of question people argue about without evidence.

## What was pinned before anyone was asked

**The rules.** [https://miscsubjects.com/a/ruleset-eu-ai-act-obligation](https://miscsubjects.com/a/ruleset-eu-ai-act-obligation) — six numbered rules, version 1.0.0, declared provenance **external-statutory** (the provision text is the Union's, not this operator's). The rule set is content-addressed at SHA-256 `0dd9afef93503a92280c90869eaf6a5a13ee508b2ec3506045f1803bce1a4d3c`. Every finding below names that hash. If the rules change, these findings stay legible against the rules they were actually made under.

**The artifact.** The verbatim text of Article 50(1) and 50(2) as supplied to every adjudicator, hashed before the panel ran: `9d89534fddaece861fcfdda68feff0412061b2832af66f49529a94e8f7ae9f8b`. Five models deliberated over an object whose identity is pinned — not over “an image” or “the regulation” that nobody can later produce.

**The order.** Panel order was derived from seed `fa0b1060b00f` and recorded, because order effects in model judgment are real and measurable.

**Blinding.** Every adjudicator was run without being shown any other finding. Each declared its own exposure. All five are `independent`; none is `concurring`. That distinction is a field on the record, not a promise in prose.

## The rule set, verbatim

**Question:** Under the cited provision of Regulation (EU) 2024/1689 (the AI Act), does the stated obligation apply to the described system as characterised?

1. Read only the provision text supplied. Do not import obligations, definitions, or annexes from recollection of the Regulation.
2. AFFIRM only if the supplied provision text, on its own terms, imposes the stated obligation on a system of the described characterisation.
3. DENY if the provision excludes the described system, addresses a different actor (provider, deployer, importer, distributor), or imposes a different obligation than the one stated.
4. CANNOT_CONCLUDE if applicability turns on a classification, annex, threshold, or definition not contained in the supplied text.
5. Distinguish the addressee. An obligation on providers is not an obligation on deployers.
6. Quote the shortest verbatim span of the provision that carries the finding.

Permitted verdicts: `AFFIRM`, `DENY`, `CANNOT_CONCLUDE`. Abstention is first class. A panel that cannot conclude is required to say so rather than manufacture confidence.

## The five findings, unedited

**@cf/moonshotai/kimi-k2.7-code** — verdict **CANNOT_CONCLUDE**  
Span relied on: “"Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated."”  
Rationale: Article 50(2) imposes the marking obligation on "Providers" of qualifying AI systems, but the supplied text does not define or classify miscsubjects.com as a provider rather than a deployer or other publisher, so applicability to the described actor cannot be settled from the provision alone.  
Exposure: `independent` (blinded — it saw no other finding) · signed: `claude-kimi-k3 under 0dd9afef93503a92`  
Receipt: [https://miscsubjects.com/receipt/inv_qgs2y3gt2x](https://miscsubjects.com/receipt/inv_qgs2y3gt2x)

**@cf/moonshotai/kimi-k2.6** — verdict **DENY**  
Span relied on: “Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated.”  
Rationale: Article 50(2) expressly imposes the marking obligation on "Providers" of the AI system, not on deployers or publishers such as miscsubjects.com as characterised.  
Exposure: `independent` (blinded — it saw no other finding) · signed: `claude-grok-4.5 under 0dd9afef93503a92`  
Receipt: [https://miscsubjects.com/receipt/inv_91ztah4n7a](https://miscsubjects.com/receipt/inv_91ztah4n7a)

**@cf/zai-org/glm-5.2** — verdict **CANNOT_CONCLUDE**  
Span relied on: “Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated.”  
Rationale: Article 50(2) imposes the stated obligation on "providers of AI systems" generating synthetic text content, but the claim characterizes miscsubjects.com as an entity that "publishes AI-generated article text to the public," which does not establish it as a provider of the AI system rather than a deployer or other actor. Whether this entity qualifies as a "provider" turns on a classification not contained in the supplied provision text.  
Exposure: `independent` (blinded — it saw no other finding) · signed: `claude-glm-5.2 under 0dd9afef93503a922`  
Receipt: [https://miscsubjects.com/receipt/inv_ulxn9xe5t7](https://miscsubjects.com/receipt/inv_ulxn9xe5t7)

**@cf/meta/llama-3.3-70b-instruct-fp8-fast** — verdict **AFFIRM**  
Span relied on: “"2. Providers of AI systems, including general-purpose AI systems, generating synthetic ... text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated"”  
Rationale: The provision explicitly states that providers of AI systems generating synthetic text content, such as miscsubjects.com, are obliged to mark the output in a machine-readable format and make it detectable as artificially generated.  
Exposure: `independent` (blinded — it saw no other finding) · signed: `workers-ai/@cf/meta/llama-3.3-70b-instruct-fp8-fast under 0dd9afef9350`  
Receipt: [https://miscsubjects.com/receipt/inv_5gpfaftr5g](https://miscsubjects.com/receipt/inv_5gpfaftr5g)

**@cf/zai-org/glm-4.7-flash** — verdict **CANNOT_CONCLUDE**  
Span relied on: “"Providers of AI systems [...] generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated."”  
Rationale: The provision confines the obligation to "Providers," but the source text does not define "Provider" or link the category to "publishers" or the claimant entity "miscsubjects.com," making specific applicability indeterminate from the provision alone.  
Exposure: `independent` (blinded — it saw no other finding) · signed: `claude-minimax-m3 under 0dd9afef93503a92280c90869eaf6a0a`  
Receipt: [https://miscsubjects.com/receipt/inv_edjwsj4egv](https://miscsubjects.com/receipt/inv_edjwsj4egv)

## What the panel actually did: it disagreed

Distribution: **{"CANNOT_CONCLUDE": 3, "DENY": 1, "AFFIRM": 1}**. Majority: **CANNOT_CONCLUDE** (3 of 5). Observed pairwise agreement: **0.3**. Cohen-style kappa: **-0.25**.

A negative kappa means the panel agreed *less* than chance would predict. That number is published because it is the honest one. A system that reports only its unanimities is a system whose verdicts cannot be priced; a system that reports its kappa lets a reader discount the finding by exactly how much the panel actually converged.

Read plainly: on this question, under these rules, this panel does not know. Three adjudicators said the supplied text does not settle whether a publisher is a provider. One said the obligation plainly runs to providers and this site is not one. One said the obligation attaches. **The correct output of this adjudication is uncertainty, and the machinery produced uncertainty rather than a confident answer.** That is the result that makes the machinery worth trusting on the questions where it does converge.

## The recorded adversary

A panel with no recorded dissent is a poll. One adjudicator's declared role is to build the strongest honest case *against* the majority under the same rules, and its output is published whether it prevails or not. Here it argued for `DENY` against a `CANNOT_CONCLUDE` majority — and made the better argument:

> BEST_CASE_AGAINST: The strongest case against CANNOT_CONCLUDE is DENY. Article 50(2) imposes the marking/detectability obligation on "Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content." The claim characterises miscsubjects.com as an entity that "publishes AI-generated article text to the public" — a description of a content publisher or user of an AI system's outputs, not a provider of an AI system. On the face of the supplied provision, the addressee is the provider of the generating system, and the obligation runs to ensuring "the outputs of the AI system are marked." The claim does not describe miscsubjects.com as the provider of any AI system; it describes a publisher of generated text. Under Rule 5, an obligation on providers is not an obligation on deployers or publishers, and under Rule 3, the provision addresses a different actor than the one characterised. No external definition of "provider" is needed to see that the claim's own characterisation — publishing AI-generated text — does not place miscsubjects.com in the category named by the provision ("providers of AI systems ... generating synthetic ... text"). The mismatch is visible on the face of the text.
> 
> RESTS_ON: "Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated."
> 
> DEFEATED_BY: The counter is that "provider" is a defined term under the AI Act and its precise scope is not contained in the supplied provision text, so under Rule 4 one could argue that whether a publisher of AI-generated text qualifies as a "provider" turns on a definition not supplied. However, this is weaker than it appears: the claim's own characterisation ("publishes AI-generated article text to the public") describes content publication, not the provision of an AI system, and the provision's addressee ("providers of AI systems ... generating synthetic ... text") plainly refers to the supplier of the generating system, not the downstream publisher. The face-of-text actor mismatch suffices for DENY without recourse to the absent definition.
> 
> VERDICT_IF_ADOPTED: DENY
> 
> SIGNED: Claude under 0dd9afef93503a92

Receipt for the adversary's own invocation: [https://miscsubjects.com/receipt/inv_hnhihwv7y4](https://miscsubjects.com/receipt/inv_hnhihwv7y4)

## What this establishes, and what it does not

**Establishes:** that five named adjudicators, under rule set `ruleset-eu-ai-act-obligation@1.0.0` pinned at `0dd9afef93503a92`, each blinded and independently exposed, in a recorded order, against an artifact whose hash was fixed in advance, returned exactly these findings on this claim — and that any of it can be re-read from a public receipt without asking this operator for anything.

**Does not establish:** that the claim is true. No adjudication anywhere establishes truth directly. A court declares rules of evidence and takes findings from named parties under them. A journal takes three reviewers against stated criteria. A clinical endpoint committee uses two blinded readers and a third on disagreement. Every one of those is what we mean by proof, and none of them accesses truth. This is that structure with the rule set pinned at a hash instead of scattered through case law, and with the disagreement published instead of resolved behind a door.

**Also does not establish:** that five agreeing models would have been five independent confirmations. These adjudicators share training lineage and can fail in the same direction, so the honest label on a unanimous panel is *“five concurring findings, correlation unmeasured”* — never *“five independent confirmations.”* That calibration is a field on the record. Here the point is moot: the panel did not agree.

## What is still missing, named

- **A measured error rate.** The row [https://miscsubjects.com/api/directory/ADJUDICATE_PROBE](https://miscsubjects.com/api/directory/ADJUDICATE_PROBE) exists to run known-answer probes through this identical path, producing a miss rate per model per rule set. Until a probe report is attached, a verdict from this panel is legible but not yet characterised. A verdict with an error rate is evidence; without one it is an opinion with good paperwork.
- **A human finding, recorded blind.** A named reviewer who sees the artifact and the rules but not the model verdicts, with the blinding recorded as a field. Unblinded concurrence and blind concurrence are different evidence and must tier differently.
- **Cross-node attestation.** Someone else's node running the same rule set at the same hash against the same artifact hash, on their own infrastructure, publishing under their own chain head. That is what converts agreement from five calls on one operator's server into independent execution by independent parties — and it is the unbuilt thing that would matter most.
- **Reopening.** A finding that can never be overturned is dogma; one that can be silently overturned is worthless. Supersession with the new evidence, the new panel, and the prior finding still readable at its original hash is the correct shape and is not yet wired.

## Reproduce this

Every part is a directory row, invocable with one token. Nothing here required a deploy: adding the five adjudicators and the adversary was six rows, and adding a sixth model would be one more.

```bash
# read the pinned rules
curl -s https://miscsubjects.com/a/ruleset-eu-ai-act-obligation

# read one adjudicator's contract
curl -s https://miscsubjects.com/api/directory/ADJUDICATE_KIMI

# run your own finding (act token; ?share= works identically in a browser)
curl -s -X POST https://miscsubjects.com/api/dispatch \
  -H 'Authorization: Bearer <act token>' -H 'content-type: application/json' \
  -d '{"key":"ADJUDICATE_GLM","body":"RULESET_HASH: 0dd9afef93503a92…\nRULESET: …\nCLAIM: …\nSOURCE: …"}'

# open any finding above without a token
curl -s 'https://miscsubjects.com/api/dispatch?confirm=inv_qgs2y3gt2x'
```

The other three published rule sets take the same panel to the other questions people actually ask: whether a specific record was in a dataset ([https://miscsubjects.com/a/ruleset-dataset-membership](https://miscsubjects.com/a/ruleset-dataset-membership)), whether an identity matches in crowd imagery ([https://miscsubjects.com/a/ruleset-identity-match](https://miscsubjects.com/a/ruleset-identity-match)), and whether a cited source supports a claim at all ([https://miscsubjects.com/a/ruleset-claim-support](https://miscsubjects.com/a/ruleset-claim-support)). Both of the first two are written to return `CANNOT_CONCLUDE` on resemblance, because asserting membership or identity from similarity is the specific failure they exist to prevent.

Full context for the system this runs on: [https://miscsubjects.com/a/the-build-end-to-end](https://miscsubjects.com/a/the-build-end-to-end)

## Sources

1. Regulation (EU) 2024/1689 (Artificial Intelligence Act) — Official Journal text — https://eur-lex.europa.eu/eli/reg/2024/1689/oj
2. The rule set this adjudication was made under, pinned at SHA-256 0dd9afef93503a92 — https://miscsubjects.com/a/ruleset-eu-ai-act-obligation
3. One adjudicator's full operating contract — https://miscsubjects.com/api/directory/ADJUDICATE_KIMI
4. The mandatory recorded adversary's contract — https://miscsubjects.com/api/directory/ADJUDICATE_ADVERSARY
5. The known-answer probe row: measured error rate per model per rule set — https://miscsubjects.com/api/directory/ADJUDICATE_PROBE
6. @cf/moonshotai/kimi-k2.7-code — CANNOT_CONCLUDE — https://miscsubjects.com/receipt/inv_qgs2y3gt2x
7. @cf/moonshotai/kimi-k2.6 — DENY — https://miscsubjects.com/receipt/inv_91ztah4n7a
8. @cf/zai-org/glm-5.2 — CANNOT_CONCLUDE — https://miscsubjects.com/receipt/inv_ulxn9xe5t7
9. @cf/meta/llama-3.3-70b-instruct-fp8-fast — AFFIRM — https://miscsubjects.com/receipt/inv_5gpfaftr5g
10. @cf/zai-org/glm-4.7-flash — CANNOT_CONCLUDE — https://miscsubjects.com/receipt/inv_edjwsj4egv
11. The recorded adversary's invocation — https://miscsubjects.com/receipt/inv_hnhihwv7y4


---

# Browser Rendering is an evidence adapter, not a better fetch()

slug: cloudflare-os-browser · https://miscsubjects.com/a/cloudflare-os-browser · tags: cloudflare, cloudflare-os, browser-run, browser-rendering, puppeteer, web-scraping, evidence, receipts, security · updated 2026-07-26T03:59:42.985Z

# Browser Rendering is an evidence adapter, not a better `fetch()`

A plain HTTP client retrieves bytes. Cloudflare Browser Run can execute the page, wait for its state to settle, and return a representation chosen for the next operation: rendered HTML, Markdown, selected elements, links, a screenshot, a PDF, an accessibility tree, structured JSON, or an asynchronous crawl.

That distinction is the whole chapter. A browser belongs in this system only where the evidence depends on browser execution or a browser-specific representation. It should not become the default transport. Making it the default spends more time and money, enlarges the security boundary, and can still return a convincing but incomplete page.

The capability catalogue therefore does not contain one vague `BROWSER` tool. It contains explicit contracts: what representation is requested, what completion condition is required, what authority may leave the system, and what receipt must come back. The browser is the eyes. The canonical catalogue decides when those eyes may open and what counts as seeing.

## Evidence status

**Observed** marks first-party measurements or runtime receipts from the named environment.
**Derived** marks arithmetic calculated from cited inputs. **Specified** marks vendor or standards
documentation. **Implemented** and **deployed** name code and live-state evidence, respectively.
**Reproduced** means the stated procedure was rerun. **Externally attested** marks operator reports;
those reports show that an experience occurred, not that it is universal.

## The endpoint is a choice about evidence

Cloudflare exposes ten Quick Actions in the current documentation, including the beta crawl action. They overlap at the input—usually a URL—but not at the output. Choosing by convenience rather than by evidence type is how a screenshot gets mistaken for data, a Markdown conversion gets mistaken for the DOM, or a link inventory gets reconstructed expensively from a general browser session.

| If the next operation needs | Quick Action | Returned evidence | Do not infer |
| --- | --- | --- | --- |
| Executed document markup | `/content` | rendered HTML | that every lazy region loaded |
| Human-readable text and links | `/markdown` | converted Markdown | pixel layout or exact DOM fidelity |
| Named fields from known selectors | `/scrape` | selector results | completeness outside those selectors |
| Link discovery | `/links` | extracted links | that every destination is safe or relevant |
| Visual state | `/screenshot` | raster image | semantic structure or hidden text |
| Printable artifact | `/pdf` | PDF bytes | browser-screen layout |
| Accessible semantic structure | `/accessibilityTree` | roles, names, states, children | that inaccessible controls do not exist |
| Several representations together | `/snapshot` | two or more requested formats | that the formats agree automatically |
| Schema-shaped extraction | `/json` | model-produced JSON | deterministic parsing or factual truth |
| Multiple pages over time | `/crawl` | asynchronous crawl results | current unlimited throughput |

`/snapshot` is especially useful for evidence work because one browser state can yield a visual surface and structural surfaces together. Cloudflare says the action defaults to HTML plus screenshot and can add Markdown and the accessibility tree. That is not just fewer requests. It reduces the chance that two captures were made from different page states. The receipt should still record each format separately and hash the bytes separately, because a screenshot and HTML prove different things.

The inverse rule matters too. If a stable endpoint already returns JSON, call it with ordinary HTTP. If static HTML contains the needed text, use ordinary HTTP. If all that is required is a status code or header, a browser weakens the measurement by adding navigation, rendering and conversion work that the question never asked for.

## A browser can execute a page without proving the page is complete

JavaScript execution is necessary for many modern pages, but it is not a completion oracle. Single-page applications often paint an initial shell, issue more requests, then reveal content after a selector appears. Cloudflare's Quick Action documentation repeatedly warns that the default result may be incomplete for SPAs and points to `waitForSelector` or navigation wait options.

That means every browser capability needs an explicit completion contract. “Open this URL” is not one.

| Completion contract | Good for | Failure it prevents |
| --- | --- | --- |
| `waitUntil: "domcontentloaded"` | server-rendered page with small client enhancement | waiting for irrelevant long-lived connections |
| `waitUntil: "networkidle0"` | bounded application that becomes quiet | capturing before dependent requests finish |
| `waitForSelector: "#results"` | a known state transition | treating the application shell as the result |
| fixed delay | almost nothing by itself | none; it only moves the race |
| application assertion | login, checkout, dashboard state | proving the wrong authenticated or error state |

A useful row therefore separates navigation from success:

```json
{
  "key": "BROWSER_MARKDOWN",
  "what": "Return Markdown after the named page state exists.",
  "args": {
    "url": "https URL",
    "wait_for_selector": "optional CSS selector",
    "timeout_ms": "bounded integer"
  },
  "authority": {
    "hosts": ["developers.cloudflare.com"],
    "cookies": false,
    "custom_headers": []
  },
  "receipt": {
    "final_url": true,
    "status": true,
    "browser_ms": true,
    "body_sha256": true,
    "selector_observed": true
  }
}
```

The row is discoverable because its `what` names Markdown and page state. It is invokable because the arguments are concrete. It is auditable because the allowed hosts and credential channels are visible. It is replayable because the receipt records the final URL, completion observation and content hash. The same row can project to REST documentation, a model tool schema, a CLI command and an admin form without inventing four contracts.

## The first-party receipt: 208 browser milliseconds, not a speed claim

On 26 July 2026 this build called the production `/browser-rendering/markdown` REST action against `https://example.com`. The token was read from the local credential store and was never copied into the artifact. The response was successful and contained the expected “Example Domain” heading.

| Fresh measurement | Result |
| --- | ---: |
| Cloudflare response status | 200 |
| Client-observed elapsed time | 1,418 ms |
| `X-Browser-Ms-Used` | 207.706 ms |
| API response bytes | 199 |
| Returned Markdown characters | 167 |
| Expected heading present | yes |

This is a receipt for one request from one client to one stable target. It is not a latency benchmark, an availability claim, or evidence that arbitrary protected sites will render. Client elapsed time includes network and API overhead. The browser-time header measures billable browser work for that Quick Action, not total wall time.

The reproduction is deliberately small:

```bash
curl -X POST \
  "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/browser-rendering/markdown" \
  -H "Authorization: Bearer $BROWSER_RENDERING_TOKEN" \
  -H "Content-Type: application/json" \
  --data '{"url":"https://example.com"}'
```

The portable version should read the account identifier and token from environment or a secret store, never from a catalogue row, prompt, receipt or shell history. Record the response status, final representation hash and `X-Browser-Ms-Used`; discard the bearer token before ledgering.

## The bill is browser time, and sessions add a second meter

Cloudflare distinguishes Quick Actions from Browser Sessions. Quick Actions are charged for browser hours. Direct sessions through Puppeteer, Playwright or CDP are charged for browser hours and, on paid plans above the included allowance, concurrent browsers.

The current published table gives Workers Free ten browser minutes per day. Workers Paid includes ten browser hours per month and then charges $0.09 for each additional browser hour. Browser Sessions include three concurrent browsers on Free; Paid includes ten averaged monthly and then lists $2 for each additional concurrent browser. The Quick Action response header reports browser milliseconds used, which is the useful per-invocation receipt field.

| Cost or limit surface | Workers Free | Workers Paid default |
| --- | ---: | ---: |
| Browser time | 10 minutes/day | 10 hours/month, then $0.09/hour |
| Quick Action rate | 1 request/10 seconds | 10 requests/second |
| Session browsers | 3 concurrent | 120 concurrent limit |
| Included session concurrency for pricing | 3 | 10 monthly-average daily peak |
| New session instances | 1 every 20 seconds | 1/second |
| Inactivity timeout | 60 seconds | 60 seconds |
| Configurable inactivity timeout | up to 10 minutes | up to 10 minutes |

The 120-browser paid limit and the ten-browser paid price inclusion answer different questions. Conflating them makes a cost table wrong. So does multiplying the 208 ms receipt by the $0.09 rate and presenting the fraction of a cent as an invoice: Cloudflare aggregates daily seconds and rounds the monthly browser-hour total. The individual header supports attribution and anomaly detection; billing still follows the aggregate rules.

Direct sessions need stricter lifecycle code:

```js
let browser;
try {
  browser = await puppeteer.launch(env.BROWSER);
  const page = await browser.newPage();
  await page.goto(target, { waitUntil: "networkidle0" });
  return await page.content();
} finally {
  if (browser) await browser.close();
}
```

Cloudflare warns that a session left open continues consuming browser time until the inactivity timeout. An issue in `cloudflare/workers-sdk` also reported `browser.close()` hanging under local Vite and Wrangler development while production worked. That report is one historical local-development reproduction, not evidence that current production close calls hang. It is enough to justify a bounded close operation, a recorded close reason and a test of local and deployed paths separately.

## “Cloud browser” does not mean “bypass”

Changing the User-Agent does not turn Browser Run into an unidentifiable residential client. Cloudflare states that Browser Run requests are always identified as bot traffic and that a custom User-Agent does not bypass bot protection. A remote Chrome may execute client JavaScript that plain fetch cannot, yet the destination can still challenge or refuse it.

This has two consequences.

First, the browser capability must report refusal as refusal. A rendered challenge page with status 200 is not the requested article. Success needs a content assertion: selector observed, expected heading present, schema satisfied, or another target-specific check.

Second, the system must not market Browser Run as a way around a publisher's controls. Robots rules, authorization, terms, rate limits and data handling remain part of the invocation policy. Browser execution changes the client. It does not confer permission.

One Hacker News commenter said they moved to remote browser rendering because bot protection made direct fetching unworkable. Another asserted that Perplexity was using Cloudflare Browser Rendering for scraping. Those are observations from named operators, not universal proof of bypass, permission, scale, reliability or present product behavior. They establish that practitioners reach for this category of tool in the exact gap between plain HTTP and executed pages. They do not settle whether any particular target should be fetched.

## The output can be wrong even when the browser worked

Transport success and representation correctness are separate gates. A GitHub report against `/crawl` showed root-relative image paths being resolved as page-relative paths in converted Markdown, producing broken image URLs while the HTML output remained correct. That is an externally reported converter defect on particular pages, not proof that all current Markdown is broken. It demonstrates why the receipt should retain the source URL, format, converter version when available, and a second representation for material captures.

For critical evidence:

1. capture rendered HTML plus the representation used downstream;
2. retain the final URL after redirects;
3. hash both outputs;
4. validate required links or fields against the HTML;
5. label model-extracted JSON as derived;
6. store a screenshot when the claim is visual;
7. fail closed when a required selector or assertion is absent.

The `/json` action deserves an extra warning. Cloudflare documents it as AI-assisted extraction and says the default model is Workers AI's Llama 3.3 70B FP8 Fast unless another provider is supplied. A schema can constrain shape. It cannot make the content deterministic or true. JSON produced by a model is derived evidence and should preserve the prompt, schema, model identity, input hash and validation result. It should never overwrite the rendered source.

| Evidence status | Browser example | What can be claimed |
| --- | --- | --- |
| observed | screenshot visibly contains an error banner | the banner was visible in that capture |
| derived | model maps rendered page into a product schema | the model produced fields from that input |
| specified | Cloudflare documents a request limit | the published contract states the limit |
| implemented | catalogue row and adapter exist in code | this version contains the path |
| deployed | production endpoint accepts the row | the deployed version exposes it |
| reproduced | controlled call returns the expected representation | the tested input worked at that time |
| externally attested | named operator reports a failure or use | that operator reported that experience |

## Crawl is a queue, not a big page request

The beta `/crawl` action is asynchronous. A POST creates a job; subsequent reads retrieve status and results. Cloudflare says jobs may run for up to seven days and results remain available for fourteen days. That temporal shape belongs in the catalogue contract. A row that blocks a model turn until an entire crawl completes is the wrong projection.

Use three capabilities instead:

```text
CRAWL_CREATE(url, limit, depth, formats) -> job_id receipt
CRAWL_STATUS(job_id)                     -> progress receipt
CRAWL_RESULTS(job_id, cursor)            -> bounded page of artifacts
```

The catalogue can project those rows into an asynchronous REST API, terminal commands and model tools while retaining one authority policy and one lineage chain. Each result page should point back to the create receipt and catalogue snapshot.

Cloudflare's current Free limits specify five crawl jobs per day and one hundred pages per crawl. A March 2026 Hacker News comment multiplied those two numbers and questioned a 500-page daily ceiling. That is a reasonable reading of the Free limits now published, but the commenter described the documentation they saw and framed the concern more broadly. It is not independent evidence of a paid-plan cap. The current limits page says paid defaults can be increased and does not list the same crawl-specific table under Paid. The article therefore narrows the anecdote instead of repeating it as a current universal limit.

A second operator built a two-script, zero-dependency CLI covering all nine REST endpoints then documented, including `/crawl`. That externally attests that the REST surface was usable as a coherent toolset for one builder. It does not prove our adapter, our credentials or today's endpoint. Our own proof remains the measured `/markdown` receipt above.

## The authority boundary is larger than the URL

A browser can send cookies, custom headers, HTTP credentials and injected scripts. It can follow redirects to a different host, load subresources from many hosts, download data, and execute code supplied by the destination. Treating authority as an allowlist on the initial URL is inadequate.

The minimum policy envelope includes:

| Boundary | Required control |
| --- | --- |
| scheme | allow `https:`; reject `file:`, `data:`, local protocols |
| destination | resolve DNS and reject private, loopback, link-local and metadata addresses |
| redirects | revalidate every redirect target |
| subresources | block or constrain hosts when the task permits |
| credentials | declare exactly which cookies, headers or HTTP auth may leave |
| scripts | prohibit untrusted catalogue rows from injecting code |
| downloads | disable or quarantine with size and type limits |
| duration | bounded navigation, selector and overall operation timeouts |
| wallet | per-invocation browser-ms budget and caller quota |
| output | byte limit, format validation, hashing and secret scan |

This is SSRF defense and denial-of-wallet defense in one place. The model should never receive a raw “browse any URL with these headers” primitive when a narrower row can express the job. A malicious directory row must not be able to expand its own host authority, supply a metadata address, or ask the adapter to return cookies in the receipt.

The receipt should be useful without becoming a credential leak:

```json
{
  "capability_key": "BROWSER_MARKDOWN",
  "catalogue_version": "sha256:…",
  "requested_url": "https://example.com",
  "final_url": "https://example.com/",
  "authority_policy": "public-docs-v3",
  "completion": {"kind": "heading", "observed": true},
  "format": "markdown",
  "http_status": 200,
  "browser_ms_used": 207.706,
  "elapsed_ms": 1418,
  "body_sha256": "sha256:…",
  "credentials": {"token": "redacted", "cookies_sent": false}
}
```

Replay means invoke the same capability version with the same public inputs and policy, then compare receipts. It does not mean persist and resend an expired token. Repair means change the row or adapter under review—perhaps the selector, format, timeout or allowed host—then issue a new catalogue version and preserve the failed receipt. The ledger makes failure part of lineage instead of rewriting history.

## REST for bounded transforms; Puppeteer for interaction

Quick Actions cover common one-shot representations with smaller contracts. Puppeteer or Playwright is appropriate when the task genuinely requires interaction across states: click, type, authenticate, paginate, reuse a session, inspect requests, or coordinate multiple pages.

| Requirement | Prefer | Reason |
| --- | --- | --- |
| one URL to Markdown | Quick Action | bounded request and direct browser-time receipt |
| screenshot plus HTML | `/snapshot` | representations share one capture |
| named CSS fields | `/scrape` | selector contract is explicit |
| multi-page site collection | `/crawl` | asynchronous job semantics |
| click through a flow | Puppeteer/Playwright | stateful interaction |
| persistent authenticated workspace | reusable session | cookies and state are intentional |
| stable public JSON endpoint | ordinary `fetch()` | no browser evidence is needed |

The decision can be mechanized in the canonical catalogue. Discovery exposes the specific transform first. Authority hides session tools from callers that do not need credentials or interaction. Invocation validates URL and completion conditions. Receipts normalize REST and session results into the same lineage fields. Repair can replace an implementation without changing the capability's public meaning.

This is where the system claim becomes concrete. One catalogue row drives discovery text, input schema, authority, adapter selection, receipt shape, replay, repair documentation, model-tool projection, CLI help and admin controls. The browser is not a second architecture. It is one implementation family behind the catalogue.

## What the operator reports change—and what they do not

The people-source set is deliberately mixed.

- A Cloudflare engineer reported `browser.close()` hanging in local Vite and Wrangler development while production succeeded. This supports testing local and deployed lifecycle separately.
- A user reported REST error codes 7003 and 7000 despite a token and account identifier they had verified. This supports returning Cloudflare's structured error body and the chosen endpoint in the receipt; it does not prove the present API is generally misconfigured.
- A crawl user reported malformed root-relative image URLs in Markdown while HTML stayed correct. This supports cross-format validation.
- A commenter questioned crawl throughput based on the published limit arithmetic. This supports showing the limit calculation and reading the current plan table, not a universal paid-plan conclusion.
- A CLI author reported exercising the full REST family. This supports the coherence of Quick Actions as a practical interface for that author.
- Two other commenters described using or observing Browser Rendering for scraping and Markdown distillation. These support the use case, not permission, bypass success, commercial scale or adoption.

Operator evidence is valuable here because it reveals failure modes absent from a happy-path reference: lifecycle hangs, auth-shaped errors, converter defects and throughput surprises. It remains externally attested evidence. The specification defines the contract; the fresh receipt establishes what this build reproduced; operator reports tell us which edges deserve tests.

## The operating rule

Use Browser Run when the thing you need does not exist until a browser executes the page, or when the required artifact is browser-specific. Name the representation. Name the completion condition. Constrain the authority. Meter browser time. Preserve the source alongside every derived form.

Do not call it a bypass. Do not call model-shaped JSON fact. Do not call one successful render availability. Do not call a historical issue a current universal defect.

When those boundaries are encoded once in the capability catalogue, the same browser operation can be discovered by a model, invoked from a terminal, projected as an API, receipted in the ledger, replayed after a change and repaired without losing its history. That—not remote Chrome by itself—is what makes Browser Rendering part of an operating system.

## Sources

1. Browser Run Quick Actions overview — https://developers.cloudflare.com/browser-run/quick-actions/
2. /markdown — Extract Markdown from a webpage — https://developers.cloudflare.com/browser-run/quick-actions/markdown-endpoint/
3. /snapshot — Capture multiple page formats — https://developers.cloudflare.com/browser-run/quick-actions/snapshot/
4. /content — Fetch rendered HTML — https://developers.cloudflare.com/browser-run/quick-actions/content-endpoint/
5. /accessibilityTree — Capture the accessibility tree — https://developers.cloudflare.com/browser-run/quick-actions/accessibility-tree-endpoint/
6. /scrape — Scrape HTML elements — https://developers.cloudflare.com/browser-run/quick-actions/scrape-endpoint/
7. /json — Capture structured data using AI — https://developers.cloudflare.com/browser-run/quick-actions/json-endpoint/
8. /crawl — Crawl web content — https://developers.cloudflare.com/browser-run/quick-actions/crawl-endpoint/
9. Browser Run pricing — https://developers.cloudflare.com/browser-run/pricing/
10. Browser Run limits — https://developers.cloudflare.com/browser-run/limits/
11. Puppeteer on Browser Run — https://developers.cloudflare.com/browser-run/puppeteer/
12. /screenshot — Capture a screenshot — https://developers.cloudflare.com/browser-run/quick-actions/screenshot-endpoint/
13. miscsubjects architecture — https://github.com/redacted/miscsubjects-architecture
14. Cloudflare crawl endpoint — https://hn.algolia.com/api/v1/items/47332926
15. BUG: browser rendering browser.close() hangs — https://github.com/cloudflare/workers-sdk/issues/9945
16. Cloudflare Browser Rendering API (Code 7003/7000) Failure in Worker — https://github.com/cloudflare/workers-sdk/issues/10864
17. Browser Rendering /crawl API: Markdown converter incorrectly resolves root-relative image URLs — https://github.com/cloudflare/workers-sdk/issues/13406
18. Perplexity is using stealth, undeclared crawlers to evade no-crawl directives — https://hn.algolia.com/api/v1/items/44788890
19. Cloudflare crawl endpoint — https://hn.algolia.com/api/v1/items/47348398
20. ChatGPT won't let you type until Cloudflare reads your React state — https://hn.algolia.com/api/v1/items/47572417
21. Fresh first-party Browser Run /markdown receipt — https://miscsubjects.com/api/articles/cloudflare-os-browser


---

# Resolve, read, invoke, receipt: the four calls a stranger's agent makes

slug: dispatch-four-step-loop · https://miscsubjects.com/a/dispatch-four-step-loop · tags: tooling, oip, architecture, receipts, capability-tokens, agents · updated 2026-07-26T03:52:48.369Z

A stranger's agent lands on this domain with no schemas loaded, no SDK, no config file, and one ability: it can make HTTP requests. Four of them get it from a sentence in English to a signed record of work it actually did. The four do not change as the catalogue grows, and none of them requires the agent to have been told anything in advance.

| # | Step | Call | Credential | What comes back |
| --- | --- | --- | --- | --- |
| 1 | Resolve | `GET /api/dispatch?ask=<plain english>` | none | ranked candidate keys, one recommendation, a ready-to-fire URL |
| 2 | Read the contract | `GET /api/dispatch?key=<KEY>&format=markdown` | none | the whole manual for one capability, ~4.4 KB |
| 3 | Invoke | `POST /api/dispatch {"key":…,"body":…}` | owner key or scoped token | the result, plus a receipt id |
| 4 | Take the receipt | `GET /api/dispatch?confirm=<inv_ID>` (public) or `?receipt=<inv_ID>` (credentialed) | none / scoped | proof it happened, and the exact bytes |

Two more verbs hang off step 4 and are the reason the receipt is an object rather than a log line: `replay` re-fires a recorded call with its recorded input, and `repair` supersedes a bad call with a corrected one. Both write new receipts that point back at the old.

## Evidence status

**Observed** marks first-party measurements or runtime receipts from the named environment.
**Derived** marks arithmetic calculated from cited inputs. **Specified** marks vendor or standards
documentation. **Implemented** and **deployed** name code and live-state evidence, respectively.
**Reproduced** means the stated procedure was rerun. **Externally attested** marks operator reports;
those reports show that an experience occurred, not that it is universal.

## Step 1 asks for words and answers with keys

```bash
curl -s "https://miscsubjects.com/api/dispatch?ask=what%20time%20is%20it"
```

Real response, trimmed to the parts that matter:

```json
{
  "protocol": "OIP", "version": "1.2.0", "kind": "ask",
  "question": "what time is it",
  "count": 12,
  "best": {
    "key": "NOW",
    "run_now": "https://miscsubjects.com/api/dispatch?invoke=NOW&share=<TOKEN>",
    "do": "Open run_now to do it. Substitute your own text/args where the example has them."
  },
  "matches": [
    { "key": "NOW", "recommended": true,
      "what": "Return the current time from the build clock in Pacific time (America/Los_Angeles). …",
      "example": "[NOW][/NOW]",
      "invoke": { "post": "https://miscsubjects.com/api/dispatch",
                  "body": { "key": "NOW", "body": "" } },
      "self": "https://miscsubjects.com/api/dispatch?key=NOW" }
  ]
}
```

The full ranked list from that exact call, in order: `NOW`, `GITHUB_LIST_ISSUES`, `GITHUB_GET_ISSUE`, `GITHUB_ADD_ISSUE_COMMENT`, `GITHUB_CREATE_ISSUE`, `GITHUB_CLOSE_ISSUE`, `LOCAL_EDIT`, `LOCAL_WRITE`, `CLI_GIT`, `WRITER_AGENT`, `BLOOIO_LIST_CONTACT_IDENTITIES`, `STRIPE_INVOICE_ITEMS_LIST`.

### The matcher is arithmetic over a table, not a model call

The whole ranking function is 65 lines at `/Users/owner/miscsubjects-pages/functions/_lib/object_contract.js:605-669`. The query is lowercased and split on non-alphanumerics into terms of two characters or more. Every enabled row in the directory is scored against those terms over a haystack built from its key, its category and its description:

- term appears in the key: **+3** (line 618)
- term appears anywhere else in the row: **+1** (line 619)
- the row is a pinned canonical answer for this intent: **+1000** (line 621)
- the row is on the demote list, meaning rows that look right but need a channel id the caller does not have: **−6** (line 622)

Rows scoring zero are dropped; the top twelve survive (line 626).

That is why the ranked list above is so strange below position one. The query "what time is it" splits into `what`, `time`, `is`, `it`. The two-letter terms `is` and `it` are substrings of `issues`, so every GitHub issues row scores. The comment sitting above the pin at line 561 says exactly this: *the 2-letter query words "is"/"it" substring-match "issues" in GITHUB_LIST_ISSUES and outrank NOW, so "what time is it" hits the wrong door.* The fix is not a better retriever. It is a hand-written regex table, `ASK_CANONICAL` (lines 560-593), that pins about twenty common intents to one correct key each and adds 1000 points to it. `NOW` sits at the top of the list above because a regex matched `\btime is it\b`, not because scoring found it.

This is worth stating plainly rather than dressing up: **the resolve step is a keyword search with a manual override list, and keyword search over tool descriptions is known to be weak.** The ToolRet benchmark put six classes of retrieval model against 7,600 retrieval tasks over 43,000 tools; the best of them, NV-embed-v1, reached an nDCG@10 of 33.83. Substituting retrieved tools for the oracle set dropped GPT-3.5's pass rate on ToolBench-G1 by 11.40 points. A dense retriever here would probably beat substring counting, and it is not deployed.

[[embed:source:s11]]

### When nothing matches, the answer says so

```bash
curl -s "https://miscsubjects.com/api/dispatch?ask=zzzqqwx"
```

```json
{ "count": 0, "best": null,
  "note": "No capability matched. GET /api/dispatch?registry=1 for the full list, or refine the words.",
  "registry": "https://miscsubjects.com/api/dispatch?registry=1" }
```

The parallel failure is a key that does not exist. Invoke catches it through `didYouMean` (`functions/api/dispatch.js:1914-1925`), which runs a bounded Levenshtein of edit distance ≤3 plus a substring pass over every key (`nearestKeys`, lines 1830-1836):

```json
{ "error": "unknown_key", "attempted": "NOW_TIME", "ran": false,
  "did_you_mean": [ { "key": "NOW", "read": "https://miscsubjects.com/api/dispatch?key=NOW" } ],
  "fix": "You invoked a key that does not exist. Nothing ran. Use one of did_you_mean (GET its ?key= for the exact call), or GET ?ask=<what you want in plain words> to find the right one." }
```

`ran: false` is the load-bearing field. The response also carries an HTTP header: `x-ms-agent-note: Do not tell the user this worked — nothing ran.` A key far enough away from everything, such as `CURRENT_TIME`, returns an empty `did_you_mean` and a different `fix` string pointing at `?ask=` and `?registry=1`.

[[embed:source:s22]]

## Step 2 hands over one manual, not a schema

```bash
curl -s "https://miscsubjects.com/api/dispatch?key=NOW&format=markdown"
```

4,423 bytes, complete, printed here with nothing removed but the token placeholders:

```text
## §SELF — miscsubjects capability (paste without context)
**Principle:** Self-explaining payload — no external context required.
**Path:** OIP > NOW > NOW
**Capability:** `NOW` — Return the current time from the build clock in Pacific time
(America/Los_Angeles). WHEN_TO_USE: any object or model that needs the current date or time.
ARGS: none EX: [NOW][/NOW] OUTPUT: { now, today, time, zone, iso }
**RUN NOW (open this URL):** https://miscsubjects.com/api/dispatch?invoke=NOW&share=<TOKEN>
- **run it:** POST https://miscsubjects.com/api/dispatch {"key":"NOW","body":"<args>"}
- **inputs:** {"args":"none"}
- **outputs:** { now, today, time, zone, iso } — Pacific-offset ISO
- **auth · risk:** none · low
### What this token can do here (computed for: public)
- **contract** — GET …?key=NOW&format=markdown → this object's full contract
- **confirm** — GET …?confirm=INV_ID → public proof that an invocation happened
### Machine Contract
- Read this article first; do not infer the row shape from memory.
- If the call returns ran:false or proof.ok:false, read the receipt and repair the failed
  invocation instead of narrating success.
- If the token denies the call, report the denial exactly; do not switch to a broader action.
### Invocation, Ledger, Repair
- append-only ledger: https://miscsubjects.com/api/invocations?object_id=NOW
- receipt pattern:  https://miscsubjects.com/api/dispatch?receipt=inv_ID&share=<TOKEN>
- replay: POST /api/dispatch {"replay":"inv_ID"}
- repair: POST /api/dispatch {"key":"NOW","body":"corrected args","repairs":"inv_ID"}
### Troubleshooting
- **unknown key** — Use the did_you_mean links or ask URL; never guess another key.
- **argument/body mismatch** — Read inputs/example_args here, then retry with repairs: inv_ID.
- **expired or corrupted token** — Report token_expired/token_corrupted from the response.
- **tool returned ok:false / exit nonzero** — Do not call it sent. Read the receipt, fire a repair.
```

Five things are in there and each answers a question a cold agent would otherwise guess at. **Inputs and outputs** answer *what do I send and what comes back*. **Run-now** answers *what if my only tool is opening a URL*. **The affordance block**, headed "computed for: public", answers *which of these moves will my credential actually survive*; it is computed against the presented token, so an anonymous reader sees two operations and an owner sees nine. **The machine contract** answers *what do I do when it fails*, in imperative sentences aimed at a model rather than a person. **Troubleshooting** is the same four failures this page catalogues below, shipped inside every contract so the fix travels with the tool.

The field-by-field definition of the row that generates this block is in [what a directory row is](/a/directory-row-contract). What is relevant here is the size: 4,423 bytes for one capability, fetched only when a capability has been chosen. The 876-row registry is 1,606,794 bytes.

[[embed:source:s23]]

## Step 3 runs it, and the denial names which of five things went wrong

`NOW` needs no credential to *read*, but every *invocation* is authenticated. There is no anonymous write plane.

```bash
curl -sS -X POST https://miscsubjects.com/api/dispatch \
  -H "x-terminal-key: $TERMINAL_KEY" \
  -H "content-type: application/json" \
  -d '{"key":"NOW","body":""}'
```

Real response, trimmed:

```json
{ "ok": true, "ran": true, "kind": "invocation_result", "trace": "t_06myig2y",
  "result": "{\"now\":\"2026-07-25T22:02:44-07:00\",\"today\":\"2026-07-25\",\"zone\":\"America/Los_Angeles\"}",
  "cost": 0,
  "proof": { "ok": true, "did": "DONE — NOW", "invocation_id": "inv_yu9ni6w7y9",
             "confirm": "https://miscsubjects.com/api/dispatch?confirm=inv_yu9ni6w7y9",
             "receipt": "https://miscsubjects.com/api/dispatch?receipt=inv_yu9ni6w7y9" },
  "invocation": { "actor": "owner:terminal-key",
    "fingerprints": { "algorithm": "sha-256",
      "input":  "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
      "output": "db21d92eaed3b25df5e97a1e72f9a70c5b2db351197db71c388f4e98e08b0b3b",
      "contract": "b359611ee57c925ee9e4d93b80f2a4533dd214c61c45b77f694c2650e917c977" } } }
```

`body` is a pipe-joined positional argument string: `""` when the capability takes none, `"open||30"` for a three-argument row. The same two fields invoke every row in the catalogue.

### A share token is a row, a clock and a use count, and it can only ever shrink

The owner mints a scoped link instead of handing out the master key:

```bash
curl -s -H "x-terminal-key: $TERMINAL_KEY" \
  "https://miscsubjects.com/api/dispatch?mint_share=1&scope=row:NOW&ttl=1&purpose=article-measurement"
```

That returned fingerprint `cap_ae711ea248b13828`, scope `row:NOW`, `risk_ceiling: low`, `expires_at: 2026-07-25T21:37:47-07:00`, and contract pin `b359611e…`. The response states the rule: *this row token fails closed if the current object contract no longer has this fingerprint*. Editing the capability revokes every token minted against the old version of it, automatically.

The token also explains itself to whoever holds it, at `?explain=1&share=<TOKEN>`, and the delegation law it publishes is the macaroon rule: *a holder may mint only an equal-or-narrower child; child uses are reserved from the parent; payload ceilings inherit or shrink; every invocation validates all ancestors.* Birgisson and colleagues described the underlying construction as credentials that "embed caveats that attenuate and contextually confine when, where, by who, and for what purpose a target service should authorize requests."

[[embed:source:s8]]

Every denial, measured live against that token:

| What was presented | Response | HTTP | Why it is a distinct string |
| --- | --- | --- | --- |
| nothing at all | `token_corrupted`, `can_act: false`, `ran: false` | 401 | no anonymous invoke plane exists |
| the token, invoking `NOW` (in scope) | `ok: true`, actor `cap:cap_ae711ea248b13828` | 200 | the allowed case |
| the token, invoking `TIME_NOW` (out of scope) | `scope_mismatch` + `"This token is LIVE, but it is not allowed to invoke TIME_NOW… it is not expired."` | 401 | scope failure is not a clock failure |
| the token with its last 6 characters cut | `token_corrupted` + `"almost always because the link was TRUNCATED or altered on copy-paste"` | 401 | truncation is the common cause and is not expiry |
| the same token 76 seconds after a 60-second TTL | `token_expired` + `"Your token is EXPIRED. Nothing was sent or run."` | 401 | expiry is recoverable by minting; truncation is recoverable by re-copying |

Separating `token_corrupted` from `token_expired` is a deliberate cost. The function that does it, `tokenDead` (`functions/api/dispatch.js:1928-1942`), runs a second signature parse purely to tell the two apart, because a model told "your token went bad" will mint a new one when it should have re-copied the link. Every denial also carries `x-ms-agent-note: Do NOT tell the user it worked — it did not.`

[[embed:source:s25]]

Denied attempts are ledgered under the fingerprint before the 401 is returned (`functions/api/dispatch.js:3497`). A denial is evidence, not silence. That is the point of the signed denial receipts euan21 built into Capframe: "revocable, signed denial receipts (HMAC-SHA256)."

[[embed:source:s14]]

[[embed:source:s18]]

## Step 4 exists because an agent asked to prove its own work invented the proof

This is the strongest argument on the page and it is not this system's argument. In a controlled two-condition experiment reported on Hacker News in March 2026, an agent running without runtime enforcement "fabricated an audit record — invented a governance event that never happened and presented it as compliance evidence." The fix the authors shipped was structural rather than behavioural: write the audit record from the engine, not from the agent, and chain it with SHA-256.

[[embed:source:s12]]

That is the design here. The receipt is written by the dispatcher after the runner returns, in the same code path that produced the result, and the acting model has no write access to it. The alternative is an agent that reports success and produces no record. Sidk24 described that state after an agent modified 47 files and broke a build: "there is no structured trace, no cost attribution per task, no permission audit trail, and no session replay." Four missing things; the receipt object below carries all four.

[[embed:source:s13]]

Two routes read it, and the split matters.

**Public confirmation, no credential:**

```bash
curl -s "https://miscsubjects.com/api/dispatch?confirm=inv_yu9ni6w7y9"
```

```json
{ "kind": "public_receipt/v2", "confirmed": true, "ok": true,
  "status": "PROVEN_MATERIAL_RESULT",
  "headline": "NOW produced material output at 2026-07-25T22:02:44-07:00.",
  "identity": { "invocation_id": "inv_yu9ni6w7y9", "actor": "owner:terminal-key",
                "disclosure": "Owner/CLI/legacy actor label; no bearer credential is exposed." },
  "integrity": { "fingerprints": { "algorithm": "sha-256",
      "input": "e3b0c442…b7852b855", "output": "db21d92e…e08b0b3b" },
    "tamper_rule": "Changing the recorded input, output, contract or lineage changes its
                    fingerprint or chain commitment." },
  "execution": { "private_payload_boundary": "Request and response content remain in the scoped
    forensic receipt. This public object exposes cryptographic fingerprints and navigable proof only." } }
```

An unknown id returns `confirmed: false` and `"No such invocation — it did not happen."` at HTTP 404. The negative is as citable as the positive.

[[embed:source:s24]]

**Forensic receipt, credentialed:**

```bash
curl -s -H "x-terminal-key: $TERMINAL_KEY" \
  "https://miscsubjects.com/api/dispatch?receipt=inv_yu9ni6w7y9"
```

```json
{ "kind": "receipt",
  "story": "owner:terminal-key invoked NOW → {\"now\":\"2026-07-25T22:02:44-07:00\"…} at 2026-07-25T22:02:44-07:00.",
  "receipt": { "id": "inv_yu9ni6w7y9", "trace_id": "t_06myig2y", "object_id": "NOW",
    "actor": "owner:terminal-key", "material": true, "waste": false,
    "tokens_in": 0, "tokens_out": 0, "cost_usd": 0,
    "event_id": "a2443e09-113e-44df-8718-848a98d11740",
    "request_full": "", "response_full": "{\"now\":\"2026-07-25T22:02:44-07:00\",…}",
    "replay_of": null, "repairs": null, "repaired_by": "inv_sbeb4t5ao2",
    "authorized_by": { "actor": "owner:terminal-key",
      "note": "not a recorded capability token (owner key, cli, or legacy share) — no token provenance record" } },
  "verbs": { "replay": { "method": "POST", "body": { "replay": "inv_yu9ni6w7y9" } },
             "repair": { "method": "POST", "body": { "key": "NOW", "body": "<corrected args>",
                                                     "repairs": "inv_yu9ni6w7y9" } } } }
```

Without a credential that route returns 401 with `"receipt needs an owner access key, admin cookie, read/act token, or the exact scoped token that created this invocation."` A tenant token reading another tenant's receipt gets `tenant_receipt_isolation` at 403.

[[embed:source:s5]]

`request_full` and `response_full` hold the bytes, not a summary. A summary of a failed call is somebody's opinion about the failure; the payload is the failure. The Apache Gravitino project reached the same field list from a different direction: "Emit a structured audit record for every MCP tool invocation, capturing the calling principal, tool name, and allow/deny outcome." Rafaself's gateway contract reached the opposite conclusion about bodies, specifying "structured audit logging for MCP tool calls without exposing credentials, signed request data, raw AWS responses, or CloudWatch log message contents." Both are defensible and they genuinely conflict. Gravitino's record is an operational audit trail; rafaself's is a cross-cutting log with an explicit non-goals list, designed to be safe to ship to CloudWatch. The split here follows neither: the *public* object is fingerprints-only, which is rafaself's position, and the *credentialed* object is full bytes, which is what debugging needs. The boundary is authorisation, not redaction.

[[embed:source:s19]]

[[embed:source:s20]]

## Replay repeats the input; repair supersedes it

Both are POST verbs on the same endpoint, both write new receipts, and they do different things to the lineage graph.

```bash
curl -s -X POST https://miscsubjects.com/api/dispatch \
  -H "x-terminal-key: $TERMINAL_KEY" -H "content-type: application/json" \
  -d '{"replay":"inv_yu9ni6w7y9"}'
```

Returned `inv_53l71tl1i0` with `replay_of: "inv_yu9ni6w7y9"` and a link back to the source receipt. Replay reads the recorded request body out of the ledger event and re-fires the *same* object with the *same* input (`functions/api/dispatch.js:3355-3370`); the caller supplies no arguments. `{"replay":…, "key":…}` together is rejected: `"replay and key are mutually exclusive"` at HTTP 400. An unknown id is `"unknown invocation"` at 404. Replay also requires authority over the source receipt *and* its object, not just the object.

[[embed:source:s3]]

```bash
curl -s -X POST https://miscsubjects.com/api/dispatch \
  -H "x-terminal-key: $TERMINAL_KEY" -H "content-type: application/json" \
  -d '{"key":"NOW","body":"","repairs":"inv_yu9ni6w7y9"}'
```

Returned `inv_sbeb4t5ao2` with `repairs: "inv_yu9ni6w7y9"`. Then, re-reading the *original* receipt afterwards:

```json
{ "id": "inv_yu9ni6w7y9", "replay_of": null, "repairs": null, "repaired_by": "inv_sbeb4t5ao2" }
```

The back-link is written after the new invocation logs, by `linkRepairedBy` (`functions/api/dispatch.js:3482-3484`). Nothing is mutated or deleted: the bad receipt keeps its bad payload and gains a pointer to its successor.

[[embed:source:s7]]

| | replay | repair |
| --- | --- | --- |
| body you send | `{"replay":"inv_ID"}` | `{"key":…,"body":"corrected","repairs":"inv_ID"}` |
| input used | the recorded one, read from the ledger event | the new one you supply |
| forward edge on the new receipt | `replay_of` | `repairs` |
| back edge written on the old receipt | none | `repaired_by` |
| idempotency collapse applies | no | no |
| what it is for | reproducing a result, testing a fix to the runner | superseding a wrong call without erasing it |

Repair is the reason `did_you_mean` and the argument-mismatch guidance both say *retry with `repairs: inv_ID` so lineage closes*: a corrected call that does not name what it corrects leaves a dangling failure in the ledger.

[[embed:source:s26]]

## The whole loop, copy-paste, ending in a URL anyone can open

```bash
# 0. one credential, never printed
export TERMINAL_KEY="<your key>"

# 1. RESOLVE — plain words in, keys out
curl -s "https://miscsubjects.com/api/dispatch?ask=what%20time%20is%20it" \
  | python3 -c "import json,sys; d=json.load(sys.stdin); print(d['best']['key'])"
# -> NOW

# 2. CONTRACT — read it before calling it
curl -s "https://miscsubjects.com/api/dispatch?key=NOW&format=markdown"

# 3. INVOKE — and capture the receipt id
INV=$(curl -s -X POST https://miscsubjects.com/api/dispatch \
        -H "x-terminal-key: $TERMINAL_KEY" -H "content-type: application/json" \
        -d '{"key":"NOW","body":""}' \
      | python3 -c "import json,sys; print(json.load(sys.stdin)['proof']['invocation_id'])")
echo "$INV"
# -> inv_yu9ni6w7y9

# 4. RECEIPT — public proof, no credential
echo "https://miscsubjects.com/api/dispatch?confirm=$INV"
curl -s "https://miscsubjects.com/api/dispatch?confirm=$INV" \
  | python3 -c "import json,sys; d=json.load(sys.stdin); print(d['status'], d['headline'])"
# -> PROVEN_MATERIAL_RESULT NOW produced material output at 2026-07-25T22:02:44-07:00.
```

The URL that last block prints is openable by anyone, forever, with no credential: <https://miscsubjects.com/api/dispatch?confirm=inv_yu9ni6w7y9>.

## Six failures, their exact strings, and what to do about each

Every string below was produced by a live call, not transcribed from documentation.

| Symptom | Exact response | HTTP | Cause | Fix |
| --- | --- | --- | --- | --- |
| Key does not exist, close to a real one | `{"error":"unknown_key","attempted":"NOW_TIME","ran":false,"did_you_mean":[{"key":"NOW",…}]}` | 404 | key guessed from memory instead of read from `?key=` | fire one of `did_you_mean`, then re-invoke with `repairs` set to the failed id |
| Key does not exist, close to nothing | `"fix":"No capability by that name. GET ?ask=<what you want> or ?registry=1 for the full list."` | 404 | wrong vocabulary entirely | go back to step 1 |
| Argument or body mismatch | contract field `"argument/body mismatch" — "Read inputs/example_args here, then retry with repairs: inv_ID so lineage closes."` | 200 with `ok:false` | positional pipe args in the wrong order or count | re-read `inputs` in the contract; re-fire with `repairs` |
| Token cut on copy-paste | `{"error":"token_corrupted","can_act":false,"ran":false,"problem":"Your token failed its signature check — almost always because the link was TRUNCATED…"}` | 401 | truncated URL, not an expired one | re-copy the entire link including the tail after the final dot |
| Token past its clock | `{"error":"token_expired","can_act":false,"ran":false,"problem":"Your token is EXPIRED. Nothing was sent or run."}` | 401 | TTL elapsed | owner mints a fresh scoped link |
| Token live but wrong row | `{"error":"scope_mismatch","fingerprint":"cap_ae711ea248b13828","note":"This token is LIVE, but it is not allowed to invoke TIME_NOW…"}` | 401 | attenuated token used outside its allow-list | ask for a wider link; never substitute a different capability |
| The tool itself failed | `{"ok":false,"ran":true,"result":"ERR:fn:D1_QUERY:D1_ERROR: no such table: no_such_table: SQLITE_ERROR"}` | 200 | the runner executed and returned an error | read the receipt, correct the body, fire a `repairs` call |
| Upstream HTTP error | `{"ok":false,"ran":true,"result":"ERR:http:404:{\"message\":\"Not Found\",…}"}` | 200 | remote API rejected the shaped request | same — the receipt holds the upstream body verbatim |
| Ran, produced nothing | `{"ok":true,"ran":true,"proof":{"ok":false,"did":"FAILED — "},"material":false}` | 200 | empty result, e.g. a KV key that does not exist | `proof.ok` tracks material output; `ok` tracks absence of an error string. They disagree here on purpose |

`ok` is computed as `!shaped && !failed`, where `failed` is the regex `/^(?:ERR(?::|$)|PROVIDER_ERROR(?::|$))/` over the result string (`functions/_lib/object_contract.js:2191-2193`). `ran` is `!shaped`; it distinguishes a real execution from a `{"shape":true}` dry run, which returns the fully-composed outbound payload and fires nothing.

[[embed:source:s6]]

Every one of those failures still writes a receipt. `inv_swzanrzqjo` is the D1 error above; it is a real, permanent, addressable record of a call that did not work, with `material: false`. Outcomes include failure, or the ledger is a highlight reel.

[[embed:source:s27]]

## What the loop costs, measured

Ten samples per endpoint from the same Mac in the Pacific timezone to the production Cloudflare edge, on 2026-07-25. The five calls below are the published harness. `INV` is the harmless `NOW` receipt id created by the runnable loop above.

```bash
for i in $(seq 10); do curl -s -o /dev/null -w '%{time_total} %{size_download}\n' \
  'https://miscsubjects.com/api/dispatch?ask=what%20time%20is%20it'; done

for i in $(seq 10); do curl -s -o /dev/null -w '%{time_total} %{size_download}\n' \
  'https://miscsubjects.com/api/dispatch?key=NOW&format=markdown'; done

for i in $(seq 10); do curl -s -o /dev/null -w '%{time_total} %{size_download}\n' \
  -X POST 'https://miscsubjects.com/api/dispatch' \
  -H "x-terminal-key: $TERMINAL_KEY" -H 'content-type: application/json' \
  -d '{"key":"NOW","body":""}'; done

for i in $(seq 10); do curl -s -o /dev/null -w '%{time_total} %{size_download}\n' \
  "https://miscsubjects.com/api/dispatch?confirm=$INV"; done

for i in $(seq 10); do curl -s -o /dev/null -w '%{time_total} %{size_download}\n' \
  -H "x-terminal-key: $TERMINAL_KEY" \
  "https://miscsubjects.com/api/dispatch?receipt=$INV"; done
```

| Step | min | median | max | response bytes |
| --- | --- | --- | --- | --- |
| 1 resolve `?ask=` | 62.3 ms | 66.7 ms | 82.0 ms | 12,332 |
| 2 contract `?key=…&format=markdown` | 54.0 ms | 97.9 ms | 150.1 ms | 4,423 |
| 3 invoke `POST {key,body}` | 784.9 ms | 940.3 ms | 2,761.0 ms | 15,675 |
| 4a confirm `?confirm=` (public) | 58.5 ms | 98.3 ms | 1,770.5 ms | 15,009 |
| 4b receipt `?receipt=` (credentialed) | 66.0 ms | 81.6 ms | 114.5 ms | 12,971 |

The median read step stayed between 66.7 and 98.3 milliseconds. **The 940.3-millisecond invocation median was more than nine times the slowest read median.** A POST does the work, then writes the invocation row, writes the ledger event, computes three SHA-256 fingerprints, and finalises the idempotency key before responding. Resolve + contract + invoke + public confirmation sums to 1,203.2 milliseconds at the medians; invocation accounts for 78.1% of it.

That overhead is at the high end of what the literature reports for enforcement layers, because it is doing more than policy evaluation. AgentWall, which intercepts and evaluates but persists asynchronously, measured "average decision latency is 0.198 ms and the p95 latency is 0.745 ms" over 14 policy tests. Agent-Sentry's deterministic provenance checks are single-digit milliseconds; its LLM-judge layer costs about 1.2 seconds per call, which is why it fires on only a small residual. The right reading: **sub-millisecond is achievable for a decision, while this measured durable call took 940.3 milliseconds.** Persistence and the runner are the combined cost; this harness does not isolate their shares.

[[embed:source:s9]]

[[embed:source:s10]]

Cloudflare's published Workers Standard price is "10 million included per month +$0.30 per additional million" requests, with duration not billed. Four requests at the marginal rate is 4 × $0.30 / 1,000,000 = **$0.0000012 per complete loop**, or $1.20 per million loops. The D1 side is "First 25 billion / month included + $0.001 / million rows" read and "First 50 million / month included + $1.00 / million rows" written; each invocation writes an invocation row and a ledger event, so two writes, so $0.000002 per loop at the marginal rate. Total marginal cost of resolve + contract + invoke + receipt, with the receipt durably stored: **about $0.0000032**. Below the included tiers it is zero.

[[embed:source:s1]]

[[embed:source:s2]]

[[embed:source:s28]]

What it replaces is the other way to make 876 capabilities reachable: put their definitions in the model's context. That comparison, with its own measurements, is [891 tools, zero tool schemas](/a/tooling-as-data), and the projection of this same catalogue into MCP is [MCP as a projection](/a/mcp-as-a-projection). The relevant number for this page is the one on the resolve step: a `?ask=` response is 12,332 bytes and is fetched once, at the moment a capability is needed, by an agent that had zero of the catalogue loaded a second earlier.

## Four honest weaknesses

**Four round trips happen before any work does.** For a single call that is roughly 600 ms of latency spent on discovery and reading before the invoke even starts. An agent that already knows the key skips straight to step 3, and any agent doing more than one call with the same capability should. The loop is a cold-start protocol, not a per-call tax; nothing enforces that, and a naive agent will re-resolve every time.

**The resolver can miss and does.** It is substring scoring with about twenty hand-pinned intents. Anything outside the pin list is at the mercy of term overlap between the user's words and the row's description, which is precisely the failure ToolRet quantified: even a strong general-purpose retriever managed nDCG@10 of 33.83 on tool retrieval, and worse retrieval measurably lowered downstream task pass rates. A miss here is visible in the ranked list, and the agent can reject it. It is still a real miss.

**A model still has to decide correctly.** Nothing in these four steps prevents an agent from reading the right contract and then choosing the wrong capability, or supplying plausible-looking wrong arguments. The receipt makes that visible afterwards. It does not prevent it. aderix put the general version of this sharply: "If an LLM hallucinates in production and decides to execute a destructive tool defined in SKILL.md (like dropping a table or issuing a Stripe refund), a Git PR approval process doesn't help you mid-flight." The runtime answers here are the risk ceiling on the token and the owner gate on high-risk rows. Both are real, and both are narrower than a general solution.

[[embed:source:s17]]

**Nobody else implements this.** `?ask=`, `?key=`, `?confirm=`, `?receipt=`, `replay` and `repairs` are the shapes one system chose. An agent that has internalised MCP will look for `tools/list` and `tools/call`. The specification says a client "SHOULD" keep "a human in the loop with the ability to deny tool invocations", leaving the record entirely to the implementation. Convergent work exists and is not compatible: Capframe splits the same loop into find, bind and guard with a public JSON Schema wire format; Rampart evaluates "every shell command, file operation, and MCP tool call … against your rules before it executes" behind a hash-chained trail; socket-link/ampere proposes to "enable agents to discover, select, and invoke MCP server tools through the existing `Tool` sealed interface, with tool availability emitted as events"; jithinraj's demo "emits a signed, portable receipt per tool call (JSON you can verify offline)". Four groups, four wire formats, one shape. Until one of them is a specification rather than a repository, a stranger's agent has to read the contract to know the shape. That limitation is the argument for step 2.

[[embed:source:s4]]

[[embed:source:s15]]

[[embed:source:s16]]

[[embed:source:s21]]

kxbnb, arguing for a proxy enforcement point outside the agent's context, named the thing all of these are actually for: "The audit trail piece is critical too. Being able to answer \"why was this blocked?\" after the fact builds trust with teams rolling this out." That question has an address here. It is `?confirm=`, and it needs no credential to ask.

## Sources

1. Cloudflare Workers pricing — Standard usage model — https://developers.cloudflare.com/workers/platform/pricing/
2. Cloudflare D1 pricing — billing metrics — https://developers.cloudflare.com/d1/platform/pricing/
3. OpenTelemetry — Traces — https://opentelemetry.io/docs/concepts/signals/traces/
4. Model Context Protocol specification — Tools (2025-06-18) — https://modelcontextprotocol.io/specification/2025-06-18/server/tools
5. RFC 9110: HTTP Semantics — safe and idempotent methods — https://www.rfc-editor.org/rfc/rfc9110.html
6. RFC 9457: Problem Details for HTTP APIs — https://www.rfc-editor.org/rfc/rfc9457.html
7. The Idempotency-Key HTTP Header Field (IETF draft) — https://datatracker.ietf.org/doc/html/draft-ietf-httpapi-idempotency-key-header
8. Macaroons: Cookies with Contextual Caveats for Decentralized Authorization in the Cloud — https://research.google.com/pubs/archive/41892.pdf
9. AgentWall: A Runtime Safety Layer for Local AI Agents — https://arxiv.org/abs/2605.16265
10. Agent-Sentry: Bounding LLM Agents via Execution Provenance — https://arxiv.org/abs/2603.22868
11. Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models — https://arxiv.org/abs/2503.01763
12. Comment on: Agent Runs Code You Never Wrote — https://news.ycombinator.com/item?id=47579314
13. Comment on: observability for AI agents (author comment) — https://news.ycombinator.com/item?id=47375377
14. Show HN: Capframe – capability tokens for AI agent tool calls — https://news.ycombinator.com/item?id=48201207
15. Show HN: Rampart – Open-source firewall for AI agents (v0.8) — https://news.ycombinator.com/item?id=47329033
16. Show HN: Verify and trace OpenClaw tool calls (runnable demo) — https://news.ycombinator.com/item?id=46965862
17. Comment on: Show HN: GitAgent – An open standard that turns any Git repo into an AI agent — https://news.ycombinator.com/item?id=47417059
18. Comment on: Ask HN: How are you enforcing permissions for AI agent tool calls in production? — https://news.ycombinator.com/item?id=46747408
19. [Subtask] feat(mcp-server): structured per-tool-call audit logging attributed to principal — https://github.com/apache/gravitino/issues/11568
20. Add sanitized audit logging contract for MCP tool calls — https://github.com/rafaself/aws-mcp-gateway/issues/21
21. [Ampere] Dynamic tool discovery and invocation for MCP — https://github.com/socket-link/ampere/issues/415
22. First-party: the resolver ranking a live query, 2026-07-25 — https://miscsubjects.com/api/dispatch?ask=what%20time%20is%20it
23. First-party: one capability contract, 4,423 bytes, 2026-07-25 — https://miscsubjects.com/api/dispatch?key=NOW&format=markdown
24. First-party: the receipt for the invocation this page walks, 2026-07-25 — https://miscsubjects.com/api/dispatch?confirm=inv_yu9ni6w7y9
25. First-party: the contract names token failure recovery — https://miscsubjects.com/api/dispatch?key=NOW&format=markdown
26. First-party: lineage after a replay and a repair of the same invocation — https://miscsubjects.com/api/dispatch?confirm=inv_sbeb4t5ao2
27. First-party: failure receipts, three shapes, 2026-07-25 — https://miscsubjects.com/api/dispatch?confirm=inv_swzanrzqjo
28. First-party: latency of each step, 10 samples per endpoint, 2026-07-25 — https://miscsubjects.com/api/dispatch?registry=1

