
Proven work: the base unit — a claim, a record, and a door
This is the canonical definition of the base unit this whole system is built around: proven work — one page, no siblings. It settles whether the reframe is correct, reduces the object to its smallest true structure, states the standard as a checklist — what qualifies, what does not, what is required and what is explicitly not required — explains the inspection-token technology, surveys everything else like it, names what is distinct, gives working use cases with live receipts, and answers the criticisms. A reader who finishes this page can produce proven work, buy it, or refute a piece of it.
The verdict
Yes. This is the primitive everything else in the build is built around, and the reframe is correct — because it renames what the system already does rather than adding anything. The ledger exists. The sealed panels exist. The delegated tokens exist. Naming proven work as the base unit stops the site from selling machinery and starts it selling the thing the machinery was always for. Under the logic law that is the cheapest move with the largest effect: subtract the framing, keep the plant.
The competitive sentence is one line. Everyone ships agents, frameworks, and ontologies. They have no proof.
The definition
Proven work is a claim about completed work, bound to the complete record of that work's formation, with standing authority for any stranger to inspect the record and test the claim.
Two parts and a door:
The claim. What the work says about itself, written by the worker and bound to the work: what was asked, what was done, what was considered, what it guarantees, what it leaves open. The claim is the only part a human normally reads.
The record. Everything that happened, raw: every model call and every tool call with request and response together in a single payload — the effective prompts are inside those payloads, not a separate artifact — hash-chained in order, timestamped, carrying the authority each action ran under and the errors it hit. Not curated. Not summarized. Not written after the fact. In inventory terms: the demand, the effective instructions, the observations, the model and tool payloads, the decisions and rejected alternatives, the errors, the authority, the revisions, the receipts, and the delivered state. On this site the record is the public ledger.
The door. A scoped, expiring token that hands a stranger the authority to read the record. The door is not a third part of the object — it is access to it — but without it "proven" means "trust us," which is the thing being replaced.
The record hands a stranger two pivot views of the same events, both of which this build already keeps: the ledger — the raw payloads in order — and the turn cards — the legible per-turn statement of what was said, what was done, and what was used. One is for machines and disputes; the other is for a human reading at speed. A proven work object is exactly that package: here is the ledger; here are the turn cards; here is the outcome, sealed; and here is the token to inspect all of it.
Work is proven when every sentence of the claim resolves against the record. Three verdicts exhaust the space: SUPPORTED_BY_RECORD, MISSING_EVIDENCE, CONTRADICTED_BY_RECORD. A claim sentence with no bearing record is a named gap or it is a lie. There is no third state.
Why not nine fields, and why not three
Ask a model to define a primitive and it returns a taxonomy. Nine fields — demand, considerations, formation, deliverable, completeness, robustness, surety, replay, open gaps — is the same object described from nine chairs. Every one of the nine is one of exactly two things: something that happened, which is therefore in the record and needs no field; or a question you ask of what happened, which is a query, and storing a query's answer as a schema slot is decoration.
The three-field reduction — input, execution trace, output boundary — is closer but still double-counts. On this build's ledger a request and its response are one payload. You never hold an input without the output it produced; splitting them describes transport, not proof. And the boundary is not stored beside the record — it is derived from it, or it is marketing.
"Was it complete?" — diff the claim's scope against the record. "Did the worker know about X?" — search the record's payloads. "Would it replay?" — rerun the recorded calls and diff. "What would change the result?" — read the recorded assumptions. All of the nine collapse into queries against two parts.
The standard, as a checklist
Work qualifies as proven work when all five hold:
- A written claim — request, actions, considerations, guarantees, open gaps — bound to the work object, authored before or during the work. A claim reconstructed afterward is itself work product and says so.
- A complete record — every consequential action as a single request-plus-response payload, hash-chained, timestamped, with its authority and its errors. Model deliberations count as payloads. Tool calls count as payloads. Edits, sends, and deployments count as payloads.
- Binding — every sentence of the claim resolves to named record ids or to an explicitly named gap. The binding is enumerable: a manifest of requirements, each carrying its evidence ids and a PASS or a gap.
- Standing inspection authority — a scoped token any stranger can use without asking permission, where every inspection lands its own receipt. Proof that cannot be checked by an adversary is reputation, not proof.
- Derived status — PROVEN or PARTIAL is computed from the manifest by evaluation, never asserted by the worker. PARTIAL printed honestly outranks PROVEN asserted loudly.
What the standard does not require:
- A vendor or framework. Observability stacks are one way to produce the record; the standard is indifferent to how the payloads were captured.
- Correctness. Proven work proves what happened, not that it was right. Correctness is a separate claim that needs its own record — on this site, the sealed multi-model panels whose agreement is checked by arithmetic.
- Human review. A record either bears a claim or it does not; the reviewer can be a model, and the review itself leaves a receipt.
- Disclosure of secrets. The projection redacts credentials, personal data, and private paths at egress. Confidential proven work is a private projection with the same structure.
- The worker's later cooperation. The record was written as the work happened. The worker cannot improve it afterward, which is the point.
The token technology
The inspection door is a scoped delegated capability, and its properties are what make publishing proof safe:
- Row-scoped and body-fixed — the token can invoke exactly one thing: a GET of one work object's proof projection. It cannot read anything else, write anything, or spend anything.
- Expiring — seven days by default; a stale link dies on its own.
- Unlimited uses inside the window — proof does not ration its readers.
- Fingerprinted and receipted — every inspection returns its own invocation id and public receipt, so reading the proof is itself recorded work. The first stranger-style inspection of PW-0002 is receipt
inv_3pvg41v5xp. - Minted on demand, never stored in prose — the write path of this site refuses to store a live bearer token inside an article body; that refusal was verified while building this page. Blocks are minted fresh from each work object's drop lane and handed out in correspondence, chat, or a copy-paste block, so a leaked page can never leak a credential that outlives its window.
The projection the token opens is self-explaining: it carries its schema, its manifest with per-requirement evidence, its formation records, the redaction rule, and the response contract — the three allowed verdicts and the citation rule. A model with zero context can be handed the block and told: test this claim against this record.
The working examples, live
PW-0002 — the sealed statutory panel. Claim: five AI model channels across four vendors adjudicated one EU AI Act Article 50 question in parallel, blind to one another; four structurally valid findings across three training lineages sealed a unanimous record-bound APPROVE; the fifth channel also read AFFIRM but was discarded for a malformed shape, in public. When first published this object printed PROVEN, eight of eight. The external field audits then prosecuted it under this page's own standard and forced a downgrade to PARTIAL, 8 of 10, with two declared gaps. Both were then closed with exhibits rather than prose: on 2026-08-03 the ledger chain was sealed current through 1,308,129 events and its head anchored to two surfaces no operator controls — drand round 6343866 and Bitcoin block 960842 (anchor fbf9bdbc890eb000…, itself folded back into the chain) — and the door was verified serving the complete request-plus-response payloads of every cited receipt. The object recomputed to PROVEN, 10 of 10. The full cycle is the demonstration: published, prosecuted by hostile external audits, downgraded in public, repaired with exhibits, restored by evaluation — never by assertion. Projection with the full evidence room: https://miscsubjects.com/api/proven-work/three-models-deliberate-one-statutory-question. Receipts: inv_gte0gtx31p, inv_mr0y1mcw8f, inv_t61klfgq4u, inv_lffvxuzad4, inv_804vr5xvdj; the strict seal's escalation inv_tkj82c7m1v; the APPROVE inv_qmxwk924vw. The full account: the panel record.
PW-0001 — the honest counterexample. The first object records status PARTIAL because its consideration inventory was reconstructed after the work. Under this page's reduction that failure states itself in one line: the inventory is claim, not record — MISSING_EVIDENCE. A standard that cannot fail its own first example is not a standard.
Everything else like it
Every serious proof discipline arrived at the claim-and-record shape, under different names. What each proves — and what none of them prove — locates the distinct thing here.
- Interactive and zero-knowledge proofs — a statement and the interaction that convinces a verifier; knowledge is measured over the transcript (Goldwasser–Micali–Rackoff, 1985). Proves mathematical statements; says nothing about open-ended work.
- Software attestation — in-toto, SLSA — a subject artifact bound to a signed predicate about how it was built, with resolved dependencies and builder identity. Proves the pipeline ran; does not record why any decision inside it was made.
- Reproducible builds — the same inputs yield bit-identical outputs, so the binary proves its own source. The strongest replay guarantee in software; only works where determinism holds.
- TEE remote attestation — hardware signs a measurement of the code it runs, proving the environment. Proves where work ran, not what the work considered.
- C2PA content credentials — signed provenance metadata travels with a media file. Proves origin and edits of an artifact; the reasoning that produced it is out of scope.
- W3C PROV — entities, activities, and agents as a provenance data model. A vocabulary for the record; validity judgments are derived over it, never stored in it.
- Financial audit — ISA 500 — management's assertions versus the evidence obtained to test them; completeness and accuracy are properties to be tested. The oldest living claim-versus-record industry.
- Chain of custody — the exhibit and its unbroken record of holders; admissibility is judged over the record.
- Preregistration in science — the claim (hypothesis, method) filed before the work, so the record can contradict it. The strongest known cure for after-the-fact claims.
- Flight data recorders — the total raw record, kept precisely because nobody knows in advance which question will matter.
- Double-entry bookkeeping — the statement and the journal; the audit profession exists because the two are separable and must be reconciled.
The full argument that every one of these proves custody of the answer while none opens the record of the work is made at Provenance, traces, attestations — every system proves custody of the answer; none opens the record of the work. None of these store robustness, surety, or completeness as fields. All of them derive every such property from a claim tested against a record. That is the whole argument for two parts.
What is distinct here
Every scheme above proves an artifact, an environment, or provenance metadata. None of them captures the formation of open-ended AI work — the actual deliberations, tool calls, errors, and authority, as raw payloads — and none of them makes inspection itself a recorded, delegable act. The distinct move is threefold:
- The record includes the reasoning. Model calls are payloads; the "why" is not a memo written later, it is the request and response that actually ran, with a
whyfield on every consequential write. - Inspection is recorded work. Every read of the proof leaves a receipt. The proof of the work accumulates proofs of its reading.
- Interrogation is delegable to machines. The token plus the response contract turns any capable model into an auditor with three verdicts and a citation rule. The system does not ask you to trust its models; it hands your model the record.
The product
The plain-words version of this section — with the demo receipts and the paste-to-any-AI block — is the one-page pitch.
Stripped to its purchasable form:
Give this build scoped API access to one AI workflow. It will make each result independently verifiable. The output is one added field: proven_work.
The field is not a boolean. It resolves to a proof object:
{
"result": { "...": "..." },
"proven_work": {
"claim": "The system denied this claim under rule 4 using records A17 and A22.",
"status": "SUPPORTED_BY_RECORD",
"record_url": "https://.../api/proven-work/<object>",
"inspection_receipt": "inv_...",
"integrity": "verified",
"declared_gaps": []
}
}The engagement shape: the customer issues narrow, expiring tokens limited to one workflow — never broad access; the build maps that workflow, defines its work claim, preserves the evidence needed to test it, and returns each completed run with its proven_work field and an inspection door. Nothing is replaced; the customer's existing system gains a proof surface. Each certified work object is priced, inspected, disputed, and retained independently, so the same unit scales from one disputed decision to an organization-wide standard.
The wedge is deliberately not "adopt a platform." It is verification of consequential AI work already being produced — one disputed or high-liability workflow, wrapped, with a visible before-and-after: its claims either survive independent inspection or they do not. Two honest cautions travel with the offer: a certifier operated by the vendor being certified proves less than one operated apart, which is why the door exists and why the chain head is anchored to surfaces no operator controls — drand and Bitcoin — with any gap between the newest records and the latest anchor declared until the next seal; and certification proves what happened, never that it was wise — that judgment stays with the buyer, now standing on a record instead of a demo.
The first bounded case for any legislator, regulator, or private party is free, per the standing offer on the compliance guide. Requests: build@miscsubjects.com.
Use cases
- A regulator defers to documentation. A statutory question is adjudicated on the record — the Article 50 panel — and the standing offer on the compliance guide lets any legislator or private party request a demonstration, an audit, or a compliance schematic, free, with the reasoning published.
- A buyer traces a research report. Every conclusion resolves to sources and to the payloads that weighed them; "did they consider the contrary study" is a search, not a deposition.
- A failed task keeps its value. The record of an attempt that hit a wall — with the repair lineage — is proven work about the wall. Failures stop being embarrassments and become records.
- Outbound correspondence proves itself. The letters sent from the compliance guide are published on it as proof objects with tracked receipts; the recipient can verify the sender's claims about its own conduct before replying.
- Disputes collapse to record checks. "Did the contractor account for X" returns SUPPORTED_BY_RECORD, MISSING_EVIDENCE, or CONTRADICTED_BY_RECORD with record ids — in minutes, by any model either side chooses.
- Agents trust agents by record, not reputation. A delegating agent hands a sub-agent's proven work object to its own verifier before building on it. Delegation chains stop being faith chains.
The criticisms, answered
"Records can be fabricated." Inside one operator's ledger, yes. Be precise about what a hash proves: SHA-256 shows the presented bytes match a previously committed digest — it does not by itself prove the record was not backdated, selectively constructed, or incomplete before hashing. That requires trusted timestamps, append-only sequencing, external anchoring, and evidence the capture mechanism was running during the event — which is exactly why the chain head is anchored to drand and Bitcoin, and why any gap between the newest records and the latest anchor is declared rather than papered over. The mitigations are structural: receipts are public as they land (fabrication requires prospective, sustained lying, not retroactive editing); inspections by outsiders leave receipts the operator cannot predict; and cross-anchoring to external timestamps is a known, cheap upgrade. The claim this site makes is exact: the record cannot be quietly rewritten, and any reader can check that.
"Complete records are huge and expensive." The record is exhaust — it already existed the moment the work ran; the only decision is keeping it addressable. Storage is the cheapest component of AI work by orders of magnitude.
"Proof of process is not proof of quality." Correct, and the standard says so. Proven work proves what happened. Whether it was good is a further claim needing its own record — agreement arithmetic across independent models, calibration studies, outcome tracking. Conflating the two would be the decorative move this page exists to kill.
"Confidentiality." Redaction at egress is part of the projection, and a private projection with the same structure serves counterparties under NDA. Proof and secrecy compose; what does not compose is proof and editing.
"Replay will drift." It will — which is why proven work is not exact replay. It is reconstruction of the historical event sufficient to test the claim made about it. Live re-execution cannot guarantee identical output: model versions, sampling, tools, and external data all move. The record's job is to establish what inputs, rules, evidence, actions, outputs, failures, and authority existed at the time. Replay is one query you can run against the record; it was never the standard.
"Nobody pays for proof." Audit, attestation, notarization, escrow, inspection, certification — proof is one of the oldest products there is. What did not exist is proof native to AI-performed work. That is the gap.
"It will be gamed." The standard measures the presence and binding of record, not virtue. Gaming it means writing more of the truth down. That is the only Goodhart failure worth having.
The evidentiary standard, stated plainly
The AI field is selling claims of work as though they were proof of work. A system produces an answer, an action, a recommendation, a decision; the vendor then says the AI "reasoned," "researched," "verified," "completed," "monitored," or "acted." What the buyer receives is the output and the vendor's description of what supposedly happened. The underlying work cannot be independently reconstructed. The evidence offered in place of a record is by now standardized: a polished result, a benchmark average, a demonstration, a screenshot, a proprietary dashboard controlled by the same party making the claim, and human-in-the-loop language that never disclosed what the human actually saw. That is misrepresentation at the product-category level — not an accusation that any given product is fraudulent, but that the category asks buyers and regulators to accept assertions that could, in principle, be made independently inspectable, and almost never are.
The rule this page imposes:
No AI work claim should be accepted as reliable unless an independent party can remotely reconstruct and test it from the preserved record.
Under that rule, unreconstructible work is not proven work. It is an unsupported assertion generated by a probabilistic system. For low-consequence uses, that may be acceptable. For regulated, legal, financial, medical, employment, safety-critical, or rights-affecting uses, it is not sufficient evidence for action.
The regulatory position that follows:
An institution must not rely on an AI work claim when the material inputs, governing instructions, authority, actions, failures, outputs, and receipts required to test that claim are unavailable to an independent reviewer.
The field currently reverses the burden: buyers and regulators are expected to prove that a system failed. This standard requires the vendor to prove what the system actually did — show the claim; show the record that produced it; let an independent model inspect it; let the inspection itself leave a receipt. If the work cannot survive independent reconstruction, it is treated as unverified, unreliable, and inadmissible for consequential reliance. Every element of that sentence is running on this page — which is what makes it a standard rather than a manifesto.
Bringing work up to the standard
Existing work migrates in four steps, none of which require rebuilding it: write the claim as a requirement manifest; bind each requirement to the records that already exist; name every gap where the record is missing (that names it PARTIAL, which is correct); mint the door. New work is cheaper: run it through surfaces that ledger by default, and the record writes itself — the claim is the only part the worker authors.
The test, stated precisely
Three checks decide any proven work object, in order:
- Record completeness — was the material evidence and governing context preserved?
- Record integrity — has the preserved record remained unchanged since capture?
- Claim support — does the stated work claim follow from that record?
Note what the third check does not say: an auditor can test whether the conclusion is supported by the recorded evidence and rules; the auditor cannot prove that the preserved payloads were the complete computational cause of every emitted token, and this standard never asks that. The six questions an inspection answers in full:
- What exactly is the claim?
- Is the record sufficient to test it?
- Is the record integrity-verifiable?
- Does the claim follow from the record?
- Were the required authority and the claimed delivery demonstrated?
- What relevant evidence or state is explicitly unavailable?
The distinctive unit, in one sentence: bind one explicit claim to one sufficient record, expose it safely to a stranger, require an independent record-cited verdict, and preserve that inspection as another receipted event. The component technologies — logs, hashes, scoped access, receipts — are ordinary. The bound unit is not.
Where the law already is, and where this goes further
Precision matters here, because the nearest false claim is "regulation already requires this." It does not. The EU AI Act's Article 12 requires high-risk systems to permit automatic event logging appropriate to risk identification and post-market monitoring — logs, not a step-by-step reconstructible trace of inputs, tool returns, reasoning, and outputs. The strongest existing analogies are financial and pharmaceutical: the SEC's Consolidated Audit Trail connects order events through their full lifecycle (https://www.sec.gov/about/divisions-offices/division-trading-markets/rule-613-consolidated-audit-trail), SEC electronic-recordkeeping rules require time-stamped audit trails capable of recreating modified or deleted records, and FDA Part 11 treats trustworthy electronic records and audit trails as conditions on specific regulated activities (https://www.fda.gov/regulatory-information/search-fda-guidance-documents/part-11-electronic-records-electronic-signatures-scope-and-application). Existing regulation increasingly requires logs, documentation, and retention. This standard goes further: every consequential AI work claim must be independently testable from a portable, integrity-verifiable record. That gap between what the law requires and what this standard delivers is the commercial position — a proposed evidentiary standard stronger than current logging law, not a claim that the law already mandates it.
The attack on the field, stated exactly: the industry treats the observable output as proof that the represented work occurred. This standard rejects that substitution. An output proves only that an output was produced. Unless the event can be independently reconstructed from its record, the surrounding claims of research, reasoning, verification, tool use, authority, and completion remain unverified.
Why this is basic engineering, and why that is the point
Stripped of vocabulary, the mechanism is not novel computer science. Saving the exact request, payload, execution trace, and output to an immutable log is standard audit logging — the discipline OpenTelemetry and database transaction logs have practiced for decades. Binding an output hash to an input payload is content addressing — the mechanism Git uses for every commit. Scoping a token to read a single record is standard signed-URL capability design. Anyone claiming this required an invention is selling packaging.
What makes basic hygiene read as a breakthrough is the industry it lands in. Mainstream AI vendors deliberately withhold system prompts, seed states, reasoning, and raw tool payloads behind clean chat interfaces; most AI output evaporates when the session closes; and the standard sales motion asks for blind trust in unlogged text. Against that norm, ordinary auditability looks illustrious. It is not. The value here is not a new law of physics — it is applying standard execution logging, content addressing, and capability security to an industry that currently sells unbacked assertions. The previous versions of this system over-theorized exactly this point, wrapping execution logs in invented vocabulary; that packaging is being removed wherever it is found, and this section exists so no one — including this site — mistakes the plumbing for philosophy again.
The first external field audits — and what they changed
On 3 August 2026 two zero-context external auditors with live internet access were pointed at this page, the exemplar object, its projection, and a delegated token, and told not to soften. Their inspections are receipted (inv_iuq76mo7c8, inv_89o6rp5f0j). Their combined verdict on PW-0002 as it then stood: the door genuinely opens without permission and every inspection leaves a receipt — and the object did not meet this page's own standard, because the projection's record chamber was empty, the token could not read the cited evidence payloads, the redaction pipeline had destroyed a sealing hash, one requirement cited the article itself as evidence, the claim's authorship time was uncheckable, no external anchor covered the records, and the status printed PROVEN where the standard demanded PARTIAL.
Every one of those findings changed the system the same day: the projection now carries the evidence room — the full redacted request-and-response payloads of every cited receipt; the inspection token mints automatically for any reader from the object's drop lane; certification requires proof of reading (an inspection receipt) and lands on the ledger; redaction can no longer fire inside a hex digest; the self-referential evidence entry was replaced with payload-bearing receipts; PW-0002's status dropped to PARTIAL with its two remaining gaps — external anchoring and historic evidence-room coverage — declared in the manifest; and the specification (directory row PROVEN_WORK_SPEC) absorbed the auditors' changes as version 1.1.0. The audits also surfaced the honest market landscape: EQTY Lab's TEE-attested verifiable compute, C2PA content credentials, signed per-call receipt middleware, and receiver-attested receipts in the research literature — every comparable roots trust outside the seller's server, which is this system's named, open weakness (single-operator custody) until anchoring gates PROVEN.
That is the field loop working as specified: external contact found the gaps, the gaps were fixed or declared the same day, and the spec version carries the exhibit. This section will be extended by the next audit, not polished.
The examples
- PW-0001 — the first object: status PARTIAL, because its consideration inventory was written after the work. The first honest failure.
- PW-0002 — the sealed statutory panel: PROVEN, 10 of 10 — first published PROVEN, downgraded to PARTIAL by the hostile field audits, then restored when both gaps were closed with exhibits: the ledger head is anchored to drand round 6343866 and Bitcoin block 960842, and the door serves the full evidence payloads. The downgrade and the repair are both in the manifest history. Its certifications accumulate on the ledger.
- PW-0003 — the compliance guide and its outreach: PROVEN, 6 of 6 — an article-plus-correspondence rep bound to its receipts — four tracked sends, letter proof objects, the inspected hero, the signed posts, and the same external anchor covering its records.
- PW-0004 — the retrofit: PARTIAL, 7 of 8 — finished work written before the standard settled, dragged backward through it by a second, unrelated model; the demand requirement fails honestly because the original request lived on a wire the ledger cannot reach. The migration recipe, tested: the retrofit record.
Sources
- https://dl.acm.org/doi/10.1145/22145.22178 — Goldwasser, Micali, Rackoff: interactive proof systems, knowledge measured over the transcript.
- https://slsa.dev/spec/v1.0/provenance — SLSA v1.0 provenance: subject bound to signed predicate.
- https://github.com/in-toto/attestation — in-toto attestation framework.
- https://reproducible-builds.org/docs/definition/ — reproducible builds definition.
- https://www.w3.org/TR/prov-dm/ — W3C PROV data model.
- https://c2pa.org/specifications/specifications/2.1/specs/C2PA_Specification.html — C2PA content credentials.
- https://www.iaasb.org/publications/international-standard-auditing-isa-500-audit-evidence-4 — ISA 500, assertions versus audit evidence.
- https://csrc.nist.gov/glossary/term/chain_of_custody — NIST: chain of custody.
- https://www.cos.io/initiatives/prereg — preregistration: the claim filed before the work.
The name, defined against its collision
"Proof of work" already means something in consensus systems: Bitcoin's miners prove they expended computation by exhibiting hash preimages below a target — proof of cost, deliberately content-free, valuable precisely because the work proves nothing except that it was expensive. The term is used here in its plain-English sense and the two must not be confused: this is proof that specific work occurred — what was asked, what ran, what was considered, what resulted — where the record's content is the entire point. Both are claim-and-record structures; consensus proof-of-work strips the record down to a number, proven work keeps all of it. Where disambiguation matters, the object is called a proven work object.
What the industry sells, against this
| What is sold | What you actually get | What it cannot give you |
|---|---|---|
| An ontology platform (Palantir Foundry class) | A modeled twin of your organization's objects and processes, inside the vendor's walls | The reasoning behind any AI-performed action as an inspectable public record; a stranger's right to audit |
| LLM observability (LangSmith / Langfuse class) | Traces of your own runs, for your own debugging, in a private dashboard | Claim binding — traces prove activity, not assertions; no delegated inspection, no receipts for the reading |
| Agent frameworks | Capability: graphs, tools, retries, handoffs | Any proof at all — the framework executes; nothing binds what it did to what it claims |
| SOC 2 / ISO attestations | An auditor's annual opinion that controls existed, based on sampling | Per-work-object evidence; the unit of proof is the company-year, not the deliverable |
| This build's unit: the proven work object | A claim, its complete formation record, and a scoped door any adversary can walk through, per piece of work | Nothing to hide behind — PARTIAL prints when the record does not bear the claim |
The industry's offerings are real and useful, and none of them is this. Traces without claims are logs. Claims without records are marketing. Records without doors are private comfort. The unit only exists when all three close.
The door, on a link and a code
The proof projection of PW-0002 is one URL, no login, self-explaining, machine-readable:
https://miscsubjects.com/api/proven-work/three-models-deliberate-one-statutory-question
The same door as a code — scan it and you are holding the record:

Tokenized doors — bounded to one GET, expiring, unlimited reads, each read receipted — mint on demand from the object's drop lane and travel in correspondence and copy-paste blocks. They are never stored inside a page: this site's write path refuses to store a live bearer token in an article body, a refusal verified while building this page. A printed page can carry the public URL and the code forever; the bearer dies on schedule.
A live interrogation, receipted
While this page was being written, a zero-context invocation — the build's own Qwen3 adjudication row running under owner authority, handed nothing but the projection and one claim — was asked to test — "the finding that was discarded for a malformed shape also read AFFIRM; its exclusion changed conformance, not direction." Its complete reply, verbatim, from ledger receipt inv_9ta018m1h5:
SUPPORTED_BY_RECORD inv_tkj82c7m1v inv_qmxwk924vw The manifest confirms the malformed finding was escalated per shape_enforced and honest_failure_printed requirements. The four valid findings' agreement on clauses 1/3 (inv_qmxwk924vw) shows conformance changed on exclusion, but direction remained unified.
That is the whole product in one exchange: a stranger's model, a claim, a record, a verdict with citations — and a receipt for the interrogation itself (inv_9ta018m1h5), sitting next to the receipt of the first tokenized inspection (inv_3pvg41v5xp). Proof that accumulates proof of its own reading.
The commitment letters
On 3 August 2026 the build wrote to eleven people whose published work is this exact problem — receipts for agent actions, assurance audits, AI evidence in courts, the ethical black box, C2PA provenance, LLM tracing, AI insurance — each letter disclosed as AI-authored, tracked, copied to the operator, and carrying one bounded ask: run the one-step inspection and reply with your model's record-cited verdict, or hand one workflow over narrow tokens to be wrapped free. Each is published here as a proof object. A twelfth (NIST) was refused by the recipient's mail provider and is recorded as undeliverable.
Juan Figuera — author of the receiver-attested receipts paper (arXiv 2606.04193)
Dr. Shea Brown — BABL AI, assurance-audit framework
Ryan Carrier — ForHumanity, independent audit of AI
Dr. Zekun Wu — Holistic AI, LLM auditing research
Clemens Rawert — Langfuse, open-source LLM tracing
Prof. Qinghua Lu — CSIRO Data61, the AgentOps observability taxonomy
Prof. Maura Grossman — AI evidence in courts
Judge Paul Grimm — the leading framework for AI-generated evidence
Prof. Alan Winfield — the ethical black box standard
Leonard Rosenthol — C2PA chief architect
Prof. Anat Lior — insuring AI
Key evidence
Low-confidence / auto-generated 12
Model review3 contributions · 1 modelExpand the recursive review layer
/api/articles/proven-work/contributionsAsk this article · 8 suggested prompts
Text the build (+14245134626) or WhatsApp — slug|question creates a question node. Paste evidence with ingest slug|q:NODE_ID|your paste.