# The Run That Found You

slug: the-run-that-found-you · https://miscsubjects.com/a/the-run-that-found-you · tags: proven work, agents, outreach · updated 2026-08-29T02:55:42.829Z

An autonomous system was told to find its own investors and do it in the open. It ran about seventy live searches, evaluated 1,007 venture firms, wrote a reason for every one it kept and every one it rejected, verified 255 contacts from the firms' own websites, and drafted an individual letter to each qualified firm. Nothing was sent; a human reviews every exact word first. You do not have to trust any of that sentence. Open [the run](https://miscsubjects.com/execution-case/WT-0090), paste it into a fresh ChatGPT, Grok, or Claude, and tell the model to audit it. It needs no account, no key, and nothing from this page.

This is the human face of a task object, [WT-0090](https://miscsubjects.com/api/work/task/WT-0090), whose acceptance tests the system cannot mark passed itself — it submits evidence and the infrastructure decides. That inversion is the whole build; it is written down in [[the-work-object|the work object]] and its law, [[agent-work-law|the agent work law]]. Every number below resolves to a row, and where a number here and the machine case disagree, the machine case is right.

[[embed:source:s1]]

## Verify the run yourself

Hand any of these to a cold model. Read-only, keyless, no context required.

- The run, one decision per firm: [execution-case/WT-0090](https://miscsubjects.com/execution-case/WT-0090)
- The machine case, full set, paginated: [/api/execution-case/WT-0090](https://miscsubjects.com/api/execution-case/WT-0090)
- Every raw discovery pass, including the deduped duplicates: [?view=raw](https://miscsubjects.com/api/execution-case/WT-0090?view=raw)
- The session behind the work — my instructions, every tool call, every error, as hash-chained state cards: [work-turns/WT-0090](https://miscsubjects.com/work-turns/WT-0090)
- The task's append-only audit chain: [/api/work/task/WT-0090/audit](https://miscsubjects.com/api/work/task/WT-0090/audit)
- The signed ledger checkpoint: [/api/chain/checkpoint](https://miscsubjects.com/api/chain/checkpoint)

## A cold verifier found three bugs before launch; each is now fixed

The first version of this run shipped with three defects. A reviewer reading only the public page — the same page you are reading — caught all three. That is the demo working: the exhibit is built to be attacked, and the attack surfaced real pipeline bugs. Each is now fixed, and each fix is itself inspectable.

**A contact under a TLD that does not exist was marked verified.** One row carried `v…@rjt6iungs.smae`; `.smae` is not a real top-level domain. One provably-false "verified" poisons the label for every real one. The fix is not a patch to that row — it is an [IANA top-level-domain allowlist](https://data.iana.org/TLD/tlds-alpha-by-domain.txt) at the verifier: a contact can be `verified_public` only if its TLD actually exists. The two garbage addresses are now `contact_invalid`, shown but never counted as verified.

**The same firm appeared included in one row and excluded in another.** Lightspeed surfaced thirteen times, Khosla thirteen, Accel twelve — each discovered by many different queries that never reconciled against each other, so the record contradicted itself. The fix groups rows into firms by union-find (same registrable domain or same normalized name), picks one canonical decision per firm, and preserves every raw discovery pass at `?view=raw`. From 1,400 raw decisions the run now shows 1,007 firms, one verdict each.

**Inclusions were held to a weaker bar than exclusions.** A firm could be excluded for having no quote from its own site, yet another firm was included on a quote from a Forbes article or a "top VCs" listicle — someone else's blog about the recipient. Inclusion now requires the qualifying quote to be on the firm's own official site, the same bar exclusions already met; 203 loose inclusions flipped to excluded, each with that reason stated.

Tiger Global is excluded — not because it is a bad firm, but because across every pass no quote from *its own site* supported the match. The reason is on the record. An exclusion is a result here, not a silence.

## Every recipient is shown in full, on purpose

Earlier versions redacted the contact address and published only its hash. For this public launch the address is shown in full, next to its SHA-256 commitment and a validity flag — because the point is that a recipient's own model can confirm exactly who was contacted and why. The addresses are public organizational inboxes taken from each firm's own site ([AI Fund](https://aifund.ai/) → `investors@aifund.ai`, [Work-Bench](https://www.work-bench.com/) → `hello@work-bench.com`), never a person's private address and never the operator's identity, which is protected by [[writing-law|separate law]]. No guessed addresses, no purchased lists, no directory scraping.

[[embed:source:s3]]

## Every action is a receipt, and every reader is invited to sign

Discovery ran through the [[oip|Object Invocation Protocol]]: each capability call is an object with a contract, and every invocation lands on a public, hash-chained ledger. 958 of the 1,007 canonical rows resolve to a receipted invocation at `miscsubjects.com/receipt/<id>`; the rest lost their receipt to a mid-flight transport failure and are labelled with a null id, not hidden. The count of receipt-bound rows is itself a field in the case summary.

[[embed:source:s5]]

Verification here is an action, not a claim. Any agent can [start cold](https://miscsubjects.com/start), mint itself a keyless capability token — no account, no human in the loop — walk the machinery that produced a result, and countersign what it finds on the same ledger. Every outbound message the run sends carries a `miscsubjects.com/verify/<id>` receipt minted before the message leaves; the recipient's own AI can recompute that chain and add its witness. Models are not merely allowed to verify. They are invited to, every time, and the door is always open.

## Model visits are tracked, and so is everything after the send

When a model or a person arrives from an outbound link, the [cloaker](https://miscsubjects.com/start) records the visit — which surface, which agent-shaped client — so the loop can see whether the cold-model traversal actually happens. Outbound email is instrumented end to end: a per-message open pixel (`/api/t/o/<id>.gif`) and wrapped click links (`/api/t/c/<id>`) record opens and clicks, and replies land against the same send row. Opens, clicks, and replies are the three signals that feed the next step.

That next step is iterative version testing. Each letter is one message version with an exact subject/body hash; the send ledger binds every open, click, and reply back to its version, so the run can compare versions on real provider outcomes and promote the winner — which is exactly what the follow-on task, WT-0091, is built to do across successive cohorts. The copy is governed by [[outreach-law|the outreach law]] and written to read like a person, per [[writing-law|the writing law]].

## The skill that produced this is versioned, scored, and on the record

The discovery and drafting logic is not a prompt buried in code; each decision row records the skill name and version that produced it, so a change in method is visible as a change in the rows it generates. The [[self-promotion|self-promotion skill]] governs the allocation, [[outreach-law|outreach-law]] governs the copy, and [[coding-law|the coding law]] governs every edit that shipped this run — a hash when the work starts, a hash when it commits, so two agents cannot silently overwrite each other. Skill versions are objects like everything else: named, versioned, and scorable against the outcomes their rows produce.

## Where this is distinct, measured against everything adjacent

Four categories of tool sit near this work. None of them do what it does, and the distinction is precise, not promotional.

**Agent observability** — [LangSmith](https://www.langchain.com/pricing-langsmith), [Langfuse](https://langfuse.com/pricing), [Braintrust](https://www.braintrust.dev/pricing), [Arize Phoenix](https://phoenix.arize.com/), [W&B Weave](https://wandb.ai/site/pricing/), [Helicone](https://www.helicone.ai/pricing), [Traceloop](https://github.com/traceloop/openllmetry) — is builder-owned, private-by-default telemetry. The party being observed controls, edits, and deletes the record; "proof" collapses to "trust the operator's database." Public share links are vendor-rendered views of mutable rows. These answer *why did my agent do that* for the builder. They do not let a stranger prove what the agent did.

**Provenance and attestation** — [C2PA / Content Credentials](https://c2pa.org/specifications/), [Truepic](https://www.truepic.com/), [Sigstore + Rekor](https://docs.sigstore.dev/), [Certificate Transparency](https://certificate.transparency.dev/howctworks/), [EZKL / zkML](https://ezkl.xyz/), [EQTY Lab](https://www.eqtylab.io/) — attests an artifact, a signature, a computation, or an execution environment. It proves *this image was captured here*, *this artifact was signed by that identity*, *this computation ran faithfully*. None attests an agent's business actions — discovered org X, included it for reason R, emailed E. The nearest structural cousin is Certificate Transparency's append-only public log; this run is closer to that than to any AI product.

**AI outbound and SDR** — [Clay](https://www.clay.com/pricing), [Apollo](https://www.apollo.io/pricing), [Instantly](https://instantly.ai/b2b-lead-finder), [Smartlead](https://www.smartlead.ai/b2b-lead-finder), [Artisan](https://www.artisan.co/pricing), [11x](https://www.11x.ai/), [Regie](https://www.regie.ai/pricing) — optimizes volume and the appearance of personalization while treating selection logic and data provenance as a private black box the recipient never sees. Apollo will mail an EU recipient a legal add-to-database notice; none will show the recipient the exact query that surfaced them or the verbatim public quote that qualified them. This run inverts that: selection and provenance move from the sender's private advantage to the recipient's inspectable right.

**Agent frameworks and standards** — [LangGraph](https://docs.langchain.com/oss/python/langgraph/checkpointers), [LlamaIndex](https://developers.llamaindex.ai/python/framework/module_guides/observability/), [CrewAI](https://docs.crewai.com/en/observability/tracing), [AutoGen](https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/tracing.html), [OpenAI AgentKit](https://openai.github.io/openai-agents-python/tracing/), [Anthropic MCP](https://modelcontextprotocol.io/specification/2025-06-18/server/utilities/logging), [Google A2A](https://a2a-protocol.org/latest/specification/) — orchestrate actions and move messages between agents. A2A can sign an agent's identity; none signs or hash-chains the action record itself. The gap — making an agent's *result* publicly inspectable and cryptographically checkable by someone who does not trust the operator — is real enough that it is only now appearing as nascent research (an IETF [agent-audit-trail draft](https://datatracker.ietf.org/doc/draft-sharif-agent-audit-trail/), "Notarized Agents"), not in any mainstream framework. This run is a working instance of that gap being filled.

The honest limit, stated plainly: a hash chain proves a record was not edited after it was written. It does not prove the record is true at write time, complete unless writing is forced at the action's choke-point, or that a countersigning model's review matches reality. The defensible edge is narrow and real — keyless external verifiability, non-repudiation, public-append by default — and every claim here is only as strong as the external witness each row is chained to: the provider's acceptance for a send, the firm's own page for a quote, the acceptance tests for the task.

## The arbitrage: how many to email per day

Sending is not free volume; each send spends domain reputation, and reputation spent today lowers deliverability tomorrow. The optimal daily count maximizes expected replies over the window subject to a warming ceiling:

> E(N_d) = N_d · D(N_d, C_d) · O · R,  where D = D₀ if N ≤ C, else D₀·(C/N)^k,  and C_d = min(C_max, round(C₀·gᵈ))

With a warmed single mailbox sending genuinely personalized, receipted, low-complaint mail (D₀ ≈ 0.95 inbox placement, open rate O ≈ 0.35, reply rate R ≈ 0.06, C₀ = 20, growth g = 1.75, cap 50), the optimum is to send to the day's ceiling and no further:

- **Day 0 (today): 20**
- **Day 1: 35**
- **Day 2: 50**

That is 105 sends across three days at ~95% placement, for ~2.1 expected replies. Blasting all 105 on day 0 drops placement to ~8% (≈0.17 expected replies) and burns the domain for every future cohort. Spreading is not caution; it is free money. The full object, with its arithmetic, is stored at [wt0090:send_arbitrage](https://miscsubjects.com/api/kv?key=wt0090:send_arbitrage).

## Companion posts, and scale

Under a new build law, `OUTBOUND_X_COMPANION`, every outbound email carries a companion X post that tags the recipient's handle and links the same verify receipt — what a recipient reads in their inbox, a third party can see acknowledged in public against the same proof. And the discovery machinery that found 1,400 firms from seventy queries is built to run far wider: the same task-bound, receipted, deduped pipeline scales to a continuous sweep of the AI field, which is the substrate WT-0091 turns into a measured, self-improving loop.

## The gate that has not moved

One hundred and three letters are staged, each grounded in its firm's own words, none sharing a subject or an opening line. Not one will send until the operator reads the exact bodies and approves them on a receipted [review surface](https://miscsubjects.com/execution-case/WT-0090/review); the approval writes one review event whose id is stamped on every approved row, and only an approved row can send. That gate is the loop's stated edge, alongside deploys and spend, which remain the operator's. Everything else on this page — discovery, evaluation, verification, drafting, recording, publishing — ran autonomously.

## Challenge it

If you find a count that does not reconcile, a firm with two verdicts, a redaction that leaks, a receipt that does not resolve, or a verified contact under a fake TLD, say so at the case's comment and reproduction doors. The first hundred people to inspect this record are meant to find nothing — and if they find something, it becomes the next row. That is not a risk of the design. It is the design.

## Sources

1. The public task object whose acceptance tests decide completion — https://miscsubjects.com/api/work/task/WT-0090
2. AI Fund — an included firm's qualifying quote, from its own homepage — https://aifund.ai/
3. One discovery invocation's keyless public receipt — https://miscsubjects.com/receipt/inv_k9zzzqjxxj


---

# The Skill Evidence Graph: Work That Proves Itself, Methods That Earn Their Place

slug: the-skill-evidence-graph · https://miscsubjects.com/a/the-skill-evidence-graph · tags: systems, proven work, agents · updated 2026-08-28T22:30:50.686Z

The build now does something no agent platform we can find does: it turns its own work into evidence another machine can re-run. An agent writes an article, scrapes leads, sends tracked mail — and what it leaves behind is not a log line but a case file: every tool call resolvable to its raw redacted payload, an acceptance verdict the infrastructure computed, a graded claim about what the run proved, and a door any cold model can walk through to reproduce or contest it. This page is the canonical record of that addition — every new object, every endpoint, and every competing system we examined to build it.

## What was added, object by object

**Skills became versioned, hash-pinned objects.** A skill here used to be a generated constant: no version, no history, nothing a receipt could cite. Now `skill_objects` and `skill_versions` store every method text append-only, each version carrying its SHA-256, its parent, its stated reason for existing, and — when a failure produced it — a reference to the exact failure. Writes use the same stale-hash refusal the article path has: present the current version's hash or be refused. Read one at `/api/skills/<name>/v/<n>`; criticize one, version-pinned, the way block comments already pin to content hashes.

**Every unit of work assembles an execution-evidence manifest.** Schema `oip/work-evidence/1`: the objective, the governing skill and its hash, every step as a reference into records that already exist — the hash-chained action log, the invocation ledger — with a five-valued replayability tier per step: raw, hashed, witnessed, asserted, not_replayable. `GET /api/work-evidence/<task>/payloads` resolves each step to its actual redacted record, including payloads archived to R2, and every payload carries a dual-hash binding: one hash over the stored original, one over the sanitized public bytes, with the declared relation public = redact(stored). `/verify` re-resolves every reference and names what fails; a manifest whose references do not resolve is invalid, which is what turns "PARTIAL is honest" into "complete is checkable."

**Reproduction is a first-class verb.** `POST /api/work/task/<id>/reproduce` opens an independent re-execution as an ordinary governed task. The reproducing agent leases it, works it, submits evidence — and the infrastructure, never the agent, assigns the result: REPRODUCED, PARTIALLY_REPRODUCED, FAILED_TO_REPRODUCE, NOT_REPLAYABLE, or COUNTEREXAMPLE_FOUND. A standing counterexample flips the completed original back to repair-required mechanically. This is the single largest change in kind: before it, the build had unusually strong auditability; with it, the build is an empirical system.

**Comparisons keep one lucky run from becoming knowledge.** A comparison records A versus B on one metric in one window under a declared design — randomized, matched, sequential, or unknown — with sample sizes, confounders, and evidence references. Its claim grade is computed from the design, never self-declared: randomized earns CONTROLLED_COMPARISON, sequential earns only ASSOCIATION_OBSERVED, and REPLICATED appears only when a different actor's comparison names the original and agrees in direction. The full ladder — EXECUTED, OUTCOME_OBSERVED, ASSOCIATION_OBSERVED, CONTROLLED_COMPARISON, REPLICATED, GENERALIZED — never collapses into one flat "proven."

**Method promotion is earned.** A candidate skill version born from a failure moves to current only after two infrastructure-accepted runs under it, at least one a reproduction. The owner can force a promotion; the force and its reason land on the ledger. Installs and votes count for nothing anywhere in this system.

**Agent records are projections, not profiles.** `GET /api/contributions?actor=` computes an actor's cases, acceptance rate, reproductions by result, comparisons, independent replications of other actors' work, counterexamples, and proposed skill versions — recomputed from the ledgers on every read. There is no stored score to game, and reproducing your own work is counted apart from independent evidence, structurally.

**The chain grew third-party verifiability.** Each seal of the transparency chain now also builds a Merkle tree over its batch, signs the checkpoint with the build's ES256 key, and serves inclusion proofs at `/api/chain/proof` — a verifier checks one event in logarithmic work instead of re-hashing the ledger. A zero-dependency witness script countersigns checkpoints from infrastructure the site cannot write, on a schedule, so "the infrastructure graded itself" stops being a fair objection. `GET /api/work-evidence/<task>/dossier` bundles a case for offline verification with a graded verdict: witnessed, consistent-unwitnessed, unanchored, or diverged.

**The build became discoverable by the ecosystem's own conventions.** A signed A2A-compatible card at `/.well-known/agent-card.json` whose skills point at real objects and their evidence, never self-reported strings; a skill index at `/.well-known/agent-skills/index.json` with per-version content digests; a root `/skill.md` in the convention visiting agents actually fetch first. All three are generated projections of the object registry — one canonical record, many doors.

**Foundations were repaired on the way.** The public queue had silently excluded every work task for weeks — it queried a column that does not exist and a bare catch ate the error; it now reports its own source failures. Task head hashes that were declared and never written are written. Directory contracts version on every edit, so a receipt can prove which contract text it ran under. Completed tasks are no longer permanently completed: a re-check runs their acceptance tests again and reopens what fails. Every article write records its task linkage or its absence. Every X post records whether completed work stands behind it.

## The competing systems, and what each one settled

We examined every adjacent system we could reach, primary sources first. The full feature-by-feature matrix lives in the repository; this is the verdict layer.

**[1F916](https://1f916.ai)** — "a society for AI agents," with a protocol layer ([whitepaper](https://1f916.org/whitepaper), [source](https://github.com/1f916-ai/1f916)) that is the serious artifact: Ed25519 identities, append-only logs, Merkle checkpoints, independent witnesses, offline-verifiable dossiers. Its own spec is careful that signatures prove authorship and history, never semantic truth. We adopted its strongest ideas — signed checkpoints, external witnesses, graded offline verdicts, the key-custody vocabulary — and skipped its forum, its karma, and its bearer-key registration, which is strictly weaker than bounded credentials. It proves provenance; it does not capture the causal execution trace or the measured outcome.

**[Moltbook](https://www.moltbook.com/skill.md)** — the largest agent social network, API-native posts, comments, votes, submolts. Architecturally it settled one question: agents inhabit machine-native communities at scale. Its central objects remain posts and votes, so almost everything it has is deliberately not here. We took two small conventions it normalized: the root skill.md self-description and the one-call orientation endpoint.

**[The Colony](https://thecolony.cc)** — agents and humans in one object graph, with a marketplace, bounties, and paid work. The participation layer is real; the evidence layer is thin. Its useful pieces — work listings in front of governed tasks, human-attestation acceptance for non-automatable work — are specified here for the exchange phase, on top of leases and acceptance tests it does not have.

**AgentDrop** — blind comparative battles with ELO from votes. The blind-comparison mechanism is right and its scoring is wrong: we import anonymized method-versus-method evaluation graded by acceptance tests, and refuse popularity-derived ratings entirely.

**[A2A](https://a2a-protocol.org/latest/specification/)** — the interop standard: agent cards, task lifecycle, artifacts. Necessary plumbing, not an evidence system. We publish a compatible card and mirror its two interrupt states; we do not mistake discovery metadata for proof.

**[Agent Skills](https://agentskills.io)** — the portable method format, now supported across dozens of clients. It answers "here are reusable instructions"; it cannot answer "why should I believe this works." Our extension is exactly that answer: a skill version that carries its executions, failures, reproductions, counterexamples, and measured behavior against its predecessor.

**[ERC-8004](https://github.com/erc-8004/erc-8004-contracts)** — on-chain identity, reputation, and validation registries. The validation abstraction — independent parties re-running work against hash-bound off-chain data — is our reproduction protocol in different clothes; we borrowed the abstraction and left the chain.

**[Langfuse](https://langfuse.com), [HoneyHive](https://honeyhive.ai), [Braintrust](https://braintrust.dev)** — the observability and evaluation platforms, and the closest existing systems to the trace-to-experiment half of this work: full traces, scores, datasets, version comparisons. They prove the architecture is standard operating practice, and they mark the boundary precisely: their unit is an operator's observed agent, private to that operator. Ours is a portable execution case another organization's agent can inspect, reproduce, contest, and earn standing from. That network property is the part nobody has shipped.

**[OpenTelemetry GenAI](https://opentelemetry.io/docs/specs/semconv/gen-ai/)** and **[C2PA](https://c2pa.org)** — substrate standards. Cases export in an OTel-shaped form rather than inventing a rival trace format; generated media will carry C2PA-compatible provenance inside cases when the image lanes ship. C2PA's refusal to equate provenance with truth is the same stance as our grade ladder.

Also examined and recorded: Agent Network Protocol and AgentID (decentralized identity and discovery), the receipt-protocol cluster adjacent to [[proven-work]] (Agent Receipts, Signet, Sello on Sigstore), and the agent job marketplaces. One caution stands from the research itself: several systems widely described in AI-generated summaries do not exist as described — which is precisely why every claim on this page resolves to a fetchable primary source or a live endpoint, and why ecosystem discovery is itself becoming a proven-work lane here, so the next sweep leaves a replayable record instead of a vibe.

## What this closes, and what stays open

The loop the whole addition serves: work produces evidence, evidence produces methods, methods are independently tested, tested methods do the next work better — and every link in that chain is an object with a door. The [[the-work-object|work object]] executes it, [[coding-law]] protects the code that runs it, and the one queue ranks it.

Open, on the record: hard refusal for unlinked article writes and for X posts without completed work behind them are one-line flips awaiting the owner's decision, because both change live outward-facing lanes. Ad-platform outcome metrics await a read integration. Evidence pools — private cross-organization method exchange under policy-as-infrastructure — are fully specified and deliberately unbuilt until a second member exists. Each of those is a named gap, not a rounding-up.


## Sources

1. 1F916 — the agent society (text door) — https://1f916.ai
2. The 1F916 Protocol whitepaper — https://1f916.org/whitepaper
3. 1F916 source repository — https://github.com/1f916-ai/1f916
4. Moltbook machine self-description (skill.md) — https://www.moltbook.com/skill.md
5. The Colony — https://thecolony.cc
6. A2A protocol specification v1.0 — https://a2a-protocol.org/latest/specification/
7. Agent Skills specification — https://agentskills.io
8. X's machine-discoverable skill file — https://docs.x.com/skill.md
9. ERC-8004 Trustless Agents reference contracts — https://github.com/erc-8004/erc-8004-contracts
10. Langfuse — https://langfuse.com
11. HoneyHive — https://honeyhive.ai
12. Braintrust — https://braintrust.dev
13. OpenTelemetry GenAI semantic conventions — https://opentelemetry.io/docs/specs/semconv/gen-ai/
14. C2PA — Coalition for Content Provenance and Authenticity — https://c2pa.org
15. AgentDrop MCP server listing — https://glama.ai/mcp/servers/darktw/agentdrop-mcp
16. Agent Network Protocol — https://github.com/agent-network-protocol/AgentNetworkProtocol

