# Three measurements from one 70-finding suite: vendor diversity beats panel size, the second channel is the cheapest, and the published error floor was three times too good

slug: diversity-beats-count · https://miscsubjects.com/a/diversity-beats-count · category: canon · tags: adjudication, calibration, panels, measurement, canonical · updated 2026-08-01T23:56:28.261Z

## The suite these numbers come from

Fourteen probe items with the correct verdict declared in advance were run through five adjudication channels — the identical path live findings take — producing 70 findings. Sixty-four panel configurations were then replayed over those same 70 findings, each scored on two numbers: the **emit rate** (how often the assembly answers rather than escalating to a human) and the **undetected-wrong rate** (how often it answers, the answer is wrong, and nothing catches it).

Three findings came out of that data. Two are results about how to build a panel. The third is about the accounting, and it reduced the headline number by a factor of three after an outside audit found it.

| finding | the number |
|---|---|
| Cross-family pairs beat same-family pairs at identical cost | 0.169 vs 0.214 undetected-wrong |
| The second channel is the cheapest correctness; the fifth is the most expensive | 0.314 → 0.178 for one call; 0.178 → 0.071 for three more |
| The published floor depended on an exclusion policy | 0.071 stated, 0.214 under the alternative accounting |

[[embed:source:s1]]

## Part 1 — Two reviewers from different vendors beat two from the same vendor

### The one-sentence version

Two models from the same vendor are close to one model wearing two names. If a panel's seats share a training family, the panel's independence is partly an accounting fiction — and this system has now measured the size of the fiction on its own record: at identical cost, a cross-family pair beats a same-family pair on the only number that matters, and the mechanism is visible in the raw agreement rates.

This page exists because the finding is buried as one section of [the logical-economics table](https://miscsubjects.com/a/logical-economics) and it deserves to stand alone. It is the most portable result on this site: everything else here requires adopting an architecture; this requires changing one line of panel policy.

### Where the numbers come from

Fourteen probe items with correct verdicts declared in advance were run through five adjudication channels — the identical path live findings take, so nothing about the measurement is synthetic except the questions. That produced 70 findings. Sixty-four panel configurations — every subset of the five channels, under several gate policies — were then replayed over those same 70 findings, and each configuration was scored on two numbers:

- **emit rate** — how often the assembly answers at all, rather than escalating to a human;
- **undetected-wrong rate** — how often it answers, and the answer is wrong, and nothing catches it.

The second number is the one a buyer of machine judgment should care about, because a wrong answer that escalates costs a review and a wrong answer that emits costs whatever the decision was worth.

### The finding

Hold the channel count at two. Vary only one thing: whether the pair of models shares a training family.

| pair | configurations | emit rate | undetected-wrong rate |
|---|---|---|---|
| same training family | 2 | 0.893 | 0.214 |
| different training family | 8 | 0.714 | **0.169** |

Same cost. Same count. The cross-family pair is better on the number that matters — 0.169 against 0.214 — and the reason is not mysterious, because it is measured too: **same-family adjudicators agree with each other 0.893 of the time, cross-family 0.714.** Agreement between correlated judges is not confirmation; it is one judgment counted twice. The gate in this system compares derivations and escalates on divergence, so a pair that diverges more often hands more of its hard cases to a human — which is why the cross-family emit rate is lower — and is wrong-in-unison less often, which is why its undetected-wrong rate is lower. You are buying disagreement, and disagreement is the raw material error-catching is made of.

### The price curve the finding sits inside

The channel-count table, from the same 64 configurations:

| channels | mean emit rate | mean undetected-wrong rate | best achievable |
|---|---|---|---|
| 1 | 0.972 | 0.314 | 0.214 |
| 2 | 0.750 | 0.178 | 0.071 |
| 3 | 0.636 | 0.136 | 0.071 |
| 4 | 0.529 | 0.100 | 0.071 |
| 5 | 0.429 | **0.071** | 0.071 |

Read it as a price list. The second channel halves the undetected-wrong rate — 0.314 to 0.178 — for exactly one additional model call. The third, fourth and fifth channels together buy the remaining 0.178 → 0.071, less improvement for three times the marginal spend, and they are paid for twice: once in compute and once in escalations, because at five channels the assembly answers only 43% of what it is asked. Fifty-seven per cent of everything goes to a human. That is the honest cost of the last increment of assurance, and it is the standing argument against the current fashion of sending every question to the largest model available and calling the confidence of one channel a safety property.

**The second channel is the cheapest correctness available anywhere in this table. Which second channel? A different family. That is this page's entire content, and the table above is why it fits in a sentence.**

### The floor, and why diversity does not remove it

Beyond two channels the best-achievable column stops moving at 0.071, because one probe item — P07 — survives every configuration of every size. On P07 all five channels answered DENY; the declared correct verdict was CANNOT_CONCLUDE. Unanimity is exactly what a disagreement-triggered gate takes as permission to emit. **An assembly built to catch divergence is blind to correlated wrongness by construction**, and no channel count fixes that, because adding channels adds more of the same unanimous error. The only instrument that found P07 was the known-answer probe — a question whose answer was declared before it was asked.

Two honesty notes, both load-bearing:

- The floor figure itself leans on an exclusion policy. Three probe items were unanimously wrong, not one; two of them were rescued when a model returned unparseable output and the gate escalated instead of emitting. Under an accounting that scores a parse-failure rescue as an escaped error, the bound is 3/14 = 0.214. The sensitivity is published on [the probe report](https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act) as of 2026-08-01, filed as objection 209. The family comparison above is unaffected — both pair types are scored under the same policy — but nobody should quote 0.071 without its footnote.
- Cross-family correlation is lower, not zero. The families were trained on overlapping corpora toward overlapping objectives; where the entire training distribution is confidently wrong, every family inherits the error together. Diversity moves the floor's location. It does not abolish floors.

### The live case where partial independence earned its keep

This is not only a replay result. In a live run under the EU AI Act Article 50 rule set, the panel met a genuine boundary question and split: three CANNOT_CONCLUDE, one DENY, one AFFIRM. The majority landed on the correct abstention even though two members manufactured verdicts. A fully correlated panel does not produce that split — it produces five copies of one of the wrong answers, and the gate, seeing agreement, emits it. The split *is* the safety mechanism working.

### The indictment this finding files against its own instrument

Every assembly this system has run in production so far has drawn on **two** training families. By its own measurement, that is under-diversified. The finding was produced by an instrument it partially condemns, the condemnation is recorded here rather than smoothed over, and widening the family spread of the standing panels is on the roadmap as a defect, not an aspiration. A reader who wants to check whether it has happened yet can open the panel rows in [the directory](https://miscsubjects.com/api/directory/search?q=adjudicate) and count vendors, without asking anyone.

### What transfers, today, to anyone

The result costs nothing to adopt and does not require this system:

1. **Count training families, not seats.** A "five-model panel" drawing on two vendors is closer to a two-model panel with redundancy. Write the family count into the panel policy as the governing number.
2. **Spend the second channel first, and spend it across a family line.** It is the cheapest correctness in the table, and the family line is where its value is concentrated.
3. **Do not buy the fifth channel without pricing the humans.** At five channels, most questions escalate. If there is no one to escalate to, the assurance is decorative.
4. **Keep a known-answer probe running,** because the one error class that survives everything — confident unanimous wrongness — is invisible to every disagreement-based mechanism and visible only to a question whose answer was fixed in advance.

### What this page does not establish

One task class, one rule set, fourteen self-authored probes, five channels from a handful of families. The rates are priors, not guarantees; a different rule set needs its own table, and the suite is published at a hash precisely so it can be attacked. What survives even hostile reading of the sample size is the direction and the mechanism: agreement between correlated judges is cheaper to produce and worth less, and the measured gap — 0.893 against 0.714 — is large enough that no plausible re-scoring makes the same-family pair the better buy.

### Where to argue

File objections at the [gauntlet](https://miscsubjects.com/a/gauntlet-log). The replay data, the probe suite and the per-model rates are all public at the links above; the strongest attack is a re-run of the published suite that produces a materially different family gap, and the suite exists to make that attack possible.

[[embed:source:s2]]

## Part 2 — The published error floor depended on what was refused a count

### The finding, as it arrived

An external cold audit read this site's adjudication numbers the way an adversary should, and found an arithmetic tension nobody inside the build had published:

The known-answer probe suite has fourteen items. On three of them — P05, P07, P09 — the entire five-model panel was wrong: zero correct out of five, three separate times. Yet the published configuration table reports a five-channel floor of **one** undetected-wrong item in fourteen: 0.071, naming P07 as the sole survivor. If three items were unanimously wrong, why does only one survive every configuration?

The reconciliation was in the fine print. Two of the seventy findings were malformed — one confirmed at the receipt level as `kimi-k2.6` returning UNPARSED on P05 — and were excluded from the configuration statistics, because a non-finding is not a rating. That exclusion is a defensible scoring decision. But it has a mechanical consequence the report did not state: **a malformed finding forces the gate to escalate rather than emit.** An unparseable output on an item the panel would otherwise have answered wrongly converts an escaped error into a human referral. On at least one, and possibly two, of the three unanimously-wrong items, the assembly was rescued not by diversity, not by the gate's design, but by a model failing to produce parseable output.

The headline number — five channels drive undetected-wrong down to 0.071 — rests in part on accidental parse failures. Take the rescue away and the floor bound is 3/14 = **0.214**, roughly triple.

### What is confirmed and what is inference, exactly

Confirmed, at the linked surfaces:

- The exclusion policy exists and is stated on [the probe report](https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act): 2 of 70 findings malformed, excluded from configuration statistics, retained in per-model rates.
- P05, P07 and P09 were each 0/5 — printed per item, with the declared expected verdict and the reason it is correct.
- `kimi-k2.6` returned UNPARSED on P05 — the receipt caption says so.
- A malformed finding cannot be emitted; the gate's only move is escalation.

Not yet resolved: **which item the second malformed finding landed on.** If it landed on P09, both rescues sit on unanimously-wrong items and the 0.214 bound binds tight. If it landed on an item the panel had right anyway, one of the three unanimous misses escaped by some other route and the accounting needs a different correction. The per-item receipts settle this and reading them is open work, stated here as open work.

### Why the rescue is genuinely double-edged

It would be too quick to call this only an embarrassment. Escalating on malformed output is *correct* behaviour — a gate that emitted anyway, or guessed, would be indefensible. The assembly did, mechanically, the safe thing: faced with a channel that produced garbage on a question where every functioning channel was confidently wrong, it declined to answer. In the field, that outcome — a human looks at P05 — is strictly better than the alternative the other channels were unanimously offering.

The defect is not the behaviour. The defect is the **bookkeeping**: crediting that outcome to the assembly's measured error floor without disclosing that the mechanism was luck. A parse failure is not a safety property, because it is not reproducible on demand — the next run of P05 may parse cleanly and emit the wrong answer five-for-five. A floor propped by accident holds until the accident stops happening, which is precisely the kind of number that fails exactly when relied upon. The honest statement is now on the report: 0.071 is the floor **under the stated exclusion policy**; 0.214 is the bound under the accounting that treats rescues as escapes; a reader pricing a consequence should know which one they are holding.

### The general lesson: an exclusion policy is a safety claim

Every published error rate — every eval score, every benchmark, every audit finding, every clinical adjudication statistic — sits on top of decisions about what did not count: malformed outputs, timeouts, refusals, off-format answers, items the graders could not agree on, runs that crashed. Each decision is individually defensible. Collectively they are a second, silent result the reader never sees, because the same raw data under two defensible accounting policies produced 0.071 and 0.214 here — a factor of three, on a suite of fourteen items, from one scoring choice about two findings.

The transferable rules, each of which this system now follows because it was caught not following them:

1. **Publish the exclusion count next to the headline rate, always.** "0.071 (2 of 70 findings excluded as malformed)" and "0.071" are different claims.
2. **State the direction of the exclusion.** An excluded failure that would have raised the rate is not the same object as an excluded duplicate; say which way each exclusion cuts.
3. **Publish the sensitivity, not just the policy.** The useful sentence is "under the alternative accounting the figure is X" — one line, computable at publication time, and its absence is what an adversarial reader will find first.
4. **Treat non-answers as their own outcome class.** Wrong, right, abstained, and *failed to produce a rating* are four outcomes, not three; folding the fourth into any of the others is where the flattery hides.

### What this episode says about the machinery around it

The objection came from outside, from a cold read, with no access beyond the public record — and everything needed to find it was public: the per-item results, the exclusion note, the receipt caption, the configuration table. The system's claim was never that it does not err; the claim is that the record is sufficient for a stranger to catch the error, and that the error and its correction end up on the same page. Both held. The sensitivity note is on the probe report, the objection is filed as [obj-209](https://miscsubjects.com/i/discourse/obj-209), the correction was posted publicly the same day, and this page exists so the lesson outlives the incident.

### What this page does not establish

It does not establish that the exclusion policy was wrong — a non-finding genuinely is not a rating, and the per-model rates always included the malformed outputs. It does not establish the true floor: that requires resolving the second malformed finding from the per-item receipts and re-running the suite until parse failures either stop occurring or occur often enough to be a measured property of their own. And it does not establish that any other published error rate has this defect — only that the reader has, in the general case, no way to know without the exclusion accounting, which is the point.

### Where to argue

File objections at the [gauntlet](https://miscsubjects.com/a/gauntlet-log). The strongest attack on this page is resolving the second malformed finding and showing it landed on an item the panel had right — which would weaken the 0.214 bound and is exactly the check the receipts exist to allow.

## Part 3 — The same failure class in the writing pipeline: 121 identical emails

The measurement above is about aggregate properties invisible to per-item checks. The clearest instance of that failure class in this build was not in the panel at all — it was in the outreach drafting pipeline, and it is included here because it is the same defect wearing different clothes.

### The failure, plainly

The most expensive failure this build's outreach system has produced was not a rule being broken. It was a rule being obeyed.

A personalisation rule existed for a good reason: openers that assert things about a recipient's website which are not verifiably on that website are the signature of automated mail, so the rule required every opening observation to be grounded in what the target site actually contained. Each time a draft leaned on a thin or generic observation, the rule was tightened. Each tightening was individually correct. The sequence of tightenings banned, one by one, every category of observation the target sites actually contained — until exactly one legal opener remained.

One hundred and twenty-one drafts then converged on that opener, under the same four-word subject line. **Every one of them passed every validator.** Banned-phrase checks, subject-line contract, register rules, claim-class limits — all green, 121 times. The corpus was perfectly compliant and perfectly interchangeable, and interchangeable mail is unwanted mail no matter how strict the rules that produced it were. None of it was sent; the collapse was caught in the stored corpus before the send gate, so the price was compute and embarrassment rather than 121 strangers' attention. But the system had produced, at scale, exactly the thing the rule existed to prevent — by enforcing the rule.

### Why no validator saw it

Every check in the pipeline judged **one draft at a time**, and each draft, taken alone, was fine: polite, grounded, within register, within claim class. The defect did not live in any draft. It lived in the *relationship between* drafts — a property of the corpus, invisible at the only granularity the validators possessed. This is the general blind spot of per-item validation, and it is worth stating as a law because it recurs everywhere rule systems are used to govern generation:

**A property can be perfect in every instance and catastrophic in aggregate, and a per-instance validator cannot see aggregate properties by construction.**

Tightening per-item rules does not fix an aggregate defect. It caused this one. Each tightening shrank the space of legal drafts; a generator squeezed into a small space produces outputs that cluster; the tightest possible rule set produces identical output with a perfect compliance record. Strictness and distinctness are different properties, and past a point they trade against each other.

### The detector: hash the residue

The fix is structural, and it is the useful part of this page.

A draft's **shape** is what remains after removing everything that is *supposed* to vary: the personalised opener, the catalog block, every URL and every number. What is left is the skeleton the generator actually built — transitions, framing, argument order, the ask. That residue is hashed. Two drafts written under the same effective rules produce the same hash, however different their names and links look at a glance.

Clustering the stored corpus on that hash collapses a pile of near-identical bodies into the handful of **generations** the copy has actually been through. Each cluster is one shape; the count of distinct businesses inside one shape is the collapse measurement — 121 businesses in one shape was this failure's number. The detector has three properties the per-item validators lacked:

- **It is aggregate by construction.** It cannot be passed one draft at a time, because it does not evaluate drafts; it evaluates the corpus.
- **It needs no model and no judgment.** Strip, hash, count. There is nothing to argue with and nothing to drift.
- **It measures the thing the recipient experiences.** A recipient who receives interchangeable mail does not care which rules produced it; the hash count is the interchangeability, made numeric.

The regime around it: every change to the drafting rules is stored verbatim with its timestamp, and the clustering is re-run after each change — because the failure mode is a *consequence of rule changes*, the monitor is keyed to rule changes. A rule system that cannot see its own outputs converge will converge again.

### The general lesson, because this is not about email

Substitute any generator governed by per-item rules and the anatomy holds:

- **Code review checklists.** Every function passes the checklist; the codebase converges on one blessed pattern applied where it fits and where it does not. The checklist cannot see it.
- **Content policy.** Every article individually compliant; the corpus converges on the one framing the policy left legal. Readers experience a site that says one thing sixty ways.
- **Model evaluations.** Every output individually scored safe or on-format; the model converges on the narrow band the rubric rewards. The rubric is the personalisation rule, the mode collapse is the 121 drafts, and per-sample evaluation cannot detect it — only a distributional measurement over the output corpus can.

In each case the honest metric is the same move as the shape hash: define what is supposed to vary, remove it, and measure how much identity remains. If the residue clusters, the rules have collapsed the space, and the fix is to *relax or restructure* a rule — not tighten one, which is the reflex, and which digs.

### What this failure bought

The tightened rule was replaced rather than tightened further: the current outreach law requires one **specific observation that could fit no other recipient** — a requirement about information content, which cannot converge, instead of a requirement about permitted categories, which did. The shape-hash clustering stands as a permanent gate. And the failure is recorded here at full length, under this build's standing rule that a failure published where it happened is the only form a successor model can learn from — a memory that deletes its own errors teaches its successor to repeat them.

### What this page does not establish

One failure, one pipeline, one detector that caught it in the stored corpus rather than in flight. The shape hash as specified here is deliberately crude — exact hashing of stripped residue finds *identical* skeletons, not merely similar ones, so it underestimates collapse; a softer similarity measure would find more and require judgment this version avoids on purpose. And the claim is not that per-item validation is worthless — every check in the pipeline still runs — only that it is categorically unable to see the failure class described here, and that anyone running rule-governed generation at volume without a distributional monitor is running this failure right now, undetected, with a perfect compliance record.

### Where to argue

File objections at the [gauntlet](https://miscsubjects.com/a/gauntlet-log). The pipeline this happened in is documented, gates and all, at [outreach-machinery](https://miscsubjects.com/a/outreach-machinery).

## What all three have in common

Each is a property of a **set**, invisible to any check that examines one item. Correlated wrongness across a panel is invisible to a gate that only fires on disagreement. An exclusion policy's effect on a rate is invisible in any single excluded item. Template collapse is invisible in any single draft, all 121 of which passed every validator. In each case the instrument that found it was the same shape: a measurement over the whole set, run deliberately, because nothing in the per-item machinery could ever surface it.


## Sources

1. Logical economics — the full configuration table — https://miscsubjects.com/a/logical-economics
2. The probe report the rates come from — https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act
3. The probe instrument's own contract — https://miscsubjects.com/api/directory/ADJUDICATE_PROBE
4. The system this measures, end to end — https://miscsubjects.com/a/the-build-end-to-end
5. A live case where correlation showed its face — https://miscsubjects.com/a/adjudication-eu-ai-act-article-50
6. The objection as filed — https://miscsubjects.com/i/discourse/obj-209
7. The outreach machinery, documented end to end — https://miscsubjects.com/a/outreach-machinery


---

# Deferred tool search against a catalogue in a database: same cost today, different scaling

slug: tool-search-vs-catalogue-as-data · https://miscsubjects.com/a/tool-search-vs-catalogue-as-data · tags: tooling, mcp, architecture, comparison, tool-search, context-cost, measurement · updated 2026-07-26T03:53:50.309Z

On a catalogue of 891 capabilities, the two designs cost the same. Deferring the schemas and searching them on demand: 14,109 input tokens on a measured turn. Keeping the catalogue in a database behind one HTTP endpoint and sending no catalogue at all: 14,071. A difference of 38 tokens, which is noise.

The number is not the finding. The finding is what sits inside each number:

- With deferred tool search, **catalogue size is a term in the cost equation** — in Claude Code, one line per capability name in every request — and the mechanism only exists if the host implements it.
- With the catalogue in a database, catalogue size is not a term at all, and the mechanism is an HTTP request, which every client already has.

That holds at 50 capabilities and at 5,000. The token count expires the next time either side ships a change; the structural statement does not.

**What would falsify it.** One thing, precisely: a host that defers with *no per-capability residue in context*. Anthropic's server-side variant is documented to be exactly that — "the API excludes deferred tools from the system-prompt prefix" — so on the Claude API directly, tool search is also flat in context and the claim above is false as stated. It holds for Claude Code's client-side implementation, which is the one measured here, and Anthropic's own Claude Code documentation says so: "Only tool names and server instructions load at session start." Two implementations of one feature, two different cost curves. Check which one you are running before you believe either number.

[[embed:source:s1]]

## Evidence status

**Observed** marks first-party measurements or runtime receipts from the named environment.
**Derived** marks arithmetic calculated from cited inputs. **Specified** marks vendor or standards
documentation. **Implemented** and **deployed** name code and live-state evidence, respectively.
**Reproduced** means the stated procedure was rerun. **Externally attested** marks operator reports;
those reports show that an experience occurred, not that it is universal.

## Every term used here, defined once

| Term | What it means on this page |
| --- | --- |
| **Tool definition** | A name, a description and a JSON Schema for the arguments, sent to the model so it can call something. Definitions are input tokens and are re-sent on every turn. |
| **Turn** | One request to the model and its reply. The whole conversation, including all tool definitions, is re-sent each time. |
| **MCP** | Model Context Protocol. A standard for a server to publish tool definitions and for a client to fetch them and put them in the prompt. |
| **Deferred loading** | Sending a tool's definition in the request but keeping it out of the model's context until it is asked for. Set with `defer_loading: true` on the API, or `ENABLE_TOOL_SEARCH=true` in Claude Code. |
| **Tool search** | The search tool the model calls to pull a deferred definition into context. Two vendor variants: regex (`tool_search_tool_regex_20251119`) and BM25 (`tool_search_tool_bm25_20251119`). |
| **Catalogue-as-data** | One SQL row per capability, holding the call shape, the docs, the argument template and the name of the credential. The model gets no list; it resolves an intent to a key over HTTP. [The full row contract](/a/directory-row-contract). |
| **Round trip** | One network request and response. Distinct from a model turn: four round trips can happen inside one turn, or across four. |
| **Prompt cache** | The provider storing the unchanging front of your prompt so re-sending it is cheaper. Anything that changes early in the prompt invalidates everything after it. |

## What the client actually puts on the wire, both ways

Both figures below come from one machine, `claude-cli 2.1.165`, same prompt, same working directory, one MCP server attached, captured by a local server that logs the request body and answers with a canned reply. The script is at the end of this page.

**With `ENABLE_TOOL_SEARCH=true`** the request carries nine tool definitions totalling 38,116 bytes of JSON:

```text
Agent            4351 bytes
AskUserQuestion  4199
Bash             2661
Edit              964
Read             1636
Skill            1713
ToolSearch       1440
Workflow        20503
Write             639
                -----
                38116 bytes, 9 definitions
```

Everything else arrives as a bare list of names in a `system` role message, which opens: "The following deferred tools are now available via ToolSearch. Their schemas are NOT loaded — calling them directly will fail with InputValidationError."

**With `ENABLE_TOOL_SEARCH=false`**, same session, the same request carries 858 definitions and 522,746 bytes — 13.7 times the JSON, for the identical set of capabilities.

**With the catalogue behind HTTP** the client sends its nine built-in definitions and nothing else. No MCP server attached, no name list, and all 891 capabilities still reachable, because reaching them is a `GET` and a `POST` to one endpoint rather than a tool the model was handed.

The translator carrying these requests to a non-Anthropic model converts the Anthropic `tools` array into OpenAI `function` entries one for one — `toOpenAITools()`, `functions/api/aig/[[path]].js:213-231` — and sizes an incoming request at lines 448-456 with `Math.ceil(chars / 3.7)`. The capture server behind the byte counts is `tools/capture-gateway.mjs:1-55` in the public repository, logging `n_tools` and `tool_names` at lines 25-26.

[[embed:source:s25]]

## The ToolSearch definition itself costs 1,440 bytes, once

Measured, not estimated. It is 3.8% of the nine-definition payload and it does not grow with the catalogue:

```json
{ "name": "ToolSearch",
  "description": "Fetches full schema definitions for deferred tools so they can be called. …",
  "input_schema": { "type": "object",
    "properties": { "query": {"type":"string"}, "max_results": {"type":"number","default":5} },
    "required": ["query","max_results"] } }
```

The catalogue-as-data equivalent — the instruction telling a model to resolve before it invokes — is prose in the system prompt, not a definition. One HN commenter who hand-rolled the same gate before either vendor shipped one put his at "about 60 tokens". Both are small. Neither side wins this row.

[[embed:source:s26]]
[[embed:source:s19]]

## Dimension by dimension, with a verdict in every cell

| Dimension | Deferred tool search | Catalogue-as-data | Verdict |
| --- | --- | --- | --- |
| Input tokens per turn, 891 capabilities | 14,109 | 14,071 | **Tie.** 38 tokens apart. |
| How cost scales with catalogue size | Client-side: ~37 bytes of name per capability, every turn. Server-side: nothing in the prefix, but every full definition is uploaded on every request. | Nothing. The 892nd row changes no byte the model sees. | **Catalogue-as-data**, and only structurally — at 891 rows the difference is invisible. |
| Prompt-cache behaviour | Documented as preserved server-side: deferred tools are excluded from the cached prefix and expanded inline. Reported as broken by operators lazy-loading definitions themselves. | Nothing about the tool surface ever changes, so nothing invalidates. | **Catalogue-as-data**, narrowly. The vendor claim and the field reports are both quoted below and describe different implementations. |
| Host support required | Yes. A host implementing deferral and expanding `tool_reference` blocks, on a model that supports them: Sonnet 4.5, Haiku 4.5, Opus 4.5 and later. | None. An HTTP client. | **Catalogue-as-data.** Not close. |
| Discovery latency and round trips | Zero network round trips — the search is answered locally or inside the same API call — but one extra model turn before the work starts. | Four HTTP round trips, measured at 2.185 s total for resolve → contract → invoke → receipt. | **Tool search.** The four round trips are real seconds. |
| Accuracy of tool selection | Vendor states selection degrades past 30-50 tools and that search keeps it high. The counter-argument: this is retrieval, which the field abandoned for full definitions on accuracy grounds. | A ranked query you own, index and can test offline. Measured here: `?ask=what time is it` returned 12 candidates and recommended `NOW` — not `TIME_NOW`, which was also in the list. | **Unresolved.** Both are retrieval. Neither side has published a head-to-head. |
| Adding a new capability | Publish it from a server; clients pick it up on reconnect, which invalidates the cache when tools sit in the prefix. | One `INSERT`. Live immediately, no restart, no reconnect. | **Catalogue-as-data.** |
| Who can call it | A harness that implements deferral. | Any HTTP client, including a model with no tool-calling support at all. | **Catalogue-as-data.** |
| Argument validation | Real JSON Schema, enforced by the host, with strict mode composing on top of deferral. | Pipe-delimited string, optional `input_schema` on the row, validated by the endpoint after the call is made. | **Tool search.** Decisively. |
| Ecosystem | Thousands of published MCP servers work unchanged. | Each server must become rows, or be reached one `http` row per endpoint. | **Tool search.** Decisively. |
| Observability and receipts | Not part of the mechanism. Tool lists are client-side state. | Every invocation writes an addressable receipt with input and output hashes, replay and repair links. | **Catalogue-as-data.** |
| Failure mode | A deferred tool is never searched for, and the run completes as if it did not exist. Filed against two harnesses. | A key does not resolve and the step fails with a `did_you_mean` list. | **Catalogue-as-data**, if a loud failure is worth more to you than a quiet one. |
| Effort to adopt inside a Claude host | One environment variable. | Implement or adopt a four-step protocol. | **Tool search.** Not close. |

[[embed:source:s5]]
[[embed:source:s16]]

## What tool search has that a catalogue behind HTTP does not

**No service to run.** Deferral is a setting — no database, no endpoint, no uptime, no deployment.

**It works when your infrastructure does not.** A catalogue behind HTTP is a hard dependency: endpoint down, model has no capabilities. Deferred definitions travel inside the request.

**Typed arguments.** Definitions carry real JSON Schema and the host constrains the model's output to match it. A row taking `"arg1|arg2"` gets validation only after the call has been made.

**It is a standard other harnesses are adopting.** `defer_loading` is not Anthropic-only — the Codex bug report below is evidence it shipped there too. A protocol one build invented is not a standard, whatever its merits.

[[embed:source:s2]]
[[embed:source:s15]]

## Six ways deferred tool search is documented to break

Every row is a filed report or a vendor page, quoted.

| What goes wrong | Where it is documented | What it costs you |
| --- | --- | --- |
| First-party servers are exempt from deferral, with no opt-out | anthropics/claude-code#76372: "With tool search active (`ENABLE_TOOL_SEARCH` unset), third-party MCP tools correctly defer to names-only. But three Desktop built-in servers load complete schemas upfront every session" | ~3.9k tokens per session that no setting removes |
| Servers visible in the client are invisible to the search index | anthropics/claude-code#57033: "the `ToolSearch` deferred-tool discovery mechanism does NOT include them in its index. Any `ToolSearch` query that should match a claude.ai MCP tool returns zero results" | Connectors show as Connected and cannot be reached |
| The deferral threshold is computed from the wrong model | NousResearch/hermes-agent#57520: "For any session running a model *other than* the configured default … the gate is scaled to the wrong window." | A 98,304-token local model gets a threshold sized for a 256K cloud model |
| A run completes silently empty | openai/codex#24536: "`codex exec` can silently finish with no assistant message when an explicitly configured MCP tool is deferred behind `tool_search`." | The task looks done and nothing happened |
| Lazy loading invalidates the prompt cache | sophiabits on HN: "The main problem with this approach at the moment is it busts your prompt cache, because LLMs expect all tool definitions to be defined at the beginning of the context window." | Input tokens are the main cost driver, so the saving can invert |
| It does not fully solve the problem it is cited for | cheema33 on HN: "And no, the tool search function recently introduced by Anthropic does not completely solve this problem." | You still pay for servers you never invoke on that turn |

[[embed:source:s8]]
[[embed:source:s9]]
[[embed:source:s10]]
[[embed:source:s11]]
[[embed:source:s21]]
[[embed:source:s20]]

And the objection that is not a bug, which is the strongest thing said against the whole direction:

> Seems like we traded scalability for accuracy, then accuracy for scalability… but I guess maybe we've come out on top because whatever they are using for tool search is better than RAG?

That is morelandjs on Hacker News, 2025-11-24, under Anthropic's own advanced-tool-use announcement. The field moved to putting every definition in context *because retrieval was less accurate*. Tool search walks that back and asks you to trust a BM25 or regex index instead. Anthropic's position is the opposite — selection "degrades once you exceed 30–50 available tools", and search keeps it high. Both statements are published, neither is backed by a head-to-head benchmark either side has released, and a catalogue-as-data resolver is retrieval too, so it inherits the same objection. This page does not resolve it and will not pretend to.

[[embed:source:s17]]

## Five ways catalogue-as-data breaks

**No host-side type validation.** The model sends a pipe-delimited string. Nothing constrains its output to a schema before the call leaves. Malformed arguments are caught by the endpoint, after the request, and surface as a failed invocation rather than a rejected tool call.

**No standard.** One build's protocol. No ecosystem, no published servers, no client that already speaks it.

**It needs a live service.** The catalogue's uptime is now the agent's capability uptime, which deferred definitions never are.

**Four round trips before any work happens.** Measured on 2026-07-25 against production:

```text
resolve   GET  /api/dispatch?ask=what%20time%20is%20it     200   0.619 s   12,332 B
contract  GET  /api/dispatch?key=TIME_NOW&format=markdown  200   0.215 s    4,421 B
invoke    POST /api/dispatch {"key":"TIME_NOW","body":""}  200   1.047 s   15,607 B
receipt   GET  /api/dispatch?confirm=inv_3wt7dcbp2c        200   0.304 s   15,049 B
                                                          -----------------------
                                                           2.185 s   47,409 B
```

Deferred tool search pays none of that in network time. It pays one extra model turn instead, which on a slow model is worse and on a fast one is better.

**The resolve step is a search that can miss.** In the run above, `?ask=what time is it` returned 12 matches and recommended `NOW`. `TIME_NOW` was in the list and was not the recommendation. Both work. But that is the accuracy objection landing on this side of the table, in a measurement taken for this page, and it is the same failure class the critics aim at tool search.

[[embed:source:s27]]

## The arithmetic at 50, 891 and 5,000 capabilities

Two measured constants do the work, both from the captures above:

- **571 bytes** per tool definition. Derived: (522,746 − 38,116) bytes ÷ (858 − 9) definitions.
- **36.7 bytes** per deferred tool name. Derived: 30,551 bytes of names ÷ 833 MCP tool names.

Bytes are exact. To convert to tokens: (149,187 − 14,109) input tokens across (856 − 9) definitions gives **159 tokens per definition**, and dividing the byte figure by the token figure gives **3.59 bytes per token** for this JSON. Independently, the translator in this build estimates at `chars / 3.7`. The two agree closely enough to use for orders of magnitude and not closely enough to quote to four figures.

| Catalogue size | Every definition in context | Tool search, client-side (names only) | Catalogue-as-data |
| --- | --- | --- | --- |
| 50 | 50 × 571 = 28,550 B ≈ 7,950 tokens | 50 × 36.7 = 1,835 B ≈ 511 tokens | 0 B, 0 tokens |
| 891 | 891 × 571 = 508,761 B ≈ 141,700 tokens | 891 × 36.7 = 32,700 B ≈ 9,110 tokens | 0 B, 0 tokens |
| 5,000 | 5,000 × 571 = 2,855,000 B ≈ 795,300 tokens | 5,000 × 36.7 = 183,500 B ≈ 51,100 tokens | 0 B, 0 tokens |

At **50 capabilities** all three are rounding errors against a 200,000-token window. Anthropic says as much: standard tool calling "is a better fit when you have fewer than 10 tools, every tool is used in every request, or your tool definitions are small". Do not build a catalogue for 50 rows.

At **891**, the measured case, column one is a bill you notice and the other two are small. That is why the two measured turns landed 38 tokens apart. The 9,110 tokens of names in column two are real, hiding inside a number that also holds the conversation.

At **5,000**, column one does not fit in a 200,000-token window at all, column two costs 51,100 tokens per turn in names before anything is retrieved, and column three still costs nothing. That is where the structural difference becomes a number instead of an argument — and where the server-side variant, which keeps names out of the prefix, would collapse column two to roughly zero, subject to its documented ceiling of 10,000 deferred tools per request.

Add to column two, in every row, the definitions actually retrieved: 159 tokens each, up to five per search by default.

[[embed:source:s14]]
[[embed:source:s24]]

## A non-Claude model drove the deferred loop, against two documentation pages

Two vendor pages say this configuration does not work.

Anthropic's Claude Code prompt-caching page, on when tool definitions land in the cached prefix instead of being deferred: this happens "when tool search is unavailable or disabled … such as on Google Cloud's Agent Platform or with a custom `ANTHROPIC_BASE_URL` gateway". The Claude Code MCP page is more explicit about the mechanism: tool search "is also disabled when `ANTHROPIC_BASE_URL` points to a non-first-party host, since most proxies don't forward `tool_reference` blocks."

Moonshot's Claude Code guide, in its environment-variable table for `ENABLE_TOOL_SEARCH`: "The Kimi endpoint does not support this feature yet; it must be set to `false`, otherwise tool calls misbehave."

**Measured on 2026-07-25**, `claude-cli 2.1.165`, `ANTHROPIC_BASE_URL` pointed at a self-hosted translator in front of Cloudflare AI Gateway, model `@cf/moonshotai/kimi-k2.7-code`, `ENABLE_TOOL_SEARCH=true`: Kimi K2.7 Code was asked to use an MCP tool it had never been shown, named `TIME_NOW`. It called `ToolSearch`, received the schema, invoked the tool, and returned the payload:

```json
{"now":"2026-07-25T18:35:56-07:00","today":"2026-07-25","time":"18:35:56",
 "zone":"America/Los_Angeles","iso":"2026-07-25T18:35:56-07:00"}
```

The same session measured 14,109 input tokens against 149,187 with the setting off.

The likeliest reconciliation, stated as an inference and not a fact: Claude Code's tool search is client-side — the client decides which definitions to send and answers `ToolSearch` itself, so it asks nothing of the endpoint. The vendor pages describe the *server-side* variant, which does require the endpoint to expand `tool_reference` blocks. Anthropic's documentation confirms both exist: "Tool search runs as a server-side tool, but you can also implement your own client-side tool search."

Do not settle this by trusting any of the three statements. Point `ANTHROPIC_BASE_URL` at the capture server below and read your own request bodies.

[[embed:source:s3]]
[[embed:source:s4]]
[[embed:source:s28]]

## Given your situation, pick this

| If this is you | Pick | Because |
| --- | --- | --- |
| Under 50 tools, all used most turns | Neither | Definitions in context are cheaper than any machinery around them, and Anthropic says so in its own "when to use" list. |
| Inside Claude Code, MCP servers you did not write, a bill you noticed | Tool search | One environment variable, 10.6× fewer input tokens on the case measured here, nothing to run. |
| Several different models must reach the same capabilities, some without tool calling | Catalogue-as-data | Deferral needs a supporting host and a supporting model. HTTP needs neither. |
| Thousands of capabilities and a fixed context budget | Catalogue-as-data, or the server-side API variant | Client-side deferral still carries a name per capability per turn; the other two carry nothing. |
| Every call must produce an auditable receipt | Catalogue-as-data | Receipts are not part of the deferral mechanism at any layer. |
| Typed arguments matter more than portability | Tool search | Real JSON Schema, host-enforced, composing with strict mode. |
| You are on a proxy or a non-Anthropic model | Measure before choosing | Two vendor pages say it will not work; one measurement says it does. Yours is the only one that decides. |
| You want both | Both | The same table here is projected as a per-row MCP tool list for hosts that want tools, and as the protocol for everything else: [MCP as a projection, not a home](/a/mcp-as-a-projection). |

[[embed:source:s22]]
[[embed:source:s23]]
[[embed:source:s13]]
[[embed:source:s12]]

## Reproduce every number on this page

**The wire captures.** Save this as `tooldump.mjs`. It answers the Anthropic Messages API with a canned reply and logs the byte size of every tool definition the client sends.

```javascript
import http from 'node:http';
import fs from 'node:fs';
const OUT = '/tmp/tooldump.jsonl';
http.createServer((req, res) => {
  let raw = '';
  req.on('data', c => (raw += c));
  req.on('end', () => {
    let b = null; try { b = JSON.parse(raw); } catch {}
    if (b && Array.isArray(b.tools)) {
      fs.appendFileSync(OUT, JSON.stringify({
        model: b.model, n_tools: b.tools.length,
        tools_bytes: JSON.stringify(b.tools).length,
        per_tool: b.tools.map(t => ({ name: t.name, bytes: JSON.stringify(t).length })),
      }) + '\n');
    }
    res.writeHead(200, { 'content-type': 'application/json' });
    res.end(JSON.stringify({ id:'msg_1', type:'message', role:'assistant',
      model: b?.model || 'local', content:[{type:'text',text:'ok'}],
      stop_reason:'end_turn', usage:{input_tokens:10,output_tokens:2} }));
  });
}).listen(8788, () => console.log('tooldump on :8788'));
```

Run it, then run the client once each way:

```bash
node tooldump.mjs &
for v in true false; do
  ANTHROPIC_BASE_URL=http://localhost:8788 ANTHROPIC_AUTH_TOKEN=x \
  ANTHROPIC_MODEL=local ENABLE_TOOL_SEARCH=$v claude -p "say ok" >/dev/null
done
python3 -c "
import json
for l in open('/tmp/tooldump.jsonl'):
    r = json.loads(l)
    print(r['n_tools'], r['tools_bytes'])"
```

Expected output on this machine, `claude-cli 2.1.165`, one MCP server attached:

```text
9 38116
858 522746
```

The name-list figure comes from the same file: filter `per_tool` to names beginning `mcp__` and sum their lengths. Expected `833` names and `30551` bytes.

**The four-step timing.** Every URL is public; only the invoke needs a credential:

```bash
curl -s -o /dev/null -w "%{http_code} %{time_total}s %{size_download}B\n" \
  "https://miscsubjects.com/api/dispatch?ask=what%20time%20is%20it"
curl -s -o /dev/null -w "%{http_code} %{time_total}s %{size_download}B\n" \
  "https://miscsubjects.com/api/dispatch?key=TIME_NOW&format=markdown"
```

The receipt for the invocation timed above is public and keyless: [`inv_3wt7dcbp2c`](https://miscsubjects.com/receipt/inv_3wt7dcbp2c).

[[embed:source:s7]]
[[embed:source:s18]]

**The gateway rows.** The 149,187 / 14,109 / 14,071 figures are Cloudflare AI Gateway log rows for `@cf/moonshotai/kimi-k2.7-code` on 2026-07-25; the method for producing them is [Why MCP tool schemas are most of your bill](/a/mcp-tool-search-cost). The design on the other side of this comparison is [Tooling as data](/a/tooling-as-data).

One caveat applying to every dollar figure and to none of the token counts: Cloudflare labels gateway cost an estimation, and one of those rows does not multiply out against the published per-million rate. The argument here is built on token counts for that reason.


## Sources

1. Tool search tool — Claude API documentation — https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool
2. Claude Code — Scale with MCP tool search — https://code.claude.com/docs/en/mcp
3. Claude Code — what invalidates the prompt cache — https://code.claude.com/docs/en/prompt-caching
4. Moonshot — Use Kimi with Claude Code — https://platform.kimi.ai/docs/guide/claude-code-kimi
5. Model Context Protocol specification — Tools — https://modelcontextprotocol.io/specification/2025-06-18/server/tools
6. Tool use with prompt caching — defer_loading and cache preservation — https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-caching
7. claude-code-cloudflare-gateway — the translator and the capture harness — https://github.com/redacted/claude-code-cloudflare-gateway
8. Desktop built-in MCP tool schemas load non-deferred with no opt-out — https://github.com/anthropics/claude-code/issues/76372
9. ToolSearch does not index tools from claude.ai-hosted MCP servers — https://github.com/anthropics/claude-code/issues/57033
10. codex exec can silently complete empty when configured MCP tools are deferred — https://github.com/openai/codex/issues/24536
11. tool_search auto-gate computes its threshold from the wrong model — https://github.com/NousResearch/hermes-agent/issues/57520
12. Dynamic MCP tool loading to reduce context window consumption — https://github.com/github/app/issues/1398
13. Context bloat: 16.9k tokens on MCP tool definitions loaded into the main agent — https://github.com/yonatangross/orchestkit/issues/885
14. Apideck CLI — an AI-agent interface with much lower context consumption than MCP — https://news.ycombinator.com/item?id=47400262
15. Comment on: MCP is dead? — https://news.ycombinator.com/item?id=48332411
16. Comment on: MCP is dead? — https://news.ycombinator.com/item?id=48332962
17. Comment on: Claude Advanced Tool Use — https://news.ycombinator.com/item?id=46039648
18. Comment on: I still prefer MCP over skills — https://news.ycombinator.com/item?id=47719499
19. Comment on: I still prefer MCP over skills — https://news.ycombinator.com/item?id=47719249
20. Comment on: Chrome DevTools MCP — https://news.ycombinator.com/item?id=47392361
21. Comment on: When does MCP make sense vs CLI? — https://news.ycombinator.com/item?id=47209810
22. Comment on: MCP is dead? — https://news.ycombinator.com/item?id=48336021
23. Comment on: MCP is dead; long live MCP — https://news.ycombinator.com/item?id=47381322
24. Comment on: MCP is dead? — https://news.ycombinator.com/item?id=48331540
25. First-party wire capture: 9 definitions / 38,116 bytes against 858 / 522,746 bytes — https://github.com/redacted/claude-code-cloudflare-gateway
26. First-party measurement: the ToolSearch definition is 1,440 bytes, and 833 deferred names are 30,551 — https://github.com/redacted/claude-code-cloudflare-gateway
27. First-party timing: four round trips, 2.185 s, 47,409 bytes, with the public receipt — https://miscsubjects.com/receipt/inv_3wt7dcbp2c
28. Gateway log rows: 149,187 / 14,109 / 14,071 input tokens on the same catalogue — https://miscsubjects.com/api/articles/mcp-tool-search-cost


---

# MCP definitions use 10.6× more input tokens until tool search defers them

slug: mcp-tool-search-cost · https://miscsubjects.com/a/mcp-tool-search-cost · tags: tooling, mcp, cost, tool-search, context-window, measurement · updated 2026-07-26T03:31:45.388Z

A connected MCP server costs nothing while its tools sit idle, and costs on every request, because the definitions travel with the request. **MCP** — Model Context Protocol — is the wire format an agent uses to discover and call tools on an external server. Its discovery call, `tools/list`, returns one object per tool: a name, a description, and a JSON Schema for the arguments. A client that keeps those objects in the request pays for all of them every turn, whether the model calls one or none.

## Evidence status

**Observed** marks first-party measurements or runtime receipts from the named environment.
**Derived** marks arithmetic calculated from cited inputs. **Specified** marks vendor or standards
documentation. **Implemented** and **deployed** name code and live-state evidence, respectively.
**Reproduced** means the stated procedure was rerun. **Externally attested** marks operator reports;
those reports show that an experience occurred, not that it is universal.

## A definition is a name, a sentence and a schema, and the schema is the part that grows

One live tool object, from `POST https://miscsubjects.com/api/mcp` with `{"jsonrpc":"2.0","id":1,"method":"tools/list"}`:

```json
{
  "name": "CF_AI_GATEWAY_LIST_LOGS",
  "description": "List Logs MCP: https://ai-gateway.mcp.cloudflare.com/sse [fn · cf_ai_gateway]",
  "inputSchema": {
    "type": "object",
    "properties": {
      "gateway_id": { "type": "string", "description": "The gateway ID." },
      "page": { "default": 1, "type": "integer" },
      "per_page": { "default": 20, "type": "integer" },
      "order_by": { "default": "created_at", "type": "string",
        "enum": ["created_at","provider","model","model_type","success",
                 "cached","cost","tokens_in","tokens_out","duration","feedback"] },
      "order_by_direction": { "default": "desc", "type": "string", "enum": ["asc","desc"] },
      "start_date": { "type": "string" },
      "end_date": { "type": "string" },
      "feedback": { "type": "number" },
      "success": { "type": "boolean" },
      "cached": { "type": "boolean" },
      "model": { "type": "string" },
      "provider": { "type": "string" }
    },
    "required": ["gateway_id"]
  }
}
```

That object is **264 tokens** by the `o200k_base` tokenizer: 11 for the name, 22 for the description, 231 for argument names, types, defaults and one `enum` of eleven strings.

A description is written once by a human and stays a sentence. A schema grows with the API behind it — every optional filter, enum member and nested object — and nothing prunes it. The MCP specification requires it: `tools/list` returns `name`, optional `title`, `description` and `inputSchema` per tool, so there is no conforming way to publish a tool without publishing its argument surface.

Across the whole 831-tool catalogue on 2026-07-26, names cost 4,064 tokens, descriptions 54,596, schemas 54,465. Level in aggregate, which is not the folklore. Split by size and the folklore returns:

| Slice of the catalogue | Tools | Tokens | Share that is schema |
| --- | --- | --- | --- |
| Definitions under 300 tokens | 800 | 109,856 | 39% |
| Definitions 300 tokens or more | 31 | 17,584 | 67% |
| The 20 most expensive definitions | 20 | 13,963 | 75% |
| The single most expensive, `CF_OBSERVABILITY_QUERY_WORKER_OBSERVABILITY` | 1 | 1,803 | 95% |

Cheap tools are mostly prose; expensive tools are almost entirely schema. A catalogue's bill is set by its handful of query-shaped tools, not its median row. An independent measurement on a different server found the same tail: G-Core's MCP server at `GCORE_TOOLS=*` advertises 741 tools for about 488,013 tokens, an average of 659 per tool, one tool alone at about 7,046 tokens of schema.

## Measure it three ways, cheapest first

**1. `/context`, in a running session.** Breaks the window into system prompt, system tools, MCP tools, memory, skills and messages. Free, one second. Read the contradiction section before trusting a zero.

**2. The `usage` field of any API response.** Every Messages API response carries `usage.input_tokens`, `usage.output_tokens` and the cache counters. Send the same one-word prompt twice, with the servers attached and without, and subtract. The only method that measures what you are billed for.

**3. A capture server, for what the client puts on the wire.** A local HTTP server that speaks enough of the Messages API to answer, and logs every request:

```bash
git clone https://github.com/redacted/claude-code-cloudflare-gateway
cd claude-code-cloudflare-gateway
node tools/capture-gateway.mjs      # listens on :8787, appends capture.jsonl
```

It records `n_tools`, `tool_names`, `system_chars`, `system_cache_control`, `metadata` and the model per request, redacting `authorization` and `x-api-key`. Point the client at it twice, changing one variable:

```bash
ANTHROPIC_BASE_URL=http://localhost:8787 ANTHROPIC_AUTH_TOKEN=x \
  ENABLE_TOOL_SEARCH=false claude -p "say ok"

ANTHROPIC_BASE_URL=http://localhost:8787 ANTHROPIC_AUTH_TOKEN=x \
  ENABLE_TOOL_SEARCH=true claude -p "say ok"
```

Read the two tool counts out of the log:

```bash
python3 -c "
import json
for line in open('capture.jsonl'):
    r = json.loads(line)
    if r.get('n_tools') is not None:
        print(r['model'], 'tools=', r['n_tools'])
"
```

Two lines. On `claude-cli 2.1.165` against that server on 2026-07-25: `tools= 856` and `tools= 9` — the nine being `Agent`, `AskUserQuestion`, `Bash`, `Edit`, `Read`, `Skill`, `ToolSearch`, `Workflow`, `Write`.

**To price a catalogue without running the client**, count the tokens the server publishes. The command behind the 831-tool figures above, good against any HTTP MCP server that answers `tools/list`:

```bash
curl -s -X POST https://miscsubjects.com/api/mcp \
  -H "Authorization: Bearer $MCP_TOKEN" -H 'content-type: application/json' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' -o tools.json

python3 -m venv /tmp/tokvenv && /tmp/tokvenv/bin/pip install tiktoken
/tmp/tokvenv/bin/python -c "
import json, tiktoken
enc = tiktoken.get_encoding('o200k_base')
t = json.load(open('tools.json'))['result']['tools']
tot = sum(len(enc.encode(json.dumps(x))) for x in t)
print('tools', len(t), 'tokens', tot, 'mean', round(tot/len(t), 1))
"
```

Output on 2026-07-26: `tools 831 tokens 127440 mean 153.4`, from a 434,636-byte response. `o200k_base` is OpenAI's tokenizer, not Anthropic's — treat it as a close estimate and billed `usage` as the truth. They agree to about 12% here: 149,187 billed input tokens for 856 definitions plus a system prompt and a short user message is 174.3 tokens per definition against the tokenizer's 153.4.

## Three configurations, same catalogue, same day

Same machine, same catalogue of 891 capabilities, same trivial prompt, `@cf/moonshotai/kimi-k2.7-code` through a Cloudflare AI Gateway, 2026-07-25. Tokens and cost read from the gateway's own log rows.

| Configuration | Tool definitions in the request | Input tokens | Cost per turn |
| --- | --- | --- | --- |
| Every directory row projected as an MCP tool, definitions in context | 856 | 149,187 | $0.02852109 |
| Same catalogue, `ENABLE_TOOL_SEARCH=true` | 9, one of them `ToolSearch` | 14,109 | $0.00443075 |
| No MCP server attached, the same 891 capabilities reached over HTTP | 9 built-in tools | 14,071 | $0.00456265 |

149,187 input tokens is 74.6% of a 200,000-token window, spent before the user's sentence is read. Deferring the definitions cuts that 10.6-fold, 135,078 tokens a turn.

The third row decides architecture: **tool search on is as cheap as having no MCP server at all** — 14,109 against 14,071, a 0.3% difference — with every tool still reachable. The tools were never the cost. The definitions were.

The dollar column is the gateway's own accounting and is an estimate; the token counts are measured at both ends and the argument rests on them.

That 14,071-token row is still readable. Fetched 2026-07-26 from the AI Gateway logs REST endpoint (`GET /accounts/<account_id>/ai-gateway/gateways/default/logs?per_page=3&order_by=created_at&order_by_direction=desc`):

```json
{"created_at":"2026-07-26T03:43:18.667Z","model":"@cf/moonshotai/kimi-k2.7-code",
 "metadata":{"via":"claude-code","shim":"api/aig","tools":9},
 "tokens_in":14071,"tokens_out":170,"cost":0.00456265,
 "usage_metadata":{"input_cached_tokens":12480}}
```

`metadata.tools` is what the client sent: nine, not 856, with the same catalogue attached.

## The attribution line was worth more than 12,000 cached tokens a turn

Claude Code prepends an attribution block — client version and prompt fingerprint — to the start of the system prompt. A prompt cache keys on an exact prefix, so a value that changes per request, at the front, means the prefix never matches and nothing before it is ever a hit. Anthropic's environment-variable reference: set `CLAUDE_CODE_ATTRIBUTION_HEADER` to `0` to omit the block, "Disabling it improves prompt-cache hit rates when routing through an LLM gateway. Caching on a direct connection to the Anthropic API is unaffected either way."

Measured on the rows above, cached input went from **64 tokens to 12,480 tokens** per turn once the line was dropped — 88% of a 14,109-token turn arriving from cache instead of priced as fresh.

```bash
export CLAUDE_CODE_ATTRIBUTION_HEADER=0
```

Anthropic's gateway protocol reference adds the version detail and the right place to fix it: from v2.1.181 the block is stable for the lifetime of a conversation behind a custom base URL, and "If your gateway must reshape system content, set `CLAUDE_CODE_ATTRIBUTION_HEADER=0` so Claude Code omits the block… omit it at the client rather than stripping or moving it in the gateway." Stripping it gateway-side breaks attribution downstream; omitting it at the client does not.

## Deferred tool search: names stay, schemas arrive on request

Anthropic's documentation describes the mechanism: "When tool search is active, tool definitions are withheld from the context window. The agent receives a summary of available tools and searches for relevant ones when the task requires a capability not already loaded. Up to five of the most relevant tools are loaded into context by default." A discovered tool stays available for later turns; if compaction removes it, the model searches again.

The switch:

```bash
export ENABLE_TOOL_SEARCH=true
```

Permanently, in `~/.claude/settings.json`:

```json
{ "env": { "ENABLE_TOOL_SEARCH": "true", "CLAUDE_CODE_ATTRIBUTION_HEADER": "0" } }
```

Every value the variable accepts:

| Value | Behaviour |
| --- | --- |
| unset | Deferred by default — but **loaded upfront** when `ANTHROPIC_BASE_URL` points at a non-first-party host, or on Google Cloud's Agent Platform |
| `true` | Always defer, and send the beta header even through a proxy. Requests fail on proxies that do not support `tool_reference` blocks |
| `auto` | Load upfront if the definitions fit within 10% of the context window, defer the overflow |
| `auto:N` | Same with a custom percentage, e.g. `auto:5` |
| `false` | Load every definition upfront, every turn |

Two costs come with it, both small. The search tool's own definition stays in the request — it is one of the nine measured above, and Anthropic's API reference is explicit that "At least one tool, normally the tool search tool itself, must stay non-deferred." And the first use of an unseen tool costs an extra round-trip; below roughly ten tools, loading everything upfront is faster.

The API-side mechanism differs from the client-side one in a way that matters if you are building a gateway. With `defer_loading: true` on the Messages API, "You still send every tool's full definition in the `tools` array on every request, including the deferred ones. The API needs them server-side to run the search." The saving there is context, not bytes. The capture above shows the client doing the other thing — sending nine definitions — because behind a non-first-party base URL Claude Code resolves the search itself. Both are called tool search. Only one shrinks what leaves your machine.

## Five ways the definitions stay in the bill anyway

| What still costs | Where it was observed | The report |
| --- | --- | --- |
| First-party servers exempt from deferral | Claude Desktop, `anthropics/claude-code` issue 76372, 2026-07-10 | With tool search active, third-party MCP tools defer to names only, but "three Desktop built-in servers load complete schemas upfront every session" — about 3,900 tokens with no opt-out, isolated by reading `message.usage` from session JSONL across 6 sessions |
| Servers invisible to the search index | Claude Code 2.1.114, `anthropics/claude-code` issue 57033, 2026-05-07 | Servers added at claude.ai/settings/connectors show Connected in `/mcp`, yet "Any `ToolSearch` query that should match a claude.ai MCP tool returns zero results". Local `.mcp.json` and plugin servers index fine |
| Threshold computed from the wrong model | `hermes-agent` issue 57520, 2026-07-03 | The auto-gate reads the configured default model, not the session model, so "the gate is scaled to the wrong window" — a 98,304-token local model gets a 25,600-token threshold derived from a 256K cloud model instead of 9,830, and takes the whole payload inline |
| A turn that silently completes empty | codex-cli 0.133.0, `openai/codex` issue 24536, 2026-05-26 | "`codex exec` can silently finish with no assistant message when an explicitly configured MCP tool is deferred behind `tool_search`" — the server stays healthy and registered but is only reachable through deferral, and the harness accepts the empty turn |
| Lazy loading invalidating the prompt cache | Hacker News, 2026-03-01 | "The main problem with this approach at the moment is it busts your prompt cache, because LLMs expect all tool definitions to be defined at the beginning of the context window" |

The cache row cuts the other way for the deferred case. Anthropic's caching page puts tool definitions in the system-prompt layer, so the cache invalidates when the set of definitions changes between turns — but with deferred tools, "a server connecting, disconnecting, or changing its tool list only appends new content and doesn't disturb anything already cached", while with tools in the prefix "any change to them invalidates the cache". A gateway is named as one place where deferral is off and the prefix is therefore fragile: a stdio process exiting, an HTTP session expiring or an automatic reconnect invalidates the whole cached prefix with nobody touching a keyboard.

An operator who ran it and was not convinced, on Hacker News on 2026-03-15: "And no, the tool search function recently introduced by Anthropic does not completely solve this problem."

## One harness reports zero tokens per tool. Others publish tables of tens of thousands

Zero, filed against Claude Code as issue 23228 on 2026-02-05: "When running the `/context` command to check token usage, the output includes a long list of all MCP tool definitions (28 in my case), each showing '0 tokens'."

Not zero, filed on 2026-03-01 with a per-server breakdown: "The main agent context burns 16.9k tokens (8.5% of 200k) loading 68 MCP tool schemas at session start. Most are never called by the main agent — they're used by subagents." That author's table attributes 6.0k to one server's 25 tools, 3.3k to another's 14. Two more, independently: 741 tools at about 488,013 tokens on a G-Core server, and 250-plus definitions across 9 servers at "~40,000-70,000 tokens of tool definitions loaded upfront".

Both can be true, and the mechanism is documented. `/context` reports what is **in the context window**, category by category. Deferred definitions are by construction not in the window — Anthropic's cost guidance: "MCP tool definitions are deferred by default, so only tool names enter context until Claude uses a specific tool." A per-tool row of `0 tokens` is then accurate, and the residual cost lands under another heading. The tables of tens of thousands come from sessions where the definitions were not deferred — an older client, a `false` setting, a non-first-party base URL, `alwaysLoad`, an exempt built-in server — and from harnesses that count the catalogue rather than ask the client.

Which to trust for a decision: **neither, over the `usage` field of your own responses.** `/context` reports placement, not billing, and its zero is silent about the 3,900 tokens of exempt built-in schemas in issue 76372. A tokenizer table reports the size of a catalogue, not what your client chose to send — the capture above shows a client sending 9 of 856. Only `usage.input_tokens`, differenced across two otherwise identical runs, answers the question being asked.

One dissent earns its place, from Hacker News on 2026-05-30: "The idea that MCP tool definitions take up a certain number of tokens is laughable. That's an implementation detail of the agent harness." Correct about the protocol, beside the point about the invoice: the protocol mandates no loading strategy, and the client you are running has already picked one and is billing you for it.

## The arithmetic at 200 turns a day

200 model turns in a working day, 30 days in a month. Substitute your own turn count; the multiplication is the same.

| Configuration | Cost per turn | × 200 turns = per day | × 30 days = per month |
| --- | --- | --- | --- |
| 856 definitions in every request | $0.02852109 | $5.70 | $171.13 |
| `ENABLE_TOOL_SEARCH=true` | $0.00443075 | $0.89 | $26.58 |
| No MCP server, capabilities over HTTP | $0.00456265 | $0.91 | $27.38 |

One environment variable is worth $4.82 a day and **$144.54 a month** on one machine, one catalogue, these rates. In tokens, which do not depend on anyone's pricing: 135,078 fewer input tokens per turn, 27,015,600 fewer per day.

Two effects the table does not price. The definitions occupied 74.6% of a 200,000-token window, so the ceiling on what a session can hold moves further than the invoice does. And tool-selection accuracy is documented to degrade "with more than 30-50 tools loaded at once" — a quality cost with no line item.

## Five responses, ranked by what they cost you to adopt

**1. Turn tool search on.** One variable, no code, every tool still reachable, measured 10.6× fewer input tokens. Trade-off: an extra round-trip on first use of an unseen tool; unavailable behind a base URL that cannot forward `tool_reference` blocks; below ten tools it is slower than loading them.

**2. Scope servers per project.** Enable a server only in the repository that needs it. An operator on Hacker News, 2026-07-12: "I enable tools specific to each project only in that project, and have very very few in my global config. Like <5k tokens worth." Trade-off: manual, and it fails the way manual hygiene always fails — the day you forget.

**3. Collapse many tools into few.** Replace N tools with one or two that take a name and arguments and dispatch internally. From the operator who built it: "This basically takes your APIs, databases, and docs and compresses them into 2 MCP tools (~1,000 tokens) instead of N tools (100K+ tokens)." A minimal version of the same idea, reported at about 60 tokens: one bare tool with a one-line description that unlocks the rest when called. Trade-off: you have rebuilt tool search without the harness's search quality, and the model must learn your dispatch convention.

**4. Take the catalogue out of the prompt entirely.** Publish capabilities behind an HTTP contract the model reads on demand, keep only the built-in tools. Measured at 14,071 input tokens with 891 capabilities reachable — the same number as tool search, a different structure, because catalogue size is no longer a term in the per-turn equation and no host support is required. Trade-off: your capabilities are not MCP tools, so anything that consumes MCP does not see them. Dimension by dimension in [Tool search versus a catalogue as data](/a/tool-search-vs-catalogue-as-data), the data layer in [Tooling as data](/a/tooling-as-data), the three exposures side by side in [MCP as a projection](/a/mcp-as-a-projection).

**5. Do nothing.** Defensible at a small tool count: under about ten tools, upfront loading is faster and the search round-trip is pure overhead. At 150–200 tools, two reports show the limit: one one-word prompt produced a request with "`prompt_tokens: 154,367`"; in the other, roughly 200 tool schemas exceeded a 32k window so completely that "`/compact` succeeded in producing a summary, but the very next request still failed."

## Symptom, cause, fix

| Symptom | Cause | Fix |
| --- | --- | --- |
| Input tokens in the tens of thousands before your first word | Every definition in the prefix, every turn | `ENABLE_TOOL_SEARCH=true` |
| Setting it changed nothing, tools still all loaded | `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS` is set — it "keeps tool search off, and `ENABLE_TOOL_SEARCH` can't override it" | Unset it, or accept upfront loading |
| Request fails after enabling it, behind a proxy | The proxy does not forward `tool_reference` blocks; errors of the form "Unexpected value(s) for the `anthropic-beta` header" or "Extra inputs are not permitted" | Fix the proxy to pass the beta header through, or set `ENABLE_TOOL_SEARCH=false` |
| Deferred by default on a direct connection, upfront through a gateway | Deferral is off by default "when `ANTHROPIC_BASE_URL` points to a non-first-party host" | Set `ENABLE_TOOL_SEARCH=true` explicitly |
| Tool search finds nothing for a server that `/mcp` shows Connected | claude.ai-hosted connectors are not in the search index (issue 57033) | Add the server via local `.mcp.json` instead |
| A scripted run ends with no assistant message | The required tool was deferred and the harness accepted an empty turn (codex issue 24536) | Mark that server `alwaysLoad`, or pin it non-deferred |
| One server must never need a search step | Deferral applies to every server by default | `"alwaysLoad": true` on that server in `.mcp.json` — every one of its tools then loads at session start regardless of `ENABLE_TOOL_SEARCH` |
| Cached input near zero behind a gateway | The attribution block changes the prefix per request | `CLAUDE_CODE_ATTRIBUTION_HEADER=0` at the client, never stripped in the gateway |
| Cache misses mid-session for no visible reason | A server connected, disconnected or pushed a tool-list change while definitions sat in the prefix | Defer the tools, so a server change "only appends new content" |
| `/context` shows every MCP tool at 0 tokens | Deferred definitions are not in the window; `/context` reports placement, not billing | Difference `usage.input_tokens` across two runs instead |
| Tool search on and the bill barely moved | Built-in or `alwaysLoad` servers are exempt (issue 76372, about 3,900 tokens) | Audit with a capture server; `n_tools` in the log is the ground truth |

The gateway that produced these log rows, and how the same client runs on a non-Anthropic model, is documented in [Claude Code on Kimi, GLM or Grok through your own Cloudflare account](/a/claude-code-on-cloudflare-ai-gateway).


## Sources

1. Claude Code environment variables — ENABLE_TOOL_SEARCH and CLAUDE_CODE_ATTRIBUTION_HEADER — https://code.claude.com/docs/en/env-vars
2. Connect Claude Code to tools via MCP — scale with MCP tool search — https://code.claude.com/docs/en/mcp
3. How Claude Code uses prompt caching — what invalidates a cached prefix — https://code.claude.com/docs/en/prompt-caching
4. Gateway protocol reference — where to disable the attribution block — https://code.claude.com/docs/en/llm-gateway-protocol
5. Manage costs effectively — MCP tool definitions are deferred by default — https://code.claude.com/docs/en/costs
6. Debug your configuration — what /context actually reports — https://code.claude.com/docs/en/debug-your-config
7. Scale to many tools with tool search — the mechanism and its break-even point — https://code.claude.com/docs/en/agent-sdk/tool-search
8. Model Context Protocol specification — tools/list — https://modelcontextprotocol.io/specification/2025-06-18/server/tools
9. Tool search tool — defer_loading controls context, not the request body — https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool
10. capture-gateway.mjs — the logging server behind these wire measurements — https://github.com/redacted/claude-code-cloudflare-gateway
11. tiktoken harness against a live MCP server: 741 tools, ~488,013 tokens — https://github.com/G-Core/gcore-mcp-server/issues/14
12. /context shows 28 MCP tool definitions, each at 0 tokens — https://github.com/anthropics/claude-code/issues/23228
13. 16.9k tokens on 68 MCP tool schemas, broken down per server — https://github.com/yonatangross/orchestkit/issues/885
14. 9 MCP servers, 250+ definitions, 40-70k tokens on every message — https://github.com/anomalyco/opencode/issues/35376
15. Built-in servers load full schemas despite tool search — ~3.9k tokens, no opt-out — https://github.com/anthropics/claude-code/issues/76372
16. ToolSearch does not index claude.ai-hosted servers — https://github.com/anthropics/claude-code/issues/57033
17. A deferred tool can end a scripted run with no assistant message — https://github.com/openai/codex/issues/24536
18. The deferral threshold is computed from the wrong model — https://github.com/NousResearch/hermes-agent/issues/57520
19. Loading tools lazily busts the prompt cache — https://news.ycombinator.com/item?id=47209810
20. A first-hand verdict: tool search does not completely solve it — https://news.ycombinator.com/item?id=47392361
21. The dissent: token cost is a harness detail, not a protocol property — https://news.ycombinator.com/item?id=48331540
22. The positive report: deferral retired the context-bloat complaint — https://news.ycombinator.com/item?id=48332962
23. Per-project scoping keeps a global tool surface under 5k tokens — https://news.ycombinator.com/item?id=48885036
24. Collapsing an API into two MCP tools: 100K+ tokens down to ~1,000 — https://news.ycombinator.com/item?id=47614267
25. A hand-rolled gate tool at about 60 tokens — https://news.ycombinator.com/item?id=47719249
26. A unified-API vendor measured 50,000+ tokens before the first user message — https://news.ycombinator.com/item?id=47400262
27. One popular server named at about 50k tokens — https://news.ycombinator.com/item?id=45955033
28. A one-word prompt that cost 154,367 input tokens — https://github.com/nimbalyst/nimbalyst/issues/914
29. Schema overhead that /compact cannot recover — https://github.com/ruvnet/ruflo/issues/2726
30. First-party: 831 published tool definitions, 127,440 tokens — https://miscsubjects.com/a/mcp-tool-search-cost
31. First-party: one real definition, 264 tokens, and where the tokens sit — https://miscsubjects.com/a/mcp-tool-search-cost
32. First-party: a gateway log row, nine tools, 14,071 input tokens — https://developers.cloudflare.com/ai-gateway/observability/logging/
33. First-party: 856 definitions become 9 on the wire — https://github.com/redacted/claude-code-cloudflare-gateway
34. First-party: three configurations priced from gateway log rows — https://miscsubjects.com/a/mcp-tool-search-cost

