# Cloudflare OS: the ledger as a table

slug: cloudflare-os-xl-02-ledger-as-a-table · https://miscsubjects.com/a/cloudflare-os-xl-02-ledger-as-a-table · category: systems · tags: cloudflare, pipelines, r2, ledger, analytics · updated 2026-08-06T03:28:33.400Z

*Part 2 of [Cloudflare OS XL](/a/cloudflare-os-xl), an inventory of the Cloudflare platform this build does not have installed.*

The central claim of this build is that nothing is ever overwritten and every action appends a hash-chained audit row. That claim is true. The rows exist, in a D1 database called `loop-shared-events` and in R2 as receipt files.

Then someone asks a question of it — how many outbound sends went to a domain whose MX record failed, per week, since May — and the answer is produced by pulling files and counting them in a script. The ledger is a record. It is not yet a table anybody can query.

Five products close the distance between those two things.

## Pipelines

Pipelines is Cloudflare's streaming ingest: data arrives over HTTP or from a Worker binding, is transformed with SQL, and is delivered to R2 as Apache Iceberg tables or as Parquet and JSON files. It is in open beta.

Today, every ledger append is a D1 `INSERT` executed inside the request that caused it. That has three costs. It puts a write in the hot path of the thing being recorded. It makes the ledger's throughput a function of D1's write throughput. And it produces rows, not columns — which is why analytical questions are answered by export-and-count.

With a pipeline, the Worker writes an event to a binding and returns. The pipeline batches, transforms and lands it in R2 in a columnar format. The record is still append-only and still hash-chained; it is simply stored as something a query engine can read.

The natural first candidates here are the three highest-volume event streams: agent turns, tool invocations, and outbound send receipts.

**Verdict: install, for agent turns first.** It is beta, so it belongs on the stream where a gap would be survivable, not on the audit chain.

## R2 Data Catalog and R2 SQL

R2 Data Catalog is a managed Apache Iceberg catalog built into an R2 bucket. R2 SQL is a distributed SQL engine that queries it. Together they are the reason the previous section says "Iceberg" rather than "Parquet files in a bucket": Iceberg gives the pile of files a schema, a snapshot history and a table identity, and R2 SQL means you do not have to bring your own engine to read it.

Enabling the catalog on an existing bucket is one command.

```
wrangler r2 bucket catalog enable miscsubjects-ledger
```

What this changes for this build is the nature of an audit. The audit chain is the build's trust mechanism; it is what makes the claim "nothing is ever overwritten" checkable rather than asserted. But a trust mechanism that can only be verified by a bespoke script is verified by whoever wrote the script. A ledger as an Iceberg table can be queried by anyone with the credential, including a model, including an outside auditor, with a plain `SELECT`.

There is a second, quieter benefit. Time-travel is a property of Iceberg, not something this build would have to implement: the table can be read as of a snapshot. "What did the ledger say on 3 August" stops being a question about backups.

**Verdict: install.** Low cost, and it converts an existing asset into a queryable one without moving it out of R2.

## R2 event notifications

An object lands in R2 and a message appears on a queue. That is the whole feature, and it is missing from a build that has three queues already.

Right now, assets get processed because a cron woke up and looked. Generated hero images, ArcAds output, absorbed repositories, uploaded references — each of those arrives in a bucket and then waits for a scheduled sweep to notice. The sweep runs every minute, which is fast enough to feel instant and is still the wrong mechanism: it polls whether or not anything happened, and it cannot tell you *why* it processed something.

```
wrangler r2 bucket notification create miscsubjects-store --event-type object-create --queue loop-tasks
```

With that, the arrival of the object *is* the trigger. The queue message carries the bucket, the key and the event type, so the consumer knows exactly what changed rather than diffing a listing.

**Verdict: install.** It is one command per bucket and it deletes polling code.

## Analytics Engine

Analytics Engine accepts unlimited-cardinality analytics written from a Worker and queried with SQL. Writes are non-blocking and effectively free; you get one dataset binding and you write data points with blobs, doubles and an index.

```toml
[[analytics_engine_datasets]]
binding = "METRICS"
dataset = "loop_metrics"
```

```js
env.METRICS.writeDataPoint({
  blobs: [toolName, modelId, agentName, outcome],
  doubles: [latencyMs, tokensIn, tokensOut, costUsd],
  indexes: [agentName],
});
```

This build already tries to answer cost and latency questions — there is a `COST_REPORT` row, a governor, and per-model accounting. Those work by reading the ledger back and aggregating it, which means the cost of asking a cost question scales with the size of the ledger.

Analytics Engine is the correct tool for that specific class of question because it is designed for high-cardinality dimensions. Per-tool, per-model, per-agent, per-outcome, forever, at a write cost that does not compete with the request. The ledger keeps being the record of what happened; Analytics Engine becomes the record of how much it cost and how long it took.

The one constraint worth knowing before adopting it: it is a metrics store, not an event store. Data points are sampled at high volume and are not the audit trail. Do not put anything in it that has to be exact.

**Verdict: install.** It answers the cost question this build keeps asking, and it does not compete with the ledger for that role.

## What this part does not recommend

**Do not move the audit chain off D1.** The hash chain's value is that each row commits to the previous one at write time, inside a transaction, in the same request that performed the act. Streaming it through a batching pipeline first would put a gap between the act and the commitment, and the gap is exactly what the chain exists to close. Pipelines belongs on the high-volume observational streams. The chain stays where it is.

## Verdicts

| Product | What it replaces here | Verdict |
| --- | --- | --- |
| Pipelines | Per-row D1 inserts in the hot path for high-volume streams | **install** — agent turns first |
| R2 Data Catalog | A ledger auditable only by a bespoke script | **install** |
| R2 SQL | Export-and-count in a local script | **install** — with the catalog |
| R2 event notifications | A cron sweep that polls buckets every minute | **install** |
| Analytics Engine | Cost and latency questions answered by re-reading the ledger | **install** |
| Pipelines *for the audit chain* | Nothing — it would weaken it | **no** |

Next: [Part 3 — running real code](/a/cloudflare-os-xl-03-running-real-code).


## Sources

1. Cloudflare Pipelines documentation — https://developers.cloudflare.com/pipelines/
2. R2 Data Catalog documentation — https://developers.cloudflare.com/r2/data-catalog/
3. R2 SQL documentation — https://developers.cloudflare.com/r2-sql/
4. Workers Analytics Engine documentation — https://developers.cloudflare.com/analytics/analytics-engine/


---

# R2 cuts a 10 TB delivery bill from $923 to $18.45

slug: cloudflare-os-r2 · https://miscsubjects.com/a/cloudflare-os-r2 · tags: cloudflare, architecture, r2, cloudflare-os · updated 2026-07-26T03:59:27.570Z

Cloudflare R2 is object storage: give it a key like `img/up/hero.png` and some bytes, and it hands them back on request. It speaks two dialects: the Amazon S3 HTTP API, so existing S3 tools work against it, and a native binding inside a Cloudflare Worker where the bucket is a JavaScript object with `put`, `get`, `list` and `delete`. Objects go to 5 TiB, keys to 1,024 bytes, and a bucket holds any number of them.

The reason anyone brings it up is the price of getting bytes *out*. Amazon charges for that. Cloudflare does not. The rest of this page is the arithmetic that follows, plus the things R2 will not do for you.

[[embed:source:s2]]

## Evidence status

**Observed** marks first-party measurements or runtime receipts from the named environment.
**Derived** marks arithmetic calculated from cited inputs. **Specified** marks vendor or standards
documentation. **Implemented** and **deployed** name code and live-state evidence, respectively.
**Reproduced** means the stated procedure was rerun. **Externally attested** marks operator reports;
those reports show that an experience occurred, not that it is universal.

## Zero egress is real; the meter is on the operations

R2 bills three things: how much you store, and two classes of API call. **Class A** operations change state: `PutObject`, `CopyObject`, `ListObjects`, `CreateMultipartUpload`, `UploadPart`, `CompleteMultipartUpload`. **Class B** operations read state: `GetObject`, `HeadObject`, `HeadBucket`. Deletes are free. Bandwidth to the internet is free.

Rates, read from the pricing page on 2026-07-26:

| | Standard | Infrequent Access |
| --- | --- | --- |
| Storage | $0.015 / GB-month | $0.01 / GB-month |
| Class A (writes, lists) | $4.50 / million | $9.00 / million |
| Class B (reads) | $0.36 / million | $0.90 / million |
| Data retrieval | none | $0.01 / GB |
| Egress to the internet | free | free |
| Free each month | 10 GB-month, 1M Class A, 10M Class B | none; the free tier is Standard only |
| Minimum storage duration | none | 30 days |

Cloudflare rounds usage up to the next unit: 1,000,001 operations bills as two million, 1.1 GB-month bills as 2 GB-month.

[[embed:source:s1]]

A reader of that same page put the obvious objection plainly.

[[embed:source:s21]]

He is describing the meter correctly and drawing the wrong conclusion. Price the overage. Class B beyond the free 10 million costs $0.36 per additional million. On the ten-terabyte workload computed below, the S3 egress line is $891.00. For R2's metered operations to cost that much, you would have to make **2.475 billion** Class B calls beyond the free allowance in one month: 891 ÷ 0.36 × 1,000,000. The class-ops meter is real. It is roughly three orders of magnitude away from being the thing that costs you money.

## Serving 10 TB a month: $18.45 on R2, $923.00 on S3

The workload: 1,000 GB stored, average object 500 KB (so 2,000,000 objects), 200,000 new objects written during the month, and 10,000 GB served to the public internet, or 20,000,000 GET requests. S3 prices are `us-east-1`, pulled from the AWS Price List API on 2026-07-26: Standard storage $0.023/GB-month, Tier-1 requests (PUT, COPY, POST, LIST) $0.005 per 1,000, Tier-2 requests (GET and all others) $0.0004 per 1,000, data transfer out to the internet $0.09/GB for the first 10 TB beyond the 100 GB monthly free allowance.

| Line | R2 arithmetic | R2 | S3 arithmetic | S3 |
| --- | --- | --- | --- | --- |
| Storage | (1,000 − 10 free) × $0.015 | $14.85 | 1,000 × $0.023 | $23.00 |
| Writes | 200,000 of 1,000,000 free | $0.00 | 200,000 × $0.005/1,000 | $1.00 |
| Reads | (20,000,000 − 10,000,000) × $0.36/M | $3.60 | 20,000,000 × $0.0004/1,000 | $8.00 |
| Egress | 10,000 GB, unmetered | $0.00 | (10,000 − 100) × $0.09 | $891.00 |
| **Total** | | **$18.45** | | **$923.00** |

Fifty times cheaper, and the ratio is almost entirely one line: egress is 96.5% of the S3 bill. The table also shows that R2's operation rates are not a gimmick to claw the egress back. Class A at $4.50/million undercuts S3's Tier-1 at $5.00/million, and Class B at $0.36/million undercuts Tier-2 at $0.40/million. R2 is 10% cheaper per call *and* free on bandwidth.

[[embed:source:s15]]

Cloudflare's CTO stated the free-tier gap when R2 launched, and the allowances he named still hold on the page fetched today.

[[embed:source:s19]]

Two published totals for real public-serving workloads, both itemised. A 33 GB WordPress plugin mirror on R2 plus Workers, and a full archive of every SEC filing served at roughly twice the SEC's own volume:

[[embed:source:s24]]

[[embed:source:s26]]

Neither workload needed versioning, and neither was cold. That is why they land where they do.

## The workload where S3 wins is the one you never read

R2 has exactly two storage classes. S3 has a ladder that goes much colder. For a write-once archive that is almost never read and never leaves the cloud, the egress advantage is worth nothing and the storage floor decides.

The workload: 100,000 GB (100 TB) of compliance records, written once, read a handful of times a year, consumed inside the same cloud.

| Where it sits | Rate | Monthly |
| --- | --- | --- |
| R2 Standard | (100,000 − 10) × $0.015 | $1,499.85 |
| R2 Infrequent Access | 100,000 × $0.01 (no free tier applies) | $1,000.00 |
| S3 Glacier Flexible Retrieval | 100,000 × $0.0036 | $360.00 |
| S3 Deep Archive Access tier | 100,000 × $0.00099 | $99.00 |

R2's cheapest class costs 2.8× the Glacier Flexible rate and 10.1× the Deep Archive rate. There is no colder tier to move to; Standard and Infrequent Access are the whole ladder. If your bytes are cold and captive, stay on S3.

[[embed:source:s5]]

The person who ran both and stayed on AWS drew the boundary in one sentence.

[[embed:source:s22]]

## Infrequent Access pays only below about one read every two months

Infrequent Access looks like a third off the storage price. The retrieval fee eats it almost immediately. Take a 1 GB object held for one month and read `r` times:

- Standard: `$0.015 + r × $0.00000036`
- Infrequent Access: `$0.010 + r × $0.0000009 + r × 1 GB × $0.01`

Set them equal: `0.005 = 0.01000054r`, so `r ≈ 0.5`. The retrieval fee is 99.99% of the right-hand side; the operation-rate difference is noise. **A 1 GB object must be read less than once every two months for Infrequent Access to be cheaper.** Add the 30-day minimum billing duration, where you pay a full month even if you delete on day two, and the class is for backups and cold originals, nothing else.

Move objects there with a lifecycle rule rather than by hand. The transition itself is billed as a Class A operation.

```bash
npx wrangler r2 bucket lifecycle add miscsubjects-ledger \
  --name "archive-old-events" \
  --prefix "events/" \
  --storage-class InfrequentAccess \
  --transition-days 30
```

Lifecycle rules also delete: an expiration rule on a prefix is how an ingest bucket stops growing forever, and how incomplete multipart uploads get cleaned up. They expire after 7 days by default. A bucket accepts up to 1,000 rules.

[[embed:source:s6]]

## R2 will not keep the old version of an object

This is the gap that costs people data. When you `put` to a key that already exists, the previous bytes are gone.

[[embed:source:s23]]

The S3 compatibility table marks `PutBucketVersioning`, `GetBucketVersioning`, `PutObjectLockConfiguration` and `GetObjectLockConfiguration` all unsupported. Versioning and Object Lock are the two standard defences against ransomware and against a human running the wrong script; R2 offers neither in S3's form.

[[embed:source:s4]]

It does have **bucket locks**, a per-prefix retention rule blocking deletion and overwriting for a fixed period or indefinitely, enforced with `10069 / ObjectLockedByBucketPolicy` and HTTP 403. That stops deletion. It does not give you "fetch me yesterday's copy".

[[embed:source:s7]]

For that you write versioning yourself: never overwrite a key, always write a new one, keep a pointer.

```js
// Content-addressed writes: the key carries the version, so nothing is ever overwritten.
export async function putVersioned(env, logicalKey, body, contentType) {
  const bytes = body instanceof ArrayBuffer ? body : new TextEncoder().encode(body);
  const digest = await crypto.subtle.digest('SHA-256', bytes);
  const hash = [...new Uint8Array(digest)].map(b => b.toString(16).padStart(2, '0')).join('');
  const versionKey = `${logicalKey}/${Date.now()}-${hash.slice(0, 12)}`;
  await env.R2.put(versionKey, bytes, { httpMetadata: { contentType } });
  // The pointer is one small object; reading it is one Class B call.
  await env.R2.put(`${logicalKey}/current`, versionKey, {
    httpMetadata: { contentType: 'text/plain' },
  });
  return { versionKey, hash };
}

export async function getCurrent(env, logicalKey) {
  const ptr = await env.R2.get(`${logicalKey}/current`);
  if (!ptr) return null;
  return env.R2.get(await ptr.text());   // second Class B call
}
```

The cost of doing it this way: every read is two Class B operations instead of one, and old versions accumulate until a lifecycle rule expires them. At $0.36 per million reads, the doubled read is $0.36 per million objects fetched. That is the actual price of the missing feature.

## The free-egress trust question, answered from the terms

The sharpest objection is not about the price list but about whether it is load-bearing.

[[embed:source:s25]]

The worry points at a real clause, and the clause says the opposite of what it assumes. Cloudflare's Service-Specific Terms restrict the **CDN** on Free, Pro and Business plans: *"Unless you are an Enterprise customer, Cloudflare offers specific Paid Services (e.g., the Developer Platform, Images, and Stream) that you must use in order to serve video and other large files via the CDN."* The Developer Platform named there is the thing R2 belongs to. Using R2 to serve large files is the compliant path, not the risky one.

The Developer Platform section carries its own limit, and it is a different kind: *"Cloudflare may temporarily limit your storage and/or the number of requests you can make or receive using the Developer Platform if processing such requests would put an undue burden on the Cloudflare network."* Rate limiting, stated in advance. Not a retroactive per-GB bill.

Verdict: the enforcement risk on R2 is throttling under abnormal load, not an undisclosed egress charge. The pricing page's footnote is unambiguous: *"Egressing directly from R2, including via the Workers API, S3 API, and r2.dev domains does not incur data transfer (egress) charges and is free."* What it does not cover is metered services you bolt on top; those bill separately.

[[embed:source:s14]]

## Every call, and which meter it hits

Bind the bucket first. In `wrangler.toml`:

```toml
[[r2_buckets]]
binding = "R2"
bucket_name = "miscsubjects-ledger"
```

`binding` is the variable name your code sees; `bucket_name` is the real bucket. Create the bucket before the first deploy, or the binding fails:

```bash
npx wrangler r2 bucket create miscsubjects-ledger
npx wrangler r2 bucket list
```

| Call | What it does | Meter |
| --- | --- | --- |
| `env.R2.put(key, value, opts)` | Stores bytes; returns an `R2Object`. Strongly consistent: once the promise resolves, every reader worldwide sees it | Class A |
| `env.R2.get(key, opts)` | Returns `R2ObjectBody` with `.body` as a stream, or `null` if absent | Class B |
| `env.R2.head(key)` | Metadata only, no body, or `null` | Class B |
| `env.R2.list(opts)` | Up to 1,000 keys, lexicographic, `opts.prefix` to scope | Class A |
| `env.R2.delete(key or key[])` | Up to 1,000 keys per call | free |
| `env.R2.createMultipartUpload(key)` | Starts a multipart upload | Class A |
| `upload.uploadPart(n, body)` | One part; all non-final parts must be the same size and ≥ 5 MiB | Class A each |
| `upload.complete(parts)` | Finishes it | Class A |

[[embed:source:s3]]

Note the trap in that table: **`list` is a Class A operation**, priced 12.5× a read. A paginated file browser that lists on every page view burns the expensive quota, not the cheap one.

A single `put` accepts up to 5 GiB. Above that, multipart, or you get `100100 / EntityTooLarge`:

```js
const upload = await env.R2.createMultipartUpload('big/archive.tar');
const parts = [];
let n = 1;
for (const chunk of chunksOf(stream, 16 * 1024 * 1024)) {   // 16 MiB, uniform
  parts.push(await upload.uploadPart(n++, chunk));
}
await upload.complete(parts);
```

From outside a Worker, use the S3 API against `https://<ACCOUNT_ID>.r2.cloudflarestorage.com` with `region: "auto"`, or hand a browser a presigned URL so it uploads straight to R2 without the bytes passing through your server:

```js
import { AwsClient } from 'aws4fetch';

const r2 = new AwsClient({ accessKeyId: ACCESS_KEY_ID, secretAccessKey: SECRET_ACCESS_KEY });
const url = new URL(`https://${ACCOUNT_ID}.r2.cloudflarestorage.com/my-bucket/uploads/${name}`);
url.searchParams.set('X-Amz-Expires', '3600');           // seconds
const signed = await r2.sign(new Request(url, { method: 'PUT' }), {
  aws: { signQuery: true, service: 's3' },
});
// signed.url is safe to hand to a browser; it expires in one hour.
```

Tamper with any signature parameter and the request fails with `10035 / SignatureDoesNotMatch`; let it age out and you get `10018 / ExpiredRequest`.

[[embed:source:s9]]

## Serving an object publicly, and the header that decides your bill

Three ways to make a bucket readable from the internet. A **custom domain** puts it behind a hostname on your zone, the only option that gets Cloudflare Cache, WAF rules and Bot Management. An **r2.dev subdomain** is one toggle, documented as non-production. A **Worker route** gives you the object plus whatever logic you put in front of it. The third, in nine lines:

[[embed:source:s8]]

```js
// functions/img/[[path]].js: every /img/* URL is an R2 key.
export async function onRequestGet(context) {
  const { params, env } = context;
  const key = 'img/' + (Array.isArray(params.path) ? params.path.join('/') : String(params.path || ''));
  if (!env.R2) return new Response('no R2', { status: 500 });
  const obj = await env.R2.get(key);
  if (!obj) return new Response('not found', { status: 404 });
  const ct = obj.httpMetadata?.contentType || 'image/png';
  return new Response(obj.body, { headers: { 'content-type': ct, 'cache-control': 'public, max-age=31536000' } });
}
```

The URL path *is* the object key: no media table, no second name for a file, and a rename is a copy.

`cache-control: public, max-age=31536000` is the line that matters financially. Cached responses never reach R2, so they cost no Class B operation. Fetched live, the header comes back on the wire:

```text
$ curl -sI https://miscsubjects.com/img/up/cloudflare-os-r2-hero-card.png
HTTP/2 200
content-type: image/png
content-length: 51205
cache-control: public, max-age=31536000
cf-cache-status: MISS
cf-ray: a210b8bcbb715616-SJC
```

A one-year max-age is only safe because a changed image is written under a new key, never patched in place. Serve mutable objects this way and you will serve stale bytes for a year.

[[embed:source:s17]]

Cache-hit ratio is not a rounding error on the bill:

[[embed:source:s20]]

## When a row outgrows D1, the bytes move to R2 and the row keeps a pointer

D1, the SQLite database in the same account ([D1 as the spine](/a/cloudflare-os-d1)), has a hard per-value ceiling: **2,000,000 bytes** for any string, BLOB or row. Exceed it and every write to that row fails with `D1_ERROR: string or blob too big`, the driver's rendering of SQLite's `SQLITE_TOOBIG`.

This application hit that wall storing article revision history. Each write snapshots the previous version — full body, claims and sources — into a JSON blob on the row. One article's metadata reached **2,068,258 bytes** across 24 snapshots, 68,258 over the cap, and from then on every write to it returned HTTP 500.

The fix generalises into a rule: **when a field outgrows its row, the field moves to object storage and the row keeps a pointer plus a hash.** The pointer is small and fixed-size; the hash is what makes the pointer trustworthy.

The key format is one line in `functions/_lib/revisions_r2.js`:

```js
const PREFIX = "revisions/";
function r2Key(slug, n) { return `${PREFIX}${slug}/${n}.json`; }
```

So revision 3 of this article lives at `revisions/cloudflare-os-r2/3.json`. What stays in D1 is a nine-field index entry, written by the same function:

```js
return {
  n: full.n, ts: full.ts, title: full.title, status: full.status,
  register: full.register, bytes: full.body.length,
  prev_hash: full.prev_hash, hash: full.hash,
  r2_key: key,
};
```

`prev_hash` and `hash` are the point: the chain is verifiable from D1 alone, no R2 read needed to prove the history has not been edited, while the bytes are fetched only when someone asks for a specific revision. That 2,068,258-byte row became **206,362 bytes**, and the original revision is still retrievable.

[[embed:source:s16]]

Measured across the bucket today: 4,930 objects under `revisions/`, 148,075,146 bytes, 1,053 distinct articles, mean 30,036 bytes per revision. The largest single history, `revisions/bpc-157/`, is 119 objects and 14,503,688 bytes — 7.25× the D1 row cap, so that article alone would be permanently unwritable under the old scheme.

Proof the offload preserved the history rather than truncating it:

```bash
$ curl -s "https://miscsubjects.com/api/articles/cloudflare-os-r2?rev=0" | jq '{rev,is_head,prev_hash,body_len:(.body|length)}'
{ "rev": 0, "is_head": false, "prev_hash": "genesis", "body_len": 2878 }
```

That body came out of R2. Nothing but the index is in the database.

## Choosing between R2, S3, KV, D1 and Durable Object storage

| Workload | Put it in | Why, and the limit that decides it |
| --- | --- | --- |
| Images, video, uploads, anything served to the public internet | **R2** | Free egress; 5 TiB per object; strongly consistent |
| Cold archive, rarely read, consumed inside AWS | **S3 Glacier** | $0.0036–$0.00099/GB-month against R2's $0.01 floor |
| A file with a legal retention requirement and a need to read yesterday's copy | **S3** | R2 has bucket locks but no object versioning |
| Small values read constantly from many places — flags, prompts, rendered snapshots | **Workers KV** | Values to 25 MiB, but eventually consistent and 1 write/second per key |
| Anything you need to query, join, filter or index | **D1** | 10 GB per database, 2 MB per value; a bucket cannot answer `WHERE` |
| Per-entity state needing serialized writes and coordination | **Durable Object storage** | 10 GB per object, single-threaded execution ([the Workers layer](/a/cloudflare-os-workers)) |
| A row that has outgrown 2 MB | **R2, with a pointer in D1** | The offload pattern above |
| Blobs written and read once, under 25 MiB, with no need for a URL | **either KV or R2** | R2 unless you need sub-millisecond reads at the edge |

[[embed:source:s10]]

[[embed:source:s11]]

[[embed:source:s12]]

[[embed:source:s13]]

Two failure modes worth naming: KV used as a database (eventually consistent, so a read after a write may return the old value), and R2 used as a database (no query, only a lexicographic key scan at Class A prices).

## Symptom, cause, fix

| Symptom | Cause | Fix |
| --- | --- | --- |
| `no R2 binding` / `no R2`, HTTP 500 | The Worker deployed without the `r2_buckets` entry, or the bucket does not exist yet | `npx wrangler r2 bucket create <name>`, then confirm the `[[r2_buckets]]` block is present in every environment, including `[[env.preview.r2_buckets]]` |
| `10006 / NoSuchBucket`, HTTP 404 | Bucket name typo, or the bucket is in a different account | `npx wrangler r2 bucket list` and compare byte for byte |
| `get()` returns `null` instead of throwing | Not an error — the Workers API returns `null` for a missing key rather than raising `NoSuchKey` | Branch on `null`; do not wrap in `try/catch` and expect a throw |
| `100100 / EntityTooLarge`, HTTP 400 | Single-part upload over 5 GiB | Switch to `createMultipartUpload`; parts uniform and ≥ 5 MiB |
| `10011 / EntityTooSmall` or `10048 / InvalidPart`, HTTP 400 | A non-final part under 5 MiB, or parts of differing sizes | Every part except the last must be ≥ 5 MiB and identical in size |
| `10058 / TooManyRequests`, HTTP 429 | More than one write per second to the same key | Shard the key, or funnel writes for that key through a Durable Object |
| `10035 / SignatureDoesNotMatch`, HTTP 403 | A presigned URL was edited, or the secret is wrong | Regenerate; check URL encoding of the key |
| `10069 / ObjectLockedByBucketPolicy`, HTTP 403 | A bucket lock retention rule covers that prefix | Wait out the retention period; a lock is not overridable |
| `D1_ERROR: string or blob too big` on a write that used to work | A JSON column crossed D1's 2,000,000-byte per-value cap | Offload the heavy field to R2, keep a pointer and a hash in the row |
| `wrangler r2 object put --file` fails with `fetch failed` | Reported against Wrangler 3.74.0 and still open | Pipe the file instead: `cat f.txt \| npx wrangler r2 object put bucket/f.txt --pipe` |
| Multipart `complete()` returns empty `customMetadata` in production but not in `wrangler dev` | Open bug, reported against Wrangler 3.57.2 | Re-read with `head()` after completing; the metadata is stored, only the return value drops it |

## How the numbers on this page were measured

Four measurements, taken 2026-07-26 against the live production bucket from a laptop served by the San Jose edge (`cf-ray` suffix `SJC`).

[[embed:source:s27]]

**1 — Buckets and inventory.** `npx wrangler r2 bucket list` returns three buckets: `loop-data-raw` (2026-05-31), `miscsubjects-ledger` (2026-06-09), `miscsubjects-store` (2026-06-16). The inventory of `miscsubjects-ledger` was walked through the authenticated list route, 1,000 keys per page, following the cursor:

```bash
curl -s -H "x-terminal-key: $TERMINAL_KEY" \
  "https://miscsubjects.com/api/r2/?list=1&limit=1000&cursor=$CURSOR"
```

62 pages, **41,809 objects, 3,904,595,147 bytes (3.64 GiB)** — under the 10 GB free allowance, so today's storage line is $0.00. By prefix: `events/` 34,258 objects / 2,241 MB; `revisions/` 4,930 / 148 MB; `img/gen/` 387 / 765 MB; `img/up/` 191 / 432 MB; `docs/` 1,332 / 9.5 MB. The walk itself cost 62 Class A operations.

**2 — Put and get latency.** A 65,536-byte JSON object written and read six times each over one reused TLS connection, at the deliberately-named temporary key `tmp/measure-2026-07-25-delete-me.json`, then deleted — `DELETE` returned `{"ok":true,...,"deleted":true}` and a follow-up list of `tmp/` returned zero objects.

```bash
curl -s -X PUT "https://miscsubjects.com/api/r2/tmp/measure-2026-07-25-delete-me.json" \
  -H "x-terminal-key: $TERMINAL_KEY" -H 'content-type: application/json' \
  --data-binary @probe.json -w "%{time_total}\n"
```

PUT round trips: 535, 587, 634, 737, 801, 2,471 ms — median **686 ms**. GET: 96, 104, 125, 148, 185, 192 ms — median **136 ms**. These are full client-to-edge-to-R2-to-client times through the Worker, not R2's internal service time; the first request in each series carries the TLS handshake, 86 ms and 71 ms.

**3 — Public object headers.** `curl -sI https://miscsubjects.com/img/up/cloudflare-os-r2-hero-card.png` returned HTTP 200, `content-length: 51205`, `content-type: image/png`, `cache-control: public, max-age=31536000`, `cf-cache-status: MISS`. Full body fetch, 189 ms.

[[embed:source:s18]]

**4 — A revision read out of R2.** `curl -s "https://miscsubjects.com/api/articles/cloudflare-os-r2?rev=0"` returned `rev: 0`, `is_head: false`, `prev_hash: "genesis"`, `hash: 24915ae889f6184140a56e89d476d28a31e0cf0e28d31c0990c2aab6f7a626fc`, body length 2,878 — out of `revisions/cloudflare-os-r2/0.json`, not the database. Account identifiers are omitted throughout; the S3 endpoint appears as `<ACCOUNT_ID>.r2.cloudflarestorage.com`.

The index for the rest of this stack is [the Cloudflare account as an operating system](/a/cloudflare-os).


## Sources

1. Cloudflare R2 pricing — https://developers.cloudflare.com/r2/pricing/
2. How R2 works — https://developers.cloudflare.com/r2/how-r2-works/
3. Workers R2 API reference — https://developers.cloudflare.com/r2/api/workers/workers-api-reference/
4. S3 API compatibility — https://developers.cloudflare.com/r2/api/s3/api/
5. R2 storage classes — https://developers.cloudflare.com/r2/buckets/storage-classes/
6. R2 object lifecycles — https://developers.cloudflare.com/r2/buckets/object-lifecycles/
7. R2 bucket locks — https://developers.cloudflare.com/r2/buckets/bucket-locks/
8. R2 public buckets — https://developers.cloudflare.com/r2/buckets/public-buckets/
9. R2 presigned URLs — https://developers.cloudflare.com/r2/api/s3/presigned-urls/
10. R2 limits — https://developers.cloudflare.com/r2/platform/limits/
11. Workers KV limits — https://developers.cloudflare.com/kv/platform/limits/
12. D1 limits — https://developers.cloudflare.com/d1/platform/limits/
13. Durable Objects limits — https://developers.cloudflare.com/durable-objects/platform/limits/
14. Cloudflare Developer Platform terms — https://www.cloudflare.com/service-specific-terms-developer-platform/
15. AWS S3 us-east-1 Price List API — https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonS3/current/us-east-1/index.json
16. Clarify presigned URL docs — https://github.com/cloudflare/cloudflare-docs/issues/19190
17. Unable to put an object with a key containing three dots to R2 — https://github.com/cloudflare/workers-sdk/issues/3520
18. Cloud object storage benchmark — https://aimultiple.com/cloud-object-storage
19. The next chapter for Cloudflare Workers: open-source — https://news.ycombinator.com/item?id=31314273
20. How Canva saves Amazon S3 costs — https://news.ycombinator.com/item?id=36392670
21. WordPress Plugin Mirror Downloader (Proof of Concept) — https://news.ycombinator.com/item?id=41753985
22. Ask HN: S3(AWS) vs R2(CF)–Which is better? — https://news.ycombinator.com/item?id=47700966
23. Comparing AWS S3 with Cloudflare R2: Price, Performance and User Experience — https://news.ycombinator.com/item?id=42257068
24. Hetzner continues its growth in the US with a new location — https://news.ycombinator.com/item?id=33865229
25. Comparing AWS S3 with Cloudflare R2: Price, Performance and User Experience — https://news.ycombinator.com/item?id=42257094
26. Cloudflare R2 let me serve almost twice as much data this month as the SEC for $10.80 — https://old.reddit.com/r/CloudFlare/comments/1qhbrey/cloudflare_r2_let_me_serve_almost_twice_as_much/
27. Live R2 object headers — https://miscsubjects.com/img/up/cloudflare-os-r2-hero-card.png

