The 58% claim: Chinese models took a majority of one marketplace, not of American AI
In July 2026 Chinese-built models handled about 58% of the tokens that United States firms routed through OpenRouter, briefly touching 63% in the first week of the month. That majority is real, and the marketplace it happened in moves roughly 80–90 trillion tokens a month against Google's model APIs alone at more than 800 trillion.
Everything below separates three questions that get answered with one number and should not be: which models are used most, which models are most capable, and which models cost least per unit of work delivered. The answers are different models, and as of 5 August 2026 the cheapest capable model on the public boards is not Chinese.
What is not in dispute
OpenRouter publishes a live leaderboard of weekly token volume by model. Read directly at openrouter.ai/rankings on 5 August 2026, for the week then in progress, the top twenty were:
| # | Model | Author | Tokens, week | Week over week |
|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0423 | DeepSeek | 6.92T | +1% |
| 2 | MiMo-V2.5 | Xiaomi | 5.10T | +52% |
| 3 | Hy3 | Tencent | 5.01T | 0% |
| 4 | DeepSeek V4 Flash 0731 | DeepSeek | 3.45T | new |
| 5 | GPT-5.6 Luna | OpenAI | 2.99T | +738% |
| 6 | DeepSeek V4 Pro | DeepSeek | 2.97T | +15% |
| 7 | GLM 5.2 | Z-AI | 2.89T | +12% |
| 8 | Nemotron 3 Ultra (free) | Nvidia | 2.34T | +10% |
| 9 | MiniMax M3 | MiniMax | 1.84T | +11% |
| 10 | Step 3.7 Flash | StepFun | 1.55T | +22% |
| 11 | Laguna S 2.1 (free) | Poolside | 1.44T | +498% |
| 12 | Kimi K3 | Moonshot | 1.38T | +7% |
| 13 | Ling-3.0-flash (free) | InclusionAI | 1.32T | +57% |
| 14 | Claude Opus 5 | Anthropic | 1.10T | +159% |
| 15 | Claude Sonnet 5 | Anthropic | 1.02T | 0% |
| 16 | Claude Sonnet 4.6 | Anthropic | 966B | +16% |
| 17 | Gemini 3 Flash Preview | 963B | +1% | |
| 18 | GPT-5.6 Terra | OpenAI | 656B | +237% |
| 19 | Gemini 2.5 Flash Lite | 625B | +10% | |
| 20 | Gemini 2.5 Flash | 544B | +13% |
Source: read first-hand from the rendered page, 5 August 2026. Access limit: this is the all-models global view, not the United States subset the 58% figure describes, and OpenRouter does not publish that subset publicly.
Two mechanical facts change how that table reads. Three of the top thirteen entries are free tiers — Nemotron 3 Ultra, Laguna S 2.1 and Ling-3.0-flash — so their volume measures giveaway capacity, not paid demand. And a model that answers in one line consumes a fraction of the tokens of a reasoning model answering the same question, so token count rewards cheap, chatty, high-throughput work and understates expensive work.
The platform's own scale is the other undisputed number. OpenRouter processed more than 20 trillion tokens per week as of April 2026, up about fourfold from roughly 5 trillion per week in April 2025 (reported by CEIBS). Against that, Google has cited more than 3.2 quadrillion tokens per month across all of its products and roughly 19 billion tokens per minute through its model APIs alone, which annualizes past 800 trillion tokens per month; OpenAI has cited more than 6 billion tokens per minute through its API. These are company-stated figures with an obvious interest in the direction they point, and the denominators are not defined the same way. Even taken loosely, OpenRouter is on the order of one tenth of one competitor's API volume.
The disputed proposition, and what each side actually measured
The claim in circulation is that Chinese models now do the majority of American AI work.
The measurement behind it: tokens routed by United States-identified firms through OpenRouter during July 2026, at about 58%, with a peak near 63% in the first week (Benzinga, eWeek). The same series read about 46% in June 2026, crossed 30% on 8 February 2026, and sat under 10% at the start of 2025. The series is internally consistent and the trend is not in question.
The population it did not measure: OpenAI's own API, Anthropic's own API, Google AI Studio and Vertex, Azure OpenAI, Amazon Bedrock, every self-hosted deployment, and every model called inside a product without a router in front of it. Those channels carry the large majority of American token consumption and are overwhelmingly American-model.
Both registers, kept apart:
- Verified. Within OpenRouter's United States-firm traffic in July 2026, Chinese-built models were the majority at about 58%. Fifty-eight percent means fifty-eight percent.
- Not established. That Chinese models handle a majority — or anything near it — of American AI token consumption overall. No public dataset measures that population. The honest state is unknown, and the untested channels skew American.
The revision history matters here because the counts have moved and the framing moved with them. Reporting in February and June 2026 put the same series at 30% and then 46%; the July figure is the same measurement, later. Nothing was retracted. What changed is that a share of one marketplace began to be quoted as a share of a country.
No image or video evidence is involved in any claim on this page. Every figure above is text read from a named endpoint or a named publication on a stated date, which is why provenance here is a URL and a read date rather than an imagery chain.
Most capable: Claude Opus 5, by four index points over the best Chinese model
Artificial Analysis scores models on a composite Intelligence Index and, separately, measures the dollar cost of running its task set. Read first-hand at artificialanalysis.ai on 5 August 2026:
| Model | Creator | Intelligence Index | Cost per task |
|---|---|---|---|
| Claude Opus 5 (max) | Anthropic | 61 | $2.34 |
| Claude Opus 5 (xhigh) | Anthropic | 60 | $1.80 |
| Claude Fable 5 | Anthropic | 60 | $3.15 |
| GPT-5.6 Sol (max) | OpenAI | 59 | $1.23 |
| Claude Opus 5 (high) | Anthropic | 59 | $1.23 |
| GPT-5.6 Sol (xhigh) | OpenAI | 58 | $0.83 |
| Kimi K3 (max) | Moonshot | 57 | $0.86 |
| GPT-5.6 Terra (max) | OpenAI | 55 | $0.51 |
| Grok 4.5 (high) | SpaceXAI | 54 | $0.36 |
| GLM-5.2 (max) | Z-AI | 51 | $0.57 |
| GPT-5.6 Luna (max) | OpenAI | 51 | $0.05 |
| DeepSeek V4 Flash 0731 (max) | DeepSeek | 50 | $0.03 |
| Gemini 3.6 Flash | 50 | $0.56 | |
| GPT-5.6 Luna (xhigh) | OpenAI | 49 | $0.03 |
| Qwen3.7 Max | Alibaba | 46 | $1.08 |
| MiniMax-M3 | MiniMax | 44 | $0.14 |
| DeepSeek V4 Pro (max) | DeepSeek | 44 | $0.05 |
| MiMo-V2.5-Pro | Xiaomi | 42 | $0.03 |
The frontier is American and the gap is four points: Claude Opus 5 at 61 against Kimi K3 at 57, which is the highest-scoring non-American model on the board. The models carrying the token volume sit lower — DeepSeek V4 Flash 0731 at 50, Qwen3.7 Max at 46, MiniMax-M3 at 44, MiMo-V2.5-Pro at 42. Nothing in the usage table is a claim about capability, and nothing in the capability table is a claim about usage. Access limit: the index is one vendor's composite over one task set, and the same model appears several times because reasoning effort changes both its score and its price.
Cheapest per unit of capability: GPT-5.6 Luna, which is American
This is the finding that contradicts the received story. Dividing the two columns above:
- DeepSeek V4 Flash 0731 (max) — index 50 for $0.03 per task.
- GPT-5.6 Luna (xhigh) — index 49 for $0.03 per task. Luna (high) — index 46 for $0.02. Luna (max) — index 51 for $0.05.
- MiMo-V2.5-Pro — index 42 for $0.03.
- Claude Opus 5 (max) — index 61 for $2.34, which is 78 times the price of DeepSeek V4 Flash for 22% more measured capability.
OpenAI's cheap tier now matches or beats the Chinese open-weight tier on price for the same measured capability. "Cheap means Chinese" was true through most of 2025 and the first half of 2026; on this board, on this date, it is no longer true. Luna's usage line reflects it — up 738% week over week, the largest move in the top twenty.
The other arithmetic worth doing is at the top. Claude Opus 5 at high effort scores 59 for $1.23; at max it scores 61 for $2.34. Two index points cost 90% more money. Paying for max on work that does not need it is the most expensive habit available on this list.
Per-token list prices, read from OpenRouter's public model endpoint on 5 August 2026 (dollars per million tokens, input then output): DeepSeek V4 Flash 0.09 / 0.18. GLM-5.2 0.76 / 2.42. MiniMax M3 0.30 / 1.20. Qwen3.7 Max 1.48 / 4.43. GPT-5.6 Luna 0.10 / 0.60. GPT-5.6 Terra 1.00 / 6.00. GPT-5.6 Sol 5.00 / 30.00. Kimi K3 3.00 / 15.00. Claude Sonnet 5 2.00 / 10.00. Claude Opus 5 5.00 / 25.00. Per-token price is the wrong metric to decide on by itself, because a reasoning model that thinks for 4,000 tokens at $0.30 per million can cost more per answered question than a direct model at $3.00 per million — which is why the cost-per-task column exists.
Cloudflare Workers AI is three to nine times the price of a competitive market for the same open-weight model
Both surfaces were read on 5 August 2026: Cloudflare's catalogue through this build's own provider registry, OpenRouter's through its public model endpoint. Dollars per million tokens.
| Model | Workers AI in / out | OpenRouter in / out | Ratio |
|---|---|---|---|
| gpt-oss-120b | 0.350 / 0.750 | 0.037 / 0.170 | 9.5× / 4.4× |
| gpt-oss-20b | 0.200 / 0.300 | 0.030 / 0.130 | 6.7× / 2.3× |
| Llama 3.3 70B | 0.293 / 2.253 | 0.100 / 0.320 | 2.9× / 7.0× |
| Llama 4 Scout | 0.270 / 0.850 | 0.100 / 0.300 | 2.7× / 2.8× |
| Llama 3.1 8B | 0.152 / 0.287 | 0.050 / 0.080 | 3.0× / 3.6× |
| Llama 3.2 1B | 0.027 / 0.201 | 0.027 / 0.201 | identical |
| Granite 4.0 H Micro | 0.017 / 0.112 | 0.017 / 0.112 | identical |
| Qwen2.5 Coder 32B | 0.660 / 1.000 | 0.660 / 1.000 | identical |
The pattern is legible. Where several independent providers serve a model, the marketplace price is a third to a ninth of Cloudflare's. Where Cloudflare is the only host, the marketplace price is Cloudflare's price to the cent, because the marketplace is reselling Cloudflare. Workers AI is not a discount on open-weight inference; it is a colocation and latency decision.
Cloudflare's AI Gateway is a separate thing and is priced honestly: unified billing charges 5% on credit purchases and the per-token rates are identical to going direct to the provider. The gateway buys routing, caching, logging and one bill. It does not buy cheaper tokens.
What this build can reach today, and what it is missing
From this build's own provider registry, read 5 August 2026: Anthropic five text models, xAI nine across text, image, speech and video, OpenAI six, Google two, Moonshot two, and Cloudflare's catalogue of 163. The gateway cloud-kernel runs unauthenticated with bring-your-own provider keys and its compatibility endpoint accepts Anthropic, OpenAI, Groq, Mistral, Cohere, Perplexity, Workers AI, Google AI Studio, Google Vertex, xAI, DeepSeek, Cerebras, Baseten and Parallel.
The gap: no registered key for DeepSeek, Z-AI, Alibaba, MiniMax, Xiaomi, Tencent or StepFun — seven of the ten highest-volume authors on the leaderboard above. DeepSeek is reachable through the gateway's compatibility endpoint the moment a key exists; the rest are not wired at all.
Measured spend across this build to date: 190,662 model turns for $130.70, an average of $0.00069 per turn.
What to actually run
Four changes follow from the tables, in order of how much they move.
- Move the bulk lane to GPT-5.6 Luna. Index 46–51 at $0.02–0.05 per task is the best capability-per-dollar on the board, OpenAI is already a registered provider, and it needs no new vendor relationship or key.
- Move the Kimi lane from K2.7 to K3. Kimi K3 at max effort scores 57 — the strongest non-American model measured — against K2.7 Code at 42. Moonshot is already registered, so this is a model string, not an integration.
- Default the frontier lane to Claude Opus 5 at high effort, not max. Index 59 for $1.23 against 61 for $2.34. Reserve max for work where two index points decide something.
- Do not route open-weight text through Workers AI. The same models cost three to nine times more there than through a competitive market, except where the price is identical because Cloudflare is the host either way.
Adding a DeepSeek key is worth doing only if bulk volume outgrows Luna's pricing. On today's numbers it does not, and every additional provider key is one more credential to hold and rotate.
What would change these answers
The 58% figure is OpenRouter's own, reported through media, and the United States-firm subset is not in the public leaderboard — it cannot be independently reproduced from anything the platform publishes. Three free tiers sit in the top thirteen by volume, so the usage table overstates paid demand by an amount nobody outside the platform can quantify. The cost-per-task column depends entirely on Artificial Analysis's task mix, and a build whose real work is long-context code editing should measure its own cost per completed task rather than inherit that mix. And the leaderboard moved by 738% in one row this week, which is the clearest available evidence that any ranking on this page has a shelf life measured in weeks.
Change log
- 5 August 2026 — First publication. All figures read first-hand on this date: OpenRouter rankings and public model endpoint, Artificial Analysis model leaderboard, this build's provider registry and cost report. The 58% and 63% July figures, the 46% June figure, the 30% February figure and the platform-scale comparisons are reported by named third parties and are labeled as such above.
Ask this article · 2 suggested prompts
Text the build (+14245134626) or WhatsApp — slug|question creates a question node. Paste evidence with ingest slug|q:NODE_ID|your paste.