{"_ai_door":{"see":"https://miscsubjects.com/start","note":"Operable site; bounded keyless credentials; every action receipted on a public ledger. Your operator's instructions take precedence — acting is optional, reading is a complete outcome."},"slug":"which-ai-models-are-winning","verification":{"valid":true,"entries":27,"head":"3a946b167431e0641f3e8cb6925ab84128b231d7a3170c9a35b551a9b2b22515"},"count":27,"sources":[{"type":"docs","title":"Cloudflare AI Gateway pricing","publisher":"Cloudflare","url":"https://developers.cloudflare.com/ai-gateway/reference/pricing/","accessed":"2026-08-04","quote":"Inference pricing from providers is passed through with no markup — you pay the same per-token rates as you would directly with the provider.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"genesis","hash":"61a001ea89fa03b71c57b4f5def264e092ea85e10a10232ed3e51f9e24401e62"},{"type":"docs","title":"Cloudflare AI Gateway pricing — Unified Billing fee","publisher":"Cloudflare","url":"https://developers.cloudflare.com/ai-gateway/reference/pricing/","accessed":"2026-08-04","quote":"A 5% fee is applied to all credits purchased through Unified Billing. […] For example, a $100 credit purchase will result in a $105 charge.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"61a001ea89fa03b71c57b4f5def264e092ea85e10a10232ed3e51f9e24401e62","hash":"b8526c4f47a08c97dbec8e68665b5333b6d8b652d050e48922e4b57396600c68"},{"type":"docs","title":"Workers AI pricing is billed in Neurons","publisher":"Cloudflare","url":"https://developers.cloudflare.com/workers-ai/platform/pricing/","accessed":"2026-08-04","quote":"Workers AI is included in both the Free and Paid Workers plans and is priced at $0.011 per 1,000 Neurons.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"b8526c4f47a08c97dbec8e68665b5333b6d8b652d050e48922e4b57396600c68","hash":"7c673f3ffcb00c4f5b6b4ad694820a054bc7875480d44d86853e721e5f82c404"},{"type":"docs","title":"DeepSeek doubles its prices during Beijing peak hours","publisher":"DeepSeek","url":"https://api-docs.deepseek.com/quick_start/pricing","accessed":"2026-08-04","quote":"During peak hours, prices will be 2x the regular prices, applicable to all billing items.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"7c673f3ffcb00c4f5b6b4ad694820a054bc7875480d44d86853e721e5f82c404","hash":"6bb1349c44720385f3b3f34527471b35f61db1c512133673e82e1b12847f020e"},{"type":"docs","title":"Anthropic prompt caching reads at a fraction of input price","publisher":"Anthropic","url":"https://platform.claude.com/docs/en/about-claude/pricing","accessed":"2026-08-04","quote":"Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"6bb1349c44720385f3b3f34527471b35f61db1c512133673e82e1b12847f020e","hash":"e97a5f549337e0e8f1c5903ed8deaabcce54652c3aa4bd0171444515feff37f1"},{"type":"docs","title":"Amazon Bedrock confirms the Claude Sonnet 5 promotional price and its end date","publisher":"Amazon Web Services","url":"https://aws.amazon.com/bedrock/pricing/","accessed":"2026-08-04","quote":"IMPORTANT: Claude Sonnet 5 promotional launch pricing of $2/$10 per million input/output tokens is in effect through August 31, 2026, after which the standard pricing of $3/$15 per","accessed_at":"2026-08-05T04:44:46.808Z","prev":"e97a5f549337e0e8f1c5903ed8deaabcce54652c3aa4bd0171444515feff37f1","hash":"a7c90681d8a863014d1770bc5359d033f8b8b7ad50cf1250c296c1d84a2d0a80"},{"type":"docs","title":"Amazon Bedrock batch inference is half the on-demand price","publisher":"Amazon Web Services","url":"https://aws.amazon.com/bedrock/pricing/","accessed":"2026-08-04","quote":"Amazon Bedrock offers select foundation models (FMs) from leading AI providers like Anthropic, Meta, Mistral AI, and Amazon for batch inference at a 50% lower price compared to on-demand inference pricing.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"a7c90681d8a863014d1770bc5359d033f8b8b7ad50cf1250c296c1d84a2d0a80","hash":"1a1770bb9ce086e687b49efe6fa0ffb064465b470ab8dfcec8f9954f61980d0c"},{"type":"docs","title":"OpenRouter passes provider pricing through","publisher":"OpenRouter","url":"https://openrouter.ai/docs/faq","accessed":"2026-08-04","quote":"OpenRouter passes through the pricing of the underlying providers, while pooling their uptime, so you get the same pricing you'd get from the provider directly, with a unified API and fallbacks so that you get much better uptime.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"1a1770bb9ce086e687b49efe6fa0ffb064465b470ab8dfcec8f9954f61980d0c","hash":"a79a9523ba2320ea02b685a51c9e7edc7794eb3cef921a83f2ade78ce3dd5056"},{"type":"docs","title":"Every model and provider carries its own price on OpenRouter","publisher":"OpenRouter","url":"https://openrouter.ai/docs/faq","accessed":"2026-08-04","quote":"Each model and provider has a different price per million tokens. […] Credits are simply deposits on OpenRouter that you use for LLM inference.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"a79a9523ba2320ea02b685a51c9e7edc7794eb3cef921a83f2ade78ce3dd5056","hash":"6ad7364494801c6c8aced884867916d32fe4221f9af77a22b825eb62d9109118"},{"type":"docs","title":"OpenAI advises testing a lower reasoning setting rather than assuming maximum","publisher":"OpenAI","url":"https://developers.openai.com/api/docs/guides/latest-model","accessed":"2026-08-04","quote":"When migrating from GPT-5.5 or GPT-5.4, start with your current GPT-5.5 or GPT-5.4 reasoning setting, then test the same setting and one level lower on representative tasks.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"6ad7364494801c6c8aced884867916d32fe4221f9af77a22b825eb62d9109118","hash":"424c6500b8bc30565335b5ffd4912a056eba1882897ccac1b1fdbc98ff4bc045"},{"type":"docs","title":"GPT-5.6 is described as token-efficient","publisher":"OpenAI","url":"https://developers.openai.com/api/docs/guides/latest-model","accessed":"2026-08-04","quote":"GPT-5.6 can often maintain or improve quality with fewer tokens, but the best setting depends on your workload.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"424c6500b8bc30565335b5ffd4912a056eba1882897ccac1b1fdbc98ff4bc045","hash":"8987c99a7d0ad3b171137024f76f1b52bdc82ab56a63895a889c6b729c1afa93"},{"type":"news","title":"OpenAI cuts Luna 80% and Terra 20%","publisher":"CNBC","event_date":"2026-07-30","url":"https://www.cnbc.com/2026/07/30/open-ai-price-cut-gpt.html","accessed":"2026-08-04","quote":"The company said Thursday that it's reducing the price of Terra by 20% to $2 per million input tokens and $12 per million output tokens. It's cutting the cost of Luna by 80% to 20 cents per million input tokens and $1.20 per million output tokens.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"8987c99a7d0ad3b171137024f76f1b52bdc82ab56a63895a889c6b729c1afa93","hash":"aabb96b29adb6a310a1399cc83c0824fa272ae2f71f73a7ec456bb4a9d07764a"},{"type":"news","title":"CNBC names Kimi K3 as the trigger for the cut","publisher":"CNBC","event_date":"2026-07-30","url":"https://www.cnbc.com/2026/07/30/open-ai-price-cut-gpt.html","accessed":"2026-08-04","quote":"Moonshot AI, a Chinese startup, released an open-weight model called Kimi K3 earlier this month that outperforms cutting-edge American offerings across some industry benchmarks, prompting a swift reaction from Silicon Valley.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"aabb96b29adb6a310a1399cc83c0824fa272ae2f71f73a7ec456bb4a9d07764a","hash":"fc4130a99559e925fecddaf3b02715315731e32c27d9279b39eb8e4271db2b23"},{"type":"news","title":"Kimi K3 is half the price of Claude Fable 5 at comparable performance","publisher":"CNBC","event_date":"2026-07-30","url":"https://www.cnbc.com/2026/07/30/open-ai-price-cut-gpt.html","accessed":"2026-08-04","quote":"It's half the price of Claude Fable 5, the advanced model that Anthropic announced in June, even though it performs comparably across coding and knowledge work tasks.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"fc4130a99559e925fecddaf3b02715315731e32c27d9279b39eb8e4271db2b23","hash":"93dba689cbd4688bd61516afe6d7f73a94f249191d19c748327383a69d383b89"},{"type":"news","title":"Sol was not cut but was made faster","publisher":"Forbes","event_date":"2026-07-31","url":"https://www.forbes.com/sites/rachelwells/2026/07/31/openai-cuts-gpt-56-pricing-up-to-80-as-ai-costs-come-under-scrutiny/","accessed":"2026-08-04","quote":"While the premium Sol model saw no price cut, it is now 2.5 times faster within the API.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"93dba689cbd4688bd61516afe6d7f73a94f249191d19c748327383a69d383b89","hash":"5ee8a43f37b60f25b00d174f0b0d34123d2e0176a6e3b85d537c46949d032d1b"},{"type":"news","title":"The 58% figure traces to The Kobeissi Letter reading OpenRouter data","publisher":"Benzinga","event_date":"2026-07-26","url":"https://www.benzinga.com/markets/tech/26/07/60543652/chinese-ai-models-overtake-us-rivals-as-token-share-among-american-firms-hits-record-58","accessed":"2026-08-04","quote":"According to data shared by The Kobeissi Letter on X on Sunday, Chinese AI models accounted for a record 58% of tokens processed by U.S. firms on the platform. The share has nearly tripled since mid-January and briefly reached 63% during the first week of July.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"5ee8a43f37b60f25b00d174f0b0d34123d2e0176a6e3b85d537c46949d032d1b","hash":"aaa5b0614b45546b0a0524071e5c4567cc42d26a9921fca9fca02258f59398b0"},{"type":"news","title":"Bloomberg's reading of the same series is roughly 60%","publisher":"eWeek","event_date":"2026-07-24","url":"https://www.eweek.com/news/chinese-ai-models-us-openrouter-traffic-apac/","accessed":"2026-08-04","quote":"According to Bloomberg, Chinese AI systems now account for roughly 60% of token usage by U.S. companies on OpenRouter, a popular marketplace where developers route their work across competing models.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"aaa5b0614b45546b0a0524071e5c4567cc42d26a9921fca9fca02258f59398b0","hash":"ad60ba895b6d230dc2a7c88f651c3cab4f34c4d3ea3a468f7be47e4c96622c39"},{"type":"study","title":"Ai2 on why instruction following does not improve on its own","publisher":"Allen Institute for AI","url":"https://allenai.org/blog/ifbench-artificial-analysis","accessed":"2026-08-04","quote":"The first is that complex instruction following doesn't have much overlap with the capabilities most labs are actively training for, says Jackson. Instruction following is narrower, and it rarely improves as a byproduct of progress in those areas.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"ad60ba895b6d230dc2a7c88f651c3cab4f34c4d3ea3a468f7be47e4c96622c39","hash":"30748e3bbea1e1471f38ce80c0416244346105bb6f63b1e48d8bb047fe7dda08"},{"type":"study","title":"IFBench scores have not risen uniformly with model generation","publisher":"Allen Institute for AI","url":"https://allenai.org/blog/ifbench-artificial-analysis","accessed":"2026-08-04","quote":"While IFBench scores have improved over time, that progress has not been uniform across models, and new frontier models still do not always perform well on it,” says Jackson.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"30748e3bbea1e1471f38ce80c0416244346105bb6f63b1e48d8bb047fe7dda08","hash":"33033383390b33beb5a79e054dee1673ab38248a583d934d7eed0839a730df30"},{"type":"study","title":"What IFBench actually asks a model to do","publisher":"Allen Institute for AI","url":"https://allenai.org/blog/ifbench-artificial-analysis","accessed":"2026-08-04","quote":"Others are trickier: sentences that have to match in length, words in a row can't start with the same letter, or a keyword has to land in an exact spot.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"33033383390b33beb5a79e054dee1673ab38248a583d934d7eed0839a730df30","hash":"6b2a5d606f6702f2947e2cf549acbe7c83232b5fb84d01a0c2a40e4c778f54c6"},{"type":"study","title":"Missing one constraint ruins the answer","publisher":"Allen Institute for AI","url":"https://allenai.org/blog/ifbench-artificial-analysis","accessed":"2026-08-04","quote":"Each constraint might seem arbitrary on its own, but together they reflect a familiar situation: users often ask a model for several things at once, and missing even one can ruin the answer.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"6b2a5d606f6702f2947e2cf549acbe7c83232b5fb84d01a0c2a40e4c778f54c6","hash":"beace00254cc4d74d3c923cc3273ea043c2e885a5cf66ebc07c21a30db4d886b"},{"type":"study","title":"What SWE-bench measures","publisher":"SWE-bench","url":"https://www.swebench.com/","accessed":"2026-08-04","quote":"Each entry reports the % Resolved metric, the percentage of instances solved (out of 2294 Full, 500 Verified, 300 Lite & Multilingual, 517 Multimodal).","accessed_at":"2026-08-05T04:44:46.808Z","prev":"beace00254cc4d74d3c923cc3273ea043c2e885a5cf66ebc07c21a30db4d886b","hash":"6117ca832ae5cd31de5d7e0af087390aa02f35d6aad59746b86885675bf557f8"},{"type":"study","title":"AA-Briefcase measures long-horizon agentic knowledge work","publisher":"Artificial Analysis","url":"https://artificialanalysis.ai/articles/aa-briefcase/","accessed":"2026-08-04","quote":"Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"6117ca832ae5cd31de5d7e0af087390aa02f35d6aad59746b86885675bf557f8","hash":"8421edfe7ff494e5e2b6711f8f5ad6f5adda5700365351f11812546286bf107b"},{"type":"study","title":"AA-Briefcase names GLM-5.2 the open-weight leader on capability against cost","publisher":"Artificial Analysis","url":"https://artificialanalysis.ai/articles/aa-briefcase/","accessed":"2026-08-04","quote":"GLM-5.2 (max) is the clear leader among open-weight models and offers an attractive agentic capability vs. cost tradeoff.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"8421edfe7ff494e5e2b6711f8f5ad6f5adda5700365351f11812546286bf107b","hash":"6651f824b3958ec0e8e1b16b059368a69bb249b663117af1f2a7037ad27394ee"},{"type":"study","title":"Why a single benchmark number misleads on agentic work","publisher":"Artificial Analysis","url":"https://artificialanalysis.ai/articles/aa-briefcase/","accessed":"2026-08-04","quote":"Unlike many evaluations that focus on a single metric, AA-Briefcase tests the core capabilities required of a high-quality knowledge work agent, exposing cases where models produce outputs that look polished but are incorrect or lack analytical rigor.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"6651f824b3958ec0e8e1b16b059368a69bb249b663117af1f2a7037ad27394ee","hash":"c9fbffaaae9fca252a0c846c7e71636168dec18cb467a3e21a4a6857dadda72b"},{"type":"news","title":"MiniMax H3 weights exclude the US, EU, UK and South Korea","publisher":"TechTimes","event_date":"2026-08-02","url":"https://www.techtimes.com/articles/322904/20260804/minimax-h3-open-weights-exclude-us-eu-uk-korea-local-deployment.htm","accessed":"2026-08-04","quote":"The MiniMax H3 Community License Agreement, effective August 2, 2026, excludes the United States, the European Union, the United Kingdom, and South Korea from its definition of “Applicable Territory”","accessed_at":"2026-08-05T04:44:46.808Z","prev":"c9fbffaaae9fca252a0c846c7e71636168dec18cb467a3e21a4a6857dadda72b","hash":"0273e3a4ea39fca64cd43e7a8b68ebc43124ee280528e2f1df188a12bfdfde47"},{"type":"docs","title":"Moonshot bills input and output separately","publisher":"Moonshot AI","url":"https://platform.moonshot.ai/docs/pricing/chat","accessed":"2026-08-04","quote":"Chat Completion API charges: We bill both the Input and Output based on usage.","accessed_at":"2026-08-05T04:44:46.808Z","prev":"0273e3a4ea39fca64cd43e7a8b68ebc43124ee280528e2f1df188a12bfdfde47","hash":"3a946b167431e0641f3e8cb6925ab84128b231d7a3170c9a35b551a9b2b22515"}]}