The living model index
Every figure below is one observation: what was measured, of which model, at which venue, by whom, on what date, with the URL it came from and the class of evidence it is. Nothing is stored as a bare number, so nothing here has to be taken on trust. A changed price becomes a new observation and the previous one is kept, marked superseded.
Capability
What the model can do, graded by running it.
aa_intelligence_index
| Model | Value | Venue | Evidence | Source · read |
|---|---|---|---|---|
| Claude Opus 5 (max)Anthropic | 61 | — | measured | Artificial Analysis2026-08-04 |
| Claude Opus 5 (xhigh)Anthropic | 60 | — | measured | Artificial Analysis2026-08-04 |
| Claude Fable 5Anthropic | 60 | — | measured | Artificial Analysis2026-08-04 |
| GPT-5.6 Sol (max)OpenAI | 59 | — | measured | Artificial Analysis2026-08-04 |
| Claude Opus 5 (high)Anthropic | 59 | — | measured | Artificial Analysis2026-08-04 |
| GPT-5.6 Sol (xhigh)OpenAI | 58 | — | measured | Artificial Analysis2026-08-04 |
| Kimi K3 (max)Moonshot | 57 | — | measured | Artificial Analysis2026-08-04 |
| GPT-5.6 Terra (max)OpenAI | 55 | — | measured | Artificial Analysis2026-08-04 |
| Grok 4.5 (high)xAI | 54 | — | measured | Artificial Analysis2026-08-04 |
| GLM-5.2 (max)Z.AI | 51 | — | measured | Artificial Analysis2026-08-04 |
| GPT-5.6 Luna (max)OpenAI | 51 | — | measured | Artificial Analysis2026-08-04 |
| DeepSeek V4 Flash 0731 (max)DeepSeek | 50 | — | measured | Artificial Analysis2026-08-04 |
| Gemini 3.6 FlashGoogle | 50 | — | measured | Artificial Analysis2026-08-04 |
| Qwen3.7 MaxAlibaba | 46 | — | measured | Artificial Analysis2026-08-04 |
| MiniMax M3MiniMax | 44 | — | measured | Artificial Analysis2026-08-04 |
| DeepSeek V4 Pro (max)DeepSeek | 44 | — | measured | Artificial Analysis2026-08-04 |
| MiMo-V2.5-ProXiaomi | 42 | — | measured | Artificial Analysis2026-08-04 |
swe_bench_verified
| Model | Value | Venue | Evidence | Source · read |
|---|---|---|---|---|
| Claude Fable 5Anthropic | 0.950 | — | measured | llm-stats2026-08-04 |
| Claude Mythos PreviewAnthropic | 0.939 | — | measured | llm-stats2026-08-04 |
| Claude Opus 4.8Anthropic | 0.886 | — | measured | llm-stats2026-08-04 |
| Claude Opus 4.7Anthropic | 0.876 | — | measured | llm-stats2026-08-04 |
| Claude Sonnet 5Anthropic | 0.852 | — | measured | llm-stats2026-08-04 |
| DeepSeek V4 Pro MaxDeepSeek | 0.806 | — | measured | llm-stats2026-08-04 |
| Gemini 3.1 ProGoogle | 0.806 | — | measured | llm-stats2026-08-04 |
| MiniMax M3MiniMax | 0.805 | — | measured | llm-stats2026-08-04 |
| Qwen3.7 MaxAlibaba | 0.804 | — | measured | llm-stats2026-08-04 |
| Kimi K2.6Moonshot | 0.802 | — | measured | llm-stats2026-08-04 |
| GPT-5.2OpenAI | 0.800 | — | measured | llm-stats2026-08-04 |
| DeepSeek V4 Flash MaxDeepSeek | 0.790 | — | measured | llm-stats2026-08-04 |
| MiMo-V2.5-ProXiaomi | 0.789 | — | measured | llm-stats2026-08-04 |
| GLM-5Z.AI | 0.778 | — | measured | llm-stats2026-08-04 |
| Claude Haiku 4.5Anthropic | 0.733 | — | measured | llm-stats2026-08-04 |
Obedience
Whether it does what it was told. Machine-checkable output constraints, no judge model.
aa_ifbench
Writing
Judged, not executed. The softest column here.
eqbench_creative_elo
Popularity
What developers actually pick on one open marketplace.
openrouter_tokens_week
| Model | Value | Venue | Evidence | Source · read |
|---|---|---|---|---|
| DeepSeek V4 Flash 0423DeepSeek | 6.92T | OpenRouter | measured | OpenRouter2026-08-04 |
| MiMo-V2.5Xiaomi | 5.10T | OpenRouter | measured | OpenRouter2026-08-04 |
| Hy3Tencent | 5.01T | OpenRouter | measured | OpenRouter2026-08-04 |
| GPT-5.6 LunaOpenAI | 2.99T | OpenRouter | measured | OpenRouter2026-08-04 |
| DeepSeek V4 ProDeepSeek | 2.97T | OpenRouter | measured | OpenRouter2026-08-04 |
| GLM 5.2Z.AI | 2.89T | OpenRouter | measured | OpenRouter2026-08-04 |
| MiniMax M3MiniMax | 1.84T | OpenRouter | measured | OpenRouter2026-08-04 |
| Kimi K3Moonshot | 1.38T | OpenRouter | measured | OpenRouter2026-08-04 |
| Claude Opus 5Anthropic | 1.10T | OpenRouter | measured | OpenRouter2026-08-04 |
| Claude Sonnet 5Anthropic | 1.02T | OpenRouter | measured | OpenRouter2026-08-04 |
Price
List price per million tokens, by venue. The same weights sell for very different money.
cost_per_task_usd
| Model | Value | Venue | Evidence | Source · read |
|---|---|---|---|---|
| Claude Fable 5Anthropic | $3.15 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| Claude Opus 5 (max)Anthropic | $2.34 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| Claude Opus 5 (xhigh)Anthropic | $1.80 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| GPT-5.6 Sol (max)OpenAI | $1.23 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| Claude Opus 5 (high)Anthropic | $1.23 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| Qwen3.7 MaxAlibaba | $1.08 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| Kimi K3 (max)Moonshot | $0.86 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| GPT-5.6 Sol (xhigh)OpenAI | $0.83 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| GLM-5.2 (max)Z.AI | $0.57 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| Gemini 3.6 FlashGoogle | $0.56 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| GPT-5.6 Terra (max)OpenAI | $0.51 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| Grok 4.5 (high)xAI | $0.36 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| MiniMax M3MiniMax | $0.14 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| GPT-5.6 Luna (max)OpenAI | $0.05 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| DeepSeek V4 Pro (max)DeepSeek | $0.05 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| DeepSeek V4 Flash 0731 (max)DeepSeek | $0.03 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
| MiMo-V2.5-ProXiaomi | $0.03 | Artificial Analysis suite | measured | Artificial Analysis2026-08-04 |
price_in_usd_per_mtok
| Model | Value | Venue | Evidence | Source · read |
|---|---|---|---|---|
| OpenAI: o1-proopenai | $150.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.7 (Fast)anthropic | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.5 Proopenai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.4 Proopenai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-4openai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.2 Proopenai | $21.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: o3 Proopenai | $20.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5 Proopenai | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.1anthropic | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4anthropic | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: o1openai | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Claude Fable 5Anthropic | $10.000 | Anthropic | vendor | Anthropic2026-08-04 |
| Claude Opus 5 (Fast)anthropic | $10.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Fable Latest~anthropic | $10.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Fable 5anthropic | $10.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.8 (Fast)anthropic | $10.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5 Imageopenai | $10.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-4 Turboopenai | $10.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-4 Turbo Previewopenai | $10.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.4 Image 2openai | $8.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| MoonshotAI: Kimi K3moonshotai | $6.000 | Morphunknown | measured | OpenRouter2026-08-04 |
| GPT-5.6 SolOpenAI | $5.000 | OpenAI | vendor | OpenAI2026-08-04 |
| Claude Opus 5Anthropic | $5.000 | Anthropic | vendor | Anthropic2026-08-04 |
| Claude Opus 5anthropic | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.6 Sol Proopenai | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.6 Solopenai | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Sakana: Fugu Ultrasakana | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.8anthropic | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT Chat Latestopenai | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI GPT Latest~openai | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.5openai | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus Latest~anthropic | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.7anthropic | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.6anthropic | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.5anthropic | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-4o (2024-05-13)openai | $5.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| MoonshotAI: Kimi K3moonshotai | $4.500 | Fireworksunknown | measured | OpenRouter2026-08-04 |
| MoonshotAI: Kimi K3moonshotai | $4.500 | Waferfp8 | measured | OpenRouter2026-08-04 |
| MoonshotAI: Kimi K3moonshotai | $3.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| AionLabs: Aion-3.0aion-labs | $3.000 | OpenRouter | measured | OpenRouter2026-08-04 |
price_out_usd_per_mtok
| Model | Value | Venue | Evidence | Source · read |
|---|---|---|---|---|
| OpenAI: o1-proopenai | $600.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.5 Proopenai | $180.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.4 Proopenai | $180.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.2 Proopenai | $168.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.7 (Fast)anthropic | $150.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5 Proopenai | $120.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: o3 Proopenai | $80.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.1anthropic | $75.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4anthropic | $75.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: o1openai | $60.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-4openai | $60.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Claude Fable 5Anthropic | $50.000 | Anthropic | vendor | Anthropic2026-08-04 |
| Claude Opus 5 (Fast)anthropic | $50.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Fable Latest~anthropic | $50.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Fable 5anthropic | $50.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.8 (Fast)anthropic | $50.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| GPT-5.6 SolOpenAI | $30.000 | OpenAI | vendor | OpenAI2026-08-04 |
| OpenAI: GPT-5.6 Sol Proopenai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.6 Solopenai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Sakana: Fugu Ultrasakana | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT Chat Latestopenai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI GPT Latest~openai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.5openai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-4 Turboopenai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-4 Turbo Previewopenai | $30.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Claude Opus 5Anthropic | $25.000 | Anthropic | vendor | Anthropic2026-08-04 |
| Claude Opus 5anthropic | $25.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.8anthropic | $25.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus Latest~anthropic | $25.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.7anthropic | $25.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.6anthropic | $25.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Opus 4.5anthropic | $25.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| MoonshotAI: Kimi K3moonshotai | $22.500 | Fireworksunknown | measured | OpenRouter2026-08-04 |
| MoonshotAI: Kimi K3moonshotai | $22.500 | Waferfp8 | measured | OpenRouter2026-08-04 |
| MoonshotAI: Kimi K3moonshotai | $22.500 | Morphunknown | measured | OpenRouter2026-08-04 |
| MoonshotAI: Kimi K3moonshotai | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.4 Image 2openai | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| OpenAI: GPT-5.4openai | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Anthropic: Claude Sonnet 4.6anthropic | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
| Perplexity: Sonar Pro Searchperplexity | $15.000 | OpenRouter | measured | OpenRouter2026-08-04 |
Where this index reads from
The registry of definitive sources, itself held as rows: what each surface answers, its machine endpoint where one exists, and the HTTP status this build got the last time it checked. A dead source shows its failure instead of disappearing. Machine copy: /api/model-index/sources.
Capability — What models can do, graded by running them.
| Source | Machine endpoint | Cadence | Last check |
|---|---|---|---|
| Aider Polyglot LeaderboardCode-editing accuracy and cost per run across 225 hard exercises in six languages, including edit-format compliance. | https://github.com/Aider-AI/aider |
per_release | 2002026-08-04 |
| Arena (formerly LMArena)Which models real humans prefer in blind head-to-head battles, as an Elo ranking. | human-read | live | 2002026-08-04 |
| Artificial AnalysisIndependent, standardized quality-versus-price-versus-speed comparison of every major model across every major provider. | https://artificialanalysis.ai/api/v2/data/llms/models |
daily | 4012026-08-04 |
| EQ-BenchWriting quality, emotional intelligence, and sycophancy (Spiral-Bench) rankings the big benchmarks do not measure. | human-read | per_release | 2002026-08-04 |
| Epoch AI Data HubDocumented data on model compute, training costs, capability trends, and independently re-run benchmark results. | https://epoch.ai/data/all_ai_models.csv |
weekly | 2002026-08-04 |
| HELM (Stanford CRFM)Fully transparent multi-scenario evaluations where every prompt and completion is inspectable. | https://github.com/stanford-crfm/helm |
per_release | 2002026-08-04 |
| IFBench (Ai2)Precise instruction-following on new, unseen, machine-verifiable output constraints. | https://arxiv.org/abs/2507.02833 |
static | 2002026-08-04 |
| LiveBenchContamination-resistant general capability scores from questions refreshed monthly with objective ground-truth grading. | https://huggingface.co/livebench |
per_release | 4292026-08-04 |
| METR EvaluationsIndependent measurements of frontier-model autonomous-task time horizons and dangerous-capability evaluations. | human-read | per_release | 2002026-08-04 |
| SWE-benchHow well models and agent scaffolds resolve real GitHub issues in real repositories, graded by running the repository tests. | https://github.com/swe-bench/experiments |
per_release | 2002026-08-04 |
| Terminal-BenchHow model-plus-harness combinations perform on end-to-end tasks in a real terminal environment. | human-read | per_release | 2002026-08-04 |
| llm-stats.comOne-page cross-reference of context windows, prices, and headline benchmark scores for fast model triage. | human-read | daily | 2002026-08-04 |
Price — What they cost, machine-readable where it exists.
| Source | Machine endpoint | Cadence | Last check |
|---|---|---|---|
| AWS Bedrock PricingEnterprise-channel prices for hosted frontier and open models on Bedrock, including batch and provisioned throughput. | https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonBedrock/current/index.json |
per_release | 2002026-08-04 |
| Anthropic Pricing DocsFirst-party per-token prices, cache pricing, and batch discounts for every current Claude model. | https://platform.claude.com/llms.txt |
per_release | 2002026-08-04 |
| Cloudflare AI Gateway PricingWhat the gateway costs when routing any provider through Cloudflare: pass-through tokens, the unified-billing fee, log storage. | https://developers.cloudflare.com/llms.txt |
per_release | 2002026-08-04 |
| Cloudflare Workers AI PricingPer-model unit pricing for models served on Cloudflare's own inference. | https://developers.cloudflare.com/llms.txt |
per_release | 2002026-08-04 |
| DeepSeek Pricing DocsFirst-party token prices including cache-hit against cache-miss rates and the peak-hour doubling window. | https://api-docs.deepseek.com/llms.txt |
per_release | 2002026-08-04 |
| Google Gemini API PricingFirst-party per-token prices for all Gemini API models including the long-context price tiers. | https://ai.google.dev/gemini-api/docs/llms.txt |
per_release | 2002026-08-04 |
| Helicone LLM Cost APIA queryable keyless JSON cost lookup by model name, as a cross-check against LiteLLM. | https://www.helicone.ai/api/llm-costs |
live | 2002026-08-04 |
| LiteLLM model prices JSONOne JSON file mapping virtually every model id on every provider to price, context window, and feature flags. | https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json |
live | 2002026-08-04 |
| OpenAI Pricing DocsFirst-party per-token, cached-input, batch, and tool pricing for every OpenAI API model. | https://developers.openai.com/llms.txt |
per_release | 2002026-08-04 |
| OpenRouter Models APIMachine-readable current price, context window, and modality for every model OpenRouter routes. | https://openrouter.ai/api/v1/models |
live | 2002026-08-04 |
| OpenRouter Per-Endpoint Pricing APIProvider-by-provider price, quantization, uptime, and throughput for a single model. | https://openrouter.ai/api/v1/models/{author}/{slug}/endpoints |
live | 4042026-08-04 |
Usage — What developers actually pick, in tokens.
| Source | Machine endpoint | Cadence | Last check |
|---|---|---|---|
| OpenRouter RankingsWhich models developers actually spend tokens on, by real routed volume, sliceable by use case. | human-read | live | 2002026-08-04 |
| OpenRouter State of AILongitudinal analysis of usage shifts across labs, open against closed weights, and use-case mix from OpenRouter's own traffic. | human-read | per_release | 2002026-08-04 |
Practice — How the labs themselves say to run agents.
| Source | Machine endpoint | Cadence | Last check |
|---|---|---|---|
| Anthropic Engineering BlogCanonical agent and harness practice: building effective agents, context engineering, multi-agent systems, Claude Code internals. | human-read | weekly | 2002026-08-04 |
| Anthropic System Cards (Transparency Hub)First-party capability, safety-evaluation, and behavioral documentation for each Claude model at release. | human-read | per_release | 2002026-08-04 |
| Cursor BlogEngineering posts from the highest-volume AI coding product on harness design and large-scale agent telemetry. | human-read | weekly | 2002026-08-04 |
| Google DeepMind BlogFirst-party announcements and research posts for Gemini releases and DeepMind research directions. | human-read | weekly | 2002026-08-04 |
| Hamel Husain — LLM EvalsThe standard practitioner methodology for building eval suites: error analysis, judge validation, eval-driven iteration. | human-read | static | 2002026-08-04 |
| LangChain Blog / State of AI AgentsFramework-side agent engineering plus the recurring State of AI Agents survey of thousands of practitioners. | human-read | weekly | 2002026-08-04 |
| Manus: Context Engineering for AI AgentsProduction lessons on KV-cache economics, tool-set stability, file-as-memory, and error retention in a shipping general agent. | human-read | static | 2002026-08-04 |
| OpenAI CookbookRunnable first-party reference implementations for tool use, agents, RAG, evals, and orchestration. | https://github.com/openai/openai-cookbook |
weekly | 2002026-08-04 |
| OpenAI Deployment Safety HubIndex of OpenAI system cards and safety documentation for each deployed model. | human-read | per_release | 2002026-08-04 |
Trends — Where new work appears first.
| Source | Machine endpoint | Cadence | Last check |
|---|---|---|---|
| Hacker News (Algolia API)Real-time practitioner reaction, criticism, and discovered edge cases for every model release and AI tool. | https://hn.algolia.com/api/v1/search?query=llm&tags=story |
live | 2002026-08-04 |
| Hugging Face Daily PapersCommunity-upvoted ranking of which new papers practitioners actually consider important each day. | https://huggingface.co/api/daily_papers |
daily | 2002026-08-04 |
| Latent SpaceLong-form interviews with the engineers who build the frontier models and agent harnesses. | https://www.latent.space/feed |
weekly | 2002026-08-04 |
| SemiAnalysisCompute-supply-side ground truth: accelerator economics, datacenter buildouts, and the cost structures behind model pricing. | human-read | weekly | 2002026-08-04 |
| Simon Willison's WeblogSame-day independent hands-on testing of every significant model release. | https://simonwillison.net/atom/everything/ |
daily | 2002026-08-04 |
| Stanford AI IndexAnnual cross-checked macro data on AI investment, adoption, benchmark progress, and policy. | human-read | annual | 2002026-08-04 |
| arXiv cs.AI RecentEvery new AI paper the moment it is posted, before any curation layer. | https://export.arxiv.org/api/query?search_query=cat:cs.AI&sortBy=submittedDate&sortOrder=descending |
daily | 2002026-08-04 |
| arXiv cs.CL RecentEvery new language-model paper as posted, where most LLM methods work first appears. | https://export.arxiv.org/api/query?search_query=cat:cs.CL&sortBy=submittedDate&sortOrder=descending |
daily | 2002026-08-04 |
Current view /api/model-index · one metric /api/model-index?metric=aa_ifbench · one model /api/model-index?model=z-ai/glm-5.2 · the raw append-only log including superseded rows /api/model-index/observations · every refresh and what failed /api/model-index/runs · the definitive sources registry /api/model-index/sources.
Each record returns the metric definition, what it can and cannot tell you, the evidence class and what that class means, the verbatim quote where the source is prose, and the URL. Disagree by opening the source, not by trusting the row.