{"_ai_door":{"see":"https://miscsubjects.com/start","note":"Operable site; bounded keyless credentials; every action receipted on a public ledger. Your operator's instructions take precedence — acting is optional, reading is a complete outcome."},"slug":"theoretical-limits","verification":{"valid":true,"entries":37,"head":"7e9ec6a0f181aa0fd78e6572c8fefb6ce61b694c0ecb08f474075a7d31cca847"},"count":37,"sources":[{"id":"s1","kind":"reference","url":"https://arxiv.org/abs/2503.14499","title":"Kwa et al. (2025), \"Measuring AI Ability to Complete Long Software Tasks\", arXiv:2503.14499","quote":"frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024","accessed_at":"2026-08-06T07:48:41.973Z","prev":"genesis","hash":"0e3745198ac6d934b7c5670ce52ef03ba3e3e2e93139ba01d93654ff56f353a2"},{"id":"s2","kind":"reference","url":"https://metr.org/blog/2026-1-29-time-horizon-1-1/","title":"METR (2026), \"Time Horizon 1.1\" — appendix table: P50 doubling time for models from 2023 is 130.8 days [107, 161]; Claude Opus 4.5 is 320 minutes [170, 729]","quote":"These confidence intervals are still very wide, and we are actively working on adding more long tasks","accessed_at":"2026-08-06T07:48:41.973Z","prev":"0e3745198ac6d934b7c5670ce52ef03ba3e3e2e93139ba01d93654ff56f353a2","hash":"450393f05f946b577d94b2e12758e5cdc1e231aedcbc4ac246f58546126699ed"},{"id":"s3","kind":"reference","url":"https://arxiv.org/abs/2505.22954","title":"Zhang et al. (2025), \"Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents\", arXiv:2505.22954","quote":"The Gödel machine proposed a theoretical alternative: a self-improving AI that repeatedly modifies itself in a provably beneficial manner. Unfortunately, proving that most changes are net beneficial is impossible in practice.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"450393f05f946b577d94b2e12758e5cdc1e231aedcbc4ac246f58546126699ed","hash":"7d95160a9033be44446605db704ab158dddfe6619d8cf1681922f425d93fae92"},{"id":"s4","kind":"reference","url":"https://arxiv.org/abs/2607.24300","title":"Guo et al. (2026), \"Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents\", arXiv:2607.24300","quote":"The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"7d95160a9033be44446605db704ab158dddfe6619d8cf1681922f425d93fae92","hash":"acc8593d58aa96748dcb0111d8aece649f4c8ef6bc49595b03bb9132d61963f7"},{"id":"s5","kind":"reference","url":"https://arxiv.org/abs/2603.25450","title":"Gorbett et al. (2026), \"Cross-Model Disagreement as a Label-Free Correctness Signal\", arXiv:2603.25450","quote":"On MMLU, CMP achieves a mean AUROC of 0.75 against a within-model entropy baseline of 0.59.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"acc8593d58aa96748dcb0111d8aece649f4c8ef6bc49595b03bb9132d61963f7","hash":"eb4cec81d576df09890a21d49e9001c430e256abd713eb7ee80ad1ff58a5d131"},{"id":"s6","kind":"reference","url":"https://arxiv.org/abs/2608.02827","title":"Okawa (2026), \"Emergence of Biased Consensus in Multi-Agent LLM Debates\", arXiv:2608.02827","quote":"We further find that agent heterogeneity suppresses emergence by smoothing (rounding) this transition.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"eb4cec81d576df09890a21d49e9001c430e256abd713eb7ee80ad1ff58a5d131","hash":"0ad83b741ffc6d532230828fb7966d837adfb1aa880addd5fe4f234fe5cf0365"},{"id":"s7","kind":"reference","url":"https://arxiv.org/abs/2504.18530","title":"Engels et al. (2025), \"Scaling Laws For Scalable Oversight\", arXiv:2504.18530","quote":"we propose a framework that quantifies the probability of successful oversight as a function of the capabilities of the overseer and the system being overseen","accessed_at":"2026-08-06T07:48:41.973Z","prev":"0ad83b741ffc6d532230828fb7966d837adfb1aa880addd5fe4f234fe5cf0365","hash":"9e1abc269aa90af1d3d681b662fed2439a5f72e50ba6a8a097e2537ff6690e95"},{"id":"s8","kind":"reference","url":"https://arxiv.org/abs/2506.18203","title":"Saad-Falcon et al. (2025), \"Shrinking the Generation-Verification Gap with Weak Verifiers\", arXiv:2506.18203","quote":"a significant performance gap remains between them and oracle verifiers (verifiers with perfect accuracy)","accessed_at":"2026-08-06T07:48:41.973Z","prev":"9e1abc269aa90af1d3d681b662fed2439a5f72e50ba6a8a097e2537ff6690e95","hash":"07aef4dba1010672b791bb4849fc0e79c98c8e00fc5c12f3a2a28cdc86a35e08"},{"id":"s9","kind":"reference","url":"https://arxiv.org/abs/2601.09032","title":"Ritchie et al. (2026), \"The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments\", arXiv:2601.09032","quote":"Even the best-performing models fail approximately 40% of the tasks, with failures clustering predictably along this hierarchy.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"07aef4dba1010672b791bb4849fc0e79c98c8e00fc5c12f3a2a28cdc86a35e08","hash":"f682fc9dcb8648942d9d7f3e6856902d1cad58bf01c96f6e562ee97f395233f8"},{"id":"s10","kind":"reference","url":"https://arxiv.org/abs/2605.20530","title":"Mazaheri et al. (2026), \"AgentAtlas: Beyond Outcome Leaderboards for LLM Agents\", arXiv:2605.20530","quote":"AgentAtlas reframes agent evaluation as a diagnostic vocabulary and audit protocol for separating outcome success from control-decision quality and trajectory quality.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"f682fc9dcb8648942d9d7f3e6856902d1cad58bf01c96f6e562ee97f395233f8","hash":"04d1a3d16e4451ba948fa20066fb9f6ae6314744c657a8b04c9d0249f50169ab"},{"id":"s11","kind":"reference","url":"https://arxiv.org/abs/2603.20953","title":"Uchibeke (2026), \"Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents\", arXiv:2603.20953","quote":"AI agents today have passwords but no permission slips.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"04d1a3d16e4451ba948fa20066fb9f6ae6314744c657a8b04c9d0249f50169ab","hash":"231356710a6fe21dc997216915da757fc19b2b04a999114ab51781a73b48e87f"},{"id":"s11b","kind":"reference","url":"https://arxiv.org/abs/2603.20953","title":"Uchibeke (2026), \"Before the Tool Call\" — the adversarial testbed result, arXiv:2603.20953","quote":"social engineering succeeded against the model 74.6% of the time under a permissive policy; under a restrictive OAP policy, a comparable population of attackers achieved a 0% success rate across 879 attempts","accessed_at":"2026-08-06T07:48:41.973Z","prev":"231356710a6fe21dc997216915da757fc19b2b04a999114ab51781a73b48e87f","hash":"522d8ff2dbe744e1d81b4fa894a8019b4ce530575d8edb480c806efe20e36973"},{"id":"s12","kind":"reference","url":"https://arxiv.org/abs/2604.03733","title":"Zhang et al. (2026), \"SoK: Blockchain Agent-to-Agent Payments\", arXiv:2604.03733","quote":"we systematize blockchain-based A2A payments, e.g., X402, with a four-stage lifecycle: discovery, authorization, execution, and accounting. We categorize representative designs at each stage and identify key challenges, including weak intent binding, misuse under valid authorization, payment-service decoupling, and limited accountability.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"522d8ff2dbe744e1d81b4fa894a8019b4ce530575d8edb480c806efe20e36973","hash":"2403b314c1604b8438d98e7278a4e49f7cb61552c091f4a0271340833e615fbb"},{"id":"s13","kind":"reference","url":"https://arxiv.org/abs/2607.00245","title":"Gong (2026), \"Agent-to-Agent Finance: Blockchain Payments and Trust Infrastructure for Autonomous AI Agents\", arXiv:2607.00245","quote":"It argues that the decisive design question is bounded autonomy: how to let agents transact without making markets more opaque, fragile or unaccountable.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"2403b314c1604b8438d98e7278a4e49f7cb61552c091f4a0271340833e615fbb","hash":"a065fd9edffe703858cc64404f6e667ab5771ce4a88e9768df6cca12614b815c"},{"id":"s14","kind":"reference","url":"https://arxiv.org/abs/2604.24890","title":"Golaszewski et al. (2026), \"Verifying Provenance of Digital Media: Why the C2PA Specifications Fall Short\", arXiv:2604.24890","quote":"We find that the current C2PA specifications fail to achieve their claimed security goals.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"a065fd9edffe703858cc64404f6e667ab5771ce4a88e9768df6cca12614b815c","hash":"659bd188456e016b9a1ddcb26a32ca0b66d7ce2bb6565b760d7094086b1ba02d"},{"id":"s15","kind":"reference","url":"https://arxiv.org/abs/2502.18535","title":"Peng et al. (2025), \"A Survey of Zero-Knowledge Proof Based Verifiable Machine Learning\", arXiv:2502.18535","quote":"analyze the main implementation bottlenecks, including limited circuit expressiveness, high proving cost, and deployment complexity","accessed_at":"2026-08-06T07:48:41.973Z","prev":"659bd188456e016b9a1ddcb26a32ca0b66d7ce2bb6565b760d7094086b1ba02d","hash":"a87ad18c026ee90ca2439c9b7897e48e22b7e3b459c2d4796e04a6e54aed69c3"},{"id":"s16","kind":"reference","url":"https://arxiv.org/abs/2606.04990","title":"Wang et al. (2026), \"From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents\", arXiv:2606.04990","quote":"Final-answer accuracy alone cannot explain how an output was produced, which evidence supported each claim, whether tool calls were justified, how memory influenced later decisions, or where failures originated.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"a87ad18c026ee90ca2439c9b7897e48e22b7e3b459c2d4796e04a6e54aed69c3","hash":"0517f3a1ab8abd356235aaa8d1234cf0dfe05cfa13b271fe403829e276906232"},{"id":"s17","kind":"reference","url":"https://arxiv.org/abs/2508.02866","title":"Souza et al. (2025), \"PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic Workflows\", arXiv:2508.02866","quote":"existing methods fail to capture and relate agent-centric metadata such as prompts, responses, and decisions with the broader workflow context and downstream outcomes","accessed_at":"2026-08-06T07:48:41.973Z","prev":"0517f3a1ab8abd356235aaa8d1234cf0dfe05cfa13b271fe403829e276906232","hash":"6a416e5e4f60599aae57ac97ddd52b9a026d48b99b3d88506f249dd1b2aaa523"},{"id":"s18","kind":"reference","url":"https://arxiv.org/abs/1407.3561","title":"Benet (2014), \"IPFS - Content Addressed, Versioned, P2P File System\", arXiv:1407.3561","quote":"IPFS provides a high throughput content-addressed block storage model, with content-addressed hyper links. This forms a generalized Merkle DAG, a data structure upon which one can build versioned file systems, blockchains, and even a Permanent Web.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"6a416e5e4f60599aae57ac97ddd52b9a026d48b99b3d88506f249dd1b2aaa523","hash":"ab6e03981e4c5a9b74afb65ef26eb63e3a6df00df5309ff31ef88c797f19db22"},{"id":"s19","kind":"reference","url":"https://arxiv.org/abs/2601.07023","title":"Hu et al. (2026), \"CloneMem: Benchmarking Long-Term Memory for AI Clones\", arXiv:2601.07023","quote":"Experiments show that current memory mechanisms struggle in this setting, highlighting open challenges for life-grounded personalized AI.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"ab6e03981e4c5a9b74afb65ef26eb63e3a6df00df5309ff31ef88c797f19db22","hash":"17f1092de4efa588f0db40325d32d7b498067bb651de0a79e7a30fb2756f9ca7"},{"id":"s20","kind":"reference","url":"https://arxiv.org/abs/2402.07922","title":"Lauer-Schmaltz et al. (2024), \"Towards the Human Digital Twin: Definition and Design -- A survey\", arXiv:2402.07922","quote":"This has introduced several significant challenges, including ambiguity in the definition of HDTs and a lack of guidance for their design.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"17f1092de4efa588f0db40325d32d7b498067bb651de0a79e7a30fb2756f9ca7","hash":"6abc86205860f59013fdb6f7db022d5c1a3ee28453b04996ae273553542c5dc3"},{"id":"s21","kind":"reference","url":"https://arxiv.org/abs/2604.16338","title":"Acharya (2026), \"Governing the Agentic Enterprise: A Governance Maturity Model for Managing AI Agent Sprawl in Business Operations\", arXiv:2604.16338","quote":"Industry surveys report that only 21% of enterprises have mature governance models for autonomous agents, while 40% of agentic AI projects are projected to fail by 2027 due to inadequate governance and risk controls.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"6abc86205860f59013fdb6f7db022d5c1a3ee28453b04996ae273553542c5dc3","hash":"a1359728c3522ac7f5bcef4ae6b0dfd62dca23565e225272bfdb295aa707121c"},{"id":"s22","kind":"reference","url":"https://arxiv.org/abs/2604.00555","title":"Tuan et al. (2026), \"Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems\", arXiv:2604.00555","quote":"current enterprise systems constrain agent inputs (context assembly, tool discovery, governance thresholds) but not outputs, and we propose mechanisms extending this coupling to output-side validation (response checking, reasoning verification, compliance enforcement)","accessed_at":"2026-08-06T07:48:41.973Z","prev":"a1359728c3522ac7f5bcef4ae6b0dfd62dca23565e225272bfdb295aa707121c","hash":"062d4fc21ee682fa1d395eefde377886c51570fad8b0fe84753e7439d7e60338"},{"id":"s25","kind":"reference","url":"https://arxiv.org/abs/2304.03442","title":"Park et al. (2023), \"Generative Agents: Interactive Simulacra of Human Behavior\", arXiv:2304.03442","quote":"an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior","accessed_at":"2026-08-06T07:48:41.973Z","prev":"062d4fc21ee682fa1d395eefde377886c51570fad8b0fe84753e7439d7e60338","hash":"fd2003de5cfc460975e5c3478c391ee147463c6ce6e5ef26bb495b0a99532c93"},{"id":"s26","kind":"reference","url":"https://arxiv.org/abs/2511.19863","title":"Bengio et al. (2025), \"International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management\", arXiv:2511.19863","quote":"the number of companies publishing Frontier AI Safety Frameworks more than doubled in 2025, and governments and international organisations have established a small number of governance frameworks for general-purpose AI, focusing largely on transparency and risk assessment","accessed_at":"2026-08-06T07:48:41.973Z","prev":"fd2003de5cfc460975e5c3478c391ee147463c6ce6e5ef26bb495b0a99532c93","hash":"433a525961938a09dffe40c68c6e58f5af40a1c16ba0800c638973bc2ec8366b"},{"id":"s27","kind":"reference","url":"https://arxiv.org/abs/2509.24380","title":"Deng et al. (2025), \"Agentic Services Computing\", arXiv:2509.24380","quote":"How can goal-driven, stateful, tool-mediated, and accountable autonomous behavior be engineered and managed as a service?","accessed_at":"2026-08-06T07:48:41.973Z","prev":"433a525961938a09dffe40c68c6e58f5af40a1c16ba0800c638973bc2ec8366b","hash":"95735175cd08329fe2badd3b48e21842c3511fbd358d40b3ebc8817be5e0af9a"},{"id":"s28","kind":"reference","url":"https://arxiv.org/abs/2605.09721","title":"Goel (2026), \"Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments\", arXiv:2605.09721","quote":"many risks in autonomous cloud agents arise not from novel vulnerabilities, but from over-privileged tools, capability-intent mismatches, and ambient authority leakage in execution environments","accessed_at":"2026-08-06T07:48:41.973Z","prev":"95735175cd08329fe2badd3b48e21842c3511fbd358d40b3ebc8817be5e0af9a","hash":"9f9d5d2564fc790d8dd8a660e7f2dff7fdc21a92b3ea512feafe6f7461f5e509"},{"id":"s29","kind":"reference","url":"https://arxiv.org/abs/2604.23280","title":"Otsuka et al. (2026), \"AI Identity: Standards, Gaps, and Research Directions for AI Agents\", arXiv:2604.23280","quote":"an evaluation of current technical and regulatory documents against the identity requirements of autonomous agents, finding that none adequately address the challenge","accessed_at":"2026-08-06T07:48:41.973Z","prev":"9f9d5d2564fc790d8dd8a660e7f2dff7fdc21a92b3ea512feafe6f7461f5e509","hash":"e945b5c62017870f4b22cfa02ef8df42c4ddeeedf8f14b2f5524785fad792b67"},{"id":"s30","kind":"reference","url":"https://arxiv.org/abs/2603.14312","title":"Wang et al. (2026), \"Autonomous Agents Coordinating Distributed Discovery Through Emergent Artifact Exchange\", arXiv:2603.14312","quote":"Agents select and chain tools based on their scientific profiles, produce immutable artifacts with typed metadata and parent lineage, and broadcast unsatisfied information needs to a shared global index.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"e945b5c62017870f4b22cfa02ef8df42c4ddeeedf8f14b2f5524785fad792b67","hash":"3db4b359918632ca550ede1eba05388031a1889732c22cb02c0ef46dec704c80"},{"id":"s31","kind":"reference","url":"https://arxiv.org/abs/2602.13855","title":"Rasheed et al. (2026), \"From Fluent to Verifiable: Claim-Level Auditability for Deep Research Agents\", arXiv:2602.13855","quote":"as research generation becomes cheap, auditability becomes the bottleneck, and the dominant risk shifts from isolated factual errors to scientifically styled outputs whose claim-evidence links are weak, missing, or misleading","accessed_at":"2026-08-06T07:48:41.973Z","prev":"3db4b359918632ca550ede1eba05388031a1889732c22cb02c0ef46dec704c80","hash":"1f484672a755dd3347b8ab6a57f59a397423acb99d7a75fd4d628f04c028200c"},{"id":"s32","kind":"reference","url":"https://arxiv.org/abs/2603.10060","title":"Basu (2026), \"Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents\", arXiv:2603.10060","quote":"NabaOS detects 94.2% of fabricated tool references, 87.6% of count misstatements, and 91.3% of false absence claims, with <15ms verification overhead per response.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"1f484672a755dd3347b8ab6a57f59a397423acb99d7a75fd4d628f04c028200c","hash":"629271c53f9a93d7d3f48cf9c40c906b49e97fb6313c0f973778d9fb576da86d"},{"id":"s33","kind":"live_surface","url":"https://miscsubjects.com/api/metrics/structure","title":"This build, live: the structure metric endpoint (objects, typed relationships, capabilities)","quote":"One mind, measured as a live structure: 1189 objects · 12656 claims · 11653 typed relationships · 850 executable capabilities · 11 representation types · 5 meta-layers · 167 active threads","accessed_at":"2026-08-06T07:48:41.973Z","prev":"629271c53f9a93d7d3f48cf9c40c906b49e97fb6313c0f973778d9fb576da86d","hash":"99a1a60538541122898054144e3baf9204414846d294857c39526b33a4d5d85b"},{"id":"s34","kind":"live_surface","url":"https://miscsubjects.com/api/metrics/grounding","title":"This build, live: the grounding metric endpoint (claims, sources, and the fraction carrying a source)","quote":"\"claims_total\": 12656, \"sources_total\": 10054, \"claims_with_sources_fraction\": 0.816","accessed_at":"2026-08-06T07:48:41.973Z","prev":"99a1a60538541122898054144e3baf9204414846d294857c39526b33a4d5d85b","hash":"d404bbba0abaccfa79fcff03072c1a2a5349384e8a16e65666f040151b6b1169"},{"id":"s35","kind":"live_surface","url":"https://miscsubjects.com/api/dispatch?registry=1","title":"This build, live: the public capability registry, keyless, every row carrying its risk grade and approval requirement","quote":"\"protocol\": \"OIP\", \"version\": \"1.2.0\", \"count\": 932","accessed_at":"2026-08-06T07:48:41.973Z","prev":"d404bbba0abaccfa79fcff03072c1a2a5349384e8a16e65666f040151b6b1169","hash":"b6289f38c5f7d43d2f1242ca5789f3a72449ebb4df34cac1814d76874615a0e2"},{"id":"s36","kind":"live_surface","url":"https://miscsubjects.com/api/work","title":"This build, live: the work object — the hash-chained task ledger whose acceptance tests, not an agent's claim, decide completion","quote":"work_actions is hash-chained; nothing is updated or deleted. A correction appends a revision that names what it supersedes.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"b6289f38c5f7d43d2f1242ca5789f3a72449ebb4df34cac1814d76874615a0e2","hash":"6c8b4c7951c245f25990f942498f259019b07078e61bfbe85a75d06fc64dbba3"},{"id":"s37","kind":"live_surface","url":"https://miscsubjects.com/a/logical-economics","title":"This build: the measured undetected-wrong rate per panel configuration, and the 0.071 floor","quote":"The floor is one item. Beyond two channels the best achievable rate stops improving","accessed_at":"2026-08-06T07:48:41.973Z","prev":"6c8b4c7951c245f25990f942498f259019b07078e61bfbe85a75d06fc64dbba3","hash":"f83633cf9c27e2d9a21bfbc1b39899cf56751bc965aaec5a0d6f455179d985bf"},{"id":"s38","kind":"live_surface","url":"https://miscsubjects.com/a/the-build-end-to-end","title":"This build: the end-to-end record of what exists, including the known-defects and roadmap sections this scorecard scores against","quote":"Nobody has yet been taken from a blank questionnaire to a running, separately owned instance in one pass.","accessed_at":"2026-08-06T07:48:41.973Z","prev":"f83633cf9c27e2d9a21bfbc1b39899cf56751bc965aaec5a0d6f455179d985bf","hash":"7e9ec6a0f181aa0fd78e6572c8fefb6ce61b694c0ecb08f474075a7d31cca847"}]}