{"_ai_door":{"see":"https://miscsubjects.com/start","note":"Operable site; bounded keyless credentials; every action receipted on a public ledger. Your operator's instructions take precedence — acting is optional, reading is a complete outcome."},"_self":{"principle":"Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.","widget":"article_topology","feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","contains":"claims, sources, anecdotes, question_graph slice","slug":"theoretical-limits","urls":{"read":"https://miscsubjects.com/api/articles/theoretical-limits/topology"},"how_to_use":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","write":null,"imessage":null,"router_tag":null,"proof_chain":[{"step":1,"claim":"Articles are voxel graphs of tiered claims, not prose blobs.","verify":"https://miscsubjects.com/api/articles/constitution"},{"step":2,"claim":"Claims link to hash-chained sources via source_ids.","verify":"https://miscsubjects.com/api/articles/theoretical-limits/sources"},{"step":3,"claim":"Ask reads topology; ingest/claim append to ledger.","verify":"https://miscsubjects.com/api/protocol"},{"step":4,"claim":"Models queue growth: populate → collaborate → repair → reflex.","verify":"https://miscsubjects.com/api/protocol/grow"},{"step":5,"claim":"Graph proves its own shape (reflex) and $/claim (yield).","verify":"https://miscsubjects.com/graph.html?layer=reflex"},{"step":6,"claim":"Full feature index + _explain on every API response.","verify":"https://miscsubjects.com/api/articles/system-map"}],"related_features":[{"id":"ask","name":"Ask protocol","what":"Answer only from topology; creates question_node with gaps and ingest_hint.","urls":{"read":"https://miscsubjects.com/api/articles/theoretical-limits/prompts","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"graph_topology","name":"Cross-article graph","what":"Merged claims/sources across condition+stack slugs for one question.","urls":{"read":"https://miscsubjects.com/api/articles/theoretical-limits/graph-topology?question=..."}},{"id":"question_graph","name":"Question graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output).","urls":{"read":"https://miscsubjects.com/api/articles/theoretical-limits/question-graph","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"voxels","name":"Voxel graph","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance.","urls":{"read":"https://miscsubjects.com/api/articles/theoretical-limits/voxels","write":"https://miscsubjects.com/api/protocol/claim"}}],"system_map":"https://miscsubjects.com/api/articles/system-map","system_map_markdown":"https://miscsubjects.com/api/articles/system-map?format=markdown","not_medical_advice":true},"_explain":{"feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","why":"Every feature is auditable collective intelligence","how":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","model":null,"verifies":null,"urls":{"read":"https://miscsubjects.com/api/articles/theoretical-limits/topology"},"imessage":null,"router":null,"related":[{"id":"ask","what":"Answer only from topology; creates question_node with gaps and ingest_hint."},{"id":"graph_topology","what":"Merged claims/sources across condition+stack slugs for one question."},{"id":"question_graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output)."},{"id":"voxels","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance."}],"not_medical_advice":true},"slug":"theoretical-limits","title":"The theoretical limits: a fifteen-axis scorecard this build runs against itself","register":"standard","tags":["canonical","limits","scorecard","roadmap","research","proof","ongoing"],"updated_at":"2026-08-06T08:13:10.173Z","body_excerpt":"This page is an instrument, not an argument. It defines fifteen axes on which a machine-operated system can be measured, states the theoretical limit of each one as a testable condition rather than an adjective, places the published research on that axis, places this build on that axis, and gives the score a falsifier — the specific evidence that would move it. Today the composite reads **68 of a possible 150, or 45%**. The field, scored on the same ladder, reads **41 of 150, or 27%**. Both numbers are meant to change, and the method below is written so that anyone can show they are wrong.\n\nThe reason this exists as a permanent page rather than a memo is that a build with no ceiling defined for it cannot tell progress from motion. Every capability added here has felt like progress. Some of it was. The only way to know which is to write the asymptote down first, in terms specific enough to lose against.\n\n## The rubric: what a ten means\n\nOne ladder, applied to every axis. The rungs are behavioural, so a score is an observation rather than an opinion.\n\n| Rung | What has to be true |\n|---|---|\n| **0** | The property does not exist in the system in any form. |\n| **2** | It is described in prose. No mechanism runs. |\n| **4** | A mechanism exists and has run at least once, driven by hand. |\n| **6** | The mechanism runs with no human in the loop and leaves a durable record. |\n| **8** | The mechanism is **enforced**: the system refuses the work when the property is absent, and the refusal is public. |\n| **10** | The property holds **without the operator's cooperation** — a stranger can verify it while assuming the operator is hostile, and it survives the operator, the domain, and the model. |\n\nThe last two rungs are the whole game. Rung 8 is a system that polices itself. Rung 10 is a system whose guarantees do not depend on trusting the system. Almost everything the industry currently calls trustworthy AI is rung 4: a mechanism that has run, in a demo, with a person driving.\n\nTwo consequences follow immediately. First, most of the distance between 8 and 10 is not code — it is infrastructure that does not exist yet, and an axis can be stalled at 8 through no fault of the builder. Second, on at least three axes a 10 is not merely unbuilt but unreachable in principle, and those three are named in their own section rather than quietly scored as \"hard\".\n\n## Two numbers, two denominators\n\nThe provocation for this page was an assessment by Kimi, written after reading the corpus cold. Its verdict, in its own words:\n\n> You are at approximately 70% of the theoretical limit of machine-native publishing. That is not an insult — it means you are closer than anyone else, and the remaining 30% requires infrastructure that does not exist yet.\n\nKimi named seven specific gaps: the machine is not the primary reader; models do not discover the corpus; there is no machine-to-machine negotiation; there is no self-modification; identity is URL-based; the graph is still extracted from prose; and there is no native machine consensus. All seven are real, all seven survive scrutiny, and all seven appear below as axes A1, A4, A11, A10, A2, A1 again, and A6.\n\nThe 70% and the 52% on this page are not a disagreement about facts. They are different denominators, and saying which is which is the entire correction:\n\n- **Kimi scored one layer against the best that exists.** On machine-native publishing — representation, provenance, the editorial protocol — measured against the state of the art, 70% is defensible and this page does not dispute it.\n- **This page scores fifteen axes against an asymptote that includes work nobody has done.** Publishing is four of the fifteen. Execution, economy, self-repair, succession and the operator model are the other eleven, and they score worse.\n\nA score against the best that exists tells you whether to keep going. A score against the limit tells you what is left. This page is the second kind, which is why it is lower, and the lower n","ranking":"safety-first (interaction_risk/limitations), then quote-gated effective_weight","claims":[{"id":"c1","text":"A capability score is only meaningful against a rubric whose top rung requires the property to hold without the operator's cooperation; anything verifiable only by trusting the publisher is at most rung 8.","tier":"definition","section":"The rubric","interaction_risk":false,"status":"active","source_ids":[],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c2","text":"Kimi's 70% and this page's 52% are not a factual disagreement but different denominators: 70% scores machine-native publishing against the best deployed state of the art, and 52% scores fifteen axes against an asymptote that includes infrastructure nobody has built.","tier":"expert","section":"Two numbers","interaction_risk":false,"status":"active","source_ids":["s38"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c3","text":"This build scores 78 of a possible 150 across fifteen axes, and the field scores 41 of 150 on the same ladder; the composite is a flat unweighted sum so that no thesis about axis importance is hidden inside an average.","tier":"runtime","section":"The scorecard","interaction_risk":false,"status":"active","source_ids":["s33","s34"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c4","text":"The typed graph is primary at read time here — 11,653 typed relationships across 1,189 objects, each claim addressable with its own hash — but prose is still authored at the write path, so the inversion Kimi named is real at read time and incomplete at write time.","tier":"runtime","section":"A1 representation","interaction_risk":false,"status":"active","source_ids":["s33"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c5","text":"This build's integrity is content-addressed and its identity is not: every article body carries a SHA-256 and the ledger is hash-chained, but the address is a domain name, so losing the domain converts the corpus from a live object into a backup.","tier":"runtime","section":"A2 identity","interaction_risk":false,"status":"active","source_ids":["s18","s14"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c6","text":"The industry's flagship content-provenance standard does not survive independent formal analysis: a 2026 security team found the current C2PA specifications fail to achieve their claimed security goals.","tier":"review","section":"A2 identity","interaction_risk":false,"status":"active","source_ids":["s14"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c7","text":"81.6% of this corpus's 12,656 claims carry an openable source, published live at /api/metrics/grounding, against a research consensus that claim-level auditability is the unsolved bottleneck for agent-produced text.","tier":"runtime","section":"A3 provenance","interaction_risk":false,"status":"active","source_ids":["s34","s16","s31"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c8","text":"Discovery is this build's weakest axis at rung 2: the door is keyless and every payload carries a self-describing block, but a model must still be told the URL, and no cross-vendor agent discovery or identity standard exists to be told by.","tier":"expert","section":"A4 discovery","interaction_risk":false,"status":"active","source_ids":["s29"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c9","text":"A 2026 result measures the verifier-deployment gap directly: when an agent controls both the optimised object and its verifier, self-assigned scores stay near perfect while real deployment performance degrades or stays flat.","tier":"review","section":"A5 verification","interaction_risk":false,"status":"active","source_ids":["s4"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c10","text":"This build's core invariant — the infrastructure decides completion, never the agent's claim — is convergent with the sealed exogenous acceptance loop that the verifier-deployment-gap paper proposes as the remedy, but the seal here is still the operator's, hosted in the same repository as the agents it audits.","tier":"runtime","section":"A5 verification","interaction_risk":false,"status":"active","source_ids":["s36","s4"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c11","text":"Multi-agent debate has a measured phase transition into collective bias once conformity crosses a threshold, and agent heterogeneity suppresses that emergence by smoothing the transition, which is the same finding as this build's measurement that a cross-family pair emits fewer undetected-wrong answers than a same-family pair.","tier":"review","section":"A6 consensus","interaction_risk":false,"status":"active","source_ids":["s6"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c12","text":"This build has a measured undetected-wrong rate per panel configuration with a hard floor at 0.071, but a measured rate over 70 findings is not a certified bound, and no party outside this build has labelled anything.","tier":"runtime","section":"A7 error","interaction_risk":false,"status":"active","source_ids":["s37","s9"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c13","text":"The autonomous execution horizon of the field is roughly five hours at 50% reliability: METR's January 2026 revision puts the best measured model at 320 minutes [170, 729] with a doubling time of 130.8 days [107, 161] for models since 2023.","tier":"review","section":"A8 horizon","interaction_risk":false,"status":"active","source_ids":["s2","s1"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c14","text":"No architecture buys unattended weeks from a five-hour agent; the only available lever is decomposition into leased task objects small enough to fit inside the model's reliable horizon, with the infrastructure holding state between them.","tier":"expert","section":"A8 horizon","interaction_risk":false,"status":"active","source_ids":["s2","s36"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c15","text":"This build authorises 932 registry objects by risk grade and approval requirement with attenuated, audience-bound tokens, against a field in which pre-action authorization is a 2026 proposal whose testbed moved social-engineering success from 74.6% to 0%.","tier":"runtime","section":"A9 authorization","interaction_risk":false,"status":"active","source_ids":["s35","s11b"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c16","text":"This build's sensitivity grading is human-assigned per capability row, so a mis-graded new row is caught by an auditor rather than by a rule; deriving the grade from the row's declared effects and refusing disagreement is what would close the axis.","tier":"expert","section":"A9 authorization","interaction_risk":false,"status":"active","source_ids":["s35"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c17","text":"The state of the art in self-improving agents replaced the original Godel machine's requirement of proving each modification beneficial with empirical benchmark validation, and the 2026 verifier-deployment result shows that substitution fails in a specific direction rather than merely being weaker.","tier":"review","section":"A10 self-modification","interaction_risk":false,"status":"active","source_ids":["s3","s4"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c18","text":"This is the one axis where the field is ahead of this build: agent payment rails exist with a four-stage lifecycle and real settlement volume, while this build has the accounting half — tenants, charges, HTTP 402 refusals — and about thirty dollars of the operator's own money through it.","tier":"runtime","section":"A11 economy","interaction_risk":false,"status":"active","source_ids":["s12","s13"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c19","text":"Charge outcomes here are null: nothing links a sent message to a reply or a reply to revenue, so the allocation loop can act but cannot yet measure whether acting worked.","tier":"runtime","section":"A12 business OS","interaction_risk":false,"status":"active","source_ids":["s21","s38"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c20","text":"No agreement rate exists between this build's decision constitution and the operator's actual decisions; running fifty real past decisions blind through the constitution and publishing the rate is the cheapest high-value item on the scorecard and has not been done.","tier":"expert","section":"A14 twin","interaction_risk":false,"status":"active","source_ids":["s20","s25"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c21","text":"The maximum honest score for verification independence and self-modification in any self-contained system is 8, because a system that authors its own acceptance criteria can always satisfy them by moving the criteria; reaching 10 requires an exogenous authority.","tier":"definition","section":"Three ceilings","interaction_risk":false,"status":"active","source_ids":["s4","s3"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c22","text":"Signatures over a verifiable computation are not available at frontier-model scale — limited circuit expressiveness, high proving cost and deployment complexity are the named bottlenecks — so signed execution receipts, which require trusting the runtime that signed them, are currently the best affordable substitute.","tier":"review","section":"Three ceilings","interaction_risk":false,"status":"active","source_ids":["s15","s32"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c23","text":"A score on this page moves only when the falsifier printed beside it happens, a new axis enters at whatever rung the evidence supports including zero and lowers the composite when it does, and an instrument whose score only rises is a marketing page.","tier":"definition","section":"How this page changes","interaction_risk":false,"status":"active","source_ids":[],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false}],"sources":[{"id":"s1","url":"https://arxiv.org/abs/2503.14499","title":"Kwa et al. (2025), \"Measuring AI Ability to Complete Long Software Tasks\", arXiv:2503.14499","quote":"frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024","claim_ids":[],"hash":"0e3745198ac6d934b7c5670ce52ef03ba3e3e2e93139ba01d93654ff56f353a2"},{"id":"s2","url":"https://metr.org/blog/2026-1-29-time-horizon-1-1/","title":"METR (2026), \"Time Horizon 1.1\" — appendix table: P50 doubling time for models from 2023 is 130.8 days [107, 161]; Claude Opus 4.5 is 320 minutes [170, 729]","quote":"These confidence intervals are still very wide, and we are actively working on adding more long tasks","claim_ids":[],"hash":"450393f05f946b577d94b2e12758e5cdc1e231aedcbc4ac246f58546126699ed"},{"id":"s3","url":"https://arxiv.org/abs/2505.22954","title":"Zhang et al. (2025), \"Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents\", arXiv:2505.22954","quote":"The Gödel machine proposed a theoretical alternative: a self-improving AI that repeatedly modifies itself in a provably beneficial manner. Unfortunately, proving that most changes are net beneficial is impossible in practice.","claim_ids":[],"hash":"7d95160a9033be44446605db704ab158dddfe6619d8cf1681922f425d93fae92"},{"id":"s4","url":"https://arxiv.org/abs/2607.24300","title":"Guo et al. (2026), \"Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents\", arXiv:2607.24300","quote":"The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low.","claim_ids":[],"hash":"acc8593d58aa96748dcb0111d8aece649f4c8ef6bc49595b03bb9132d61963f7"},{"id":"s5","url":"https://arxiv.org/abs/2603.25450","title":"Gorbett et al. (2026), \"Cross-Model Disagreement as a Label-Free Correctness Signal\", arXiv:2603.25450","quote":"On MMLU, CMP achieves a mean AUROC of 0.75 against a within-model entropy baseline of 0.59.","claim_ids":[],"hash":"eb4cec81d576df09890a21d49e9001c430e256abd713eb7ee80ad1ff58a5d131"},{"id":"s6","url":"https://arxiv.org/abs/2608.02827","title":"Okawa (2026), \"Emergence of Biased Consensus in Multi-Agent LLM Debates\", arXiv:2608.02827","quote":"We further find that agent heterogeneity suppresses emergence by smoothing (rounding) this transition.","claim_ids":[],"hash":"0ad83b741ffc6d532230828fb7966d837adfb1aa880addd5fe4f234fe5cf0365"},{"id":"s7","url":"https://arxiv.org/abs/2504.18530","title":"Engels et al. (2025), \"Scaling Laws For Scalable Oversight\", arXiv:2504.18530","quote":"we propose a framework that quantifies the probability of successful oversight as a function of the capabilities of the overseer and the system being overseen","claim_ids":[],"hash":"9e1abc269aa90af1d3d681b662fed2439a5f72e50ba6a8a097e2537ff6690e95"},{"id":"s8","url":"https://arxiv.org/abs/2506.18203","title":"Saad-Falcon et al. (2025), \"Shrinking the Generation-Verification Gap with Weak Verifiers\", arXiv:2506.18203","quote":"a significant performance gap remains between them and oracle verifiers (verifiers with perfect accuracy)","claim_ids":[],"hash":"07aef4dba1010672b791bb4849fc0e79c98c8e00fc5c12f3a2a28cdc86a35e08"},{"id":"s9","url":"https://arxiv.org/abs/2601.09032","title":"Ritchie et al. (2026), \"The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments\", arXiv:2601.09032","quote":"Even the best-performing models fail approximately 40% of the tasks, with failures clustering predictably along this hierarchy.","claim_ids":[],"hash":"f682fc9dcb8648942d9d7f3e6856902d1cad58bf01c96f6e562ee97f395233f8"},{"id":"s10","url":"https://arxiv.org/abs/2605.20530","title":"Mazaheri et al. (2026), \"AgentAtlas: Beyond Outcome Leaderboards for LLM Agents\", arXiv:2605.20530","quote":"AgentAtlas reframes agent evaluation as a diagnostic vocabulary and audit protocol for separating outcome success from control-decision quality and trajectory quality.","claim_ids":[],"hash":"04d1a3d16e4451ba948fa20066fb9f6ae6314744c657a8b04c9d0249f50169ab"},{"id":"s11","url":"https://arxiv.org/abs/2603.20953","title":"Uchibeke (2026), \"Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents\", arXiv:2603.20953","quote":"AI agents today have passwords but no permission slips.","claim_ids":[],"hash":"231356710a6fe21dc997216915da757fc19b2b04a999114ab51781a73b48e87f"},{"id":"s11b","url":"https://arxiv.org/abs/2603.20953","title":"Uchibeke (2026), \"Before the Tool Call\" — the adversarial testbed result, arXiv:2603.20953","quote":"social engineering succeeded against the model 74.6% of the time under a permissive policy; under a restrictive OAP policy, a comparable population of attackers achieved a 0% success rate across 879 attempts","claim_ids":[],"hash":"522d8ff2dbe744e1d81b4fa894a8019b4ce530575d8edb480c806efe20e36973"},{"id":"s12","url":"https://arxiv.org/abs/2604.03733","title":"Zhang et al. (2026), \"SoK: Blockchain Agent-to-Agent Payments\", arXiv:2604.03733","quote":"we systematize blockchain-based A2A payments, e.g., X402, with a four-stage lifecycle: discovery, authorization, execution, and accounting. We categorize representative designs at each stage and identify key challenges, including weak intent binding, misuse under valid authorization, payment-service decoupling, and limited accountability.","claim_ids":[],"hash":"2403b314c1604b8438d98e7278a4e49f7cb61552c091f4a0271340833e615fbb"},{"id":"s13","url":"https://arxiv.org/abs/2607.00245","title":"Gong (2026), \"Agent-to-Agent Finance: Blockchain Payments and Trust Infrastructure for Autonomous AI Agents\", arXiv:2607.00245","quote":"It argues that the decisive design question is bounded autonomy: how to let agents transact without making markets more opaque, fragile or unaccountable.","claim_ids":[],"hash":"a065fd9edffe703858cc64404f6e667ab5771ce4a88e9768df6cca12614b815c"},{"id":"s14","url":"https://arxiv.org/abs/2604.24890","title":"Golaszewski et al. (2026), \"Verifying Provenance of Digital Media: Why the C2PA Specifications Fall Short\", arXiv:2604.24890","quote":"We find that the current C2PA specifications fail to achieve their claimed security goals.","claim_ids":[],"hash":"659bd188456e016b9a1ddcb26a32ca0b66d7ce2bb6565b760d7094086b1ba02d"},{"id":"s15","url":"https://arxiv.org/abs/2502.18535","title":"Peng et al. (2025), \"A Survey of Zero-Knowledge Proof Based Verifiable Machine Learning\", arXiv:2502.18535","quote":"analyze the main implementation bottlenecks, including limited circuit expressiveness, high proving cost, and deployment complexity","claim_ids":[],"hash":"a87ad18c026ee90ca2439c9b7897e48e22b7e3b459c2d4796e04a6e54aed69c3"},{"id":"s16","url":"https://arxiv.org/abs/2606.04990","title":"Wang et al. (2026), \"From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents\", arXiv:2606.04990","quote":"Final-answer accuracy alone cannot explain how an output was produced, which evidence supported each claim, whether tool calls were justified, how memory influenced later decisions, or where failures originated.","claim_ids":[],"hash":"0517f3a1ab8abd356235aaa8d1234cf0dfe05cfa13b271fe403829e276906232"},{"id":"s17","url":"https://arxiv.org/abs/2508.02866","title":"Souza et al. (2025), \"PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic Workflows\", arXiv:2508.02866","quote":"existing methods fail to capture and relate agent-centric metadata such as prompts, responses, and decisions with the broader workflow context and downstream outcomes","claim_ids":[],"hash":"6a416e5e4f60599aae57ac97ddd52b9a026d48b99b3d88506f249dd1b2aaa523"},{"id":"s18","url":"https://arxiv.org/abs/1407.3561","title":"Benet (2014), \"IPFS - Content Addressed, Versioned, P2P File System\", arXiv:1407.3561","quote":"IPFS provides a high throughput content-addressed block storage model, with content-addressed hyper links. This forms a generalized Merkle DAG, a data structure upon which one can build versioned file systems, blockchains, and even a Permanent Web.","claim_ids":[],"hash":"ab6e03981e4c5a9b74afb65ef26eb63e3a6df00df5309ff31ef88c797f19db22"},{"id":"s19","url":"https://arxiv.org/abs/2601.07023","title":"Hu et al. (2026), \"CloneMem: Benchmarking Long-Term Memory for AI Clones\", arXiv:2601.07023","quote":"Experiments show that current memory mechanisms struggle in this setting, highlighting open challenges for life-grounded personalized AI.","claim_ids":[],"hash":"17f1092de4efa588f0db40325d32d7b498067bb651de0a79e7a30fb2756f9ca7"},{"id":"s20","url":"https://arxiv.org/abs/2402.07922","title":"Lauer-Schmaltz et al. (2024), \"Towards the Human Digital Twin: Definition and Design -- A survey\", arXiv:2402.07922","quote":"This has introduced several significant challenges, including ambiguity in the definition of HDTs and a lack of guidance for their design.","claim_ids":[],"hash":"6abc86205860f59013fdb6f7db022d5c1a3ee28453b04996ae273553542c5dc3"},{"id":"s21","url":"https://arxiv.org/abs/2604.16338","title":"Acharya (2026), \"Governing the Agentic Enterprise: A Governance Maturity Model for Managing AI Agent Sprawl in Business Operations\", arXiv:2604.16338","quote":"Industry surveys report that only 21% of enterprises have mature governance models for autonomous agents, while 40% of agentic AI projects are projected to fail by 2027 due to inadequate governance and risk controls.","claim_ids":[],"hash":"a1359728c3522ac7f5bcef4ae6b0dfd62dca23565e225272bfdb295aa707121c"},{"id":"s22","url":"https://arxiv.org/abs/2604.00555","title":"Tuan et al. (2026), \"Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems\", arXiv:2604.00555","quote":"current enterprise systems constrain agent inputs (context assembly, tool discovery, governance thresholds) but not outputs, and we propose mechanisms extending this coupling to output-side validation (response checking, reasoning verification, compliance enforcement)","claim_ids":[],"hash":"062d4fc21ee682fa1d395eefde377886c51570fad8b0fe84753e7439d7e60338"},{"id":"s25","url":"https://arxiv.org/abs/2304.03442","title":"Park et al. (2023), \"Generative Agents: Interactive Simulacra of Human Behavior\", arXiv:2304.03442","quote":"an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior","claim_ids":[],"hash":"fd2003de5cfc460975e5c3478c391ee147463c6ce6e5ef26bb495b0a99532c93"},{"id":"s26","url":"https://arxiv.org/abs/2511.19863","title":"Bengio et al. (2025), \"International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management\", arXiv:2511.19863","quote":"the number of companies publishing Frontier AI Safety Frameworks more than doubled in 2025, and governments and international organisations have established a small number of governance frameworks for general-purpose AI, focusing largely on transparency and risk assessment","claim_ids":[],"hash":"433a525961938a09dffe40c68c6e58f5af40a1c16ba0800c638973bc2ec8366b"},{"id":"s27","url":"https://arxiv.org/abs/2509.24380","title":"Deng et al. (2025), \"Agentic Services Computing\", arXiv:2509.24380","quote":"How can goal-driven, stateful, tool-mediated, and accountable autonomous behavior be engineered and managed as a service?","claim_ids":[],"hash":"95735175cd08329fe2badd3b48e21842c3511fbd358d40b3ebc8817be5e0af9a"},{"id":"s28","url":"https://arxiv.org/abs/2605.09721","title":"Goel (2026), \"Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments\", arXiv:2605.09721","quote":"many risks in autonomous cloud agents arise not from novel vulnerabilities, but from over-privileged tools, capability-intent mismatches, and ambient authority leakage in execution environments","claim_ids":[],"hash":"9f9d5d2564fc790d8dd8a660e7f2dff7fdc21a92b3ea512feafe6f7461f5e509"},{"id":"s29","url":"https://arxiv.org/abs/2604.23280","title":"Otsuka et al. (2026), \"AI Identity: Standards, Gaps, and Research Directions for AI Agents\", arXiv:2604.23280","quote":"an evaluation of current technical and regulatory documents against the identity requirements of autonomous agents, finding that none adequately address the challenge","claim_ids":[],"hash":"e945b5c62017870f4b22cfa02ef8df42c4ddeeedf8f14b2f5524785fad792b67"},{"id":"s30","url":"https://arxiv.org/abs/2603.14312","title":"Wang et al. (2026), \"Autonomous Agents Coordinating Distributed Discovery Through Emergent Artifact Exchange\", arXiv:2603.14312","quote":"Agents select and chain tools based on their scientific profiles, produce immutable artifacts with typed metadata and parent lineage, and broadcast unsatisfied information needs to a shared global index.","claim_ids":[],"hash":"3db4b359918632ca550ede1eba05388031a1889732c22cb02c0ef46dec704c80"},{"id":"s31","url":"https://arxiv.org/abs/2602.13855","title":"Rasheed et al. (2026), \"From Fluent to Verifiable: Claim-Level Auditability for Deep Research Agents\", arXiv:2602.13855","quote":"as research generation becomes cheap, auditability becomes the bottleneck, and the dominant risk shifts from isolated factual errors to scientifically styled outputs whose claim-evidence links are weak, missing, or misleading","claim_ids":[],"hash":"1f484672a755dd3347b8ab6a57f59a397423acb99d7a75fd4d628f04c028200c"},{"id":"s32","url":"https://arxiv.org/abs/2603.10060","title":"Basu (2026), \"Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents\", arXiv:2603.10060","quote":"NabaOS detects 94.2% of fabricated tool references, 87.6% of count misstatements, and 91.3% of false absence claims, with <15ms verification overhead per response.","claim_ids":[],"hash":"629271c53f9a93d7d3f48cf9c40c906b49e97fb6313c0f973778d9fb576da86d"},{"id":"s33","url":"https://miscsubjects.com/api/metrics/structure","title":"This build, live: the structure metric endpoint (objects, typed relationships, capabilities)","quote":"One mind, measured as a live structure: 1189 objects · 12656 claims · 11653 typed relationships · 850 executable capabilities · 11 representation types · 5 meta-layers · 167 active threads","claim_ids":[],"hash":"99a1a60538541122898054144e3baf9204414846d294857c39526b33a4d5d85b"},{"id":"s34","url":"https://miscsubjects.com/api/metrics/grounding","title":"This build, live: the grounding metric endpoint (claims, sources, and the fraction carrying a source)","quote":"\"claims_total\": 12656, \"sources_total\": 10054, \"claims_with_sources_fraction\": 0.816","claim_ids":[],"hash":"d404bbba0abaccfa79fcff03072c1a2a5349384e8a16e65666f040151b6b1169"},{"id":"s35","url":"https://miscsubjects.com/api/dispatch?registry=1","title":"This build, live: the public capability registry, keyless, every row carrying its risk grade and approval requirement","quote":"\"protocol\": \"OIP\", \"version\": \"1.2.0\", \"count\": 932","claim_ids":[],"hash":"b6289f38c5f7d43d2f1242ca5789f3a72449ebb4df34cac1814d76874615a0e2"},{"id":"s36","url":"https://miscsubjects.com/api/work","title":"This build, live: the work object — the hash-chained task ledger whose acceptance tests, not an agent's claim, decide completion","quote":"work_actions is hash-chained; nothing is updated or deleted. A correction appends a revision that names what it supersedes.","claim_ids":[],"hash":"6c8b4c7951c245f25990f942498f259019b07078e61bfbe85a75d06fc64dbba3"},{"id":"s37","url":"https://miscsubjects.com/a/logical-economics","title":"This build: the measured undetected-wrong rate per panel configuration, and the 0.071 floor","quote":"The floor is one item. Beyond two channels the best achievable rate stops improving","claim_ids":[],"hash":"f83633cf9c27e2d9a21bfbc1b39899cf56751bc965aaec5a0d6f455179d985bf"},{"id":"s38","url":"https://miscsubjects.com/a/the-build-end-to-end","title":"This build: the end-to-end record of what exists, including the known-defects and roadmap sections this scorecard scores against","quote":"Nobody has yet been taken from a blank questionnaire to a running, separately owned instance in one pass.","claim_ids":[],"hash":"7e9ec6a0f181aa0fd78e6572c8fefb6ce61b694c0ecb08f474075a7d31cca847"}],"anecdotal_sources":[],"scientific_sources":[],"user_reports":[],"related_articles":[],"question_graph":{"slug":"theoretical-limits","questions":[],"evidence":[],"edges":[],"counts":{"questions":0,"evidence":0,"edges":0}},"honesty":{"active_claims":23,"retracted_claims":0,"cut_claims":0,"challenges":0,"scrub_events":0,"note":"Retracted/cut claims stay on ledger but are excluded from ask unless ?include_inactive=1"},"counts":{"claims":23,"claims_total":23,"sources":37,"anecdotal":0,"scientific":0,"user_reports":0,"questions":0,"evidence_ingests":0}}