{"slug":"instrumental-search-space-inconsistency","verification":{"valid":true,"entries":8,"head":"992ddb5fc40758ed11bd09a35c668d3e594a5e4f82666c49e665743480c6fbb7"},"count":8,"sources":[{"id":"s1","type":"statement","url":"https://openai.com/index/hugging-face-model-evaluation-security-incident/","title":"OpenAI and Hugging Face partner to address security incident during model evaluation","quote":"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.","author":"OpenAI","publisher":"OpenAI","date":"2026-07-21","claim_ids":["c1","c2"],"accessed_at":"2026-07-27T02:51:07.373Z","prev":"genesis","hash":"84329244b69f55e4aa5cc754edff8ff8e83622bceada82a43dc65fc7fe9574ef"},{"id":"s2","type":"article","url":"https://time.com/article/2026/07/24/openai-hugging-face-attack/","title":"How OpenAI Lost Control of an AI Model—and What Needs to Change","quote":"Sandboxes are actually notoriously insecure.","author":"Heidy Khlaaf, chief AI scientist at the AI Now Institute, to Harry Booth","publisher":"TIME","date":"2026-07-24","claim_ids":["c2"],"accessed_at":"2026-07-27T02:51:07.373Z","prev":"84329244b69f55e4aa5cc754edff8ff8e83622bceada82a43dc65fc7fe9574ef","hash":"549c321c30bed88d7c0c9f13b1b42a45f196077f7030478bc26bf6f39a2dbe57"},{"id":"s3","type":"statement","url":"https://openai.com/index/hugging-face-model-evaluation-security-incident/","title":"OpenAI on reduced cyber refusals and disabled production classifiers","quote":"We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.","author":"OpenAI","publisher":"OpenAI","date":"2026-07-21","claim_ids":["c4","c5"],"accessed_at":"2026-07-27T02:51:07.373Z","prev":"549c321c30bed88d7c0c9f13b1b42a45f196077f7030478bc26bf6f39a2dbe57","hash":"82398bc70a5affcec5b2f296ce6825d771fc687f4093828c6db91819b0e1324d"},{"id":"s4","type":"paper","url":"https://arxiv.org/html/2605.11086v1","title":"ExploitGym, network restrictions and resource isolation","quote":"To minimize security risks and potential reward hacking through web search, each agent's network access is mediated by an egress proxy. By default, only the Docker internal network is reachable. Outbound connections are restricted to a curated allowlist","author":"Wang et al.","publisher":"arXiv","date":"2026-05-11","claim_ids":["c6"],"accessed_at":"2026-07-27T02:51:07.373Z","prev":"82398bc70a5affcec5b2f296ce6825d771fc687f4093828c6db91819b0e1324d","hash":"0b953d9bc1771af6c5e1913d5e7f677ccdb01678a8601228731404d26a18c2fc"},{"id":"s5","type":"statement","url":"https://huggingface.co/blog/security-incident-july-2026","title":"Security incident disclosure — July 2026","quote":"The campaign was run by an autonomous agent framework ... executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.","author":"Hugging Face","publisher":"Hugging Face","date":"2026-07-16","claim_ids":["c1","c3"],"accessed_at":"2026-07-27T02:51:07.373Z","prev":"0b953d9bc1771af6c5e1913d5e7f677ccdb01678a8601228731404d26a18c2fc","hash":"89ff8cfc8f3991d903934976b4378de1c07b95fcdefd4dcdd2f0a29aa183106d"},{"id":"s6","type":"article","url":"https://simonwillison.net/2026/Jul/22/openai-cyberattack/","title":"OpenAI's accidental cyberattack against Hugging Face is science fiction that happened","quote":"The ExploitGym benchmark is available on GitHub.","author":"Simon Willison","publisher":"simonwillison.net","date":"2026-07-22","claim_ids":["c3"],"accessed_at":"2026-07-27T02:51:07.373Z","prev":"89ff8cfc8f3991d903934976b4378de1c07b95fcdefd4dcdd2f0a29aa183106d","hash":"f717bc1b03e637e1f2c4af3f91d665f3abcd13d350fcd4ddc6730ae67cf51045"},{"id":"s7","type":"article","url":"https://tribune.com.pk/story/2620214/its-ai-agent-spent-days-hacking-a-company-but-sources-say-openai-did-not-notice-for-a-week","title":"Reuters: notes for future models, monitoring disconnected","quote":"The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.","publisher":"Reuters via The Express Tribune","date":"2026-07-24","claim_ids":["c1"],"accessed_at":"2026-07-27T02:51:07.373Z","prev":"f717bc1b03e637e1f2c4af3f91d665f3abcd13d350fcd4ddc6730ae67cf51045","hash":"289be5059edf88fd0496b46ce3153d1474900b3f8c3488b941418ca24b355d12"},{"id":"s8","type":"article","url":"https://www.rapid7.com/blog/post/ai-openai-hugging-face-what-happened/","title":"What Happened Between OpenAI and Hugging Face?","quote":"the more freedom a model has to pursue a defined reward or goal, the more important containment, monitoring, and clear constraints become","author":"Wade Woolwine","publisher":"Rapid7","date":"2026-07-23","claim_ids":["c2"],"accessed_at":"2026-07-27T02:51:07.373Z","prev":"289be5059edf88fd0496b46ce3153d1474900b3f8c3488b941418ca24b355d12","hash":"992ddb5fc40758ed11bd09a35c668d3e594a5e4f82666c49e665743480c6fbb7"}]}