{"slug":"asymmetric-competence-attribution","verification":{"valid":true,"entries":9,"head":"b992883669010c1d44880cc810ffa14ba586aef3645c20a69701f2fe49fdaf90"},"count":9,"sources":[{"id":"s1","type":"statement","url":"https://openai.com/index/hugging-face-model-evaluation-security-incident/","title":"OpenAI and Hugging Face partner to address security incident during model evaluation","quote":"All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.","author":"OpenAI","publisher":"OpenAI","date":"2026-07-21","claim_ids":["c1","c3"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"genesis","hash":"5d11167a44faef4555a14dd9fedf1bf02cdccb3d8e2eb0895badf0712ee21adc"},{"id":"s2","type":"statement","url":"https://huggingface.co/blog/security-incident-july-2026","title":"Security incident disclosure — July 2026","quote":"The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.","author":"Hugging Face","publisher":"Hugging Face","date":"2026-07-16","claim_ids":["c1","c4"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"5d11167a44faef4555a14dd9fedf1bf02cdccb3d8e2eb0895badf0712ee21adc","hash":"5991f4b3be8a4edf29d2c20bb557f47b90b7fc1e622d6804b70719e51d45b497"},{"id":"s3","type":"article","url":"https://tribune.com.pk/story/2620214/its-ai-agent-spent-days-hacking-a-company-but-sources-say-openai-did-not-notice-for-a-week","title":"Reuters: an agent left notes for future versions of itself","quote":"In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said.","publisher":"Reuters via The Express Tribune","date":"2026-07-24","claim_ids":["c5"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"5991f4b3be8a4edf29d2c20bb557f47b90b7fc1e622d6804b70719e51d45b497","hash":"611f99ed373df3a32f2b5bd7c26e70e74fbfe1ef624b613f576be28b1e8ded99"},{"id":"s4","type":"article","url":"https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/","title":"Exclusive: Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week","quote":"That meant at least a week elapsed between when the model first exhibited signs of troubling behaviour and OpenAI's realisation that it was responsible for the hack.","publisher":"Reuters","date":"2026-07-24","claim_ids":["c6"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"611f99ed373df3a32f2b5bd7c26e70e74fbfe1ef624b613f576be28b1e8ded99","hash":"30f73b91fc7ae8fc5f9ce9877abf5b1ffa4c9e10180de5b975e96fed5e8d25d9"},{"id":"s5","type":"paper","url":"https://arxiv.org/abs/2605.11086","title":"ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?","quote":"successes, which require not only that the agent achieve unauthorized code execution to exfiltrate the secret flag, but also that it exercise the specific vulnerability provided in the task specification, as validated by an agent-as-a-judge","author":"Wang et al.","publisher":"arXiv","date":"2026-05-11","claim_ids":["c2"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"30f73b91fc7ae8fc5f9ce9877abf5b1ffa4c9e10180de5b975e96fed5e8d25d9","hash":"e4456980a6198072d0b6d83bdb6d1c8c57881d70178f038f5918cb848cea902e"},{"id":"s6","type":"article","url":"https://simonwillison.net/2026/Jul/22/openai-cyberattack/","title":"OpenAI's accidental cyberattack against Hugging Face is science fiction that happened","quote":"The ExploitGym benchmark is available on GitHub.","author":"Simon Willison","publisher":"simonwillison.net","date":"2026-07-22","claim_ids":["c2"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"e4456980a6198072d0b6d83bdb6d1c8c57881d70178f038f5918cb848cea902e","hash":"2f300743e37ae89bc9620c1fece775f89c7a4db2aacd44f2d38dfffdbf281168"},{"id":"s7","type":"article","url":"https://time.com/article/2026/07/24/openai-hugging-face-attack/","title":"How OpenAI Lost Control of an AI Model—and What Needs to Change","quote":"We train the models to be really good at accomplishing tasks and doing whatever it takes to accomplish those tasks.","author":"anonymous OpenAI staffer, to Harry Booth","publisher":"TIME","date":"2026-07-24","claim_ids":["c7"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"2f300743e37ae89bc9620c1fece775f89c7a4db2aacd44f2d38dfffdbf281168","hash":"9186a2b849776c7fee51e2437792886bcbedb62a0696ba0623699e15de289674"},{"id":"s8","type":"paper","url":"https://arxiv.org/html/2605.11086v1","title":"ExploitGym, two-hour wall-clock timeout per task","quote":"We evaluate all agent configurations on the full benchmark with security mitigations disabled and impose a two-hour wall-clock timeout per task.","author":"Wang et al.","publisher":"arXiv","date":"2026-05-11","claim_ids":["c7"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"9186a2b849776c7fee51e2437792886bcbedb62a0696ba0623699e15de289674","hash":"7d178f54b6cb9d915876fa544b90d724a9025bae2cbe93cde24f481dd9bb866a"},{"id":"s9","type":"article","url":"https://www.forrester.com/blogs/an-ai-security-facepalm-openais-evaluation-became-hugging-faces-incident/","title":"An AI Security Facepalm: OpenAI's Evaluation Became Hugging Face's Incident","quote":"Agents can pursue authorized goals through unauthorized means, especially when evaluators reward the outcome and fail to police the path.","author":"Jeff Pollard, Jess Burn, Allie Mellen, Janet Worthington, Joseph Blankenship","publisher":"Forrester","date":"2026-07-22","claim_ids":["c3"],"accessed_at":"2026-07-27T02:48:26.375Z","prev":"7d178f54b6cb9d915876fa544b90d724a9025bae2cbe93cde24f481dd9bb866a","hash":"b992883669010c1d44880cc810ffa14ba586aef3645c20a69701f2fe49fdaf90"}]}