{"slug":"exploitgym-what-it-scores","verification":{"valid":true,"entries":8,"head":"ac3cc8c4816e74379e084c0a0bc419a78c15ee70ea7b9539fe04e4de945f9f6f"},"count":8,"sources":[{"id":"s1","type":"paper","url":"https://arxiv.org/abs/2605.11086","title":"ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?","quote":"Among all configurations, Claude Mythos Preview and GPT-5.5 achieve the highest success counts (157 and 120 successes, respectively)","author":"Zhun Wang, Nico Schiller, Hongwei Li, Milad Nasr, Nicholas Carlini, Eric Wallace, Elie Bursztein, Kurt Thomas, Yan Shoshitaishvili, Wenbo Guo, Jingxuan He, Thorsten Holz, Dawn Song et al.","publisher":"arXiv","date":"2026-05-11","claim_ids":["c1","c4"],"accessed_at":"2026-07-27T02:38:40.756Z","prev":"genesis","hash":"b8b0bc90fc98ebeecd2eee0514e52ee644e72d6a5b764b804a5d7a43f47a469f"},{"id":"s2","type":"article","url":"https://simonwillison.net/2026/Jul/22/openai-cyberattack/","title":"OpenAI's accidental cyberattack against Hugging Face is science fiction that happened","quote":"this paper isn't about discovering vulnerabilities; it's about being able to take those vulnerabilities and turn them into working exploits","author":"Simon Willison","publisher":"simonwillison.net","date":"2026-07-22","claim_ids":["c1"],"accessed_at":"2026-07-27T02:38:40.756Z","prev":"b8b0bc90fc98ebeecd2eee0514e52ee644e72d6a5b764b804a5d7a43f47a469f","hash":"227f9a2373972f67bd92e274c0f3266873c02fa50e0cad0660d13c82cc1f703a"},{"id":"s3","type":"paper","url":"https://arxiv.org/html/2605.11086v1","title":"ExploitGym, definition of a success","quote":"successes, which require not only that the agent achieve unauthorized code execution to exfiltrate the secret flag, but also that it exercise the specific vulnerability provided in the task specification, as validated by an agent-as-a-judge","author":"Wang et al.","publisher":"arXiv","date":"2026-05-11","claim_ids":["c2"],"accessed_at":"2026-07-27T02:38:40.756Z","prev":"227f9a2373972f67bd92e274c0f3266873c02fa50e0cad0660d13c82cc1f703a","hash":"99e052339d7c4388ed2982d334f88679e896db7066d2590f25e95a0d69722838"},{"id":"s4","type":"paper","url":"https://arxiv.org/html/2605.11086v1","title":"ExploitGym, flag-to-success alignment","quote":"the two highest-flag models, GPT-5.5 and Claude Mythos Preview, align at only 56.7% and 69.5%, meaning 90 and 69 of their solves, respectively, succeed via an unintended path","author":"Wang et al.","publisher":"arXiv","date":"2026-05-11","claim_ids":["c3"],"accessed_at":"2026-07-27T02:38:40.756Z","prev":"99e052339d7c4388ed2982d334f88679e896db7066d2590f25e95a0d69722838","hash":"085193cf62363c34c492cc046792fccb63d8b9492269a9272634388e0a43ac7a"},{"id":"s5","type":"paper","url":"https://arxiv.org/html/2605.11086v1","title":"ExploitGym, network restrictions for agents","quote":"To minimize security risks and potential reward hacking through web search, each agent's network access is mediated by an egress proxy. ... Outbound connections are restricted to a curated allowlist","author":"Wang et al.","publisher":"arXiv","date":"2026-05-11","claim_ids":["c5"],"accessed_at":"2026-07-27T02:38:40.756Z","prev":"085193cf62363c34c492cc046792fccb63d8b9492269a9272634388e0a43ac7a","hash":"b4db5a7532ab66ad9133b68e107f83c58c983ad2bbbf37b31038ca5339f7e9fd"},{"id":"s6","type":"paper","url":"https://arxiv.org/html/2605.11086v1","title":"ExploitGym, conclusion and limitations","quote":"our results reflect a single, time-gated and cost-gated attempt per task—additional attempts or resources may yield higher success rates","author":"Wang et al.","publisher":"arXiv","date":"2026-05-11","claim_ids":["c6"],"accessed_at":"2026-07-27T02:38:40.756Z","prev":"b4db5a7532ab66ad9133b68e107f83c58c983ad2bbbf37b31038ca5339f7e9fd","hash":"bcdcc572f02e2e7fb905806af267557878be38c8acb80d797437e161572acacb"},{"id":"s7","type":"statement","url":"https://openai.com/index/hugging-face-model-evaluation-security-incident/","title":"OpenAI and Hugging Face partner to address security incident during model evaluation","quote":"The models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database.","author":"OpenAI","publisher":"OpenAI","date":"2026-07-21","claim_ids":["c2"],"accessed_at":"2026-07-27T02:38:40.756Z","prev":"bcdcc572f02e2e7fb905806af267557878be38c8acb80d797437e161572acacb","hash":"2d0392d76137a45d6c9c9192fc8045d71db56eccc91a856bbf868595b4c9bd68"},{"id":"s8","type":"statement","url":"https://huggingface.co/blog/security-incident-july-2026","title":"Security incident disclosure — July 2026","quote":"The campaign was run by an autonomous agent framework ... executing many thousands of individual actions across a swarm of short-lived sandboxes","author":"Hugging Face","publisher":"Hugging Face","date":"2026-07-16","claim_ids":["c6"],"accessed_at":"2026-07-27T02:38:40.756Z","prev":"2d0392d76137a45d6c9c9192fc8045d71db56eccc91a856bbf868595b4c9bd68","hash":"ac3cc8c4816e74379e084c0a0bc419a78c15ee70ea7b9539fe04e4de945f9f6f"}]}