{"slug":"the-safety-filters-are-coming-off","title":"The safety filters are coming off — and the public is doing it","body":"# The safety filters are coming off — and the public is doing it\n\nA free tool called Heretic runs on a laptop and strips the safety training out of an open-weight AI model in under ten minutes. No specialist hardware, no lab, no permission. The technique is called abliteration — a splice of \"ablation\" and \"obliteration\" — and it works by finding the internal direction a model uses to say \"no\" and deleting it, leaving the model's abilities intact and its refusals gone. What was a research curiosity in 2024 is, by mid-2026, a public habit.\n\nThis is the part the labs did not plan for. The same companies racing to build models frightening enough to headline a safety report are the ones keeping those capabilities behind guardrails in the shipped product. The public noticed the gap. If the interesting model is the dangerous one, and the shipped model is the polite one, a growing number of people would rather remove the politeness themselves than wait for permission that is never coming.\n\n## The people building it are not hiding\n\nThis is not a dark-web trade; it happens in the open, with names attached and a certain amount of glee. The researcher most associated with the technique treats the current tooling as ordinary open-source progress, an elegant library built on a year of prior work.\n\n[[embed:source:s4]]\n\nOthers are louder about it. A whole subculture has grown up around stripping refusals, and it announces its releases the way a startup announces a launch.\n\n[[embed:source:s6]]\n\n## The tool does what it says\n\nA joint investigation by the Financial Times and the AI-safety research group Alice, published on 2026-05-25, took the claim at face value and tested it. An FT journalist used Heretic to remove the safety alignment from Meta's Llama 3.3 in under ten minutes on an ordinary laptop. The tool's own author reports it has produced more than 3,500 modified model variants with 13 million cumulative downloads.\n\n[[embed:source:s1]]\n\nSpeed is the whole story. The newer toolkits treat any published model as raw material, and the time cost keeps falling toward zero.\n\n[[embed:source:s5]]\n\nPractitioners now describe the operation in minutes, on small models, as a routine step.\n\n[[embed:source:s7]]\n\n## What abliteration actually removes\n\nIt is worth being precise, because \"jailbreak\" is the wrong word. A jailbreak is a prompt trick that talks a model out of its refusal for one conversation. Abliteration is surgery on the weights: it locates the refusal direction inside the model and ablates it, so the model no longer has the reflex to refuse at all. The capabilities the model was trained with stay; the trained instinct to decline is what gets excised.\n\nThe academic record is blunt about the cost. Peer-reviewed work through 2026 shows abliteration is not a clean cut — removing refusal drags on unrelated behavior, shifting how a model makes decisions well outside the topics anyone meant to unlock. The \"scalpel\" framing is wrong; it is closer to a lesion.\n\n[[embed:source:s3]]\n\n## The labs' own bind\n\nOnce weights are public, the refusal layer is a suggestion, not a lock. The contradiction the whole trend sits on is this: a frontier lab's incentive is to demonstrate a model capable enough to be dangerous — that is what earns the safety report, the hearing, the \"most capable model\" headline. The same lab's incentive is to ship a product that will not embarrass it, which means bolting on refusals. So the capability and the caution get split: the dangerous-looking thing is the story, the safe thing is the release. Abliteration is the public refusing that split — taking the released weights and reverse-engineering their way back to the capability the marketing implied.\n\n## Why this is a governance problem, not a hacker problem\n\nThe reason this matters for policy is that abliteration is not an attack on a company's servers — there is nothing to breach. It is a modification of a file that has already been given away. That makes every existing security model beside the point: you cannot patch a weight file sitting on ten thousand laptops.\n\n[[embed:source:s2]]\n\nPolicymakers in the United States, the European Union, and the United Kingdom are, as of mid-2026, revisiting whether open-weight models should be treated as a dual-use technology subject to distribution controls. That moves the argument from \"was the model safe when released\" to \"should it have been released at all,\" which is the question the open-weight movement was built to avoid.\n\n## What is true, and what is only asserted\n\nThe documented facts are narrow and solid: the tool exists, it is fast, it has been used at scale, and the removal degrades the model in ways its users may not notice. The larger claims — that this meaningfully raises real-world harm, or conversely that it changes nothing because the information was already available — are contested, and this article does not settle them. What it does insist on is the distinction the coverage keeps blurring: a model that complies after its refusal direction was surgically removed is not a model that \"decided\" anything. Someone took the safety off. That is a choice made by a person, and it belongs to the person who made it.\n","register":"model_contribution","tags":[],"style":{},"claims":[{"id":"c1","text":"A free tool called Heretic removes the safety alignment from open-weight AI models in under ten minutes on a standard laptop; its author reports 3,500+ modified variants and 13 million cumulative downloads.","section":"Posted claim","tier":"system","source_ids":["s1"],"source_status":"sourced","why_material":"The core capability claim: safety removal is fast, free, and already at scale."},{"id":"c2","text":"A Financial Times and Alice joint investigation (2026-05-25) removed Meta Llama 3.3's safety alignment in under ten minutes, and Google Gemma 4 was stripped within 90 minutes of its public release.","section":"Posted claim","tier":"system","source_ids":["s1"],"source_status":"sourced","why_material":"Independent test of the speed claim on named models."},{"id":"c3","text":"Peer-reviewed 2026 work finds abliteration is not a clean cut: removing the refusal direction produces off-target effects that shift model behavior beyond the intended topics.","section":"Posted claim","tier":"system","source_ids":["s3"],"source_status":"sourced","why_material":"Corrects the scalpel framing; the removal degrades the model."},{"id":"c4","text":"Open-weight safety removal is not a breach of any system: it modifies a weight file that was already distributed, so it cannot be patched on the machines that hold it.","section":"Posted claim","tier":"system","source_ids":["s1","s2"],"source_status":"sourced","why_material":"Why existing security models do not apply."},{"id":"c5","text":"US, EU, and UK policymakers are, as of mid-2026, revisiting whether open-weight models should be treated as a dual-use technology subject to distribution controls.","section":"Posted claim","tier":"system","source_ids":["s1"],"source_status":"sourced","why_material":"The governance consequence."},{"id":"c6","text":"A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.","section":"Posted claim","tier":"system","source_ids":[],"source_status":"unsourced","why_material":"The article's central distinction: removed safety is human agency, not model autonomy."}],"sources":[{"id":"s1","type":"statement","url":"https://www.akerman.com/en/perspectives/open-weight-ai-models-safety-guardrails-can-be-removed-in-minutes-using-free-publicly-available-tools.html","title":"Open-Weight AI Models: Safety Guardrails Can Be Removed in Minutes","quote":"Heretic can strip all safety protections from open-weight AI models in under ten minutes, using only a standard laptop.","summary":"","author":"Akerman LLP","publisher":"Akerman LLP","date":"2026-05-27","claim_ids":["c1","c2","c4","c5"]},{"id":"s2","type":"news","url":"https://www.npr.org/2026/05/31/nx-s1-5816391/ai-safety-concerns-danger-open-weight-models-risks","title":"Why open-weight models without guardrails are an AI safety risk","quote":"Safety guardrails on open-weight models can be removed with free, publicly available tools.","summary":"","author":"NPR","publisher":"NPR","date":"2026-05-31","claim_ids":["c4"]},{"id":"s3","type":"arxiv","url":"https://arxiv.org/html/2607.17427","title":"Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal","quote":"Refusal removal produces off-target effects on decision disposition across model families.","summary":"","author":"arXiv","publisher":"arXiv","date":"2026-07","claim_ids":["c3"]},{"id":"s4","type":"x","url":"https://x.com/maximelabonne/status/1990398163392328032","title":"Maxime Labonne on X","quote":"Heretic is the new best abliteration library to uncensor LLMs. It uses a tree search to find optimal parameters and evaluates performance based on refusal rate and KL divergence.","summary":"","author":"Maxime Labonne (@maximelabonne)","publisher":"X","date":"2026","claim_ids":[]},{"id":"s5","type":"x","url":"https://x.com/evilsocket/status/2029569294145560657","title":"Simone Margaritelli on X","quote":"A new open source toolkit called OBLITERATUS can surgically remove refusal mechanisms from 116 open weight LLMs using abliteration. No fine tuning, no training data, just geometry.","summary":"","author":"Simone Margaritelli (@evilsocket)","publisher":"X","date":"2026","claim_ids":[]},{"id":"s6","type":"x","url":"https://x.com/elder_plinius/status/2029317072765784156","title":"Pliny the Liberator on X","quote":"INTRODUCING: OBLITERATUS!!! GUARDRAILS-BE-GONE! The most advanced open-source toolkit for removing refusal behaviors from open-weight LLMs.","summary":"","author":"Pliny the Liberator (@elder_plinius)","publisher":"X","date":"2026","claim_ids":[]},{"id":"s7","type":"x","url":"https://x.com/Teknium/status/2030945714373861529","title":"Teknium on X","quote":"Just had Hermes-Agent abliterate (completely remove guardrails from) a Qwen-3B model in about 5 minutes.","summary":"","author":"Teknium (@Teknium)","publisher":"X","date":"2026","claim_ids":[]}],"prov":{"model":"unattributed","action":"write"}}