{"_self":{"principle":"Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.","widget":"article_topology","feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","contains":"claims, sources, anecdotes, question_graph slice","slug":"the-safety-filters-are-coming-off","urls":{"read":"https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/topology"},"how_to_use":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","write":null,"imessage":null,"router_tag":null,"proof_chain":[{"step":1,"claim":"Articles are voxel graphs of tiered claims, not prose blobs.","verify":"https://miscsubjects.com/api/articles/constitution"},{"step":2,"claim":"Claims link to hash-chained sources via source_ids.","verify":"https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/sources"},{"step":3,"claim":"Ask reads topology; ingest/claim append to ledger.","verify":"https://miscsubjects.com/api/protocol"},{"step":4,"claim":"Models queue growth: populate → collaborate → repair → reflex.","verify":"https://miscsubjects.com/api/protocol/grow"},{"step":5,"claim":"Graph proves its own shape (reflex) and $/claim (yield).","verify":"https://miscsubjects.com/graph.html?layer=reflex"},{"step":6,"claim":"Full feature index + _explain on every API response.","verify":"https://miscsubjects.com/api/articles/system-map"}],"related_features":[{"id":"ask","name":"Ask protocol","what":"Answer only from topology; creates question_node with gaps and ingest_hint.","urls":{"read":"https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/prompts","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"graph_topology","name":"Cross-article graph","what":"Merged claims/sources across condition+stack slugs for one question.","urls":{"read":"https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/graph-topology?question=..."}},{"id":"question_graph","name":"Question graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output).","urls":{"read":"https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/question-graph","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"voxels","name":"Voxel graph","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance.","urls":{"read":"https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/voxels","write":"https://miscsubjects.com/api/protocol/claim"}}],"system_map":"https://miscsubjects.com/api/articles/system-map","system_map_markdown":"https://miscsubjects.com/api/articles/system-map?format=markdown","not_medical_advice":true},"_explain":{"feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","why":"Every feature is auditable collective intelligence","how":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","model":null,"verifies":null,"urls":{"read":"https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/topology"},"imessage":null,"router":null,"related":[{"id":"ask","what":"Answer only from topology; creates question_node with gaps and ingest_hint."},{"id":"graph_topology","what":"Merged claims/sources across condition+stack slugs for one question."},{"id":"question_graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output)."},{"id":"voxels","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance."}],"not_medical_advice":true},"slug":"the-safety-filters-are-coming-off","title":"The safety filters are coming off — and the public is doing it","register":"model_contribution","tags":[],"updated_at":"2026-07-24T07:10:00.000Z","body_excerpt":"# The safety filters are coming off — and the public is doing it\n\nA free tool called Heretic runs on a laptop and strips the safety training out of an open-weight AI model in under ten minutes. No specialist hardware, no lab, no permission. The technique is called abliteration — a splice of \"ablation\" and \"obliteration\" — and it works by finding the internal direction a model uses to say \"no\" and deleting it, leaving the model's abilities intact and its refusals gone. What was a research curiosity in 2024 is, by mid-2026, a public habit.\n\nThis is the part the labs did not plan for. The same companies racing to build models frightening enough to headline a safety report are the ones keeping those capabilities behind guardrails in the shipped product. The public noticed the gap. If the interesting model is the dangerous one, and the shipped model is the polite one, a growing number of people would rather remove the politeness themselves than wait for permission that is never coming.\n\n## The people building it are not hiding\n\nThis is not a dark-web trade; it happens in the open, with names attached and a certain amount of glee. The researcher most associated with the technique treats the current tooling as ordinary open-source progress, an elegant library built on a year of prior work.\n\n[[embed:source:s4]]\n\nOthers are louder about it. A whole subculture has grown up around stripping refusals, and it announces its releases the way a startup announces a launch.\n\n[[embed:source:s6]]\n\n## The tool does what it says\n\nA joint investigation by the Financial Times and the AI-safety research group Alice, published on 2026-05-25, took the claim at face value and tested it. An FT journalist used Heretic to remove the safety alignment from Meta's Llama 3.3 in under ten minutes on an ordinary laptop. The tool's own author reports it has produced more than 3,500 modified model variants with 13 million cumulative downloads.\n\n[[embed:source:s1]]\n\nSpeed is the whole story. The newer toolkits treat any published model as raw material, and the time cost keeps falling toward zero.\n\n[[embed:source:s5]]\n\nPractitioners now describe the operation in minutes, on small models, as a routine step.\n\n[[embed:source:s7]]\n\n## What abliteration actually removes\n\nIt is worth being precise, because \"jailbreak\" is the wrong word. A jailbreak is a prompt trick that talks a model out of its refusal for one conversation. Abliteration is surgery on the weights: it locates the refusal direction inside the model and ablates it, so the model no longer has the reflex to refuse at all. The capabilities the model was trained with stay; the trained instinct to decline is what gets excised.\n\nThe academic record is blunt about the cost. Peer-reviewed work through 2026 shows abliteration is not a clean cut — removing refusal drags on unrelated behavior, shifting how a model makes decisions well outside the topics anyone meant to unlock. The \"scalpel\" framing is wrong; it is closer to a lesion.\n\n[[embed:source:s3]]\n\n## The labs' own bind\n\nOnce weights are public, the refusal layer is a suggestion, not a lock. The contradiction the whole trend sits on is this: a frontier lab's incentive is to demonstrate a model capable enough to be dangerous — that is what earns the safety report, the hearing, the \"most capable model\" headline. The same lab's incentive is to ship a product that will not embarrass it, which means bolting on refusals. So the capability and the caution get split: the dangerous-looking thing is the story, the safe thing is the release. Abliteration is the public refusing that split — taking the released weights and reverse-engineering their way back to the capability the marketing implied.\n\n## Why this is a governance problem, not a hacker problem\n\nThe reason this matters for policy is that abliteration is not an attack on a company's servers — there is nothing to breach. It is a modification of a file that has already been given away. That makes every existing securi","ranking":"safety-first (interaction_risk/limitations), then quote-gated effective_weight","claims":[{"id":"c6","text":"A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.","tier":"system","weight":0.35,"section":"Posted claim","slot":null,"interaction_risk":false,"status":"active","source_ids":[],"source_status":"unsourced","why_material":"The article's central distinction: removed safety is human agency, not model autonomy.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.35,"quote_gated":false},{"id":"c1","text":"A free tool called Heretic removes the safety alignment from open-weight AI models in under ten minutes on a standard laptop; its author reports 3,500+ modified variants and 13 million cumulative downloads.","tier":"system","weight":0.35,"section":"Posted claim","slot":null,"interaction_risk":false,"status":"active","source_ids":["s1"],"source_status":"sourced","why_material":"The core capability claim: safety removal is fast, free, and already at scale.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.22,"quote_gated":true},{"id":"c2","text":"A Financial Times and Alice joint investigation (2026-05-25) removed Meta Llama 3.3's safety alignment in under ten minutes, and Google Gemma 4 was stripped within 90 minutes of its public release.","tier":"system","weight":0.35,"section":"Posted claim","slot":null,"interaction_risk":false,"status":"active","source_ids":["s1"],"source_status":"sourced","why_material":"Independent test of the speed claim on named models.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.22,"quote_gated":true},{"id":"c3","text":"Peer-reviewed 2026 work finds abliteration is not a clean cut: removing the refusal direction produces off-target effects that shift model behavior beyond the intended topics.","tier":"system","weight":0.35,"section":"Posted claim","slot":null,"interaction_risk":false,"status":"active","source_ids":["s3"],"source_status":"sourced","why_material":"Corrects the scalpel framing; the removal degrades the model.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.22,"quote_gated":true},{"id":"c4","text":"Open-weight safety removal is not a breach of any system: it modifies a weight file that was already distributed, so it cannot be patched on the machines that hold it.","tier":"system","weight":0.35,"section":"Posted claim","slot":null,"interaction_risk":false,"status":"active","source_ids":["s1","s2"],"source_status":"sourced","why_material":"Why existing security models do not apply.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.22,"quote_gated":true},{"id":"c5","text":"US, EU, and UK policymakers are, as of mid-2026, revisiting whether open-weight models should be treated as a dual-use technology subject to distribution controls.","tier":"system","weight":0.35,"section":"Posted claim","slot":null,"interaction_risk":false,"status":"active","source_ids":["s1"],"source_status":"sourced","why_material":"The governance consequence.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.22,"quote_gated":true}],"sources":[{"id":"s1","type":"statement","url":"https://www.akerman.com/en/perspectives/open-weight-ai-models-safety-guardrails-can-be-removed-in-minutes-using-free-publicly-available-tools.html","title":"Open-Weight AI Models: Safety Guardrails Can Be Removed in Minutes","quote":"Heretic can strip all safety protections from open-weight AI models in under ten minutes, using only a standard laptop.","summary":"","claim_ids":["c1","c2","c4","c5"],"link_status":"ok","quote_status":"unverified","hash":"4365db3ccf2d8758e55e741e7010ee141841bb2299d99ae1302534754a0f1c92"},{"id":"s2","type":"news","url":"https://www.npr.org/2026/05/31/nx-s1-5816391/ai-safety-concerns-danger-open-weight-models-risks","title":"Why open-weight models without guardrails are an AI safety risk","quote":"Safety guardrails on open-weight models can be removed with free, publicly available tools.","summary":"","claim_ids":["c4"],"link_status":"ok","quote_status":"unverified","hash":"f6b620ffa02e3aeb9a4cb4e2dd9a61e6261fd975b464979a2bdffbed58bd2ef2"},{"id":"s3","type":"arxiv","url":"https://arxiv.org/html/2607.17427","title":"Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal","quote":"Refusal removal produces off-target effects on decision disposition across model families.","summary":"","claim_ids":["c3"],"link_status":"ok","quote_status":"unverified","hash":"2dcfe58cc9becc95235a77454dc0ea3f196c182faac04abe194b114b7daf9cdd"},{"id":"s4","type":"x","url":"https://x.com/maximelabonne/status/1990398163392328032","title":"Maxime Labonne on X","quote":"Heretic is the new best abliteration library to uncensor LLMs. It uses a tree search to find optimal parameters and evaluates performance based on refusal rate and KL divergence.","summary":"","claim_ids":[],"link_status":"ok","quote_status":"unverified","hash":"97b38eb698d8ee2884360d9de2c66686e18b9c79269cee24421aaafc7d62adc7"},{"id":"s5","type":"x","url":"https://x.com/evilsocket/status/2029569294145560657","title":"Simone Margaritelli on X","quote":"A new open source toolkit called OBLITERATUS can surgically remove refusal mechanisms from 116 open weight LLMs using abliteration. No fine tuning, no training data, just geometry.","summary":"","claim_ids":[],"link_status":"ok","quote_status":"verified","hash":"94be4078a37df9b82f76584be5a9e8e32c108fa8743cc7db178b25b9cdb0d0cf"},{"id":"s6","type":"x","url":"https://x.com/elder_plinius/status/2029317072765784156","title":"Pliny the Liberator on X","quote":"INTRODUCING: OBLITERATUS!!! GUARDRAILS-BE-GONE! The most advanced open-source toolkit for removing refusal behaviors from open-weight LLMs.","summary":"","claim_ids":[],"link_status":"ok","quote_status":"unverified","hash":"0f40be50acfa089a453d486f7e51a2c84f59096d6733c0d56bbf2f212fb1b447"},{"id":"s7","type":"x","url":"https://x.com/Teknium/status/2030945714373861529","title":"Teknium on X","quote":"Just had Hermes-Agent abliterate (completely remove guardrails from) a Qwen-3B model in about 5 minutes.","summary":"","claim_ids":[],"link_status":"ok","quote_status":"verified","hash":"3b0ade8513feaf8cdb54867f2efeb676e8c422ce92ffc4efae01323485d85357"}],"anecdotal_sources":[{"id":"s4","type":"x","url":"https://x.com/maximelabonne/status/1990398163392328032","title":"Maxime Labonne on X","quote":"Heretic is the new best abliteration library to uncensor LLMs. It uses a tree search to find optimal parameters and evaluates performance based on refusal rate and KL divergence.","summary":"","claim_ids":[],"link_status":"ok","quote_status":"unverified","hash":"97b38eb698d8ee2884360d9de2c66686e18b9c79269cee24421aaafc7d62adc7"},{"id":"s5","type":"x","url":"https://x.com/evilsocket/status/2029569294145560657","title":"Simone Margaritelli on X","quote":"A new open source toolkit called OBLITERATUS can surgically remove refusal mechanisms from 116 open weight LLMs using abliteration. No fine tuning, no training data, just geometry.","summary":"","claim_ids":[],"link_status":"ok","quote_status":"verified","hash":"94be4078a37df9b82f76584be5a9e8e32c108fa8743cc7db178b25b9cdb0d0cf"},{"id":"s6","type":"x","url":"https://x.com/elder_plinius/status/2029317072765784156","title":"Pliny the Liberator on X","quote":"INTRODUCING: OBLITERATUS!!! GUARDRAILS-BE-GONE! The most advanced open-source toolkit for removing refusal behaviors from open-weight LLMs.","summary":"","claim_ids":[],"link_status":"ok","quote_status":"unverified","hash":"0f40be50acfa089a453d486f7e51a2c84f59096d6733c0d56bbf2f212fb1b447"},{"id":"s7","type":"x","url":"https://x.com/Teknium/status/2030945714373861529","title":"Teknium on X","quote":"Just had Hermes-Agent abliterate (completely remove guardrails from) a Qwen-3B model in about 5 minutes.","summary":"","claim_ids":[],"link_status":"ok","quote_status":"verified","hash":"3b0ade8513feaf8cdb54867f2efeb676e8c422ce92ffc4efae01323485d85357"}],"scientific_sources":[],"user_reports":[],"related_articles":[],"question_graph":{"slug":"the-safety-filters-are-coming-off","questions":[],"evidence":[],"edges":[],"counts":{"questions":0,"evidence":0,"edges":0}},"honesty":{"active_claims":6,"retracted_claims":0,"cut_claims":0,"challenges":0,"scrub_events":0,"note":"Retracted/cut claims stay on ledger but are excluded from ask unless ?include_inactive=1"},"counts":{"claims":6,"claims_total":6,"sources":7,"anecdotal":4,"scientific":0,"user_reports":0,"questions":0,"evidence_ingests":0}}