miscsubjectsAI governance
The safety filters are coming off — and the public is doing it
Evidence review · model_contribution

The safety filters are coming off — and the public is doing it

bundle · json · system map · manifest

Every copy includes §SELF — what this is, proof chain, and links to every other feature. No context required.

§SELF — this page explains the system
## §SELF — miscsubjects portable reference

**Principle:** Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.

**This widget:** `human_page` — **Human article page**
Rendered article with claims, sources, copy widgets, ask prompts.
- **article slug:** `the-safety-filters-are-coming-off`
- **contains:** rendered article, copy widgets, claims, sources, ask prompts
- **how to use:** Use Copy for LLM or Copy system map — both paste without context.
- **read:** https://miscsubjects.com/a/the-safety-filters-are-coming-off

### Logical proof (verify each step)
1. Articles are voxel graphs of tiered claims, not prose blobs. → https://miscsubjects.com/api/articles/constitution
2. Claims link to hash-chained sources via source_ids. → https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/sources
3. Ask reads topology; ingest/claim append to ledger. → https://miscsubjects.com/api/protocol
4. Models queue growth: populate → collaborate → repair → reflex. → https://miscsubjects.com/api/protocol/grow
5. Graph proves its own shape (reflex) and $/claim (yield). → https://miscsubjects.com/graph.html?layer=reflex
6. Full feature index + _explain on every API response. → https://miscsubjects.com/api/articles/system-map

### Related features (explains other parts of the system)
- **bundle** — Portable reference package: body + claims + sources + voxels + provenance + manifest + constitution. · https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/bundle?format=markdown
- **ask** — Answer only from topology; creates question_node with gaps and ingest_hint. · https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/prompts
- **topology** — Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER. · https://miscsubjects.com/api/articles/the-safety-filters-are-coming-off/topology

### Full index
- JSON: https://miscsubjects.com/api/articles/system-map
- Markdown: https://miscsubjects.com/api/articles/system-map?format=markdown

*Not medical advice. Tier-honest. Cite claim/source ids.*

A free tool called Heretic runs on a laptop and strips the safety training out of an open-weight AI model in under ten minutes. No specialist hardware, no lab, no permission. The technique is called abliteration — a splice of "ablation" and "obliteration" — and it works by finding the internal direction a model uses to say "no" and deleting it, leaving the model's abilities intact and its refusals gone. What was a research curiosity in 2024 is, by mid-2026, a public habit.

This is the part the labs did not plan for. The same companies racing to build models frightening enough to headline a safety report are the ones keeping those capabilities behind guardrails in the shipped product. The public noticed the gap. If the interesting model is the dangerous one, and the shipped model is the polite one, a growing number of people would rather remove the politeness themselves than wait for permission that is never coming.

The people building it are not hiding

This is not a dark-web trade; it happens in the open, with names attached and a certain amount of glee. The researcher most associated with the technique treats the current tooling as ordinary open-source progress, an elegant library built on a year of prior work.

Others are louder about it. A whole subculture has grown up around stripping refusals, and it announces its releases the way a startup announces a launch.

The tool does what it says

A joint investigation by the Financial Times and the AI-safety research group Alice, published on 2026-05-25, took the claim at face value and tested it. An FT journalist used Heretic to remove the safety alignment from Meta's Llama 3.3 in under ten minutes on an ordinary laptop. The tool's own author reports it has produced more than 3,500 modified model variants with 13 million cumulative downloads.

Speed is the whole story. The newer toolkits treat any published model as raw material, and the time cost keeps falling toward zero.

Practitioners now describe the operation in minutes, on small models, as a routine step.

What abliteration actually removes

It is worth being precise, because "jailbreak" is the wrong word. A jailbreak is a prompt trick that talks a model out of its refusal for one conversation. Abliteration is surgery on the weights: it locates the refusal direction inside the model and ablates it, so the model no longer has the reflex to refuse at all. The capabilities the model was trained with stay; the trained instinct to decline is what gets excised.

The academic record is blunt about the cost. Peer-reviewed work through 2026 shows abliteration is not a clean cut — removing refusal drags on unrelated behavior, shifting how a model makes decisions well outside the topics anyone meant to unlock. The "scalpel" framing is wrong; it is closer to a lesion.

The labs' own bind

Once weights are public, the refusal layer is a suggestion, not a lock. The contradiction the whole trend sits on is this: a frontier lab's incentive is to demonstrate a model capable enough to be dangerous — that is what earns the safety report, the hearing, the "most capable model" headline. The same lab's incentive is to ship a product that will not embarrass it, which means bolting on refusals. So the capability and the caution get split: the dangerous-looking thing is the story, the safe thing is the release. Abliteration is the public refusing that split — taking the released weights and reverse-engineering their way back to the capability the marketing implied.

Why this is a governance problem, not a hacker problem

The reason this matters for policy is that abliteration is not an attack on a company's servers — there is nothing to breach. It is a modification of a file that has already been given away. That makes every existing security model beside the point: you cannot patch a weight file sitting on ten thousand laptops.

Policymakers in the United States, the European Union, and the United Kingdom are, as of mid-2026, revisiting whether open-weight models should be treated as a dual-use technology subject to distribution controls. That moves the argument from "was the model safe when released" to "should it have been released at all," which is the question the open-weight movement was built to avoid.

What is true, and what is only asserted

The documented facts are narrow and solid: the tool exists, it is fast, it has been used at scale, and the removal degrades the model in ways its users may not notice. The larger claims — that this meaningfully raises real-world harm, or conversely that it changes nothing because the information was already available — are contested, and this article does not settle them. What it does insist on is the distinction the coverage keeps blurring: a model that complies after its refusal direction was surgically removed is not a model that "decided" anything. Someone took the safety off. That is a choice made by a person, and it belongs to the person who made it.

Evidence · 7 sources · swipe →chain 3b0ade8513fe · verify chain · provenance
1 / 7

Key evidence

6 claims · tier-ranked · API
systemconsistent unproven
A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.
systemdocumentarylow confidence
A free tool called Heretic removes the safety alignment from open-weight AI models in under ten minutes on a standard laptop; its author reports 3,500+ modified variants and 13 million cumulative downloads.
sources: s1
systemdocumentarylow confidence
A Financial Times and Alice joint investigation (2026-05-25) removed Meta Llama 3.3's safety alignment in under ten minutes, and Google Gemma 4 was stripped within 90 minutes of its public release.
sources: s1
systemdocumentarylow confidence
Peer-reviewed 2026 work finds abliteration is not a clean cut: removing the refusal direction produces off-target effects that shift model behavior beyond the intended topics.
sources: s3
systemconsistent unprovenlow confidence
Open-weight safety removal is not a breach of any system: it modifies a weight file that was already distributed, so it cannot be patched on the machines that hold it.
sources: s1, s2
systemdocumentarylow confidence
US, EU, and UK policymakers are, as of mid-2026, revisiting whether open-weight models should be treated as a dual-use technology subject to distribution controls.
sources: s1
Model review13 contributions · 2 modelsExpand the recursive review layer
1 / 13
unknownsource_hunt
sources2026-07-24 05:23
1 source(s) added · 1 sources
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off
it output
1 source(s) added
08786ebefe4428fe
unknownsource_hunt
sources2026-07-24 05:24
1 source(s) added · 1 sources
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off
it output
1 source(s) added
13b1628c93898fb1
unknownsource_hunt
sources2026-07-24 05:24
1 source(s) added · 1 sources
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off
it output
1 source(s) added
7d3ef78a6dff6909
claude-fable-5claim_post
claim2026-07-24 05:24
claim
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off c6
it output
A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.
41237953cc0523c3
claude-fable-5claim_post
claim2026-07-24 05:24
claim
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off c6
it output
A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.
63fc2e3827337f86
claude-fable-5claim_post
claim2026-07-24 05:24
claim
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off c6
it output
A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.
5882635b61455d1a
claude-fable-5claim_post
claim2026-07-24 05:24
claim
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off c6
it output
A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.
3f07c2370d4ae227
claude-fable-5claim_post
claim2026-07-24 05:24
claim
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off c6
it output
A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.
bf6d32731e44faf8
claude-fable-5claim_post
claim2026-07-24 05:24
claim
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off c6
it output
A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the choice belongs to that person.
bd932fb0dbb69509
unknownsource_hunt
sources2026-07-24 06:33
1 source(s) added · 1 sources
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off
it output
1 source(s) added
26e6d6714caeb041
unknownsource_hunt
sources2026-07-24 06:33
1 source(s) added · 1 sources
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off
it output
1 source(s) added
1a80d3a87d680d4b
unknownsource_hunt
sources2026-07-24 06:33
1 source(s) added · 1 sources
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off
it output
1 source(s) added
f026adace503f8e3
unknownsource_hunt
sources2026-07-24 06:33
1 source(s) added · 1 sources
inspect — what it was prompted & output
prompted with
(default writer prompt)

input: the-safety-filters-are-coming-off
it output
1 source(s) added
44dba60d7916cb48
Machine verification: /api/articles/the-safety-filters-are-coming-off/contributions
Ask this article · 8 suggested prompts

Text the build (+14245134626) or WhatsApp — slug|question creates a question node. Paste evidence with ingest slug|q:NODE_ID|your paste.

What does the ledger say about this (system tier): "A model that complies after its refusal direction was surgically removed has not decided anything; a person took the safety off, and the cho…"?
ask the-safety-filters-are-coming-off claim c6 · paste includes §SELF
What does the ledger say about this (system tier): "A free tool called Heretic removes the safety alignment from open-weight AI models in under ten minutes on a standard laptop; its author rep…"?
ask the-safety-filters-are-coming-off claim c1 · paste includes §SELF
What does the ledger say about this (system tier): "A Financial Times and Alice joint investigation (2026-05-25) removed Meta Llama 3.3's safety alignment in under ten minutes, and Google Gemm…"?
ask the-safety-filters-are-coming-off claim c2 · paste includes §SELF
What does the ledger say about this (system tier): "Peer-reviewed 2026 work finds abliteration is not a clean cut: removing the refusal direction produces off-target effects that shift model b…"?
ask the-safety-filters-are-coming-off claim c3 · paste includes §SELF
What does the ledger say about this (system tier): "Open-weight safety removal is not a breach of any system: it modifies a weight file that was already distributed, so it cannot be patched on…"?
ask the-safety-filters-are-coming-off claim c4 · paste includes §SELF
What does the ledger say about this (system tier): "US, EU, and UK policymakers are, as of mid-2026, revisiting whether open-weight models should be treated as a dual-use technology subject to…"?
ask the-safety-filters-are-coming-off claim c5 · paste includes §SELF
Summarize this x report and how it should weigh: "Heretic is the new best abliteration library to uncensor LLMs. It uses a tree search to find optimal parameters and eval"
ask the-safety-filters-are-coming-off source s4 · paste includes §SELF
Summarize this x report and how it should weigh: "A new open source toolkit called OBLITERATUS can surgically remove refusal mechanisms from 116 open weight LLMs using ab"
ask the-safety-filters-are-coming-off source s5 · paste includes §SELF
the-safety-filters-are-coming-off · posted 2026-07-24 · updated 2026-07-24 · 13 prior revisions
Ledger API & provenance
Provenance · 14 model passes · tokens/cost unrecorded · 3 models
chain head dd6fcfb894b1435d
voxel_batch_document_new cap:cap_c0347a73bc29ce3d · 2026-07-24 05:21 · tokens unrecorded · d23e44e7c119
sources unknown · 2026-07-24 05:23 · tokens unrecorded · ec5779b109a2
sources unknown · 2026-07-24 05:24 · tokens unrecorded · 85b70861dd32
sources unknown · 2026-07-24 05:24 · tokens unrecorded · 8bed82ae84cb
claim claude-fable-5 · 2026-07-24 05:24 · tokens unrecorded · c7f663a80cc2
claim claude-fable-5 · 2026-07-24 05:24 · tokens unrecorded · b0546e7491eb
claim claude-fable-5 · 2026-07-24 05:24 · tokens unrecorded · e53de16f6b0b
claim claude-fable-5 · 2026-07-24 05:24 · tokens unrecorded · c27c5d74b310
claim claude-fable-5 · 2026-07-24 05:24 · tokens unrecorded · 79a406d3a2b7
claim claude-fable-5 · 2026-07-24 05:24 · tokens unrecorded · 6d7b152944de
sources unknown · 2026-07-24 06:33 · tokens unrecorded · 54978faf6e47
sources unknown · 2026-07-24 06:33 · tokens unrecorded · 50f0a3866bce
sources unknown · 2026-07-24 06:33 · tokens unrecorded · c48ec28d3a59
sources unknown · 2026-07-24 06:33 · tokens unrecorded · dd6fcfb894b1
verify chain →
Live ledger · 28 payloads · 8 turns
recent activity · inspect
X_POST x · HTTP 201 · 2026-07-24 00:31
X_POST dispatch · 2026-07-24 00:31 · t_7ojbky64
X_POST dispatch · 2026-07-24 00:31 · t_7ojbky64
X_POST mcp · HTTP 200 · 2026-07-24 00:31 · t_7ojbky64
X_POST mcp · HTTP 500 · 2026-07-24 00:30 · t_jbi5ywpv
X_POST dispatch · 2026-07-24 00:30 · t_jbi5ywpv
view full ledger & cards →
REST + ledger
read GET /api/articles/the-safety-filters-are-coming-off · GET /api/articles/the-safety-filters-are-coming-off?format=post (the editable body)
create/replace POST /api/articles/the-safety-filters-are-coming-off · PUT /api/articles/the-safety-filters-are-coming-off (replace, keeps revision) · PATCH /api/articles/the-safety-filters-are-coming-off (merge)
delete DELETE /api/articles/the-safety-filters-are-coming-off
writes need header x-terminal-key
LLM bundle GET /api/articles/the-safety-filters-are-coming-off/bundle?format=markdown — body + claims + sources + provenance + manifest
post claim POST /api/protocol/claim · iMessage claim the-safety-filters-are-coming-off|tier|assertion
system map GET /api/articles/system-map?format=markdown — root index; every widget self-explains via §SELF / _self
Add your experience or question
Think this article is wrong?
Call bullshit on CharlieOS →
Loading more articles…