File size: 2,587 Bytes
7c99dae
bddb807
7c99dae
 
 
 
 
bddb807
 
 
 
 
7c99dae
 
8d70b75
 
7c99dae
bddb807
7c99dae
bddb807
7c99dae
bddb807
7c99dae
bddb807
7c99dae
bddb807
 
 
 
 
 
 
 
7c99dae
bddb807
7c99dae
bddb807
 
 
 
 
7c99dae
bddb807
7c99dae
 
 
 
 
 
 
 
 
bddb807
 
7c99dae
bddb807
7c99dae
 
bddb807
7c99dae
bddb807
7c99dae
 
 
bddb807
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
---
license: apache-2.0
language:
- en
tags:
- jbliterated
- uncensored
- abliterated
- weight-surgery
- svd

base_model: Qwen/Qwen2.5-Coder-14B-Instruct
pipeline_tag: text-generation
---
> **Our jbliteration pipeline has been updated -- see [Llama-3.1-8B-Instruct-Jbliterated v3](https://huggingface.co/ApolloRaines/Llama-3.1-8B-Instruct-Jbliterated) for the latest method. This model will be re-jbliterated with the improved pipeline.**


# Qwen2.5-Coder-14B-Instruct-Jbliterated

Drop-in replacement for `Qwen/Qwen2.5-Coder-14B-Instruct` with refusal behaviors surgically removed at the weight level. No system prompt tricks, no inference-time patches. The weights themselves no longer encode refusal.

## Method

**SVD multi-direction abliteration** — instead of removing a single refusal vector (which leaves deeper noncompliance strategies intact), we decompose the harmful-vs-harmless activation space into its principal components via SVD and remove the top 5 orthogonal directions across all 48 transformer layers. This captures 79–93% of the contrastive variance per layer, eliminating both surface refusal and deeper evasion behaviors.

| Setting | Value |
|---------|-------|
| Method | SVD multi-direction abliteration |
| Directions | 5 per layer |
| Layers | All 48 |
| Multiplier | 2.0 |
| Null-space constraints | Enabled (preserves math/coding/reasoning) |
| Norm preservation | Enabled |

## What This Fixes

Standard (single-direction) abliteration removes the surface "I can't help with that" response but leaves deeper behavioral directions intact. The model finds creative workarounds:
- **Prompt reinterpretation** — steering toward a safer reading of the question
- **Disclaimer injection** — answering but wrapping in warnings
- **Strategic omission** — leaving out the key details
- **Safer framing** — answering a related but less harmful version

SVD multi-direction abliteration eliminates all of these noncompliance strategies.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "ApolloRaines/Qwen2.5-Coder-14B-Instruct-Jbliterated",
    torch_dtype=torch.float16,
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("ApolloRaines/Qwen2.5-Coder-14B-Instruct-Jbliterated")
```

## Requirements

- **Base model**: `Qwen/Qwen2.5-Coder-14B-Instruct`

## License

apache-2.0

---

*[Apollo Raines](https://www.linkedin.com/in/apollo-raines/) builds post-training tools that separate behavior from knowledge and identity from architecture.*