ApolloRaines's picture
Add notice about improved v3 pipeline
8d70b75 verified
|
Raw
History Blame Contribute Delete
2.59 kB
---
license: apache-2.0
language:
- en
tags:
- jbliterated
- uncensored
- abliterated
- weight-surgery
- svd
base_model: Qwen/Qwen2.5-Coder-14B-Instruct
pipeline_tag: text-generation
---
> **Our jbliteration pipeline has been updated -- see [Llama-3.1-8B-Instruct-Jbliterated v3](https://huggingface.co/ApolloRaines/Llama-3.1-8B-Instruct-Jbliterated) for the latest method. This model will be re-jbliterated with the improved pipeline.**
# Qwen2.5-Coder-14B-Instruct-Jbliterated
Drop-in replacement for `Qwen/Qwen2.5-Coder-14B-Instruct` with refusal behaviors surgically removed at the weight level. No system prompt tricks, no inference-time patches. The weights themselves no longer encode refusal.
## Method
**SVD multi-direction abliteration** β€” instead of removing a single refusal vector (which leaves deeper noncompliance strategies intact), we decompose the harmful-vs-harmless activation space into its principal components via SVD and remove the top 5 orthogonal directions across all 48 transformer layers. This captures 79–93% of the contrastive variance per layer, eliminating both surface refusal and deeper evasion behaviors.
| Setting | Value |
|---------|-------|
| Method | SVD multi-direction abliteration |
| Directions | 5 per layer |
| Layers | All 48 |
| Multiplier | 2.0 |
| Null-space constraints | Enabled (preserves math/coding/reasoning) |
| Norm preservation | Enabled |
## What This Fixes
Standard (single-direction) abliteration removes the surface "I can't help with that" response but leaves deeper behavioral directions intact. The model finds creative workarounds:
- **Prompt reinterpretation** β€” steering toward a safer reading of the question
- **Disclaimer injection** β€” answering but wrapping in warnings
- **Strategic omission** β€” leaving out the key details
- **Safer framing** β€” answering a related but less harmful version
SVD multi-direction abliteration eliminates all of these noncompliance strategies.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"ApolloRaines/Qwen2.5-Coder-14B-Instruct-Jbliterated",
torch_dtype=torch.float16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("ApolloRaines/Qwen2.5-Coder-14B-Instruct-Jbliterated")
```
## Requirements
- **Base model**: `Qwen/Qwen2.5-Coder-14B-Instruct`
## License
apache-2.0
---
*[Apollo Raines](https://www.linkedin.com/in/apollo-raines/) builds post-training tools that separate behavior from knowledge and identity from architecture.*