Instructions to use cgus/Qwen2-7B-Instruct-abliterated-exl2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cgus/Qwen2-7B-Instruct-abliterated-exl2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="cgus/Qwen2-7B-Instruct-abliterated-exl2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("cgus/Qwen2-7B-Instruct-abliterated-exl2") model = AutoModelForCausalLM.from_pretrained("cgus/Qwen2-7B-Instruct-abliterated-exl2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use cgus/Qwen2-7B-Instruct-abliterated-exl2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "cgus/Qwen2-7B-Instruct-abliterated-exl2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cgus/Qwen2-7B-Instruct-abliterated-exl2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/cgus/Qwen2-7B-Instruct-abliterated-exl2
- SGLang
How to use cgus/Qwen2-7B-Instruct-abliterated-exl2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "cgus/Qwen2-7B-Instruct-abliterated-exl2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cgus/Qwen2-7B-Instruct-abliterated-exl2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "cgus/Qwen2-7B-Instruct-abliterated-exl2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cgus/Qwen2-7B-Instruct-abliterated-exl2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use cgus/Qwen2-7B-Instruct-abliterated-exl2 with Docker Model Runner:
docker model run hf.co/cgus/Qwen2-7B-Instruct-abliterated-exl2
Qwen2-7B-Instruct-abliterated-exl2
Model: Qwen2-7B-Instruct-abliterated
Made by: natong19
Based on original model: Qwen2-7B-Instruct
Created by: Qwen
| Quant | VRAM/4k | VRAM/8k | VRAM/16k | VRAM/32k |
|---|---|---|---|---|
| 4bpw h6 (main) | 5.3GB | 5.6GB | 5.9GB | 6.8GB |
| 4.25bpw h6 | 5.5GB | 5.8GB | 6.2GB | 7.1GB |
| 4.65bpw h6 | 5.8GB | 6.1GB | 6.5GB | 7.3GB |
| 5bpw h6 | 6GB | 6.4GB | 6.7GB | 7.7GB |
| 6bpw h6 | 6.8GB | 7.2GB | 7.5GB | 8.4GB |
| 8bpw h8 | 8.2GB | 8.6GB | 8.9GB | 9.8GB |
Quantization notes
Made with Exllamav2 0.1.5 and the default dataset.
Doesn't seem to work with 4 or 8bit cache with Exllamav2-0.1.5, maybe it could change in the future.
I'm quite impressed with its ability to process a non-English text at 32k context with usable results with my 12GB GPU, with 8bpw precision at that.
How to run
This quantization uses GPU and requires Exllamav2 loader, model files have to be fully loaded in VRAM to work.
It should work well either with Nvidia RTX cards on Windows/Linux or AMD on Linux. For other hardware it's better to use GGUF models instead.
This model can be loaded with in following applications:
Text Generation Webui
KoboldAI
ExUI, etc.
Original model card
Qwen2-7B-Instruct-abliterated
Introduction
Abliterated version of Qwen2-7B-Instruct using failspy's notebook. The model's strongest refusal directions have been ablated via weight orthogonalization, but the model may still refuse your request, misunderstand your intent, or provide unsolicited advice regarding ethics or safety.
Quickstart
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "natong19/Qwen2-7B-Instruct-abliterated"
device = "cuda" # the device to load the model onto
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
prompt = "Give me a short introduction to large language model."
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(device)
generated_ids = model.generate(
model_inputs.input_ids,
max_new_tokens=256
)
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)
Evaluation
Evaluation framework: lm-evaluation-harness 0.4.2
| Datasets | Qwen2-7B-Instruct | Qwen2-7B-Instruct-abliterated |
|---|---|---|
| ARC (25-shot) | 62.5 | 62.5 |
| GSM8K (5-shot) | 73.0 | 72.2 |
| HellaSwag (10-shot) | 81.8 | 81.7 |
| MMLU (5-shot) | 70.7 | 70.5 |
| TruthfulQA (0-shot) | 57.3 | 55.0 |
| Winogrande (5-shot) | 76.2 | 77.4 |
- Downloads last month
- 6
Model tree for cgus/Qwen2-7B-Instruct-abliterated-exl2
Base model
natong19/Qwen2-7B-Instruct-abliterated