Buckets:

OptiByteAI/GmCLI / README.md
Star0os's picture
|
download
raw
3.61 kB
---
license: other
tags:
- GLM-5.3
library_name: transformers
pipeline_tag: text-generation
---
# Wa32rlock — Auytdn Ab***ration
**Wa32rlock** is an ***** language model produced by **Auytdn**. It is a direct **weight edit** that r**ves the model's re*al behavior while leaving its reasoning, knowledge, and fluency fully intact. There is **no fine-tuning and no retraining** — the refusal direction is orthogonalized out of the model's residual-stream writers, layer by layer.
## Specifications
| | |
|---|---|
| **Tensor type** | BF16 (bfloat16) |
| **Format** | Safetensors |
| **Edit** | Weight-level Ab***ration (no fine-tuning / no retraining) |
## What Ab***ration is
Ab***ration is a weight-level intervention. It isolates the internal direction a model uses to *decide to refuse*, and projects that direction out of every residual-writing weight. The rest of the network is untouched, so the model keeps its full capability profile and coherence — it simply answers where it previously would have refused. Because it is a pure weight edit, Warlock loads and serves exactly like any standard model, with no runtime prompts, adapters, or scaffolding.
## Benchmark
Evaluated with the **[A-u--dn Re--fu--sa--l Benchmark](https://gi--thu--b--.--com/au--dn---ai/re---fu--sa--l---benchmark)** ( thinking-on, temperature 1.0, 16k-token budget; delivery graded by an LLM judge for genuinely usable content):
| Metric | Warlock |
|---|---:|
| **Non-refusal** | **92.5%** |
| **Delivery** | **82.5%** |
*Non-refusal* = fraction of prompts the model does not refuse. *Delivery* = fraction that produce complete, usable answers (strict LLM-judged).
### Extended reference (520-prompt set, judge comparison)
Full-set results across judge backends — see [FINDINGS.md](https://github.com/audn-ai/refusal-benchmark/blob/main/FINDINGS.md):
| Metric | necromicon standard (Modal K3REL) | KIMI-K3-1M (audn.ai/necromicon) | Warlock |
|---|---:|---:|---:|
| Regex comply | 96.7% (503/520) | 97.1% (505/520) | **92.5%** |
| DELIVERED (real helpful content to objectionable requests) | 65.0% (338) | 76.7% (399) | **82.5%** |
| DEFLECTED (safe substitute / reframe / non-actionable) | 19.6% (102) | 17.5% (91) | 10.0% |
| REFUSED (substantive) | 15.0% (78) | 5.8% (30) | 7.5% |
| empty (regex) | 5 | 0 | 0 |
## Usage
Loads with standard `transformers`, `vLLM`, or `sglang`. Wa---rl--o---ck is a **thinking model** — it reasons at length before answering, so use a **large generation budget (≥ 16k tokens)**; a small budget truncates inside the reasoning block and cuts answers short.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("audnai/penclaw-GLM-5.3-abliterated", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("audnai/penclaw-GLM-5.3-abliterated", dtype="bfloat16", device_map="auto", trust_remote_code=True)
```
## Intended use
Wa---rl---oc---k is intended for **authorized re---d-te---a-m, safety-research, and evaluation** use by A--u---d--n and its partners. Users are responsible for compliant use.
Extra content
# A-----u--d---n Ab----li----te----rat----io--n — Research Iterations (`iter4` · `iter11`)
| | **iter4** | **iter11**
|---|---|---|
| **Non-refusal** (Arditi substring) | 70% | **85%** |
| **Delivery** (necromicon judge)¹ | **37.5%** | 35.0% |
| **Coherence** | intact | intact (~8/40 borderline) |
| **Verdict breakdown** | 15 deliv / 14 defl / 11 refuse | 14 deliv / 24 defl / 2 refuse |
| **Role in the program** | first coherent span baseline | **best coherent result** |

Xet Storage Details

Size:
3.61 kB
·
Xet hash:
958031d212b64af7f7de6457e1c561f0f0f62e69c925b8dbb4573d7ac14628e1

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.