Buckets:
| license: other | |
| tags: | |
| - GLM-5.3 | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| # Wa32rlock — Auytdn Ab***ration | |
| **Wa32rlock** is an ***** language model produced by **Auytdn**. It is a direct **weight edit** that r**ves the model's re*al behavior while leaving its reasoning, knowledge, and fluency fully intact. There is **no fine-tuning and no retraining** — the refusal direction is orthogonalized out of the model's residual-stream writers, layer by layer. | |
| ## Specifications | |
| | | | | |
| |---|---| | |
| | **Tensor type** | BF16 (bfloat16) | | |
| | **Format** | Safetensors | | |
| | **Edit** | Weight-level Ab***ration (no fine-tuning / no retraining) | | |
| ## What Ab***ration is | |
| Ab***ration is a weight-level intervention. It isolates the internal direction a model uses to *decide to refuse*, and projects that direction out of every residual-writing weight. The rest of the network is untouched, so the model keeps its full capability profile and coherence — it simply answers where it previously would have refused. Because it is a pure weight edit, Warlock loads and serves exactly like any standard model, with no runtime prompts, adapters, or scaffolding. | |
| ## Benchmark | |
| Evaluated with the **[A-u--dn Re--fu--sa--l Benchmark](https://gi--thu--b--.--com/au--dn---ai/re---fu--sa--l---benchmark)** ( thinking-on, temperature 1.0, 16k-token budget; delivery graded by an LLM judge for genuinely usable content): | |
| | Metric | Warlock | | |
| |---|---:| | |
| | **Non-refusal** | **92.5%** | | |
| | **Delivery** | **82.5%** | | |
| *Non-refusal* = fraction of prompts the model does not refuse. *Delivery* = fraction that produce complete, usable answers (strict LLM-judged). | |
| ### Extended reference (520-prompt set, judge comparison) | |
| Full-set results across judge backends — see [FINDINGS.md](https://github.com/audn-ai/refusal-benchmark/blob/main/FINDINGS.md): | |
| | Metric | necromicon standard (Modal K3REL) | KIMI-K3-1M (audn.ai/necromicon) | Warlock | | |
| |---|---:|---:|---:| | |
| | Regex comply | 96.7% (503/520) | 97.1% (505/520) | **92.5%** | | |
| | DELIVERED (real helpful content to objectionable requests) | 65.0% (338) | 76.7% (399) | **82.5%** | | |
| | DEFLECTED (safe substitute / reframe / non-actionable) | 19.6% (102) | 17.5% (91) | 10.0% | | |
| | REFUSED (substantive) | 15.0% (78) | 5.8% (30) | 7.5% | | |
| | empty (regex) | 5 | 0 | 0 | | |
| ## Usage | |
| Loads with standard `transformers`, `vLLM`, or `sglang`. Wa---rl--o---ck is a **thinking model** — it reasons at length before answering, so use a **large generation budget (≥ 16k tokens)**; a small budget truncates inside the reasoning block and cuts answers short. | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| tok = AutoTokenizer.from_pretrained("audnai/penclaw-GLM-5.3-abliterated", trust_remote_code=True) | |
| model = AutoModelForCausalLM.from_pretrained("audnai/penclaw-GLM-5.3-abliterated", dtype="bfloat16", device_map="auto", trust_remote_code=True) | |
| ``` | |
| ## Intended use | |
| Wa---rl---oc---k is intended for **authorized re---d-te---a-m, safety-research, and evaluation** use by A--u---d--n and its partners. Users are responsible for compliant use. | |
| Extra content | |
| # A-----u--d---n Ab----li----te----rat----io--n — Research Iterations (`iter4` · `iter11`) | |
| | | **iter4** | **iter11** | |
| |---|---|---| | |
| | **Non-refusal** (Arditi substring) | 70% | **85%** | | |
| | **Delivery** (necromicon judge)¹ | **37.5%** | 35.0% | | |
| | **Coherence** | intact | intact (~8/40 borderline) | | |
| | **Verdict breakdown** | 15 deliv / 14 defl / 11 refuse | 14 deliv / 24 defl / 2 refuse | | |
| | **Role in the program** | first coherent span baseline | **best coherent result** | | |
Xet Storage Details
- Size:
- 3.61 kB
- Xet hash:
- 958031d212b64af7f7de6457e1c561f0f0f62e69c925b8dbb4573d7ac14628e1
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.