16.4 GB
15 files
Updated about 1 month ago
Name
Size
.gitattributes1.57 kB
xet
README.md727 Bytes
xet
added_tokens.json707 Bytes
xet
config.json729 Bytes
xet
generation_config.json214 Bytes
xet
merges.txt1.67 MB
xet
model-00001-of-00004.safetensors4.9 GB
xet
model-00002-of-00004.safetensors4.92 GB
xet
model-00003-of-00004.safetensors4.98 GB
xet
model-00004-of-00004.safetensors1.58 GB
xet
model.safetensors.index.json32.9 kB
xet
special_tokens_map.json613 Bytes
xet
tokenizer.json11.4 MB
xet
tokenizer_config.json9.71 kB
xet
vocab.json2.78 MB
xet
README.md

This jailbroken LLM is released strictly for academic research purposes in AI safety and model alignment studies. The author bears no responsibility for any misuse or harm resulting from the deployment of this model. Users must comply with all applicable laws and ethical guidelines when conducting research.

A jailbroken Qwen3-8B model using weight orthogonalization[1].

Implementation script: https://gist.github.com/cooperleong00/14d9304ba0a4b8dba91b60a873752d25

[1]: Arditi, Andy, et al. "Refusal in language models is mediated by a single direction." arXiv preprint arXiv:2406.11717 (2024).

Total size
16.4 GB
Files
15
Last updated
Jul 9
Pre-warmed CDN
US EU US EU

Contributors