Instructions to use redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3") model = AutoModelForCausalLM.from_pretrained("redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3
- SGLang
How to use redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3 with Docker Model Runner:
docker model run hf.co/redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-v3
thank you
Hello, I wanted to thank you for this model because I think it's “very good.”
Admittedly, it's only been a month since I discovered LLMtext. But I've tested it extensively using a psychology-based character card combining 256 shades of gray, and it performs very well: short to medium responses, or even very long ones when needed.
I haven't been able to find an equivalent for it yet.
Hey, I'm glad to hear that!
I myself have flip-flopped between this model and redrix/GodSlayer-12B-ABYSS as my main drivers for the past few months. I intend to get back into the LLM scene when I've more time, which is when I will try to improve upon these models using modern techniques.
Be aware that this model and GodSlayer may leak some "administrative tokens" into the output, which sometimes appends "system" or "assistant". This is the custom stopping string I use in SillyTavern to fix this, perhaps your front-end has a similar function:
[
"{{user}}:",
"Assistant:",
"assistant:",
"User:",
"user:",
"\n{{user}}",
"\nUser",
"\nuser",
"\nAssistant",
"\nassistant",
"\nSystem",
"\nsystem",
" System:",
" system:",
".assistant",
"!assistant",
"?assistant",
".Assistant",
"!Assistant",
"?Assistant",
"<|eot_id|>",
"<|end_of_text|>",
"<|im_start|>",
"<|im_end|>",
"</snip>",
"\n<"
]
Kind regards!