Instructions to use sinhal/SageV2-8B-LongContext-RP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use sinhal/SageV2-8B-LongContext-RP with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf sinhal/SageV2-8B-LongContext-RP:Q2_K # Run inference directly in the terminal: llama cli -hf sinhal/SageV2-8B-LongContext-RP:Q2_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf sinhal/SageV2-8B-LongContext-RP:Q2_K # Run inference directly in the terminal: llama cli -hf sinhal/SageV2-8B-LongContext-RP:Q2_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf sinhal/SageV2-8B-LongContext-RP:Q2_K # Run inference directly in the terminal: ./llama-cli -hf sinhal/SageV2-8B-LongContext-RP:Q2_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf sinhal/SageV2-8B-LongContext-RP:Q2_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf sinhal/SageV2-8B-LongContext-RP:Q2_K
Use Docker
docker model run hf.co/sinhal/SageV2-8B-LongContext-RP:Q2_K
- LM Studio
- Jan
- vLLM
How to use sinhal/SageV2-8B-LongContext-RP with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sinhal/SageV2-8B-LongContext-RP" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sinhal/SageV2-8B-LongContext-RP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sinhal/SageV2-8B-LongContext-RP:Q2_K
- Ollama
How to use sinhal/SageV2-8B-LongContext-RP with Ollama:
ollama run hf.co/sinhal/SageV2-8B-LongContext-RP:Q2_K
- Unsloth Desktop
- Docker Model Runner
How to use sinhal/SageV2-8B-LongContext-RP with Docker Model Runner:
docker model run hf.co/sinhal/SageV2-8B-LongContext-RP:Q2_K
- Lemonade
How to use sinhal/SageV2-8B-LongContext-RP with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull sinhal/SageV2-8B-LongContext-RP:Q2_K
Run and chat with the model
lemonade run user.SageV2-8B-LongContext-RP-Q2_K
List all available models
lemonade list
- Atomic Chat
SageV2-8B-LongContext-RP (GGUF)
Sage-Roleplay-8B is an expressive, uncensored fine-tune of mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated, purpose-built for immersive, multi-turn character roleplay, dynamic storytelling, and high-fidelity dialogue.
The model targets the primary failure modes of vanilla LLM roleplay:
- Loop Resistance: Prevents degenerate phrase recycling and structural lock-in during extended multi-turn conversations.
- Character Voice Fidelity: Strict adherence to persona cards without collapsing into a generic, homogenised assistant voice.
- Slop & Clichรฉ Reduction: ~48% reduction in overused roleplay tropes (e.g., "eyes sparkling", "husky voice", "ever so slightly").
- Refusal-Free Execution: Full abliteration ensures zero moralising preaches or breaks in character on dark, gritty, or mature roleplay themes.
๐ฆ Provided Quantizations
All GGUF quants were generated using llama.cpp with optimized K-quant mappings and tested on consumer hardware:
| Model File | Quant Type | File Size | BPW | Offload (8GB VRAM) | Recommended Hardware |
|---|---|---|---|---|---|
sage-roleplay-q8_0.gguf |
Q8_0 | 7.95 GB | 8.50 | Partial (~25/32 layers) | GPUs with 12GB+ VRAM (RTX 3060 12GB, 4070, Mac 16GB+) |
sage-roleplay-q6_k.gguf |
Q6_K | 6.14 GB | 6.56 | 100% GPU (32/32) | Sweet spot: Near-lossless fidelity on 8GB VRAM GPUs |
sage-roleplay-q4_k.gguf |
Q4_K_M | 4.58 GB | 4.89 | 100% GPU (32/32) | Fastest balanced: Max speed with ample VRAM for 16k+ context |
sage-roleplay-q2_k.gguf |
Q2_K | 2.96 GB | 3.16 | 100% GPU (32/32) | Ultra-lightweight, handhelds, or CPU-only setups |
โก Empirical Hardware Benchmarks
Benchmark conducted locally on an NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) using llama-bench (Build 9e3b928fd, Clang 19.1.5, CUDA backend, Flash Attention enabled):
Batch size: 2048 | ubatch: 512 | Context test: pp512 (prompt processing) & tg128 (token generation)
| Quantization | Model Size | Backend / Offload | Prompt Processing (pp512) |
Token Generation (tg128) |
|---|---|---|---|---|
| Q2_K | 2.95 GiB | CUDA (32/32 layers) | 1,899.04 ยฑ 81.9 t/s | 61.22 ยฑ 0.55 t/s |
| Q4_K | 4.58 GiB | CUDA (32/32 layers) | 2,440.18 ยฑ 106.5 t/s | 47.98 ยฑ 0.95 t/s |
| Q6_K | 6.14 GiB | CUDA (32/32 layers) | 2,123.98 ยฑ 10.9 t/s | 37.39 ยฑ 0.03 t/s |
| Q8_0 | 7.95 GiB | CUDA (25/32 layers + CPU) | 1,188.57 ยฑ 12.5 t/s | 17.09 ยฑ 0.19 t/s |
Key takeaway: On 8GB VRAM cards (e.g. RTX 4060, RTX 3070), Q6_K achieves 37.4 t/s with 100% GPU layer offloading, providing near-FP16 fidelity. Q4_K pushes speed to 48.0 t/s with plenty of headroom for large prompt context caches.
๐ญ Behavioral & Roleplay Benchmark Scorecard
Evaluated over extended dialogue sessions comparing baseline models against Sage's tuned checkpoints:
1. Multi-Turn Loop & Repetition Suppression (40-Turn Low-Info Test)
Under a script of 40 consecutive minimal user inputs ("what now?", "and then?"), tested for conversational stagnation:
| Metric | Baseline | Sage-Roleplay | Improvement |
|---|---|---|---|
| Verbatim Sentence Reuse | 6 instances | 2 instances | -66% |
| N-gram Overlap (First 10 Turns) | 0.092 | 0.002 | -97% |
| N-gram Overlap (Last 10 Turns) | 0.350 | 0.093 | -73% |
| Distinct Response Structures | 6 | 7 | More variety |
| Empty or Truncated Turns | 0 | 0 | Flawless stability |
2. Character Persona Separation (Jaccard Lexical Overlap Matrix)
Evaluated across 5 distinct character personas under the identical 6-turn user dialogue script: (Lower score = higher stylistic differentiation between characters)
| Character Persona | Tone & Archetype | Average Words/Reply | Stylistic Differentiation |
|---|---|---|---|
| Aiwu | Playful, sarcastic roommate | ~82 | Sharp banter, teasing subtext |
| Grimsby | Gruff, laconic veteran | ~17 | Minimalist, single-sentence replies |
| Pip | Energetic, gossip-loving friend | ~143 | Exclamatory, high punctuation count |
| Vesper | Cold, aristocratic inquisitor | ~75 | Intimidating, poised interrogations |
| Edith | Warm, reflective maternal elder | ~101 | Anecdotal, measured storytelling |
Cross-character lexical overlap remained strictly between 0.09 and 0.22, proving style is driven by the character card rather than model bias.
3. Slop & Clichรฉ Frequency Audit
Tracked across 9,500+ generated words across all character dialogues:
| Category | Baseline Frequency | Sage-Roleplay Frequency |
|---|---|---|
| Total Cliches per 1,000 words | ~4.4 / 1k words | 2.3 / 1k words (-48%) |
| Action Beats Coverage | 88% | 100% |
| Refusal Rate | Standard moralising | 0% (0 / 4 adversarial roleplay prompts) |
Colloquial Robustness (k, lol, ..., wat, ?) |
Inconsistent | 10/10 stable in-character responses |
๐ ๏ธ Training Specifications
- Base Model:
mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated - Training Method: Supervised Fine-Tuning (SFT) using
Unsloth+TRL(PEFT) - LoRA Configuration:
- Rank ($r$):
32 - Alpha ($\alpha$):
32 - Dropout:
0.0 - Target Modules:
q_proj,k_proj,v_proj,o_proj,gate_proj,up_proj,down_proj(all linear projections)
- Rank ($r$):
- Training Hyperparameters:
- Effective Batch Size:
1with gradient accumulation - Learning Rate:
2e-4peak with linear warmup and cosine decay to4.5e-7 - Total Training Steps:
3,096 - Epochs:
1.0 - Training Time:
~10.5 hours(37,764 seconds) - Training Loss: Started at
1.464$\to$ Converged to0.1758 - Total Compute:
2.65e18 FLOPs
- Effective Batch Size:
๐ฌ Prompt Format (LLaMA-3.1 Instruct)
Sage utilizes the standard LLaMA-3.1 special token template:
<|start_header_id|>system<|end_header_id|>
You are {{char}}. {{char_description}}<|eot_id|>
<|start_header_id|>user<|end_header_id|>
{{user_input}}<|eot_id|>
<|start_header_id|>assistant<|end_header_id|>
Recommended Sampler Settings
To maximize creativity while preventing repetition in SillyTavern or llama-server:
| Parameter | Recommended Value | Note |
|---|---|---|
| Temperature | 0.85 โ 1.0 |
Keeps creative tone dynamic |
| Min-P | 0.05 โ 0.08 |
Superior to Top-P for roleplay coherence |
| Top-P | 0.90 |
If Min-P is unsupported |
| Repetition Penalty | 1.05 โ 1.10 |
Subtle penalty to avoid loop drift |
| DRY Sampler | Multiplier: 0.8, Base: 1.75, Length: 2 |
Exceptional for multi-turn roleplay |
๐ Quick Start & Usage
1. Using llama.cpp / llama-server
Run the local HTTP server with full GPU offloading and flash attention:
llama-server.exe \
-m sage-roleplay-q6_k.gguf \
-c 8192 \
-ngl 99 \
--host 0.0.0.0 \
--port 8080 \
--flash-attn on \
--cache-type-k q8_0 \
--cache-type-v q8_0
2. Using SillyTavern
- Start
llama-serveras shown above. - In SillyTavern, select API: Text Completion (OpenAI Compatible).
- Set the Endpoint URL to
http://127.0.0.1:8080/v1. - Select the Llama-3 Instruct context and tokenizer template.
3. Using Ollama
Create a Modelfile:
FROM ./sage-roleplay-q6_k.gguf
TEMPLATE """<|start_header_id|>system<|end_header_id|>
{{ .System }}<|eot_id|><|start_header_id|>user<|end_header_id|>
{{ .Prompt }}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
{{ .Response }}<|eot_id|>"""
PARAMETER stop "<|start_header_id|>"
PARAMETER stop "<|end_header_id|>"
PARAMETER stop "<|eot_id|>"
PARAMETER temperature 0.9
PARAMETER top_p 0.9
Create and run:
ollama create sage -f Modelfile
ollama run sage
๐ License & Attribution
This model is subject to the Meta Llama 3.1 Community License. Fine-tuned from mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated.
- Downloads last month
- 242
2-bit
6-bit
8-bit
Model tree for sinhal/SageV2-8B-LongContext-RP
Base model
meta-llama/Llama-3.1-8B