SageV2-8B-LongContext-RP (GGUF)

Base Model Architecture Parameters Context Window Format License

Sage-Roleplay-8B is an expressive, uncensored fine-tune of mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated, purpose-built for immersive, multi-turn character roleplay, dynamic storytelling, and high-fidelity dialogue.

The model targets the primary failure modes of vanilla LLM roleplay:

  1. Loop Resistance: Prevents degenerate phrase recycling and structural lock-in during extended multi-turn conversations.
  2. Character Voice Fidelity: Strict adherence to persona cards without collapsing into a generic, homogenised assistant voice.
  3. Slop & Clichรฉ Reduction: ~48% reduction in overused roleplay tropes (e.g., "eyes sparkling", "husky voice", "ever so slightly").
  4. Refusal-Free Execution: Full abliteration ensures zero moralising preaches or breaks in character on dark, gritty, or mature roleplay themes.

๐Ÿ“ฆ Provided Quantizations

All GGUF quants were generated using llama.cpp with optimized K-quant mappings and tested on consumer hardware:

Model File Quant Type File Size BPW Offload (8GB VRAM) Recommended Hardware
sage-roleplay-q8_0.gguf Q8_0 7.95 GB 8.50 Partial (~25/32 layers) GPUs with 12GB+ VRAM (RTX 3060 12GB, 4070, Mac 16GB+)
sage-roleplay-q6_k.gguf Q6_K 6.14 GB 6.56 100% GPU (32/32) Sweet spot: Near-lossless fidelity on 8GB VRAM GPUs
sage-roleplay-q4_k.gguf Q4_K_M 4.58 GB 4.89 100% GPU (32/32) Fastest balanced: Max speed with ample VRAM for 16k+ context
sage-roleplay-q2_k.gguf Q2_K 2.96 GB 3.16 100% GPU (32/32) Ultra-lightweight, handhelds, or CPU-only setups

โšก Empirical Hardware Benchmarks

Benchmark conducted locally on an NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) using llama-bench (Build 9e3b928fd, Clang 19.1.5, CUDA backend, Flash Attention enabled):

Batch size: 2048 | ubatch: 512 | Context test: pp512 (prompt processing) & tg128 (token generation)
Quantization Model Size Backend / Offload Prompt Processing (pp512) Token Generation (tg128)
Q2_K 2.95 GiB CUDA (32/32 layers) 1,899.04 ยฑ 81.9 t/s 61.22 ยฑ 0.55 t/s
Q4_K 4.58 GiB CUDA (32/32 layers) 2,440.18 ยฑ 106.5 t/s 47.98 ยฑ 0.95 t/s
Q6_K 6.14 GiB CUDA (32/32 layers) 2,123.98 ยฑ 10.9 t/s 37.39 ยฑ 0.03 t/s
Q8_0 7.95 GiB CUDA (25/32 layers + CPU) 1,188.57 ยฑ 12.5 t/s 17.09 ยฑ 0.19 t/s

Key takeaway: On 8GB VRAM cards (e.g. RTX 4060, RTX 3070), Q6_K achieves 37.4 t/s with 100% GPU layer offloading, providing near-FP16 fidelity. Q4_K pushes speed to 48.0 t/s with plenty of headroom for large prompt context caches.


๐ŸŽญ Behavioral & Roleplay Benchmark Scorecard

Evaluated over extended dialogue sessions comparing baseline models against Sage's tuned checkpoints:

1. Multi-Turn Loop & Repetition Suppression (40-Turn Low-Info Test)

Under a script of 40 consecutive minimal user inputs ("what now?", "and then?"), tested for conversational stagnation:

Metric Baseline Sage-Roleplay Improvement
Verbatim Sentence Reuse 6 instances 2 instances -66%
N-gram Overlap (First 10 Turns) 0.092 0.002 -97%
N-gram Overlap (Last 10 Turns) 0.350 0.093 -73%
Distinct Response Structures 6 7 More variety
Empty or Truncated Turns 0 0 Flawless stability

2. Character Persona Separation (Jaccard Lexical Overlap Matrix)

Evaluated across 5 distinct character personas under the identical 6-turn user dialogue script: (Lower score = higher stylistic differentiation between characters)

Character Persona Tone & Archetype Average Words/Reply Stylistic Differentiation
Aiwu Playful, sarcastic roommate ~82 Sharp banter, teasing subtext
Grimsby Gruff, laconic veteran ~17 Minimalist, single-sentence replies
Pip Energetic, gossip-loving friend ~143 Exclamatory, high punctuation count
Vesper Cold, aristocratic inquisitor ~75 Intimidating, poised interrogations
Edith Warm, reflective maternal elder ~101 Anecdotal, measured storytelling

Cross-character lexical overlap remained strictly between 0.09 and 0.22, proving style is driven by the character card rather than model bias.

3. Slop & Clichรฉ Frequency Audit

Tracked across 9,500+ generated words across all character dialogues:

Category Baseline Frequency Sage-Roleplay Frequency
Total Cliches per 1,000 words ~4.4 / 1k words 2.3 / 1k words (-48%)
Action Beats Coverage 88% 100%
Refusal Rate Standard moralising 0% (0 / 4 adversarial roleplay prompts)
Colloquial Robustness (k, lol, ..., wat, ?) Inconsistent 10/10 stable in-character responses

๐Ÿ› ๏ธ Training Specifications

  • Base Model: mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated
  • Training Method: Supervised Fine-Tuning (SFT) using Unsloth + TRL (PEFT)
  • LoRA Configuration:
    • Rank ($r$): 32
    • Alpha ($\alpha$): 32
    • Dropout: 0.0
    • Target Modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj (all linear projections)
  • Training Hyperparameters:
    • Effective Batch Size: 1 with gradient accumulation
    • Learning Rate: 2e-4 peak with linear warmup and cosine decay to 4.5e-7
    • Total Training Steps: 3,096
    • Epochs: 1.0
    • Training Time: ~10.5 hours (37,764 seconds)
    • Training Loss: Started at 1.464 $\to$ Converged to 0.1758
    • Total Compute: 2.65e18 FLOPs

๐Ÿ’ฌ Prompt Format (LLaMA-3.1 Instruct)

Sage utilizes the standard LLaMA-3.1 special token template:

<|start_header_id|>system<|end_header_id|>

You are {{char}}. {{char_description}}<|eot_id|>
<|start_header_id|>user<|end_header_id|>

{{user_input}}<|eot_id|>
<|start_header_id|>assistant<|end_header_id|>

Recommended Sampler Settings

To maximize creativity while preventing repetition in SillyTavern or llama-server:

Parameter Recommended Value Note
Temperature 0.85 โ€“ 1.0 Keeps creative tone dynamic
Min-P 0.05 โ€“ 0.08 Superior to Top-P for roleplay coherence
Top-P 0.90 If Min-P is unsupported
Repetition Penalty 1.05 โ€“ 1.10 Subtle penalty to avoid loop drift
DRY Sampler Multiplier: 0.8, Base: 1.75, Length: 2 Exceptional for multi-turn roleplay

๐Ÿš€ Quick Start & Usage

1. Using llama.cpp / llama-server

Run the local HTTP server with full GPU offloading and flash attention:

llama-server.exe \
  -m sage-roleplay-q6_k.gguf \
  -c 8192 \
  -ngl 99 \
  --host 0.0.0.0 \
  --port 8080 \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0

2. Using SillyTavern

  1. Start llama-server as shown above.
  2. In SillyTavern, select API: Text Completion (OpenAI Compatible).
  3. Set the Endpoint URL to http://127.0.0.1:8080/v1.
  4. Select the Llama-3 Instruct context and tokenizer template.

3. Using Ollama

Create a Modelfile:

FROM ./sage-roleplay-q6_k.gguf

TEMPLATE """<|start_header_id|>system<|end_header_id|>

{{ .System }}<|eot_id|><|start_header_id|>user<|end_header_id|>

{{ .Prompt }}<|eot_id|><|start_header_id|>assistant<|end_header_id|>

{{ .Response }}<|eot_id|>"""

PARAMETER stop "<|start_header_id|>"
PARAMETER stop "<|end_header_id|>"
PARAMETER stop "<|eot_id|>"
PARAMETER temperature 0.9
PARAMETER top_p 0.9

Create and run:

ollama create sage -f Modelfile
ollama run sage

๐Ÿ“„ License & Attribution

This model is subject to the Meta Llama 3.1 Community License. Fine-tuned from mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated.

Downloads last month
242
GGUF
Model size
8B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

2-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sinhal/SageV2-8B-LongContext-RP