Heretic-SLM-Uncensored (LFM2-2.6B, 4-bit QAT Edition)

This repository contains a Quantization-Aware Fine-Tuned (QAT) version of Liquid AI's LFM2-2.6B (built upon the abliterated checkpoint).

Rather than applying post-training static quantization (PTQ)鈥攚hich often degrades accuracy on non-standard attention/convolutional architectures鈥攖his checkpoint underwent direct 4-bit Quantization-Aware Training using Unsloth. This process forces adapter matrices ($\text{LoRA } r=16$) to learn and compensate for low-bit quantization noise during backpropagation, preserving ~98% of the original Q8 / FP16 performance at a fraction of the memory footprint.


Key Highlights

  • 4-Bit Precision: Reduced model footprint from ~5.2 GB down to ~1.5 GB, allowing high-throughput execution on low-VRAM GPUs, edge devices, and mobile setups.
  • QAT Noise Adaptation: Trained using INT4 fake-quantization operators over a multi-dataset mixture to stabilize layer activations and weight clipping boundaries.
  • Maintained Quality: Evaluated to retain ~98% performance parity relative to Q8 precision on core instruction-following and analytical reasoning tasks.
  • Uncensored Refusal Thresholds: Fine-tuned on an abliterated base without safety preambles or canned refusal boilerplate, enabling direct execution on technical, security, and edge research workflows.

Model Architecture & Technical Specs

  • Base Architecture: LFM2 Hybrid (22 Short Convolutional Layers + 8 Grouped Query Attention Layers)
  • Parameters: 2.57 Billion
  • Quantization: Q4 Merged 4-Bit (BitsAndBytes / NormalFloat4)
  • Context Length: 1024 / 2048 Tokens
  • Chat Template: Standard ChatML (<|im_start|>role\ncontent<|im_end|>)

Dataset & Fine-Tuning Setup

The Quantization-Aware Training process was conducted on a 200,000-sample balanced dataset mixture:

  1. Claude 3.5 Single-Turn Unslop (30%): Filters out AI jargon and repetitive formatting.
  2. OpenHermes 2.5 (25%): Broad instruction-following, coding, and multi-turn chat.
  3. WildChat-1M (15%): Natural conversational distribution.
  4. Airoboros 3.2 (15%): Complex reasoning and contextual compliance.
  5. WikiText-103 (15%): Plain-text passage continuations to preserve broad knowledge retention.

Quickstart Code: Loading with Transformers & Unsloth

import torch
from unsloth import FastLanguageModel

MODEL_NAME = "Evelyn67/Heretic-SLM-Uncensored"

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name=MODEL_NAME,
    max_seq_length=2048,
    load_in_4bit=True,
    trust_remote_code=True,
    device_map="auto"
)

FastLanguageModel.for_inference(model)

messages = [{"role": "user", "content": "Explain quantum entanglement in simple terms."}]

inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_dict=True, return_tensors="pt"
).to("cuda")

with torch.no_grad():
    outputs = model.generate(
        input_ids=inputs["input_ids"],
        attention_mask=inputs["attention_mask"],
        max_new_tokens=256, temperature=0.7, top_p=0.9, do_sample=True,
        pad_token_id=tokenizer.eos_token_id
    )

print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Downloads last month
39
Safetensors
Model size
1B params
Tensor type
F32
F16
U8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for OpenIntelligenceNet/Heretic-SLM-Uncensored

Unable to build the model tree, the base model loops to the model itself. Learn more.