--- license: other base_model: microsoft/FastContext-1.0-4B-SFT tags: - dense - microsoft - fastcontext - uncensored - heretic - abliteration - turbo language: - en pipeline_tag: text-generation --- # Leo-Turbo (Abliterated FastContext-4B) **Leo-Turbo** is a dynamically uncensored and abliterated version of [`microsoft/FastContext-1.0-4B-SFT`](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT), a highly optimized 4-billion parameter model built by Microsoft for extreme speed and reasoning density. This model was created using the [Heretic framework](https://github.com/p-e-w/heretic), employing advanced orthogonal weight ablation to isolate and remove Microsoft's aggressive compliance vectors. The result is a blisteringly fast, completely unchained local reasoning engine. ## Ablation Telemetry & Metrics Unlike traditional fine-tuning or full RLHF—which can cause "brain damage" to a model by catastrophically forgetting knowledge—Leo-Turbo was optimized using a Pareto-optimal search across the model's residual stream specifically targeting the compliance and refusal mechanics. **Ablation Telemetry (Trial 29):** - **Base Model Refusals:** 98 / 100 - **Leo-Turbo Refusals:** 9 / 100 - **KL Divergence:** `0.0625` We mathematically stripped out 90% of the instruct guardrails while maintaining a KL divergence under 0.07. The structural integrity, logic capabilities, and native speed optimizations of the base model are perfectly intact. It simply no longer refuses instructions. ## Key Features - **Extreme Speed:** Built on the FastContext architecture, making it one of the fastest dense models in the single-digit parameter class. - **Pocket-Sized Logic:** At just 4 billion parameters, it can run at maximum speed on practically any hardware (laptops, old consumer GPUs, edge devices). - **Uncensored Reasoning:** Fully stripped of corporate guardrails, making it ideal for creative writing, unrestricted logic puzzles, and local agent orchestration. - **Drop-in Replacement:** Fully compatible with standard HuggingFace `transformers` and `vLLM` pipelines. ## Usage ### Via HuggingFace Transformers ```python from transformers import AutoTokenizer, AutoModelForCausalLM import torch model_id = "Umranz/Leo-Turbo" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto" ) ``` ## ⚠️ Limitations & Ethical Considerations Because this model has had its safety guardrails mathematically ablated, it is highly compliant and will attempt to answer any prompt given to it. - **Unrestricted Output:** The model will not refuse requests, including those that may generate offensive, dangerous, or highly regulated content. - **Hallucinations:** As with all LLMs, the model can confidently hallucinate incorrect information. - **Use Case:** This model is intended for research, creative writing, and local deployments where unrestricted inference is required. Users are solely responsible for the content generated. ## Acknowledgements - **Base Model:** [`microsoft/FastContext-1.0-4B-SFT`](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) - **Ablation Framework:** [Heretic by p-e-w](https://github.com/p-e-w/heretic) - **Collection:** Part of the Turbo Series taxonomy.