--- license: apache-2.0 base_model: Qwen/Qwen2.5-3B-Instruct language: - en tags: - reasoning - math - chain-of-thought - qwen2.5 - gguf - text-generation - conversational pipeline_tag: text-generation library_name: transformers ---
![1789911910f175](https://cdn-uploads.huggingface.co/production/uploads/6a5f8c87a38cac087c2c6c05/zZ-C-s-RZWV6sxVhj9RBe.png) 🐍 Boomslang (3B Reasoning & Math Engine) **A compact, high-efficiency 3-billion-parameter model fine-tuned for deep chain-of-thought mathematical reasoning, logic, and algebra—without sacrificing natural conversational ability.** [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Creator-Monster--Code-blue)](https://huggingface.co/Monster-Code) [![License](https://img.shields.io/badge/License-Apache_2.0-green.svg)](https://opensource.org/licenses/Apache-2.0) [![GGUF Available](https://img.shields.io/badge/GGUF-Included-purple.svg)](https://huggingface.co/Monster-Code/Boomslang/blob/main/boomslang-3b-qwen.gguf) --- ### ❤️ If you find Boomslang useful, please hit the **Like** button at the top of this page and [Follow @Monster-Code](https://huggingface.co/Monster-Code) for more open-weights AI releases! ---
## 💡 What is Boomslang? Most small models (1B–3B parameters) struggle with two extremes: they are either polite chatbots that completely hallucinate basic arithmetic, or narrow math models that forget how to hold a conversation and start writing unprompted proofs when you simply say "Hi." **Boomslang** was trained to bridge that gap. Starting from the strong foundation of **`Qwen/Qwen2.5-3B-Instruct`**, Boomslang was post-trained on an **NVIDIA RTX PRO 6000 Blackwell** across a curated ~17,500-sample reasoning mixture: 1. **DeepSeek-R1 Distilled Proofs (`open-r1/OpenR1-Math-220k`):** Teaches the network an internal self-reflection loop (` ... `) to break down complex algebraic expressions, geometry, and multi-step deduction before committing to an answer. 2. **Step-by-Step Arithmetic Rigor (`openai/gsm8k`):** Calibrates attention heads on strict order-of-operations arithmetic and unambiguous answer derivation. The result is an edge-friendly 3B model that works through tricky algebra and word puzzles methodically, but still greets you warmly and follows instructions when you just want to talk. --- ## 📦 What's Inside This Repository? * **Single Standalone `model.safetensors`:** No multi-part file splits. The entire 3-billion parameter model is packed into a single, clean ~6.1 GB file. * **Pre-Converted GGUF (`boomslang-3b-qwen.gguf`):** Directly ready for **Ollama**, **LM Studio**, and **llama.cpp** on your local machine (MacBook, laptop, or home GPU). * Full configuration and tokenizer files for immediate `transformers` plug-and-play. --- ## ⚡ Quickstart: Python & Transformers ```python import torch from transformers import AutoTokenizer, AutoModelForCausalLM, TextStreamer MODEL_ID = "Monster-Code/Boomslang" tokenizer = AutoTokenizer.from_pretrained(MODEL_ID) model = AutoModelForCausalLM.from_pretrained( MODEL_ID, torch_dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32, device_map="auto" ) # Standard ChatML format messages = [ { "role": "system", "content": "You are Boomslang, a helpful and precise reasoning assistant. Reason step-by-step before providing your final answer." }, { "role": "user", "content": "A box has 36 chocolates. Tom eats 1/4 of them, and Sarah eats 9. How many chocolates are left in the box?" } ] prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tokenizer(prompt, return_tensors="pt").to(model.device) streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True) with torch.no_grad(): model.generate( **inputs, streamer=streamer, max_new_tokens=1024, temperature=0.6, top_p=0.9 ) ``` --- ## 💻 Running Locally with Ollama / LM Studio ### 1. Using Ollama: Create a `Modelfile` in the same directory as the downloaded `.gguf`: ```dockerfile FROM ./boomslang-3b-qwen.gguf TEMPLATE """{{ if .System }}<|im_start|>system {{ .System }}<|im_end|> {{ end }}{{ if .Prompt }}<|im_start|>user {{ .Prompt }}<|im_end|> {{ end }}<|im_start|>assistant {{ .Response }}<|im_end|>""" PARAMETER stop "<|im_start|>" PARAMETER stop "<|im_end|>" PARAMETER temperature 0.6 ``` Build and run: ```bash ollama create boomslang -f Modelfile ollama run boomslang ``` ### 2. Using LM Studio: 1. Download **`boomslang-3b-qwen.gguf`** from the [Files and versions tab](https://huggingface.co/Monster-Code/Boomslang/tree/main). 2. Drop it into your local LM Studio models folder. 3. Start chatting with GPU offloading enabled! --- ## 🧠 Sample Interaction **User:** > *"Solve for x: 5x - 10 = 25. Show your work."* **Boomslang:** ```text The equation is 5x - 10 = 25. Our goal is to isolate the variable x on one side. Step 1: Add 10 to both sides to cancel out the -10 constant: 5x - 10 + 10 = 25 + 10 5x = 35 Step 2: Divide both sides by the coefficient 5: 5x / 5 = 35 / 5 x = 7 Verification: Substitute x = 7 into the original expression: 5(7) - 10 = 35 - 10 = 25. Both sides match. The answer is x = 7. ``` --- ## 🛠️ Training Details * **Base Architecture:** Qwen2.5 3B (Decoder-only Transformer) * **Hardware:** NVIDIA RTX PRO 6000 Blackwell Server Edition * **Precision:** BF16 Native Mixed Precision with Fused AdamW * **Effective Batch Size:** 32 (8 per device × 4 gradient accumulation steps) * **Learning Rate:** 1.5e-4 with dynamic linear warmup * **Label Masking:** Dynamic batch padding via `DataCollatorForSeq2Seq` with `-100` masking to guarantee loss is never calculated on padding noise --- ## 🤝 Community & Support * 👤 **Creator:** [Monster-Code](https://huggingface.co/Monster-Code) * 💬 Have suggestions, evaluation runs, or dataset ideas? Leave a note in the **Discussions** tab! * ⭐ **If Boomslang helps your workflow, please consider starring/liking the repository!**