--- language: - en license: apache-2.0 tags: - mamba - ssm - causal-lm - neuralai - base-model - sft-queued pipeline_tag: text-generation base_model: state-spaces/mamba-2.8b --- # ๐Ÿง  NeuralAI ยท Mamba K3 NeuralAI's **2.8B Mamba SSM** โ€” the third and largest model in the NeuralAI line (K1 130M โ†’ K2 793M โ†’ **K3 2.8B**), fine-tuned from `state-spaces/mamba-2.8b` for assistant-style chat with worked reasoning, code, math, writing, and safety behaviors. ## Model Details - **Architecture:** Mamba SSM (`state-spaces/mamba-2.8b`) - **Parameters:** 2.8B - **Status:** SFT queued โ€” weights will land here when training completes - **Fine-tune:** LoRA r=16, ฮฑ=32 on Mamba `in_proj`, bf16, cosine LR 2e-4 - **Dataset:** NeuralAI curated set โ€” 3,068 examples (computed-answer math, reasoning puzzles, code, safety, writing, factual, multi-turn) - **Chat template:** NeuralAI `[INST]` (same as Mamba K1) - **Creator:** De'Andrew Preston Harris, NeuralAI ## Status - This page is **created ahead of the SFT run** (per the model-line order in [github.com/Subject-Emu-5259/NeuralAI](https://github.com/Subject-Emu-5259/NeuralAI)). - Training is **paused pending GPU allocation** on the Colab runtime; the dataset and training script are ready to launch the moment a GPU is available. - When the SFT run completes, the merged weights + tokenizer + config are pushed here automatically. ## Training Command (when GPU is available) ```bash BASE_MODEL=state-spaces/mamba-2.8b HF_REPO=Subject-Emu-5259/NeuralAI-Mamba-K3 \ python training/finetune_run.py ``` ## Prompt Format Same as Mamba K1 โ€” the NeuralAI `[INST]` chat template: ```text <|endoftext|>[INST] {user message} [/INST] ``` ## License Apache 2.0