--- license: apache-2.0 library_name: transformers pipeline_tag: text-generation base_model: - mossez-systems/Mossez-100M-Coder-Base tags: - causal-lm - conversational - code - fill-in-the-middle - instruct - llama - research - experimental --- # Mossez-100M-Coder-Instruct Mossez-100M-Coder-Instruct is an experimental 100M-parameter coding instruction model with this weight lineage: `Mossez-100M-Base -> Mossez-100M-Coder-Base -> Mossez-100M-Coder-Instruct`. The general [`Mossez-100M-Instruct`](https://huggingface.co/mossez-systems/Mossez-100M-Instruct) was used only as a tokenizer, chat-template, release, and inference reference; its weights were not used as source weights for this model. ## Model details | Property | Value | |---|---:| | Parameters | 100,098,048 | | Architecture | Llama-compatible decoder-only Transformer | | Layers / hidden size | 12 / 768 | | Query / KV heads | 12 / 4 | | Context length | 1,024 tokens | | Vocabulary | 32,007 | | Objective | Assistant-only SFT loss | | Weight format | Safetensors, FP32 | | License | Apache-2.0 | ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "mossez-systems/Mossez-100M-Coder-Instruct" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto") messages = [{"role": "user", "content": "Write a short Python function that adds two integers."}] prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tokenizer(prompt, return_tensors="pt").to(model.device) output = model.generate(**inputs, do_sample=False, max_new_tokens=96) new_tokens = output[0, inputs.input_ids.shape[1]:] print(tokenizer.decode(new_tokens, skip_special_tokens=True)) ``` ## Training and evaluation The model was fine-tuned for one bounded epoch: 660 optimizer steps over 2,640 project-authored examples, using assistant-only loss. Immutable validation and test sets contain 330 examples each across 11 balanced task types. See [TRAINING_REPORT.md](TRAINING_REPORT.md), [EVALUATION.md](EVALUATION.md), and [DATASET_ATTRIBUTION.md](DATASET_ATTRIBUTION.md). The released `model.safetensors` SHA-256 is `0aade7d070122633abccd70cc5500e5bc36c7f69fafeadd9f5ee0b5a3e0766bf`. ## Limitations This is a small research model, not a reliable or safe production coding assistant. The authored SFT corpus is balanced but narrow and template-heavy, so held-out loss may overstate general-world capability. Expect repetition, incorrect constants, malformed code, hallucinated APIs, weak instruction following, and early EOS. Validate, test, and sandbox every output.