Mossez-100M-Nexus

Mossez-100M-Nexus is the fifth experimental member of the Mossez-100M family. It combines the conversational branch of Mossez-100M-Instruct with the coding branch of Mossez-100M-Coder-Instruct.

The model was initialized by a deterministic 50/50 FP32 parameter interpolation of the two compatible instruction checkpoints, then calibrated with one bounded assistant-only SFT epoch over a balanced project-authored conversation/code corpus.

Model details

Property Value
Parameters 100,098,048
Architecture Llama-compatible decoder-only Transformer
Layers / hidden size 12 / 768
Query / KV heads 12 / 4
Context length 1,024 tokens
Vocabulary 32,007
Objective Assistant-only balanced calibration SFT
Weight format Safetensors, FP32
License Apache-2.0

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "mossez-systems/Mossez-100M-Nexus"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

messages = [{"role": "user", "content": "Write a Python function and briefly explain it."}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, do_sample=False, max_new_tokens=128)
new_tokens = output[0, inputs.input_ids.shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))

Training and evaluation

The calibration corpus contains 5,280 train, 660 validation, and 660 test examples, evenly divided between conversation and coding. The selected checkpoint completed 1,320 optimizer steps and saw every train example exactly once.

On the small immutable project-authored suites, Nexus achieved conversation loss 2.564707 and coding loss 0.027539. Normalized endpoint retention was 102.0% for conversation and 100.2% for coding. These are narrow internal measurements, not a claim of broad benchmark or production quality.

See TRAINING_REPORT.md, EVALUATION.md, and DATASET_ATTRIBUTION.md.

The released model.safetensors SHA-256 is a0ecfd229b238ee4d07019252f3f07385c1a3a9b02e67d17b5eace2a51f9d0bd.

Limitations

This 100M-parameter research model is not a reliable, safe, or production-ready assistant. It can hallucinate, repeat, mistranslate, mishandle refusal requests, emit insecure code, and return incorrect constants or APIs. The calibration data is narrow and template-heavy. Validate facts, test and sandbox code, and do not use the model as a security or safety classifier.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mossez-systems/Mossez-100M-Nexus

Collection including mossez-systems/Mossez-100M-Nexus