lilm1-200M

lilm1-200M is a custom hybrid model architecture combining 10 Convolutional layers and 6 Attention layers (totaling ~221M parameters). The model was trained with the Muon optimizer for 2D weight matrices and AdamW for other parameters.

Model Details

  • Architecture: Hybrid LFM (10 Convolutional + 6 Attention layers)
  • Parameters: 262.2M
  • Vocab Size: 49152 (using tokenizer HuggingFaceTB/SmolLM2-135M)
  • Sequence Length: 2048
  • Dataset: glouriousgautam/lilm-pretrain-12B (FineWeb, FineWeb-Edu, FineMath, DCLM, FinePDFs)
  • Total Tokens: ~10.5B tokens
  • Precision: bfloat16

Training Configuration

  • Effective Batch Size: 294,912 tokens/step
  • Total Training Steps: 32000
  • Warmup Steps: 640
  • Muon Optimizer: LR = 0.01, Momentum = 0.95, WD = 0.01
  • AdamW Optimizer: LR = 0.0003, WD = 0.1, Betas = (0.9, 0.95)

Metrics & Tracking

  • Final Training Loss: 2.7235
  • Final Perplexity: 15.23
  • Tokens Seen: 9,692,774,400
  • Weights & Biases Project: lilm-pretraining
Downloads last month
6
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train glouriousgautam/lilm1-200M