metadata
license: mit
library_name: pytorch
pipeline_tag: text-generation
tags:
- text-generation
- pretraining
- decoder-only
- from-scratch
- 11m
- 1096-context
- rope
- rmsnorm
- swiglu
FrontD.11m
FrontD.11m is an approximately 11-million-parameter decoder-only language model trained completely from scratch.
Architecture
- Parameters: 10,996,608
- Context: 1096
- Vocabulary: 16000
- Hidden size: 256
- Layers: 6
- Attention heads: 16
- FFN size: 1152
- RoPE
- RMSNorm
- SwiGLU
- Tied input/output embeddings
Pretraining dataset
Plans11/Organized_PreTrain_1k_Context
Training
Training is performed in 20,000-example sessions.
A persistent SHA-256 example ledger records only examples whose training session has successfully completed.
The model is initialized randomly and pretrained from scratch.